Temperature monitoring method of server and electronic equipment
By carefully dividing the heat dissipation areas in the server and dynamically adjusting the sampling frequency, the shortcomings of area division and dynamic sampling management in traditional server temperature control technology are solved, and accurate temperature monitoring and rapid fault response are achieved in high-density servers.
Patent Information
- Application Number
- CN202511190995.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-08-25
AI Technical Summary
Traditional server temperature control technology has problems in high-density scenarios, such as weak regional division and anomaly location capabilities, and extensive dynamic sampling and threshold management, making it difficult to balance monitoring accuracy, resource consumption, and fault response speed.
By carefully dividing the server plane into multiple heat dissipation areas, deploying temperature monitoring points according to the hardware layout and heat dissipation characteristics, and dynamically adjusting the sampling frequency based on the temperature change rate, combined with the temperature mean standard deviation to judge anomalies, it is possible to accurately identify single-point and regional temperature anomalies.
It improves temperature monitoring accuracy, shortens troubleshooting time, reduces resource consumption, and improves server operation stability and operation and maintenance efficiency.
Smart Images

Figure CN120687328A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of servers, and in particular to a temperature monitoring method and electronic equipment for a server. Background Art
[0002] As the hardware integration of storage and computing servers increases, the risk of temperature anomalies in the planar heat dissipation area increases significantly. In addition, the hardware thermal coupling problem is prominent in multi-disk storage servers or multi-processor computing servers, and single-point heat dissipation anomalies may cause chain failures.
[0003] Traditional server temperature control solutions have the following limitations when dealing with such scenarios: (1) Weak regional division and anomaly location capabilities The plane area division is rough and not finely segmented according to the hardware layout. The monitoring points are sparse and it is difficult to identify single-point anomalies. Some technologies increase monitoring points but lack temperature gradient time analysis, resulting in delayed anomaly identification.
[0004] (2) Dynamic sampling and threshold management are rough The sampling strategy does not take into account regional thermal sensitivity, resulting in delayed monitoring of key areas or waste of resources; the temperature threshold setting does not take real-time variables into consideration, making complex working conditions prone to misjudgment.
[0005] The above limitations make it difficult to balance monitoring accuracy, resource consumption, and fault response speed in high-density server scenarios. Summary of the Invention
[0006] The present invention provides a temperature monitoring method and electronic equipment for a server, which at least solves the problems of weak regional division and abnormality positioning capabilities, and rough dynamic sampling and threshold management in the current planar temperature control technology of servers.
[0007] The present invention provides a temperature monitoring method for a server, comprising the following steps: obtaining a first temperature change rate of at least one temperature monitoring point in the server; if the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate, determining a target heat dissipation area in which any of the temperature monitoring points is located, and adjusting a sampling frequency of the temperature monitoring points in the target heat dissipation area to a target sampling frequency; based on the target sampling frequency, detecting a second temperature change rate of the temperature monitoring points in the target heat dissipation area within a preset time length; if there is a target temperature monitoring point in the target heat dissipation area whose second temperature change rate within the preset time length is greater than the second preset change rate, determining that the temperature of the target temperature monitoring point is abnormal; otherwise, obtaining a temperature monitoring result of the target heat dissipation area based on a temperature mean standard deviation of the target heat dissipation area within the preset time length.
[0008] The present invention also provides a temperature monitoring system for a server, comprising: an acquisition module for acquiring a first temperature change rate of at least one temperature monitoring point in the server; an adjustment module for determining a target heat dissipation area in which any temperature monitoring point is located if the first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate, and adjusting the sampling frequency of the temperature monitoring points in the target heat dissipation area to a target sampling frequency; a monitoring module for detecting a second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time period based on the target sampling frequency; if the second temperature change rate of a target temperature monitoring point in the target heat dissipation area within the preset time period is greater than the second preset change rate, determining that the temperature of the target temperature monitoring point is abnormal; otherwise, obtaining a temperature monitoring result of the target heat dissipation area based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time period.
[0009] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned server temperature monitoring method when executing the computer program.
[0010] The present invention also provides a non-volatile computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the temperature monitoring method of the server are implemented.
[0011] The present invention also provides a computer program product, including a computer program, which implements the above-mentioned server temperature monitoring method when executed by a processor.
[0012] Through the present invention, if the first temperature change rate of any temperature monitoring point in the server is greater than or equal to the first change rate, the target heat dissipation area of any temperature monitoring point is determined, and the sampling frequency of the temperature monitoring point in the target heat dissipation area is adjusted to the target sampling frequency; based on the target sampling frequency, the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time period is detected; if the second temperature change rate of the target temperature monitoring point in the target heat dissipation area within the preset time period is greater than the second change rate, the temperature of the target temperature monitoring point is determined to be abnormal; otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time period. Thus, the problems of the current plane temperature control technology of the server, such as weak regional division and abnormality location capabilities, and extensive dynamic sampling and threshold management, are solved. The plane region division is refined, and the sampling frequency is automatically adjusted according to the temperature change amplitude, so that the single point abnormality in the region can be accurately identified, the monitoring accuracy is improved, and the monitoring accuracy, resource consumption and response speed can be balanced in high-density server scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0014] Figure 1 A flow chart of a method for monitoring the temperature of a server according to an embodiment of the present invention; Figure 2 A schematic diagram of server area division according to an embodiment of the present invention; Figure 3 A schematic flow chart of a method for monitoring server temperature according to an embodiment of the present invention; Figure 4 is a schematic diagram of a temperature monitoring system for a server according to an embodiment of the present invention; Figure 5 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0016] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0017] As storage and computing servers increase hardware integration, the risk of temperature anomalies in planar cooling areas increases significantly. In multi-bay storage servers or multi-processor computing servers, hardware thermal coupling within the same planar cooling area becomes prominent, and single-point cooling anomalies can trigger cascading failures. Traditional temperature control solutions have significant limitations in addressing such scenarios: static sampling strategies struggle to capture sudden temperature rises, fixed temperature thresholds fail to adapt to hardware thermal characteristics under varying operating conditions (such as the temperature difference between a computing server's training tasks and its idle state), and the coarse division of planar areas (e.g., treating the entire hard drive backplane as a single monitoring unit) easily obscures abnormal temperature points by regional averages, making it difficult to accurately locate the source of the fault. While the BMC (Baseboard Management Controller), the server's out-of-band management core, provides basic temperature collection capabilities, existing technologies often rely on single-point data recording and fixed-strategy alarms, lacking dynamic analysis of temperature gradients within the planar area and adaptive response mechanisms.
[0018] In planar temperature control technology, existing solutions have the following main problems: (1) Weak regional division and anomaly location capabilities: Most solutions simply divide the server plane into several large areas without fine-grained segmentation based on the hardware layout, resulting in sparse temperature monitoring points within the area and difficulty in identifying single-point anomalies. Although some technologies increase the number of monitoring points, they lack temperature gradient analysis in the time dimension, resulting in significant delays in anomaly identification.
[0019] (2) Dynamic sampling and threshold management are crude: Sampling strategies often use fixed frequencies or simple threshold switching, without considering regional thermal sensitivity differences, resulting in delayed monitoring of key areas or waste of resources. Temperature thresholds are often set based on hardware nominal values, without incorporating variables such as real-time load and historical temperature, making misjudgment prone to occur under complex operating conditions.
[0020] The core bottleneck of current flat-panel temperature control technology for storage and computing servers lies in the lack of granularity in zone division, resulting in low anomaly location accuracy and the inability of static policies to adapt to the dynamic thermal characteristics of the hardware. These limitations make it difficult for traditional solutions to balance monitoring accuracy, resource consumption, and fault response speed in high-density server scenarios.
[0021] In order to solve the above problems, an embodiment of the present invention provides a temperature monitoring method for a server, specifically as follows: Figure 1 shown.
[0022] like Figure 1 As shown, the temperature monitoring method of the server includes the following steps: Step S101: Acquire a first temperature change rate of at least one temperature monitoring point in a server.
[0023] Optionally, in some embodiments, before obtaining the first temperature change rate of at least one temperature monitoring point in the server, it includes: dividing the server into multiple heat dissipation areas based on the heat dissipation characteristics of the hardware area of the server, wherein the multiple heat dissipation areas are a backplane heat dissipation area, a power supply heat dissipation area, a central processing unit heat dissipation area, and an interface card heat dissipation area; determining the number of temperature monitoring points in the backplane heat dissipation area according to the area and heat dissipation characteristics of the backplane heat dissipation area, determining the number of temperature monitoring points in the power supply heat dissipation area according to the area and heat dissipation characteristics of the power supply heat dissipation area, determining the number of temperature monitoring points in the central processing unit heat dissipation area according to the area and heat dissipation characteristics of the central processing unit heat dissipation area, and determining the number of temperature monitoring points in the interface card heat dissipation area according to the area and heat dissipation characteristics of the interface card heat dissipation area.
[0024] The server includes multiple heat dissipation areas divided by hardware layout, and each heat dissipation area is provided with at least one temperature monitoring point. The at least one temperature monitoring point may be one or more.
[0025] Specific as Figure 2 As shown, the system first meticulously divides the server plane into multiple independent cooling zones based on the hardware layout. Taking a storage server as an example, the hard drive backplanes in a storage server have different cooling conditions due to their front-to-back positioning. These zones are divided into front-back hard drive cooling zones, internal-back hard drive cooling zones, and rear-back hard drive cooling zones. The front-back hard drive cooling zone is close to the server's air inlet, offering relatively smooth airflow, but may be affected by dust accumulation in the external environment. The internal-back hard drive zone is located inside the server, relying primarily on internal air ducts for heat dissipation. The rear-back hard drive zone is close to the air outlet, allowing for faster heat dissipation, but the higher outlet temperature may also affect it.
[0026] In addition, the power supply unit (PSU) cooling area is separately divided because it generates a lot of heat and its stability is related to the overall operation of the server. The central processing unit (CPU) cooling area, as the server's computing core, generates concentrated heat and is extremely sensitive to temperature, and is also separately divided. The interface card cooling area contains various network and storage interface cards, which are densely packed and generate heat during operation, and is also divided as an independent area.
[0027] Within each cooling zone, the system strategically deploys multiple temperature monitoring points based on factors such as area and hardware heat generation. For example, in the front backplane hard drive cooling area, temperature sensors are evenly spaced along the direction of the hard drives' arrangement to ensure coverage of temperature changes across the entire area. This creates a high-density monitoring network, enabling comprehensive and accurate capture of temperature changes across all server components.
[0028] Through the above technical solution, the number of monitoring points is dynamically determined according to the area and heat dissipation characteristics of each heat dissipation area (such as air duct design, equipment heat density, heat conduction efficiency, etc.), achieving efficient allocation of monitoring resources. Through partitioned independent monitoring systems, heat dissipation anomalies in different areas can be responded to in real time to improve temperature monitoring accuracy.
[0029] Specifically, after the system is started, based on the division of each heat dissipation area of the server and the deployment of monitoring points, an initial sampling frequency will be set for each temperature monitoring point. All temperature monitoring points in all heat dissipation areas will perform temperature sampling at the initial sampling frequency, and the first temperature change rate of all temperature monitoring points will be monitored.
[0030] In the embodiment of the present invention, the initial sampling frequency is once every 10 seconds.
[0031] In step S102 , if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate, the target heat dissipation area where the temperature monitoring point is located is determined, and the sampling frequencies of all temperature monitoring points in the target heat dissipation area are adjusted to the target sampling frequency.
[0032] The first preset change rate is 0.5° C. / minute, and the second preset change rate is 5° C. / minute.
[0033] When the server's overall temperature is stable—that is, the rate of change of the first temperature at all temperature monitoring points in all cooling zones is within ±0.5°C / minute—the system uniformly sets the temperature monitoring points to collect temperature data from each zone at an initial sampling frequency of once every 10 seconds. This unified, low-frequency data collection strategy not only meets basic temperature monitoring requirements during normal server operation, but also effectively reduces resource consumption on the baseboard management controller, ensuring stable system operation.
[0034] If the first temperature change rate of any temperature monitoring point among all temperature monitoring points within a preset time period (for example, 5 minutes) is greater than or equal to 0.5°C / minute, the target heat dissipation area where any temperature monitoring point is located is determined, and the sampling frequency of all temperature monitoring points in the target heat dissipation area is adjusted to the target sampling frequency.
[0035] Optionally, in some embodiments, the sampling frequency of all temperature monitoring points in the target heat dissipation area is adjusted to the target sampling frequency, including: determining the target temperature change interval in which the first temperature change rate of any temperature monitoring point is located, and determining the target sampling frequency based on the target temperature change interval.
[0036] Optionally, in some embodiments, the target temperature change interval in which the first temperature change rate of any temperature monitoring point is located is determined, and the target sampling frequency is determined based on the target temperature change interval, including: judging whether the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate; if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, then the target temperature change interval is determined to be the first interval, and the first sampling frequency is used as the target sampling frequency of the temperature monitoring point in the target heat dissipation area; if the first temperature change rate of any temperature monitoring point is greater than the second preset change rate, then the target temperature change interval is determined to be the second interval, and the second sampling frequency is used as the target sampling frequency of the temperature monitoring point in the target heat dissipation area.
[0037] In some embodiments, the first preset change rate is 0.5°C / minute, the second preset change rate is 5°C / minute, the first interval is greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, the second interval is greater than 5°C / minute, the first sampling frequency is 1 time per second, and the second sampling frequency is 10 times per second.
[0038] Among them, the temperature change interval in the embodiment of the present invention includes a first interval and a second interval, the temperature change rate in the first interval is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, and the temperature change rate in the second interval is greater than the second preset change rate, that is, the first interval is greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, and the second interval is greater than 5°C / minute.
[0039] If the first temperature change rate of any temperature monitoring point is greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, the target temperature change interval is determined to be the first interval. If the first temperature change rate of any temperature monitoring point is greater than 5°C / minute, the target temperature change interval is determined to be the second interval, and the target sampling frequency is determined based on the target temperature change interval.
[0040] In the embodiment of the present invention, the first sampling frequency is 1 time / second, which is medium frequency sampling in the present invention, and the second sampling frequency is 0.1 times / second, which is high frequency sampling in the present invention.
[0041] The specific adjustment method of the sampling frequency is: Let f be the temperature sampling frequency (unit: times / second), ΔTi be the temperature change rate of the i-th temperature monitoring point in the region (unit: °C / minute), and ΔTmax=max(ΔT1, ΔT2, …, ΔTn) be the maximum temperature change rate among all temperature monitoring points in the region. The sampling frequency rule can be expressed as: (1) Initial sampling frequency: If and only if the temperature change rate of all temperature monitoring points in the heat dissipation area satisfies |ΔTi|≤0.5℃ / minute (for all i): Then maintain the initial sampling frequency of the temperature monitoring point: f = 0.1 (i.e., sampling once every 10 seconds); (2) IF sampling When the temperature change rate of at least one temperature monitoring point in the heat dissipation area satisfies 0.5°C / minute < ΔTi ≤ 5°C / minute: Adjust the sampling frequency of all temperature detection points in the heat dissipation area to the target sampling frequency: f = 1 (i.e., 1 sample per second); (3) High-frequency sampling trigger When there is at least one monitoring point in the area with a temperature change rate that satisfies ΔTi>5°C / minute: Adjust the sampling frequency of all temperature detection points in the heat dissipation area to the target sampling frequency: f = 10 (i.e., one sample every 0.1 second).
[0042] For example, in a certain heat dissipation area, when the first temperature change rate of a temperature monitoring point rises to the range of 0.5℃ / minute-5℃ / minute, the system will give priority to increasing the sampling frequency of all temperature monitoring points in the heat dissipation area where the temperature monitoring point is located to the first sampling frequency, that is, 1 sampling per second. Figure 2 For example, if the temperature change rate of hard disk 1 in the front backplane hard disk area reaches 0.8℃ / minute, the system will immediately increase the sampling frequency of all temperature monitoring points of hard disk 1-hard disk 6 in the front backplane hard disk area to the first sampling frequency, that is, the target sampling frequency of all temperature monitoring points is 1 sample per second, so as to obtain the temperature change data of the area more timely and accurately.
[0043] If the first temperature change rate of a temperature monitoring point in a certain heat dissipation zone exceeds 5°C / minute, the system will determine that the zone is at risk of emergency overheating and will immediately initiate a high-frequency sampling mode of 0.1 times / second for all temperature monitoring points in the area where the temperature monitoring point is located. For example, if the temperature change rate of a temperature monitoring point in the CPU zone reaches 7°C / minute, all temperature monitoring points in the CPU zone will enter a high-frequency sampling state, sampling once every 0.1 seconds, capturing temperature data at an extremely high frequency to ensure that the rapidly changing temperature conditions in the zone can be accurately grasped.
[0044] Through this technical solution, when the temperature change rate of at least one temperature monitoring point within a specific heat dissipation zone falls between 0.5°C / minute and 5°C / minute, the system increases the sampling frequency for all monitoring points in that zone to once per second. This adjustment allows temperature data to be acquired at an appropriate frequency, avoiding data redundancy caused by oversampling while promptly capturing subtle temperature changes. This provides accurate and timely data support for subsequent analysis, helping to accurately identify temperature trends and potential problems. If the temperature change rate of any temperature monitoring point in the zone exceeds 5°C / minute, the system immediately initiates a high-frequency sampling mode of once every 0.1 seconds. In the event of an emergency overheating risk, this ultra-high-frequency sampling can accurately capture rapid temperature fluctuations in real time, capturing every moment of temperature change and ensuring the most detailed and accurate temperature data, providing critical information for emergency response.
[0045] Optionally, in some embodiments, after obtaining the first temperature change rate of the monitoring point of each heat dissipation area, it includes: controlling the temperature monitoring points in the heat dissipation area to maintain the current sampling frequency when there is no temperature monitoring point whose first temperature change rate within a preset time period is greater than or equal to the first preset change rate.
[0046] The preset time length may be a threshold value pre-set by the user, a threshold value obtained through a limited number of experiments, or a threshold value obtained through a limited number of computer simulations, and is not specifically limited here.
[0047] It can be understood that after obtaining the first temperature change rate of all temperature monitoring points in each heat dissipation area, if the first temperature change rate of all temperature monitoring points within the preset time length is less than the first preset change rate, all temperature monitoring points in all heat dissipation areas are controlled to maintain the current sampling frequency, with sampling once every 10 seconds.
[0048] The first preset change rate set by the system is 0.5°C / minute, the preset duration is 5 minutes, and the initial sampling frequency of all temperature monitoring points is once every 10 seconds.
[0049] After the system starts running, it samples the temperature at each temperature monitoring point every 10 seconds and records the temperature. Every 5 minutes (the preset duration), the system calculates the temperature change rate of each temperature monitoring point during these 5 minutes. For example, if the temperature at temperature monitoring point A is 25°C at the start and 26°C after 5 minutes, then its temperature change rate within 5 minutes is: (26-25)÷5=0.2℃ / minute; The temperature at the temperature monitoring point B is 23°C at the beginning and 24°C after 5 minutes. The temperature change rate is: (24-23)÷5=0.2℃ / minute; Similarly, the temperature change rates of all monitoring points in each heat dissipation area are calculated.
[0050] The system checks the temperature change rates calculated for all temperature monitoring points and finds that none of them have a rate of change greater than or equal to 0.5°C / minute (the first preset rate of change). In this case, all temperature monitoring points continue to maintain their current sampling frequency of once every 10 seconds and will not increase or decrease their sampling frequency.
[0051] The above technical solution determines whether to adjust the sampling frequency based on the temperature change rate of all monitoring points in each heat dissipation zone within a preset time period. When the temperature change rate of all monitoring points is less than a first preset change rate (e.g., 0.5°C / minute), the current sampling frequency is maintained. This dynamic adjustment mechanism ensures that the sampling frequency matches the actual temperature changes, avoiding the unnecessary resource consumption caused by high-frequency sampling when temperature changes are gentle, such as data storage space, transmission bandwidth, and processor computing resources, thereby achieving efficient resource utilization.
[0052] Step S103: Based on the target sampling frequency, the second temperature change rate of the temperature monitoring point in the target heat dissipation area within the preset time length is detected. If the second temperature change rate of the target temperature monitoring point in the target heat dissipation area within the preset time length is greater than the second preset change rate, it is determined that the temperature of the target temperature monitoring point is abnormal. Otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time length.
[0053] It should be understood that the temperature monitoring points in the target heat dissipation area perform temperature sampling based on the target sampling frequency, and calculate the second temperature change rate within the preset time period in real time during the continuous collection of temperature data, so as to obtain the temperature anomaly monitoring result of the server based on the second temperature change rate of the temperature monitoring points in the target heat dissipation area within the preset time period.
[0054] Specifically, while continuously collecting temperature data, the system calculates the second temperature change rate at each temperature monitoring point in real time. This second temperature change rate is calculated by calculating the temperature rise per unit time. For example, if the temperature at a monitoring point rises from 25°C to 35°C within 1 minute, the second temperature change rate is (35°C - 25°C) ÷ 1 minute = 10°C / minute. This calculation provides a key basis for determining whether temperature changes are abnormal.
[0055] Specifically, the system continuously compares the second temperature change rate of all temperature monitoring points, and detects whether there is a target temperature monitoring point among all temperature monitoring points whose second temperature change rate is greater than the second preset change rate. If there is a target temperature monitoring point, it means that the temperature rise speed of the target temperature monitoring point is faster than that of other temperature monitoring points in the neighborhood, and the target temperature monitoring point is marked as an abnormal temperature point.
[0056] Optionally, in some embodiments, after determining that the temperature of the target temperature monitoring point is abnormal, the method includes: performing a temperature abnormality warning based on the target temperature monitoring point, and sending an alarm message of the temperature abnormality of the target temperature monitoring point to a preset terminal.
[0057] Based on the abnormal temperature point, the temperature abnormality warning is carried out, and the alarm information of the abnormal temperature of the target temperature monitoring point in the server is sent to the preset terminal.
[0058] For example, if the temperature of a hard drive's temperature monitoring point rises significantly faster than other points in the neighborhood, and the difference in temperature increase per minute exceeds 5°C, the point will be marked as an abnormal temperature point.
[0059] If the temperature of a temperature monitoring point rises from 25°C to 35°C within 1 minute, its second temperature change rate is (35°C-25°C) ÷ 1 minute = 10°C / minute. 10°C / minute is greater than 5°C / minute. At this time, the temperature monitoring point is judged to be an abnormal temperature point. If the second temperature change rate of all temperature monitoring points is less than 5°C / minute within 1 minute, all temperature monitoring points are judged to be normal temperature points.
[0060] With this technical solution, once the second temperature change rate at a temperature monitoring point exceeds a second preset rate, it is quickly flagged as an abnormal temperature point. This real-time monitoring mechanism promptly detects abnormal temperature changes, buying valuable time for subsequent warnings and processing, and preventing the abnormal situation from further deteriorating. Once an abnormal temperature point is detected, the system immediately issues a temperature anomaly warning based on the abnormal point and sends an alarm message to a preset terminal. This allows relevant personnel, whether on-site operations and maintenance personnel or remote monitoring personnel, to be notified of temperature anomalies immediately, allowing them to take timely measures to reduce the probability of failure.
[0061] Optionally, in some embodiments, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within a preset time period, and also includes: calculating the standard deviation of the temperature mean of the temperature monitoring points in the target heat dissipation area within a preset time period; if the standard deviation of the temperature mean is greater than a preset threshold value corresponding to the target heat dissipation area, it is determined that the temperature of the target heat dissipation area is abnormal, and a temperature abnormality warning is issued, and an alarm message of the temperature abnormality of the target heat dissipation area is sent to a preset terminal.
[0062] It should be understood that if there is no target temperature monitoring point whose second temperature change rate is greater than the second preset change rate, it means that all temperature monitoring points are normal temperature points. At this time, it is necessary to determine whether the entire target heat dissipation area is an area with abnormal temperature based on the standard deviation of the temperature mean of the temperature monitoring points.
[0063] Specifically, the system evaluates the temperature distribution of the target heat dissipation area by calculating the mean temperature and standard deviation of the mean temperature for all temperature monitoring points within the target heat dissipation area. If the standard deviation of the mean temperature of the target heat dissipation area exceeds a preset threshold, it indicates that the target heat dissipation area temperature is abnormal, requiring a temperature anomaly warning. The system then sends an alarm message about the abnormal temperature of the target heat dissipation area to a preset terminal on the server.
[0064] It's important to note that different cooling zones have varying normal temperature fluctuation ranges due to their hardware composition and operating characteristics. Therefore, the system sets unique standard deviation thresholds based on each zone's hardware characteristics and historical operating data. For example, in the PSU zone, where heat generation is stable and temperature fluctuations are minimal during normal operation, the preset threshold is 3°C. In the interface card zone, where equipment operating conditions vary and temperature fluctuations are significant, the preset threshold is 4°C. When the calculated standard deviation for a zone exceeds the preset threshold, the system immediately triggers a warning for abnormal temperature in that zone.
[0065] For example, if the target heat dissipation area is the PSU area, the standard deviation of the mean temperature of the PSU area within the preset time period is calculated. If the standard deviation of the mean temperature of the PSU area within the preset time period is 2°C, and 2°C is less than 3°C, it means that the temperature of the PSU area is normal. If the standard deviation of the mean temperature of the PSU area within the preset time period is 5°C, and 5°C is greater than 3°C, it means that the temperature of the PSU area is abnormal. At this time, a temperature abnormality warning is issued, and an alarm message of the abnormal temperature of the PSU area is sent to the preset terminal.
[0066] The above technical solution, by comprehensively considering the data from all monitoring points, can fully understand the overall temperature distribution characteristics within the target heat dissipation area. This helps to identify potential temperature anomalies and temperature change trends, providing comprehensive and accurate information for subsequent temperature management and fault prevention. When the standard deviation of the mean temperature in the target heat dissipation area is greater than the preset threshold, the system immediately determines that the temperature in that area is abnormal and triggers an early warning mechanism, ensuring that relevant operation and maintenance personnel are informed of the abnormal situation immediately. Different heat dissipation areas have different normal temperature distribution fluctuation ranges due to different hardware composition and operating characteristics. The system sets a unique standard deviation threshold based on the hardware characteristics of each area, fully considering the heating characteristics and heat dissipation requirements of different hardware devices, and better adapting to the characteristics of regional temperature changes.
[0067] Optionally, in some embodiments, after sending the alarm information of the abnormal temperature at the target temperature monitoring point to the preset terminal, the method includes: generating a fault diagnosis report and processing suggestions according to the alarm information of the abnormal temperature at the target temperature monitoring point.
[0068] During operation, if the temperature monitoring system detects an abnormal temperature at a target temperature monitoring point, it immediately triggers an alarm mechanism and sends an alarm message containing key information about the abnormal temperature at the target temperature monitoring point to a pre-defined terminal device. This pre-defined terminal can be a mobile phone, computer, or other device used by relevant operation and maintenance personnel, ensuring that they are notified of abnormal conditions in a timely manner.
[0069] After successfully sending an alarm, the system uses the acquired alarm information regarding the temperature anomaly at the target temperature monitoring point and its built-in fault diagnosis algorithms and knowledge base to conduct a comprehensive and in-depth analysis and assessment of the anomaly. By analyzing multiple factors, including the degree of temperature anomaly, its changing trends, and historical data comparisons, the system generates a detailed fault diagnosis report. This report clearly identifies the possible causes of the temperature anomaly, such as device heat dissipation failure, excessively high ambient temperature, or sensor failure.
[0070] The system also provides targeted action suggestions to maintenance personnel based on the fault diagnosis results, pre-set handling strategies, and empirical data. These suggestions may include specific steps, such as checking the proper operation of the device's cooling fan, adjusting ambient temperature control parameters, or replacing a faulty sensor. These suggestions also include priority and estimated processing time, helping maintenance personnel quickly and effectively resolve temperature anomalies.
[0071] Through the above technical solution, by immediately generating a fault diagnosis report and processing suggestions after sending the alarm information, operation and maintenance personnel do not need to spend a lot of time to investigate the cause of the fault and formulate a processing plan by themselves. They can quickly carry out maintenance work based on the reports and suggestions provided by the system, which greatly shortens the time for fault handling and improves the operating efficiency and stability of the entire system.
[0072] Optionally, in some embodiments, after detecting the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time period based on the target sampling frequency, it includes: determining whether the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate; when the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate, reducing the target sampling frequency of the temperature monitoring point in the target heat dissipation area to the initial sampling frequency.
[0073] It can be understood that after all temperature monitoring points in the target heat dissipation area detect the second temperature change rate within the preset time period based on the target sampling frequency, if the second temperature change rate of all temperature monitoring points in the target heat dissipation area is less than the first preset change rate, the target sampling frequency of all temperature monitoring points is reduced to the initial sampling frequency.
[0074] Specifically, the system continuously monitors the temperature change trends of all temperature monitoring points in the target heat dissipation area. When the second temperature change rate of all temperature monitoring points in the target heat dissipation area falls back to the stable range (that is, the temperature change rate is less than the first preset change rate), the sampling frequency of all temperature monitoring points will be reduced to the initial sampling frequency, that is, once every 10 seconds.
[0075] For example, if the PSU area previously entered high-frequency sampling due to an increase in load causing the temperature change rate to exceed the standard, when the load decreases and the temperature change rate of all temperature monitoring points in the PSU area stabilizes, the sampling frequency will gradually return to the initial sampling frequency, avoiding unnecessary high-frequency sampling and reducing waste of system resources.
[0076] Through the above technical solution, high-frequency adoption and medium-frequency sampling will put the system in a high-load operation state, which may easily lead to system overload, freezes, and other problems. When the temperature changes at the temperature monitoring points tend to stabilize, reducing the sampling frequency can reduce the system's working pressure, avoid unnecessary high-frequency sampling, and reduce system resource waste.
[0077] To enable those skilled in the art to further understand the temperature monitoring method of the server in the embodiment of the present application, the following is described in detail with reference to specific embodiments. Figure 3 shown.
[0078] In step S301: Layout of heat dissipation areas and deployment of temperature monitoring points: The system first divides the server hardware layout into multiple independent heat dissipation areas such as CPU, hard disk backplane, PSU, etc., and deploys monitoring points in each heat dissipation area.
[0079] In step S302: the initial sampling frequency of the temperature monitoring point is set. During the initial operation, the temperature is collected at a frequency of once every 10 seconds.
[0080] In step S303: sampling starts.
[0081] In step S304, it is determined whether the first temperature change rate of any temperature monitoring point is greater than or equal to 0.5°C / minute. If so, step S305 is executed. If the first temperature change rates of all temperature monitoring points are less than 0.5°C / minute, step S306 is executed.
[0082] In step S305: determine whether the first temperature change rate of the temperature monitoring point is greater than 5°C / minute. If so, execute step S307; if greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, execute step S308.
[0083] In step S306: the initial sampling frequency is maintained.
[0084] In step S307 : the sampling frequencies of all temperature monitoring points in the target heat dissipation area where the temperature monitoring point is located are adjusted to a high-frequency sampling frequency, and step S309 is executed.
[0085] In step S308: the sampling frequencies of all temperature monitoring points in the target heat dissipation area where the temperature monitoring point is located are adjusted to the intermediate frequency sampling frequency, and step S309 is executed.
[0086] In step S309: calculate whether there is a single high temperature point among all temperature monitoring points in the target heat dissipation area. If there is a single high temperature point, execute step S310; if not, execute step S311.
[0087] In step S310: a single point is marked as an abnormal temperature point, and a single point temperature abnormality warning is triggered.
[0088] In step S311: calculate the temperature mean standard deviation of all temperature monitoring points in the target heat dissipation area, and execute S312.
[0089] In step S312: determine whether the standard deviation of the temperature mean exceeds a preset threshold. If so, execute step S313; if not, execute step S314.
[0090] In step S313: triggering a target heat dissipation area temperature abnormality warning.
[0091] In step S314: continue with the next monitoring.
[0092] In summary, the technical effects brought about by the embodiments of the present invention are as follows.
[0093] (1) Significantly improve server temperature monitoring performance: Based on the heating characteristics of hardware regions, the system collects and calculates the temperature change rate in seconds, quickly capturing subtle temperature changes. For example, in the hard drive backplane area, it can promptly detect local temperature fluctuations caused by a single hard drive failure, significantly improving monitoring accuracy compared to traditional monitoring methods. (2) In terms of fault diagnosis, both single-point and regional anomaly judgment are used simultaneously. The monitoring point neighborhood is dynamically delineated for comparison gradients to accurately locate single-point anomalies. Regional anomalies are judged using a dedicated standard deviation threshold. In the interface card area, overall overheating caused by dense equipment or failures can be detected in advance, greatly shortening troubleshooting time and effectively reducing the risk of server downtime. (3) To save energy and reduce consumption, the system adopts a dynamic sampling strategy. When the temperature is stable, low-frequency sampling is used to reduce BMC resource consumption; when temperature fluctuations increase, the frequency is increased to avoid invalid data collection, significantly reducing system operating energy consumption and extending the service life of server hardware. In addition, the fault diagnosis reports and treatment suggestions output by the system effectively improve operation and maintenance efficiency, and have high practical value and economic benefits in scenarios such as data centers.
[0094] According to the server temperature monitoring method proposed in an embodiment of the present invention, if the first temperature change rate of any temperature monitoring point in the server is greater than or equal to the first change rate, the target heat dissipation area of the temperature monitoring point is determined, and the sampling frequency of the temperature monitoring point in the target heat dissipation area is adjusted to the target sampling frequency; based on the target sampling frequency, the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time period is detected; if the second temperature change rate of the target temperature monitoring point in the target heat dissipation area within the preset time period is greater than the second change rate, the target temperature monitoring point is determined to have a temperature anomaly; otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time period. This solves the problems of current server planar temperature control technology, such as weak regional division and anomaly location capabilities, and extensive dynamic sampling and threshold management. It refines the planar regional division, automatically adjusts the sampling frequency according to the temperature change amplitude, realizes accurate identification of single-point anomalies in the area, improves monitoring accuracy, and can balance monitoring accuracy, resource consumption, and response speed in high-density server scenarios.
[0095] Next, a temperature monitoring system for a server according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0096] Figure 4 4 is a schematic diagram of a temperature monitoring system for a server according to an embodiment of the present invention.
[0097] like Figure 4 As shown, the server temperature monitoring system 10 includes: an acquisition module 100 , an adjustment module 200 and a monitoring module 300 .
[0098] Among them, the acquisition module 100 is used to obtain the first temperature change rate of at least one temperature monitoring point in the server; the adjustment module 200 is used to determine the target heat dissipation area where any temperature monitoring point is located if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate, and adjust the sampling frequency of the temperature monitoring point in the target heat dissipation area to the target sampling frequency; the monitoring module 300 is used to detect the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time length based on the target sampling frequency. If the second temperature change rate of the target temperature monitoring point in the target heat dissipation area within the preset time length is greater than the second preset change rate, it is determined that the temperature of the target temperature monitoring point is abnormal. Otherwise, the temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time length.
[0099] Optionally, in some embodiments, before obtaining the first temperature change rate of at least one temperature monitoring point in the server, the acquisition module is further used to: divide the server into multiple heat dissipation areas based on the heat dissipation characteristics of the hardware area of the server, wherein the multiple heat dissipation areas are a backplane heat dissipation area, a power supply heat dissipation area, a central processing unit heat dissipation area, and an interface card heat dissipation area; determine the number of temperature monitoring points in the backplane heat dissipation area according to the area and heat dissipation characteristics of the backplane heat dissipation area, determine the number of temperature monitoring points in the power supply heat dissipation area according to the area and heat dissipation characteristics of the power supply heat dissipation area, determine the number of temperature monitoring points in the central processing unit heat dissipation area according to the area and heat dissipation characteristics of the central processing unit heat dissipation area, and determine the number of temperature monitoring points in the interface card heat dissipation area according to the area and heat dissipation characteristics of the interface card heat dissipation area.
[0100] Optionally, in some embodiments, after determining that the temperature of the target temperature monitoring point is abnormal, the monitoring module 300 is further used to: perform temperature abnormality warning based on the target temperature monitoring point, and send alarm information of the temperature abnormality of the target temperature monitoring point to a preset terminal.
[0101] Optionally, in some embodiments, the monitoring module 300 is further used to: calculate the standard deviation of the temperature mean of the temperature monitoring points in the target heat dissipation area within a preset time period; if the standard deviation of the temperature mean is greater than a preset threshold value corresponding to the target heat dissipation area, it is determined that the temperature of the target heat dissipation area is abnormal, and a temperature abnormality warning is issued, and an alarm message of the abnormal temperature of the target heat dissipation area is sent to a preset terminal.
[0102] Optionally, in some embodiments, the adjustment module 200 is further configured to: determine a target temperature change interval in which the first temperature change rate of any temperature monitoring point is located, and determine a target sampling frequency according to the target temperature change interval.
[0103] Optionally, in some embodiments, the adjustment module 200 is further used to: determine whether the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate; if the first temperature change rate of any temperature monitoring point is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, then determine that the target temperature change interval is the first interval, and use the first sampling frequency as the target sampling frequency of the temperature monitoring point in the target heat dissipation area; if the first temperature change rate of any temperature monitoring point is greater than the second preset change rate, then determine that the target temperature change interval is the second interval, and use the second sampling frequency as the target sampling frequency of the temperature monitoring point in the target heat dissipation area.
[0104] Optionally, in some embodiments, the first preset change rate is 0.5°C / minute, the second preset change rate is 5°C / minute, the first interval is greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, the second interval is greater than 5°C / minute, the first sampling frequency is 1 time per second, and the second sampling frequency is 10 times per second.
[0105] Optionally, in some embodiments, after obtaining the first temperature change rate of at least one temperature monitoring point in the server, the acquisition module 100 is further used to: control the temperature monitoring points in the heat dissipation area to maintain the current sampling frequency when there is no temperature monitoring point whose first temperature change rate within a preset time period is greater than or equal to the first preset change rate.
[0106] Optionally, in some embodiments, after sending the alarm information of the abnormal temperature at the target temperature monitoring point to the preset terminal, the monitoring module 300 is further used to: generate a fault diagnosis report and processing suggestions based on the alarm information of the abnormal temperature at the target temperature monitoring point.
[0107] Optionally, in some embodiments, after detecting the second temperature change rate of the temperature monitoring point in the target heat dissipation area within a preset time period based on the target sampling frequency, the monitoring module 300 is further used to: determine whether the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate; if the second temperature change rate of the temperature monitoring point in the target heat dissipation area is less than the first preset change rate, reduce the target sampling frequency of the temperature monitoring point in the target heat dissipation area to the initial sampling frequency.
[0108] It should be noted that, for the description of the features in the embodiment corresponding to the temperature monitoring system of the server, reference can be made to the relevant description of the embodiment corresponding to the temperature monitoring method of the server, which will not be repeated here.
[0109] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device may include: Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .
[0110] When the processor 502 executes the program, the temperature monitoring method of the server provided in the above embodiment is implemented.
[0111] Furthermore, the electronic device further includes: The communication interface 503 is used for communication between the memory 501 and the processor 502 .
[0112] The memory 501 is used to store computer programs that can be run on the processor 502 .
[0113] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0114] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0115] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.
[0116] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0117] An embodiment of the present invention further provides a non-volatile computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned server temperature monitoring method embodiments when running.
[0118] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0119] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the above-mentioned server temperature monitoring method when executed by a processor.
[0120] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0121] The above is a detailed introduction to the temperature monitoring method and electronic device for a server provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A temperature monitoring method for a server, characterized in that: The following steps are involved: Obtaining a first temperature change rate of at least one temperature monitoring point in the server; If the first temperature change rate of any of the temperature monitoring points is greater than or equal to the first preset change rate, determining the target heat dissipation area where any of the temperature monitoring points is located, and adjusting the sampling frequency of the temperature monitoring points in the target heat dissipation area to the target sampling frequency; Based on the target sampling frequency, a second temperature change rate of a temperature monitoring point in the target heat dissipation area within a preset time length is detected. If the second temperature change rate of a target temperature monitoring point in the target heat dissipation area within the preset time length is greater than the second preset change rate, it is determined that the temperature of the target temperature monitoring point is abnormal. Otherwise, a temperature monitoring result of the target heat dissipation area is obtained based on the standard deviation of the temperature mean of the target heat dissipation area within the preset time length.
2. The temperature monitoring method of a server according to claim 1, characterized in that: Before obtaining a first temperature change rate of at least one temperature monitoring point in the server, the method includes: Based on the heat dissipation characteristics of the hardware area of the server, the server is divided into multiple heat dissipation areas, wherein the multiple heat dissipation areas are respectively a backplane heat dissipation area, a power supply heat dissipation area, a central processing unit heat dissipation area and an interface card heat dissipation area; The number of temperature monitoring points in the backplane heat dissipation area is determined based on the area and heat dissipation characteristics of the backplane heat dissipation area, and the number of temperature monitoring points in the power supply heat dissipation area is determined based on the area and heat dissipation characteristics of the power supply heat dissipation area, and the number of temperature monitoring points in the central processing unit heat dissipation area is determined based on the area and heat dissipation characteristics of the central processing unit heat dissipation area, and the number of temperature monitoring points in the interface card heat dissipation area is determined based on the area and heat dissipation characteristics of the interface card heat dissipation area.
3. The temperature monitoring method of a server according to claim 1, characterized in that: After determining that the temperature of the target temperature monitoring point is abnormal, the method includes: A temperature anomaly warning is performed based on the target temperature monitoring point, and an alarm message of the temperature anomaly of the target temperature monitoring point is sent to a preset terminal.
4. The temperature monitoring method of a server according to claim 1, characterized in that: The method further includes: obtaining a temperature monitoring result of the target heat dissipation area based on a temperature mean standard deviation of the target heat dissipation area within a preset time period; Calculate the standard deviation of the temperature mean of the temperature monitoring points in the target heat dissipation area within the preset time period; If the temperature mean standard deviation is greater than the preset threshold corresponding to the target heat dissipation area, the temperature of the target heat dissipation area is determined to be abnormal, and a temperature abnormality warning is issued, and an alarm message of the temperature abnormality of the target heat dissipation area is sent to a preset terminal.
5. The temperature monitoring method of a server according to claim 1, characterized in that: The step of adjusting the sampling frequencies of all temperature monitoring points in the target heat dissipation area to the target sampling frequency includes: A target temperature change interval in which the first temperature change rate of any of the temperature monitoring points is located is determined, and the target sampling frequency is determined according to the target temperature change interval.
6. The temperature monitoring method of a server according to claim 5, characterized in that: Determining the target temperature change interval in which the first temperature change rate of any of the temperature monitoring points is located, and determining the target sampling frequency according to the target temperature change interval, includes: Determining whether a first temperature change rate of any of the temperature monitoring points is greater than or equal to a first preset change rate and less than or equal to a second preset change rate; If the first temperature change rate of any of the temperature monitoring points is greater than or equal to the first preset change rate and less than or equal to the second preset change rate, the target temperature change interval is determined to be the first interval, and the first sampling frequency is used as the target sampling frequency of the temperature monitoring points in the target heat dissipation area; If the first temperature change rate of any of the temperature monitoring points is greater than the second preset change rate, the target temperature change interval is determined to be the second interval, and the second sampling frequency is used as the target sampling frequency of the temperature monitoring points in the target heat dissipation area.
7. The temperature monitoring method of a server according to claim 6, characterized in that: The first preset change rate is 0.5°C / minute, the second preset change rate is 5°C / minute, the first interval is greater than or equal to 0.5°C / minute and less than or equal to 5°C / minute, the second interval is greater than 5°C / minute, the first sampling frequency is 1 time per second, and the second sampling frequency is 10 times per second.
8. The temperature monitoring method of a server according to claim 1, characterized in that: After obtaining the first temperature change rate of the monitoring point of each heat dissipation area, the method includes: In a case where none of the temperature monitoring points has a first temperature change rate greater than or equal to the first preset change rate within a preset time period, the temperature monitoring points in the heat dissipation area are controlled to maintain a current sampling frequency.
9. The temperature monitoring method of a server according to claim 3, characterized in that: After sending the alarm information of abnormal temperature of the target temperature monitoring point to the preset terminal, the method includes: Generate a fault diagnosis report and processing suggestions based on the alarm information of the abnormal temperature of the target temperature monitoring point.
10. The temperature monitoring method of a server according to claim 1, characterized in that: After detecting a second temperature change rate of a temperature monitoring point in the target heat dissipation area within a preset time period based on the target sampling frequency, the method includes: Determine whether the second temperature change rates of the temperature monitoring points in the target heat dissipation area are all less than the first preset change rate; When the second temperature change rates of the temperature monitoring points in the target heat dissipation area are all less than the first preset change rates, the target sampling frequency of the temperature monitoring points in the target heat dissipation area is reduced to the initial sampling frequency.
11. An electronic device, characterized in that: The system comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the temperature monitoring method for a server according to any one of claims 1 to 10.
Citation Information
Patent Citations
Server heat dissipation abnormity detection method and device, equipment and medium
CN118093290A
Optical module heat dissipation method and device based on 800G transmission, computer equipment and storage medium
CN119012637A
Heat dissipation control method for network server rack
CN119828870A
Computer equipment heat dissipation method, device and equipment
CN120143951A
Cooling equipment control method and device, equipment and storage medium
CN120353134A
Cited By
Multi-sensor temperature data sampling method and device, electronic equipment and medium
CN121577192A