Temperature monitoring device, temperature monitoring method, and temperature monitoring program

The temperature monitoring device addresses the challenge of detecting heat conduction abnormalities in CPU cooling systems by analyzing temperature changes across multiple sensors, enabling precise identification and remediation of cooling issues.

JP7683973B1Active Publication Date: 2025-05-27NEC PLATFROMS LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024034546
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-05-27
Estimated Expiration
2044-03-07

AI Technical Summary

Technical Problem

Existing methods for monitoring CPU temperature in information processing devices cannot detect abnormalities in heat conduction from the CPU to the cooling device, leading to incomplete diagnosis of cooling issues.

Method used

A temperature monitoring device and method that utilize multiple temperature sensors and processors to measure and analyze temperature changes of both the CPU and the cooling device over predetermined periods, calculating temperature changes and determining the occurrence and cause of cooling abnormalities.

Benefits of technology

Effectively identifies when electronic components are not being cooled properly and determines the underlying cause, enabling targeted interventions to prevent performance degradation or system failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683973000001_ABST
    Figure 0007683973000001_ABST
Patent Text Reader

Abstract

This helps determine when electronic components and modules that require heat dissipation are not being cooled properly and identifies the root cause. [Solution] A temperature monitoring device of the present invention acquires the temperatures of an electronic component and its cooling device at a first time, a second time, and a third time included in each of a plurality of measurement periods, calculates a first temperature change between the temperature of the electronic component acquired at the first time and the temperature of the electronic component acquired at the second time, calculates a second temperature change between the temperature of the cooling device acquired at the second time and the temperature of the cooling device acquired at the third time, and performs a process of determining whether a cooling abnormality has occurred in the electronic component based on the calculated first temperature change and second temperature change.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a temperature monitoring device, a temperature monitoring method, and a temperature monitoring program for monitoring the temperature of electronic components and the like. [Background technology]

[0002] In devices such as computers, electronic components such as a central processing unit (CPU) and modules such as a hard disk drive (HDD) are used, and since these electronic components and modules consume power and generate heat, appropriate cooling is required. In the following description, the electronic components and modules are referred to as "electronic components." In particular, in information processing devices such as computers and servers, the CPU has the greatest impact on information processing performance, and generally consumes the most power among electronic components, so generates a lot of heat. When the CPU becomes excessively hot, its internal protection function is activated, suppressing or stopping operation. In this way, when the CPU becomes too hot, the processing performance of the information processing device decreases, or even processing becomes impossible.

[0003] In order to prevent such a situation and maintain normal operation, a cooling device using a cooling component such as a heat sink (heat sink) or a cooling fan is usually attached to the CPU via heat dissipation grease that enhances heat transfer. In addition, attaching a component to a circuit board or the like in this manner is also generally described as "mounting". When a cooling device using a cooling fan is mounted on a CPU, the rotation speed of the cooling fan may be controlled according to the measured CPU temperature. Furthermore, in order to diagnose whether the CPU cooling device is operating normally, for example, a method is known for detecting a failure of the cooling device by comparing the temperature when the CPU load is maintained constant with the temperature measured in advance when the CPU is being cooled normally (Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2009-187347 A Summary of the Invention [Problem to be solved by the invention]

[0005] The disclosures of the above-mentioned prior art documents are incorporated herein by reference. The following analysis was conducted by the present inventors.

[0006] However, the method disclosed in Patent Document 1 can detect an abnormality in the cooling device of the CPU, but cannot detect whether the heat conduction from the CPU to the cooling device is normal or not. If the heat conduction from the CPU to the cooling device is not normal, the fundamental cause of the CPU cooling being hindered may not be known even if only the failure of the cooling device is detected. In other words, for example, if the fundamental cause of the CPU cooling being hindered is a decrease in the efficiency of heat conduction between the cooling device and the surrounding air, the CPU will not be cooled normally even if the cooling device is replaced.

[0007] In view of the above-mentioned problems, the present disclosure aims to determine when electronic components and modules that require heat dissipation in devices such as information processing devices are not being cooled properly and to contribute to identifying the underlying cause. [Means for solving the problem]

[0008] In a first aspect of the present disclosure, there is provided a temperature monitoring device that includes one or more processors and monitors the temperature of an electronic device, the temperature monitoring device comprising: an electronic component; a cooling device attached to the electronic component; a first temperature sensor attached to the electronic component and measuring a temperature of the electronic component; and a second temperature sensor attached to the cooling device and measuring a temperature of the cooling device. The one or more processors of the temperature monitoring device are configured to acquire a temperature of the electronic component measured by at least the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods, acquire a temperature of the electronic component measured by the first temperature sensor and a temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods, acquire a temperature of the cooling device measured by at least the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods, calculate a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the plurality of predetermined measurement periods and the temperature of the electronic component acquired at the second time, calculate a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the plurality of predetermined measurement periods and the temperature of the cooling device acquired at the third time, determine an occurrence of a cooling abnormality of the electronic component based on the calculated first temperature change and the second temperature change, and determine and report a cause of the occurrence of the cooling abnormality of the electronic component.

[0009] In a second aspect of the present disclosure, there is provided a temperature monitoring method for monitoring a temperature of an electronic device including an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component to measure a temperature of the electronic component, and a second temperature sensor attached to the cooling device to measure a temperature of the cooling device. The temperature monitoring method includes a first acquisition step, a second acquisition step, a third acquisition step, a first calculation step, a second calculation step, and a determination / reporting step. The first acquisition step acquires the temperature of the electronic component measured by at least the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods. The second acquisition step acquires the temperature of the electronic component measured by the first temperature sensor and the temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods. The third acquisition step acquires the temperature of the cooling device measured by at least the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods. The first calculation step calculates a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the predetermined plurality of measurement periods and the temperature of the electronic component acquired at the second time. The second calculation step calculates a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the predetermined plurality of measurement periods and the temperature of the cooling device acquired at the third time. The judging and reporting step judges an occurrence of a cooling abnormality of the electronic component based on the calculated first temperature change and the calculated second temperature change, and judges and reports a cause of the occurrence of the cooling abnormality of the electronic component.

[0010] In a third aspect of the present disclosure, there is provided a temperature monitoring program for causing one or more processors of a temperature monitoring device for monitoring the temperature of an electronic device including an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component for measuring the temperature of the electronic component, and a second temperature sensor attached to the cooling device for measuring the temperature of the cooling device to execute a first acquisition process, a second acquisition process, a third acquisition process, a first calculation process, a second calculation process, and a determination / report process. The first acquisition process acquires at least the temperature of the electronic component measured by the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods. The second acquisition process acquires the temperature of the electronic component measured by the first temperature sensor and the temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods. The third acquisition process acquires the temperature of the cooling device measured by at least the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods. The first calculation process calculates a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the predetermined plurality of measurement periods and the temperature of the electronic component acquired at the second time. The second calculation process calculates a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the predetermined plurality of measurement periods and the temperature of the cooling device acquired at the third time. The determination and reporting process determines the occurrence of a cooling abnormality in the electronic component based on the calculated first temperature change and the second temperature change, and determines and reports a cause of the occurrence of the cooling abnormality in the electronic component. The program may be recorded in a computer-readable storage medium. The storage medium may be a non-transitory medium such as a semiconductor memory, a hard disk, a magnetic recording medium, or an optical recording medium. The present disclosure may be embodied as a computer program product. Effect of the Invention

[0011] Each aspect of the present disclosure can contribute to determining that electronic components, modules, etc. that require heat dissipation in devices such as electronic equipment are not being cooled normally and to identifying the root cause of this. [Brief description of the drawings]

[0012] [Figure 1A] FIG. 1A is a diagram illustrating an example of the hardware configuration of a server device that can form the basis of the present disclosure. [Figure 1B] FIG. 1B is a diagram illustrating a cross section of the CPU portion shown in FIG. 1A. [Diagram 2] FIG. 2 is a diagram illustrating an example of a functional block of a temperature monitoring function realized by a monitoring device of a server device. [Figure 3A] FIG. 3A is a diagram illustrating, in graph form, the change over time of the CPU temperature TC and the heat sink temperature TH during the initial measurement period including times t1, t2, and t3. [Figure 3B] FIG. 3B is a graph illustrating the change over time of the CPU temperature TC' and the heat sink temperature TH' during the period including times t1, t2, and t3 in the second and subsequent measurement periods. [Figure 3C] FIG. 3C is a graph illustrating the change over time of the CPU temperature TC' and the heat sink temperature TH' during the period including times t1, t2, and t3 in the second and subsequent measurement periods. [Figure 3D] FIG. 3D is a graph illustrating the change over time of the CPU temperature TC' and the heat sink temperature TH' during the period including times t1, t2, and t3 in the second and subsequent measurement periods. [Figure 3E] FIG. 3E is a diagram illustrating, in the form of a graph, changes over time in the CPU temperature TC' and the heat sink temperature TH' during a period including times t1, t2, and t3 in the second and subsequent measurement periods. [Figure 4A]FIG. 4A is a flowchart showing the process of the temperature monitoring function shown in FIG. [Figure 4B] FIG. 4B is a flowchart showing the process of the temperature monitoring function shown in FIG. [Figure 4C] FIG. 4C is a flowchart showing the process of the temperature monitoring function shown in FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. However, the present disclosure is not limited by the embodiments described below. In addition, in each drawing, the same or corresponding elements are appropriately assigned the same reference numerals, and the same or corresponding processes and communications are appropriately assigned the same reference numerals. Furthermore, it is necessary to note that the drawings are schematic, and the dimensions between the elements and their ratios may differ from reality. In addition, the threshold values ​​shown below are examples, and may be changed as appropriate depending on the application and configuration of the embodiment according to the present disclosure. Furthermore, the threshold values ​​shown below are rough values ​​and should be understood as values ​​with some range, and there is no substantial difference between the descriptions "above threshold" and "greater than threshold", and there is no substantial difference between the descriptions "below threshold" and "below threshold".

[0014] FIG. 1A is a diagram illustrating an example of a hardware configuration of a server device 1 that can be the basis of this disclosure. FIG. 1B is a diagram illustrating a cross section of the CPU unit 10 shown in FIG. 1A. FIG. 2 is a diagram illustrating an example of a functional block of the temperature monitoring function unit 2 realized by the monitoring device 144 of the server device 1. The server device 1 is an example of an electronic device (electronic device) that uses a CPU (electronic component), such as a computer, a signal processing device, a sequencer, and a control device. As illustrated in FIG. 1A, the server device 1 adopts a configuration in which the CPU unit 10, the main storage device 142, the monitoring device 144, the auxiliary storage device 146, and the interface (IF (InterFace)) device 148 mounted on the motherboard 120 are connected to each other via a bus inside the server device 1 so that data can be input and output. However, the configuration shown in FIG. 1A can be applied not only to so-called server devices, but also to general devices that can process information, such as PCs (Personal Computers), PDAs (Personal Digital Assistants), and smartphones.

[0015] As shown in FIG. 1B, the CPU section 10 includes a CPU 100, grease 102, and a heat sink 104. The grease 102, which has high thermal conductivity and improves the efficiency of heat conduction from the CPU 100 to the heat sink 104, is applied to the upper surface of the CPU 100 mounted on the motherboard 120. That is, the CPU 100 contacts the heat sink 104 via the grease 102, and heat is efficiently conducted from the CPU 100 to the heat sink 104. As the CPU 100, one having an arrangement configuration shown in FIG. 1B can be used as an example, but the CPU 100 can also be mounted via a socket rather than directly mounted on the motherboard 120. Also, in the CPU section 10, instead of the heat sink 104, an air-cooled cooling system using a cooling fan or Peltier element, or a liquid-cooled cooling system using water as a refrigerant, may be used. That is, the heat sink 104 is an example of a cooling system for the CPU 100.

[0016] 1A and 1B, a temperature sensor 152 used to measure the temperature of the CPU 100 is attached, and a temperature sensor 150 used to measure the temperature of the heat sink 104 is attached. Output terminals of the temperature sensors 150, 152 are connected to an interface device 148 via wiring or the like, and the temperature sensors 150, 152 output the measured temperatures of the heat sink 104 and the CPU 100 to the interface device 148. A notification device 154 is further connected to the interface device 148 via a cable or the like, and notifies the user of the server device 1 of the occurrence of an abnormality, such as a decrease in the cooling efficiency of the CPU 100 in the CPU section 10 (hereinafter referred to as a "cooling abnormality") and displays the cause to the user.

[0017] The notification device 154 includes at least one of a light-emitting device and a sound output device, and a display (neither of which are shown). The light-emitting device includes a light-emitting element such as an LED that outputs the occurrence of a cooling abnormality in the CPU part 10 as a light signal, and the sound output device includes a speaker and a sound synthesizer that output the occurrence of a cooling abnormality in the CPU part 10 as a sound signal. The display device displays the cause of the cooling abnormality in the CPU part 10 to the user. The notification device 154 can be configured using the light-emitting element, sound output device, and display included in the server device 1.

[0018] 1A includes a volatile memory element such as a RAM (Random Access Memory) and a non-volatile memory element such as a ROM (Read Only Memory) and a flash memory. The non-volatile memory element of the main memory 142 stores, on a medium-term or long-term basis, programs including instructions for implementing the functions of the OS (operating system) 106, the functions of applications that run on the OS (operating system) 106, the functions of a boot loader, and the like, as well as data required for executing these programs.

[0019] 1A, OS106 includes a load execution function for increasing the utilization rate of CPU100 by executing a given program and applying a load, and a function for measuring the utilization rate of CPU100, which indicates the proportion of the processing time occupied by a program being executed by CPU100. The boot loader function includes a function for starting OS106 when the power supply of server device 1 is switched from an OFF state to an ON state. The volatile memory element of main storage device 142 temporarily stores data required for CPU100 to execute programs such as OS106, application programs, and the boot loader.

[0020] The CPU section 10 can include one or more CPUs (processors) and executes programs stored in the main storage device 142, the auxiliary storage device 146, etc. However, in the following description, for the sake of specificity and clarification, as shown in FIG. 1B, a specific example will be given in which the CPU section 10 includes only one CPU 100.

[0021] The auxiliary storage device 146 includes a nonvolatile storage device such as an HDD and an SSD (Solid State Drive). Like the nonvolatile storage element of the main storage device 142, the auxiliary storage device 146 can also store programs and data required for executing these programs. The auxiliary storage device 146 also includes a USB interface used for connecting a USB (Universal Serial Bus) memory or the like, and can write data to a USB device such as a USB memory and read data from the USB device.

[0022] The interface device 148 performs processing for interfacing between an input device such as a keyboard that accepts user operations and an output device such as a display that outputs information to the user, and the server device 1. The interface device 148 further accepts the values ​​of the temperatures of the heat sink 104 and the CPU 100 measured by the temperature sensors 150, 152, and outputs them to the monitoring device 144. The interface device 148 also displays on the display the occurrence of a cooling abnormality in the CPU 100 detected by the monitoring device 144 and its cause, and notifies the user.

[0023] The monitoring device 144 is also called a BMC (Baseboard Management Controller) and may be, for example, a so-called general-purpose one-chip microcomputer mounted on the motherboard 120. The monitoring device 144 includes, for example, one or more processors that perform processing independent of the CPU 100, and storage elements such as a ROM and a RAM (neither of which are shown in the figure). The ROM built into the monitoring device 144 can store a program including instructions for implementing the temperature monitoring function unit 2 shown in FIG. 2 and the function as a BMC.

[0024] The processor built in the monitoring device 144 executes a program stored in the ROM, and as a function of the BMC, manages the server device 1 and monitors events such as failures occurring inside the server device 1. The processor built in the monitoring device 144 also realizes the function of the temperature monitoring function unit 2 shown in FIG. 2. However, a program including instructions for realizing the function of the BMC and the temperature monitoring function unit 2 may be stored in the main storage device 142 or the auxiliary storage device 146. In this case, the monitoring device 144 realizes the function of the BMC and the temperature monitoring function unit 2 by executing a program stored in the monitoring device 144 or the auxiliary storage device 146. Note that the function of the BMC and the function of the temperature monitoring function unit 2 can be realized by dedicated hardware that does not involve the execution of a program. Also, the function of the BMC may be included in the OS 106, and the function of the temperature monitoring function unit 2 may be additionally included in the OS 106. In this case, the monitoring device 144 is omitted.

[0025] 2, the temperature monitoring function unit 2 includes a temperature information collecting unit 200, a temperature monitoring unit 202, a temperature change monitoring table database (temperature change monitoring table DB (Data Base)) 204, a cooling anomaly determining unit 212, and a cooling anomaly reporting unit 214. The temperature monitoring unit 202 includes a first factor counter 206, a second factor counter 208, and a third factor counter 210. With these components, the temperature monitoring function unit 2 measures the temperatures of the heat sink 104 and the CPU 100 using the temperature sensors 150, 152, and detects the occurrence of a cooling anomaly in the CPU 100.

[0026] The temperature monitoring function unit 2 determines which of the first to third factors is causing the occurrence of the cooling abnormality, and notifies the user of the server device 1 of the occurrence of the cooling abnormality and which of the first to third factors caused the cooling abnormality. When the server device 1 includes multiple CPU units 10, temperature sensors 150, 152 are connected to the CPUs 100 and heat sinks 104 of the multiple CPU units 10, respectively. In this case, the temperature monitoring function unit 2 measures the temperature using the temperature sensors 152, 150 attached to the heat sinks 104 and CPUs 100 included in each of the multiple CPU units 10, and detects the occurrence of a cooling abnormality in the CPUs 100.

[0027] The first cause of the cooling abnormality is a decrease in the efficiency of heat conduction due to an abnormality in the heat conduction from the CPU 100 to the heat sink 104. The second cause is a decrease in the efficiency of heat dissipation from the heat sink 104 to the surrounding air even though the heat conduction from the CPU 100 to the heat sink 104 is normal. The third cause of the cooling abnormality is the simultaneous occurrence of both the first and second causes.

[0028] First, the conditions under which it is determined that no cooling abnormality has occurred in the CPU 100 and the conditions under which it is determined that a cooling abnormality has occurred in the CPU 100 due to any of the first to third causes will be described with reference to Figs. 3A to 3E and Tables 1 to 4. The processes of measuring the temperatures of the CPU 100 and the heat sink 104 and determining whether or not there is a cooling abnormality by the temperature monitoring function unit 2 executed in the monitoring device 144 are performed during each of a number of predetermined measurement periods. Fig. 3A is a diagram illustrating, in the form of a graph, the changes over time in the temperature TC (unit: °C) of the CPU 100 and the temperature TH of the heat sink 104 during the period including times t1, t2, and t3 (first to third times) during the initial measurement period.

[0029] 3A includes temperature changes ΔTC (=TC1-TC2) and ΔTH (=TH1-TH2) of the CPU 100 and the heat sink 104 calculated in the first measurement period. The temperature changes ΔTC and ΔTH (first temperature change, second temperature change) calculated in the first measurement period are used as reference values ​​in the process of determining whether or not the cooling abnormality of the CPU 100 has occurred due to the first to third factors in the second and subsequent measurement periods. However, these reference temperature changes do not need to be calculated in only one measurement period, and for example, the average value of the temperature changes ΔTC and ΔTH calculated in a predetermined measurement period from the first to several times may be used as the reference.

[0030] 3B to 3E are diagrams illustrating, in the form of graphs, changes over time in temperature TC' of CPU 100 and temperature TH' of heat sink 104 during a period including times t1, t2, and t3 during the second or subsequent measurement period. Note that Figs. 3B to 3E further include temperature changes ΔTC and ΔTH used as references, and temperature changes ΔTC' (=TC1'-TC2') and ΔTH' (=TH1'-TH2') calculated during the second or subsequent measurement period.

[0031] In each of FIG. 3B to FIG. 3E, the dotted lines indicate the changes in temperatures TC', TH' when a cooling abnormality occurs in the CPU 100 due to any of the first to third factors, and the solid lines indicate the temperatures TC, TH shown in FIG. 3A. Needless to say, the times t1, t2, and t3 do not have to coincide in each of the multiple measurement periods, but the time intervals between the times t1 and t2 coincide, and the time intervals between the times t2 and t3 coincide. Here, a specific example is given in which the temperatures of the CPU 100 and the heat sink 104 are measured at three times t1, t2, and t3 in one measurement period. Meanwhile, the temperature measurement may be performed according to the time difference between the temperature change of the CPU 100 and the temperature change of the heat sink 104, and for example, the measurement of temperatures at four or more times and the processing based on the measured temperatures are within the scope of the present disclosure. More specifically, the temperature TC2 may be measured at the time t2, and the temperature TH2 may be measured at the time t2' slightly shifted from the time t2.

[0032] 3A is, for example, the timing when the power supply of the server device 1 installed at a predetermined location is first turned from OFF to ON and the OS 106 is started, and the CPU 100 is in a low usage and unloaded state. Alternatively, the initial measurement period is the timing when the power supply of the server device 1 is first turned from OFF to ON after maintenance of the server device 1 is completed, and the OS 106 is started, and the CPU 100 is in a unloaded state.

[0033] 3B to 3E are times when the CPU 100 is in an almost unloaded state, for example, late at night when the company in which the server device 1 is installed is not in business, and there are no or very few users using the server device 1. Alternatively, the second and subsequent measurement periods are times when the power of the server device 1 is turned from OFF to ON for some reason, such as after maintenance of the server device 1 is completed, and the OS 106 is started, and the CPU 100 is in an unloaded state.

[0034] The predetermined time in the first and second and subsequent measurement periods is time t1. Time t2 in these measurement periods is time t2 when the OS 106 increases the utilization rate of the CPU 100 after time t1, and a time Δt (unit: seconds) has elapsed since a load was applied. Time t3 is time when a predetermined time has elapsed since the OS 106 put the CPU 100 into an unloaded state after time t2.

[0035] Tables 1 to 4 show temperature change management tables corresponding to Figs. 3B to 3E, respectively. Each entry in the temperature change management tables shown in Tables 1 to 4 includes, in association with each other, reference temperature changes ΔTC, ΔTH, temperature changes ΔTC', ΔTH' calculated at each of the second and subsequent measurement timings, and differences α, β between the temperature changes ΔTC, ΔTH and the temperature changes ΔTC', ΔTH'. Note that α=ΔTC'-ΔTC and β=ΔTH'-ΔTH. Table 1 shows the above numerical values ​​when no cooling abnormality occurs in CPU 100. Tables 2 to 4 show the above numerical values ​​when a cooling abnormality occurs in CPU 100 due to each of the first to third factors.

[0036] In the calculation of the temperature changes ΔTC, ΔTH, ΔTC', ΔTH', it is not essential which temperature is subtracted from which other temperature. In other words, to give a specific example, it is not essential whether to subtract the temperature TC2 from the temperature TC1 or the temperature TC1 from the temperature TC2 to calculate the temperature difference ΔTC. In the description here, the reason for performing the former subtraction is to concretize and clarify the description of the invention. Similarly, it is also not essential whether to subtract the temperature changes ΔTC, ΔTH from the temperature changes ΔTC', ΔTH' or to subtract the temperature changes ΔTC', ΔTH' from the temperature changes ΔTC, ΔTH to calculate the differences α, β. Even if any other temperature is subtracted from any other temperature, the positive and negative of the threshold value described later used to determine the cooling abnormality of the CPU 100 is appropriately changed according to the relationship between these temperatures.

[0037] First, with reference to FIG. 3B and Table 1, a condition under which the temperature monitoring function unit 2 determines that no cooling abnormality has occurred in the CPU 100 will be described. As shown in FIG. 3A, when no cooling abnormality has occurred in the CPU 100, the temperature TC of the CPU 100 rises and reaches a peak between time t1 and time t2, and then drops between time t2 and time t3. The temperature of the heat sink 104 starts to rise with a delay from time t1, reaches a peak between time t2 and time t3, and then drops after time t3. In this way, a delay occurs between the change in the temperature TC of the CPU 100 and the change in the temperature TH of the heat sink 104.

[0038] [Table 1]

[0039] As shown by the dotted and solid lines in FIG. 3B, if the temperature TC of the CPU 100 and the temperature changes ΔTC', ΔTH' of the heat sink 104 are substantially the same as the changes shown in FIG. 3A (ΔTC ≒ ΔTC = TC1' - TC2', ΔTH ≒ = ΔTH' = TH2 - TH3') during the second and subsequent measurement periods, it can be determined that no cooling abnormality has occurred in the CPU 100. The temperature changes ΔTC', ΔTH' being substantially the same as the changes shown in FIG. 3A also means that the differences α, β between the reference temperature changes ΔTC, ΔTH and the temperature changes ΔTC', ΔTH' are within a range that can be considered to be within the error range (for example, about ±1°). As described above, when the differences α, β between the reference temperature changes ΔTC, ΔTH and the temperature changes ΔTC', ΔTH' during the second and subsequent measurement periods are small enough to be considered to be within the error range, the temperature monitoring function unit 2 can determine that no cooling abnormality has occurred in the CPU 100.

[0040] Next, with reference to Fig. 3C and Table 2, a condition will be described in which the temperature monitoring function unit 2 determines that a cooling abnormality of the CPU 100 has occurred due to an abnormality (first cause) in the heat conduction from the CPU 100 to the heat sink 104 via the grease 102. In the second and subsequent measurement periods, as shown in Fig. 3C, the temperature TC' of the CPU 100 shown by the dotted line is higher than the reference temperature TC shown by the solid line. On the other hand, the temperature TH' of the heat sink 104 shown by the dotted line is lower than the reference temperature TC shown by the solid line.

[0041] These mean that the temperature of the heat sink 104 remains low despite the steep temperature gradient from the CPU 100 to the heat sink 104, that is, an abnormality has occurred in the heat conduction from the CPU 100 via the grease 102, causing a cooling abnormality. This also means that the difference α exceeds the error range and becomes large on the positive side, and the difference β exceeds the error range and becomes large on the negative side. In such a case, the temperature monitoring function unit 2 can determine that a cooling abnormality of the CPU 100 has occurred due to the first factor.

[0042] [Table 2]

[0043] Next, referring to FIG. 3D and Table 3, the conditions under which the temperature monitoring function unit 2 determines that a cooling abnormality of the CPU 100 has occurred due to an abnormality (second cause) in which the efficiency of heat dissipation from the heat sink 104 to the surrounding air decreases will be described. In the second and subsequent measurement periods, as shown in FIG. 3D, the temperatures TC', TH' of the CPU 100 shown by the dotted lines are higher than the reference temperatures TC, TH shown by the solid lines. These indicate that the temperature gradient from the CPU 100 to the heat sink 104 is normal and there is no abnormality in the thermal conduction from the CPU 100 to the heat sink 104 via the grease 102, but the temperatures of the CPU 100 and the heat sink 104 are high, that is, the efficiency of heat dissipation from the heat sink 104 to the surrounding air has decreased. This also means that both the differences α and β exceed the error range and become large on the positive side. In such a case, the temperature monitoring function unit 2 can determine that a cooling abnormality of the CPU 100 has occurred due to the second cause.

[0044] [Table 3]

[0045] Next, with reference to FIG. 3E and Table 4, a condition under which the temperature monitoring function unit 2 determines that a cooling abnormality in the CPU 100 has occurred due to both the first and second factors (third factors) will be described. In the second and subsequent measurement periods, as shown in FIG. 3E, the temperature TC of the CPU 100 shown by the dotted line is higher than the reference temperature TC shown by the solid line, exceeding the margin of error. Meanwhile, the temperature TH of the heat sink 104 shown by the dotted line is the same as the reference temperature TH shown by the solid line, within the margin of error. This also means that the difference α exceeds the margin of error and becomes large on the positive side, while the difference β becomes 0 within the margin of error.

[0046] These indicate the possibility that an abnormality has occurred in the thermal conduction from the CPU 100 to the heat sink 104 via the grease 102, and that the efficiency of heat dissipation from the heat sink 104 to the surrounding air has decreased. In such a case, the temperature monitoring function unit 2 can determine that a cooling abnormality in the CPU 100 has occurred due to the third factor.

[0047] [Table 4]

[0048] Referring again to FIG. 2, the temperature information collecting unit 200 of the temperature monitoring function unit 2 acquires the utilization rate of the CPU 100 via the OS 106 when it is time to start the second or subsequent measurement period from the first. When the temperature information collecting unit 200 confirms that the utilization rate of the CPU 100 has continued to be sufficiently low and unloaded, it applies a load to the CPU 100 via the OS 106 for a period of time Δt. The temperature information collecting unit 200 acquires the utilization rate of the CPU 100 from the OS 106 to confirm that a load is being applied to the CPU 100 during this period. Furthermore, when this period ends, the temperature information collecting unit 200 reduces the utilization rate of the CPU 100 via the OS 106, returning it to an unloaded state.

[0049] As described above, the temperature information collecting unit 200 controls the load of the CPU 100 during each of the first, second and subsequent measurement periods. Furthermore, during each of the first, second and subsequent measurement periods, the temperature information collecting unit 200 acquires the temperatures TC1-TC3 (at least TC1, TC2), TC1'-TC3' (at least TC1', TC2') of the CPU 100 and the temperatures TH1-TH3 (at least TH2, TH3), TH1'-TH3' (at least TH2', TH3') of the heat sink 104 from the temperature sensor 150. The temperature information collecting unit 200 outputs the temperatures TC1-TC3, TC1'-TC3', TH1-TH3, TH1'-TH3' acquired during each of the first, second and subsequent measurement periods to the temperature monitoring unit 202.

[0050] The temperature monitoring unit 202 calculates temperature changes ΔTC, ΔTC', ΔTH, ΔTH' from temperatures TC1-TC3, TC1'-TC3', TH1-TH3, TH1'-TH3' input from the temperature information collecting unit 200 during each of the first to second and subsequent measurement periods, and further calculates their differences α and β. The temperature monitoring unit 202 associates the calculated temperature changes ΔTC, ΔTC', ΔTH, ΔTH' and differences α and β as shown in Tables 1-4, and stores them in the respective entries of the temperature change management table stored in the temperature change monitoring table DB204.

[0051] Furthermore, the temperature monitoring unit 202 processes the temperature change monitoring table stored in the temperature change monitoring table DB 204, and determines whether or not a cooling abnormality in the CPU 100 caused by the first to third factors has occurred. If the temperature monitoring unit 202 determines that a cooling abnormality in the CPU 100 caused by each of the first to third factors has occurred, the temperature monitoring unit 202 increments the count values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 (adds 1 to the measured value).

[0052] At the start of the first measurement period and when the temperature monitoring unit 202 notifies the user of the server device 1 of a cooling anomaly in the CPU 100, the temperature monitoring unit 202 sets the measurement values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 to 0. At the end of each of the first to second and subsequent measurement periods, the temperature monitoring unit 202 outputs the measurement values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 to the cooling anomaly determination unit 212.

[0053] The cooling anomaly determination unit 212 processes the measured values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210, and outputs information indicating whether a cooling anomaly has occurred in the CPU 100 and which of the first to third factors is the cause to the cooling anomaly notification unit 214. Note that first to third thresholds are set for the measured values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210, respectively. When any of the measured values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 reaches any of the first to third thresholds, the cooling anomaly determination unit 212 notifies the cooling anomaly notification unit 214 of the occurrence of a cooling anomaly in the CPU 100 and outputs the cause (any of the first to third thresholds) to the cooling anomaly notification unit 214.

[0054] When the cooling anomaly notification unit 214 is notified by the cooling anomaly determination unit 212 that a cooling anomaly has occurred in the CPU 100 and the cause thereof is input, it controls the notification device 154 to notify a user, such as the administrator of the server device 1, of the occurrence of the cooling anomaly by optical and audio signals. The cooling anomaly notification unit 214 further displays the cause of the occurrence of the cooling anomaly (any of the first to third causes) on a display to inform the user.

[0055] When the temperature monitoring function unit 2 notifies the user of the occurrence of a cooling abnormality and indicates the cause, the user performs work to eliminate the cause of the occurrence of the cooling abnormality. When the cooling abnormality of the CPU 100 caused by the first factor is notified, the user performs work such as bringing the CPU 100 and the heat sink 104 into closer contact with each other, changing the mounting position of the heat sink 104, and reapplying the grease 102. When the cooling abnormality of the CPU 100 caused by the second factor is notified, the user performs work to improve the cooling effect of the heat sink 104, such as cleaning the periphery of the heat sink 104. When the cooling abnormality of the CPU 100 caused by the third factor is notified, the user performs work (work to eliminate the first and second factors) such as bringing the CPU 100 and the heat sink 104 into closer contact with each other, changing the mounting position of the heat sink 104, reapplying the grease 102, and cleaning the periphery of the heat sink 104.

[0056] Hereinafter, the processing of the temperature monitoring function unit 2 shown in Fig. 2 will be described with reference to a flowchart. Fig. 4A to Fig. 4C are flowcharts showing the processing of the temperature monitoring function unit 2 shown in Fig. 2. As shown in Fig. 4A, in S100, after the user of the server device 1 sets the server device 1 to a company or the like, first, the power supply of the server device 1 is switched from an OFF state to an ON state.

[0057] In S102, a boot loader stored in the main memory device 142 (FIG. 1A) of the server device 1 reads out an OS similarly stored in the monitoring device 144, loads it into the main memory device 142, and starts the OS 106. The OS 106 starts the temperature monitoring function unit 2, and the temperature change monitoring table DB 204 of the temperature monitoring function unit 2 clears the count values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 to zero (sets them to 0). The temperature monitoring function unit 2 proceeds to the processing of S2 shown in FIGS. 4B and 4C.

[0058] As shown in Fig. 4B, in S200, the temperature information collecting unit 200 of the temperature monitoring function unit 2 acquires the utilization rate of the CPU 100 via the OS 106, and compares the acquired utilization rate with a predetermined threshold value. If the temperature information collecting unit 200 determines as a result of this comparison that the acquired utilization rate is equal to or less than the predetermined threshold value, the CPU 100 is in an unloaded state, and the first or i (2 <= i)th measurement period has started, the process proceeds to S202 (Y in S200). If the temperature information collecting unit 200 determines that the acquired utilization rate is greater than the predetermined threshold value, and the CPU 100 is not in an unloaded state, the process remains at S200 (N in S200).

[0059] In S202, when a predetermined time has elapsed after the start of the first measurement period and time t1 has arrived, the temperature information collecting unit 200 reads out temperatures TC1, TH1 of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. When a predetermined time has elapsed after the start of the i-th measurement period and time t1 has arrived, the temperature information collecting unit 200 reads out temperatures TC1', TH1' of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. However, reading out temperature TH1' of the heat sink 104 is not essential. In S204, the temperature information collecting unit 200 causes the CPU 100 to execute a given program via the OS 106, and increases the utilization rate of the CPU 100 to apply a load between times t1 and t2, that is, for a period of time Δt.

[0060] In S206, at time t2 during the first measurement period, the temperature information collection unit 200 reads out temperatures TH2, TC2 of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. At time t2 during the i-th measurement period, the temperature information collection unit 200 reads out temperatures TH2', TC2' of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. Furthermore, the temperature information collection unit 200 causes the CPU 100 to stop execution of a given program via the OS 106, and brings the utilization rate of the CPU 100 close to 0 between times t2 and t3, placing it in an unloaded state.

[0061] In S208, at time t3 during the first measurement period, the temperature information collection unit 200 reads out temperatures TH3, TC3 of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. Also, at time t3 during the i-th measurement period, the temperature information collection unit 200 reads out temperatures TH3', TC3' of the heat sink 104 and the CPU 100 from the temperature sensors 150, 152. However, reading out temperature TC3' from the CPU 100 is not essential.

[0062] In S210, the temperature monitoring unit 202 calculates the temperature changes ΔTC (=TC1-TC2) and ΔTH (=TH1-TH2) during the first measurement period (see FIG. 3A). Also, during the i-th measurement period, the temperature monitoring unit 202 calculates the temperature changes ΔTC' (=TC1'-TC2') and ΔTH' (=TC1'-TC2'). In S212, the temperature monitoring unit 202 calculates the differences α (=ΔTC'-ΔTC) and β (=ΔTH'-ΔTH) between the temperature changes ΔTC, ΔTH and the temperature changes ΔTC', ΔTH'. The temperature monitoring unit 202 associates the temperature changes ΔTC, ΔTC', ΔTH, ΔTH' and the differences α and β as shown in Tables 1 to 4, creates one entry to be included in the temperature change monitoring table, and stores it in the temperature change monitoring table DB 204.

[0063] In S214, the temperature information collecting unit 200 judges whether the processes of S200 to S212 have been performed in the first measurement period. If the processes of S200 to S212 have been performed in the first measurement period (Y in S214), the temperature monitoring function unit 2 proceeds to S220. If the processes of S200 to S212 have been performed in the i-th measurement period (N in S214), the temperature monitoring function unit 2 proceeds to S230 (FIG. 4C).

[0064] In S220, the temperature information collecting unit 200 judges whether the second measurement period has started. When the second measurement period has started (Y in the process of S220), the temperature monitoring function unit 2 returns to the process of S200. When the second measurement period has not started, the temperature monitoring function unit 2 remains in the process of S220.

[0065] As shown in Fig. 4C, in S230, the temperature monitoring unit 202 judges whether the difference α calculated in the i-th measurement period is higher than a predetermined threshold value of +3°C (α≦+3). When the difference α calculated in the i-th measurement period is +3°C or less, the temperature monitoring unit 202 judges that no cooling abnormality has occurred in the CPU 100, as explained with reference to Figs. 3A, 3B and Tables 1 and 2, and returns to the process of S200 (see Fig. 3B and Table 1). When the difference α is higher than +3°C, the temperature monitoring function unit 2 proceeds to the process of S232.

[0066] In S232, the temperature monitoring unit 202 judges whether the difference β calculated in the i-th measurement period is lower than a predetermined threshold value of −2° C. If the difference α calculated in the i-th measurement period is lower than −2° C. (Y in the process of S232), the temperature monitoring function unit 2 proceeds to the process of S234. If the difference α is higher than −2° C. (N in the process of S232), the temperature monitoring function unit 2 proceeds to the process of S240.

[0067] In S234, the temperature monitoring unit 202 determines that the cooling abnormality of the CPU 100 has occurred due to the first cause described with reference to Fig. 3C and Table 2, that is, a decrease in efficiency of heat conduction caused by an abnormality in heat conduction from the CPU 100 to the heat sink 104. In S236, the temperature monitoring unit 202 increments the count value of the first cause counter 206.

[0068] In S240, the temperature monitoring unit 202 judges whether the difference β calculated in the i-th measurement period is higher than a predetermined threshold value of +2° C. If the difference β is higher than +2° C. (Y in the process of S240), the temperature monitoring function unit 2 proceeds to the process of S242. If the difference β is equal to or lower than +2° C. (N in the process of S240), the temperature monitoring function unit 2 proceeds to the process of S250.

[0069] In S242, the temperature monitoring unit 202 determines that the cooling abnormality of the CPU 100 has occurred due to the second cause described with reference to Fig. 3D and Table 3, that is, a decrease in efficiency of heat dissipation from the heat sink 104 to the surrounding air despite normal heat conduction from the CPU 100 to the heat sink 104. In S244, the temperature monitoring unit 202 increments the count value of the second cause counter 208.

[0070] In S250, the temperature monitoring unit 202 determines that the cooling abnormality of the CPU 100 has occurred due to the third cause described with reference to Fig. 3E and Table 4, i.e., due to both the first and second causes. In S252, the temperature monitoring unit 202 increments the count value of the third cause counter 210.

[0071] In S260, the cooling anomaly determination unit 212 determines whether or not one or more of the count values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 exceed a predetermined threshold value for each of the first factor counter 206, the second factor counter 208, and the third factor counter 210. When one or more of the count values ​​of the first factor counter 206, the second factor counter 208, and the third factor counter 210 exceed a threshold value, the cooling anomaly determination unit 212 determines that the cooling anomaly of the CPU 100 has actually occurred, rather than that it has temporarily occurred due to some cause (Y in the process of S260). Furthermore, the temperature monitoring function unit 2 proceeds to the process of S262. When the cooling anomaly determination unit 212 determines that the cooling anomaly of the CPU 100 has not actually occurred (N in the process of S260), the temperature monitoring function unit 2 returns to the process of S200.

[0072] In S262, the cooling abnormality notification unit 214 notifies the user of the server device 1 of the occurrence of a cooling abnormality in the CPU 100 and its cause by means of an optical signal, an audio signal, and a display on the display.

[0073] According to the processing of the temperature monitoring function unit 2 described above, when a cooling abnormality occurs in the CPU 100, the fact and the cause can be appropriately notified to the user of the server device 1. Therefore, the user of the server device 1 can appropriately and fundamentally eliminate the cause of the cooling abnormality in the CPU 100, and can prevent the server device 1 from slowing down or stopping its operation. In addition, the temperature monitoring function unit 2 can be applied to a CPU unit 10 that uses a cooling device that cools the CPU 100 using a fan or the like instead of the heat sink 104. Even in such a case, the temperature monitoring function unit 2 prevents the noise and power consumption of the cooling device from increasing compared to when a cooling abnormality in the CPU 100 is prevented by constantly rotating a cooling fan at high speed.

[0074] A part or all of the above-described embodiments may be described as, but is not limited to, the following supplementary notes. [Appendix 1] A temperature monitoring device that includes one or more processors and monitors the temperature of an electronic device comprising: an electronic component; a cooling device attached to the electronic component; a first temperature sensor attached to the electronic component and measuring the temperature of the electronic component; and a second temperature sensor attached to the cooling device and measuring the temperature of the cooling device. The one or more processors of the temperature monitoring device are configured to acquire a temperature of the electronic component measured by at least the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods, acquire a temperature of the electronic component measured by the first temperature sensor and a temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods, acquire a temperature of the cooling device measured by at least the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods, calculate a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the plurality of predetermined measurement periods and the temperature of the electronic component acquired at the second time, calculate a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the plurality of predetermined measurement periods and the temperature of the cooling device acquired at the third time, determine an occurrence of a cooling abnormality of the electronic component based on the calculated first temperature change and the second temperature change, and determine and report a cause of the occurrence of the cooling abnormality of the electronic component. [Appendix 2] 2. The temperature monitoring device according to claim 1, which reports the occurrence of a determined cooling abnormality in the electronic component and the cause of the determined cooling abnormality in the electronic component. [Appendix 3] 3. The temperature monitoring device of claim 1, wherein at the first time, the electronic component is in an unloaded state, from the first time until the second time which is a predetermined first time, the electronic component is in a loaded state, and from the second time until the third time which is a predetermined second time, the electronic component is in an unloaded state, and during each of the multiple measurement periods, the temperature of the electronic component is highest between the first time and the second time, and the temperature of the cooling device is highest between the second time and the third time. [Appendix 4] A temperature monitoring device as described in any of Appendices 1 to 3, which determines that a cooling abnormality has occurred in the electronic component when a value obtained by subtracting a reference first temperature change obtained from the first temperature change acquired in one or more predetermined measurement periods from the first temperature change obtained in each of the predetermined plurality of measurement periods is higher than a predetermined first threshold value. [Appendix 5] A temperature monitoring device as described in any of Appendices 1 to 4, which determines that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the electronic component to the cooling device when a value obtained by subtracting a reference second temperature change obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the predetermined plurality of measurement periods is lower than a predetermined second threshold value. [Appendix 6] A temperature monitoring device as described in any of Appendices 1 to 5, which determines that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the cooling device to the surroundings when a value obtained by subtracting a reference second temperature change obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the predetermined plurality of measurement periods is lower than a predetermined second threshold value. [Appendix 7] A temperature monitoring device as described in any of Appendices 1 to 6, which determines that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the electronic component to the cooling device and a decrease in efficiency of heat conduction from the cooling device to the surroundings when a value obtained by subtracting the reference second temperature change obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the predetermined plurality of measurement periods is higher than a predetermined second threshold value. [Appendix 8] A temperature monitoring device as described in any one of appendices 1 to 7, which determines that a cooling anomaly has occurred in the electronic component due to a decrease in efficiency of heat conduction from the electronic component to the cooling device, which reports the occurrence of the determined cooling anomaly in the electronic component and the cause of the determined cooling anomaly in the electronic component when any of the following occurs more than a predetermined number of times: determining that a cooling anomaly in the electronic component has occurred due to a decrease in efficiency of heat conduction from the electronic component to the cooling device, which has occurred due to a decrease in efficiency of heat conduction from the cooling device to the surroundings, and which has occurred due to both a decrease in efficiency of heat conduction from the electronic component to the cooling device and a decrease in efficiency of heat conduction from the cooling device to the surroundings. [Appendix 9] A temperature monitoring method for monitoring the temperature of an electronic device including an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component for measuring a temperature of the electronic component, and a second temperature sensor attached to the cooling device for measuring a temperature of the cooling device. The temperature monitoring method includes a first acquisition step, a second acquisition step, a third acquisition step, a first calculation step, a second calculation step, and a determination / notification step. The first acquisition step acquires at least the temperature of the electronic component measured by the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods. The second acquisition step acquires the temperature of the electronic component measured by the first temperature sensor and the temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods. The third acquisition step acquires at least the temperature of the cooling device measured by the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods. The first calculation step calculates a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the predetermined plurality of measurement periods and the temperature of the electronic component acquired at the second time. The second calculation step calculates a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the predetermined plurality of measurement periods and the temperature of the cooling device acquired at the third time. The determination and notification step determines the occurrence of a cooling abnormality in the electronic component based on the calculated first temperature change and the second temperature change, and determines and notifies a cause of the occurrence of the cooling abnormality in the electronic component. [Appendix 10] A temperature monitoring program that causes one or more processors of a temperature monitoring device that monitors the temperature of an electronic device, the temperature monitoring device including an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component for measuring a temperature of the electronic component, and a second temperature sensor attached to the cooling device for measuring a temperature of the cooling device, to execute a first acquisition process, a second acquisition process, a third acquisition process, a first calculation process, a second calculation process, and a determination / notification process. The first acquisition process acquires at least the temperature of the electronic component measured by the first temperature sensor at a first time included in each of a plurality of predetermined measurement periods. The second acquisition process acquires the temperature of the electronic component measured by the first temperature sensor and the temperature of the cooling device measured by the first temperature sensor at a second time after the first time included in each of the plurality of predetermined measurement periods. The third acquisition process acquires the temperature of the cooling device measured by at least the second temperature sensor at a third time after the second time included in each of the plurality of predetermined measurement periods. The first calculation process calculates a first temperature change between the temperature of the electronic component acquired at least at the first time in each of the predetermined plurality of measurement periods and the temperature of the electronic component acquired at the second time. The second calculation process calculates a second temperature change between the temperature of the cooling device acquired at least at the second time in each of the predetermined plurality of measurement periods and the temperature of the cooling device acquired at the third time. The determination and notification process determines the occurrence of a cooling abnormality in the electronic component based on the calculated first temperature change and the second temperature change, and determines and notifies a cause of the occurrence of the cooling abnormality in the electronic component. It goes without saying that combinations of the various aspects of the appendix of the present disclosure, or arbitrary combinations of the various elements described in the various aspects and embodiments (including non-selection of some elements) can be made by those skilled in the art at any time in accordance with the basic concept of the present disclosure. Alternatively, the present disclosure also includes the temperature monitoring method and the temperature monitoring program described in appendix 9 and 10 being implemented in the configuration of the temperature monitoring device described in appendix 2 to 8.

[0075] The disclosures of the above cited patent documents are incorporated herein by reference. Within the framework of the entire disclosure of the present invention (including the scope of claims), modifications and adjustments of the embodiments or examples are possible based on the basic technical ideas. Furthermore, within the framework of the entire disclosure of the present invention, various combinations or selections (including partial deletions) of various disclosed elements (including each element of each claim, each element of each embodiment or example, each element of each drawing, etc.) are possible. In other words, the present invention naturally includes various modifications and corrections that a person skilled in the art would be able to make in accordance with the entire disclosure and technical ideas, including the scope of claims. In particular, the numerical ranges described in this specification should be interpreted as specifically describing any numerical value or small range included in the range, even if not otherwise specified. Furthermore, the disclosures of the above cited documents are deemed to be included in the disclosures of this application, which may be used in part or in whole in combination with the disclosures of this specification as part of the disclosure of the present invention, in accordance with the spirit of the present invention, as necessary. [Explanation of symbols]

[0076] 1. Server device 10 CPU section 100 CPU 102 Grease 104 Heat sink 106 OS 120 Motherboard 142 Main storage 144 Monitoring equipment 146 Auxiliary storage 148 Interface Device 150,152 Temperature Sensor 154 Notification device 2 Temperature monitoring function section 200 Temperature Information Collection Department 202 Temperature monitoring section 204 Temperature change monitoring table database (DB) 206 First factor counter 208 Second factor counter 210 Third factor counter 212 Cooling abnormality determination section 214 Cooling Abnormality Reporting Department

Claims

1. A temperature monitoring device for monitoring a temperature of an electronic device, the temperature monitoring device including one or more processors, the electronic device including: an electronic component; a cooling device attached to the electronic component; a first temperature sensor attached to the electronic component for measuring a temperature of the electronic component; and a second temperature sensor attached to the cooling device for measuring a temperature of the cooling device, The one or more processors: acquiring a temperature of the electronic component measured by at least the first temperature sensor at a first time point included in each of a plurality of predetermined measurement periods and during which the electronic component is in an unloaded state; At a second time when a predetermined time has elapsed since a load was applied to the electronic component after the first time included in each of the plurality of predetermined measurement periods, a temperature of the electronic component measured by the first temperature sensor and a temperature of the cooling device measured by the second temperature sensor are acquired; At a third time when a predetermined time has elapsed since the electronic component was placed in an unloaded state after the second time included in each of the plurality of predetermined measurement periods, a temperature of the cooling device measured by at least the second temperature sensor is acquired; calculating a first temperature change between at least the temperature of the electronic component acquired at the first time and the temperature of the electronic component acquired at the second time during each of the predetermined plurality of measurement periods; calculating a second temperature change between at least the temperature of the cooling device acquired at the second time and the temperature of the cooling device acquired at the third time during each of the predetermined plurality of measurement periods; calculating a first difference between the first temperature change and a temperature change of a first reference, calculating a second difference between the second temperature change and a temperature change of a second reference, determining whether a cooling abnormality has occurred in the electronic component based on the calculated first difference and second difference, and determining and reporting a cause of the occurrence of the cooling abnormality in the electronic component. A temperature monitoring device configured to perform a process.

2. Reporting the occurrence of the determined cooling abnormality of the electronic component and the cause of the determined cooling abnormality of the electronic component.

2. The temperature monitoring device of claim 1.

3. In each of the plurality of measurement periods, the temperature of the electronic component is highest between the first time and the second time, and the temperature of the cooling device is highest between the second time and the third time.

3. The temperature monitoring device of claim 2.

4. When a value obtained by subtracting a reference first temperature change obtained from the first temperature change acquired in one or more predetermined measurement periods from the first temperature change obtained in each of the plurality of predetermined measurement periods is higher than a predetermined first threshold value, it is determined that a cooling abnormality has occurred in the electronic component.

4. The temperature monitoring device of claim 3.

5. When a value obtained by subtracting a reference second temperature change obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the plurality of predetermined measurement periods is lower than a predetermined second threshold value, it is determined that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the electronic component to the cooling device.

5. The temperature monitoring device of claim 4.

6. When a value obtained by subtracting a reference second temperature change obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the plurality of predetermined measurement periods is higher than a predetermined second threshold value, it is determined that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the cooling device to the surroundings.

5. The temperature monitoring device of claim 4.

7. When a value obtained by subtracting the second temperature change of a reference obtained from the second temperature change acquired in one or more predetermined measurement periods from the second temperature change obtained in each of the plurality of predetermined measurement periods is lower than a predetermined second threshold value, it is determined that a cooling abnormality of the electronic component has occurred due to a decrease in efficiency of heat conduction from the electronic component to the cooling device and a decrease in efficiency of heat conduction from the cooling device to the surroundings.

7. The temperature monitoring device of claim 6.

8. When any one of the following occurs more than a predetermined number of times: determining that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the electronic component to the cooling device; that a cooling abnormality has occurred in the electronic component due to a decrease in efficiency of heat conduction from the cooling device to the surroundings; and that a cooling abnormality has occurred in the electronic component due to both a decrease in efficiency of heat conduction from the electronic component to the cooling device and a decrease in efficiency of heat conduction from the cooling device to the surroundings, reporting the occurrence of the determined cooling abnormality in the electronic component and the cause of the determined cooling abnormality in the electronic component.

8. The temperature monitoring device of claim 7.

9. A temperature monitoring method using a temperature monitoring device that includes one or more processors and monitors the temperature of an electronic device having an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component and measuring a temperature of the electronic component, and a second temperature sensor attached to the cooling device and measuring a temperature of the cooling device, comprising: The one or more processors, a first acquisition step of acquiring a temperature of the electronic component measured by at least the first temperature sensor at a first time point included in each of a plurality of predetermined measurement periods and during which the electronic component is in an unloaded state; a second acquisition step of acquiring a temperature of the electronic component measured by the first temperature sensor and a temperature of the cooling device measured by the second temperature sensor at a second time when a predetermined time has elapsed since a load was applied to the electronic component after the first time included in each of the plurality of predetermined measurement periods; a third acquisition step of acquiring a temperature of the cooling device measured by at least the second temperature sensor at a third time when a predetermined time has elapsed since the electronic component was placed in an unloaded state after the second time included in each of the plurality of predetermined measurement periods; a first calculation step of calculating a first temperature change between at least the temperature of the electronic component acquired at the first time and the temperature of the electronic component acquired at the second time during each of the predetermined plurality of measurement periods; a second calculation step of calculating a second temperature change between at least the temperature of the cooling device acquired at the second time and the temperature of the cooling device acquired at the third time during each of the predetermined plurality of measurement periods; a determining / reporting step of calculating a first difference between the first temperature change and a temperature change of a first reference, calculating a second difference between the second temperature change and a temperature change of a second reference, determining whether a cooling abnormality has occurred in the electronic component based on the calculated first difference and second difference, and determining and reporting a cause of the occurrence of the cooling abnormality in the electronic component; A temperature monitoring method that performs the above.

10. a temperature monitoring device for monitoring a temperature of an electronic device, the electronic device including an electronic component, a cooling device attached to the electronic component, a first temperature sensor attached to the electronic component for measuring a temperature of the electronic component, and a second temperature sensor attached to the cooling device for measuring a temperature of the cooling device; a first acquisition process for acquiring a temperature of the electronic component measured by at least the first temperature sensor at a first time when the electronic component is in an unloaded state, the first acquisition process being included in each of a plurality of predetermined measurement periods; a second acquisition process for acquiring a temperature of the electronic component measured by the first temperature sensor and a temperature of the cooling device measured by the second temperature sensor at a second time that is a predetermined time after a load is applied to the electronic component and that is subsequent to the first time included in each of the plurality of predetermined measurement periods; a third acquisition process for acquiring a temperature of the cooling device measured by at least the second temperature sensor at a third time that is a predetermined time elapsed after the electronic component is placed in an unloaded state and that is subsequent to the second time included in each of the plurality of predetermined measurement periods; a first calculation process for calculating a first temperature change between at least the temperature of the electronic component acquired at the first time and the temperature of the electronic component acquired at the second time during each of the predetermined measurement periods; a second calculation process for calculating a second temperature change between at least the temperature of the cooling device acquired at the second time and the temperature of the cooling device acquired at the third time during each of the predetermined measurement periods; a determination and reporting process for calculating a first difference between the first temperature change and a temperature change of a first reference, calculating a second difference between the second temperature change and a temperature change of a second reference, determining the occurrence of a cooling abnormality of the electronic component based on the calculated first difference and second difference, and determining and reporting a cause of the occurrence of the cooling abnormality of the electronic component; A temperature monitoring program that runs

Citation Information

Patent Citations

  • Information processor and cooling performance determination method

    JP2010152740A

  • Electronic circuit cooling unit abnormality inspection system

    JP2012064975A

  • Motor driving device detecting abnormality of heat radiation performance of heat sink and detection method

    JP2017028833A

  • Vehicle control device

    JP2022114802A

  • Information processor, and determination method

    JP2022164131A