Computer temperature anomaly processing method and system
By building a joint optimization model in the computer system, the temperature, load and heat dissipation capacity data of the computing components are collected and analyzed in real time, and the coordinated adjustment of task migration and the heat dissipation system is achieved, which solves the problem of the independence of load balancing and temperature control and improves the stability and reliability of the system.
Patent Information
- Application Number
- CN202510776289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
AI Technical Summary
In existing computer systems, load balancing and temperature control are independent of each other. As a result, when the temperature of computing components is abnormal, the load balancing system cannot be adjusted in time, and the cooling system is also difficult to adjust accurately, affecting the stability and reliability of the system.
By collecting the temperature, load and heat dissipation capacity data of computing components in real time, a joint optimization model is built to achieve coordinated adjustment of task migration and cooling system, dynamically optimize load and heat dissipation to ensure stable operation of the system under different working conditions.
It achieves stable operation of the computer system under temperature and load balance conditions, improves the system's performance, stability and reliability, optimizes the configuration of heat dissipation resources, and reduces energy consumption and equipment loss.
Smart Images

Figure CN120669828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and system for processing temperature anomalies of a computer. Background Art
[0002] In today's digital age, computer systems are widely used in various fields, especially in scenarios such as server clusters and high-performance computing systems. The requirements for computing resource load and temperature control are becoming increasingly stringent. These systems often need to run at high load for a long time to process massive amounts of data and complex computing tasks. As the performance of computing components continues to improve, the heat they generate also increases sharply. Temperature control has become a key factor in ensuring stable system operation.
[0003] At present, the load balancing strategy and temperature control in computer systems are independent of each other. Load balancing mainly focuses on reasonably distributing computing tasks to various computing components to improve the overall computing efficiency and resource utilization of the system. Its core lies in the reasonable distribution of tasks, while the temperature conditions of computing components are less considered. Temperature control mainly relies on heat dissipation devices, such as cooling fans, heat sinks, etc., to maintain the temperature of computing components within an acceptable range through passive cooling. This separation mode has many disadvantages. When the temperature of a computing component is abnormal, the load balancing system cannot perceive and make adjustments in time, resulting in tasks still being concentrated on components with too high temperatures, further exacerbating the temperature increase. At the same time, due to the lack of effective coordination with load balancing, the heat dissipation system can only perform passive cooling according to preset parameters. It is difficult to make precise adjustments according to the actual temperature and load conditions, resulting in poor heat dissipation effect, which not only reduces the system's operating efficiency, but also increases the risk of component damage due to high temperature, seriously affecting the stability and reliability of the system. Summary of the Invention
[0004] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide a computer temperature anomaly processing method and system. It can build a joint optimization model by collecting the temperature, load and heat dissipation capacity data of the computing components in real time. When a temperature anomaly occurs, the task is migrated according to the model, and the task and the heat dissipation system are adjusted in coordination to achieve dynamic optimization of load and heat dissipation. Through cyclic monitoring and continuous optimization, it ensures that the computer system is always in a stable operating state with balanced temperature and load, effectively solving the problem caused by the independence of load balancing and temperature control in traditional technologies, and improving the performance, stability and reliability of the system.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a method and system for processing temperature anomalies in a computer, the method comprising the following specific steps:
[0006] Real-time data collection: Deploy temperature sensors at key computer locations, and utilize load monitoring modules and heat dissipation capacity detection equipment to collect temperature, load, and heat dissipation capacity data in real time.
[0007] Joint optimization model construction: Based on the collected data, comprehensive consideration of various factors of computing components, with system performance and temperature balance as the goal, a joint optimization model describing task allocation and temperature control strategies is constructed, and task migration rules and heat dissipation control parameters are formulated;
[0008] Temperature anomaly determination and task migration: The collected component temperature is compared with the preset threshold. After determining the component with abnormal temperature, the system migrates some tasks of the abnormal component to appropriate components based on model rules and the load and heat dissipation of other components, and records the relevant information.
[0009] Coordinated adjustment of tasks and heat dissipation: Monitors component temperature changes after task migration, adjusts the priority and number of unmigrated tasks based on model strategies, and coordinates the cooling system's workload based on component temperature and load to achieve load and heat dissipation optimization.
[0010] Loop monitoring and continuous optimization: Update system information and evaluate model parameters based on new data, dynamically monitoring and optimizing component temperature and load.
[0011] Furthermore, in the step of constructing the joint optimization model, based on the acquired temperature, load and heat dissipation capacity data of each computing component, the temperature threshold, load bearing capacity and heat dissipation capacity factors of each computing component are comprehensively considered, and by analyzing the correlation between the data, the operating status and requirements of the system under different data combinations are determined, and a dynamic load balancing and temperature control joint optimization model is constructed. During the construction process, mathematical relationships are established to describe the task allocation and temperature control strategies, and corresponding task migration rules are formulated for different temperature and load conditions, and the heat dissipation intensity adjustment coefficient of the heat dissipation system is determined.
[0012] Furthermore, in the joint optimization model construction step, corresponding task migration rules are formulated according to different temperature and load conditions and the heat dissipation intensity adjustment coefficient of the heat dissipation system is determined. The task migration rule determination formula is: Among them, P ij It represents the tendency to migrate tasks from the i-th temperature abnormal component to the j-th target component. The larger the value, the more inclined to migrate tasks from the i-th component to the j-th component. max It is the upper limit of the unified temperature threshold preset for all computing components. It is a fixed value set according to the hardware characteristics of the components and the requirements for safe operation. i is the real-time temperature of the i-th computing component, S j is the heat dissipation capacity of the jth target computing component, L jis the real-time load of the jth target computing component, ∈ is a very small positive number used to avoid the denominator being zero, and is a fixed constant set artificially. The heat dissipation intensity adjustment coefficient of the heat dissipation system is: Among them, R i is the heat dissipation intensity adjustment coefficient of the i-th computing component, which is used to control the working intensity of the heat dissipation system. i is the real-time temperature of the i-th computing component, T max is the upper temperature threshold preset for all computing components, T min It is the preset unified lower temperature threshold for all computing components. It is a fixed value set according to the minimum temperature requirement for normal operation of the components. α and β are adjustment coefficients, which are fixed constants determined through experiments and system debugging. They are used to adjust the sensitivity and base value of the heat dissipation intensity regulation.
[0013] Furthermore, in the temperature anomaly determination and task migration step, part of the tasks of the abnormal component are migrated to the appropriate component and the relevant information is recorded. Specifically, according to the established task migration rules, combined with the load conditions and heat dissipation capacity data of other computing components in the system, the target components with relatively low load and strong heat dissipation capacity are screened out, and the number of migration tasks is calculated through the task migration amount formula. The tasks are migrated to the selected target component for execution according to the task segmentation and transmission mechanism, and the task information is recorded.
[0014] Furthermore, in the temperature anomaly determination and task migration step, the number of migration tasks is calculated using a task migration amount formula, which is: Among them, M i is the number of tasks that need to be migrated for the i-th temperature abnormal component, T i is the real-time temperature of the i-th computing component, T threshold,i is the preset temperature threshold of the i-th computing component, C i is the current task processing capability coefficient of the i-th computing component, which is evaluated based on the component performance and current operating status. max It is the upper limit of the temperature threshold preset by all computing components, a fixed value, T j is the real-time temperature of the jth computing component, S j is the heat dissipation capacity of the jth computing component, N total is the total number of tasks currently performed by the i-th computing component.
[0015] Furthermore, in the task and heat dissipation coordinated adjustment step, the temperature changes of each computing component are continuously monitored in real time. If it is found that the temperature of some components rises too fast due to task migration, or the temperature is close to its preset temperature threshold, the priority of the non-migrated tasks is dynamically adjusted according to the joint optimization model, and tasks with lower priorities are given priority for subsequent migration, and the total amount of migrated tasks is controlled. At the same time, according to the current temperature and load conditions of each computing component, the working intensity of the heat dissipation system is adjusted. For components with higher temperatures and heavier loads, the corresponding cooling fan speed is increased, or the cooling power of the heat dissipation equipment is increased. For components with lower temperatures and lighter loads, the working intensity of the heat dissipation system is reduced to reduce energy consumption.
[0016] Furthermore, in the task and heat dissipation coordinated adjustment step, the priority of the non-migrated task is dynamically adjusted according to the joint optimization model, and the priority calculation formula is: Among them, Q k,i is the priority adjustment coefficient of the kth unmigrated task in the i-th computing component, which is used to adjust the task priority. The larger the value, the higher the priority. max is the upper temperature threshold preset for all computing components, T i is the real-time temperature of the i-th computing component, T i,k is the temperature increment expected to be generated by the execution of the kth task in the i-th computing component, m is the total number of unmigrated tasks in the i-th computing component, and γ is the priority adjustment sensitivity coefficient.
[0017] Furthermore, in the task and heat dissipation coordinated adjustment step, the working intensity of the heat dissipation system is adjusted, and the adjustment formula is: Among them, R i ′ is the heat dissipation intensity adjustment coefficient of the i-th computing component after adjustment, R i is the heat dissipation intensity adjustment coefficient of the i-th computing component before adjustment, T i ′ is the real-time temperature of the i-th computing component after task migration, T i is the real-time temperature of the i-th computing component before task migration, T max is the upper temperature threshold preset for all computing components, T min It is the unified temperature threshold lower limit preset for all computing components, and δ is the dynamic adjustment coefficient of heat dissipation intensity.
[0018] In another aspect, a computer temperature anomaly processing system is provided, the system comprising the following components:
[0019] Monitoring module: used to obtain real-time data on the temperature, load and heat dissipation capacity of each computing component of the computer;
[0020] Modeling module: Builds a dynamic load balancing and temperature control joint optimization model based on the data obtained by the monitoring module;
[0021] Judgment module: compares the temperature of each computing component obtained by the monitoring module with the temperature threshold preset in the joint optimization model to determine whether the computing component has a temperature anomaly;
[0022] Task migration module: When the judgment module detects a temperature anomaly in a computing component, it migrates some tasks on the abnormal component to other suitable computing components based on the joint optimization model and records relevant information about the task migration.
[0023] Collaborative Adjustment Module: After task migration is completed, it dynamically adjusts the priority and number of unmigrated tasks based on the temperature changes of each computing component. It also coordinates the working intensity of the cooling system based on the component temperature and load.
[0024] Control module: controls the operation of the entire temperature anomaly handling system, updates system information based on new data, evaluates model parameters, and dynamically monitors and optimizes component temperature and load.
[0025] Compared with the prior art, this computer temperature anomaly processing method and system has the following beneficial effects:
[0026] 1. The present invention deeply couples dynamic load balancing and temperature control to build a joint optimization model. Based on the real-time collected temperature, load and heat dissipation capacity data of each computing component, it realizes dynamic and precise adjustment of task allocation and heat dissipation intensity. When the load fluctuates or the heat dissipation environment changes, the system can quickly perceive the changes in the status of the computing components and automatically adjust the task migration rules and heat dissipation system control parameters according to the model to ensure that the computer system can operate stably and efficiently under different working conditions, significantly improving the system's adaptability and effectively avoiding performance degradation or system failure due to temperature anomalies.
[0027] 2. The present invention rationally coordinates the working intensity of the heat dissipation system through a joint optimization model. For computing components with lower temperatures and smaller loads, the working intensity of the heat dissipation system is appropriately reduced to avoid wasting heat dissipation resources. For components with higher temperatures and larger loads, the heat dissipation intensity is accurately improved. This not only achieves a balance between system performance and energy consumption, reduces system operating costs, but also reduces the loss of heat dissipation equipment and extends the service life of the equipment. It has good economic and environmental benefits and is in line with the development concept of green energy conservation.
[0028] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0030] Figure 1 The present invention is a flowchart of a method for processing abnormal temperature of a computer;
[0031] Figure 2 The figure is a structural diagram of a temperature anomaly processing system for a computer. DETAILED DESCRIPTION
[0032] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0033] Example 1
[0034] In the data center, each server is used as a basic unit and is equipped with a high-precision digital temperature sensor to monitor the temperature of each component in real time at intervals of 100ms to obtain temperature data T i Through the built-in management module of the server, the task queue depth, process CPU occupancy and memory usage are continuously monitored, and the server load status is comprehensively evaluated to obtain the load data L j At the same time, the heat dissipation capacity detection equipment integrated in the server chassis is used to collect data such as the speed of the cooling fan and the surface temperature of the aluminum heat sink in real time, and normalized using a standardized method to obtain the heat dissipation capacity data S of each server. j .
[0035] Based on the collected temperature T of all servers i 、Load L i and heat dissipation capacity S i Based on the data, combined with the processor safe operating temperature range (0-100℃) and memory chip operating temperature requirements (0-85℃) provided in the Dell server hardware manual, and taking into account the actual operating environment and energy-saving requirements of the data center, a unified temperature threshold upper limit T is set. max =85℃, lower limit T min =30℃, such as Figure 1 As shown, on this basis, a dynamic load balancing and temperature control joint optimization model is constructed. In order to quantify the rationality of migrating tasks from server i to server j, the task migration tendency formula is used where ∈ is set to 10 -6 , used to avoid the situation where the denominator is zero. In this formula, (T max -T i ) reflects the temperature pressure of the source server i, S j Reflects the heat dissipation capacity of the target server j, L j Represents the load of target server j, P ij The larger the value, the more inclined to migrate tasks from server i to server j. At the same time, the heat dissipation intensity adjustment coefficient formula Calculate the working intensity of the cooling system. In the initial stage of system construction, set α to 0.8 and β to 0.2. The values will be adjusted according to the actual operation conditions. In this formula, The basic adjustment ratio is determined based on the position of the real-time temperature of server i within the threshold range. α and β are used to adjust the sensitivity and basic value of the heat dissipation intensity adjustment to accurately match the heat dissipation power with the temperature risk.
[0036] The management system continuously compares the real-time collected temperature of each server component with the preset temperature threshold. When the CPU temperature of a server is detected to reach 86°C, exceeding the preset threshold of 85°C, the system immediately determines that the server has a temperature anomaly. Then, it performs task migration based on the joint optimization model. First, the task migration amount formula is used. Calculate the number of tasks that need to be migrated, where C i According to the current performance and operation status of the server, it is evaluated as 0.7, which represents the current task processing capacity coefficient of the server. total is the current total number of tasks of the server, set to 20, n is the total number of servers in the data center, then, among the other servers whose load is less than 40% and whose cooling fan speed is less than 70% of the maximum speed (indicating strong cooling capacity), by calculating the P of each target server ij value, filter out P ij The three servers with the largest values are used as target servers. Non-critical tasks such as image cache updates and log analysis running on the abnormal servers are reasonably allocated and migrated to these three target servers according to task priority and resource usage. The type, start time, end time, and target server information of the task migration are recorded in detail.
[0037] After the task migration is completed, the system continuously monitors the temperature changes of all servers at intervals of 500ms. If it is found that the CPU temperature of one of the target servers rises rapidly from 50°C to 78°C within 10 minutes after receiving the task, approaching the threshold, the coefficient formula is immediately adjusted according to the task priority. Dynamically adjust the priority of the unmigrated tasks on the original abnormal server, where T i,kIt is obtained by predicting the historical task execution data and the current component status. It represents the temperature increment generated by the expected execution of the kth task in the i-th server. m is the total number of unmigrated tasks in the i-th server. γ is initially set to 1.2 and will be adjusted according to actual conditions to control the sensitivity of priority adjustment. According to the calculation results, high-priority tasks such as database queries are retained on the original server for continued execution. Low-priority tasks such as web page static resource loading are prioritized for migration, and the total amount of migrated tasks is controlled by the flow control algorithm to avoid new temperature anomalies in other servers due to excessive task migration. At the same time, the heat dissipation intensity is dynamically adjusted according to the temperature and load of each server. Coordinated adjustment of the cooling system, where R i is the heat dissipation intensity adjustment coefficient before adjustment, T′ i is the new real-time temperature of the server after task migration. δ is initially set to 0.2 to control the amplitude of heat dissipation intensity adjustment. For servers with a temperature exceeding 75°C and a load greater than 60%, R′ is increased. i , increase the cooling fan speed to 90% of the maximum speed. For servers with a temperature below 40°C and a load less than 30%, reduce R′ i , reduce the cooling fan speed to 40% of the maximum speed, and achieve reasonable allocation of cooling resources and energy consumption optimization while ensuring stable operation of the server.
[0038] The system dynamically optimizes and adjusts the model and parameters based on newly collected data. Every day between 2:00 and 4:00 a.m., when the data center business is at its lowest point, the system automatically starts a deep optimization program, using this time to retrain the joint optimization model. Combined with the actual operating data of the past 24 hours, it adjusts the task migration rules and cooling system control parameters to further improve the model's accuracy and adaptability.
[0039] Example 2
[0040] In the high-performance computing center, high-precision temperature sensors are deployed on the core chips and storage module surfaces of the computing cards to collect temperature data of each component in real time at intervals of 80ms to ensure high-frequency data updates. With the help of load monitoring tools, the complex scientific computing tasks performed by the computing nodes are monitored in all directions, and key load indicators such as the number of parallel computing threads, GPU memory usage, and storage I / O throughput of the tasks are obtained in real time. Through load assessment technology, when the GPU utilization rate exceeds 90% for 15 consecutive minutes, the node is determined to be in a high-load state. At the same time, through monitoring equipment integrated in the liquid cooling system of the computing node, data such as coolant flow rate and coolant inlet and outlet temperature are collected in real time. These data are normalized using standardized methods to obtain accurate heat dissipation capacity data for each computing node.
[0041] Based on the collected computing node temperature, load and heat dissipation capacity data, combined with the safe operating temperature range (0-90℃) in the computing node hardware data and the operating temperature requirements of the storage module (0-80℃), and taking into account the strict requirements of scientific research tasks for computing stability and accuracy, the temperature threshold suitable for computing nodes is set to 82℃ and 25℃ respectively. On this basis, a dynamic load balancing and temperature control joint optimization model is constructed, and the model is trained using the operating data of past scientific computing tasks to determine the task migration rules and heat dissipation control strategies under different temperature and load scenarios. For example, when the computing node temperature is between 75-82℃ and the GPU load exceeds 85%, the model will quickly start the task migration mechanism, and based on factors such as the heat dissipation capacity of each node and the remaining computing resources, accurately calculate the task share that needs to be migrated and the appropriate target node, providing reliable protection for the stable operation of scientific computing tasks.
[0042] The real-time collected temperatures of each computing node component are compared with the preset temperature threshold. When the computing card temperature of a computing node is detected to reach 83°C, exceeding the preset threshold of 82°C, the node is immediately determined to have a temperature anomaly. Then, based on the joint optimization model, two computing nodes are selected as target nodes from other computing nodes with GPU loads below 30% and coolant flow rates above the average level (indicating good heat dissipation capabilities). By comprehensively evaluating the computing performance, current task queue length, and temperature carrying capacity of each target node, the system selects two computing nodes as target nodes. Then, some computing subtasks of the mid-term climate model simulation task being carried out on the abnormal node are rationally divided and migrated to these two target nodes according to the task dependencies and resource consumption. The specific content of the task migration, migration time, and target node information are recorded in detail to facilitate subsequent tracking and management of task execution.
[0043] After task migration is complete, the system continuously monitors the temperature changes of all computing nodes at 400ms intervals. If the GPU temperature of one target node rapidly rises from 45°C to 78°C within 8 minutes after receiving a task, approaching the threshold, the system immediately adjusts the priority of the unmigrated tasks on the original abnormal node based on the model. High-priority tasks, such as real-time analysis of scientific experimental data, are prioritized to ensure their running resources on the original node. Lower-priority tasks, such as experimental data preprocessing, are prioritized for migration. Traffic shaping technology is used to control the data transmission rate and total amount of migrated tasks to prevent temperature anomalies on other nodes caused by excessive task migration. Simultaneously, the cooling system is coordinated and adjusted based on the temperature and load of each node. For nodes with temperatures exceeding 70°C and GPU loads greater than 70%, the coolant flow rate is increased to 95% of the maximum flow rate. For nodes with temperatures below 35°C and GPU loads less than 20%, the coolant flow rate is reduced to 35% of the maximum flow rate. This ensures efficient computing tasks while optimizing the allocation of cooling resources and controlling energy consumption.
[0044] The system dynamically optimizes and adjusts the model and parameters based on newly collected data. Every Sunday between 1:00 and 3:00 a.m., when there are relatively few scientific computing tasks, the system automatically starts a comprehensive optimization program, using this time to conduct in-depth training on the joint optimization model. Combined with the actual operating data of the past week, it adjusts the task migration rules and cooling system control parameters to further improve the model's adaptability to complex load changes in scientific computing tasks.
[0045] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for handling abnormal temperature of a computer, characterized in that: The method comprises the following specific steps: Real-time data collection: Deploy temperature sensors at key computer locations, and utilize load monitoring modules and heat dissipation capacity detection equipment to collect temperature, load, and heat dissipation capacity data in real time. Joint optimization model construction: Based on the collected data, comprehensive consideration of various factors of computing components, with system performance and temperature balance as the goal, a joint optimization model describing task allocation and temperature control strategies is constructed, and task migration rules and heat dissipation control parameters are formulated; Temperature anomaly determination and task migration: The collected component temperature is compared with the preset threshold. After determining the component with abnormal temperature, the system migrates some tasks of the abnormal component to appropriate components based on model rules and the load and heat dissipation of other components, and records the relevant information. Coordinated adjustment of tasks and heat dissipation: Monitors component temperature changes after task migration, adjusts the priority and number of unmigrated tasks based on model strategies, and coordinates the cooling system's workload based on component temperature and load to achieve load and heat dissipation optimization. Loop monitoring and continuous optimization: Update system information and evaluate model parameters based on new data, dynamically monitoring and optimizing component temperature and load.
2. The method for processing abnormal temperature of a computer according to claim 1, characterized in that: In the joint optimization model construction step, based on the acquired temperature, load and heat dissipation capacity data of each computing component, the temperature threshold, load bearing capacity and heat dissipation capacity factors of each computing component are comprehensively considered. By analyzing the correlation between the data, the operating status and requirements of the system under different data combinations are determined, and a dynamic load balancing and temperature control joint optimization model is constructed. During the construction process, mathematical relationships are established to describe the task allocation and temperature control strategies. For different temperature and load conditions, corresponding task migration rules are formulated and the heat dissipation intensity adjustment coefficient of the heat dissipation system is determined.
3. The method for processing temperature anomalies of a computer according to claim 1, characterized in that: In the joint optimization model construction step, corresponding task migration rules are formulated according to different temperature and load conditions, and the heat dissipation intensity adjustment coefficient of the heat dissipation system is determined. The task migration rule determination formula is: Among them, P ij It represents the tendency to migrate tasks from the i-th temperature abnormal component to the j-th target component. The larger the value, the more inclined to migrate tasks from the i-th component to the j-th component. max It is the upper limit of the unified temperature threshold preset for all computing components. It is a fixed value set according to the hardware characteristics of the components and the requirements for safe operation. i is the real-time temperature of the i-th computing component, S j is the heat dissipation capacity of the jth target computing component, L j is the real-time load of the jth target computing component, ∈ is a very small positive number used to avoid the denominator being zero, and is a fixed constant set artificially. The heat dissipation intensity adjustment coefficient of the heat dissipation system is: Among them, R i is the heat dissipation intensity adjustment coefficient of the i-th computing component, which is used to control the working intensity of the heat dissipation system. i is the real-time temperature of the i-th computing component, T max is the upper temperature threshold preset for all computing components, T min It is the preset unified lower temperature threshold for all computing components. It is a fixed value set according to the minimum temperature requirement for normal operation of the components. α and β are adjustment coefficients, which are fixed constants determined through experiments and system debugging. They are used to adjust the sensitivity and base value of the heat dissipation intensity regulation.
4. The method for processing abnormal temperature of a computer according to claim 1, wherein: In the temperature anomaly determination and task migration step, part of the tasks of the abnormal component are migrated to the appropriate component and the relevant information is recorded. Specifically, based on the established task migration rules and combined with the load conditions and heat dissipation capacity data of other computing components in the system, the target components with relatively low load and strong heat dissipation capacity are screened out, and the number of migration tasks is calculated through the task migration amount formula. The tasks are migrated to the selected target components for execution according to the task segmentation and transmission mechanism, and the task information is recorded.
5. The method for processing temperature anomalies of a computer according to claim 3, characterized in that: In the temperature anomaly determination and task migration step, the number of migration tasks is calculated using the task migration amount formula, which is: Among them, M i is the number of tasks that need to be migrated for the i-th temperature abnormal component, T i is the real-time temperature of the i-th computing component, T threshold,i is the preset temperature threshold of the i-th computing component, C i is the current task processing capability coefficient of the i-th computing component, which is evaluated based on the component performance and current operating status. max It is the upper limit of the temperature threshold preset by all computing components, a fixed value, T j is the real-time temperature of the jth computing component, S j is the heat dissipation capacity of the jth computing component, N total is the total number of tasks currently performed by the i-th computing component.
6. The method for processing abnormal temperature of a computer according to claim 1, characterized in that: In the task and heat dissipation coordinated adjustment step, the temperature changes of each computing component are continuously monitored in real time. If it is found that the temperature of some components rises too fast due to task migration, or the temperature is close to its preset temperature threshold, the priority of the unmigrated tasks is dynamically adjusted according to the joint optimization model, and tasks with lower priorities are preferentially selected for subsequent migration, and the total amount of migrated tasks is controlled. At the same time, according to the current temperature and load conditions of each computing component, the working intensity of the heat dissipation system is adjusted. For components with higher temperatures and heavier loads, the corresponding cooling fan speed is increased, or the cooling power of the heat dissipation equipment is increased. For components with lower temperatures and lighter loads, the working intensity of the heat dissipation system is reduced to reduce energy consumption.
7. The method for processing abnormal temperature of a computer according to claim 3, characterized in that: In the task and heat dissipation coordinated adjustment step, the priority of the non-migrated task is dynamically adjusted according to the joint optimization model. The priority calculation formula is: Among them, Q k,i is the priority adjustment coefficient of the kth unmigrated task in the i-th computing component, which is used to adjust the task priority. The larger the value, the higher the priority. max is the upper temperature threshold preset for all computing components, T i is the real-time temperature of the i-th computing component, T i,k is the temperature increment expected to be generated by the execution of the kth task in the i-th computing component, m is the total number of unmigrated tasks in the i-th computing component, and γ is the priority adjustment sensitivity coefficient.
8. The method for processing abnormal temperature of a computer according to claim 3, characterized in that: In the task and heat dissipation coordinated adjustment step, the working intensity of the heat dissipation system is adjusted, and the adjustment formula is: Among them, R i ′ is the heat dissipation intensity adjustment coefficient of the i-th computing component after adjustment, R i is the heat dissipation intensity adjustment coefficient of the i-th computing component before adjustment, T i ′ is the real-time temperature of the i-th computing component after task migration, T i is the real-time temperature of the i-th computing component before task migration, T max is the upper temperature threshold preset for all computing components, T min It is the unified temperature threshold lower limit preset for all computing components, and δ is the dynamic adjustment coefficient of heat dissipation intensity.
9. A computer temperature anomaly processing system, the system being used in a computer temperature anomaly processing method according to any one of claims 1 to 8, characterized in that: The system comprises the following components: Monitoring module: used to obtain real-time data on the temperature, load and heat dissipation capacity of each computing component of the computer; Modeling module: Builds a dynamic load balancing and temperature control joint optimization model based on the data obtained by the monitoring module; Judgment module: compares the temperature of each computing component obtained by the monitoring module with the temperature threshold preset in the joint optimization model to determine whether the computing component has a temperature anomaly; Task migration module: When the judgment module detects a temperature anomaly in a computing component, it migrates some tasks on the abnormal component to other suitable computing components based on the joint optimization model and records relevant information about the task migration. Collaborative Adjustment Module: After task migration is completed, it dynamically adjusts the priority and number of unmigrated tasks based on the temperature changes of each computing component. It also coordinates the working intensity of the cooling system based on the component temperature and load. Control module: Controls the operation of the entire temperature anomaly handling system, updates system information based on new data, evaluates model parameters, and dynamically monitors and optimizes component temperature and load.