Chip aging test equipment temperature control method and system
Patent Information
- Application Number
- CN202611334555.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于提供一种芯片老化测试设备温度控制方法及系统,旨在解决现有技术中由于横向热传导导致高负载工位隐蔽过热与相邻低负载工位被动升温共存,进而破坏老化温度场均匀性的技术问题,能够基于热传导拓扑网络对横向传热进行量化补偿并实现动态分区控制,有效消除隐蔽过热与被动升温,保障老化温度场的一致性
[0008]本发明提供的芯片老化测试设备温度控制方法,具有能够有效识别隐蔽过热与被动升温的真实状态,消除横向热传导对温度感知信号的干扰,从而提升芯片老化测试过程中温度场均匀性与老化应力一致性的优点。
Smart Images

Figure CN122837534A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated circuit aging test technology, and more specifically, to a temperature control method and system for chip aging test equipment. Background Technology
[0002] Chip aging tests typically do not involve a constant load throughout, but rather include alternating high-stress and low-stress phases. In existing technologies, cooling systems are usually configured with a rated cooling capacity for each fixed zone. Within a zone, each station shares the cooling resources of that loop, and the zones operate largely independently, maintaining a stable average temperature within each zone by adjusting valve openings or medium temperature. During the execution of the same aging batch, the test task scheduling system often automatically switches the chip's operating mode according to the aging stage, which can lead to significant variations in power consumption at several stations.
[0003] When dynamic scheduling causes multiple workstations within a partition to simultaneously switch to high-power mode, while most workstations in adjacent partitions switch to low-power mode, the load center of gravity shifts spatially across partitions. If these high-power workstations happen to be located near the edge of adjacent partitions, the enormous heat they generate not only needs to be carried away by the local cooling loop but will also be laterally conducted to adjacent partitions under low load conditions through the aging board substrate, cooling structure, and interface materials. Since the adjacent partitions have extremely low thermal loads at this time, their cooling loops are operating at low flow rates or even partially shut down, and the laterally conducted heat causes the temperature of the low-power workstations in the adjacent partitions to passively rise. At the same time, this lateral heat conduction effect dilutes the temperature sensing signal at the edge of the high-power workstations, causing the temperature reading received by the local controller to be lower than expected, thus mistakenly believing that no additional cooling is needed, resulting in the actual junction temperature of the high-power chip being too high.
[0004] Existing fixed-zone independent control mechanisms fail to account for the interference of lateral heat conduction on temperature sensing signals. When faced with load center shifts across zones, they cannot effectively identify the true state of hidden overheating and passive temperature rise. Therefore, in chip aging test equipment, when dynamic task scheduling causes high-load stations to concentrate in a certain cooling zone near the edge of an adjacent zone, and the adjacent zone is in a low-load state, the independent cooling control of fixed zones and the dilution of temperature sensing signals in the high-temperature zone by lateral heat conduction lead to the coexistence of hidden overheating in the high-load station and passive temperature rise in the adjacent low-load station, disrupting the uniformity of the aging temperature field and the consistency of aging stress.
[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0006] The purpose of this invention is to provide a temperature control method and system for chip aging test equipment, which aims to solve the technical problem in the prior art where hidden overheating at high-load stations and passive heating at adjacent low-load stations coexist due to lateral heat conduction, thereby disrupting the uniformity of the aging temperature field. The invention can quantitatively compensate for lateral heat transfer based on the heat conduction topology network and realize dynamic zoning control, effectively eliminating hidden overheating and passive heating, and ensuring the consistency of the aging temperature field.
[0007] In a first aspect, the present invention provides a temperature control method for a chip aging test device, comprising the following steps: S1. Obtain a pre-constructed thermal conduction topology network for the chip aging test equipment; the thermal conduction topology network is constructed based on the target information of each test station in the chip aging test equipment. S2. Based on the heat conduction topology network, identify the test station that belongs to the real heat source and use it as the real heat source station, and identify the test station that belongs to the passive heat source and use it as the passive heat source station. Then, by compensating the measured temperature of the real heat source station, the reconstructed real temperature of the real heat source station is obtained. S3. Take the cooling zone to which the actual heat source station belongs as the target zone; based on the heat conduction topology network, determine whether the target zone is in a cross-zone boundary mismatch state according to the reconstructed real temperature; S4. In the cross-zone boundary mismatch state, the cooling resources of the target zone are forcibly allocated so that more cooling resources are allocated to the edge area, and the cooling circuits near the boundary side in the cooling zones physically adjacent to the target zone are controlled to increase the cooling output to form a thermal barrier wall. S5. Based on the difference between the reconstructed true temperature and the preset physical damage safety threshold, assess the remaining heat carrying capacity of the edge region, and according to the remaining heat carrying capacity, perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region. S6. After performing the control in the preceding steps, perform temperature field uniformity verification on the real heat source station and the passive heat-receiving station, and correct the heat conduction topology network based on the verification results.
[0008] The temperature control method for chip aging test equipment provided by this invention has the advantages of effectively identifying the true state of hidden overheating and passive heating, eliminating the interference of lateral heat conduction on temperature sensing signals, thereby improving the uniformity of temperature field and consistency of aging stress during chip aging test.
[0009] In a second aspect, the present invention provides a temperature control system for a chip aging test device, comprising: The data acquisition unit is used to acquire a pre-constructed thermal conduction topology network for the chip aging test equipment; the thermal conduction topology network is constructed based on the target information of each test station in the chip aging test equipment. The temperature compensation unit is used to identify test stations belonging to real heat sources and use them as real heat source stations based on the heat conduction topology network, and to identify test stations belonging to passive heating and use them as passive heating stations. The unit also compensates for the measured temperature of the real heat source stations to obtain the reconstructed real temperature of the real heat source stations. The status judgment unit is used to take the cooling zone to which the real heat source station belongs as the target zone; based on the heat conduction topology network and the reconstructed real temperature, it determines whether the target zone is in a cross-zone boundary mismatch state. The first control unit is used to forcibly allocate cooling resources to the target partition under the cross-zone boundary mismatch state, so that more cooling resources are allocated to the edge area, and control the cooling circuits near the boundary side in the cooling partitions physically adjacent to the target partition to increase the cooling output to form a thermal barrier wall. The second control unit is used to assess the remaining heat carrying capacity of the edge region based on the difference between the reconstructed real temperature and the preset physical damage safety threshold, and to perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region according to the remaining heat carrying capacity. The verification and correction unit is used to perform temperature field uniformity verification on the real heat source station and the passive heat-receiving station after performing the control of the preceding steps, and to correct the heat conduction topology network according to the verification results.
[0010] As can be seen from the above, the temperature control method for chip aging test equipment provided by this invention, by constructing a heat conduction topology network, can accurately identify the lateral heat conduction paths and their thermal resistance characteristics between each test station, solving the problem of hidden overheating in high-power stations and passive heating in adjacent areas caused by neglecting lateral heat flow in existing technologies. By reconstructing the real temperature, the controller can capture the real thermal load state under high load in real time; through the judgment of cross-zone boundary mismatch state and forced cooling resource allocation, this application implements a thermal barrier at the physical layer, effectively preventing heat from overflowing laterally across zones; at the same time, based on the task scheduling strategy of remaining heat carrying capacity, the test layout and power distribution are optimized from the source, ultimately achieving a highly balanced temperature field and consistency in high-stress testing during the chip aging process.
[0011] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing embodiments of the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0012] Figure 1 This is a flowchart of a temperature control method for a chip aging test device provided in an embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of a temperature control system for a chip aging test device provided in an embodiment of the present invention.
[0014] Label Explanation: 100. Data acquisition unit; 200. Temperature compensation unit; 300. Status judgment unit; 400. First control unit; 500. Second control unit; 600. Verification and correction unit. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0016] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0017] In traditional chip aging test equipment, the cooling system has a rated cooling capacity for each fixed zone, and each station within a zone shares a cooling circuit. The zones operate essentially independently. When dynamic task scheduling causes the center of gravity of the load to shift across zones in space, the heat generated by the high-load station will be laterally conducted through the aging board substrate to adjacent zones under low load. This lateral heat conduction effect dilutes the temperature sensing signal at the edge of the high-power station, causing the controller to receive a lower temperature reading. This leads to the controller mistakenly believing that no additional cooling is needed, resulting in an excessively high actual junction temperature of the high-power chip. This causes hidden overheating in the high-load station to coexist with passive heating in adjacent low-load stations, disrupting the uniformity of the aging temperature field and the consistency of aging stress.
[0018] For example, in a chip aging test scenario, the test task scheduling system automatically switches the chip's operating mode according to the aging stage. This causes multiple workstations in one partition to simultaneously switch to high-power mode, while most workstations in adjacent partitions switch to low-power mode. Since these high-power workstations are located near the edges of adjacent partitions, the heat they generate is conducted laterally to the adjacent partitions through the substrate. This results in the temperature of the low-power workstations passively rising, even when the cooling circuits in the adjacent partitions are operating at lower flow rates or even partially shut down. Simultaneously, because the heat conduction effect dilutes the temperature sensing signal at the edges of the high-power workstations, the controller in this area cannot identify the true junction temperature state, leading to an imbalance in cooling resource allocation and significant deviations in the local temperature field.
[0019] If the above problems are not resolved, lateral thermal conduction interference will continue to cause distortion of the temperature sensing signal, making it impossible for the control system to accurately identify the true thermal state of the chip. This will lead to uneven temperature field distribution during the aging test, seriously affecting the consistency of chip aging stress, and may even cause physical damage to the chip due to local overheating, reducing the reliability of the test equipment and the validity of the test results.
[0020] For reference, see the appendix. Figure 1 This invention provides a temperature control method for a chip aging test device, comprising the following steps: S1. Obtain the pre-constructed heat conduction topology network for the chip aging test equipment; the heat conduction topology network is constructed based on the target information of each test station in the chip aging test equipment; the target information includes the spatial location information of each test station, the identification of its cooling zone, the task switching instruction sequence, and the real-time temperature sampling sequence. S2. Based on the heat conduction topology network, identify test stations belonging to real heat sources and designate them as real heat source stations, and identify test stations belonging to passive heating and designate them as passive heating stations. Then, by compensating for the measured temperature of the real heat source stations, the reconstructed real temperature of the real heat source stations is obtained; specifically including steps S21-S24: S21. Obtain the physical arrangement distance of each test station from the spatial location information, and determine the partition boundary relationship between the cooling partitions to which each test station belongs based on the identification of the respective cooling partition. S22. Based on the heat conduction topology network, set the cross-zone thermal impedance evaluation value for each test station according to the physical arrangement distance and partition boundary relationship; S23. Extract the temperature change rate from the real-time temperature sampling sequence, and identify the real heat source station and the passively heated station by matching the target time when the temperature change rate exceeds the preset change rate threshold with the timestamp in the task switching instruction sequence. S24. Based on the temperature rise of the passively heated station and the cross-zone thermal impedance assessment value corresponding to the passively heated station, calculate the cross-zone heat flux loss, and obtain the reconstructed true temperature by adding the cross-zone heat flux loss as a compensation term to the measured temperature of the real heat source station. S3. Take the cooling zone to which the actual heat source station belongs as the target zone; based on the heat conduction topology network, determine whether the target zone is in a cross-zone boundary mismatch state according to the reconstructed actual temperature; specifically including steps S31-S32: S31. Using the difference between the reconstructed real temperature and the target aging temperature as the weight, the spatial coordinates of each real heat source station in the target zone are calculated by weighted average to obtain the dynamic centroid of the heat load space in the target zone. Based on the Euclidean distance between the dynamic centroid of the heat load space and the geometric center of the target zone, the centroid offset is determined. S32. When the center of gravity offset exceeds the preset offset threshold and the edge high load space density of the target partition reaches the preset risk threshold, the target partition is determined to be in a cross-zone boundary mismatch state; the edge high load space density is calculated based on the number of actual heat source stations in the edge region and the area of the edge region. S4. In the case of cross-zone boundary mismatch, the cooling resources of the target zone are forcibly allocated so that more cooling resources are allocated to the edge area, and the cooling circuits near the boundary side in the cooling zone that is physically adjacent to the target zone are controlled to increase the cooling output to form a thermal barrier wall. S5. Based on the difference between the reconstructed real temperature and the preset physical damage safety threshold, assess the remaining heat carrying capacity of the edge region, and according to the remaining heat carrying capacity, perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region. S6. After performing the control steps in the preceding steps, verify the temperature field uniformity of the real heat source station and the passive heat-receiving station, and correct the heat conduction topology network based on the verification results.
[0021] For ease of understanding, the following explains some key terms in this embodiment: Thermal conduction topology network: This network is an abstract representation of the heat transfer paths within the chip aging test equipment. It treats each test station in the equipment as a node, and the heat transfer paths between stations and between stations and the cooling system as edges. Each node contains information such as its spatial location, its associated cooling zone identifier, task status, and real-time temperature. Each edge is assigned a thermal resistance assessment value to quantify the ease of heat transfer. By constructing this network, the system can comprehensively and dynamically grasp the heat flow distribution and changes within the equipment, providing a data foundation for precise temperature control.
[0022] Real heat source workstations: These refer to test workstations that actively generate a large amount of heat due to their own task switching (e.g., switching from a low-power mode to a high-power mode). Identifying real heat source workstations is one of the core aspects of this method, as it directly relates to the accurate assessment of heat load and the rational allocation of cooling resources.
[0023] Passively heated workstations: These are test workstations that, while not switching to high-power mode, experience a passive temperature increase due to heat conduction laterally from nearby real heat source workstations. Identifying passively heated workstations helps quantify the impact of lateral heat conduction and provides a basis for temperature compensation of real heat source workstations.
[0024] Inter-zone thermal resistance assessment value: This value quantifies the resistance to heat transfer between different cooling zones or between different workstations within the same zone. A higher thermal resistance assessment value indicates more difficult heat transfer; conversely, a lower value indicates easier heat transfer. Accurately assessing inter-zone thermal resistance is crucial for calculating inter-zone heat flux loss and thus compensating for the actual heat source workstation temperature.
[0025] Reconstructing the true temperature: This refers to a temperature value that is closer to the actual junction temperature of the chip, obtained by compensating for the measured temperature at the actual heat source location. Since lateral heat conduction can dilute the temperature sensing signal in the edge region, resulting in a lower measured temperature, reconstructing the true temperature is crucial for avoiding hidden overheating.
[0026] Cross-zone boundary mismatch state: This refers to an abnormal thermal state where high-power tasks are concentrated in a cooling zone near the edge of an adjacent zone, and the adjacent zone is under low load. Because the cooling control of each zone is independent and lateral heat conduction dilutes the temperature sensing signal of the high-temperature zone, an abnormal thermal state arises where hidden overheating of the high-load station coexists with passive heating of the adjacent low-load station. Identifying and handling this state is key to solving the core technical problem of this method.
[0027] Edge High-Load Space Density: This metric quantifies the concentration of high-power workstations within the edge region of a cooling zone. When the edge high-load space density reaches a certain threshold, it indicates a high thermal risk in that area, requiring special attention.
[0028] Remaining thermal capacity refers to the ability of an edge region to withstand additional thermal loads under current cooling conditions. Assessing remaining thermal capacity helps the system determine whether it can continue to perform high-power tasks in that region, thereby avoiding physical damage to the chip.
[0029] This application proposes a temperature control method for chip aging test equipment, aiming to solve the problem that in chip aging test equipment, the load center of gravity shifts across zones due to dynamic task scheduling, causing lateral heat conduction interference with temperature sensing under the fixed partition independent control mechanism, resulting in hidden overheating of high-load stations and passive heating of adjacent low-load stations, thereby destroying the uniformity of the aging temperature field and stress consistency.
[0030] To achieve the above objectives, this method first acquires a pre-constructed thermal conductivity topology network for the chip aging test equipment. This thermal conductivity topology network is constructed based on the target information of each test station in the chip aging test equipment. The target information includes the spatial location information of each test station, the identification of its corresponding cooling zone, the task switching instruction sequence, and the real-time temperature sampling sequence. For example, the system can acquire the most basic temperature status information through a distributed thermal sensor network. At the equipment hardware level, high-precision thermocouples or NTC thermistors are independently deployed on the bottom of the aging board slot corresponding to each test station or on the fixture surface that directly contacts the chip under test. The data acquisition module synchronously reads the real-time temperature readings of all stations at a set high-frequency sampling period (e.g., five to ten times per second). The purpose of high-frequency sampling is to accurately capture the tiny temperature step response caused by the instant of task switching in subsequent processing, thereby avoiding distortion of the temperature rise signal due to sampling delay. At the same time, the system accesses the data stream of the test task scheduling system in real time through the internal communication bus of the equipment to obtain the task status information currently being executed by each station. This information includes, but is not limited to: the unique identifier of the workstation, the current test program stage, the chip operating voltage and injection current set for this stage, the expected power consumption mode (divided into high-stress full-load mode, medium-load mode, and low-stress sleep mode), the duration of this power consumption mode, and the timestamp of the expected next mode switch. After acquiring the above two independent types of information—physical temperature and logical task—the system organizes and associates them into a global spatiotemporal topology matrix. This topology matrix uses all test workstations within the device as network nodes. Each node is assigned a multi-dimensional state vector, which encapsulates the workstation's three-dimensional physical coordinates, its fixed cooling zone number, the current real-time temperature reading, the historical temperature time series within the previously set time window, the current power consumption mode flag, and the precise timestamp of the most recent task switch. Between nodes, the system establishes adjacency edges representing heat conduction paths based on factors such as the physical arrangement distance of the workstations on the aging board, whether they cross cooling zone boundaries, and whether there are mechanical barriers in between. Each edge is assigned an initial cross-zone thermal impedance assessment value. The thermal impedance of adjacent workstations within the same zone is lower, while the thermal impedance of adjacent workstations across zones is differentiated based on the degree of isolation of the physical structure. In particular, the cross-zone connecting edges of edge workstations located at the junction of two fixed cooling zones are specially marked as "high-sensitivity heat conduction channels." Through the acquisition, organization, and correlation of the above information, the system aligns the originally isolated temperature sensing data with dynamic test task commands in the same spatiotemporal coordinate system, forming a dynamic topology graph containing complete contextual information.
[0031] Furthermore, based on the heat conduction topology network, the system identifies test stations belonging to real heat sources and designates them as such, as well as test stations passively heated and designates them as such. By compensating for the measured temperature of the real heat source stations, the reconstructed true temperature of each station is obtained. Specifically, the physical arrangement distance of each test station is obtained from spatial location information, and the boundary relationships between the cooling zones to which each test station belongs are determined according to the cooling zone identifier. For example, the precise three-dimensional coordinates of each station can be obtained through manual measurement or by importing CAD models, and the system automatically identifies which stations belong to the same zone and which stations are located at the zone boundaries according to preset cooling zone division rules. Based on the heat conduction topology network, the cross-zone thermal impedance evaluation value corresponding to each test station is set according to the physical arrangement distance and the zone boundary relationship. For example, the initial thermal impedance evaluation value can be set based on empirical values or by simulating the heat conduction process using finite element analysis software. The system extracts the temperature change rate from the real-time temperature sampling sequence and identifies real heat source workstations and passively heated workstations by matching the target moment when the temperature change rate exceeds a preset threshold with the timestamp in the task switching instruction sequence. For example, the system first traverses each node in the network, calculates the derivative of its historical temperature time series, and extracts the temperature change rate. When the temperature change rate of a node exceeds a preset fluctuation shielding threshold, the node is marked as an "active temperature rise node" and enters the attribution determination logic. The system first performs time correlation verification to extract the task switching timestamp of the active temperature rise node and checks whether it has just received a scheduling instruction to switch from low power to high power within the set thermal response time window. If the starting point of the temperature rise closely matches the timestamp of the high power task, and the slope of the temperature change rate matches the estimated heat jump amplitude of the high power mode, then it is determined that the temperature rise is caused by the heat generated by the workstation itself, and the node is identified as a "real heat source workstation". Conversely, if a node has not recently undergone a task switch with increased power consumption, or is currently in a low-stress sleep state, but its temperature still shows an abnormally significant rise, the system searches for its physical neighbors along the topology network. If a node that has been identified as a "real heat source" is found among its neighbors (especially those crossing partition boundaries), and the temperature rise curve of this low-power node lags significantly behind that of the real heat source on the time axis, the system determines that the node has received heat through lateral conduction and identifies it as a "passively heated node."
[0032] Subsequently, based on the temperature rise of the passively heated station and the corresponding cross-zone thermal impedance assessment value, the cross-zone heat flux loss is calculated. This cross-zone heat flux loss is then added as a compensation term to the measured temperature of the actual heat source station to obtain the reconstructed true temperature. The specific calculation logic is as follows: cross-zone heat flux loss equals the reciprocal of the cross-zone thermal impedance assessment value multiplied by the temperature rise of the passively heated station. Then, all cross-zone heat flux loss to adjacent zones is accumulated (i.e., the cross-zone heat flux loss of all passively heated stations physically adjacent to the actual heat source station is accumulated to obtain the total cross-zone heat flux loss), and this is added as a temperature compensation term (i.e., the total cross-zone heat flux loss) to the current sensor reading of the actual heat source station to calculate the reconstructed true temperature. After object differentiation, the system performs temperature rise signal dilution and recovery processing on the actual heat source stations located at the edge of the partition. The system extracts the temperature rise of all adjacent passively heated workstations and, combined with the cross-regional thermal impedance assessment values recorded in the topology network, calculates the heat flux lost through lateral conduction across regions. Specifically, the equivalent value of the lost heat flux equals the reciprocal of the cross-regional thermal impedance multiplied by the temperature rise of the passively heated workstation. Subsequently, the system accumulates the equivalent values of all heat flux lost to neighboring regions and adds them as a temperature compensation term to the current sensor reading of the actual heat source workstation, thereby calculating a crucial intermediate result: reconstructing the true temperature.
[0033] Next, the cooling zone to which the actual heat source station belongs is designated as the target zone. Based on the heat conduction topology network, the system determines whether the target zone is in a cross-zone boundary mismatch state based on the reconstructed actual temperature. Specifically, the spatial coordinates of each actual heat source station within the target zone are weighted and averaged using the difference between the reconstructed actual temperature and the target aging temperature as the weight. This yields the dynamic centroid of the heat load within the target zone, and the centroid offset is determined based on the Euclidean distance between the dynamic centroid of the heat load and the geometric center of the target zone. For example, the system first extracts the physical coordinates of all actual heat source stations within a fixed cooling zone, and uses the difference between the reconstructed actual temperature and the target aging set temperature of each station as the weight to calculate the "dynamic centroid of the heat load" within that zone. The coordinates of this centroid reflect the core location within the current zone that truly requires a large amount of cooling resources. Subsequently, the system extracts the geometric center coordinates of the fixed cooling zone in its physical structure, calculates the Euclidean distance between the dynamic centroid of the heat load and the geometric center of the zone, and obtains the "centroid offset." Simultaneously, the system counts the number of actual heat source stations located in the edge region of the partition (i.e., the geometric boundary zone near adjacent partitions) and calculates the "edge high-load space density" based on the area. The system compares the center-of-gravity offset with a preset effective coverage radius threshold and the edge high-load space density with a preset aggregation risk threshold. This leads to different processing branches. If the center-of-gravity offset is less than the effective coverage radius threshold and the edge high-load space density is less than the aggregation risk threshold, it indicates that the current high-load tasks are mainly concentrated in the central region of the partition. At this time, the center-facing cooling flow field of the fixed partition can effectively cover the heat source, and the risk of lateral heat conduction is extremely low. The system determines that there is no boundary mismatch and outputs an instruction to maintain the existing independent control mode of the partition, continuing to rely on the average feedback within the partition for routine cooling resource adjustment. When the center-of-gravity offset exceeds the preset offset threshold and the edge high-load space density of the target partition reaches the preset risk threshold, the target partition is determined to be in a cross-region boundary mismatch state. If the center-of-gravity offset is greater than or equal to the effective coverage radius value and the edge high-load space density exceeds the aggregation risk threshold, it indicates that a large number of high-power stations are densely clustered at the edge of the cooling partition due to dynamic task scheduling. Because the edge region is at the end of the cooling flow field, conventional center-enhanced cooling cannot effectively reach this area, and the system determines that it has entered a "cross-zone boundary mismatch" problem state. To address this state, the system further examines the state vectors of adjacent zones. If most stations in adjacent zones are marked as low-power mode or passively heated stations, it fully confirms the extreme instability conditions in the recommended technology scenario. At this point, the system generates a high-priority "cross-zone boundary mismatch flag" and a diagnostic report containing the specific physical coordinates of the mismatch boundary, a list of high-load stations at the edge of this zone, a list of passively heated stations in neighboring zones, and the thermal load centroid offset vector.
[0034] Under the cross-region boundary mismatch condition, forced distribution is performed on the cooling resources of the target partition, so that more cooling resources are allocated to edge areas, and the cooling loop near the boundary side in the cooling partition physically adjacent to the target partition (that is, the cooling partition to which the passive heated station belongs) is controlled to increase cooling output, so as to form a thermal blocking wall. For example, the system first calculates the required "total compensation cooling equivalent" for cooling these edge stations to the safe range according to the deviation between the reconstructed real temperature of the real heat source stations at the local edge and the target aging temperature, combined with the heat capacity parameter of the chip. Subsequently, instead of having the local area bear all of this total compensation, the system splits the total compensation cooling equivalent into the share borne by the local area and the share borne by the adjacent area according to dynamic weights based on the cross-region thermal impedance and the spatial radiation model of the cooling flow field. For the share borne by the adjacent area, the system issues a "blocking and absorbing" control command to the adjacent area cooling controller. Specifically, the system instructs the adjacent area to appropriately turn on or increase the output of the local cooling loop on the side near the mismatched boundary (for example, increase the opening degree of the air valve on this side or increase the flow rate of the cooling liquid branch). The primary purpose of this extra cooling provided by the adjacent area is not to cool the low-load stations of the adjacent area itself, but to form a "thermal blocking wall" at the physical boundary. It can directly absorb the huge amount of heat transversely conducted from the edge of the local area, thereby quickly pressing the temperature of the passive heated station back to the reference line. Meanwhile, through the conduction effect of the cooling substrate, it helps to take away the heat from the edge of the local area from the side. To prevent overcooling of low-load stations in the adjacent area due to extra cooling, the system strictly limits the upper limit of the compensation output of the adjacent area, so that it just offsets the calculated transverse conduction heat flux. For the share borne by the local area, the system issues a "directional tilting" control command to the local cooling controller. Under the premise of keeping the total cooling amount of the local area from drastic sudden changes (to protect the low-load stations in the center of the local area from overcooling), the system forces the existing cooling resources of the local area to tilt and gather towards the mismatched boundary by adjusting the angle of the guide vanes inside the device or changing the distribution ratio of the multi-way proportional valves. Through this cross-region linkage control, the distribution of cooling resources is no longer limited to fixed geometric partitions, but is dynamically reshaped closely following the spatial center of gravity of the real heat load. This treatment process directly blocks the malignant conduction in the contradiction generation mechanism that "the control output cannot effectively cover the real high heat source and cannot maintain the stability of low-load stations". Finally, the system outputs the new parameter setting values of the cooling actuators of the local area and the adjacent area, and issues them to the underlying hardware drive module for execution. At the same time, the system transmits the current cooling resource distribution saturation state as an intermediate result to the next step, so as to introduce higher-dimensional task scheduling intervention when the cooling capacity reaches the physical limit.
[0035] Based on the difference between the reconstructed true temperature and the preset physical damage safety threshold, the system assesses the remaining heat carrying capacity of the edge region. Based on this remaining heat carrying capacity, it performs delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region. For example, the system first assesses the "remaining heat carrying capacity" of the mismatch boundary region based on the cooling resource allocation saturation status output in the previous step and the current reconstructed true temperature of each edge workstation. The system calculates the estimated time required for the reconstructed true temperature to reach the upper limit of the chip's physical damage safety threshold if high-power tasks are added to this region under the current maximum linkage cooling compensation. If this estimated time is extremely short, it indicates that the heat carrying capacity of the edge region is close to absolute saturation and is in an extremely high thermal risk state. Subsequently, the system intercepts and parses the next-stage task queue pre-scheduled by the task scheduling system. The system checks the workstation IDs in the scheduling queue that are planned to switch from low-power to high-power in the near future. If it finds that the scheduling system is preparing to continue issuing high-power tasks to workstations that happen to be located in high-thermal-risk edge regions, the system will immediately trigger a task carrying capacity correction judgment. Entering the high-risk interception branch: The system forcibly sends a "task delay interception instruction" to the task scheduling system. This instruction requires the scheduling system to freeze the high-power switching actions of these edge workstations, forcing them to maintain a low-power sleep state for the next few test cycles until the thermal load center of gravity in this area shifts away from the edge due to the completion of tasks at other workstations, or the reconstructed real temperature steadily falls below the safe baseline under the effect of cooling compensation. If the overall progress of the test batch does not allow for a long delay, the system enters the spatial reallocation branch: outputting a "task spatial reallocation suggestion" to the task scheduling system. Based on the current global thermal risk map, the system selects low-load workstations located in the center of the current or neighboring areas, with low current reconstructed real temperatures and abundant cooling resources, and suggests that the scheduling system move the high-power tasks originally scheduled to be executed at the edge to these central workstations. The strong direct-facing cooling capacity of the central area can easily absorb these high loads and is less likely to cause cross-area heat loss. As the final output, the system writes the updated "maximum power consumption limit" that each test workstation is allowed to carry in the current time period into the interactive database shared with the scheduling system in real time. When generating the next test sequence, the task scheduling system must read and verify this limit as a hard constraint.
[0036] Finally, after executing the control measures of the preceding steps, the temperature field uniformity of the actual heat source stations and passively heated stations is verified, and the heat conduction topology is corrected based on the verification results. For example, within the set verification time window after the linkage control command and task correction command take effect, the system continuously collects the latest temperature data of all relevant stations on both sides of the mismatch boundary at high frequency and recalculates their reconstructed true temperature. The system first performs a hidden overheating elimination verification, focusing on monitoring the reconstructed true temperature trajectory of the actual heat source stations at the edge of the local area. If it is found that the temperature does not show a decreasing trend as expected, or even remains close to the upper limit of the safety threshold, it indicates that the cross-regional lateral heat flux calculated in the preceding steps may be too small, resulting in an underestimation of the actual heat load, or the directional cooling compensation allocated to this area is still insufficient. At this time, the system enters the adaptive parameter correction branch, automatically increasing the heat conduction coefficient of the corresponding cross-regional connection edge in the spatiotemporal topology network according to a preset fixed step size. This correction will directly lead to a larger equivalent value of the recalculated heat flux loss, thereby increasing the reconstructed true temperature. This forces the system to further increase the weighting of cooling resources in this area in the next control cycle until the temperature of the high-load edge workstations is completely suppressed to the target aging range. Subsequently, the system performs a passive temperature rise blocking verification, focusing on monitoring the temperature trajectory of the passively heated workstations in the adjacent area. If it is found that although the temperature of these low-load workstations has dropped, the drop is too large, or even falls below the lower limit of the aging stress requirement, it indicates that the compensation cooling capacity activated by the adjacent area to build a thermal barrier is too large, resulting in over-cooling. At this time, the system enters the reverse correction branch, proportionally reducing the share of compensation cooling equivalent borne by the adjacent area, and fine-tuning the opening of the adjacent area's air valve or flow valve, so that the adjacent area can maintain the basic temperature environment required by the low-power workstations while blocking lateral heat conduction. During the verification process, if the system detects an extreme abnormal state, such as a large-scale disconnection of the temperature sensor causing input loss, or a state conflict caused by the task scheduling system forcibly rejecting the task interception suggestion due to a special highest priority instruction, the system will immediately abandon the fine-grained linkage control and trigger the device-level hardware degradation protection mechanism. The degradation protection mechanism unconditionally cuts off or reduces the operating voltage and injection current of all high-load stations within the entire testing equipment, sacrificing the acceleration efficiency of the current batch of aging tests to absolutely ensure that the tested chips do not suffer mass thermal breakdown and burnout. Under normal stable operation, the system records the final topology parameters (such as corrected thermal impedance and optimal cooling allocation weights) and corresponding task load distribution patterns of each successfully verified and temperature field balanced test in the system operation log as high-value historical experience data. This data is not only used for continuous optimization of the current batch, but will also be directly called upon in the future when replacing a new batch of aging boards of different specifications as the benchmark model for initial cooling resource pre-allocation and topology network initialization, thus forming a continuously self-evolving, increasingly accurate intelligent temperature control closed loop.
[0037] The following example will provide a more detailed explanation of the above technical solution: Suppose a chip aging test apparatus contains two adjacent cooling zones, A and B. Zone A has multiple test stations on its edge, which are scheduled by a task scheduling system to perform high-power tasks, while the stations in Zone B are in a low-power sleep state. Traditional fixed-zone independent control mechanisms fail to consider the interference of lateral heat conduction on temperature sensing signals. When the load center shifts across zones, they cannot effectively identify the true state of hidden overheating and passive heating. Therefore, when dynamic task scheduling causes high-load stations to concentrate in the edge region of Zone A near Zone B, and Zone B is in a low-load state, the independent cooling control of the fixed zones and the dilution of the temperature sensing signal in the high-temperature area of Zone A by lateral heat conduction lead to the coexistence of hidden overheating in the high-load stations of Zone A and passive heating in the adjacent low-load stations of Zone B, disrupting the uniformity of the aging temperature field and the consistency of aging stress.
[0038] To address this issue, this application proposes a temperature control method for chip aging test equipment. First, the system acquires a pre-constructed thermal conduction topology network for the chip aging test equipment. This network includes the spatial location information of all test stations in partitions A and B, their respective cooling partition identifiers, task switching command sequences, and real-time temperature sampling sequences. For example, the system acquires basic temperature status information through a distributed thermal sensor network. At the equipment hardware level, high-precision thermocouples or NTC thermistors are independently deployed on the bottom of the aging board slot corresponding to each test station or on the surface of the fixture that directly contacts the chip under test. The data acquisition module synchronously reads the real-time temperature readings of all stations at a set high-frequency sampling period. The purpose of high-frequency sampling is to accurately capture the minute temperature step response caused by task switching in subsequent processing, thereby avoiding temperature rise signal distortion due to sampling delay. Simultaneously, the system accesses the data stream of the test task scheduling system in real time through the internal communication bus of the equipment to obtain the task status information currently being executed by each station. This information includes, but is not limited to: the unique identifier of the workstation, the current test program stage, the chip operating voltage and injection current set for this stage, the expected power consumption mode, the duration of this power consumption mode, and the timestamp of the expected next mode switch. After acquiring the above two independent types of information—physical temperature and logical task—the system organizes and associates them into a global spatiotemporal topology matrix. This topology matrix uses all test workstations within the device as network nodes. Each node is assigned a multi-dimensional state vector, which encapsulates the workstation's three-dimensional physical coordinates, its fixed cooling zone number, the current real-time temperature reading, the historical temperature time series within the previously set time window, the current power consumption mode flag, and the precise timestamp of the most recent task switch. Between nodes, the system establishes adjacent edges representing heat conduction paths based on factors such as the physical arrangement distance of the workstations on the aging board, whether they cross cooling zone boundaries, and whether there are mechanical barriers in between. Each edge is assigned an initial cross-zone thermal impedance assessment value. The thermal impedance of adjacent workstations within the same zone is smaller, while the thermal impedance of adjacent workstations across zones is differentiated based on the degree of physical structural isolation. In particular, the edge workstations located at the junction of two fixed cooling zones have their cross-zone connection edges specially marked as "high-sensitivity heat conduction channels." Through the acquisition, organization, and correlation of the above information, the system aligns the originally isolated temperature sensing data with dynamic test task instructions in the same spatiotemporal coordinate system, forming a dynamic topology graph containing complete contextual information.
[0039] Next, based on the heat conduction topology network, the system identifies high-power workstations at the edge of partition A as actual heat source workstations and low-power workstations in partition B near partition A as passively heated workstations. The system reconstructs the true temperature by compensating for the measured temperature of the actual heat source workstations. Specifically, the system obtains the physical arrangement distance of each test workstation from spatial location information and determines the partition boundary relationship between the cooling partitions to which each test workstation belongs based on the partition identifier. Based on the heat conduction topology network, the system sets the cross-zone thermal impedance evaluation value corresponding to each test workstation according to the physical arrangement distance and partition boundary relationship. The system extracts the temperature change rate from the real-time temperature sampling sequence and identifies actual heat source workstations and passively heated workstations by matching the target moment when the temperature change rate exceeds a preset change rate threshold with the timestamp in the task switching instruction sequence. For example, the system first traverses each node in the network, calculates the derivative of its historical temperature time series, and extracts the temperature change rate. When the temperature change rate of a node exceeds a preset fluctuation shielding threshold, the node is marked as an "active temperature rise node" and enters the attribution determination logic. The system first performs a time correlation check to extract the task switching timestamp of the node with active temperature rise, and checks whether it has just received a scheduling instruction to switch from low power to high power within the set thermal response time window. If the starting point of the temperature rise closely matches the timestamp of the high-power task, and the slope of the temperature change rate matches the estimated heat jump amplitude of the high-power mode, then the temperature rise is determined to be caused by the heat generated by the station itself, and the node is identified as a "real heat source station". Conversely, if the node has not recently experienced a task switching with increased power, or is even currently in a low-stress sleep state, but its temperature still shows an abnormally significant rise, the system searches for its physical neighbors along the topology network. If a node already identified as a "real heat source station" is found among its neighbors (especially neighbors that cross partition boundaries), and the temperature rise curve of the low-power node lags significantly behind the temperature rise curve of the real heat source station on the time axis, the system determines that the node has received heat through lateral conduction, and identifies it as a "passively heated station". Subsequently, based on the temperature rise of the passively heated station and the corresponding cross-zone thermal impedance assessment value, the cross-zone heat flux loss is calculated. This cross-zone heat flux loss is then added as a compensation term to the measured temperature of the actual heat source station to obtain the reconstructed true temperature. The specific calculation logic is as follows: the cross-zone heat flux loss equals the reciprocal of the cross-zone thermal impedance assessment value multiplied by the temperature rise of the passively heated station. Then, all cross-zone heat flux loss to adjacent zones is accumulated (i.e., the cross-zone heat flux loss of all passively heated stations physically adjacent to the actual heat source station is accumulated to obtain the total cross-zone heat flux loss), and this is added as a temperature compensation term (i.e., the total cross-zone heat flux loss) to the current sensor reading of the actual heat source station to calculate the reconstructed true temperature.After classifying the objects, the system performs temperature rise signal dilution and recovery processing on the actual heat source stations located at the edge of the partition. The system extracts the temperature rise amplitude of all adjacent passively heated stations and, combined with the cross-regional thermal impedance assessment value recorded in the topology network, calculates the heat flux lost through lateral conduction across the region. Specifically, the calculation logic is that the equivalent value of the lost heat flux equals the reciprocal of the cross-regional thermal impedance multiplied by the temperature rise amplitude of the passively heated station. Subsequently, the system accumulates the equivalent values of all heat flux lost to neighboring regions and adds them as a temperature compensation term to the current sensor reading of the actual heat source station, thereby calculating a key intermediate result: the reconstructed true temperature.
[0040] Subsequently, the system selects partition A as the target partition. Based on the heat conduction topology and the reconstructed true temperature, the system determines whether partition A is in a cross-boundary mismatch state. Specifically, the system uses the difference between the reconstructed true temperature and the target aging temperature as the weight to perform a weighted average calculation on the spatial coordinates of each real heat source station within partition A, calculating the dynamic centroid of the heat load space within partition A. Based on the Euclidean distance between the dynamic centroid of the heat load space and the geometric center of partition A, the centroid offset is determined. Simultaneously, the system counts the number of real heat source stations located within the edge region of partition A and calculates the edge high-load space density based on the region area. When the centroid offset exceeds a preset offset threshold and the edge high-load space density reaches a preset risk threshold, the system determines that partition A is in a cross-boundary mismatch state.
[0041] After determining that partition A is in a cross-region boundary mismatch state, the system forcibly allocates cooling resources to partition A, ensuring that more cooling resources are distributed to the edge region of partition A. Simultaneously, the system controls the cooling loops near the boundary in partition B, which is physically adjacent to partition A, to increase their cooling output, thus forming a thermal barrier. For example, the system first calculates the "total compensation cooling equivalent" required to cool these edge stations to a safe range based on the deviation between the reconstructed true temperature and the target aging temperature of the actual heat source stations at the edge of the partition, combined with the chip's thermal capacity parameters. Subsequently, the system no longer assigns all of this total compensation to the partition itself, but instead, based on the spatial radiation model of cross-region thermal impedance and cooling flow field, dynamically divides the total compensation cooling equivalent into a share borne by the partition and a share borne by neighboring partitions. For the share borne by neighboring partitions, the system issues "blocking and absorption" control commands to the neighboring partition's cooling controller. Specifically, the system instructs the neighboring partition to appropriately open or increase the output of local cooling loops near the mismatch boundary. The primary purpose of this additional cooling provided by the neighboring partition is not to cool the low-load stations within the neighboring partition itself, but to form a "thermal barrier" at the physical boundary. It can directly absorb the enormous heat conducted laterally from the edge of the zone, thus quickly bringing the temperature of the passively heated workstations back to the baseline. Simultaneously, through the conduction effect of the cooling substrate, it helps to remove heat from the edge of the zone. To prevent overcooling of low-load workstations in adjacent zones due to additional cooling, the system strictly limits the upper limit of the compensation output of adjacent zones, ensuring it exactly offsets the calculated lateral heat flux. For the share of cooling load borne by the zone itself, the system issues a "directional tilt" control command to the zone's cooling controller. While maintaining the total cooling load of the zone without drastic changes, the system forces the existing cooling resources in the zone to tilt and concentrate towards the mismatch boundary by adjusting the angle of the guide vanes inside the equipment or changing the distribution ratio of the multi-way proportional valve. Through this cross-zone linkage control, the distribution of cooling resources is no longer limited to fixed geometric partitions but dynamically reshaped according to the spatial center of gravity of the actual heat load. This process directly blocks the vicious cycle in the contradiction mechanism where "control output cannot effectively cover the actual high heat source and cannot maintain the stability of low-load workstations." Finally, the system outputs new parameter settings for the cooling actuators in the local and neighboring zones and sends them to the underlying hardware driver module for execution. Simultaneously, the system uses the current cooling resource allocation saturation status as an intermediate result and passes it to the next step, allowing for higher-dimensional task scheduling intervention when cooling capacity reaches its physical limit.
[0042] Based on this, the system assesses the remaining thermal carrying capacity of the edge region of partition A by comparing the reconstructed actual temperature with the preset physical damage safety threshold. According to the remaining thermal carrying capacity, the system performs delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region of partition A. For example, the system first assesses the "remaining thermal carrying capacity" of the mismatch boundary region based on the cooling resource allocation saturation status output in the previous step and the current reconstructed actual temperature of each edge workstation. The system calculates the estimated time required for the reconstructed actual temperature to reach the upper limit of the chip's physical damage safety threshold if high-power tasks continue to be added to this region under the current maximum coordinated cooling compensation. If this estimated time is extremely short, it indicates that the thermal carrying capacity of the edge region is close to absolute saturation and is in a state of extremely high thermal risk. Subsequently, the system intercepts and parses the next-stage task queue pre-scheduled by the task scheduling system. The system checks the workstation IDs in the scheduling queue that are planned to switch from low-power to high-power in the near future. If it finds that the scheduling system is preparing to continue issuing high-power tasks to workstations that happen to be located in high-thermal-risk edge regions, the system will immediately trigger a task carrying capacity correction judgment. Entering the high-risk interception branch: The system forcibly sends a "task delay interception command" to the task scheduling system. This command requires the scheduling system to freeze the high-power switching actions of these edge workstations, forcing them to maintain a low-power sleep state for the next few test cycles until the center of thermal load in this area moves out of the edge due to the completion of tasks at other workstations, or the reconstructed real temperature steadily falls back below the safe baseline under the effect of cooling compensation. If the overall progress of the test batch does not allow for a long delay, the system enters the spatial reallocation branch: outputting a "task spatial reallocation suggestion" to the task scheduling system. Based on the current global thermal risk map, the system selects low-load workstations located in the center of this area or neighboring areas, with low current reconstructed real temperatures and abundant cooling resources, and suggests that the scheduling system move the high-power tasks originally scheduled to be executed at the edge to these central workstations. The powerful direct-facing cooling capacity of the central area can easily absorb these high loads and is less likely to cause cross-area heat loss. As the final output, the system writes the updated "maximum power consumption limit" that each test workstation is allowed to carry in the current time period into the interactive database shared with the scheduling system in real time. When generating the next test sequence, the task scheduling system must read and verify this upper limit value as a hard constraint.
[0043] Finally, after executing the preceding control steps, the system performs temperature field uniformity verification on the actual heat source stations and passively heated stations, and corrects the heat conduction topology network based on the verification results. For example, within the set verification time window after the linkage control command and task correction command take effect, the system continuously collects the latest temperature data of all relevant stations on both sides of the mismatch boundary at high frequency, and recalculates their reconstructed true temperature. The system first performs a hidden overheating elimination verification, focusing on monitoring the reconstructed true temperature trajectory of the actual heat source stations at the edge of the local area. If it is found that the temperature does not show a decreasing trend as expected, or even remains close to the upper limit of the safety threshold, it indicates that the cross-regional lateral heat conduction flux calculated in the preceding steps may be too small, resulting in an underestimation of the actual heat load, or the directional cooling compensation allocated to the local area is still insufficient. At this time, the system enters the adaptive parameter correction branch, automatically increasing the heat conduction coefficient of the corresponding cross-regional connection edge in the spatiotemporal topology network according to a preset fixed step size. This correction will directly lead to a larger equivalent value of the recalculated heat flux loss, thereby increasing the reconstructed true temperature. This forces the system to further increase the weighting of cooling resources in this area during the next control cycle, until the temperature of the high-load edge workstations is completely suppressed to the target aging range. Subsequently, the system performs passive temperature rise blocking verification, focusing on monitoring the temperature trajectory of passively heated workstations in neighboring areas. If it is found that although the temperature of these low-load workstations has dropped, the drop is too large, or even falls below the lower limit of the aging stress requirement, it indicates that the compensation cooling capacity activated by the neighboring area to build a thermal barrier is too large, resulting in over-cooling. At this time, the system enters the reverse correction branch, proportionally reducing the share of compensation cooling equivalent borne by the neighboring area, and fine-tuning the opening of the neighboring area's air valve or flow valve, so that the neighboring area can maintain the basic temperature environment required by the low-power workstations while blocking lateral heat conduction. Under normal stable operation, the system records the final topology parameters and corresponding task load distribution patterns of each successful verification and temperature field equilibrium in the system operation log as high-value historical experience data. This data is not only used for continuous optimization of the current batch, but will also be directly called as the benchmark model for initial cooling resource pre-allocation and topology network initialization when a new batch of aging boards of different specifications is replaced in the future, thus forming a continuously self-evolving and increasingly accurate intelligent temperature control closed loop.
[0044] As can be seen from the above examples, the technical solution of this application, by constructing a heat conduction topology network, organically combines physical spatial location, cooling partitions, and task scheduling information, achieving accurate modeling of the complex thermal flow state during chip aging testing. Compared with the traditional fixed-partition independent control mechanism, this application constructs a heat conduction topology network by acquiring target information, providing basic data support for subsequent identification of heat sources and passively heated workstations. By determining the cross-region thermal impedance assessment value through the physical arrangement distance and partition boundary relationship, the heat conduction characteristics between different workstations can be quantified, thus providing a physical basis for heat flux calculation. By matching the temperature change rate with task switching commands, the actual heat source and passively heated workstation are identified, effectively distinguishing between active heating and passive heating, solving the problem of not being able to identify hidden overheating in traditional control. By calculating the cross-region lost heat flux and compensating for the measured temperature, the true temperature is reconstructed, eliminating the dilution interference of lateral heat conduction on the temperature sensing signal, enabling the system to obtain the true chip junction temperature. By calculating the dynamic centroid of the heat load space and judging the cross-region boundary mismatch state, real-time monitoring of load offset risk is achieved. In a mismatched state, by forcibly allocating cooling resources to the edge areas and forming thermal barrier walls, the lateral diffusion of heat to adjacent zones is effectively blocked, ensuring the uniformity of the temperature field. High-power tasks are scheduled based on remaining heat capacity, avoiding the risk of localized overheating at its source. Finally, through temperature field uniformity verification and network correction, a closed-loop adaptive control system is formed, ensuring the consistency of aging test stress and the safety of equipment operation.
[0045] In some embodiments, the specific steps in step S22 include: S221. Obtain the thermal conductivity of the target object in the chip aging test equipment; the target object includes the aging board substrate, cooling structure, and interface material; S222. Obtain the geometric dimensions between each test station; the geometric dimensions are determined based on the physical arrangement distance and the boundary relationship of the partitions. S223. Using the physical formulas of heat conduction, calculate the theoretical thermal impedance value of the heat conduction path corresponding to each test station based on thermal conductivity and geometric dimensions, and set the theoretical thermal impedance value as the cross-zone thermal impedance evaluation value corresponding to each test station.
[0046] To accurately determine the cross-zone thermal impedance evaluation value for each test station, this application proposes a calculation method based on physical parameters. First, the thermal conductivity of the target object in the chip aging test equipment needs to be obtained. Thermal conductivity is an important physical parameter characterizing the thermal conductivity of a material, and its value directly reflects the material's ability to transfer heat. In the chip aging test equipment, heat transfer between different test stations passes through various media, such as the aging board substrate, cooling structures (e.g., heat sinks, coolant channel walls), and interface materials between the chip and the aging board (e.g., thermal grease, solder). The thermal conductivity of these key materials can be obtained by consulting the technical specifications and standard material databases provided by material suppliers, or through experimental measurements, such as using professional testing methods like heat flow metering or laser scintillation. Second, the geometric dimensions between each test station need to be obtained. These geometric dimensions refer to spatial parameters such as the length and cross-sectional area of the heat transfer path between different test stations; they are key factors affecting heat conduction efficiency and resistance. The determination of geometric dimensions is based on the physical arrangement distance and partition boundary relationship of each test station. For example, information such as the center distance between each workstation, the thickness of the aging plate, and the cross-sectional area of the cooling channel can be extracted from the equipment's design drawings and CAD model data, or precise measurements can be taken using tools such as high-precision calipers and 3D scanners. Finally, using physical formulas for heat conduction, the theoretical thermal impedance value of the heat conduction path corresponding to each test workstation is calculated based on the obtained thermal conductivity and geometric dimensions, and this theoretical thermal impedance value is set as the cross-regional thermal impedance evaluation value for each test workstation. Physical formulas for heat conduction, such as Fourier's law, can quantitatively describe the laws of heat transfer. Thermal impedance is the resistance encountered by heat flow through a certain path, and its calculation usually involves the length of the heat transfer path, the heat transfer cross-sectional area, and the thermal conductivity of the material. For example, for one-dimensional steady-state heat conduction, the thermal impedance R can be expressed as R=L / (k*A), where L is the length of the heat transfer path, k is the thermal conductivity of the material, and A is the heat transfer cross-sectional area. In this way, the abstract cross-regional thermal impedance evaluation value is transformed into a theoretical calculation value based on actual physical properties, thereby improving the scientificity and accuracy of temperature compensation.
[0047] This application's solution introduces physical-level heat conduction parameters, transforming abstract cross-region thermal impedance assessment values into theoretical calculations based on actual physical properties, thereby improving the scientific rigor and accuracy of temperature compensation. Specifically, in the temperature control method of chip aging test equipment, to accurately identify the actual heat source station and the passively heated station, and to compensate for the measured temperature of the actual heat source station to obtain a reconstructed true temperature, it is necessary to set cross-region thermal impedance assessment values for each test station. This solution, by obtaining the thermal conductivity of the aging board substrate, cooling structure, and interface materials, can characterize the ease of heat transfer between different stations from the perspective of material thermal properties, ensuring that the assessment values have a solid physical basis. Simultaneously, by obtaining the geometric dimensions between each test station and combining the physical arrangement distance and partition boundary relationships, the geometric characteristics of the heat conduction path can be accurately defined, providing spatial dimension data support for subsequent thermal impedance calculations. Based on this, by correlating thermal conductivity with geometric dimensions using the physical formulas of heat conduction, the theoretical thermal impedance value of the heat conduction path between each test station can be obtained. This calculation method based on physical principles avoids errors caused by empirical estimations, ensuring that the cross-zone thermal impedance assessment value accurately reflects the internal heat transfer characteristics of the equipment. Therefore, when calculating the cross-zone heat flux loss, the compensation term accurately matches the actual heat loss, effectively solving the temperature reconstruction deviation problem caused by inaccurate thermal impedance assessment. By providing more accurate cross-zone thermal impedance assessment values, this solution makes temperature compensation for the actual heat source station more accurate. This provides more reliable input for subsequent steps such as determining whether the target zone is in a cross-zone boundary mismatch state, forcibly allocating cooling resources, and assessing the remaining heat carrying capacity of the edge area, thereby improving the overall accuracy and effectiveness of the temperature control method for the entire chip aging test equipment.
[0048] As a specific implementation method, when obtaining the thermal conductivity of the target object in the chip aging test equipment, FR-4 epoxy resin fiberglass board can be used for the aging board substrate, with a thermal conductivity of approximately 0.25 W / (m·K); the aluminum heat sink in the cooling structure has a thermal conductivity of approximately 205 W / (m·K); and the interface material between the chip and the aging board, if thermally conductive silicone grease is used, can have a thermal conductivity of 1.5 W / (m·K). When obtaining the geometric dimensions between each test station, the distance between the chip centers of adjacent test stations can be determined to be 15 mm, the thickness of the aging board substrate to be 2 mm, and the effective heat transfer cross-sectional area of the coolant channel to be 5 square millimeters, based on the equipment design drawings. Subsequently, the theoretical thermal resistance value of the heat conduction path corresponding to each test station is calculated using physical formulas for heat conduction, such as Fourier's law. For example, for a path of lateral heat conduction between two adjacent workstations via an aging board substrate, its thermal resistance can be approximately calculated as R = L / (k*A), where L is the distance between workstations, k is the thermal conductivity of the aging board substrate, and A is the effective heat transfer cross-sectional area. These calculated theoretical thermal resistance values are set as the cross-zone thermal resistance evaluation values for each test workstation. For example, the theoretical thermal resistance of heat transfer between adjacent workstations via the substrate might be 0.5 K / W, while the theoretical thermal resistance of a workstation crossing the cooling zone boundary might be 1.2 K / W due to the possible presence of additional partition structures.
[0049] Through the above technical solution, this application can accurately quantify the thermal conduction resistance caused by differences in physical structure between different test stations, avoiding deviations between the cross-regional thermal impedance assessment value and the actual physical environment. The theoretical thermal impedance value calculated based on physical parameters and thermal conduction physical formulas can truly reflect the heat transfer characteristics inside the equipment, thereby ensuring that the compensation term value can accurately match the actual heat loss situation when calculating the cross-regional heat flux. This effectively solves the temperature reconstruction deviation problem caused by inaccurate thermal impedance assessment, making the reconstructed true temperature of the actual heat source station more accurate, thereby improving the accuracy of the temperature control method of chip aging test equipment in identifying hidden overheating and passive heating, and helping to maintain the uniformity of the aging temperature field and the consistency of aging stress.
[0050] Traditional temperature control methods for chip aging test equipment, when calculating the heat flux lost across zones and adding it as a compensation term to the measured temperature of the actual heat source station to reconstruct the true temperature, struggle to accurately distinguish and quantify the superimposed effect of multi-source heat conduction when the passively heated station may be affected by multiple actual heat source stations or when the heat conduction path is complex. This leads to inaccurate calculation of the heat flux lost across zones, which in turn affects the accuracy of reconstructing the true temperature of the actual heat source station. Ultimately, these methods cannot effectively solve the problems of hidden overheating in high-load stations and passive temperature rise in adjacent low-load stations.
[0051] In some embodiments, the specific steps in step S24 include: S241. For each passively heated workstation, identify the actual heat source workstation that is physically adjacent to the passively heated workstation and use it as the target workstation; S242. For the temperature rise of the passively heated station, calculate the influence weight of each target station on the temperature rise of the passively heated station based on the cross-zone thermal resistance assessment value between the passively heated station and each target station. S243. Based on the influence weight, the temperature rise of the passively heated station is allocated to each target station to obtain the temperature rise of each target station relative to the passively heated station. S244. For each real heat source station, the shunting heat flux lost from the real heat source station to each passive heat source station is calculated based on the temperature rise of the passive heat source station that is physically adjacent to the real heat source station and the cross-zone thermal impedance assessment value corresponding to the passive heat source station. S245. The heat flux lost from the actual heat source station to all passively heated stations physically adjacent to the actual heat source station is summed to obtain the total cross-zone heat flux lost from the actual heat source station. S246. By adding the total heat flux lost across zones as a compensation term to the measured temperature of the actual heat source station, the reconstructed true temperature is obtained.
[0052] In the above technical solution, for each passively heated workstation, the actual heat source workstations physically adjacent to the passively heated workstation are identified as target workstations. This aims to clarify which actual heat source workstations may have a thermal impact on a specific passively heated workstation. This can be done based on a pre-constructed heat conduction topology network, by traversing the adjacent nodes of each passively heated workstation and selecting the nodes marked as actual heat source workstations as target workstations. Alternatively, a thermal influence radius can be set by calculating the physical distance between the passively heated workstation and all actual heat source workstations, and actual heat source workstations within this radius can be identified as target workstations.
[0053] For the temperature rise of the passively heated workstation, the influence weight of each target workstation on the temperature rise of the passively heated workstation is calculated based on the cross-zone thermal impedance assessment value between the passively heated workstation and each target workstation. This aims to quantify the contribution of different real heat source workstations to the temperature rise of the same passively heated workstation. The influence weight can be calculated by normalizing the inverse of the cross-zone thermal impedance assessment value; that is, the smaller the thermal impedance, the greater the influence weight. For example, if there are multiple target workstations, the proportion of the inverse of the thermal impedance between each target workstation and the passively heated workstation to the sum of the inverses of the thermal impedance of all target workstations can be calculated. Alternatively, the influence weight can also be obtained by establishing a heat conduction model, combining the heating power of each target workstation and its distance from the passively heated workstation, through simulation or empirical formulas.
[0054] Based on influence weights, the temperature rise of a passively heated workstation is allocated to each target workstation, resulting in the attributed temperature rise of each target workstation to the passively heated workstation. The aim is to decompose the total temperature rise of a passively heated workstation into the individual real heat source workstations according to their respective degrees of influence. This can be achieved by directly multiplying the temperature rise of the passively heated workstation by the influence weight of each target workstation. Alternatively, an iterative algorithm can be used to calculate the specific contribution of each target workstation to the temperature rise of the passively heated workstation based on the influence weights and a heat conduction physics model.
[0055] For each actual heat source station, the shunt heat flux lost from the actual heat source station to each passively heated station is calculated based on the assigned temperature rise of the physically adjacent passively heated station and the corresponding inter-zone thermal impedance assessment value of the passively heated station. This aims to calculate the specific amount of heat lost from each actual heat source station to each adjacent passively heated station via heat conduction. The shunt heat flux can be calculated using the physical formulas of heat conduction by dividing the assigned temperature rise by the corresponding inter-zone thermal impedance assessment value. Alternatively, it can be calculated by multiplying the assigned temperature rise by the inter-zone thermal conductivity (i.e., the reciprocal of the inter-zone thermal impedance assessment value).
[0056] The total heat flux lost by the actual heat source station to all physically adjacent passively heated stations is summed to obtain the total cross-zone heat flux loss of the actual heat source station. This aims to summarize the total heat lost from a single actual heat source station to all its affected passively heated stations. This can be achieved by simply summing the heat flux lost from the actual heat source station to all its adjacent passively heated stations.
[0057] By adding the total heat flux lost across zones as a compensation term to the measured temperature of the actual heat source station, a reconstructed true temperature is obtained. This aims to more accurately reflect the actual heating state of the actual heat source station by compensating for the temperature signal diluted by heat conduction. This can be achieved by directly adding the total heat flux lost across zones (after unit conversion) to the measured temperature, or by converting it to an equivalent temperature rise and then adding it to the measured temperature.
[0058] This application's solution, through a refined heat allocation logic, achieves a shift from macroscopic compensation to microscopic source tracing, ensuring the accuracy of cross-regional heat flux loss calculation and thus improving the accuracy of reconstructed true temperature. First, by identifying multiple target stations around the passively heated station, an analytical foundation for multi-source heat conduction is established. Next, the influence weights are calculated using the cross-regional thermal impedance assessment values between the passively heated station and each target station. The key to this step is the introduction of thermal impedance as the allocation basis, which objectively reflects the differences in the influence of different heat sources on the same heated point, thereby reasonably decomposing the temperature rise amplitude into the assigned temperature rise amplitude of each target station. Subsequently, the shunt heat flux is calculated based on the assigned temperature rise amplitude and the thermal impedance assessment value, ensuring that the heat loss calculation conforms to the laws of physical conduction. Finally, by accumulating the shunt heat fluxes to obtain the total cross-regional heat flux loss and performing temperature compensation, the reconstructed true temperature accurately reflects the actual heating state of the real heat source station under multi-source thermal interference. This solution effectively solves the problem of heat distribution under multiple heat source coupling, improves the accuracy of temperature sensing signal reconstruction, and provides reliable data support for subsequent precise scheduling of cooling resources.
[0059] As a specific implementation method, a concrete example is given below. Assume that there is a passively heated station P1 in the chip aging test equipment, and its physically adjacent real heat source stations include R1 and R2.
[0060] First, for the passively heated station P1, the system identifies R1 and R2 as its physically adjacent real heat source stations and uses them as target stations.
[0061] Next, assuming the temperature rise of the passively heated station P1 is ΔT_P1, the system calculates the influence weights W_R1_P1 and W_R2_P1 of R1 on the temperature rise of P1 based on the pre-set inter-regional thermal impedance assessment values recorded in the heat conduction topology network. For example, if the thermal impedance assessment value between R1 and P1 is low, then W_R1_P1 will be relatively high.
[0062] Then, based on these influence weights, the system allocates the temperature rise amplitude ΔT_P1 of the passively heated station P1 to R1 and R2. For example, the temperature rise amplitude assigned to P1 by R1 is ΔT_R1_P1 = ΔT_P1 * W_R1_P1, and the temperature rise amplitude assigned to P1 by R2 is ΔT_R2_P1 = ΔT_P1 * W_R2_P1.
[0063] Subsequently, for the actual heat source station R1, the system calculates the shunt heat flux Q_R1_P1 = ΔT_R1_P1 / Z_R1_P1, based on the assigned temperature rise ΔT_R1_P1 that flows from R1 to the passively heated station P1, and the estimated inter-regional thermal impedance Z_R1_P1 between R1 and P1. Similarly, if R1 is physically adjacent to another passively heated station P2, the shunt heat flux Q_R1_P2 that flows from R1 to P2 will also be calculated.
[0064] Finally, the system accumulates the shunting heat flux that flows from the real heat source station R1 to all physically adjacent passively heated stations (e.g., P1 and P2) to obtain the total cross-zone heat flux loss of R1, Q_R1_total = Q_R1_P1 + Q_R1_P2.
[0065] Finally, the system uses this total cross-regional heat flux loss Q_R1_total as a compensation term, adding it to the measured temperature T_R1_measured of the actual heat source station R1, thus obtaining the reconstructed true temperature of R1: T_R1_reconstructed = T_R1_measured + Q_R1_total. In this way, even if a passively heated station is affected by multiple actual heat source stations, the actual heat loss of each actual heat source station can be accurately calculated, thereby reconstructing its true temperature more accurately.
[0066] Through the above technical solution, this application achieves accurate calculation of cross-regional heat flux loss by finely identifying multi-source heat conduction paths and quantifying the contribution of each real heat source station to the temperature rise of the passively heated station. This accuracy significantly improves the accuracy of reconstructing the real temperature at the real heat source stations, enabling the chip aging test equipment to more realistically reflect the actual thermal load state of the chip. Therefore, this solution can effectively solve the problems of hidden overheating at high-load stations and passive heating at adjacent low-load stations, thereby ensuring the uniformity of the aging temperature field and the consistency of aging stress, and improving the reliability and efficiency of chip aging tests.
[0067] In some of the solutions mentioned above in this application, a forced allocation of cooling resources to the target zone is proposed to form a thermal barrier. However, in actual implementation, if the cooling output of adjacent zones is blindly increased, the passively heated workstations in the adjacent zones that were originally at the normal aging temperature may be over-cooled, thereby destroying the temperature consistency and stress stability required for aging tests, or causing energy waste.
[0068] In some embodiments, step S4, which involves controlling the cooling circuit near the boundary side of a cooling zone physically adjacent to the target zone to increase its cooling output to form a thermal barrier wall, includes the following specific steps: S41. Obtain the target aging temperature of the passively heated station in the cooling zone that is physically adjacent to the target zone; S42. Based on the total cross-zone heat flux lost from the actual heat source station to all cooling zones physically adjacent to the target zone, determine the upper limit of the cooling output of the cooling zones physically adjacent to the target zone for forming a thermal barrier wall. S43. Monitor the actual temperature of the passively heated workstation in the cooling zone that is physically adjacent to the target zone, and adjust the cooling output of the cooling circuit near the boundary side in the cooling zone that is physically adjacent to the target zone based on the difference between the actual temperature and the target aging temperature, and in combination with the upper limit of cooling output, so as to form a thermal barrier wall while ensuring that the actual temperature of the passively heated workstation is not lower than the target aging temperature.
[0069] The purpose of obtaining the target aging temperature for passively heated stations within the cooling zone physically adjacent to the target zone is to set a reference temperature for these stations. This ensures that, during the formation of the thermal barrier, these stations will not deviate from their intended aging conditions due to overcooling. This target aging temperature can be pre-stored in the test equipment's configuration database and set according to the type of chip under test and the aging test specifications; alternatively, it can be provided in real-time by the test task scheduling system as a specific temperature requirement for the current aging stage; or, the temperature can be dynamically obtained from a preset lookup table based on the chip model and aging profile.
[0070] Based on the total cross-zone heat flux lost from the actual heat source station to all cooling zones physically adjacent to the target zone, the upper limit of cooling output required for the cooling zones physically adjacent to the target zone to form a thermal barrier is determined. This step quantifies the maximum cooling capacity required by adjacent cooling zones to counteract lateral heat conduction, thereby avoiding unnecessary energy consumption and potential overcooling risks. This upper limit of cooling output can be calculated using a thermodynamic model based on the total cross-zone heat flux calculated in the previous step, combined with the heat exchange efficiency of the cooling loop and the thermophysical properties of the cooling medium; alternatively, it can be converted into a corresponding upper limit of cooling output based on a mapping function trained by a machine learning algorithm using historical operating data and experience; or, this upper limit value can be estimated and calibrated based on the design parameters of the cooling system and the theoretical thermal impedance of the heat conduction path.
[0071] Monitoring the actual temperature of passively heated stations within cooling zones physically adjacent to the target zone aims to obtain real-time temperature data for these stations, providing accurate feedback signals for subsequent closed-loop control. The actual temperature can be continuously sampled by deploying high-precision temperature sensors (e.g., thermocouples, platinum resistance thermometers, or NTC thermistors) below or near each passively heated station; alternatively, an infrared thermal imager can be used to scan the entire area, and image processing techniques can be employed to extract the surface temperature of each station; or, a distributed temperature sensing network integrated within the aging board substrate can be used to achieve real-time monitoring of the passively heated station temperature.
[0072] Based on the difference between the actual temperature and the target aging temperature, and in conjunction with the upper limit of cooling output, the cooling output of the cooling loops near the boundary in the cooling zones physically adjacent to the target zone is adjusted. This forms a thermal barrier while ensuring that the actual temperature of the passively heated workstations does not fall below the target aging temperature. This step is the core of intelligent cooling control, ensuring that temperature stability of adjacent low-power workstations is maintained while effectively blocking heat conduction. The cooling output of the cooling loops can be dynamically adjusted by a proportional-integral-derivative (PID) controller based on the difference between the actual temperature and the target aging temperature, while limiting the cooling output within a preset upper limit to prevent overcooling. Alternatively, a fuzzy logic controller can be used to intelligently adjust the flow rate of the cooling medium or the fan speed based on the temperature difference and the upper limit of cooling output. Furthermore, a model predictive control (MPC) algorithm can be used to comprehensively consider the current temperature, target temperature, upper limit of cooling output, and system dynamic response to optimize the cooling output strategy for a future period.
[0073] This application's solution, by introducing consideration of the target temperature of the passively heated workstations, quantification of the upper limit of cooling output, and a dynamic adjustment mechanism based on real-time temperature feedback, ensures that the formation of the thermal barrier wall effectively blocks heat conduction while maintaining temperature stability of adjacent low-power workstations, avoiding over-cooling. Specifically, after identifying that the target partition is in a cross-region boundary mismatch state, in order to respond to the instruction to control the cooling loops near the boundary side in the physically adjacent cooling partition to increase cooling output to form a thermal barrier wall, this solution first obtains the target aging temperature of the passively heated workstations in the physically adjacent cooling partition. This sets a lower limit for subsequent cooling adjustment, ensuring that the construction of the thermal barrier wall does not come at the expense of the testing environment of adjacent partitions. Secondly, based on the total cross-region heat flux lost from the actual heat source workstations to all physically adjacent cooling partitions, the upper limit of cooling output for the physically adjacent cooling partitions to form the thermal barrier wall is determined. This feature utilizes the heat loss data calculated in the preceding steps to ensure that the cooling output effectively offsets the heat transferred laterally without causing resource waste or temperature runaway due to overcooling. Finally, by monitoring the actual temperature of the passively heated workstations in the cooling zones physically adjacent to the target zone, and dynamically adjusting based on the difference between the actual temperature and the target aging temperature, combined with the upper limit of the cooling output, closed-loop control of the cooling loop is achieved. This approach not only effectively blocks heat transfer across zones but also ensures, through a real-time feedback mechanism, that the actual temperature of the passively heated workstations remains above the target aging temperature while forming a thermal barrier. This ensures the heat dissipation needs of high-power workstations while also considering the temperature uniformity and stress consistency of aging tests in adjacent zones. This solution, through fine-grained control of the cooling output of adjacent cooling zones, avoids the overcooling problem that may result from blindly increasing cooling output. This allows the strategies for forced allocation of cooling resources and the formation of thermal barriers proposed in the preceding steps to be implemented more accurately and efficiently, thereby improving the robustness and intelligence of the temperature control method of the entire chip aging test equipment.
[0074] The following is a concrete example. In a chip aging test apparatus, when a target partition is determined to be in a cross-zone boundary mismatch state due to high-power tasks concentrated in the edge area, and there is a passively heated station in its physically adjacent cooling partition, the system first obtains the target aging temperature of the passively heated station from a preset test configuration file, for example, set to 85°C. Next, the system calculates the total cross-zone heat flux lost from the actual heat source station to the adjacent cooling partition based on the previous steps, for example, 50W, and combines this with the heat exchange efficiency of the cooling loop to determine the upper limit of the cooling output of the cooling loop near the boundary side in the adjacent cooling partition as 60W. Subsequently, the system continuously monitors the current actual temperature of each passively heated station using NTC thermistors deployed below it. Suppose that at a certain moment, the actual temperature of a passively heated station is 86°C, higher than the target aging temperature. At this time, a PID controller will send a command to the actuator of the cooling loop based on a 1°C temperature difference and considering the 60W upper limit of cooling output. The actuator can be an electrically operated proportional valve that increases the coolant flow rate by adjusting its opening, thereby increasing cooling output. If the actual temperature drops and approaches 85°C, the controller will correspondingly reduce the valve opening to prevent the temperature from dropping further below the target aging temperature. In this way, the adjacent cooling zones form an effective thermal barrier at the boundary, absorbing lateral heat from the target zone while ensuring that the temperature of the passively heated station is maintained at or slightly above 85°C, preventing overcooling.
[0075] Through the above technical solution, this application effectively addresses the problem in chip aging test equipment where dynamic task scheduling leads to high-load stations concentrated in the edge area of a cooling zone near an adjacent zone, while the adjacent zone is under low load. Blindly increasing the cooling output of the adjacent zone could cause over-cooling of passively heated stations within that zone, which were originally at normal aging temperatures. This could disrupt the temperature consistency and stress stability required for aging testing, or result in energy waste. This solution precisely obtains the target aging temperature of the passively heated stations and determines the upper limit of cooling output based on the cross-zone heat flux loss. By real-time monitoring and dynamic adjustment of the cooling loop output, it achieves intelligent construction of the thermal barrier. This not only ensures that the thermal barrier can efficiently absorb laterally conducted heat and effectively block heat flow, but more importantly, it avoids over-cooling of adjacent low-load stations, thus maintaining temperature uniformity and stress stability during the aging test and ensuring the accuracy of the test results. Simultaneously, this refined cooling control avoids unnecessary energy waste and improves the operating efficiency and economy of the equipment.
[0076] In some embodiments, step S5, which involves assessing the remaining thermal carrying capacity of the edge region based on the difference between the reconstructed true temperature and a preset physical damage safety threshold, includes the following specific steps: S51. Obtain the upper limit of the cooling system heat dissipation power in the edge region; S52. Obtain the real-time heat dissipation requirements of the edge region; the real-time heat dissipation requirements are determined based on the difference between the reconstructed true temperature and the physical damage safety threshold. S53. Obtain the heat capacity of the edge region and the preset target temperature rise rate; S54. Calculate the remaining heat carrying capacity of the edge region by combining the difference between the upper limit of the cooling system's heat dissipation power and the real-time heat dissipation demand with the heat capacity and the target temperature rise rate.
[0077] Obtaining the upper limit of the cooling system's heat dissipation power in the edge region refers to determining the maximum heat dissipation load that the edge region can withstand. This upper limit can be pre-calibrated or calculated in real time based on parameters such as the cooling system's design specifications, the flow rate of the cooling medium, the efficiency of the heat exchanger, and the power of the fan or pump. For example, it can be obtained by referring to the maximum cooling capacity specified in the cooling system's design documents, or by testing and calibrating the actual heat dissipation capacity of the cooling system under full-load operation. Another approach is to estimate the current maximum heat dissipation capacity of the cooling system in real time based on sensor data inside the cooling system (such as coolant inlet and outlet temperatures, flow rate, fan speed, etc.) combined with a thermodynamic model.
[0078] Obtaining the real-time heat dissipation requirement of an edge region refers to dynamically quantifying the required heat dissipation based on the difference between the current thermal state of that region and its safety limits. This real-time heat dissipation requirement is determined by the difference between the actual reconstructed temperature and the physical damage safety threshold. For example, when the actual reconstructed temperature is much lower than the physical damage safety threshold, the real-time heat dissipation requirement is low; while when the actual reconstructed temperature is close to or even approaches the physical damage safety threshold, the real-time heat dissipation requirement increases significantly. This can be achieved through a preset function or lookup table that maps the temperature difference to the corresponding heat dissipation requirement. For example, by taking a weighted average or the maximum value of these differences and multiplying it by a preset heat load conversion coefficient, the additional heat that the edge region needs to dissipate to prevent chip damage is obtained, i.e., the real-time heat dissipation requirement.
[0079] The system acquires the heat capacity of the edge region and the preset target temperature rise rate. The heat capacity of the edge region refers to the heat absorbed by all materials within that region (such as the aging board substrate, chip, and fixtures) when the temperature increases by one degree Celsius. Heat capacity can be calculated using the material's density, specific heat capacity, and volume, or obtained through experimental measurements, or by consulting equipment design documents or material databases. The system acquires the mass and specific heat capacity of the main constituent materials within the edge region, such as the aging board substrate, chip packaging, and heat sink, and then sums them to obtain the total heat capacity of the edge region. This total heat capacity is used for subsequent calculations of the remaining heat load capacity. The preset target temperature rise rate refers to the maximum rate at which the temperature of the edge region is allowed to rise under safe operating conditions. This rate can be set according to the chip's thermal resistance characteristics, the acceleration factor requirements of the aging test, and the system's temperature stability requirements. For example, it can be set to no more than 0.1°C per second to ensure that the system has sufficient response time to intervene during sudden high-power tasks.
[0080] The remaining heat handling capacity of the edge region is calculated by combining the difference between the upper limit of the cooling system's heat dissipation power and the real-time heat dissipation demand, along with the heat capacity and the target temperature rise rate. This calculation aims to quantify how much additional heat load the edge region can withstand under the current cooling capacity and thermal conditions, or in other words, how long it can maintain high-power operation without triggering the physical damage safety threshold. Specifically, the heat dissipation demand is divided by (heat capacity multiplied by the target temperature rise rate) to obtain a time value, which represents the time required for the edge region temperature to reach the target rise rate under the current additional heat dissipation capacity. This time value, or its reciprocal (representing the additional heat that can be handled per unit time), is the remaining heat handling capacity.
[0081] This application's solution constructs a quantitative residual heat carrying capacity assessment model by introducing key physical parameters such as the upper limit of cooling system heat dissipation power, real-time heat dissipation requirements, heat capacity, and target temperature rise rate, achieving a leap from qualitative assessment to quantitative calculation. By obtaining the upper limit of cooling system heat dissipation power in the edge region, the maximum heat dissipation boundary that this region can provide at the physical level is clarified, setting a benchmark for subsequent heat carrying capacity calculations. By obtaining the real-time heat dissipation requirements of the edge region and linking them to the difference between the reconstructed true temperature and the physical damage safety threshold, the urgency between the current thermal state of the workstation and the safety limit can be dynamically reflected, ensuring the real-time nature and accuracy of heat dissipation requirement calculations. By obtaining the heat capacity of the edge region and the preset target temperature rise rate, thermal inertia and temperature rise control constraints are introduced, so that the assessment process not only considers the current heat dissipation balance but also predicts the temperature change trend when performing high-power tasks, avoiding the risk of overheating due to heat accumulation. By combining the difference between the upper limit of the cooling system's heat dissipation power and the real-time heat dissipation demand with the heat capacity and the target temperature rise rate, we can accurately determine the additional heat load space that the edge area can bear under the premise of ensuring safety, thus providing a scientific basis for decision-making on mission delay interception or spatial reallocation scheduling.
[0082] In one specific implementation, the system first reads the upper limit of the cooling system's heat dissipation power in the edge region from a preset equipment parameter database. For example, the maximum heat dissipation capacity of the cooling module in this region is 50W. Simultaneously, the system monitors the reconstructed true temperature of each real heat source station within the edge region in real time and compares it with a preset physical damage safety threshold (e.g., 125℃). Assuming the current reconstructed true temperature is 110℃, the real-time heat dissipation requirement can be calculated using a linear function. For example, if an additional 2W of heat dissipation is required for every 1℃ increase, then the real-time heat dissipation requirement is (125-110)*2=30W. Next, the system obtains the heat capacity of the edge region from a material property database. For example, the heat capacity of this region is 10J / ℃. The preset target temperature rise rate is 0.5℃ / s. Based on this, the system calculates the difference between the upper limit of the cooling system's heat dissipation power and the real-time heat dissipation requirement (50W-30W=20W), indicating that the cooling system currently has an additional 20W of heat dissipation capacity. Then, combining the heat capacity and the target temperature rise rate, the remaining heat load capacity is calculated. For example, it can be calculated how long it would take for the region to reach the upper limit of the chip's physical damage safety threshold if high-power tasks are continuously added under the current maximum combined cooling compensation. If this estimated time is extremely short, for example, 20W / (10J / ℃*0.5℃ / s)=4s, it indicates that the thermal carrying capacity of this edge region is close to absolute saturation and is in a state of extremely high thermal risk. Conversely, if the estimated time is long, it indicates that the region still has the capacity to handle a certain amount of high-power tasks.
[0083] Through the above technical solution, the system can accurately quantify and assess the remaining heat load capacity of the edge region, thus avoiding the blindness of judging solely based on temperature differences. This assessment method comprehensively considers the physical limits of the cooling system, the current real-time heat load, and thermal inertia effects, enabling the task scheduling system to make decisions based on more scientific and comprehensive data when facing high-power tasks. This effectively improves the thermal safety management level during chip aging testing, reduces the risk of chip damage due to overheating, and optimizes the efficiency and reliability of task scheduling.
[0084] Reference Appendix Figure 2 This invention provides a temperature control system for a chip aging test equipment (the temperature control system adopts the chip aging test equipment temperature control method of the above embodiment, and the specific process is described in the corresponding steps above), including: The data acquisition unit 100 is used to acquire a pre-constructed thermal conduction topology network for the chip aging test equipment; the thermal conduction topology network is constructed based on the target information of each test station in the chip aging test equipment. The temperature compensation unit 200 is used to identify test stations that belong to real heat sources and use them as real heat source stations based on the heat conduction topology network, and to identify test stations that belong to passive heating and use them as passive heating stations. It also compensates for the measured temperature of the real heat source stations to obtain the reconstructed real temperature of the real heat source stations. The status judgment unit 300 is used to take the cooling zone to which the real heat source station belongs as the target zone; based on the heat conduction topology network, it determines whether the target zone is in a cross-zone boundary mismatch state according to the reconstructed real temperature. The first control unit 400 is used to forcibly allocate cooling resources to the target zone in the cross-zone boundary mismatch state, so that more cooling resources are allocated to the edge area, and control the cooling circuits near the boundary side in the cooling zone that is physically adjacent to the target zone to increase the cooling output to form a thermal barrier wall. The second control unit 500 is used to assess the remaining heat carrying capacity of the edge region based on the difference between the reconstructed real temperature and the preset physical damage safety threshold, and to perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region according to the remaining heat carrying capacity. The verification and correction unit 600 is used to verify the temperature field uniformity of the real heat source station and the passive heat-receiving station after performing the control of the preceding steps, and to correct the heat conduction topology network based on the verification results.
[0085] In this context, the units described as separate components may or may not be physically separate. Similarly, the components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0086] Furthermore, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0087] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0088] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A temperature control method for a chip aging test device, characterized in that, Includes the following steps: S1. Obtain a pre-constructed thermal conduction topology network for the chip aging test equipment; the thermal conduction topology network is constructed based on the target information of each test station in the chip aging test equipment. S2. Based on the heat conduction topology network, identify the test station that belongs to the real heat source and use it as the real heat source station, and identify the test station that belongs to the passive heat source and use it as the passive heat source station. Then, by compensating the measured temperature of the real heat source station, the reconstructed real temperature of the real heat source station is obtained. S3. Take the cooling zone to which the actual heat source station belongs as the target zone; based on the heat conduction topology network, determine whether the target zone is in a cross-zone boundary mismatch state according to the reconstructed real temperature; S4. In the cross-zone boundary mismatch state, the cooling resources of the target zone are forcibly allocated so that more cooling resources are allocated to the edge area, and the cooling circuits near the boundary side in the cooling zones physically adjacent to the target zone are controlled to increase the cooling output to form a thermal barrier wall. S5. Based on the difference between the reconstructed true temperature and the preset physical damage safety threshold, assess the remaining heat carrying capacity of the edge region, and according to the remaining heat carrying capacity, perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region. S6. After performing the control in the preceding steps, perform temperature field uniformity verification on the real heat source station and the passive heat-receiving station, and correct the heat conduction topology network based on the verification results.
2. The temperature control method for the chip aging test equipment according to claim 1, characterized in that, The target information includes the spatial location information of each test station, the identification of its corresponding cooling zone, the task switching instruction sequence, and the real-time temperature sampling sequence.
3. The temperature control method for the chip aging test equipment according to claim 2, characterized in that, The specific steps in step S2 include: S21. Obtain the physical arrangement distance of each test station from the spatial location information, and determine the partition boundary relationship between the cooling partitions to which each test station belongs based on the corresponding cooling partition identifier; S22. Based on the heat conduction topology network, and according to the physical arrangement distance and the partition boundary relationship, set the cross-zone thermal impedance evaluation value corresponding to each of the test stations; S23. Extract the temperature change rate from the real-time temperature sampling sequence, and identify the real heat source station and the passively heated station by matching the target time when the temperature change rate exceeds the preset change rate threshold with the timestamp in the task switching instruction sequence; S24. Based on the temperature rise of the passively heated station and the cross-zone thermal impedance assessment value corresponding to the passively heated station, calculate the cross-zone heat flux loss, and obtain the reconstructed true temperature by adding the cross-zone heat flux loss as a compensation term to the measured temperature of the real heat source station.
4. The temperature control method for the chip aging test equipment according to claim 3, characterized in that, The specific steps in step S22 include: S221. Obtain the thermal conductivity of the target object in the chip aging test equipment; S222. Obtain the geometric dimensions between each of the test stations; the geometric dimensions are determined based on the physical arrangement distance and the partition boundary relationship; S223. Using the physical formula of thermal conduction, based on the thermal conductivity and the geometric dimensions, calculate the theoretical thermal impedance value of the thermal conduction path corresponding to each of the test stations, and set the theoretical thermal impedance value as the cross-zone thermal impedance evaluation value corresponding to each of the test stations.
5. The temperature control method for the chip aging test equipment according to claim 3, characterized in that, The specific steps in step S24 include: S241. For each of the passively heated workstations, identify the actual heat source workstations that are physically adjacent to the passively heated workstations and designate them as target workstations; S242. For the temperature rise of the passively heated station, calculate the influence weight of each target station on the temperature rise of the passively heated station based on the cross-zone thermal resistance assessment value between the passively heated station and each target station. S243. Based on the influence weight, the temperature rise of the passively heated station is allocated to each of the target stations to obtain the temperature rise of each target station relative to the passively heated station. S244. For each of the real heat source stations, the shunting heat flux lost from the real heat source station to each of the passive heat receiving stations is calculated based on the assigned temperature rise of the passive heat receiving station that is physically adjacent to the real heat source station and the cross-zone thermal impedance assessment value corresponding to the passive heat receiving station. S245. The heat flux diverted from the actual heat source station to all the passively heated stations physically adjacent to the actual heat source station is summed to obtain the total cross-zone heat flux of the actual heat source station. S246. The reconstructed true temperature is obtained by adding the total cross-regional heat flux loss as a compensation term to the measured temperature of the actual heat source station.
6. The temperature control method for the chip aging test equipment according to claim 1, characterized in that, The specific steps in step S3 include: S31. Using the difference between the reconstructed true temperature and the target aging temperature as weights, calculate the dynamic centroid of the heat load space within the target partition, and determine the centroid offset based on the Euclidean distance between the dynamic centroid of the heat load space and the geometric center of the target partition. S32. When the center of gravity offset exceeds a preset offset threshold and the edge high load space density of the edge region of the target partition reaches a preset risk threshold, the target partition is determined to be in the cross-region boundary mismatch state; the edge high load space density is calculated based on the number of actual heat source stations in the edge region and the area of the edge region.
7. The temperature control method for the chip aging test equipment according to claim 6, characterized in that, In step S31, the specific steps for calculating the dynamic centroid of the heat load space within the target zone, using the difference between the reconstructed true temperature and the target aging temperature as weights, include: Using the difference between the reconstructed real temperature and the target aging temperature as weights, the spatial coordinates of each real heat source station within the target partition are calculated by weighted averaging to obtain the dynamic centroid of the heat load space within the target partition.
8. The temperature control method for the chip aging test equipment according to claim 1, characterized in that, In step S4, the specific steps of controlling the cooling circuit near the boundary side of the cooling zone physically adjacent to the target zone to increase the cooling output to form a thermal barrier wall include: S41. Obtain the target aging temperature of the passively heated station within the cooling zone that is physically adjacent to the target zone; S42. Based on the total cross-zone heat flux lost from the actual heat source station to all cooling zones physically adjacent to the target zone, determine the upper limit of the cooling output of the cooling zones physically adjacent to the target zone for forming a thermal barrier wall. S43. Monitor the actual temperature of the passively heated station in the cooling zone physically adjacent to the target zone, and adjust the cooling output of the cooling circuit near the boundary side in the cooling zone physically adjacent to the target zone based on the difference between the actual temperature and the target aging temperature, and in conjunction with the upper limit of cooling output, so as to form a thermal barrier wall while ensuring that the actual temperature of the passively heated station is not lower than the target aging temperature.
9. The temperature control method for the chip aging test equipment according to claim 1, characterized in that, In step S5, the specific steps for evaluating the remaining thermal carrying capacity of the edge region based on the difference between the reconstructed true temperature and the preset physical damage safety threshold include: S51. Obtain the upper limit of the heat dissipation power of the cooling system in the edge region; S52. Obtain the real-time heat dissipation requirement of the edge region; the real-time heat dissipation requirement is determined based on the difference between the reconstructed true temperature and the physical damage safety threshold; S53. Obtain the heat capacity of the edge region and the preset target temperature rise rate; S54. Calculate the remaining heat carrying capacity of the edge region by taking the difference between the upper limit of the cooling system's heat dissipation power and the real-time heat dissipation requirement, combined with the heat capacity and the target temperature rise rate.
10. A temperature control system for a chip aging test device, characterized in that, include: The data acquisition unit is used to acquire the heat conduction topology network pre-built for the chip aging test equipment; The heat conduction topology network is constructed based on the target information of each test station in the chip aging test equipment; The temperature compensation unit is used to identify test stations belonging to real heat sources and use them as real heat source stations based on the heat conduction topology network, and to identify test stations belonging to passive heating and use them as passive heating stations. The unit also compensates for the measured temperature of the real heat source stations to obtain the reconstructed real temperature of the real heat source stations. The status determination unit is used to identify the cooling zone to which the actual heat source station belongs as the target zone. Based on the heat conduction topology network, and according to the reconstructed true temperature, it is determined whether the target partition is in a cross-region boundary mismatch state; The first control unit is used to forcibly allocate cooling resources to the target partition under the cross-zone boundary mismatch state, so that more cooling resources are allocated to the edge area, and control the cooling circuits near the boundary side in the cooling partitions physically adjacent to the target partition to increase the cooling output to form a thermal barrier wall. The second control unit is used to assess the remaining heat carrying capacity of the edge region based on the difference between the reconstructed real temperature and the preset physical damage safety threshold, and to perform delayed interception or spatial reallocation scheduling for high-power tasks planned to be executed in the edge region according to the remaining heat carrying capacity. The verification and correction unit is used to perform temperature field uniformity verification on the real heat source station and the passive heat-receiving station after performing the control of the preceding steps, and to correct the heat conduction topology network according to the verification results.