A server backplane transmission link loss evaluation and optimization method and system

CN122507599APending Publication Date: 2026-08-04SHENZHEN SUNSHINE GOOD CIRCUIT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN SUNSHINE GOOD CIRCUIT CO LTD
Filing Date
2026-05-13
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0005]鉴于此,本发明提出了一种服务器背板传输链路损耗评估与优化方法及系统,旨在解决现有服务器背板传输链路损耗评估方式与通信调度过程缺少关联,难以根据链路动态损耗状态调整通信任务,导致高负载链路损耗评估和通信调度匹配性不足的问题

Benefits of technology

[0054]Compared with existing technologies, the advantages of this invention are as follows: By collecting historical communication data, historical task topology data, historical temperature data, and link parameters to establish a bandwidth utilization prediction model, link loss assessment no longer relies solely on the fixed design parameters of the backplane link, but can be predicted by combining the communication load change patterns in historical distributed training tasks; by acquiring the current task communication topology information, current bandwidth utilization, and current temperature data during the execution of the current distributed training task, and predicting the bandwidth utilization of each link within a future time window, the load changes of each link in subsequent communication stages can be obtained in advance; by combining the predicted bandwidth utilization value, current temperature data, and server backplane thermal conductivity parameters to calculate the predicted temperature, and by performing temperature correction on conductor loss and dielectric loss based on the predicted temperature, temperature changes caused by link load changes can be incorporated into the link loss prediction. The link loss assessment process improves the matching between predicted link loss and actual operating status. By calculating the allowable link loss based on data transmission rate, encoding method, load balancing configuration, and error rate requirements, and determining the loss margin by the difference between the allowable link loss and the predicted link loss, a unified quantitative judgment basis can be provided for the loss status of different links under different operating configurations. By scheduling communication tasks based on the loss margin of each link, the communication scheduling scheme can be generated in conjunction with the link loss status, reducing the situation where links with predicted link loss exceeding the allowable link loss continue to carry high-load communication tasks. By collecting actual bandwidth utilization, actual temperature data, and communication performance data after the execution of communication tasks, and updating the bandwidth utilization prediction model and temperature correction coefficient, the subsequent link load prediction and link loss correction processes can be continuously adjusted according to the actual operating data of the server. Therefore, this invention can realize continuous processing of link load prediction, link temperature prediction, link loss assessment, communication task scheduling, and feedback updates during the operation of distributed training tasks, improving the accuracy of server backplane transmission link loss assessment and enhancing the matching degree between loss assessment results and communication scheduling processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122507599A_ABST
    Figure CN122507599A_ABST
Patent Text Reader

Abstract

This invention relates to the field of transmission link loss assessment technology, and discloses a method and system for assessing and optimizing transmission link loss on a server backplane. The method includes: collecting historical communication data, historical task topology data, historical temperature data, and link parameters to establish a bandwidth utilization prediction model; predicting the bandwidth utilization of each link within a future time window in the current distributed training task; calculating the predicted temperature based on the predicted bandwidth utilization, current temperature data, and thermal conductivity parameters; correcting conductor loss and dielectric loss based on the predicted temperature to obtain the predicted link loss; determining the loss margin based on the allowable link loss and generating a communication scheduling scheme; and collecting actual operating data after execution to update the prediction model and temperature correction coefficients. This invention improves the accuracy of link loss assessment and the matching performance of communication scheduling through closed-loop processing of load prediction, temperature correction assessment, and feedback scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of transmission link loss assessment technology, and more specifically, to a method and system for assessing and optimizing transmission link loss on server backplanes. Background Technology

[0002] As the scale of large-scale model training continues to increase, AI servers typically employ a system architecture where multiple computing nodes and accelerator cards work together. The server backplane, as a crucial interconnecting medium for high-speed data exchange between computing nodes, accelerator cards, switching chips, and storage units, needs to handle a large number of high-bandwidth, low-latency data transmission tasks. The transmission loss of the backplane link directly affects the amplitude attenuation of high-speed signals in the link, the signal quality at the receiving end, and the stability of data transmission. Therefore, assessing the loss of the server backplane transmission link is a critical aspect of server system design and operation maintenance.

[0003] Current server backplane link loss assessments are typically based on link length, material parameters, signal frequency, connector parameters, and simulation or test results from the design phase. This type of assessment is primarily geared towards backplane design verification or factory testing scenarios. It usually uses link parameters under fixed operating conditions, fixed temperatures, or uniform temperatures to estimate the transmission loss of each link in order to determine whether the backplane links meet the signal integrity requirements at the corresponding data rates.

[0004] However, when AI servers perform large-scale distributed training tasks, the communication patterns differ significantly across training stages, and the amount of data exchanged between different computing nodes also varies with the task topology, parallel strategies, and training iterations. This results in an uneven load on the transmission links within the server backplane; some links may bear high communication pressure for extended periods, while others operate under relatively low load. The backplane operating environment also fluctuates with the overall server load, placing different links under varying actual operating conditions. Summary of the Invention

[0005] In view of this, the present invention proposes a method and system for evaluating and optimizing the transmission link loss of a server backplane, aiming to solve the problem that the existing server backplane transmission link loss evaluation method lacks correlation with the communication scheduling process, making it difficult to adjust communication tasks according to the dynamic loss status of the link, resulting in insufficient matching between high-load link loss evaluation and communication scheduling.

[0006] In one aspect, this invention proposes a method for evaluating and optimizing server backplane transmission link loss, comprising:

[0007] Collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and establish a bandwidth utilization prediction model.

[0008] During the execution of the current distributed training task, the communication topology information of the current task and the current bandwidth utilization and current temperature data of each link are obtained. The current communication topology information, current bandwidth utilization and current temperature data are then input into the bandwidth utilization prediction model to obtain the predicted bandwidth utilization value of each link within the future time window.

[0009] Based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane, the predicted temperature of each link within the future time window is calculated.

[0010] Based on the predicted temperature and the link parameters corresponding to each link, the conductor loss and dielectric loss of each link at the reference temperature are pre-calculated and temperature-corrected to obtain the predicted link loss of each link within the future time window.

[0011] Based on the data transmission rate, encoding method, equalization configuration and error rate requirements of each link, calculate the allowable link loss for each link, and determine the loss margin for each link based on the difference between the allowable link loss and the predicted link loss.

[0012] Based on the loss margin of each link, communication tasks are scheduled to obtain a communication scheduling scheme.

[0013] The communication task is executed based on the communication scheduling scheme, and the actual bandwidth utilization, actual temperature data and communication performance data are collected after the communication task is executed.

[0014] Based on the actual bandwidth utilization, actual temperature data, and communication performance data, update the bandwidth utilization prediction model and the temperature correction coefficient for temperature correction.

[0015] Furthermore, when establishing a bandwidth utilization prediction model, the following are included:

[0016] The historical communication data is divided according to the training iteration rounds and communication stages, and the communication data volume, communication duration, participating computing nodes and communication type of each communication stage are extracted.

[0017] The communication connection relationships between participating computing nodes are determined based on the historical task topology data, and the communication connection relationships are mapped to the corresponding links in the server backplane;

[0018] Calculate the historical bandwidth utilization rate of each link based on the amount of communication data carried, the duration of communication, and the rated bandwidth of the link during the corresponding communication phase.

[0019] The bandwidth utilization prediction model is trained by aligning the historical bandwidth utilization, historical temperature data, link parameters, and task topology data of the corresponding communication stage for each link according to time.

[0020] Furthermore, when obtaining the predicted bandwidth utilization of each link within the future time window, the following are included:

[0021] Analyze the current task communication topology information to obtain the communication node pairs, communication data volume, and communication phase order in the current distributed training task;

[0022] Based on the correspondence between server backplane ports and computing nodes, the communication node pairs are mapped to corresponding backplane links;

[0023] Input the current bandwidth utilization rate, current temperature data, and corresponding communication stage data of each link into the bandwidth utilization prediction model to obtain the predicted bandwidth utilization rate of each link in the continuous sampling period within the future time window.

[0024] When the same link carries multiple communication tasks within the same sampling period, the bandwidth utilization prediction values ​​corresponding to the multiple communication tasks are merged.

[0025] Furthermore, when calculating the predicted temperature corresponding to each link within the future time window, the following steps are included:

[0026] Based on the predicted bandwidth utilization, rated bandwidth, and parameters of each link, calculate the heat generation of each link within the future time window.

[0027] Based on the thermal conduction parameters of the server backplane, determine the thermal coupling relationship between each link and the heat exchange relationship between each link and the heat dissipation boundary;

[0028] Using the current temperature data of each link as the initial temperature, the predicted temperature of each link is calculated periodically based on the heat generation, thermal coupling relationship, and heat exchange relationship of the link.

[0029] Furthermore, when obtaining the predicted link loss, the following are included:

[0030] Based on the link length, conductor material parameters, dielectric material parameters, and signal operating frequency of each link, calculate the conductor loss and dielectric loss of each link at the reference temperature.

[0031] The conductor loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of resistance of the conductor material.

[0032] The dielectric loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of the dielectric material.

[0033] The predicted link loss is obtained by adding the conductor loss, dielectric loss, conductor loss temperature correction, and dielectric loss temperature correction at the reference temperature.

[0034] Furthermore, when determining the loss margin for each link, the following is included:

[0035] Based on the data transmission rate, encoding method, and error rate of each link, determine the signal quality requirements of the receiving end of the corresponding link;

[0036] Based on the balanced configuration of each link, determine the balanced compensation amount for the corresponding link;

[0037] Calculate the allowable link loss for each link based on the output signal quality of the transmitter, the signal quality requirements of the receiver, and the equalization compensation amount.

[0038] The link loss margin is obtained by subtracting the predicted link loss from the allowed link loss of the same link.

[0039] Furthermore, when obtaining the communication scheduling scheme, it includes:

[0040] The link to be adjusted is determined based on the loss margin of each link, and the communication task carried by the link to be adjusted is determined.

[0041] Based on the communication phase sequence and computational dependencies of the communication tasks, select adjustable communication tasks;

[0042] Among the candidate links connected to the sending node and receiving node of the adjustable communication task, the candidate link with a loss margin greater than that of the link to be adjusted is selected.

[0043] The adjustable communication task is adjusted from the link to be adjusted to the selected candidate link to obtain a communication scheduling scheme that includes the communication path adjustment results.

[0044] Furthermore, when the communication path cannot be adjusted, this includes:

[0045] Determine the collective communication phase in which the link to be adjusted participates;

[0046] When there are optional communication topologies in the aggregated communication phase, select the communication topology that reduces the amount of communication data carried by the link to be adjusted, and update the communication connection relationship between the participating computing nodes.

[0047] When the data fragmentation method in the aggregated communication phase is adjustable, the amount of data fragments allocated to the communication node pairs corresponding to the links to be adjusted is reduced, and the reduced amount of data fragments is allocated to the communication node pairs corresponding to the links with larger loss margins;

[0048] The communication scheduling scheme is obtained based on the updated communication connection relationship or the adjusted data fragmentation amount.

[0049] Furthermore, updating the bandwidth utilization prediction model and the temperature correction coefficient includes:

[0050] Calculate the actual bandwidth utilization of each link based on the actual amount of data transmitted after the communication task is executed, the actual communication duration, and the rated bandwidth of the link.

[0051] The bandwidth utilization prediction model is updated based on the deviation between the actual bandwidth utilization of each link and the corresponding predicted bandwidth utilization.

[0052] Based on actual temperature data and communication performance data, the actual link loss of each link within the corresponding sampling period is inferred.

[0053] Based on the deviation between the actual link loss and the corresponding predicted link loss, update the conductor loss temperature correction factor and the dielectric loss temperature correction factor.

[0054] Compared with existing technologies, the advantages of this invention are as follows: By collecting historical communication data, historical task topology data, historical temperature data, and link parameters to establish a bandwidth utilization prediction model, link loss assessment no longer relies solely on the fixed design parameters of the backplane link, but can be predicted by combining the communication load change patterns in historical distributed training tasks; by acquiring the current task communication topology information, current bandwidth utilization, and current temperature data during the execution of the current distributed training task, and predicting the bandwidth utilization of each link within a future time window, the load changes of each link in subsequent communication stages can be obtained in advance; by combining the predicted bandwidth utilization value, current temperature data, and server backplane thermal conductivity parameters to calculate the predicted temperature, and by performing temperature correction on conductor loss and dielectric loss based on the predicted temperature, temperature changes caused by link load changes can be incorporated into the link loss prediction. The link loss assessment process improves the matching between predicted link loss and actual operating status. By calculating the allowable link loss based on data transmission rate, encoding method, load balancing configuration, and error rate requirements, and determining the loss margin by the difference between the allowable link loss and the predicted link loss, a unified quantitative judgment basis can be provided for the loss status of different links under different operating configurations. By scheduling communication tasks based on the loss margin of each link, the communication scheduling scheme can be generated in conjunction with the link loss status, reducing the situation where links with predicted link loss exceeding the allowable link loss continue to carry high-load communication tasks. By collecting actual bandwidth utilization, actual temperature data, and communication performance data after the execution of communication tasks, and updating the bandwidth utilization prediction model and temperature correction coefficient, the subsequent link load prediction and link loss correction processes can be continuously adjusted according to the actual operating data of the server. Therefore, this invention can realize continuous processing of link load prediction, link temperature prediction, link loss assessment, communication task scheduling, and feedback updates during the operation of distributed training tasks, improving the accuracy of server backplane transmission link loss assessment and enhancing the matching degree between loss assessment results and communication scheduling processes.

[0055] On the other hand, this application also provides a server backplane transmission link loss assessment and optimization system for implementing the above-mentioned server backplane transmission link loss assessment and optimization method, including:

[0056] The model training module is used to collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and to establish a bandwidth utilization prediction model.

[0057] The bandwidth prediction module is used to acquire the current task communication topology information, the current bandwidth utilization rate and the current temperature data of each link during the execution of the current distributed training task, and input the current task communication topology information, current bandwidth utilization rate and current temperature data into the bandwidth utilization prediction model to obtain the predicted value of bandwidth utilization rate of each link within the future time window.

[0058] The temperature prediction module is used to calculate the predicted temperature of each link within the future time window based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane.

[0059] The loss correction module is used to perform temperature correction on the conductor loss and dielectric loss of each link at the reference temperature based on the predicted temperature of each link and the link parameters, so as to obtain the predicted link loss of each link in the future time window.

[0060] The loss margin calculation module is used to calculate the allowable link loss for each link based on the data transmission rate, encoding method, equalization configuration and error requirements of each link, and to determine the loss margin of each link based on the difference between the allowable link loss and the predicted link loss.

[0061] The communication scheduling module is used to schedule communication tasks based on the loss margin of each link to obtain a communication scheduling scheme.

[0062] The execution acquisition module is used to execute communication tasks based on the communication scheduling scheme and to collect actual bandwidth utilization, actual temperature data and communication performance data after the execution of the communication tasks.

[0063] The feedback update module is used to update the bandwidth utilization prediction model and the temperature correction coefficient based on the actual bandwidth utilization, actual temperature data and communication performance data.

[0064] It is understandable that the aforementioned server backplane transmission link loss assessment and optimization system and method have the same beneficial effects, and will not be elaborated further here. Attached Figure Description

[0065] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0066] Figure 1 A flowchart illustrating a method for evaluating and optimizing server backplane transmission link loss according to an embodiment of the present invention;

[0067] Figure 2 This is a schematic diagram of communication scheduling adjustment provided in an embodiment of the present invention;

[0068] Figure 3 This is a functional block diagram of a server backplane transmission link loss assessment and optimization system provided in an embodiment of the present invention. Detailed Implementation

[0069] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While the accompanying drawings show exemplary embodiments of the present disclosure, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0070] See Figure 1 As shown, this application proposes a method for evaluating and optimizing server backplane transmission link loss, including:

[0071] S1: Collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and establish a bandwidth utilization prediction model.

[0072] S2: During the execution of the current distributed training task, obtain the current task communication topology information, as well as the current bandwidth utilization and current temperature data of each link. Input the current task communication topology information, current bandwidth utilization and current temperature data into the bandwidth utilization prediction model to obtain the predicted bandwidth utilization value of each link within the future time window.

[0073] S3: Calculate the predicted temperature of each link within the future time window based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane.

[0074] S4: Based on the predicted temperature and link parameters corresponding to each link, perform temperature correction on the pre-calculated conductor loss and dielectric loss of each link at the reference temperature to obtain the predicted link loss of each link in the future time window.

[0075] S5: Calculate the allowable link loss for each link based on the data transmission rate, encoding method, equalization configuration and error rate requirements of each link, and determine the loss margin of each link based on the difference between the allowable link loss and the predicted link loss.

[0076] S6: Based on the loss margin of each link, schedule the communication tasks to obtain a communication scheduling scheme;

[0077] S7: Execute communication tasks based on the communication scheduling scheme, and collect actual bandwidth utilization, actual temperature data and communication performance data after the execution of the communication tasks;

[0078] S8: Update the bandwidth utilization prediction model and the temperature correction coefficient based on the actual bandwidth utilization, actual temperature data and communication performance data.

[0079] Specifically, this embodiment applies to servers comprising multiple compute nodes, accelerator cards, switching chips, and server backplanes. Each link on the server backplane refers to a high-speed interconnect channel within the backplane that handles data transmission between compute nodes. Distributed training tasks can be data-parallel, model-parallel, pipeline-parallel, or hybrid-parallel training tasks. Historical communication data can be obtained from training framework communication records, network card port statistics, switching chip port statistics, or server operation logs. This data includes communication occurrence time, communication data volume, communication duration, sending compute node, receiving compute node, and communication type. Communication types include parameter synchronization, gradient convergence, activation value transmission, or model sharding exchange. Historical task topology data includes the communication connection relationships between compute nodes, the order of communication stages, and communication dependencies. Historical temperature data can be collected by temperature sensors, onboard monitoring chips, or server management controllers deployed on the server backplane. When a temperature sampling point corresponds to multiple adjacent links, weights can be allocated according to the reciprocal of the distance from the temperature sampling point to the centerline of each link, or the temperature sampling value can be converted into the temperature value of the corresponding link according to the thermal coupling coefficient obtained from server thermal testing. Link parameters include link length, rated link bandwidth, conductor material parameters, dielectric material parameters, connector insertion loss, signal operating frequency, rated link speed, the correspondence between backplane ports and computing nodes, and the relationship between adjacent links in the backplane. When establishing a bandwidth utilization prediction model, historical communication data is first organized according to the training iteration rounds and communication stages. The amount of communication data passing through the same backplane link within the same communication stage is accumulated. For bidirectional links, the amount of communication data can be counted separately for the sending and receiving directions, or it can be counted based on the total rated bandwidth of both directions. Then, the amount of communication data carried by the link in the corresponding communication stage is divided by the product of the communication duration and the rated bandwidth of the link to obtain the bandwidth utilization rate of the link in the corresponding communication stage. When the communication duration spans multiple sampling periods, the amount of communication data is allocated to the corresponding sampling period according to the actual transmission time of each sampling period. Subsequently, bandwidth utilization, historical temperature data, link parameters, and corresponding task topology data are aligned according to sampling time to form model training samples. When the sampling period of historical temperature data is longer than that of communication data, linear interpolation can be performed on the temperature values ​​between two adjacent temperature sampling points to ensure that the temperature data and bandwidth utilization are in the same time series. The bandwidth utilization prediction model can employ a time-series regression model, a recurrent neural network model, a graph neural network model, or a gradient boosting tree model. Its inputs include communication node pairs, communication data volume, communication stage sequence, current bandwidth utilization, current temperature data, and link parameters. Its output is the predicted bandwidth utilization value of each link within a future time window in a continuous sampling period.The future time window is determined based on the communication scheduling cycle and the temperature sampling cycle. For example, if the communication scheduling cycle is 100ms to 500ms and the temperature sampling cycle is 500ms to 1s, the future time window can be 1s to 10s and divided into multiple consecutive sampling cycles. The current task communication topology information includes the communication node pairs, communication data volume, communication stage sequence, and communication dependencies in the current distributed training task. The current bandwidth utilization is calculated from the actual transmitted data volume, actual communication duration, and link rated bandwidth within the current sampling cycle. The current temperature data is the temperature sampling value corresponding to each link within the current sampling cycle, or the link temperature calculated from multiple temperature sampling values ​​according to distance weight or thermal coupling coefficient. After obtaining the bandwidth utilization prediction value, the bandwidth utilization prediction value is first multiplied by the link rated bandwidth to obtain the predicted transmission load of each link within the future time window. Then, based on the heat generation power per unit time calibrated in the full-load operation test, the ratio between the predicted transmission load and the link rated bandwidth, and the basic heat generation of the link during idle operation, the link heat generation of each link within the corresponding sampling cycle is calculated. The thermal conductivity parameters of a server backplane include the thermal conductivity of the backplane material, interlayer thermal resistance, link spacing, heat dissipation boundary temperature, airflow cooling capacity, and heat capacity. These parameters can be obtained from server thermal design data or through calibration tests under no-load, rated load, and full-load conditions. For example, during calibration, the target link continuously transmits test data at its rated rate, recording the temperature changes of the target link and adjacent links. Based on the temperature rise curve, the thermal resistance of the target link itself, the thermal coupling coefficient of adjacent links, and the heat exchange coefficient at the heat dissipation boundary are calculated. When calculating the predicted temperature, the current temperature data is used as the initial temperature. The temperature rise caused by the heat generated by each link itself, the heat transferred from adjacent links to the link, and the heat released from the link to the heat dissipation boundary are sequentially included in the temperature update process of the same sampling period. The predicted temperature of each link within the future time window is obtained period by period. The reference temperature can be taken as the reference temperature for backplane design testing, such as 25℃ or 40℃, or it can be taken as the average of multiple temperature sampling values ​​of each link after the server has been running stably under no load for 30 minutes. The conductor loss and dielectric loss of each link at the reference temperature can be obtained by backplane channel simulation based on the link length, conductor material parameters, dielectric material parameters, signal operating frequency and connector parameters, or it can be obtained by decomposing the insertion loss at the reference temperature after testing with a vector network analyzer.When performing temperature correction, first calculate the temperature difference between the predicted temperature and the reference temperature. For conductor loss, multiply the temperature coefficient of resistance of the conductor material by the temperature difference to obtain the relative change in conductor resistance. Then, multiply this relative change in conductor resistance by the conductor loss at the reference temperature to obtain the conductor loss temperature correction. For dielectric loss, read or interpolate the dielectric constant and loss tangent at the predicted temperature from the dielectric material test data table. Calculate the ratios of the dielectric constant and loss tangent at the predicted temperature to the dielectric constant and loss tangent at the reference temperature, and then correct the dielectric loss at the reference temperature accordingly to obtain the dielectric loss temperature correction. The predicted link loss is the sum of the conductor loss at the reference temperature, the dielectric loss at the reference temperature, the conductor loss temperature correction, the dielectric loss temperature correction, and the connector insertion loss. When the connector insertion loss is already included in the link loss test results at the reference temperature, the connector insertion loss is not added again. The allowable link loss is calculated based on the current link rate level, coding method, equalization configuration, and error rate requirement. Specifically, the link rate level is first determined according to the data transmission rate, and the corresponding transmitter output amplitude, receiver minimum input amplitude, receiver minimum eye height, or receiver minimum signal-to-noise ratio are read from the server SerDes parameter table or link training parameter table. Then, the corresponding receiver margin is selected according to the error rate requirement. For example, a first receiver margin is used when the error rate requirement is no higher than 10^-12, and a second receiver margin greater than the first receiver margin is used when the error rate requirement is no higher than 10^-15. The first and second receiver reserve margins can be obtained from the bit error rate test curves of the same type of backplane at the reference temperature. Then, according to the coding method, the signal margin corresponding to the coding gain or coding overhead is read from the coding configuration table, and the compensation amount corresponding to the transmitter pre-emphasis, continuous-time linear equalization, and decision feedback equalization is read according to the equalization configuration. Finally, the available signal budget corresponding to the transmitter output amplitude is added to the signal margin corresponding to the coding method and the compensation amount corresponding to the equalization configuration, and then the reception requirement corresponding to the minimum input amplitude of the receiver and the receiver reserve margin are subtracted to obtain the allowable link loss of the link under the current rate and configuration.The loss margin is the difference between the allowed link loss and the predicted link loss. When the loss margin is greater than zero, it is determined that the predicted link loss of the corresponding link in the future time window does not exceed the allowed link loss, and the link can continue to carry the original communication task. Only if the loss margin remains greater than zero after adding communication tasks to be migrated will it be considered a candidate link. When the loss margin is equal to zero, it is determined that the predicted link loss of the corresponding link is equal to the allowed link loss, and the link maintains the original communication task but is not assigned any new communication tasks. When the loss margin is less than zero, it is determined that the predicted link loss of the corresponding link exceeds the allowed link loss, and the link is identified as a link to be adjusted. When the loss margin of multiple links is less than zero, the scheduling order is determined according to the order of loss margin from low to high. During communication scheduling, the communication tasks carried by the links to be adjusted are first determined. Then, adjustable communication tasks are selected based on the communication stage sequence and computational dependencies of the communication tasks. For communication tasks that can change communication paths, the predicted link loss and loss margin after adjustment are calculated one by one among the candidate links connected to the sending and receiving nodes. The candidate link with the largest adjusted loss margin that is greater than zero is selected. If there is no candidate link with a loss margin greater than zero, path migration is not performed, and the process proceeds to communication topology adjustment or data fragmentation adjustment. When performing communication topology adjustment, the amount of communication data carried by each link under each topology in the current set of communication stages is calculated. The predicted link loss and loss margin are calculated for each topology. The communication topology that makes the loss margin of the link to be adjusted greater than zero is selected. When multiple communication topologies meet this condition, the communication topology with the largest adjusted loss margin value is selected. When no communication topology can make the loss margin of the link to be adjusted greater than zero, the communication topology that reduces the predicted link loss of the link to be adjusted by the largest amount is selected, and the loss margin is recalculated in the next sampling period. When adjusting data fragmentation, the amount of data fragments allocated to the corresponding communication node pairs of the link to be adjusted is reduced, and the reduced data fragments are allocated to other communication node pairs with a loss margin greater than zero after adjustment. When multiple communication node pairs meet this condition, the reduced data fragments are allocated in descending order of the loss margin after adjustment. When the loss margin of all communication node pairs is not greater than zero, the fragmentation method that maximizes the reduction in predicted link loss of the link to be adjusted is selected, and the loss margin is recalculated in the next sampling period. The communication scheduling scheme includes one or more of the following: communication path, communication topology, and data fragmentation method. After the communication task is executed, the actual bandwidth utilization is calculated based on the actual amount of data transmitted, the actual communication duration, and the rated bandwidth of the link. The actual temperature data is collected by a temperature sampling device. Communication performance data includes bit error rate, number of retransmissions, number of link training failures, communication latency, and throughput.The actual link loss can be calculated using the link training register, SerDes receiver eye diagram detection results, transmitter output configuration, receiver equalization convergence parameters, and bit error statistics. Specifically, the actual link loss after the communication task is executed is deduced from the actual output amplitude of the transmitter, the actual eye height or actual input amplitude of the receiver, the current equalization compensation amount, and the signal margin corresponding to the coding method. Finally, the actual bandwidth utilization is compared with the predicted bandwidth utilization value within the corresponding sampling period to obtain the bandwidth prediction deviation. When the bandwidth prediction deviation deviates in the same direction for multiple consecutive sampling periods, the model parameters of the bandwidth utilization prediction model are updated. When the bandwidth prediction deviation does not deviate in the same direction continuously, the data of that sampling period is added to subsequent training samples. The actual link loss is compared with the predicted link loss to obtain the loss deviation. When the actual link loss corresponding to the actual temperature data is higher than the predicted link loss at the same temperature, the temperature correction factor for conductor loss or dielectric loss of the corresponding link is increased. When the actual link loss corresponding to the actual temperature data is lower than the predicted link loss at the same temperature, the temperature correction factor for conductor loss or dielectric loss of the corresponding link is decreased. When the actual link loss is consistent with the predicted link loss, the corresponding temperature correction factor remains unchanged.

[0080] See Figure 2 As shown, Figure 2 This is a schematic diagram of communication scheduling adjustment provided in an embodiment of the present invention. Figure 2 This illustrates the process of path adjustment for communication tasks driven by link loss margin, where... Figure 2 (a) in the diagram is the schematic diagram before adjustment. Communication task T1 is sent from sending node Node A to receiving node Node D. The original communication path is transmitted through link L1. According to the link loss assessment results, the loss margin of link L1 is less than zero, indicating that the predicted link loss of link L1 in the corresponding sampling period exceeds the allowable link loss. Therefore, link L1 is determined as the link to be adjusted. Figure 2 (b) in the diagram shows the adjusted path. After screening the candidate links connected to the sending and receiving nodes of communication task T1, it is determined that the loss margin of link L2 is greater than zero, which meets the migration condition. Therefore, communication task T1 is adjusted from the original carrying link L1 to the candidate link L2, and the corresponding path adjustment result and communication scheduling scheme are obtained. Figure 2 The right side also shows the rules for determining the loss margin, which is the difference between the allowed link loss and the predicted link loss. When the loss margin is less than zero, the corresponding link is identified as a link to be adjusted; when the loss margin is equal to zero, the original communication task is maintained and no new communication task is assigned; when the loss margin is greater than zero, the corresponding link can participate in communication scheduling as a candidate link.

[0081] In some embodiments of this application, establishing a bandwidth utilization prediction model includes:

[0082] Historical communication data is divided according to training iteration rounds and communication stages, and the amount of communication data, communication duration, participating computing nodes and communication type of each communication stage are extracted;

[0083] Based on historical task topology data, determine the communication connection relationships between participating computing nodes and map these relationships to the corresponding links in the server backplane.

[0084] Calculate the historical bandwidth utilization rate of each link based on the amount of communication data carried, the duration of communication, and the rated bandwidth of the link during the corresponding communication phase.

[0085] By aligning the historical bandwidth utilization, historical temperature data, link parameters, and task topology data of the corresponding communication stages of each link by time, a bandwidth utilization prediction model is trained.

[0086] Specifically, historical communication data is collected according to the iteration rounds of the training task. Each training round can be further divided into forward propagation communication phase, backpropagation communication phase, gradient synchronization phase, parameter update phase, or model fragmentation and exchange phase. For different training frameworks, they can also be divided according to the names of communication operators recorded in the communication log, such as AllReduce, AllGather, ReduceScatter, or point-to-point transmission. The amount of communication data in each communication phase can be obtained by counting the number of bytes sent and received, and the communication duration can be calculated based on the communication start timestamp and communication end timestamp. The participating computing nodes are the computing nodes that send or receive data in that communication phase, and the communication type is determined according to the data interaction method corresponding to that communication phase. Historical task topology data includes the logical communication relationships between computing nodes and the physical connection relationships within the server. The logical communication relationships record the sending computing node, receiving computing node, and communication direction, while the physical connection relationships record the correspondence between computing nodes, backplane ports, switching chip ports, and backplane links. When mapping communication connections to corresponding links in the server backplane, the corresponding backplane ports are first determined based on the sending and receiving computing nodes. Then, the backplane links through which the communication data passes are determined based on the backplane port connection table. When a communication node pair corresponds to one backplane link, the communication data volume of that communication node pair is included in that backplane link. When a communication node pair passes through multiple backplane links, the communication data volume of that communication node pair is included in each backplane link on the path. When a communication node pair has multiple parallel links, the actual carrying link is determined according to the server communication scheduling record. If the communication scheduling record is missing, the communication data volume can be allocated according to the port counter increment of each parallel link in the corresponding communication phase. When calculating historical bandwidth utilization, the communication data volume carried by the corresponding link in a communication phase is divided by the product of the communication duration of that communication phase and the rated bandwidth of the link. If the communication phase spans multiple sampling periods, the communication data volume is allocated to the corresponding sampling period based on the duration of the communication phase in each sampling period, and the bandwidth utilization is calculated separately for each period. Historical temperature data and historical bandwidth utilization are aligned by sampling time. When the temperature sampling period differs from the communication sampling period, the communication sampling period is used as the reference, and linear interpolation is performed between two adjacent temperature samples. Alternatively, the temperature sample closest to the current communication sampling time is used as the temperature data for that sampling period. Link parameters include link length, rated link bandwidth, backplane layer of the link, adjacent link numbers, conductor material parameters, dielectric material parameters, and signal operating frequency. Training samples can consist of historical bandwidth utilization, historical temperature data, link parameters, communication phase type, and task topology data from multiple consecutive sampling periods. The sample label is the actual bandwidth utilization of the corresponding link in one or more subsequent sampling periods.Bandwidth utilization prediction models can be trained using time-series regression models, recurrent neural network models, graph neural network models, or gradient boosting tree models. When emphasizing the load patterns of links changing over time, time-series regression models or recurrent neural network models can be used; when emphasizing the topology of computing nodes and the connection relationships of backplane links, graph neural network models can be used. During training, the difference between the predicted bandwidth utilization and the actual bandwidth utilization is used as the basis for updating model parameters, enabling the bandwidth utilization prediction model to output predicted bandwidth utilization values ​​for each link within a future time window based on changes in task topology, communication stages, and historical link operating states.

[0087] In some embodiments of this application, obtaining the predicted bandwidth utilization of each link within a future time window includes:

[0088] Analyze the current task communication topology information to obtain the communication node pairs, communication data volume, and communication phase order in the current distributed training task;

[0089] Based on the correspondence between server backplane ports and compute nodes, the communication node pairs are mapped to the corresponding backplane links;

[0090] Input the current bandwidth utilization rate, current temperature data and corresponding communication stage data of each link into the bandwidth utilization rate prediction model to obtain the predicted bandwidth utilization rate of each link in the continuous sampling period within the future time window.

[0091] When the same link carries multiple communication tasks within the same sampling period, the bandwidth utilization prediction values ​​corresponding to the multiple communication tasks are merged.

[0092] Specifically, the current task communication topology information can be provided by the training framework's task scheduling records, communication operator execution plans, cluster scheduler records, or the server communication management module. This includes communication node pairs between computing nodes in the current distributed training task, the amount of communication data corresponding to each communication node pair, the communication direction, the order of communication stages, and communication dependencies. A communication node pair refers to a sending computing node and a receiving computing node that transmit data within the same communication stage. For example, node A sends gradient fragments to node B, or node C and node D perform parameter synchronization. The order of communication stages refers to the sequential execution relationship of each communication stage within the same training iteration. For example, activation value transmission is performed first, followed by gradient backpropagation, and then gradient synchronization. After parsing the current task communication topology information, the corresponding sending and receiving ports for each communication node pair are determined based on the correspondence between computing nodes and backplane ports. Then, the actual backplane links traversed by the communication node pair are determined according to the server backplane port connection table or the switching chip forwarding table. When a communication node pair traverses multiple backplane links, the communication data volume corresponding to that communication node pair is mapped to each backplane link on the path. When multiple available paths exist, the actual carrying link is determined first according to the current communication scheduling record. If the current communication scheduling record has not yet been generated, it is determined according to the default routing table or the actual carrying link of the previous sampling period. The current bandwidth utilization rate can be calculated based on the actual data transmission volume, actual communication duration, and rated bandwidth of each link in the current sampling period. The current temperature data can be directly obtained from the temperature sensor sampling values ​​corresponding to each link, or it can be calculated based on the distance weight or thermal coupling coefficient between adjacent temperature sampling points and the link. The corresponding communication stage data includes the communication stage type, communication stage sequence, communication data volume, number of communication node pairs, and communication dependencies. After inputting the above data into the bandwidth utilization prediction model, the model outputs the bandwidth utilization prediction value for each link in each sampling period according to the continuous sampling period within the future time window. For example, when the future time window is 5s and the sampling period is 500ms, the model outputs the bandwidth utilization prediction value for each link in 10 consecutive sampling periods. For the case where the same link carries multiple communication tasks in the same sampling period, the bandwidth utilization prediction value for each communication task on the link is calculated separately first, and then the multiple bandwidth utilization prediction values ​​are added together to obtain the combined bandwidth utilization prediction value for the link in the sampling period. When the combined bandwidth utilization prediction value is greater than the utilization limit corresponding to the available bandwidth of the link, the combined bandwidth utilization prediction value is corrected to the utilization limit, and the excess part is recorded as the queued communication load in the sampling period. The utilization limit can be calculated based on the link's rated bandwidth, protocol overhead, and reserved management bandwidth. For example, the available bandwidth is obtained by deducting the protocol overhead and management reserved bandwidth from the link's rated bandwidth, and then the utilization limit is obtained by dividing the available bandwidth by the link's rated bandwidth.

[0093] In some embodiments of this application, calculating the predicted temperature corresponding to each link within a future time window includes:

[0094] Based on the predicted bandwidth utilization, rated bandwidth, and parameters of each link, calculate the heat generation of each link within the future time window.

[0095] Based on the thermal conduction parameters of the server backplane, determine the thermal coupling relationship between each link and the heat exchange relationship between each link and the heat dissipation boundary;

[0096] Using the current temperature data of each link as the initial temperature, the predicted temperature of each link is calculated periodically based on the link's heat generation, thermal coupling relationship, and heat exchange relationship.

[0097] Specifically, the predicted bandwidth utilization rate of each link is the proportion of the link load in the corresponding sampling period within a future time window. First, this predicted bandwidth utilization rate is multiplied by the link's rated bandwidth to obtain the predicted transmission load of the link in the corresponding sampling period. Then, the heat generation power per unit time of the link under different transmission loads is determined based on the link parameters. These parameters include at least the link length, conductor material parameters, dielectric material parameters, signal operating frequency, number of connectors, and the backplane layer on which the link is located. The heat generation power per unit time can be obtained through calibration. During calibration, links of the same type are operated under no-load, half-rated bandwidth, and rated bandwidth conditions, respectively. The stable temperature rise under each operating condition is recorded, and the base heat generation power and the incremental heat generation power varying with the transmission load are calculated using the server backplane's thermal capacity parameters. When calculating the heat generation of the link in a certain sampling period, the heat generation power per unit time corresponding to that sampling period is multiplied by the sampling period duration to obtain the heat generated by the link in that sampling period. The thermal conductivity parameters of a server backplane include the thermal conductivity of the backplane material, interlayer thermal resistance, link spacing, backplane stack-up structure, thermal coupling coefficient between adjacent links, thermal resistance from the link to the heat dissipation boundary, airflow heat transfer coefficient, and heat dissipation boundary temperature. The heat dissipation boundary can be the backplane edge, connector mounting area, airflow heat transfer area, or location thermally connected to the chassis. The thermal coupling relationship between links can be initially determined using backplane structure data and then corrected through thermal testing. For example, the target link can be continuously transmitting test data at its rated bandwidth while other links remain idle. Temperature changes of the target link and adjacent links can be recorded. When the temperature of adjacent links increases with the temperature of the target link, the thermal coupling coefficient between them can be determined based on the ratio of the temperature rise of the adjacent link to that of the target link. The heat exchange relationship between each link and the heat dissipation boundary can be determined using the temperature difference between the link temperature and the heat dissipation boundary temperature, the airflow heat transfer coefficient, and the thermal resistance from the link to the heat dissipation boundary. For example, under test conditions with a fixed server fan speed, the cooling curve of the link after stopping high-load transmission can be recorded, and the coefficient of heat released from the link to the heat dissipation boundary can be determined based on the cooling curve. When calculating the predicted temperature, the current temperature data of each link is used as the initial temperature for the first sampling period. If a link does not have a directly corresponding temperature sensor, the current temperature of the link is calculated based on the distance weights from adjacent temperature sampling points to the link, or based on the calibrated thermal coupling coefficient. Subsequently, calculations are performed sequentially according to the sampling period: first, the link's own temperature rise is calculated based on the heat generated by the link in the current sampling period; then, based on the temperature difference between the link and the link and the thermal coupling coefficient, the heat transferred from adjacent links to the link or from the link to adjacent links is calculated; then, based on the temperature difference between the link temperature and the heat dissipation boundary temperature and the heat exchange relationship, the heat released by the link to the heat dissipation boundary is calculated; finally, the temperature rise corresponding to its own heat generation, the temperature change corresponding to heat transfer from adjacent links, and the temperature change corresponding to heat exchange at the heat dissipation boundary are combined to obtain the predicted temperature of the link at the end of the sampling period.The predicted temperature at the end of the current sampling period is used as the initial temperature for the next sampling period. This calculation is repeated until the temperature prediction for all sampling periods within the future time window is completed. For cases where the same link carries multiple communication tasks within the same sampling period, the predicted bandwidth utilization of the link is first merged, and then the link's heat generation is calculated. For cases where multiple adjacent links are simultaneously at high bandwidth utilization, the heat generation of each link is calculated separately, and the mutual heat transfer between adjacent links is calculated through thermal coupling relationships, thereby obtaining the predicted temperature of each link within the future time window.

[0098] In some embodiments of this application, obtaining the predicted link loss includes:

[0099] Based on the link length, conductor material parameters, dielectric material parameters, and signal operating frequency of each link, calculate the conductor loss and dielectric loss of each link at the reference temperature.

[0100] The conductor loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of resistance of the conductor material.

[0101] The dielectric loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of the dielectric material.

[0102] The predicted link loss is obtained by adding the conductor loss, dielectric loss, conductor loss temperature correction, and dielectric loss temperature correction at the reference temperature.

[0103] Specifically, the reference temperature is used to calculate the initial link loss. It can be the test temperature used during server backplane design verification, such as 25°C or 40°C, or the average temperature obtained by continuously collecting multiple temperature values ​​under stable no-load operation of the server. The conductor loss and dielectric loss of each link at the reference temperature can be obtained through channel simulation during the backplane design phase, or through testing with a vector network analyzer during the prototype testing phase. Conductor loss is related to link length, conductor material resistivity, conductor cross-sectional structure, conductor surface roughness, and signal operating frequency, while dielectric loss is related to link length, dielectric constant of the dielectric material, dielectric loss tangent, and signal operating frequency. In the specific calculation, first determine the link length, backplane layer, conductor material type, dielectric material type, and signal operating frequency for each link. Then, obtain the reference values ​​for conductor loss and dielectric loss for that link at the reference temperature. Both can be recorded in decibels. After obtaining the predicted temperature of each link within the future time window, the temperature difference between the predicted temperature and the reference temperature is calculated for each sampling period of each link. When the predicted temperature is higher than the reference temperature, the temperature difference is positive; when the predicted temperature is equal to the reference temperature, the temperature difference is zero; and when the predicted temperature is lower than the reference temperature, the temperature difference is negative. For the conductor loss temperature correction, the resistance change ratio corresponding to a unit temperature change is first determined based on the temperature coefficient of resistance of the conductor material. This resistance change ratio is then multiplied by the temperature difference to obtain the relative change in conductor resistance at the current predicted temperature. Subsequently, the conductor loss reference value is corrected based on the relative change in conductor resistance to obtain the conductor loss temperature correction. For example, for copper conductors, the temperature coefficient of resistance corresponding to copper material can be used. If the predicted temperature is higher than the reference temperature, the conductor loss temperature correction is positive; if the predicted temperature is lower than the reference temperature, the conductor loss temperature correction is negative. For the dielectric loss temperature correction, the changes in dielectric constant and loss tangent at the predicted temperature are first determined based on the temperature coefficient of the dielectric parameter of the dielectric material. The temperature coefficient of the dielectric parameter can be obtained from the temperature characteristic table provided by the dielectric material supplier, or it can be obtained by fitting the dielectric performance of the same dielectric material after testing at different temperatures. Then, the dielectric constant and loss tangent at the predicted temperature are compared with those at the reference temperature to determine the change in dielectric loss relative to the reference temperature, and the dielectric loss temperature correction is calculated accordingly. If the loss tangent corresponding to the predicted temperature is higher than that corresponding to the reference temperature, the dielectric loss temperature correction is positive; if they are the same, the dielectric loss temperature correction is zero; if the loss tangent corresponding to the predicted temperature is lower than that corresponding to the reference temperature, the dielectric loss temperature correction is negative.Finally, the conductor loss baseline value, dielectric loss baseline value, conductor loss temperature correction, and dielectric loss temperature correction for the same link in the same sampling period are added together to obtain the predicted link loss for that link in that sampling period. After performing the above calculations for each sampling period in the future time window, the predicted link loss sequence for that link in the future time window is obtained.

[0104] In some embodiments of this application, determining the loss margin for each link includes:

[0105] Based on the data transmission rate, encoding method, and error rate of each link, determine the signal quality requirements of the receiving end of the corresponding link;

[0106] Based on the balanced configuration of each link, determine the balanced compensation amount for the corresponding link;

[0107] Calculate the allowable link loss for each link based on the output signal quality of the transmitter, the signal quality requirements of the receiver, and the equalization compensation amount.

[0108] The link loss margin is obtained by subtracting the predicted link loss from the allowed link loss of the same link.

[0109] Specifically, when determining the loss margin, first, for each link, read its current data transmission rate, encoding method, load balancing configuration, and error rate requirement. The data transmission rate can be the current operating speed of the link, such as 25Gbps, 56Gbps, or 112Gbps; the encoding method can be the current line coding or error correction coding method used by the link; the error rate requirement can be determined by the communication reliability requirements corresponding to the current distributed training task, such as a bit error rate not exceeding 10%. -12 Or 10 -15The receiver signal quality requirements are not arbitrarily set, but determined based on the link rate level, coding method, and error rate requirements from the server's SerDes parameter table, link protocol specification table, or error rate test curves of the same model backplane. Specifically, the minimum input amplitude, minimum eye height, or minimum signal-to-noise ratio of the receiver at that data transmission rate can be determined first. Then, the required receiver signal quality to achieve that error rate can be found in the error rate test curves according to the error rate requirements. When error correction coding is enabled on the link, the required receiver signal quality is adjusted according to the error correction capability corresponding to the coding method. For example, if error correction coding can relax the receiver eye height requirement under the same error rate requirement, the relaxed receiver eye height is used as the receiver signal quality requirement. When error correction coding is not enabled, the receiver signal quality requirement corresponding to the error rate requirement in the uncoded state is adopted. The quality of the transmitter output signal can be obtained from the SerDes transmitter configuration parameters, including the transmitter output amplitude, transmitter pre-emphasis level, or transmitter eye diagram amplitude. When the transmitter output signal quality is expressed in amplitude, the transmitter output amplitude and the minimum input amplitude at the receiver are converted into a decibel-based available loss budget. The equalization compensation amount is determined based on the current equalization configuration of the link, which includes at least one of transmitter pre-emphasis, continuous-time linear equalization, and decision feedback equalization. The compensation amount corresponding to each equalization configuration can be read from the link training parameter table, the SerDes register configuration table, or the link calibration test results. When multiple equalizations are enabled simultaneously, the compensation amounts of each equalization are added together, and the summed equalization compensation amount does not exceed the maximum equalization compensation amount verified in the calibration test for this link. When calculating the allowable link loss, the available loss budget corresponding to the transmitter's output signal quality is added to the equalization compensation amount, and then the reserve amount corresponding to the receiver's signal quality requirements is subtracted. This yields the allowable link loss under the current data transmission rate, coding method, equalization configuration, and error rate requirements. When the transmitter's output signal quality, receiver's signal quality requirements, and equalization compensation amount are all expressed in decibels, the addition and subtraction are performed directly using decibel values. Subsequently, the allowable link loss for the same link in the same future sampling period is subtracted from the predicted link loss to obtain the link's loss margin in that sampling period. If the loss margin is greater than zero, it indicates that the predicted link loss is lower than the allowable link loss, and the difference is the remaining loss margin that the link can still withstand. If the loss margin is equal to zero, it indicates that the predicted link loss is equal to the allowable link loss, and no new communication load will be allocated to this link. If the loss margin is less than zero, it indicates that the predicted link loss exceeds the allowable link loss, and this link will be considered for adjustment in subsequent communication scheduling. For cases where the future time window contains multiple sampling periods, calculate the loss margin for each sampling period; when the loss margin of the same link is less than zero in any sampling period, add the link to the set of links to be adjusted in the corresponding sampling period.

[0110] In some embodiments of this application, obtaining the communication scheduling scheme includes:

[0111] The links to be adjusted are determined based on the loss margin of each link, and the communication tasks carried by the links to be adjusted are determined.

[0112] Based on the communication phase sequence and computational dependencies of the communication tasks, select adjustable communication tasks;

[0113] Among the candidate links connected to the sending and receiving nodes of the adjustable communication task, select the candidate link with a loss margin greater than that of the link to be adjusted.

[0114] Adjustable communication tasks are moved from the link to be adjusted to the selected candidate link, resulting in a communication scheduling scheme that includes the communication path adjustment results.

[0115] Specifically, when obtaining the communication scheduling scheme, the links to be adjusted are first determined based on the loss margin of each link within the future time window. Specifically, if the loss margin of a link is less than zero in any sampling period, that link is identified as a link to be adjusted within that sampling period; if the loss margins of multiple links are all less than zero, they are processed sequentially in ascending order of loss margin; if the loss margin of a link is equal to zero, the communication tasks already carried by that link are maintained, and no new or migrated communication tasks are assigned to that link; if the loss margin of a link is greater than zero, that link can participate in subsequent path selection as a candidate link, but its loss margin still needs to be recalculated after a communication task is migrated. After determining the links to be adjusted, the communication tasks carried by the link to be adjusted within the corresponding sampling period are determined based on the communication scheduling records, link port counters, and task communication topology information. The communication tasks include the sending node, receiving node, communication data volume, communication stage, and communication start time. Subsequently, adjustable communication tasks are selected based on the communication phase sequence and computational dependencies. Communication tasks that have already started transmission but would cause training iteration failure upon interruption are not considered adjustable. Communication tasks that have not yet started transmission or whose transmission paths can be changed without altering the sending node, receiving node, or training results are considered adjustable. For adjustable communication tasks, candidate links connecting the sending and receiving nodes are first determined. These candidate links can be derived from the server backplane port connection table, switch chip forwarding table, or communication routing table. Candidate links that have been disabled within the target sampling period, whose training has failed, or whose loss margin is zero are not considered migration targets. For the remaining candidate links, the bandwidth utilization, predicted temperature, predicted link loss, and loss margin after migrating the adjustable communication task to that candidate link are calculated. If a candidate link has a loss margin greater than zero after migration, the candidate link with the largest loss margin after migration is selected, and the adjustable communication task is moved from the link to be adjusted to the selected candidate link. If multiple candidate links have the same loss margin after migration, the candidate link with the lowest communication latency is selected. If the communication latency is still the same, the candidate link with the lowest current bandwidth utilization is selected. If the loss margin of all candidate links is not greater than zero after migration, no communication path adjustment is performed, and the correspondence between the adjustable communication task and the link to be adjusted is retained, so as to enter the communication topology adjustment or data fragmentation adjustment process. The resulting communication scheduling scheme includes at least the task identifier of the adjusted communication task, the original bearer link, the adjusted bearer link, the sampling period at which the adjustment takes effect, and the corresponding communication stage.

[0116] In some embodiments of this application, when the communication path cannot be adjusted, the following applies:

[0117] Determine the collective communication phase in which the link to be adjusted will participate;

[0118] When there are alternative communication topologies available during the aggregated communication phase, select the communication topology that reduces the amount of communication data carried by the link to be adjusted, and update the communication connection relationships between the participating computing nodes.

[0119] When the data fragmentation method in the aggregated communication phase is adjustable, reduce the amount of data fragments allocated to the corresponding communication node pairs of the links to be adjusted, and allocate the reduced amount of data fragments to the corresponding communication node pairs of links with larger loss margins;

[0120] Based on the updated communication connection relationships or the adjusted data fragmentation amount, a communication scheduling scheme is obtained.

[0121] Specifically, when communication path adjustment cannot meet scheduling requirements, the process of communication topology adjustment or data fragmentation adjustment is initiated. Communication paths cannot be adjusted in the following situations: there are no candidate links connecting the sending and receiving nodes of the adjustable communication task; the candidate link is disabled, training has failed, or it is locked by other tasks; or after migrating the adjustable communication task to a candidate link, the loss margin of the candidate link within the target sampling period is less than or equal to zero. In this case, the collective communication stage in which the link to be adjusted participates is first determined based on the task communication topology information and the communication stage sequence. The collective communication stage can be a communication stage requiring the participation of multiple computing nodes, such as AllReduce, AllGather, ReduceScatter, or Broadcast. After determining the collective communication stage, the optional communication topologies supported by that stage are read, such as ring topology, tree topology, hierarchical topology, or switch chip aggregation topology. The communication data volume carried by each backplane link under each optional communication topology, the predicted bandwidth utilization, the predicted temperature, the predicted link loss, and the loss margin are calculated respectively. If a communication topology exists that can change the loss margin of the link to be adjusted from less than zero to greater than zero, that topology is selected. If multiple communication topologies meet this condition, the topology with the highest adjusted loss margin value is selected. If no communication topology exists that makes the loss margin of the link to be adjusted greater than zero, the topology that reduces the predicted link loss of the link to be adjusted by the largest margin is selected, and the loss margin of the link to be adjusted is recalculated in the next sampling period. After selecting the communication topology, the communication connection relationships between the participating nodes are updated according to the selected topology, so that some communication data that was originally transmitted through the link to be adjusted is now transmitted through the updated communication connection relationships. If the data fragmentation method during the aggregated communication phase can be adjusted, then the amount of data fragments is further allocated according to the loss margin of each communication node pair corresponding to the link. Specifically, the amount of data fragments allocated to the communication node pair corresponding to the link to be adjusted is reduced, and the reduced amount of data fragments is allocated to other communication node pairs with a loss margin greater than zero after adjustment. When the loss margin of multiple communication node pairs is greater than zero after receiving new data fragments, the reduced amount of data fragments is allocated in descending order of the loss margin after adjustment. When the loss margin of a certain communication node pair is equal to zero after receiving new data fragments, no more new data fragments are allocated to that communication node pair. When the loss margin of all communication node pairs that can receive data fragments is less than or equal to zero after adjustment, the fragmentation method that reduces the predicted link loss of the link to be adjusted by the greatest extent is selected, and this fragmentation method is used as the temporary scheduling result of the current sampling period. A communication scheduling scheme is generated based on the updated communication connection relationship or the adjusted data fragmentation amount. The communication scheduling scheme includes the communication stage identifier, the communication topology before adjustment, the communication topology after adjustment, the data fragmentation amount before adjustment, the data fragmentation amount after adjustment, the sampling period for adjustment to take effect, and the backplane links involved.

[0122] In some embodiments of this application, updating the bandwidth utilization prediction model and the temperature correction coefficient includes:

[0123] Calculate the actual bandwidth utilization of each link based on the actual amount of data transmitted after the communication task is executed, the actual communication duration, and the rated bandwidth of the link.

[0124] The bandwidth utilization prediction model is updated based on the deviation between the actual bandwidth utilization of each link and the corresponding predicted bandwidth utilization.

[0125] Based on actual temperature data and communication performance data, the actual link loss of each link within the corresponding sampling period is inferred.

[0126] Based on the deviation between the actual link loss and the corresponding predicted link loss, update the conductor loss temperature correction factor and the dielectric loss temperature correction factor.

[0127] Specifically, after a communication task is executed, the actual data transmission volume and actual communication duration of each link within the corresponding sampling period are read from the communication scheduling record, network card port counter, switch chip port counter, or server operation log. The actual data transmission volume is then divided by the product of the actual communication duration and the link's rated bandwidth to obtain the actual bandwidth utilization rate of each link. When the same link carries multiple communication tasks within the same sampling period, the actual data transmission volumes of the multiple communication tasks are combined to calculate the actual bandwidth utilization rate of that link. Subsequently, the actual bandwidth utilization rate of each link within the same sampling period is compared with the corresponding predicted bandwidth utilization rate to obtain the bandwidth utilization rate prediction deviation. When the actual bandwidth utilization rate is greater than the predicted bandwidth utilization rate, the bandwidth utilization rate prediction deviation is positive; when the actual bandwidth utilization rate is equal to the predicted bandwidth utilization rate, the bandwidth utilization rate prediction deviation is zero; and when the actual bandwidth utilization rate is less than the predicted bandwidth utilization rate, the bandwidth utilization rate prediction deviation is negative. When updating the bandwidth utilization prediction model based on the bandwidth utilization prediction deviation, the current task communication topology information, current temperature data, current bandwidth utilization, and actual bandwidth utilization corresponding to the current sampling period are used as new training samples. If the bandwidth utilization prediction deviation of the same link is in the same direction for three or more consecutive sampling periods, or if the absolute value of the bandwidth utilization prediction deviation exceeds the allowable error range, the model parameters of the bandwidth utilization prediction model are incrementally updated. When the bandwidth utilization prediction deviation is zero, or if the absolute value of the bandwidth utilization prediction deviation does not exceed the allowable error range, the data for that sampling period is stored in the sample cache, and the model parameters are not updated immediately. The allowable error range is determined based on the error distribution between the predicted bandwidth utilization and the actual bandwidth utilization in historical training samples; for example, the sum of the average absolute value of historical errors and twice the standard deviation is taken as the upper limit of the allowable error. Actual temperature data is collected by a temperature sensor, onboard monitoring chip, or server management controller after the communication task is executed. Communication performance data includes one or more of the following: bit error rate, number of retransmissions, number of link training failures, communication latency, throughput, receiver eye height, receiver eye width, receiver input amplitude, transmitter output amplitude, and equalization convergence parameters. To infer the actual link loss, first read the actual output amplitude of the transmitter, the actual input amplitude of the receiver, or the eye diagram detection result of the receiver. Then read the signal margin corresponding to the current equalization compensation amount and the coding method. Subsequently, based on the difference between the actual output amplitude of the transmitter and the actual input amplitude of the receiver, subtract the signal margin corresponding to the current equalization compensation amount and the coding method to obtain the actual signal attenuation generated by the link within the sampling period, and use the signal attenuation as the actual link loss.When a link provides a link training register or a SerDes diagnostic register, the actual link loss can also be read from the link loss calibration table based on the equalization level, receiver eye diagram margin, and bit error statistics recorded in the register. When multiple sources of actual link loss exist for the same link, the link loss value corresponding to the link training register or SerDes diagnostic register is used first. If the register data is missing, the link loss value obtained by back-calculating the actual output amplitude of the transmitter and the actual input amplitude of the receiver is used. The actual link loss is compared with the predicted link loss within the same link and the same sampling period to obtain the loss prediction deviation. When the actual link loss is greater than the predicted link loss, the loss prediction deviation is positive; when the actual link loss is equal to the predicted link loss, the loss prediction deviation is zero; and when the actual link loss is less than the predicted link loss, the loss prediction deviation is negative. When updating the conductor loss temperature correction coefficient and the dielectric loss temperature correction coefficient, the temperature difference between the actual temperature data and the reference temperature is first calculated. When the temperature difference is greater than zero, the temperature correction coefficient corresponding to the high-temperature range is updated; when the temperature difference is equal to zero, the temperature correction coefficient is not updated; and when the temperature difference is less than zero, the temperature correction coefficient corresponding to the low-temperature range is updated. For cases where the temperature difference is not zero, the adjustment amount of the temperature correction coefficient is determined based on the ratio of the loss prediction deviation to the temperature difference. When the loss prediction deviation is positive, the temperature correction coefficient for the corresponding temperature range is increased by the adjustment amount; when the loss prediction deviation is negative, the temperature correction coefficient for the corresponding temperature range is decreased by the adjustment amount; when the loss prediction deviation is zero, the corresponding temperature correction coefficient remains unchanged. If the bit error rate in the communication performance data is higher than the communication task requirements, the number of retransmissions is higher than the previous sampling period, or the eye diagram margin at the receiver is lower than the link training requirements, and the actual temperature is higher than the reference temperature, then the dielectric loss temperature correction coefficient is updated first. If the output amplitude and equalization configuration at the transmitter remain unchanged, but the input amplitude at the receiver is lower than the input amplitude at the receiver corresponding to the previous sampling period, and the actual temperature is higher than the reference temperature, then the conductor loss temperature correction coefficient is updated first. The updated bandwidth utilization prediction model and temperature correction coefficient are used for bandwidth prediction, temperature prediction, and link loss prediction in the next future time window.

[0128] In another preferred embodiment based on the above embodiments, see [reference] Figure 3 As shown, this embodiment provides a server backplane transmission link loss assessment and optimization system, including:

[0129] The model training module is used to collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and to establish a bandwidth utilization prediction model.

[0130] The bandwidth prediction module is used to obtain the current task communication topology information, the current bandwidth utilization rate and the current temperature data of each link during the execution of the current distributed training task. The current task communication topology information, current bandwidth utilization rate and current temperature data are input into the bandwidth utilization prediction model to obtain the predicted value of the bandwidth utilization rate of each link within the future time window.

[0131] The temperature prediction module is used to calculate the predicted temperature of each link within a future time window based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane.

[0132] The loss correction module is used to perform temperature correction on the conductor loss and dielectric loss of each link at the reference temperature based on the predicted temperature and link parameters corresponding to each link, so as to obtain the predicted link loss of each link in the future time window.

[0133] The loss margin calculation module is used to calculate the allowable link loss for each link based on the data transmission rate, encoding method, equalization configuration and error requirements of each link, and to determine the loss margin of each link based on the difference between the allowable link loss and the predicted link loss.

[0134] The communication scheduling module is used to schedule communication tasks based on the loss margin of each link to obtain a communication scheduling scheme.

[0135] The execution acquisition module is used to execute communication tasks based on the communication scheduling scheme, and to collect the actual bandwidth utilization, actual temperature data and communication performance data after the execution of the communication tasks;

[0136] The feedback update module is used to update the bandwidth utilization prediction model and the temperature correction coefficient based on the actual bandwidth utilization, actual temperature data, and communication performance data.

[0137] Understandably, by using the model training module to correlate and model historical communication data, historical task topology data, historical temperature data, and link parameters, the system can predict the bandwidth utilization of each link within a future time window before and after the execution of the current distributed training task, combining the task communication topology and the current state of the links, rather than relying solely on fixed link design parameters for loss assessment. The temperature prediction module combines the predicted bandwidth utilization with the server backplane's thermal conductivity parameters to obtain the predicted temperature of each link as the communication load changes, allowing the link loss assessment process to consider the thermal coupling effects between different links in the backplane. The loss correction module corrects the conductor and dielectric losses at the reference temperature based on the predicted temperature, resulting in a more accurate prediction of link losses closer to the operating state, reducing the bias caused by estimating link losses solely based on reference or uniform temperature conditions. The loss margin calculation module compares the predicted link losses with the allowable link losses, clarifying the loss margin of each link within the future time window and providing a quantitative basis for subsequent communication scheduling. The communication scheduling module adjusts communication tasks based on loss margins, reducing the communication load on links where predicted link loss exceeds allowable loss, or altering the data volume carried by links through adjustments to communication paths, topologies, and data fragmentation methods. The execution acquisition and feedback update modules collect actual bandwidth utilization, actual temperature data, and communication performance data after scheduling execution, updating the bandwidth utilization prediction model and temperature correction coefficient accordingly. This allows subsequent prediction processes to continuously incorporate actual server operating data for correction. Therefore, this implementation achieves closed-loop processing between link load prediction, link temperature prediction, link loss assessment, communication scheduling, and feedback updates during distributed training task execution, improving the matching between server backplane transmission link loss assessment and communication scheduling.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope defined by the present invention.

Claims

1. A method for evaluating and optimizing transmission link loss on a server backplane, characterized in that, include: Collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and establish a bandwidth utilization prediction model. During the execution of the current distributed training task, the communication topology information of the current task and the current bandwidth utilization and current temperature data of each link are obtained. The current communication topology information, current bandwidth utilization and current temperature data are then input into the bandwidth utilization prediction model to obtain the predicted bandwidth utilization value of each link within the future time window. Based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane, the predicted temperature of each link within the future time window is calculated. Based on the predicted temperature and the link parameters corresponding to each link, the conductor loss and dielectric loss of each link at the reference temperature are pre-calculated and temperature-corrected to obtain the predicted link loss of each link within the future time window. Based on the data transmission rate, encoding method, equalization configuration and error rate requirements of each link, calculate the allowable link loss for each link, and determine the loss margin for each link based on the difference between the allowable link loss and the predicted link loss. Based on the loss margin of each link, communication tasks are scheduled to obtain a communication scheduling scheme. The communication task is executed based on the communication scheduling scheme, and the actual bandwidth utilization, actual temperature data and communication performance data are collected after the communication task is executed. Based on the actual bandwidth utilization, actual temperature data, and communication performance data, update the bandwidth utilization prediction model and the temperature correction coefficient for temperature correction.

2. The server backplane transmission link loss assessment and optimization method according to claim 1, characterized in that, When establishing a bandwidth utilization prediction model, the following should be included: The historical communication data is divided according to the training iteration rounds and communication stages, and the communication data volume, communication duration, participating computing nodes and communication type of each communication stage are extracted. The communication connection relationships between participating computing nodes are determined based on the historical task topology data, and the communication connection relationships are mapped to the corresponding links in the server backplane; Calculate the historical bandwidth utilization rate of each link based on the amount of communication data carried, the duration of communication, and the rated bandwidth of the link during the corresponding communication phase. The bandwidth utilization prediction model is trained by aligning the historical bandwidth utilization, historical temperature data, link parameters, and task topology data of the corresponding communication stage for each link according to time.

3. The server backplane transmission link loss assessment and optimization method according to claim 2, characterized in that, When obtaining the predicted bandwidth utilization of each link within a future time window, the following is included: Analyze the current task communication topology information to obtain the communication node pairs, communication data volume, and communication phase order in the current distributed training task; Based on the correspondence between server backplane ports and computing nodes, the communication node pairs are mapped to corresponding backplane links; Input the current bandwidth utilization rate, current temperature data, and corresponding communication stage data of each link into the bandwidth utilization prediction model to obtain the predicted bandwidth utilization rate of each link in the continuous sampling period within the future time window. When the same link carries multiple communication tasks within the same sampling period, the bandwidth utilization prediction values ​​corresponding to the multiple communication tasks are merged.

4. The server backplane transmission link loss assessment and optimization method according to claim 3, characterized in that, Calculating the predicted temperature for each link within the future time window includes: Based on the predicted bandwidth utilization, rated bandwidth, and parameters of each link, calculate the heat generation of each link within the future time window. Based on the thermal conduction parameters of the server backplane, determine the thermal coupling relationship between each link and the heat exchange relationship between each link and the heat dissipation boundary; Using the current temperature data of each link as the initial temperature, the predicted temperature of each link is calculated periodically based on the heat generation, thermal coupling relationship, and heat exchange relationship of the link.

5. The server backplane transmission link loss assessment and optimization method according to claim 4, characterized in that, When obtaining the predicted link loss, the following are included: Based on the link length, conductor material parameters, dielectric material parameters, and signal operating frequency of each link, calculate the conductor loss and dielectric loss of each link at the reference temperature. The conductor loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of resistance of the conductor material. The dielectric loss temperature correction is calculated based on the temperature difference between the predicted temperature and the reference temperature of each link, and based on the temperature coefficient of the dielectric material. The predicted link loss is obtained by adding the conductor loss, dielectric loss, conductor loss temperature correction, and dielectric loss temperature correction at the reference temperature.

6. The server backplane transmission link loss assessment and optimization method according to claim 5, characterized in that, When determining the loss margin for each link, the following should be included: Based on the data transmission rate, encoding method, and error rate of each link, determine the signal quality requirements of the receiving end of the corresponding link; Based on the balanced configuration of each link, determine the balanced compensation amount for the corresponding link; Calculate the allowable link loss for each link based on the output signal quality of the transmitter, the signal quality requirements of the receiver, and the equalization compensation amount. The link loss margin is obtained by subtracting the predicted link loss from the allowed link loss of the same link.

7. The server backplane transmission link loss assessment and optimization method according to claim 6, characterized in that, When obtaining the communication scheduling scheme, it includes: The link to be adjusted is determined based on the loss margin of each link, and the communication task carried by the link to be adjusted is determined. Based on the communication phase sequence and computational dependencies of the communication tasks, select adjustable communication tasks; Among the candidate links connected to the sending node and receiving node of the adjustable communication task, the candidate link with a loss margin greater than that of the link to be adjusted is selected. The adjustable communication task is adjusted from the link to be adjusted to the selected candidate link to obtain a communication scheduling scheme that includes the communication path adjustment results.

8. The server backplane transmission link loss assessment and optimization method according to claim 7, characterized in that, When the communication path cannot be adjusted, including: Determine the collective communication phase in which the link to be adjusted participates; When there are optional communication topologies in the aggregated communication phase, select the communication topology that reduces the amount of communication data carried by the link to be adjusted, and update the communication connection relationship between the participating computing nodes. When the data fragmentation method in the aggregated communication phase is adjustable, the amount of data fragments allocated to the communication node pairs corresponding to the links to be adjusted is reduced, and the reduced amount of data fragments is allocated to the communication node pairs corresponding to the links with larger loss margins; The communication scheduling scheme is obtained based on the updated communication connection relationship or the adjusted data fragmentation amount.

9. The server backplane transmission link loss assessment and optimization method according to claim 8, characterized in that, When updating the bandwidth utilization prediction model and temperature correction coefficient, the following is included: Calculate the actual bandwidth utilization of each link based on the actual amount of data transmitted after the communication task is executed, the actual communication duration, and the rated bandwidth of the link. The bandwidth utilization prediction model is updated based on the deviation between the actual bandwidth utilization of each link and the corresponding predicted bandwidth utilization. Based on actual temperature data and communication performance data, the actual link loss of each link within the corresponding sampling period is inferred. Based on the deviation between the actual link loss and the corresponding predicted link loss, update the conductor loss temperature correction factor and the dielectric loss temperature correction factor.

10. A server backplane transmission link loss assessment and optimization system, used to implement the server backplane transmission link loss assessment and optimization method as described in any one of claims 1-9, characterized in that, include: The model training module is used to collect historical communication data, historical task topology data, historical temperature data of each link on the server backplane, and link parameters of each link when the server executes historical distributed training tasks, and to establish a bandwidth utilization prediction model. The bandwidth prediction module is used to acquire the current task communication topology information, the current bandwidth utilization rate and the current temperature data of each link during the execution of the current distributed training task, and input the current task communication topology information, current bandwidth utilization rate and current temperature data into the bandwidth utilization prediction model to obtain the predicted value of bandwidth utilization rate of each link within the future time window. The temperature prediction module is used to calculate the predicted temperature of each link within the future time window based on the predicted bandwidth utilization of each link, the current temperature data, and the heat conduction parameters of the server backplane. The loss correction module is used to perform temperature correction on the conductor loss and dielectric loss of each link at the reference temperature based on the predicted temperature of each link and the link parameters, so as to obtain the predicted link loss of each link in the future time window. The loss margin calculation module is used to calculate the allowable link loss for each link based on the data transmission rate, encoding method, equalization configuration and error requirements of each link, and to determine the loss margin of each link based on the difference between the allowable link loss and the predicted link loss. The communication scheduling module is used to schedule communication tasks based on the loss margin of each link to obtain a communication scheduling scheme. The execution acquisition module is used to execute communication tasks based on the communication scheduling scheme and to collect actual bandwidth utilization, actual temperature data and communication performance data after the execution of the communication tasks. The feedback update module is used to update the bandwidth utilization prediction model and the temperature correction coefficient based on the actual bandwidth utilization, actual temperature data and communication performance data.