GPU (Graphics Processing Unit) hot plug method, device and equipment and storage medium
By using distributed phase synchronization circuits, reconfigurable impedance matching networks, core-particle interconnection protocols, hardware-level prediagnosis systems and intelligent prediction models in GPU hot-swap technology, the problems of power supply phase mismatch, insufficient signal reflection suppression, poor heterogeneous integration compatibility and lack of pre-check mechanisms in the prior art are solved, and more efficient and safer GPU hot-swap operation is achieved, reducing the hardware damage rate.
Patent Information
- Application Number
- CN202510225326.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
The existing GPU hot-swap technology has problems such as power supply phase mismatch, insufficient signal reflection suppression, poor heterogeneous integration compatibility and lack of pre-checking mechanisms, resulting in high hardware damage rate.
The power phase difference between the motherboard and the GPU is detected by using the preset distributed phase synchronization circuit in the GPU slot power path, and a reverse compensation signal is generated to process the surge current peak; the signal reflection coefficient is optimized using the reconfigurable impedance matching network and the preset GPU impedance characteristic database; the protocol compatibility processing is performed based on the core interconnection protocol; the hardware-level prediagnostic system and intelligent prediction model are used for diagnosis and prediction; and finally, the four-stage plug-in and unplug operation is used for hot-swap operation with the GPS synchronization clock.
Improves the efficiency and security of GPU hot-swap operation, reduces hardware damage rate, and improves user experience.
Smart Images

Figure CN120104541A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a GPU hot-plug method, device, equipment and storage medium. Background Art
[0002] Currently, the existing GPU (Graphics Processing Unit) hot-swap technology has the following four problems:
[0003] First, power phase mismatch: At the moment of hot plugging, the surge current caused by the phase difference between the GPU and the motherboard power supply exceeds 150A / μs;
[0004] Second, insufficient signal reflection suppression: During hot plugging, high-speed interfaces may experience impedance mutation, which may lead to signal integrity degradation, such as eye closure greater than 30%.
[0005] Third, heterogeneous integration has poor compatibility: Under the chiplet architecture, GPU interconnection protocols and power supply specifications of multiple vendors are not unified, which makes it impossible for traditional slots to perform adaptive matching operations;
[0006] Fourth, the pre-check mechanism is missing: Due to the lack of a pre-plugging diagnostic system based on hardware parameters such as parasitic capacitance and contact resistance, the hardware damage rate is greater than 2%.
[0007] As can be seen from the above, how to improve the efficiency and safety of executing GPU hot-plug operations during the execution of GPU hot-plug operations is a problem that needs to be solved urgently. Summary of the invention
[0008] In view of this, the purpose of the present invention is to provide a GPU hot-plug method, device, equipment and storage medium, which can improve the efficiency and safety of GPU hot-plug operation during the execution of GPU hot-plug operation, and reduce the damage rate of hardware involved in the production process. The specific scheme is as follows:
[0009] In a first aspect, the present application provides a GPU hot-plug method, comprising:
[0010] Using a preset distributed phase synchronization circuit in a GPU slot power path to detect an initial power phase difference between a motherboard to be processed and a GPU to be processed, and using a reverse compensation signal generated based on the initial power phase difference to process an initial surge current peak value corresponding to the initial power phase difference, and determining a network signal to be optimized based on the obtained target surge current peak value and the initial network signal;
[0011] The initial signal reflection coefficient corresponding to the network signal to be optimized is adjusted by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database, and the network signal to be optimized is optimized by using the obtained target signal reflection coefficient to obtain the target network signal;
[0012] Based on the target network signal and using the chip interconnection protocol, the to-be-processed motherboard and the to-be-processed GPU are subjected to protocol compatibility processing to obtain a target GPU that is compatible with the to-be-processed motherboard protocol, and the target GPU is diagnosed using a preset hardware-level pre-diagnosis system, and after the obtained diagnostic result indicates that the diagnosis has passed, a preset intelligent prediction model is used to predict a temperature prediction result, a load prediction result, and an inrush current peak prediction result of the target GPU;
[0013] A hot plug operation is performed using a preset plug-in protocol and a GPS synchronized clock and based on the temperature prediction result, the load prediction result and the surge current peak prediction result.
[0014] Optionally, the method uses a preset distributed phase synchronization circuit in a GPU slot power path to detect an initial power phase difference between a motherboard to be processed and a GPU to be processed, and uses a reverse compensation signal generated based on the initial power phase difference to process an initial surge current peak value corresponding to the initial power phase difference, and determines a network signal to be optimized based on the obtained target surge current peak value and the initial network signal, including:
[0015] Deploy a preset distributed phase synchronization circuit into a GPU slot power path, and use an inner loop in a preset dual-loop feedback control mechanism in the preset distributed phase synchronization circuit to call a preset programmable logic device, so as to use a preset zero-crossing switching algorithm to detect in real time an initial power phase difference between a to-be-processed motherboard and a to-be-processed GPU;
[0016] Determine whether the value corresponding to the initial power supply phase difference is greater than a preset phase difference threshold value, and if the value corresponding to the initial power supply phase difference is greater than the preset phase difference threshold value, use the outer loop in the preset dual-loop feedback control mechanism to call a preset dynamic adjustment algorithm and generate a reverse compensation signal based on the initial power supply phase difference;
[0017] The reverse compensation signal is used to reduce the initial surge current peak value corresponding to the initial power supply phase difference to no more than a preset safety threshold value, to obtain a target surge current peak value, and the network signal to be optimized is determined based on the target surge current peak value.
[0018] Optionally, the adjusting the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and a preset GPU impedance characteristic database includes:
[0019] Generate a reconfigurable impedance matching network based on the PCIe interface, integrate the reconfigurable impedance matching network at the PCIe interface, and determine whether the current mode is a preset low power consumption mode;
[0020] If the current mode is the preset low power mode, a preset series inductance compensation mechanism is activated, and a programmable LC array in the reconfigurable impedance matching network at the PCIe interface is called, and the initial signal reflection coefficient is reduced based on a preset GPU impedance characteristic database to obtain a target signal reflection coefficient;
[0021] If the current mode is not the preset low power consumption mode, the current mode is set to the preset high bandwidth mode, and the target high frequency noise corresponding to the network signal to be optimized is suppressed to obtain the suppressed network signal, and then the programmable LC array is called and the initial signal reflection coefficient corresponding to the suppressed network signal is increased based on the preset GPU impedance characteristic database to obtain the target signal reflection coefficient.
[0022] Optionally, performing protocol compatibility processing on the motherboard to be processed and the GPU to be processed based on the target network signal and using a chiplet interconnection protocol to obtain a target GPU that is compatible with the protocol of the motherboard to be processed includes:
[0023] After receiving the target network signal, calling a protocol conversion engine corresponding to the core grain interconnection protocol, and determining a protocol conversion delay time between the to-be-processed mainboard and the to-be-processed GPU based on a preset hardware state machine implementation protocol, to obtain a protocol conversion delay time;
[0024] Wherein, the protocol conversion delay time is less than a preset delay time threshold; the chiplet interconnection protocol includes the preset hardware state machine implementation protocol and a dynamic function area; the dynamic function area includes a power function area and a signal function area;
[0025] Calling the power function area designed based on the preset enhanced contact design technology to increase the contact pressure corresponding to the GPU to be processed to no less than a preset pressure safety threshold, thereby obtaining a target contact pressure;
[0026] Calling the signal function area to process the initial differential pair spacing corresponding to the GPU to be processed to obtain a target differential pair spacing;
[0027] The to-be-processed motherboard and the to-be-processed GPU are subjected to protocol compatibility processing based on the protocol conversion delay time, the target contact pressure and the target differential pair spacing to obtain a target GPU that is protocol-compatible with the to-be-processed motherboard.
[0028] Optionally, the diagnosing the target GPU by using a preset hardware-level pre-diagnosis system includes:
[0029] Using a preset hardware-level pre-diagnosis system and based on a preset test signal, detecting a capacitance deviation corresponding to the target GPU to obtain a capacitance deviation diagnosis result;
[0030] Using a preset measurement method and a preset contact resistance scanning technology to identify abnormal contacts in the target GPU, and obtain abnormal contact diagnosis results;
[0031] The target GPU is verified based on a preset ESD protection verification technology to obtain a protection diagnosis result, and the target GPU is verified using a preset protocol handshake test technology to obtain a protocol handshake diagnosis result.
[0032] Optionally, after the obtained diagnostic result indicates that the diagnosis has passed, a preset intelligent prediction model is used to predict a temperature prediction result, a load prediction result, and a surge current peak prediction result of the target GPU, including:
[0033] sequentially determining whether the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result, and the protocol handshake diagnosis result all indicate that the diagnosis has passed;
[0034] If the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result, and the protocol handshake diagnosis result all indicate that the diagnosis is passed, the temperature and load of the GPU to be processed are predicted using a preset main model in a preset intelligent prediction model deployed in the AI chip to obtain a temperature prediction result and a load prediction result;
[0035] Using a preset auxiliary model in the preset intelligent prediction model to perform an inrush current peak value preset operation on the GPU to be processed, to obtain an inrush current peak value prediction result;
[0036] Determine whether the temperature prediction result, the load prediction result, and the surge current peak prediction result exceed the corresponding prediction result thresholds respectively; if so, trigger a model retraining operation of the preset intelligent prediction model.
[0037] Optionally, the hot plug operation is performed by using a preset plug-in protocol and a GPS synchronized clock and based on the temperature prediction result, the load prediction result, and the surge current peak prediction result, including:
[0038] Using the GPS synchronized clock in the preset hot-swap execution module, and based on the pre-discharge phase plug-in protocol, the physical unlock phase plug-in protocol, the signal isolation phase plug-in protocol, and the power disconnection phase plug-in protocol and the corresponding preset time thresholds, a corresponding plug-in timing matrix is established to implement the hot-swap operation based on the plug-in timing matrix;
[0039] The plug-in timing matrix includes a pre-discharge time, a physical unlocking time, a signal isolation time and a power disconnection time corresponding to the pre-discharge stage, the physical unlocking stage, the signal isolation stage and the power disconnection stage respectively.
[0040] In a second aspect, the present application provides a GPU hot-swap device, comprising:
[0041] A power phase difference determination module, configured to detect an initial power phase difference between a motherboard to be processed and a GPU to be processed by using a preset distributed phase synchronization circuit in a power path of a GPU slot, and to process an initial surge current peak value corresponding to the initial power phase difference by using a reverse compensation signal generated based on the initial power phase difference, and to determine a network signal to be optimized based on the obtained target surge current peak value and the initial network signal;
[0042] A network signal optimization module, used to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database, and optimize the network signal to be optimized by using the obtained target signal reflection coefficient to obtain a target network signal;
[0043] A GPU diagnostic module, configured to perform protocol compatibility processing on the to-be-processed motherboard and the to-be-processed GPU based on the target network signal and using a core particle interconnection protocol to obtain a target GPU that is compatible with the protocol of the to-be-processed motherboard, and to diagnose the target GPU using a preset hardware-level pre-diagnosis system, and after the obtained diagnostic result indicates that the diagnosis has passed, to predict the temperature prediction result, load prediction result, and surge current peak prediction result of the target GPU using a preset intelligent prediction model;
[0044] The operation time determination module is used to perform hot plug operation based on the temperature prediction result, the load prediction result and the surge current peak prediction result by using a preset plug-in protocol and a GPS synchronized clock.
[0045] In a third aspect, the present application provides an electronic device, including:
[0046] Memory, used to store computer programs;
[0047] The processor is used to execute the computer program to implement the aforementioned GPU hot plug method.
[0048] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned GPU hot-plug method when executed by a processor.
[0049] As can be seen from the above, before hot-swapping the GPU, the present application needs to use the preset distributed phase synchronization circuit in the GPU slot power path to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed, and use the reverse compensation signal generated based on the initial power phase difference to process the initial surge current peak corresponding to the initial power phase difference, and determine the network signal to be optimized based on the obtained target surge current peak and the initial network signal; use the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and use the obtained target signal reflection coefficient to optimize the network signal to be optimized to obtain the target network signal; based on The target network signal is used to perform protocol compatibility processing on the processing motherboard and the processing GPU by using the protocol conversion engine and the dynamic function area in the adaptive interconnection protocol center of the core grain architecture, so as to obtain a protocol-compatible motherboard and a protocol-compatible GPU, and the protocol-compatible GPU is diagnosed by using a preset hardware-level pre-diagnosis system, and after the obtained diagnosis result characterizes that the diagnosis has passed, the temperature prediction result, the load prediction result and the surge current peak prediction result of the protocol-compatible GPU are predicted by using a preset intelligent prediction model; the four-stage plug-in protocol and the GPS synchronous clock in the preset hot-swap execution module are used, and the operation time is determined based on the temperature prediction result, the load prediction result and the predicted surge current peak, and the hot-swap operation is realized by using the operation time.
[0050] It can be seen that the present application firstly needs to use the preset distributed phase synchronization circuit in the GPU slot power path to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed, and use the reverse compensation signal generated based on the initial power phase difference to process the initial surge current peak corresponding to the initial power phase difference, and determine the network signal to be optimized based on the obtained target surge current peak and the initial network signal; secondly, use the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and use the obtained target signal reflection coefficient to optimize the network signal to be optimized to obtain the target network signal; then, based on the target network The network signal is used to perform protocol compatibility processing on the processing motherboard and the processing GPU using the protocol conversion engine and dynamic function area in the core grain architecture adaptive interconnection protocol center, so as to obtain a protocol-compatible motherboard and a protocol-compatible GPU, and the protocol-compatible GPU is diagnosed using the preset hardware-level pre-diagnosis system. Furthermore, after the obtained diagnostic result characterizes that the diagnosis has passed, the temperature prediction result, load prediction result and surge current peak prediction result of the protocol-compatible GPU are predicted using the preset intelligent prediction model; finally, the four-stage plug-in protocol and GPS synchronization clock in the preset hot-swap execution module are used, and the operation time is determined based on the temperature prediction result, load prediction result and predicted surge current peak value, and the hot-swap operation is realized using the operation time. In this way, the efficiency and safety of executing GPU hot-swap operations are improved, the damage rate of the hardware involved in the production process is reduced, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0052] Figure 1 A flowchart of a GPU hot-swap method disclosed in this application;
[0053] Figure 2 A schematic diagram of a specific process of hot-swapping a GPU disclosed in the present application;
[0054] Figure 3 A schematic diagram of operation contents and time constraints corresponding to a specific preset plug-in and unplug-out protocol disclosed in this application;
[0055] Figure 4 A schematic diagram of the structure of a GPU hot-swap device disclosed in this application;
[0056] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] At present, the existing GPU hot-swap technology has the following four problems: power phase mismatch, insufficient signal reflection suppression, poor heterogeneous integration compatibility, and lack of pre-check mechanism. To this end, the present application provides a GPU hot-swap method that can improve the efficiency and safety of performing GPU hot-swap operations, thereby reducing the damage rate of hardware involved in the production process.
[0059] See also Figure 1 As shown, an embodiment of the present invention discloses a GPU hot plug method, comprising:
[0060] Step S11, using a preset distributed phase synchronization circuit in the GPU slot power path to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed, and using a reverse compensation signal generated based on the initial power phase difference to process the initial surge current peak corresponding to the initial power phase difference, and determining the network signal to be optimized based on the obtained target surge current peak and the initial network signal.
[0061] In this embodiment, the process of hot plugging the GPU will go through five stages, such as Figure 2 As shown: the above five stages are initialization and hardware self-check stage, real-time monitoring and signal processing stage, intelligent decision-making and hardware control stage, hot-swap execution stage, and state feedback and self-optimization.
[0062] In a specific implementation, before the real-time monitoring and signal processing stage, the embodiment of the present application may adopt a dual-channel 12V power input architecture, and configure an independent MOSFET (Metal-Oxide-SemiconductorField-Effect Transistor Switch) switch for each channel in the architecture, and the on-resistance is less than 1mΩ, and in addition, it supports a maximum peak current of 200A. Subsequently, the preset distributed phase synchronization circuit in the GPU slot power path is used to detect the first phase corresponding to the motherboard to be processed and the second phase corresponding to the GPU to be processed, and then the initial power phase difference is determined based on the first phase and the second phase.
[0063] Specifically, the preset distributed phase synchronization circuit in the GPU slot power path is used to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed, and the reverse compensation signal generated based on the initial power phase difference is used to process the initial surge current peak corresponding to the initial power phase difference, and the network signal to be optimized is determined based on the obtained target surge current peak and the initial network signal, which may include: deploying the preset distributed phase synchronization circuit to the GPU slot power path, and using the inner loop in the preset dual-loop feedback control mechanism in the preset distributed phase synchronization circuit to call the preset programmable logic device, so as to use the preset zero-crossing switching algorithm to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed in real time; judging whether the numerical value corresponding to the initial power phase difference is greater than the preset phase difference threshold, if the numerical value corresponding to the initial power phase difference is greater than the preset phase difference threshold, using the outer loop in the preset dual-loop feedback control mechanism to call the preset dynamic adjustment algorithm and generate a reverse compensation signal based on the initial power phase difference; using the reverse compensation signal to reduce the initial surge current peak corresponding to the initial power phase difference to no more than the preset safety threshold, to obtain the target surge current peak, and determining the network signal to be optimized based on the target surge current peak.
[0064] In this embodiment, the embodiment of the present application first needs to integrate a four-wire temperature measurement circuit in the GPU core area, and in a specific implementation, the sampling rate of the temperature measurement circuit is 1MHz, and the sampling accuracy range is -0.1℃~+0.1℃. Secondly, a Hall current sensor needs to be deployed in the power path, and the bandwidth is 1MHz. Subsequently, the inner loop in the preset dual-loop feedback control mechanism is adopted, and nanosecond phase tracking is achieved based on FPGA (Field-Programmable Gate Array), and the response time is less than 5ns, while the outer loop dynamically adjusts the compensation amount through the PID (Proportional-Integral-Derivative) algorithm to adapt to the power range of 10-1000W, so as to use the Hall current sensor and detect the power phase difference at the moment of plugging and unplugging based on the preset zero-crossing switching algorithm, and trigger the parallel switching operation when the power phase difference is not greater than 5°, so as to control the voltage fluctuation within the preset fluctuation range and reduce the surge current peak to less than 20A / μs.
[0065] Furthermore, the embodiment of the present application needs to implement pre-emphasis intensity determination and equalization compensation operations on high-speed signal lines, such as PCIe (Peripheral Component Interconnect Express Clock Line, i.e., peripheral component interconnect high-speed bus clock line) clock lines, and the calculation formula is as follows:
[0066] Pre-emphasis = 20×log(f_cutoff / f_signal);
[0067] Among them, Pre-emphasis is the pre-emphasis intensity, log is the logarithmic operation, f_cutoff is the channel frequency determined based on the PCB (Printed Circuit Board) material, trace length and number of vias, which is used to respond to the cutoff frequency of 3dB attenuation; f_signal is the main frequency component of the actual transmission signal.
[0068] Step S12: using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and using the obtained target signal reflection coefficient to optimize the network signal to obtain the target network signal.
[0069] In this embodiment, a reconfigurable impedance matching network including a programmable LC (Inductor-Capacitor Array) array and a preset GPU impedance feature database needs to be generated at the PCIe interface to call the reconfigurable impedance matching network to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, so as to use the adjusted signal reflection coefficient to perform an optimization operation on the network signal to be optimized. In the process of adjusting the initial signal reflection coefficient corresponding to the network signal to be optimized, the embodiment of the present application needs to perform a corresponding coefficient adjustment operation based on the current mode.
[0070] Specifically, adjusting the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance feature database can include: generating a reconfigurable impedance matching network based on a PCIe interface, integrating the reconfigurable impedance matching network at the PCIe interface, and judging whether the current mode is a preset low power mode; if the current mode is the preset low power mode, activating the preset series inductance compensation mechanism, calling the programmable LC array in the reconfigurable impedance matching network at the PCIe interface, and reducing the initial signal reflection coefficient based on the preset GPU impedance feature database to obtain a target signal reflection coefficient; if the current mode is not the preset low power mode, setting the current mode to the preset high bandwidth mode, suppressing the target high-frequency noise corresponding to the network signal to be optimized, obtaining a suppressed network signal, and then calling the programmable LC array and increasing the initial signal reflection coefficient corresponding to the suppressed network signal based on the preset GPU impedance feature database to obtain a target signal reflection coefficient.
[0071] In a specific implementation, the embodiment of the present application deploys a differential signal isolator in the PCIe interface circuit of the GPU slot, wherein the differential signal isolator uses an electromagnetic isolation chip with a common mode rejection ratio of not less than 80dB to reduce the crosstalk between the signal line and the power line to below -120dB (Decibel). In addition, after designing the reconfigurable impedance matching network, the embodiment of the present application needs to use the reconfigurable impedance matching network to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and the adjustment formula is as follows:
[0072] Z_match = (Z_GPU × Z_motherboard)^0.5;
[0073] Among them, Z_GPU is the input impedance of the target frequency band generated by the GPU board based on the gold finger contacts, package leads and chip pin parasitic parameters; Z_motherboard is the equivalent output impedance of the motherboard slot, which includes the transmission line characteristic impedance, termination matching network impedance and power plane impedance.
[0074] It is worth mentioning that when generating a reconfigurable impedance matching network at the PCIe interface, the embodiment of the present application needs to integrate a T-type programmable LC array, and the above-mentioned T-type programmable LC array supports impedance matching in the 0.1-10GHz frequency band. In addition, when the current mode is a low-power mode, that is, Z is not less than 50Ω, the embodiment of the present application needs to activate series inductance compensation and reduce the signal reflection coefficient to no more than -30dB; when the current mode is a high-bandwidth mode, that is, Z is not more than 25Ω, the embodiment of the present application needs to call a parallel capacitor group to absorb high-frequency noise to increase the eye opening to more than 80%. It is worth mentioning that the preset GPU impedance feature database pre-stores S parameter curves of GPUs from different manufacturers.
[0075] Step S13, based on the target network signal and using the chip interconnection protocol, protocol compatibility processing is performed on the to-be-processed motherboard and the to-be-processed GPU to obtain a target GPU that is compatible with the protocol of the to-be-processed motherboard, and the target GPU is diagnosed using a preset hardware-level pre-diagnosis system, and after the obtained diagnostic result indicates that the diagnosis has passed, a preset intelligent prediction model is used to predict the temperature prediction result, load prediction result and surge current peak prediction result of the target GPU.
[0076] In this embodiment, after obtaining the target network signal, the embodiment of the present application needs to determine the protocol corresponding to the motherboard to be processed, so as to perform protocol compatibility processing on the GPU to be processed based on the protocol corresponding to the motherboard to be processed and using the chip interconnection protocol, and obtain a target GPU that is compatible with the protocol of the motherboard to be processed. Specifically, based on the target network signal and using the core grain interconnection protocol, protocol compatibility processing is performed on the processing motherboard and the processing GPU to obtain a target GPU compatible with the protocol of the processing motherboard, which may include: after receiving the target network signal, calling the protocol conversion engine corresponding to the core grain interconnection protocol, and determining the protocol conversion delay time for the processing motherboard and the processing GPU based on the preset hardware state machine implementation protocol to obtain the protocol conversion delay time; wherein the protocol conversion delay time is less than the preset delay time threshold; the core grain interconnection protocol includes a preset hardware state machine implementation protocol and a dynamic function area; the dynamic function area includes a power function area and a signal function area; calling the power function area designed based on the preset enhanced contact design technology to increase the contact pressure corresponding to the processing GPU to not less than the preset pressure safety threshold to obtain the target contact pressure; calling the signal function area to process the initial differential pair spacing corresponding to the processing GPU to obtain the target differential pair spacing; performing protocol compatibility processing on the processing motherboard and the processing GPU based on the protocol conversion delay time, the target contact pressure and the target differential pair spacing to obtain the target GPU compatible with the protocol of the processing motherboard.
[0077] In a specific implementation, the protocol conversion engine is compatible with heterogeneous bus protocols such as OpenCAPI (Open Coherent Accelerator Processor Interface), CXL (Compute Express Link), and AXI (Advanced eXtensible Interface), and implements protocol conversion through a hardware state machine, and the protocol conversion delay time is no more than 10ns. In addition, the embodiment of the present application chooses to divide the dynamic functional area including the power functional area and the signal functional area in the slot gold finger area, wherein the power functional area is used to increase the contact pressure to 2.5N / point using a W-shaped double-bend contact, and the signal functional area is used to optimize the differential pair spacing to 0.3mm, and supports 56Gbps (Gigabits per second, a data transmission rate unit) PAM4 (Pulse Amplitude Modulation 4-level, four-level pulse amplitude modulation) signal transmission.
[0078] In this embodiment, after obtaining a target GPU that is compatible with the motherboard protocol to be processed, the embodiment of the present application needs to use a preset hardware-level pre-diagnosis system to diagnose the target GPU and obtain a diagnosis result, so as to perform subsequent operations after several diagnosis results indicate that the diagnosis has passed. Specifically, diagnosing the target GPU using a preset hardware-level pre-diagnosis system may include: using a preset hardware-level pre-diagnosis system and based on a preset test signal to detect the capacitance deviation corresponding to the target GPU, and obtaining a capacitance deviation diagnosis result; using a preset measurement method and a preset contact resistance scanning technology to identify abnormal contacts in the target GPU, and obtaining an abnormal contact diagnosis result; verifying the target GPU based on a preset ESD (Electrostatic Discharge) protection verification technology to obtain a protection diagnosis result, and verifying the target GPU using a preset protocol handshake test technology to obtain a protocol handshake diagnosis result.
[0079] In a specific implementation, when performing a parasitic capacitance calibration operation on a target GPU, an embodiment of the present application selects to inject a 1MHz test signal to detect the parasitic capacitance deviation between the slot and the GPU. If the parasitic capacitance deviation is no more than 5pF, the capacitance deviation diagnosis result obtained by characterization passes the diagnosis, and the resistance value of each contact is measured using a preset four-wire method to identify abnormal contacts. If the accuracy range is -0.1mΩ~+0.1mΩ, the abnormal contact diagnosis result obtained by characterization passes the diagnosis, and an 8kV contact discharge pulse is applied for ESD protection verification. If the response time is no more than 1ns, the protection diagnosis result obtained by characterization passes the diagnosis, and the CXL bus handshake protocol is simulated to verify the compatibility of the electrical and logic layers. If the electrical and logic layers are compatible, the protocol handshake diagnosis result obtained by characterization passes the diagnosis.
[0080] In this embodiment, the embodiment of the present application needs to use a preset intelligent prediction model to predict the temperature prediction results, load prediction results and surge current peak prediction results corresponding to the target GPU after the above-mentioned capacitor deviation diagnosis results, abnormal contact diagnosis results, protection diagnosis results and protocol handshake diagnosis results have all passed the diagnosis, so as to perform a hot plug operation based on the obtained temperature prediction results, load prediction results and surge current peak prediction results. Specifically, after the obtained diagnostic result indicates that the diagnosis has passed, the preset intelligent prediction model is used to predict the temperature prediction result, load prediction result and surge current peak prediction result of the target GPU, which may include: judging in turn whether the capacitor deviation diagnosis result, abnormal contact diagnosis result, protection diagnosis result and protocol handshake diagnosis result all indicate that the diagnosis has passed; if the capacitor deviation diagnosis result, abnormal contact diagnosis result, protection diagnosis result and protocol handshake diagnosis result all indicate that the diagnosis has passed, the preset main model in the preset intelligent prediction model deployed on the AI chip is used to predict the temperature and load of the GPU to be processed, and the temperature prediction result and the load prediction result are obtained; the preset auxiliary model in the preset intelligent prediction model is used to perform a surge current peak preset operation on the GPU to be processed, and the surge current peak prediction result is obtained; and it is judged whether the temperature prediction result, load prediction result and surge current peak prediction result exceed the corresponding prediction result thresholds, and if so, the model retraining operation of the preset intelligent prediction model is triggered.
[0081] In a specific implementation, the embodiment of the present application deploys the preset intelligent prediction model on a dedicated AI acceleration chip, such as TPU v4 (Tensor Processing Unit version 4, the fourth generation of tensor processing unit), and adopts 8-bit fixed-point quantization technology, so that the model inference delay time corresponding to the preset intelligent prediction model is no more than 5ms, and the calculation formula is as follows:
[0082] W_int8 = round(W_float32 × 127 / max(|W_float32|));
[0083] Among them, W_float32 is the neural network weight matrix of FP32 precision, which is stored in the chip SRAM. max(|W_float32|) is the maximum value of the absolute value of the weight in the neural network weight matrix, which is used to normalize the weight to the range of [-1,1]. W_int8 is the quantized INT8 weight, which implements 8-bit fixed-point operation through the hardware accelerator.
[0084] In addition, the embodiment of the present application adopts a preset multi-model collaborative prediction mechanism to run three prediction models in parallel: the main model (LSTM) (Long Short-Term Memory, i.e., long short-term memory network), the auxiliary model (ARIMA) (AutoRegressive Integrated Moving Average, i.e., autoregressive integrated moving average model) and the verification model (SVM) (Support Vector Machine, i.e., support vector machine). Among them, the main model (LSTM) is used to predict the GPU temperature and load within the next 30 seconds, the auxiliary model (ARIMA) is used to predict the peak value of the power path surge current, and the verification model (SVM) is used to cross-check the credibility of the above prediction results. When the deviation of the prediction result is not less than 15%, the operation of retraining the model is triggered.
[0085] Step S14: using a preset plug-in protocol and a GPS synchronized clock, and performing a hot-plug operation based on the temperature prediction result, the load prediction result, and the surge current peak prediction result.
[0086] In this embodiment, after obtaining the temperature prediction result, load prediction result and inrush current peak prediction result corresponding to the target GPU, the embodiment of the present application needs to use the preset plug-in protocol and GPS (Global Positioning System) to synchronize the clock, and perform an operation time determination operation based on the temperature prediction result, load prediction result and inrush current peak prediction result, so as to perform a hot plug operation based on the operation time. Among them, the operation content and time constraints corresponding to the preset plug-in protocol are as follows: Figure 3 As shown. Specifically, using a preset plug-in protocol and a GPS synchronized clock, and performing a hot-swap operation based on a temperature prediction result, a load prediction result, and a surge current peak prediction result, may include: using a preset GPS synchronized clock in a hot-swap execution module, and establishing a corresponding plug-in timing matrix based on a pre-discharge phase plug-in protocol, a physical unlock phase plug-in protocol, a signal isolation phase plug-in protocol, and a power disconnect phase plug-in protocol and corresponding preset time thresholds, so as to implement a hot-swap operation based on the plug-in timing matrix; wherein the plug-in timing matrix includes a pre-discharge time, a physical unlock time, a signal isolation time, and a power disconnect time corresponding to the pre-discharge phase, the physical unlock phase, the signal isolation phase, and the power disconnect phase, respectively.
[0087] It is worth mentioning that the GPS synchronized clock is used to coordinate multi-node operations to establish a plug-in timing matrix, and the calculation formula is as follows:
[0088] T_matrix = [t_precharge, t_unlock, t_isolation, t_poweroff];
[0089] Among them, t_precharge is the charging time of the GPU power pin before plugging and unplugging, t_unlock is the release time of the mechanical lock, and it matches the response speed of the servo motor; t_isolation is the electrical isolation delay time, which is used for the high-speed optocoupler to cut off the signal path; t_poweroff is the soft shutdown time of the power module, which is used to discharge the energy storage capacitor to a safe voltage, and the time deviation of each stage is no more than 100ns.
[0090] It can be seen that the embodiment of the present application firstly detects the initial power phase difference between the motherboard to be processed and the GPU to be processed by using the preset distributed phase synchronization circuit in the power path of the GPU slot, and uses the reverse compensation signal generated based on the initial power phase difference to process the initial surge current peak corresponding to the initial power phase difference, and determines the network signal to be optimized based on the obtained target surge current peak and the initial network signal; secondly, the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database are used to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and the obtained target signal reflection coefficient is used to optimize the network signal to be optimized to obtain the target network signal; then, based on the target The network signal is processed by the protocol conversion engine and dynamic function area in the core grain architecture adaptive interconnection protocol center to perform protocol compatibility processing on the processing motherboard and the processing GPU to obtain a protocol compatible motherboard and a protocol compatible GPU, and the protocol compatible GPU is diagnosed by using a preset hardware-level pre-diagnosis system. Furthermore, after the obtained diagnostic result characterizes that the diagnosis has passed, the temperature prediction result, load prediction result and surge current peak prediction result of the protocol compatible GPU are predicted by using a preset intelligent prediction model; finally, the four-stage plug-in protocol and GPS synchronization clock in the preset hot-swap execution module are used, and the operation time is determined based on the temperature prediction result, load prediction result and predicted surge current peak value, and the hot-swap operation is realized by using the operation time. In this way, the efficiency and safety of executing GPU hot-swap operations are improved, and the damage rate of hardware involved in the production process is reduced.
[0091] Accordingly, see Figure 4 As shown, the present application also provides a GPU hot-swap device, comprising:
[0092] A power phase difference determination module 11 is used to detect the initial power phase difference between the motherboard to be processed and the GPU to be processed by using a preset distributed phase synchronization circuit in the power path of the GPU slot, and to process the initial surge current peak value corresponding to the initial power phase difference by using a reverse compensation signal generated based on the initial power phase difference, and to determine the network signal to be optimized based on the obtained target surge current peak value and the initial network signal;
[0093] The network signal optimization module 12 is used to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database, and optimize the network signal to be optimized by using the obtained target signal reflection coefficient to obtain the target network signal;
[0094] The GPU diagnosis module 13 is used to perform protocol compatibility processing on the to-be-processed motherboard and the to-be-processed GPU based on the target network signal and using the chip interconnection protocol to obtain a target GPU that is compatible with the protocol of the to-be-processed motherboard, and to diagnose the target GPU using a preset hardware-level pre-diagnosis system, and after the obtained diagnosis result indicates that the diagnosis has passed, to predict the temperature prediction result, load prediction result and surge current peak prediction result of the target GPU using a preset intelligent prediction model;
[0095] The operation time determination module 14 is used to perform hot plug operation based on the temperature prediction result, the load prediction result and the surge current peak prediction result by using a preset plug-in protocol and a GPS synchronized clock.
[0096] As can be seen from the above, before performing GPU hot-plugging, the embodiment of the present application first needs to use the preset distributed phase synchronization circuit in the GPU slot power path to detect the initial power phase difference between the to-be-processed motherboard and the to-be-processed GPU, and use the reverse compensation signal generated based on the initial power phase difference to process the initial surge current peak corresponding to the initial power phase difference, and determine the network signal to be optimized based on the obtained target surge current peak and the initial network signal; secondly, use the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized, and use the obtained target signal reflection coefficient to optimize the network signal to be optimized to obtain the target network signal; then After that, based on the target network signal and using the protocol conversion engine and dynamic function area in the core grain architecture adaptive interconnection protocol center, the protocol-compatible processing is performed on the motherboard to be processed and the GPU to be processed, so as to obtain the protocol-compatible motherboard and the protocol-compatible GPU, and the protocol-compatible GPU is diagnosed using the preset hardware-level pre-diagnosis system. Furthermore, after the obtained diagnostic result characterizes that the diagnosis has passed, the temperature prediction result, load prediction result and surge current peak prediction result of the protocol-compatible GPU are predicted using the preset intelligent prediction model; finally, the four-stage plug-in protocol and GPS synchronization clock in the preset hot-swap execution module are used, and the operation time is determined based on the temperature prediction result, load prediction result and predicted surge current peak value, and the hot-swap operation is realized using the operation time. In this way, the efficiency and safety of executing GPU hot-swap operations are improved, and the damage rate of hardware involved in the production process is reduced.
[0097] In some specific implementations, the power phase difference determining module 11 may specifically include:
[0098] A power phase difference detection unit, used to deploy a preset distributed phase synchronization circuit into a GPU slot power path, and use an inner loop in a preset dual-loop feedback control mechanism in the preset distributed phase synchronization circuit to call a preset programmable logic device, so as to use a preset zero-crossing switching algorithm to detect in real time an initial power phase difference between a to-be-processed motherboard and a to-be-processed GPU;
[0099] A compensation signal generating unit, configured to determine whether a value corresponding to the initial power supply phase difference is greater than a preset phase difference threshold value, and if the value corresponding to the initial power supply phase difference is greater than the preset phase difference threshold value, using the outer loop in the preset dual-loop feedback control mechanism to call a preset dynamic adjustment algorithm and generate a reverse compensation signal based on the initial power supply phase difference;
[0100] A network signal determination unit is used to use the reverse compensation signal to reduce the initial surge current peak corresponding to the initial power supply phase difference to no more than a preset safety threshold, obtain a target surge current peak, and determine the network signal to be optimized based on the target surge current peak.
[0101] In some specific implementations, the network signal optimization module 12 may specifically include:
[0102] A mode determination unit, configured to generate a reconfigurable impedance matching network based on a PCIe interface, integrate the reconfigurable impedance matching network at the PCIe interface, and determine whether a current mode is a preset low power consumption mode;
[0103] A first reflection coefficient adjustment unit is used to activate a preset series inductance compensation mechanism if the current mode is the preset low power consumption mode, and call a programmable LC array in the reconfigurable impedance matching network at the PCIe interface, and reduce the initial signal reflection coefficient based on a preset GPU impedance characteristic database to obtain a target signal reflection coefficient;
[0104] The second reflection coefficient adjustment unit is used to set the current mode to the preset high bandwidth mode if the current mode is not the preset low power consumption mode, and suppress the target high-frequency noise corresponding to the network signal to be optimized to obtain the suppressed network signal, and then call the programmable LC array and increase the initial signal reflection coefficient corresponding to the suppressed network signal based on the preset GPU impedance characteristic database to obtain the target signal reflection coefficient.
[0105] In some specific implementations, the GPU diagnostic module 13 may specifically include:
[0106] A delay time determination unit is used to call a protocol conversion engine corresponding to a chip interconnection protocol after receiving the target network signal, and determine a protocol conversion delay time for the to-be-processed motherboard and the to-be-processed GPU based on a preset hardware state machine implementation protocol to obtain a protocol conversion delay time; wherein the protocol conversion delay time is less than a preset delay time threshold; the chip interconnection protocol includes the preset hardware state machine implementation protocol and a dynamic function area; the dynamic function area includes a power function area and a signal function area;
[0107] A contact pressure determination unit, configured to call the power function area designed based on a preset enhanced contact design technology to increase the contact pressure corresponding to the GPU to be processed to no less than a preset pressure safety threshold, thereby obtaining a target contact pressure;
[0108] A differential pair spacing determination unit, configured to call the signal function area to process the initial differential pair spacing corresponding to the GPU to be processed to obtain a target differential pair spacing;
[0109] The target GPU determination unit is used to perform protocol compatibility processing on the motherboard to be processed and the GPU to be processed based on the protocol conversion delay time, the target contact pressure and the target differential pair spacing to obtain a target GPU that is compatible with the protocol of the motherboard to be processed.
[0110] In some specific implementations, the GPU diagnostic module 13 may specifically include:
[0111] a capacitance deviation diagnosis result determination unit, configured to detect a capacitance deviation corresponding to the target GPU using a preset hardware-level pre-diagnosis system and based on a preset test signal, and obtain a capacitance deviation diagnosis result;
[0112] an abnormal contact diagnosis result determination unit, configured to identify abnormal contacts in the target GPU using a preset measurement method and a preset contact resistance scanning technology, and obtain an abnormal contact diagnosis result;
[0113] The protocol handshake diagnosis result determination unit is used to verify the target GPU based on a preset ESD protection verification technology to obtain a protection diagnosis result, and to verify the target GPU using a preset protocol handshake test technology to obtain a protocol handshake diagnosis result.
[0114] In some specific implementations, the GPU diagnostic module 13 may specifically include:
[0115] A diagnostic result judgment unit, used to judge in sequence whether the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result and the protocol handshake diagnosis result all indicate that the diagnosis has passed;
[0116] A first prediction result determination unit is used to predict the temperature and load of the GPU to be processed by using a preset main model in a preset intelligent prediction model deployed in the AI chip to obtain a temperature prediction result and a load prediction result if the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result, and the protocol handshake diagnosis result all indicate that the diagnosis is passed;
[0117] A second prediction result determination unit, configured to perform an inrush current peak value preset operation on the GPU to be processed by using a preset auxiliary model in the preset intelligent prediction model to obtain an inrush current peak value prediction result;
[0118] The model retraining unit is used to determine whether the temperature prediction result, the load prediction result and the surge current peak prediction result exceed the corresponding prediction result thresholds respectively. If they exceed, the model retraining operation of the preset intelligent prediction model is triggered.
[0119] In some specific implementations, the operation time determination module 14 may specifically include:
[0120] A hot-swap operation execution unit is used to use the GPS synchronized clock in a preset hot-swap execution module, and to establish a corresponding plug-in timing matrix based on a pre-discharge phase plug-in protocol, a physical unlock phase plug-in protocol, a signal isolation phase plug-in protocol, and a power disconnect phase plug-in protocol and corresponding preset time thresholds, so as to implement a hot-swap operation based on the plug-in timing matrix; wherein the plug-in timing matrix includes a pre-discharge time, a physical unlock time, a signal isolation time, and a power disconnect time corresponding to the pre-discharge phase, the physical unlock phase, the signal isolation phase, and the power disconnect phase, respectively.
[0121] Furthermore, the present application also discloses an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input and output interface 25 and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the GPU hot plug method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0122] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0123] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0124] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the GPU hot-plug method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0125] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the GPU hot-plugging method disclosed above. The specific steps of the method can refer to the corresponding contents disclosed in the above embodiments, and will not be repeated here.
[0126] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0127] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0128] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0129] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0130] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A GPU hot-swap method, characterized in that: include: Using a preset distributed phase synchronization circuit in a GPU slot power path to detect an initial power phase difference between a motherboard to be processed and a GPU to be processed, and using a reverse compensation signal generated based on the initial power phase difference to process an initial surge current peak value corresponding to the initial power phase difference, and determining a network signal to be optimized based on the obtained target surge current peak value and the initial network signal; The initial signal reflection coefficient corresponding to the network signal to be optimized is adjusted by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database, and the network signal to be optimized is optimized by using the obtained target signal reflection coefficient to obtain the target network signal; Based on the target network signal and using the chip interconnection protocol, the to-be-processed motherboard and the to-be-processed GPU are subjected to protocol compatibility processing to obtain a target GPU that is compatible with the to-be-processed motherboard protocol, and the target GPU is diagnosed using a preset hardware-level pre-diagnosis system, and after the obtained diagnostic result indicates that the diagnosis has passed, a preset intelligent prediction model is used to predict a temperature prediction result, a load prediction result, and an inrush current peak prediction result of the target GPU; A hot plug operation is performed using a preset plug-in protocol and a GPS synchronized clock and based on the temperature prediction result, the load prediction result and the surge current peak prediction result.
2. The GPU hot-swap method according to claim 1, wherein: The method uses a preset distributed phase synchronization circuit in a GPU slot power path to detect an initial power phase difference between a motherboard to be processed and a GPU to be processed, and uses a reverse compensation signal generated based on the initial power phase difference to process an initial surge current peak value corresponding to the initial power phase difference, and determines a network signal to be optimized based on the obtained target surge current peak value and the initial network signal, including: Deploy a preset distributed phase synchronization circuit into a GPU slot power path, and use an inner loop in a preset dual-loop feedback control mechanism in the preset distributed phase synchronization circuit to call a preset programmable logic device, so as to use a preset zero-crossing switching algorithm to detect in real time an initial power phase difference between a to-be-processed motherboard and a to-be-processed GPU; Determine whether the value corresponding to the initial power supply phase difference is greater than a preset phase difference threshold value, and if the value corresponding to the initial power supply phase difference is greater than the preset phase difference threshold value, use the outer loop in the preset dual-loop feedback control mechanism to call a preset dynamic adjustment algorithm and generate a reverse compensation signal based on the initial power supply phase difference; The reverse compensation signal is used to reduce the initial surge current peak value corresponding to the initial power supply phase difference to no more than a preset safety threshold value, to obtain a target surge current peak value, and the network signal to be optimized is determined based on the target surge current peak value.
3. The GPU hot-swap method according to claim 1, characterized in that: The method of adjusting the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database comprises: Generate a reconfigurable impedance matching network based on the PCIe interface, integrate the reconfigurable impedance matching network at the PCIe interface, and determine whether the current mode is a preset low power consumption mode; If the current mode is the preset low power mode, a preset series inductance compensation mechanism is activated, and a programmable LC array in the reconfigurable impedance matching network at the PCIe interface is called, and the initial signal reflection coefficient is reduced based on a preset GPU impedance characteristic database to obtain a target signal reflection coefficient; If the current mode is not the preset low power consumption mode, the current mode is set to the preset high bandwidth mode, and the target high frequency noise corresponding to the network signal to be optimized is suppressed to obtain the suppressed network signal, and then the programmable LC array is called and the initial signal reflection coefficient corresponding to the suppressed network signal is increased based on the preset GPU impedance characteristic database to obtain the target signal reflection coefficient.
4. The GPU hot-swap method according to claim 1, characterized in that: The method of performing protocol compatibility processing on the motherboard to be processed and the GPU to be processed based on the target network signal and using the chiplet interconnection protocol to obtain a target GPU that is compatible with the protocol of the motherboard to be processed includes: After receiving the target network signal, calling a protocol conversion engine corresponding to the core grain interconnection protocol, and determining a protocol conversion delay time between the to-be-processed mainboard and the to-be-processed GPU based on a preset hardware state machine implementation protocol, to obtain a protocol conversion delay time; Wherein, the protocol conversion delay time is less than a preset delay time threshold; the chiplet interconnection protocol includes the preset hardware state machine implementation protocol and a dynamic function area; the dynamic function area includes a power function area and a signal function area; Calling the power function area designed based on the preset enhanced contact design technology to increase the contact pressure corresponding to the GPU to be processed to no less than a preset pressure safety threshold, thereby obtaining a target contact pressure; Calling the signal function area to process the initial differential pair spacing corresponding to the GPU to be processed to obtain a target differential pair spacing; The to-be-processed motherboard and the to-be-processed GPU are subjected to protocol compatibility processing based on the protocol conversion delay time, the target contact pressure and the target differential pair spacing to obtain a target GPU that is protocol-compatible with the to-be-processed motherboard.
5. The GPU hot-swap method according to claim 1, characterized in that: The diagnosing the target GPU by using a preset hardware-level pre-diagnosis system includes: Using a preset hardware-level pre-diagnosis system and based on a preset test signal, detecting a capacitance deviation corresponding to the target GPU to obtain a capacitance deviation diagnosis result; Using a preset measurement method and a preset contact resistance scanning technology to identify abnormal contacts in the target GPU, and obtain abnormal contact diagnosis results; The target GPU is verified based on a preset ESD protection verification technology to obtain a protection diagnosis result, and the target GPU is verified using a preset protocol handshake test technology to obtain a protocol handshake diagnosis result.
6. The GPU hot-swap method according to claim 5, characterized in that: After the obtained diagnostic result indicates that the diagnosis is passed, a preset intelligent prediction model is used to predict the temperature prediction result, the load prediction result and the surge current peak prediction result of the target GPU, including: sequentially determining whether the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result, and the protocol handshake diagnosis result all indicate that the diagnosis has passed; If the capacitance deviation diagnosis result, the abnormal contact diagnosis result, the protection diagnosis result, and the protocol handshake diagnosis result all indicate that the diagnosis is passed, the temperature and load of the GPU to be processed are predicted using a preset main model in a preset intelligent prediction model deployed in the AI chip to obtain a temperature prediction result and a load prediction result; Using a preset auxiliary model in the preset intelligent prediction model to perform an inrush current peak value preset operation on the GPU to be processed, to obtain an inrush current peak value prediction result; Determine whether the temperature prediction result, the load prediction result, and the surge current peak prediction result exceed the corresponding prediction result thresholds respectively; if so, trigger a model retraining operation of the preset intelligent prediction model.
7. The GPU hot-swap method according to claim 1, characterized in that: The hot plug operation is performed by using a preset plug-in protocol and a GPS synchronized clock and based on the temperature prediction result, the load prediction result and the surge current peak prediction result, including: Using the GPS synchronized clock in the preset hot-swap execution module, and based on the pre-discharge phase plug-in protocol, the physical unlock phase plug-in protocol, the signal isolation phase plug-in protocol, and the power disconnection phase plug-in protocol and the corresponding preset time thresholds, a corresponding plug-in timing matrix is established to implement the hot-swap operation based on the plug-in timing matrix; The plug-in timing matrix includes a pre-discharge time, a physical unlocking time, a signal isolation time and a power disconnection time corresponding to the pre-discharge stage, the physical unlocking stage, the signal isolation stage and the power disconnection stage respectively.
8. A GPU hot-swap device, characterized in that: include: A power phase difference determination module, used to detect the initial power phase difference between the to-be-processed motherboard and the to-be-processed GPU using a preset distributed phase synchronization circuit in the power path of the GPU slot, and to process the initial surge current peak value corresponding to the initial power phase difference using a reverse compensation signal generated based on the initial power phase difference, and to determine the network signal to be optimized based on the obtained target surge current peak value and the initial network signal; A network signal optimization module, used to adjust the initial signal reflection coefficient corresponding to the network signal to be optimized by using the programmable LC array in the reconfigurable impedance matching network and the preset GPU impedance characteristic database, and optimize the network signal to be optimized by using the obtained target signal reflection coefficient to obtain a target network signal; A GPU diagnostic module, configured to perform protocol compatibility processing on the to-be-processed motherboard and the to-be-processed GPU based on the target network signal and using a core particle interconnection protocol to obtain a target GPU that is compatible with the protocol of the to-be-processed motherboard, and to diagnose the target GPU using a preset hardware-level pre-diagnosis system, and after the obtained diagnostic result indicates that the diagnosis has passed, to predict the temperature prediction result, load prediction result, and surge current peak prediction result of the target GPU using a preset intelligent prediction model; The operation time determination module is used to perform hot plug operation based on the temperature prediction result, the load prediction result and the surge current peak prediction result by using a preset plug-in protocol and a GPS synchronized clock.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the GPU hot-plug method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the GPU hot-plug method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method for adaptively matching impedance according to protocol and connector
CN120705104A
A method for adapting impedance to protocol and a connector
CN120705104B
Regulation and control method, device and equipment of load power supply loop, medium and program product
CN120743076A