Cloud mobile phone server live migration method and related equipment

By obtaining the temperature, load and real-time cooling capacity parameters of the server node and selecting the target migration node based on the multi-protect weight algorithm, the problem of low cooling efficiency of traditional cloud mobile phone servers and the inability to respond to dynamic thermal load changes in real time is solved, and efficient cooling resource utilization and service continuity are achieved.

CN120336027APending Publication Date: 2025-07-18启朔(深圳)科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510492956.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional air-cooled cooling solutions are inefficient in high-frequency CPU scenarios. Immersed liquid cooling has problems such as attenuation of cooling liquid dielectric performance and complex maintenance. The existing migration technology cannot respond to dynamic thermal load changes in real time, resulting in low utilization of heat dissipation resources, making it difficult to ensure the high availability and low latency requirements of cloud mobile phone services.

Method used

By obtaining the temperature, load and real-time cooling capacity parameters of the server node, a candidate node score collection is generated based on the multi-factor weight algorithm, the target migration node is selected, and the thermal migration operation is performed. Combined with software-defined network and incremental synchronization technology, dynamic cooling resource matching and service continuity are achieved.

Benefits of technology

The dynamic matching of heat dissipation capabilities and task allocation is achieved, the utilization rate of heat dissipation resources and system energy efficiency are improved, the node temperature is stable within the safety threshold, and the problems of heat dissipation lag, high energy consumption and service delay in traditional technologies are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336027A_ABST
    Figure CN120336027A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud mobile phone server thermal migration method and related equipment, and relates to the technical field of data thermal management, and the method comprises the steps: obtaining temperature data, load data and real-time heat dissipation capability parameters of server nodes; determining a thermal state evaluation value based on the temperature data, the load data and the real-time heat dissipation capability parameter; when the thermal state evaluation value is greater than a preset evaluation threshold value, generating a candidate node score set based on a multi-dimensional weight algorithm; selecting a target migration node according to a sorting result of the candidate node score set; and executing a live migration operation on the target migration node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data thermal management, and particularly to a method for hot migration of cloud mobile phone servers and related devices. Background Art

[0002] With the rapid development of cloud mobile phone technology, the power density and computing intensity of server clusters have increased. The efficiency of traditional air-cooled heat dissipation solutions drops sharply in high-frequency CPU scenarios, and the proportion of heat dissipation energy consumption can reach more than 40% of the total power consumption, making it difficult to meet the energy-saving requirements. Although immersion liquid cooling can improve the heat dissipation capacity, it faces problems such as attenuation of the dielectric properties of the coolant and complex maintenance, and it is difficult to achieve a balance between heat dissipation efficiency and system reliability. In addition, existing migration technologies are mainly based on load balancing strategies and cannot respond to dynamic thermal load changes in real time, resulting in low utilization rate of heat dissipation resources and making it difficult to guarantee the high availability and low latency requirements of cloud mobile phone services. Therefore, there is an urgent need for a method for hot migration of cloud mobile phone servers to solve the above-mentioned technical problems. Summary of the Invention

[0003] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further elaborated in detail in the Detailed Description section. The Summary of the Invention section of this application does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0004] In a first aspect, this application provides a method for hot migration of cloud mobile phone servers, the method including:

[0005] Obtain temperature data, load data, and real-time heat dissipation capacity parameters of a server node;

[0006] Based on the temperature data, load data, and real-time heat dissipation capacity parameters, determine a thermal state evaluation value;

[0007] When the thermal state evaluation value is greater than a preset evaluation threshold, generate a candidate node score set based on a multi-dimensional weight algorithm;

[0008] Select a target migration node according to the sorting result of the candidate node score set;

[0009] Perform a hot migration operation on the target migration node.

[0010] In some embodiments, obtaining temperature data, load data, and real-time heat dissipation capacity parameters of a server node includes:

[0011] Collect temperature data of the server node through a temperature sensor, where the temperature data includes first computing unit temperature data and second computing unit temperature data;

[0012] Obtain the load data of the server node through the hardware performance counter, where the load data includes the load data of the first computing unit, the load data of the second computing unit, and the communication bandwidth load data;

[0013] Based on the coolant flow rate, coolant physical property parameters, and the fouling attenuation coefficient of the liquid cooling plate, calculate the real-time heat dissipation capacity parameters of the server node through the dynamic heat dissipation model, where the fouling attenuation coefficient of the liquid cooling plate is determined based on the attenuation curve trained from historical operation and maintenance data.

[0014] In some embodiments, based on the temperature data, load data, and real-time heat dissipation capacity parameters, determine the thermal state evaluation value, including:

[0015] Perform weight assignment on the temperature data of the first computing unit, the temperature data of the second computing unit, the load data of the first computing unit, the load data of the second computing unit, and the communication bandwidth load data respectively to generate dynamic weight coefficients;

[0016] Based on the dynamic weight coefficients, perform weighted fusion on the temperature data and the load data to obtain the initial heat load value;

[0017] Based on the real-time heat dissipation capacity parameters and the preset heat dissipation requirement threshold, determine the heat dissipation capacity margin;

[0018] Based on the initial heat load value and the heat dissipation capacity margin, determine the thermal state evaluation value.

[0019] In some embodiments, when the thermal state evaluation value is greater than the preset evaluation threshold, generate a candidate node score set based on the multi-dimensional weight algorithm, including:

[0020] When the thermal state evaluation value is greater than the preset evaluation threshold, perform the following steps:

[0021] Based on the network topology of the server cluster, determine the hop count constraint conditions of the candidate nodes, where the hop count constraint conditions include the first hop count threshold and the second hop count threshold;

[0022] Based on the temperature margin of the candidate node and the first preset temperature threshold, calculate the temperature score item;

[0023] According to the heat dissipation capacity margin of the candidate node and the preset heat dissipation requirement threshold, calculate the heat dissipation score item;

[0024] Perform weighted summation of the temperature score item and the heat dissipation score item through the preset weight ratio to obtain the comprehensive score of the candidate node;

[0025] Screen the candidate nodes that meet the hop count constraint conditions from the candidate nodes to form a candidate node set;

[0026] Sort the candidate node set in descending order based on the comprehensive score of the candidate nodes to generate a sorted candidate node score set.

[0027] In some embodiments, selecting a target migration node according to the sorting result of the candidate node score set includes:

[0028] Based on the sorting result of the candidate node score set, determine the candidate node with the highest score;

[0029] Verify whether the candidate node with the highest score meets the preset network bandwidth constraint;

[0030] When the preset network bandwidth constraint is met, determine the candidate node with the highest score as the target migration node;

[0031] When the preset network bandwidth constraint is not met, sequentially select the sub-optimal candidate nodes from the candidate node score set in descending order of scores for verification until a target migration node that meets the preset network bandwidth constraint is determined.

[0032] In some embodiments, performing a hot migration operation on the target migration node includes:

[0033] Reserve a dedicated bandwidth channel for the target migration node through a software-defined network controller, and allocate an accelerated forwarding priority to the dedicated bandwidth channel;

[0034] Pre-allocate virtualized resources on the target migration node and establish a shadow process synchronized with the source node's business logic;

[0035] Migrate the service status data to the target migration node based on the incremental synchronization algorithm;

[0036] After the migration is completed, switch the service request to the target migration node.

[0037] In some embodiments, after performing the hot migration operation on the target migration node, it further includes:

[0038] Real-time monitor the temperature change rate and coolant flow state parameters of the target migration node;

[0039] Dynamically adjust the coolant flow rate and heat exchanger operation parameters based on the model predictive control framework so that the temperature of the target migration node does not exceed the second preset temperature threshold.

[0040] In a second aspect, the present application proposes a cloud phone server hot migration device, including:

[0041] A node data acquisition unit for acquiring the temperature data, load data, and real-time heat dissipation capacity parameters of the server node;

[0042] A thermal state evaluation unit determines a thermal state evaluation value based on temperature data, load data, and real-time heat dissipation capacity parameters;

[0043] A candidate node determination unit is configured to generate a candidate node score set based on a multi-dimensional weight algorithm when the thermal state evaluation value is greater than a preset evaluation threshold;

[0044] A migration node selection unit is configured to select a target migration node according to the sorting result of the candidate node score set;

[0045] A thermal migration operation unit is configured to perform a thermal migration operation on the target migration node.

[0046] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to implement the steps of the method for hot migration of a cloud mobile phone server according to any one of the first aspects when executing the computer program stored in the memory.

[0047] In a fourth aspect, the present application provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it implements the method for hot migration of a cloud mobile phone server according to any one of the first aspects.

[0048] In summary, the present application dynamically obtains the temperature, load, and real-time heat dissipation capacity parameters of server nodes, accurately evaluates the thermal state, and generates a candidate node score set. By combining network topology constraints, it selects the optimal target node for migration. This method realizes the dynamic matching of heat dissipation capacity and task allocation, effectively avoids the imbalance of high-heat task allocation, and at the same time ensures service continuity through a reserved bandwidth channel and an incremental synchronization mechanism. In addition, based on the collaborative optimization of the multi-dimensional weight algorithm and real-time parameters, it improves the utilization rate of heat dissipation resources and system energy efficiency, ensures that the node temperature is stable within the safety threshold, and solves the technical problems of heat dissipation lag, high energy consumption, and service delay in traditional technologies.

[0049] For the method for hot migration of a cloud mobile phone server proposed by the present application, other advantages, objectives, and features of the present application will be partially reflected by the following description, and partially will be understood by those skilled in the art through the research and practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to limit this specification. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0051] Figure 1 It is a schematic flowchart of a method for hot migration of a cloud mobile phone server provided by an embodiment of the present application;

[0052] Figure 2 Schematic structural diagram of a cloud mobile phone server hot migration device provided by an embodiment of the present application;

[0053] Figure 3 Schematic structural diagram of an electronic device for cloud mobile phone server hot migration provided by an embodiment of the present application. Detailed implementation manners

[0054] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0055] Please refer to Figure 1 , which is a schematic flowchart of a cloud mobile phone server hot migration method provided by an embodiment of the present application, and specifically may include:

[0056] S110. Obtain temperature data, load data and real-time heat dissipation capacity parameters of a server node;

[0057] Exemplarily, obtaining temperature data, load data and real-time heat dissipation capacity parameters of a server node is a basic link for hot migration decision-making. The temperature data is collected in real time by high-precision sensors deployed on computing units (such as CPUs and GPUs), reflecting the real-time thermal state of the node; the load data is dynamically obtained through hardware performance counters, covering key indicators such as computing resource occupancy rate and communication bandwidth load, and is used to quantify the thermal load intensity of the node. The collaborative analysis of the two can accurately identify overheating risk nodes and provide multi-dimensional data support for subsequent heat dissipation capacity evaluation.

[0058] The calculation of real-time heat dissipation capacity parameters is based on the dynamic operating state of the cooling system, including coolant flow rate, physical properties parameters, and the fouling attenuation coefficient of the liquid cooling plate, etc. The fouling attenuation coefficient is dynamically corrected by the attenuation curve trained with historical operation and maintenance data, which characterizes the degree of heat dissipation efficiency attenuation caused by dirt deposition in the microchannels. Combining this parameter with physical quantities such as real-time flow rate and temperature difference, a dynamic heat dissipation model is constructed to quantify the instantaneous heat dissipation capacity margin of the node, providing key inputs for the heat migration strategy and ensuring the real-time and accuracy of heat dissipation resource allocation.

[0059] S120. Determine the heat state evaluation value based on the temperature data, load data, and real-time heat dissipation capacity parameters;

[0060] Exemplarily, the determination of the heat state evaluation value depends on the comprehensive analysis of the temperature data, load data, and real-time heat dissipation capacity parameters. The temperature data reflects the real-time heat load intensity of the node, and the load data quantifies the occupancy of computing resources and communication bandwidth. Combining the two can identify the heat risk level of the node; the real-time heat dissipation capacity parameters are dynamically calculated through the coolant flow rate, physical properties parameters, and fouling attenuation coefficient, which characterize the remaining capacity of the current heat dissipation system of the node. Through multi-dimensional parameter fusion, the system can accurately evaluate whether the heat state of the node exceeds the safety threshold.

[0061] The core logic of the heat state evaluation value is to balance the heat load demand and heat dissipation capacity supply of the node. An initial heat load value is generated based on the weighted fusion of temperature and load data, and then compared with the real-time heat dissipation capacity margin. Finally, the heat state evaluation value is determined through proportional relationship or difference calculation. This evaluation value directly reflects the thermal stability of the node, providing a quantitative basis for subsequent migration decisions and ensuring that high-heat tasks are preferentially allocated to nodes with sufficient heat dissipation capacity.

[0062] S130. When the heat state evaluation value is greater than the preset evaluation threshold, generate a candidate node score set based on the multi-dimensional weight algorithm;

[0063] Exemplarily, when the heat state evaluation value exceeds the preset evaluation threshold, it indicates that the current node has an overheating risk and insufficient local heat dissipation capacity, and the candidate node scoring mechanism needs to be triggered. The system determines the hop count constraint of the candidate nodes based on the network topology of the server cluster, and filters out the candidate node set that meets the network delay requirements, providing an initial screening range for subsequent scoring.

[0064] For the candidate node set, the temperature margin and heat dissipation capacity margin are weighted and calculated through the multi-dimensional weight algorithm to generate the comprehensive score of the candidate nodes. Among them, the weight of the temperature margin is higher than that of the heat dissipation capacity margin, giving priority to avoiding the migration risk of high-heat nodes. Finally, the candidate node score set is generated in descending order of scores, providing a priority sorting basis for the selection of migration targets.

[0065] S140. Select a target migration node according to the sorting result of the candidate node score set;

[0066] Exemplarily, the selection of the target migration node is based on the sorting result of the candidate node score set, and the node with the highest comprehensive score is preferentially selected as the migration target. The score set is generated by a multi-dimensional weighting algorithm, comprehensively considering the temperature margin, heat dissipation capacity margin and network topology constraints of the candidate nodes, ensuring that the selected node has sufficient heat dissipation resources while meeting the low-latency network transmission requirements, thereby optimizing service continuity and thermal stability.

[0067] In the application of the sorting result, the system first verifies whether the node with the highest score meets the preset network bandwidth constraint. If it meets, it is directly selected as the target migration node; if it does not meet, the sub-optimal nodes are verified in descending order of scores until a candidate node with both high heat dissipation capacity and network performance requirements is screened out. This mechanism balances the heat dissipation efficiency and network resource availability, avoiding migration failures or service degradations caused by overemphasis on a single indicator.

[0068] S150. Perform a live migration operation on the target migration node.

[0069] Exemplarily, the core goal of performing a live migration operation on the target migration node is to achieve seamless service switching and dynamic adaptation of heat dissipation resources. The software-defined network controller reserves a dedicated bandwidth channel for the migration operation and realizes sub-second state saving based on an improved checkpoint recovery tool to ensure the efficient synchronization of service data; at the same time, virtualized resources are pre-allocated on the target node and a shadow process is established, enabling the source node and the target node to process requests in parallel, minimizing the service interruption time and ensuring business continuity.

[0070] During the live migration process, the system is linked with the liquid cooling control module in real time, and the coolant flow rate and heat exchanger operation parameters are dynamically adjusted according to the temperature change rate of the target node. The migration task allocation and liquid cooling parameters are jointly optimized through a model predictive control framework to ensure that the temperature of the target node is stable within the safety threshold while minimizing the energy consumption index. This collaborative mechanism effectively solves the problems of lagging heat dissipation response and energy efficiency imbalance in traditional live migration, achieving a double improvement in heat dissipation efficiency and service quality.

[0071] In summary, in the embodiments of the present application, by constructing a multi-dimensional dynamic collaborative regulation mechanism, improvements have been achieved in terms of thermal management efficiency, system stability, and service quality. This method innovatively integrates real-time physical environment parameters with virtualized resource scheduling, and through the collaborative monitoring of the temperature sensing network and hardware performance counters, accurately captures the dynamic thermal load characteristics of server nodes. Based on the dynamic heat dissipation model of the liquid cooling system's physical property parameters and fouling attenuation curve, online evaluation and prediction of the heat dissipation capacity are realized, effectively overcoming the problem of response lag in traditional static heat dissipation planning. By introducing the hop count constraint condition of network topology awareness and the multi-dimensional weight scoring algorithm, heat dissipation margin, temperature gradient distribution, and network transmission efficiency are taken into account simultaneously in candidate node screening, avoiding sub-optimal migration decisions caused by simple load balancing strategies. In the migration execution phase, software-defined network channel reservation and incremental state synchronization technologies are adopted to reduce service switching delay and ensure the high availability of the cloud phone service. After migration, based on the closed-loop regulation mechanism of model predictive control, through the real-time feedback of coolant flow state parameters and temperature change rate, dynamic optimization of the heat dissipation system operation parameters is achieved, forming a full-link closed-loop management from thermal state warning to resource migration and then to heat dissipation regulation. Compared with traditional solutions, the present application shows advantages in terms of heat dissipation energy consumption ratio, node temperature balance, and service continuity, providing an innovative solution for thermal management in high-density computing scenarios.

[0072] In some examples, obtaining temperature data, load data, and real-time heat dissipation capacity parameters of a server node includes:

[0073] Collecting temperature data of the server node through a temperature sensor, where the temperature data includes first computing unit temperature data and second computing unit temperature data;

[0074] Obtaining load data of the server node through a hardware performance counter, where the load data includes first computing unit load data, second computing unit load data, and communication bandwidth load data;

[0075] Based on the coolant flow rate, coolant physical property parameters, and liquid cooling plate fouling attenuation coefficient, calculating the real-time heat dissipation capacity parameters of the server node through a dynamic heat dissipation model, where the liquid cooling plate fouling attenuation coefficient is determined based on the attenuation curve trained from historical operation and maintenance data.

[0076] Exemplarily, the acquisition of temperature data is achieved through high-precision temperature sensors deployed on server nodes, specifically including the temperature data of the first computing unit (such as CPU temperature) and the temperature data of the second computing unit (such as GPU or NPU temperature). The temperature sensor uses a PT1000 platinum resistance sensor with a measurement accuracy of ±0.1°C, which monitors the thermal state of the computing unit in real time and transmits it to the central processing unit through the data bus. The acquisition frequency of temperature data is synchronized with the heat migration control cycle (for example, once every 5 seconds) to ensure the real-time nature of the thermal state assessment. The temperature data of the first computing unit and the second computing unit respectively reflect the heat load distribution of different computing modules, providing key inputs for the subsequent multi-dimensional weight algorithm.

[0077] The acquisition of load data is dynamically collected through hardware performance counters, including the load data of the first computing unit (CPU utilization rate), the load data of the second computing unit (GPU / NPU utilization rate), and the communication bandwidth load data (such as the throughput of PCIe channels or network interfaces). The hardware performance counter records the occupancy rate of computing resources at a preset sampling period and statistically calculates the communication bandwidth load through a bandwidth monitoring chip (such as an sFlow sampler) with an accuracy of up to ±0.5%. The load data is analyzed in collaboration with the temperature data to quantify the heat load generation rate of the node and identify the different requirements of compute-intensive tasks or communication-intensive tasks for the cooling system.

[0078] The calculation of real-time heat dissipation capacity parameters is based on a dynamic heat dissipation model, the inputs of which include the coolant flow rate (0.5 - 2.5 L / min), the physical properties of the coolant (density ρ, specific heat capacity c), and the fouling attenuation coefficient η(t) of the liquid cooling plate. The physical properties of the coolant are determined according to the physical property table of the selected cooling medium (such as fluorinated liquid Novec 7100); the fouling attenuation coefficient η(t) of the liquid cooling plate is dynamically corrected through an exponential decay model (η(t) = η0·e^(-kt)) trained with historical operation and maintenance data, where the decay rate k is obtained by fitting the microchannel flow resistance monitoring data. The fouling attenuation coefficient characterizes the degree of heat dissipation efficiency attenuation of the liquid cooling plate caused by dirt deposition, with a value range of 0.85 - 1.0, and its periodic update mechanism (such as once every 24 hours) ensures the accuracy of the heat dissipation capacity calculation.

[0079] The dynamic heat dissipation model calculates the real-time heat dissipation capacity of the node through the formula Q = ρ·c·v·ΔT·η(t), where ΔT is the temperature difference between the inlet and outlet of the liquid cooling plate, and v is the coolant flow rate. This model combines real-time flow rate regulation and fouling coefficient correction to dynamically quantify the heat dissipation capacity margin of the node. The comparison between the heat dissipation capacity margin and the heat load value directly determines whether to trigger the migration decision. Through the multi-dimensional analysis integrating temperature, load, and heat dissipation capacity parameters, the system realizes the precise allocation of heat dissipation resources and avoids the problem of misallocation of heat dissipation resources caused by single-parameter evaluation in traditional solutions.

[0080] In some instances, based on temperature data, load data, and real-time heat dissipation capacity parameters, a thermal state evaluation value is determined, including:

[0081] Weight assignments are respectively made to the temperature data of the first computing unit, the temperature data of the second computing unit, the load data of the first computing unit, the load data of the second computing unit, and the communication bandwidth load data to generate dynamic weight coefficients;

[0082] Based on the dynamic weight coefficients, the temperature data and the load data are weighted and fused to obtain an initial heat load value;

[0083] Based on the real-time heat dissipation capacity parameters and a preset heat dissipation requirement threshold, a heat dissipation capacity margin is determined;

[0084] Based on the initial heat load value and the heat dissipation capacity margin, the thermal state evaluation value is determined.

[0085] Exemplarily, the determination of the thermal state evaluation value first performs weighted fusion on multi-source parameters through a dynamic weight assignment mechanism. The temperature data of the first computing unit (such as CPU temperature) and the temperature data of the second computing unit (such as GPU temperature) are respectively given a first temperature weight and a second temperature weight; the load data of the first computing unit (CPU utilization rate), the load data of the second computing unit (GPU utilization rate), and the communication bandwidth load data (such as PCIe channel throughput) are respectively assigned a first load weight, a second load weight, and a communication weight. The initial weight coefficients are set using a composite weight algorithm. For example, the first load weight (CPU load weight) accounts for 0.6, the second load weight (GPU load weight) accounts for 0.3, and the communication weight (PCIe bandwidth load weight) accounts for 0.1, and the temperature weights are dynamically assigned according to the heat source distribution characteristics. The weight coefficients are periodically optimized through a genetic algorithm, and the optimization period is a preset time interval (such as once every 24 hours) to adapt to the heat load characteristics under different service scenarios. The fitness function of the genetic algorithm is constructed based on the historical heat load distribution characteristics, migration success rate, and heat dissipation efficiency index, aiming to minimize the thermal state evaluation error, dynamically adjust the weight ratios of each parameter (such as the dynamic range of the CPU weight is 0.5 - 0.7, the GPU weight is 0.2 - 0.4, and the PCIe weight is 0.05 - 0.15), and ensure that the sum of the first load weight, the second load weight, and the communication weight is 1, ensuring the objectivity and adaptability of the weight assignment.

[0086] During the weighted fusion process, the initial heat load value is calculated through the following steps: Multiply the temperature data of the first calculation unit by the first temperature weight, and multiply the temperature data of the second calculation unit by the second temperature weight to obtain weighted temperature values; Multiply the load data of the first calculation unit by the first load weight, multiply the load data of the second calculation unit by the second load weight, and multiply the communication bandwidth load data by the communication weight to obtain weighted load values. Subsequently, add the weighted temperature value and the weighted load value to generate the initial heat load value. This value comprehensively reflects the real-time heat load intensity of the node, quantifies the correlation between computing resource occupancy and temperature rise, and provides a benchmark input for heat dissipation capacity matching.

[0087] The calculation of the heat dissipation capacity margin is based on the difference between the real-time heat dissipation capacity parameter and the preset heat dissipation requirement threshold. The real-time heat dissipation capacity parameter is obtained through a dynamic heat dissipation model, specifically, the coolant density is multiplied by the coolant specific heat capacity, then multiplied by the product of the coolant flow rate and the temperature difference between the inlet and outlet of the liquid cooling plate, and finally multiplied by the fouling attenuation coefficient of the liquid cooling plate. The fouling attenuation coefficient is dynamically corrected based on the attenuation curve trained by historical operation and maintenance data, and represents the attenuation of heat dissipation efficiency caused by microchannel fouling deposition. The preset heat dissipation requirement threshold is set according to the node type and business scenario. For example, the requirement threshold for high-frequency computing nodes is higher than that for low-frequency storage nodes. The heat dissipation capacity margin is the real-time heat dissipation capacity minus the preset requirement threshold, which directly reflects the remaining capacity of the node's current heat dissipation system.

[0088] The thermal state evaluation value is determined by the ratio of the initial heat load value to the heat dissipation capacity margin. Specifically, divide the initial heat load value by the heat dissipation capacity margin to obtain a normalized evaluation index. When this ratio is greater than 1, it indicates that the heat load exceeds the heat dissipation capacity margin and the node has an overheating risk; when the ratio is less than or equal to 1, the node is in a thermally stable state. The thermal state evaluation value serves as the core basis for migration decisions, ensuring that high-heat tasks are preferentially migrated to nodes with sufficient heat dissipation capacity margins. At the same time, it avoids subjective judgment errors through quantitative indicators, improving the response accuracy and reliability of the thermal management system.

[0089] In some instances, when the thermal state evaluation value is greater than the preset evaluation threshold, a candidate node score set is generated based on the multi-dimensional weight algorithm, including:

[0090] When the thermal state evaluation value is greater than the preset evaluation threshold, the following steps are executed:

[0091] Based on the network topology of the server cluster, determine the hop count constraint conditions for candidate nodes, where the hop count constraint conditions include the first hop count threshold and the second hop count threshold;

[0092] Based on the temperature margin of the candidate node and the first preset temperature threshold, calculate the temperature score item;

[0093] According to the heat dissipation capacity margin of the candidate node and the preset heat dissipation requirement threshold, calculate the heat dissipation score item;

[0094] The temperature scoring item and the heat dissipation scoring item are weighted and summed through a preset weight ratio to obtain the comprehensive score of the candidate node;

[0095] Filter the candidate nodes that meet the hop count constraint conditions from the candidate nodes to form a candidate node set;

[0096] Based on the comprehensive score of the candidate nodes, the candidate node set is sorted in descending order to generate a sorted candidate node score set.

[0097] Exemplarily, first, the hop count constraint conditions of the candidate nodes are determined based on the network topology of the server cluster. Specifically, the system sets a first hop count threshold and a second hop count threshold. The first hop count threshold defines the physical connection level between the candidate node and the source node. For example, the hop count of the nodes within a single cabinet does not exceed 1; the second hop count threshold extends to the nodes across cabinets but in the same computer room, usually set to not exceed 3 hops. This hop count constraint mechanism effectively balances the network transmission delay and the heat dissipation resource distribution, ensuring that the candidate nodes have reasonable physical proximity within the reach of the network topology and avoiding the problem of a sharp increase in service delay caused by long-distance migration.

[0098] The calculation of the temperature scoring item is quantified based on the difference relationship between the temperature margin of the candidate node and the first preset temperature threshold. The temperature margin is defined as the first preset temperature threshold minus the current temperature value of the candidate node. The larger the difference, the higher the heat load margin that the candidate node can bear. The system uses a piecewise linear function to normalize the temperature margin. When the margin is higher than the preset safety interval, the highest scoring weight is given; when the margin is in the critical interval, it is converted proportionally; when the margin is lower than the lowest threshold, the candidate qualification is directly excluded. This scoring mechanism preferentially screens the nodes with the greatest heat dissipation potential in the temperature gradient distribution to provide thermal stability guarantee for high-heat task migration.

[0099] The generation of the heat dissipation scoring item depends on the dynamic comparison between the real-time heat dissipation capacity margin of the candidate node and the preset heat dissipation demand threshold. The heat dissipation capacity margin is calculated through a dynamic heat dissipation model, specifically the real-time heat dissipation capacity parameter minus the preset heat dissipation demand threshold. Among them, the real-time heat dissipation capacity parameter is obtained based on the product of the coolant flow rate, the specific heat capacity physical property parameter, and the fouling attenuation coefficient of the liquid cooling plate. The fouling attenuation coefficient is periodically corrected according to the exponential decay curve trained by historical operation and maintenance data. The heat dissipation scoring item uses a logarithmic function to perform a non-linear mapping on the margin value. When the margin exceeds the benchmark demand, the scoring growth rate slows down, avoiding overemphasis on a single heat dissipation index and ignoring other constraint conditions.

[0100] The comprehensive score of candidate nodes is obtained by weighted summation of the temperature score item and the heat dissipation score item through a preset weight ratio. The system defaults that the weight ratio of the temperature score item is 60%, and the weight ratio of the heat dissipation score item is 40%. This weight allocation scheme has been verified as the optimal solution set through Monte Carlo simulation. When forming the candidate node set, the system first filters the nodes that meet the hop count constraint condition, and then sorts them in descending order according to the comprehensive score. During the score sorting process, a stability verification mechanism is introduced. When the score difference of three consecutive candidate nodes is less than the preset tolerance threshold, the dynamic adjustment process of the weight coefficient is automatically triggered to ensure the discrimination and decision reliability of the score result. The finally generated candidate node score set provides a basis for priority sorting for migration target selection, realizing the collaborative optimization of heat dissipation resources and constraints.

[0101] In some instances, the target migration node is selected according to the sorting result of the candidate node score set, including:

[0102] Based on the sorting result of the candidate node score set, determine the candidate node with the highest score;

[0103] Verify whether the candidate node with the highest score meets the preset network bandwidth constraint;

[0104] When the preset network bandwidth constraint is met, determine the candidate node with the highest score as the target migration node;

[0105] When the preset network bandwidth constraint is not met, select the sub-optimal candidate nodes in descending order of score from the candidate node score set for verification until the target migration node that meets the preset network bandwidth constraint is determined.

[0106] Exemplarily, in the implementation process of selecting the target migration node, first lock the candidate node with the highest score according to the descending order result of the candidate node score set. The score set is generated by a multi-dimensional weight algorithm, in which the weight ratio of the temperature margin score item is 60%, and the weight ratio of the heat dissipation capacity margin score item is 40%. This weight allocation scheme has been verified as the optimal solution set through Monte Carlo simulation. A stability verification mechanism is introduced during the score sorting process. When the score difference between adjacent candidate nodes is less than the preset tolerance threshold, the dynamic weight coefficient adjustment process is triggered to ensure the discrimination and objectivity of the sorting result. The determination of the node with the highest score comprehensively considers the heat dissipation resource margin, temperature gradient distribution, and network topology proximity, providing an initial preferred target for the migration operation.

[0107] After selecting the candidate node with the highest score, the system executes the preset network bandwidth constraint verification process. The network bandwidth constraint conditions include the minimum guaranteed bandwidth and end-to-end transmission delay threshold of the migration dedicated channel, such as 10G bit / s bandwidth and delay not exceeding 20ms. During the verification process, the current bandwidth occupancy rate and available channel capacity of the target node are detected in real time through the software-defined network controller, and the possible fluctuations in the migration process are predicted based on historical traffic characteristics. The bandwidth constraint verification adopts a double verification mechanism. First, the coarse-grained data is obtained by polling the port counter based on the simple network management protocol, and then the fine-grained traffic analysis is performed through the sFlow sampler to ensure the accuracy and real-time performance of the verification results.

[0108] When the candidate node with the highest score meets the preset network bandwidth constraint, the system determines it as the target migration node and triggers the resource reservation instruction. The resource reservation process allocates a dedicated bandwidth channel for the migration operation through the software-defined network controller, sets the accelerated forwarding priority identifier, and ensures that the migration data flow obtains service quality assurance. The bandwidth capacity of the dedicated channel is not less than 120% of the preset demand threshold, and the reservation duration covers the entire life cycle of the migration. At the same time, the system pre-allocates the virtualized resource pool at the target node and establishes a shadow process synchronized with the business logic of the source node. The memory pre-loading mechanism is used to reduce the service interruption time of the migration process and achieve business continuity assurance.

[0109] If the candidate node with the highest score does not meet the network bandwidth constraint, the system selects the suboptimal candidate node in descending order of score for iterative verification. The suboptimal node verification process introduces a bandwidth preemption strategy, which allows temporary occupation of non-critical business bandwidth quotas, and the preemption ratio does not exceed 15% of the total bandwidth. During the verification process, the network jitter parameters are monitored in real time. When the jitter value of three consecutive samples exceeds 500 microseconds, the path rerouting mechanism is automatically triggered. If no node that meets the constraints is found after traversing the candidate node set, the system executes an emergency strategy, including starting the cross-computer room migration process and linking the aerosol cooling system to enhance the heat dissipation capacity, while forcibly enabling the transport layer security protocol encryption channel to ensure the reliable execution of the migration operation within the extended topology.

[0110] In some examples, performing a hot migration operation on a target migration node includes:

[0111] A dedicated bandwidth channel is reserved for the target migration node through a software-defined network controller, and an accelerated forwarding priority is assigned to the dedicated bandwidth channel;

[0112] Pre-allocate virtualized resources on the target migration node and establish a shadow process that is synchronized with the business logic of the source node;

[0113] Migrate service status data to the target migration node based on the incremental synchronization algorithm;

[0114] After the migration is completed, switch the service request to the target migration node.

[0115] Exemplarily, the execution of the live migration operation first reserves a dedicated bandwidth channel for the target migration node through the software-defined network controller and assigns an accelerated forwarding priority to this channel. The software-defined network controller implements the network slicing function based on the OpenFlow protocol, assigns an independent virtual channel for the migration data stream, with a bandwidth capacity of 10Gbps ± 0.5Gbps, and the priority is marked as the EF (accelerated forwarding) traffic class (DSCP = 46). By real-time monitoring the network traffic status, dynamically adjust the channel bandwidth occupancy rate not exceeding 15%, ensuring the real-time and stability of the migration data transmission, and at the same time avoiding interference with existing services.

[0116] Pre-allocate virtualized resources on the target migration node, including computing resources, memory resources, and storage resources, and establish a shadow process that synchronizes with the source node's business logic through hardware virtualization technology (such as Intel VT-x). The shadow process loads the process state image of the source node, but does not allocate actual computing resources for the time being, and is only used to receive and cache incremental data. The resource pre-allocation process adopts a memory locking mechanism (mlock system call) to prevent latency caused by page swapping, ensuring that the service interruption time during migration ≤ 50ms, meeting the reliability requirements.

[0117] The migration of service status data is implemented based on the incremental synchronization algorithm. Specifically, the xdelta algorithm is used to perform differential comparison and compressed transmission on the memory state and storage data. Each synchronization uses 256KB as the data block unit, and the CRC-32 checksum is used to ensure data integrity. When the checksum fails, a three-time retransmission mechanism is triggered. During the migration process, the shadow processes of the source node and the target node process requests in parallel, and achieve sub-second status synchronization (typical latency 300 - 800ms) through RDMA (Remote Direct Memory Access) technology. At the same time, the TCP / IP header compression technology is used to reduce protocol overhead, and the compression rate ≥ 12%.

[0118] After the migration is completed, the system seamlessly switches the service request to the target migration node. The switching process updates the traffic routing table through the software-defined network controller, redirects the original request pointing to the source node to the target node, and releases the virtualized resources of the source node. During the switching period, the system real-time monitors the service response time and error rate of the target node. If the request response time exceeds 20ms or the error rate ≥ 0.1% for three consecutive times, the rollback mechanism is triggered and the system is switched back to the source node. Finally, the target node takes over all service loads, and the source node enters the low-power standby state, realizing the dual optimization of heat dissipation resource dynamic adaptation and service continuity.

[0119] In some instances, after performing the live migration operation on the target migration node, it further includes:

[0120] Monitor the temperature change rate of the target migration node and the coolant flow state parameters in real time;

[0121] Based on the model predictive control framework, dynamically adjust the coolant flow rate and the operating parameters of the heat exchanger so that the temperature of the target migration node does not exceed the second preset temperature threshold.

[0122] Exemplarily, after the heat migration operation is completed, the system monitors the temperature change rate in real time through a high-precision temperature sensor deployed on the target migration node. Specifically, a PT1000 platinum resistance sensor is used, with a measurement accuracy of ±0.1°C, and the sampling frequency is synchronized with the heat migration control period (for example, once every 5 seconds). The coolant flow state parameters are collected by a Coriolis mass flowmeter and a capacitive liquid level sensor, including the coolant flow rate (0.5 - 2.5 L / min), the microchannel Reynolds number, and the temperature difference between the inlet and outlet of the liquid cooling plate (ΔT ≤ 6.2°C). The flow state parameters are calibrated in real time through a preset optimization curve of a fluid mechanics simulation model (such as the ANSYS Fluent k-ε model) to ensure data accuracy and provide reliable input for dynamic adjustment.

[0123] Based on the model predictive control framework, the system constructs a multi-objective optimization function with a preset control period (for example, 5 seconds). The objectives are to minimize the power usage effectiveness (PUE) index, the highest node temperature, and the migration cost. The control variables include the coolant flow rate, the rotational speed of the heat exchanger fan (500 - 2000 rpm), and the microchannel flow state parameters of the liquid cooling plate. The constraint conditions include that the temperature of the target node does not exceed the second preset temperature threshold (for example, 85°C), the coolant flow rate mutation rate ≤ 20% / s, and the network bandwidth occupancy rate ≤ 15%. During the optimization process, the rolling horizon algorithm is used to solve the optimal solution in each control period, generating a flow rate adjustment instruction and a fan rotational speed instruction.

[0124] During the dynamic adjustment process, the system adjusts the coolant flow rate in real time according to the temperature change rate (dT / dt). When it is detected that the temperature rise rate exceeds 2°C / s, the flow rate is increased to the preset upper limit (2.5 L / min) through closed-loop PID control, while restricting the flow rate mutation rate to avoid the water hammer effect (pressure fluctuation ≤ 10 kPa). The rotational speed of the heat exchanger fan is dynamically adjusted based on the temperature difference between the inlet and outlet. When ΔT ≥ 5°C, the rotational speed is increased to 2000 rpm to enhance the heat dissipation efficiency. The gradient density microchannel structure arranged in a Fibonacci pattern optimizes the fluid distribution, and together with the redundant pumping system (switching time < 200 ms), it ensures the stability of the flow rate adjustment.

[0125] When the model predictive control fails to solve or the delay exceeds 2 seconds, the system automatically switches to the preset PID backup control strategy. The parameters of the backup strategy are tuned based on historical operation data. For example, the proportional coefficient Kp = 0.5, the integral coefficient Ki = 0.1, and the derivative coefficient Kd = 0.05. The coolant flow rate adjustment instruction is sent to the Grundfos CR series pump through the RS485 communication protocol for execution. If the temperature of the target node continues to approach the second preset temperature threshold, the aerosol cooling mode is triggered, and the fluorinated liquid is atomized and sprayed through the titanium alloy sintered capillary blocking device to quickly absorb the residual heat. At the same time, the system records the adjustment log and updates the parameters of the digital twin model to achieve the closed-loop optimization of the heat dissipation strategy and long-term reliability guarantee.

[0126] In some instances, it also includes:

[0127] In the heat migration operation, an improved CRIU (Checkpoint / Restore in Userspace) tool is used to achieve sub-second service state preservation and restoration, and the state preservation delay is controlled within the range of 300 - 800 ms. By integrating an incremental synchronization algorithm (such as the xdelta algorithm), the memory and storage states are compared for differences and compressed and transmitted in units of 256 KB data blocks. The CRC-32 check is used to ensure data integrity, and a three-time retransmission mechanism is triggered when the check fails. Virtualized resources are pre-allocated on the target node and a shadow process is established. This process preloads the service logic image but does not occupy actual computing resources. The parallel processing and state synchronization between the source node and the target node are achieved through RDMA (Remote Direct Memory Access) technology. The dedicated migration channel reserves 10% bandwidth through the SDN controller (QoS guaranteed bandwidth 10 ± 0.5 Gbps, DSCP = 46). Combined with hardware virtualization technology, the memory pages are locked (mlock system call), and the service interruption time is compressed to within 50 ms to meet the SLA requirements of the cloud mobile phone service.

[0128] The liquid cooling module adopts a main / backup dual-pump redundant design. The pump model is Grundfos CR15-6. The redundant switching time is <200 ms, the flow rate control range is 0.5 - 2.5 L / min, and the steady-state error is ±0.1 L / min. Each rack is configured with an independent closed-loop circuit, which is linked with the intelligent control module through the RS485 communication protocol to receive the flow rate adjustment instruction in real time. When the main pump fails, the backup pump is automatically activated based on the monitoring signal of the capacitive liquid level sensor (response time <50 ms), and at the same time, it switches to the PID backup control mode (proportional coefficient Kp = 0.5, integral coefficient Ki = 0.1, derivative coefficient Kd = 0.05) to ensure that the coolant flow rate mutation rate ≤20% / s under emergency conditions and avoid the pipeline stress fluctuation caused by the water hammer effect (pressure fluctuation <10 kPa). The dual-pump system, the plate heat exchanger (made of 316L stainless steel, heat transfer efficiency ≥85%), and the external cooling tower form a two-stage heat dissipation architecture, making the coolant return water temperature stable at 30 ± 2 °C (under the condition of wet bulb temperature of 28 °C).

[0129] The microchannel layer of the liquid cooling plate is processed with copper microchannels by laser etching technology. The channel width is 50 μm (tolerance ±2 μm), the depth is 300 μm, and the surface roughness Ra ≤ 0.8 μm. The microchannels expand outward from the center according to the Fibonacci sequence law. The center reference spacing is 0.5 mm, and the edge expands to 0.75 mm. The ratio of the channel spacing increment to the previous spacing follows the golden ratio. Through ANSYS Fluent fluid mechanics simulation verification, in the flow rate range of 0.5 - 2.5 L / min, the Reynolds number fluctuation range of the Fibonacci arrangement scheme is reduced to [500, 1500] (the traditional uniform arrangement is [400, 2000]), the pressure drop is reduced by 15%, and the inlet and outlet temperature difference ΔT ≤ 6.2 °C (measurement error ±0.3 °C). The design of this application reduces flow instability by optimizing the flow state distribution, and combines a 0.5 mm aluminum nitride ceramic substrate (thermal conductivity ≥170 W / m·K) with an active metal brazing process (vacuum welding at 780 - 820 °C) to achieve efficient heat conduction and structural reliability.

[0130] A titanium alloy sintered capillary structure blocking device is set at the outlet end of the liquid cooling plate, which is installed 2 - 3 mm away from the outlet end face (tolerance ±0.1 mm), with a porosity of 35% - 40% and a capillary pressure ≥2.5 kPa. This device blocks the reverse flow of the coolant caused by the siphon effect through the chemical compatibility of titanium alloy and fluorinated liquid (such as 3M Novec 7100). When the coolant flow rate exceeds the limit (>2.5 L / min) or the leakage rate >50 ml / min, it automatically switches to the aerosol cooling mode, and quickly absorbs the residual heat through the atomizing spray function of the titanium alloy capillary structure. In addition, the anti-freezing protection mechanism injects propylene glycol-based additives (volume ratio ≤5%) when the coolant temperature <0 °C to avoid fluid solidification under low-temperature conditions and ensure the all-weather operation stability of the system.

[0131] Please refer to Figure 2 , which is a schematic structural diagram of a cloud mobile phone server hot migration device provided by an embodiment of the present application, including:

[0132] A node data acquisition unit 21, configured to acquire temperature data, load data, and real-time heat dissipation capacity parameters of a server node;

[0133] A thermal state evaluation unit 22, which determines a thermal state evaluation value based on the temperature data, load data, and real-time heat dissipation capacity parameters;

[0134] A candidate node determination unit 23, configured to generate a candidate node score set based on a multi-dimensional weight algorithm when the thermal state evaluation value is greater than a preset evaluation threshold;

[0135] A migration node selection unit 24, configured to select a target migration node according to the sorting result of the candidate node score set;

[0136] A hot migration operation unit 25, configured to perform a hot migration operation on the target migration node.

[0137] Please refer to Figure 3 , an embodiment of the present application further provides an electronic device 300, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements the steps of any method for hot migration of a cloud mobile phone server.

[0138] Since the electronic device introduced in this embodiment is the device used to implement a cloud mobile phone server hot migration device in an embodiment of the present application, based on the method introduced in an embodiment of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in an embodiment of the present application will not be described in detail here. As long as the device used by those skilled in the art to implement the method in an embodiment of the present application belongs to the scope protected by the present application.

[0139] In a specific implementation process, when the computer program 311 is executed by the processor, it can implement any implementation manner in the corresponding embodiment of the first aspect.

[0140] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] Those skilled in the art should understand that the embodiments of the present application may provide a method, a system, or a computer program product. Therefore, the present application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media that contain computer-readable program code.

[0142] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0145] The embodiments of the present application also provide a computer program product, which includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute Figure 1 the process of a cloud mobile phone server hot migration method corresponding to the embodiment.

[0146] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium, an optical medium, or a semiconductor medium, etc.

[0147] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0148] In several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0149] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0150] In addition, each functional unit in various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of a hardware and / or software functional unit.

[0151] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device to execute all or part of the steps of the methods of various embodiments of this application.

[0152] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

[0153] Although the preferred embodiments of this specification have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of this specification.

[0154] Obviously, those skilled in the art can make various changes and deformations to this specification without departing from the spirit and scope of this specification. Thus, if these modifications and deformations of this specification fall within the scope of the claims of this specification and their equivalent technologies, this specification is also intended to include these modifications and deformations.

Claims

1. A method for hot migration of cloud mobile phone servers, characterized in that, Including: Obtain the temperature data, load data, and real-time heat dissipation capacity parameters of the server node; Determine the thermal state evaluation value based on the temperature data, the load data, and the real-time heat dissipation capacity parameters; When the thermal state evaluation value is greater than the preset evaluation threshold, generate a candidate node score set based on the multi-dimensional weight algorithm; Select the target migration node according to the sorting result of the candidate node score set; Perform a hot migration operation on the target migration node.

2. The method according to claim 1, characterized in that The obtaining of the temperature data, load data, and real-time heat dissipation capacity parameters of the server node includes: Collect the temperature data of the server node through a temperature sensor, where the temperature data includes the temperature data of the first computing unit and the temperature data of the second computing unit; Obtain the load data of the server node through a hardware performance counter, where the load data includes the load data of the first computing unit, the load data of the second computing unit, and the communication bandwidth load data; Based on the coolant flow rate, coolant physical property parameters, and liquid cooling plate fouling attenuation coefficient, calculate the real-time heat dissipation capacity parameters of the server node through a dynamic heat dissipation model, where the liquid cooling plate fouling attenuation coefficient is determined based on the attenuation curve trained from historical operation and maintenance data.

3. The method according to claim 2, wherein The determining of the thermal state evaluation value based on the temperature data, the load data, and the real-time heat dissipation capacity parameters includes: Perform weight assignment on the temperature data of the first computing unit, the temperature data of the second computing unit, the load data of the first computing unit, the load data of the second computing unit, and the communication bandwidth load data respectively to generate dynamic weight coefficients; Based on the dynamic weight coefficients, perform weighted fusion on the temperature data and the load data to obtain an initial heat load value; Determine the heat dissipation capacity margin based on the real-time heat dissipation capacity parameters and the preset heat dissipation requirement threshold; Determine the thermal state evaluation value based on the initial heat load value and the heat dissipation capacity margin.

4. The method according to claim 1, wherein The generating of the candidate node score set based on the multi-dimensional weight algorithm when the thermal state evaluation value is greater than the preset evaluation threshold includes: When the thermal state evaluation value is greater than the preset evaluation threshold, perform the following steps: Based on the network topology of the server cluster, determine the hop count constraint conditions of the candidate nodes, where the hop count constraint conditions include a first hop count threshold and a second hop count threshold; Calculate the temperature score item based on the temperature margin of the candidate node and the first preset temperature threshold; Calculate the heat dissipation score item according to the heat dissipation capacity margin of the candidate node and the preset heat dissipation requirement threshold; Perform weighted summation of the temperature score item and the heat dissipation score item through a preset weight ratio to obtain the comprehensive score of the candidate node; Screen the candidate nodes that meet the hop count constraint conditions from the candidate nodes to form a candidate node set; Perform a descending order sorting on the candidate node set based on the comprehensive score of the candidate node to generate a sorted candidate node score set.

5. The method according to claim 1, wherein The selecting of the target migration node according to the sorting result of the candidate node score set includes: Based on the sorting result of the candidate node score set, determine the candidate node with the highest score; Verify whether the candidate node with the highest score meets the preset network bandwidth constraint; When the preset network bandwidth constraint is met, determine the candidate node with the highest score as the target migration node; When the preset network bandwidth constraint is not met, sequentially select sub-optimal candidate nodes from the candidate node score set in descending order of scores for verification until a target migration node that meets the preset network bandwidth constraint is determined.

6. The method according to claim 1, wherein The performing a live migration operation on the target migration node includes: Reserve a dedicated bandwidth channel for the target migration node through a software-defined network controller, and allocate an accelerated forwarding priority to the dedicated bandwidth channel; Pre-allocate virtualized resources to the target migration node and establish a shadow process synchronized with the service logic of the source node; Migrate the service state data to the target migration node based on the incremental synchronization algorithm; After the migration is completed, switch the service request to the target migration node.

7. The method according to claim 1, characterized in that, After performing the live migration operation on the target migration node, it further includes: Real-time monitor the temperature change rate and coolant flow state parameters of the target migration node; Dynamically adjust the coolant flow rate and heat exchanger operation parameters based on the model predictive control framework so that the temperature of the target migration node does not exceed the second preset temperature threshold.

8. A cloud mobile phone server hot migration device, characterized in that, It includes: A node data acquisition unit for acquiring temperature data, load data, and real-time heat dissipation capacity parameters of a server node; A thermal state evaluation unit for determining a thermal state evaluation value based on the temperature data, the load data, and the real-time heat dissipation capacity parameters; A candidate node determination unit for generating a candidate node score set based on a multi-dimensional weight algorithm when the thermal state evaluation value is greater than a preset evaluation threshold; A migration node selection unit for selecting a target migration node according to the sorting result of the candidate node score set; A live migration operation unit for performing a live migration operation on the target migration node.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the cloud mobile phone server live migration method as described in any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program, when executed by the processor, implements the cloud mobile phone server live migration method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Load migration method and system for balancing energy efficiency of server cluster

    CN121037369A