Dynamic elastic resource expansion power station computing power server

By constructing a multi-module collaborative dynamic elastic resource expansion mechanism, the node and link status of hydropower scheduling tasks are collected and evaluated in real time. The scheduling strategy is dynamically adjusted using a deep learning model, which solves the problems of resource status perception delay and task recovery failure during hot migration, and improves the scheduling stability and security of hydropower station computing servers.

CN120821567BActive Publication Date: 2026-05-01SANXIA JINSHAJIANG YUNCHUAN HYDROPOWER DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SANXIA JINSHAJIANG YUNCHUAN HYDROPOWER DEV CO LTD
Filing Date
2025-07-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In multi-node heterogeneous computing environments, the hot migration process under the existing scheduling mechanism suffers from resource status perception delays, inaccurate target node status, and abnormal migration task context mapping, leading to task recovery failures. This is especially problematic in critical control links of hydropower stations or hydrological prediction links, causing scheduling response delays, data processing interruptions, or security policy gaps, resulting in serious consequences.

Method used

By constructing a multi-module collaborative dynamic elastic resource expansion mechanism, including a distributed state-aware acquisition module, a data preprocessing and structuring module, a feature extraction and risk quantification analysis module, a deep learning intelligent assessment module, and a dynamic risk control and cycle management module, node and link state data are collected in real time, key indicators are extracted, risk assessment is performed using a deep learning model, and scheduling strategies are dynamically adjusted to reduce the time sensitivity of migration tasks and the priority of resource allocation.

Benefits of technology

It enables intelligent control of the entire process of hydropower scheduling tasks during hot migration, improves the scheduling stability and resource adaptability of the system in high-concurrency and heterogeneous node environments, ensures task continuity and operational safety, and ensures the real-time performance of key hydropower scheduling tasks and the overall reliability of system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120821567B_ABST
    Figure CN120821567B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic elastic resource expansion power station computing power server, it is related to computing resource management technical field, including distributed state perception acquisition module, data pre-processing and structured construction module, feature extraction and risk quantification analysis module, deep learning intelligent evaluation module and dynamic risk regulation and rhythm management module:Distributed state perception acquisition module, in the process of water power dispatching task live migration, through the distributed state perception system deployed in the central dispatching control platform of power station computing power server, continuously, real-time acquisition the running state of each key node in migration link and link quality parameter data.The application realizes risk identification and intelligent regulation in the process of water power dispatching task live migration by the dynamic elastic resource expansion mechanism of multi-module cooperation, improves the perception ability to node state and link exception, discriminates migration loss risk in combination with deep model, and dynamically adjusts dispatching strategy, guarantees task continuity and operation safety.
Need to check novelty before this filing date? Find Prior Art

Description

A dynamic, elastic resource-expanding computing server for hydropower stations Technical Field

[0001] This invention relates to the field of computing resource management technology, specifically to a dynamic, elastic resource expansion computing server for hydropower stations. Background Technology

[0002] The "Dynamically Elastic Resource-Expanded Hydropower Station Computing Server" refers to a high-performance computing platform designed for intelligent scheduling and operation management scenarios in hydropower stations. It possesses the ability to dynamically adjust computing resources based on actual load demands and task changes. By integrating virtualization management, resource awareness, and scheduling optimization technologies, this server can automatically expand CPU, memory, storage, or concurrent node resources during sudden increases in workload or critical computationally intensive applications, achieving real-time replenishment of computing power. When the task load decreases or idles, it automatically releases redundant resources, improving resource utilization and reducing energy consumption. This computing server is particularly suitable for hydropower station environments requiring high reliability and elastic resource control capabilities. It can support complex computing tasks such as large-scale hydrological data analysis, unit operation prediction, and risk warning model training, providing real-time, efficient, and intelligent computing support for hydropower stations.

[0003] Existing technologies have the following shortcomings: In multi-node heterogeneous computing environments, to meet the high-concurrency and high-load operation requirements of real-time tasks in hydropower stations, a scheduling engine is typically used to dynamically manage computing resources. A "live migration" mechanism is used to migrate running tasks from heavily loaded nodes to less loaded, idle nodes without interrupting execution, thereby optimizing overall resource utilization and improving task processing efficiency. However, under existing scheduling mechanisms, the live migration process carries risks such as delayed resource status awareness, inaccurate target node status, and abnormal task context mapping. Specifically, the scheduling engine may make migration decisions based on lagging node resource information, mistakenly migrating tasks to nodes that are actually under high load or unavailable; or during the migration to the target node, task recovery may fail due to incorrect mounting or complete synchronization of task memory images, data handles, and session states. The aforementioned problems will directly lead to the instantaneous loss of tasks in the migration link. In particular, for high-priority tasks in critical control links or hydrological prediction links, it is very easy to cause scheduling response delays, data processing interruptions or security policy gaps, which may further lead to serious consequences such as scheduling failures, delayed risk warnings or abnormal operation of hydropower units.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a dynamic elastic resource expansion hydropower station computing server. Through a multi-module collaborative dynamic elastic resource expansion mechanism, it realizes risk identification and intelligent control during the hot migration process of hydropower scheduling tasks, improves the ability to perceive node status and link anomalies, combines deep model to identify migration loss risks, and dynamically adjusts scheduling strategies to ensure task continuity and operational safety, thereby solving the problems in the background technology mentioned above.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a dynamic elastic resource expansion hydropower station computing server, comprising a distributed state awareness acquisition module, a data preprocessing and structured construction module, a feature extraction and risk quantification analysis module, a deep learning intelligent assessment module, and a dynamic risk control and cycle time management module.

[0007] During the hot migration of hydropower dispatching tasks, the distributed state awareness acquisition module continuously and in real time collects the operating status and link quality parameter data of each key node in the migration link through the distributed state awareness system deployed on the central dispatching and control platform of the hydropower station's computing power server.

[0008] The data preprocessing and structuring module preprocesses the multi-dimensional raw data collected in real time. After preprocessing, it constructs a data set with unified field definitions and time index markers, and stores it in a structured format according to time sequence.

[0009] The feature extraction and risk quantification analysis module, based on the preprocessed structured dataset, uses feature engineering methods to extract key indicators that sensitively reflect the potential instantaneous loss risk of the task in the migration link, and performs comprehensive analysis on the extracted key indicators to quantify the severity of the instantaneous loss risk during the task migration process.

[0010] The deep learning intelligent evaluation module inputs key feature indicators that have undergone comprehensive analysis and risk quantification into a deep learning model that has been pre-trained based on historical data. The model performs multi-dimensional dynamic evaluation of the task migration link status and intelligently determines whether there is a risk of instantaneous loss of the current task.

[0011] The dynamic risk control and takt management module automatically triggers a dynamic risk control mechanism for migration operations when it identifies a risk of instantaneous loss of a task in the migration chain. This reduces the time sensitivity of migration tasks, extends the allowed execution time of migration operations, and dynamically adjusts the takt of task scheduling and the priority of resource allocation. It proactively reduces the frequency of high-frequency migration operations, thereby achieving "frequency reduction and risk avoidance".

[0012] Preferably, during the hot migration of hydropower dispatching tasks, the link status is collected through a distributed status awareness system deployed on the central dispatching and control platform of the hydropower station's computing power server. The specific steps are as follows:

[0013] Resource monitoring probes and communication quality acquisition components are deployed at key locations on each computing node and network link involved in hot migration to ensure independent collection of local resource status and link information;

[0014] Each sensing unit continuously collects data parameters from the node at a set frequency;

[0015] All collected operational status and link quality data are transmitted back to the central dispatch and control platform in real time via the internal data bus or monitoring channel, where they are uniformly received and partitioned by the dispatch engine.

[0016] Preferably, based on the preprocessed structured dataset, feature engineering methods are used to extract key indicators that sensitively reflect the potential instantaneous loss risk of tasks in the migration link. The extracted indicators include the degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node, and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node. The degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node are comprehensively analyzed under the detection window to generate cache paging interference indicators and mapping space misalignment indicators, respectively. The severity of the instantaneous loss risk during the task migration process is quantified by the cache paging interference indicators and mapping space misalignment indicators.

[0017] Preferably, the specific steps for generating a cache paging interference index by comprehensively analyzing the degree of access conflict between migration tasks and locally running tasks in the memory page table cache area under the detection window are as follows:

[0018] Conflict detection is performed on the access behavior of migration tasks and locally running tasks in the memory page table cache area of ​​hot migration nodes. The access trajectories of migration tasks and local tasks to page table cache resources are recorded, and the number of overlapping accesses within the same memory page mapping area is calculated. The frequency of conflicts between migration tasks and local tasks during cache access is defined as the conflict frequency factor. The expression for calculating the conflict frequency factor is as follows:

[0019] ,

[0020] In the formula, N conflict N is the number of conflicts when migration tasks and local tasks access the same cache page table entry within the detection window. total It is the sum of the cumulative number of accesses to the current memory page mapping region in the page table cache by the migration task and the local task, C f It is the conflict frequency factor;

[0021] Weighting factors are set to determine the impact of different memory page mapping region accesses on the migration task, and then the local conflict degree is calculated using the following formula: L conflict =C f ·, where w is the regional weighting factor, L conflict It refers to the degree of local conflict;

[0022] Based on the obtained local conflict degree L of each memory region conflict Calculate the overall cache paging interference index. The cache paging interference index is used to characterize the intensity of resource interference experienced by the migration task in the entire page table cache space. The calculation formula is as follows:

[0023] ,

[0024] In the formula, PPI is the cache paging interference metric. It is the local conflict degree of the i-th memory page mapping region. Let be the number of accesses to the i-th memory page mapping region, N be the total number of memory page mapping regions, and ln be a logarithmic function.

[0025] Preferably, the specific steps for generating a mapping space misalignment index by comprehensively analyzing the address offset fluctuation trend of the task's virtual memory space mapped to the target node under the detection window are as follows:

[0026] During task hot migration, obtain the virtual address segment sequence V of the current task on the source node. s and the address segment sequence V after remapping in the target node d ,in, It is the j-th virtual address in the source node, and n is the total number of virtual addresses. Let j be the j-th virtual address in the target node. For each pair of mapped addresses, calculate the structure offset using the following expression:

[0027] ,

[0028] In the formula, h(·) is the address hash mapping function, Δ j It is the address mapping offset;

[0029] Based on the obtained address mapping offset sequence, a mapping space misalignment index is generated using the following formula:

[0030] ,

[0031] In the formula, Δ j-1 It is the address mapping offset of the previous virtual address unit, min(Δj Δ j-1 ) is the minimum value of the address mapping offset between the j-th virtual address unit and the (j-1)-th virtual address unit, and MDI is the mapping space misalignment index.

[0032] Preferably, the cache paging interference index and mapping space misalignment index, which have undergone comprehensive analysis and risk quantification, are input into a deep learning model that has been pre-trained based on historical data. The model generates a task instantaneous loss risk coefficient, and the task migration link status is dynamically evaluated in multiple dimensions based on the task instantaneous loss risk coefficient, so as to intelligently determine whether there is an instantaneous loss risk for the current task.

[0033] Preferably, the instantaneous task loss risk coefficient generated by a deep learning model pre-trained based on historical data for multi-dimensional dynamic evaluation of the task migration link status is compared and analyzed with a pre-set reference threshold for the instantaneous task loss risk coefficient to intelligently determine whether the current task has an instantaneous loss risk. The judgment logic is as follows:

[0034] If the instantaneous loss risk coefficient of a task is greater than the preset reference threshold for instantaneous loss risk coefficient of a task, then the current task is determined to have an instantaneous loss risk; if the instantaneous loss risk coefficient of a task is less than or equal to the preset reference threshold for instantaneous loss risk coefficient of a task, then the current task is determined not to have an instantaneous loss risk.

[0035] Preferably, when a risk of momentary loss of a task is detected in the migration chain, a dynamic risk control mechanism for the migration operation is automatically triggered to reduce the time sensitivity of the migration task, extend the allowed execution time of the migration operation, and dynamically adjust the task scheduling takt time and resource allocation priority to proactively reduce the frequency of high-frequency migration operations, thereby achieving "frequency reduction and risk avoidance". The specific steps are as follows:

[0036] Once a risk of momentary task loss is identified during the migration process, the time sensitivity attribute of the scheduled task is first adjusted. By calculating the extended execution time multiplier for the expanded migration operation, the task migration time window is dynamically adjusted. The expression for calculating the extended execution time multiplier for the migration operation is as follows:

[0037] ,

[0038] In the formula, λ1 is the time window expansion sensitivity coefficient, and R th It is the reference threshold for the instantaneous loss risk coefficient, α w This is the multiple by which the migration operation is allowed to extend its execution time.

[0039] The allowed execution time extension factor α based on the acquired migration operation. wThe clock speed reduction rate of the scheduled task in the scheduling system is calculated, and the scheduling trigger frequency of the task is dynamically adjusted. The calculation expression is as follows:

[0040] ,

[0041] In the formula, λ2 is the clock attenuation gain coefficient, and δ max It is the upper limit of the cycle deceleration ratio, δ p It is the deceleration rate of the scheduling cycle;

[0042] Based on the instantaneous loss risk coefficient R d Migration operation allows for a multiplier α to extend execution time. w With the deceleration ratio δ of the scheduling cycle p The scheduling resource compression priority coefficient of the target task is jointly calculated and used to adjust the sorting priority in the global resource contention queue. The calculation expression is as follows:

[0043] ,

[0044] In the formula, λ3 is the nonlinear curvature factor of the compression mapping, e is the natural base, and π is the coefficient of performance. r It is the resource compression priority coefficient.

[0045] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0046] This invention achieves intelligent full-process control over the potential instantaneous loss risk of hydropower scheduling tasks during hot migration by constructing a dynamic and elastic resource expansion mechanism with multi-module collaborative capabilities. This effectively improves the system's scheduling stability and resource adaptability in high-concurrency, heterogeneous node environments. The system acquires the real-time operating status and link quality parameters of key nodes through a distributed state awareness module. Combined with data preprocessing and structured construction operations, it ensures the temporal consistency and analytical adaptability of multi-source data. Based on this, it extracts and quantifies key indicators that accurately reflect migration risks, such as the degree of cache paging interference and the trend of mapping space misalignment. A deep learning model is used to intelligently determine the migration link status of the current task, identifying potential loss risks during migration in advance. Finally, based on the risk identification results, the scheduling system dynamically adjusts the migration duration, scheduling rhythm, and task priority, proactively avoiding unstable resource states and achieving flexible control of the scheduling rhythm. This significantly improves the safety and continuity of the task migration process, ensuring the real-time performance of critical hydropower scheduling tasks and the overall reliability of the system operation. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0048] Figure 1 is a schematic diagram of a dynamic elastic resource expansion hydropower station computing server according to the present invention. Detailed Implementation

[0049] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0050] This invention provides a dynamic, elastic resource expansion computing server for hydropower stations, as shown in Figure 1, comprising a distributed state-aware acquisition module, a data preprocessing and structuring module, a feature extraction and risk quantification analysis module, a deep learning intelligent assessment module, and a dynamic risk control and cycle time management module.

[0051] During the hot migration of hydropower dispatching tasks, the distributed state awareness acquisition module continuously and in real time collects the operating status and link quality parameter data of each key node in the migration link through the distributed state awareness system deployed on the central dispatching and control platform of the hydropower station's computing power server.

[0052] Throughout the hydropower station's computing server system, particularly within the central control platform responsible for task scheduling and resource allocation, a "distributed state awareness system" is embedded. This system comprises multiple awareness components distributed across different computing nodes, communication links, and resource interfaces. These components continuously monitor operational status information closely related to the hot migration process (such as CPU utilization, memory usage, and task execution load) and migration link quality parameters (such as network latency, packet loss rate, and bandwidth utilization). This system's deployment enables the system to monitor the health status and resource change trends of nodes and links in real time before, during, and after task migration, providing a data foundation and decision support for deciding whether to perform migration and how to mitigate migration risks. Its main functions include three points: first, sensing node operating pressure and performance status to avoid migrating tasks to overloaded nodes; second, detecting link stability to identify potential transmission failure risks in advance; and third, providing high-precision real-time data input for subsequent migration strategy decisions, risk assessments, and control mechanisms. This mechanism is an indispensable awareness foundation for ensuring the safety and continuity of the hot migration process in the hydropower dispatching system.

[0053] During task hot migration, the system continuously collects operational data related to task execution and migration link stability at millisecond or second-level frequencies through status awareness modules embedded in each node and communication channel. This ensures that the scheduling and control platform always has a grasp of the current system's operational status. Specifically, the collected data includes resource operation status parameters such as CPU utilization, memory usage, disk I / O rate, and task execution load of each node, as well as communication quality parameters at the link level such as network transmission latency, packet loss rate, bandwidth utilization, link jitter, and connection stability. The purpose of this data collection is to provide real-time data support for key decisions such as whether a task is suitable for migration, whether there are link bottlenecks during migration, and whether the target node has the capacity to handle the load. This effectively avoids risks such as task migration failure, context mounting errors, or task interruption caused by link instability or node overload, improving the stability, intelligence, and migration success rate of the entire scheduling system.

[0054] The purpose of this process is to comprehensively and in real-time capture operational status information directly related to the link status during the task hot migration process, so as to ensure that subsequent analysis and evaluation can be based on high-quality and comprehensive raw data.

[0055] During the hot migration of hydropower dispatching tasks, the link status is collected through a distributed status awareness system deployed on the central dispatching and control platform of the hydropower station's computing server. The specific steps are as follows:

[0056] The first step is to deploy sensing nodes by setting up resource monitoring probes and communication quality acquisition components at key locations on each computing node and network link involved in hot migration, to ensure that local resource status and link information can be collected independently.

[0057] The second step is to collect status parameters in real time. Each sensing unit continuously collects key parameters such as CPU utilization, memory usage, load changes, task running status, transmission delay, packet loss rate, bandwidth usage, and link fluctuations in the network link at a set frequency.

[0058] The third step involves data reporting and synchronization with the scheduling platform. All collected operational status and link quality data are transmitted back to the central scheduling and control platform in real time via the internal data bus or monitoring channel. The scheduling engine receives and records this data in designated areas, providing timely and comprehensive operational support information for subsequent task hot migration strategy decisions. This process enables dynamic monitoring of the migration environment throughout the entire process and is a fundamental step in ensuring the stable migration capability of the scheduling system.

[0059] The data preprocessing and structuring module preprocesses the multi-dimensional raw data collected in real time. After preprocessing, it constructs a data set with unified field definitions and time index markers, and stores it in a structured format according to time sequence.

[0060] Preprocessing of the real-time acquired multi-dimensional raw data mainly includes the following aspects: First, outlier removal: Abnormal data caused by acquisition failures, instantaneous node anomalies, or network jitter is identified and eliminated by setting threshold ranges, using sliding window statistics, or outlier detection algorithms (such as Z-score, IQR). Second, missing values ​​are filled: Data gaps caused by network latency or sensor fluctuations are repaired using methods such as linear interpolation, time series prediction, or neighbor mean imputation to maintain data continuity. Finally, data standardization and normalization are performed: Data of different dimensions and scales are mapped to a unified interval through Z-score standardization or Min-Max normalization to eliminate model bias caused by numerical differences. The core function of this series of operations is to improve the consistency, integrity, and comparability of the data, avoid interference from outlier data in feature extraction and model judgment, and ensure a high-precision and high-stability data foundation for subsequent migration risk analysis, indicator modeling, and deep learning evaluation processes.

[0061] After preprocessing, a data set with unified field definitions and time index markers is constructed. This means that after cleaning, completing, and standardizing the original data, various collected multi-dimensional parameters (such as CPU utilization, link latency, task load, etc.) are uniformly named and categorized according to preset field specifications. Simultaneously, each data sample is bound with an accurate timestamp, forming a standardized data set with a unified field structure and strict time series marking. Subsequently, the system stores this data set in a structured format (such as a time series database (TSDB) or an ordered log structure). Its core function is to achieve alignment and unification of different nodes and data types in the time domain, enabling subsequent feature extraction, state modeling, and risk trend analysis to be accurately compared and extrapolated within the same time dimension, improving the computational efficiency, data consistency, and model adaptability of the entire hot migration risk identification process.

[0062] The purpose of this process is to eliminate data noise and bias, establish an efficient and stable dataset, and make subsequent feature extraction and risk analysis more accurate and reliable, thereby reducing the error and volatility of subsequent model predictions.

[0063] The feature extraction and risk quantification analysis module, based on the preprocessed structured dataset, uses feature engineering methods to extract key indicators that sensitively reflect the potential instantaneous loss risk of the task in the migration link, and performs comprehensive analysis on the extracted key indicators to quantify the severity of the instantaneous loss risk during the task migration process.

[0064] Based on the preprocessed structured dataset, feature engineering methods are used to extract key indicators that sensitively reflect the potential instantaneous loss risk of tasks during the migration process. The extracted indicators include the degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node, and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node. The degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node, and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node are comprehensively analyzed under the detection window to generate cache paging interference indicators and mapping space misalignment indicators. The severity of the instantaneous loss risk during the task migration process is quantified by the cache paging interference indicators and mapping space misalignment indicators.

[0065] When a severe access conflict occurs between the migration task and the locally running task in the memory page table cache of a hot migration node, it clearly indicates a potential risk of transient loss of the hydropower scheduling task during the migration chain. This is because during the hot migration process, the migration task needs to quickly rebuild its virtual memory mapping relationship in the target node, including critical operations such as page table loading, TLB (Translation Lookaside Buffer) filling, and cache synchronization. If a locally running task is already frequently occupying page table cache resources at this time, it will cause frequent cache replacement, page table invalidation, and TLB conflicts, severely interfering with the state recovery process of the migration task. This can lead to incorrect context reconstruction, thread stack misalignment, or instruction anomalies, resulting in "transient loss" phenomena such as task mounting failure and state chain breakage. For high real-time and high-stability tasks like hydropower scheduling, a single hot migration interruption can break the scheduling instruction chain or cause the loss of hydrological control logic, resulting in serious operational deviations and potential risks. Therefore, page table cache conflicts are not only a system performance issue but also a high-risk signal during the migration of critical tasks, and should be used as an important underlying indicator for identifying potential loss risks.

[0066] The specific steps for generating a cache paging interference index by comprehensively analyzing the degree of access conflict between migration tasks and locally running tasks in the memory page table cache area of ​​a hot-migrating node under the detection window are as follows:

[0067] Conflict detection is performed on the access behavior of migration tasks and locally running tasks in the memory page table cache area of ​​hot migration nodes. The access trajectories of migration tasks and local tasks to page table cache resources are recorded, and the number of overlapping accesses within the same memory page mapping area is calculated. The frequency of conflicts between migration tasks and local tasks during cache access is defined as the conflict frequency factor. The expression for calculating the conflict frequency factor is as follows:

[0068] ,

[0069] In the formula, N conflictThis refers to the number of conflicts when migration tasks and local tasks access the same cache page table entry within the detection window. It is statistically determined by the cache access monitoring unit and is a direct manifestation of capturing resource sharing conflict behavior. N total It is the sum of the cumulative accesses of migration tasks and local tasks to the current memory page mapping region in the page table cache, reflecting the overall access activity level and providing a standardized reference benchmark for the degree of conflict. f It is the conflict frequency factor, which represents the proportion of access conflicts between migration tasks and local running tasks on page table cache resources within the detection window. The value is 0-1. The larger the factor is, the more intense the resource contention and the more serious the cache conflict. It is a risk precursor to the interruption or failure of the migration task status loading process.

[0070] The conflict frequency factor is used to characterize the intensity of resource contention for page table cache resources within the same time period.

[0071] Weighting factors are set to determine the impact of access to different memory page mapping regions on the migration task. These weighting factors are determined based on the activity level of the accessed region or the system scheduling priority, and then the local conflict degree is calculated using the following formula: L conflict =C f ·w, where w is the region weight factor, representing the importance weight of the memory region during task migration. It can be determined by the following factors: region access frequency (e.g., stack segment, shared library mapping segment), data type (e.g., state context page, IO cache page), mapping timing priority (i.e., whether it is a region that must be loaded in the early stages of migration). The larger the weight factor value, the greater the impact on the stable recovery of the migration task if the region is disturbed. The value ranges from 0 to 1. conflict It is the local conflict degree, which represents the overall impact intensity of resource conflicts caused by shared access behavior between migration tasks and local tasks in a certain memory page table region or cache segment. The larger the value, the more significant the interference risk that the region poses to the migration state recovery process.

[0072] This local conflict level is used to reflect the intensity of interference caused by cache resource contention in a specific page table region to the migration task, providing a quantitative basis for the calculation of subsequent interference risk indicators.

[0073] Based on the obtained local conflict degree L of each memory region conflict Calculate the overall cache paging interference index. The cache paging interference index is used to characterize the intensity of resource interference experienced by the migration task in the entire page table cache space. The calculation formula is as follows:

[0074] ,

[0075] In the formula, PPI is the cache paging interference index, which characterizes the overall interference intensity experienced by the migration task in the memory page table cache area of ​​the target node during hot migration. The value ranges from 0 to 1. The larger the value, the more likely the context reconstruction of the migration task is to be hindered by cache contention, the lower the migration success rate, and the higher the risk of instantaneous task loss. It is the local conflict degree of the i-th memory page mapping region. is the number of accesses to the i-th memory page mapping region, obtained through the system-level event counter or hardware cache hit / miss log, reflecting the access frequency of this region, i.e. the "active background" of conflict events. A high access frequency means that once a conflict occurs in this region, it will have a greater impact on system performance and task stability. N is the total number of memory page mapping regions, and ln is a logarithmic function.

[0076] The main purpose of introducing the logarithmic function in the above steps is to non-linearly amplify the access frequency of memory page mapping regions, thereby enhancing the ability to identify cache interference risks in high-activity regions. Specifically, when a cache region is accessed frequently, its access conflicts will have a much greater destructive impact on the task hot migration process than low-frequency regions. Directly using linear weights may lead to the average effect of high-access and low-access regions, reducing the model's sensitivity to "high-risk hotspots." By taking the logarithm of the access frequency, the contribution weight of medium- and high-frequency access regions to the final cache paging interference index can be effectively increased while compressing the fluctuations in extreme access values. This more accurately reflects their importance in causing instantaneous loss risks such as context mounting failure and cache reconstruction obstacles during migration. Therefore, the logarithmic function here is not only a mathematical transformation tool but also a key mechanism for achieving "risk amplification" of interference intensity and "robust expression" of data structures.

[0077] By introducing a logarithmic function, the ability to respond to the impact of conflicts in high-frequency access areas can be enhanced, and the interference level in critical areas during migration can be nonlinearly amplified, thereby more sensitively capturing the precursors of task state recovery interruptions.

[0078] The cache paging interference index is generated by comprehensively analyzing the degree of access conflict between the migration task and the locally running task in the memory page table cache area within the detection window. A higher index indicates a higher potential risk of transient data loss for the hydropower scheduling task during the migration process; conversely, a lower index value suggests a stable migration environment and a lower risk of task loss. This index is calculated by analyzing the degree of access conflict between the migration task and the locally running task in the memory page table cache area (such as TLB, page table mapping area, PageCache) within the detection window. When the migration task arrives at the target node and attempts to rebuild its memory structure, if the local task frequently accesses and occupies critical cache resources, it will cause an increase in TLB hit failure rate and intensify page table replacement, thus hindering the migration task from completing context loading and state mounting. This can easily lead to "transient data loss" phenomena such as recovery failure and execution interruption in the final stage of the migration.

[0079] The significant fluctuations in the address offset of the task's virtual memory space mapping on the target node clearly indicate a potential risk of momentary loss of the hydropower scheduling task during the migration process. During hot migration, the complete runtime state of the task depends on whether it can construct a virtual memory space layout on the target node that is highly consistent with the source node, particularly concerning thread stacks, shared memory segments, pointer reference structures, and the memory mapping relationship between kernel and user modes. When the mapped address deviates drastically in the target node, it indicates that the address references in the task context will experience a global or partial misalignment, which can easily lead to pointer resolution errors, memory access out-of-bounds errors, or page table failures during task recovery, resulting in task state recovery failure or logical collapse. Furthermore, hydropower scheduling tasks often have high real-time and state consistency requirements; any momentary mapping error can cause scheduling link breaks or prediction logic interruptions, ultimately leading to the momentary loss phenomenon of the task "appearing to have completed migration but actually failing to run" during the hot migration process. Therefore, drastic fluctuations in the mapping offset are a key early warning signal reflecting memory structure-level anomalies and possess a high degree of capability in characterizing migration risks.

[0080] The specific steps for generating a mapping space misalignment index by comprehensively analyzing the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node under the detection window are as follows:

[0081] During task hot migration, obtain the virtual address segment sequence V of the current task on the source node. s and the address segment sequence V after remapping in the target node d ,in, It is the j-th virtual address in the source node, and n is the total number of virtual addresses. Let j be the j-th virtual address in the target node. For each pair of mapped addresses, calculate the structure offset using the following expression:

[0082] ,

[0083] In the formula, This represents the mapping address in virtual memory of the j-th virtual address corresponding to the task in the target node after the task migration, used to... By comparing whether their structures are consistent, it is possible to determine whether there is an address misalignment or mapping drift. This indicates that when a task is running on the source node before hot migration, the j-th virtual address in its virtual address space (such as a page, a segment table entry, a stack address, etc.) provides a reference for constructing the original address structure before migration. h(·) is an address hash mapping function used to map different address structures to a unified and comparable logical identifier. Δ j It is the address mapping offset, which represents the offset of the j-th virtual address unit in the migration task in the address mapping structure of the target node relative to the source node.

[0084] Address hash mapping functions are used to convert virtual addresses into comparable, standardized logical representations. Their core function is to uniformly map the virtual address structures of different computing nodes to a normalized identifier space, enabling symmetry analysis and structural offset calculation of the address mapping state between the source and target nodes. In hot migration scenarios, due to differences in memory allocation strategies, page table organization methods, or segmented management mechanisms among different nodes, directly comparing virtual addresses often loses consistency. Therefore, hash mapping functions are introduced to extract structural features of addresses (such as page frame numbers, segment table indexes, cache slot numbers, etc.) to achieve equivalent reduction of addresses in the logical dimension. This function can be existing functions such as page number extraction functions (e.g., right-shifting the virtual address by page length), modulo operation mapping functions (e.g., address modulo a predefined number of page frames), or segment number / segment offset extraction functions based on page table structures. Through this mapping, the system can perform unified and unambiguous structural offset analysis of mapped addresses in different nodes, which is a key foundation for supporting mapping consistency measurement and misalignment index generation during migration.

[0085] The main purpose of this step is to generate a continuous sequence of address offsets {Δ1, Δ2, ..., Δ...} n This serves as the foundational data support for measuring the stability of the mapping structure during the task migration process.

[0086] Based on the obtained address mapping offset sequence, a mapping space misalignment index is generated to reflect the relative perturbation intensity of offset fluctuations between adjacent addresses in the virtual address mapping. The generation formula is as follows:

[0087] ,

[0088] In the formula, Δ j-1 It is the address mapping offset of the previous virtual address unit, min(Δ j Δ j-1 ) is the minimum address mapping offset between the j-th virtual address unit and the (j-1)-th virtual address unit. MDI is the Mapping Space Misalignment Index, which represents the quantitative index of the address offset fluctuation trend relative to the original mapping structure of the source node when the migrated task re-establishes the virtual memory space mapping in the target node during the task hot migration process. By detecting the mapping offset jump rate between adjacent virtual address segments, it reflects whether the task has phenomena such as mapping structure misalignment, continuity disruption or local disturbance aggregation in the target node. The value range is 0-1.

[0089] The MDI index value obtained through the above steps can be used to reflect whether there are drastic fluctuations in the virtual address mapping of the task in the target node. The larger the index value, the higher the degree of mapping discontinuity and structural misalignment, thus characterizing the greater the risk of instantaneous state loss of the task during hot migration.

[0090] The larger the address offset fluctuation trend of the task's virtual memory space mapped to the target node within the detection window, the higher the potential risk of transient loss in the migration link of the hydropower scheduling task. Conversely, the lower the risk, or even the less risk it may be considered. This indicator quantifies and evaluates the address offset and its fluctuation trend during the remapping of the virtual memory space of the migration task on the target node within the detection window. Essentially, it reflects whether the task can achieve a complete restoration of the original memory structure after migration. A high indicator value indicates a large offset or frequent jumps in the mapping, meaning that pointer references, page table structures, and data addresses in the task context may be corrupted. This can easily lead to incorrect loading of the migrated state, triggering access exceptions, interruptions, or logical crashes, thus constituting a typical "transient loss" risk. Conversely, when the mapping space misalignment indicator is at a low level, it indicates that the virtual address mapping is stable and highly consistent with the source node, and the task state can be smoothly restored, indicating that the migration process is safe.

[0091] The deep learning intelligent evaluation module inputs key feature indicators that have undergone comprehensive analysis and risk quantification into a deep learning model (such as LSTM, Transformer or CNN-LSTM hybrid network) that has been pre-trained based on historical data. The model performs multi-dimensional dynamic evaluation of the task migration link status and intelligently determines whether there is a risk of instantaneous loss of the current task.

[0092] The cache paging interference index and mapping space misalignment index, which have undergone comprehensive analysis and risk quantification, are input into a deep learning model that has been pre-trained based on historical data. The model generates a task instantaneous loss risk coefficient, and the task migration link status is dynamically evaluated in multiple dimensions based on the task instantaneous loss risk coefficient, so as to intelligently determine whether there is an instantaneous loss risk for the current task.

[0093] A deep learning model pre-trained on historical data refers to a system that, before the actual hot migration operation, has built a high-quality training dataset using a large amount of historical task migration process runtime status data, link quality parameters, and final migration results (such as success / failure, recovery time, state loss level, etc.). Deep learning algorithms are then used to learn the temporal characteristics, nonlinear variation patterns, and risk evolution modes within this dataset, ultimately resulting in a model capable of predicting task migration risks. This model may be based on architectures such as LSTM (Long Short-Term Memory), Transformer, or a CNN-LSTM hybrid structure to capture the time dependencies and nonlinear interaction patterns between multidimensional indicators during the migration process. During training, the model uses cache paging interference indicators, mapping space misalignment indicators, and other auxiliary indicators (such as link congestion indicators, scheduling delay indicators, etc.) as input features, combined with migration result labels, to perform supervised learning. This allows the model to understand under what combination of indicators or dynamic trends the task migration is more prone to instantaneous loss risk. Through a large number of training samples and multiple rounds of iterative optimization, the model gradually optimizes the weight parameters and has the ability to identify "high-risk migration states". It can predict potential failures in advance before the actual migration is triggered, and plays the role of "predictor" in the scheduling decision chain.

[0094] Once trained, the model can be run as an online evaluator after system deployment. When the system's real-time monitoring module extracts cache paging interference indicators and mapping space misalignment indicators during the current task migration process, it inputs these into the deep learning model. Based on historically learned feature weights and risk patterns, the model quickly outputs a quantitative result representing the current migration safety, namely the "task instantaneous loss risk coefficient." This coefficient reflects the probability risk level of failure or interruption faced by the current migration operation. Based on the magnitude, rate of change, and relative threshold of this coefficient, the system can achieve a multi-dimensional dynamic evaluation of the current migration state, assisting in determining whether the current state falls within an unmigratable risk window, whether delayed scheduling is necessary, and whether the migration target should be adjusted or risk avoidance mechanisms activated. Because this model has been trained and optimized on a large scale using historical task data, it possesses the ability to generalize to complex state patterns and can quickly capture potential loss risks under various subtle features such as link anomalies, resource conflicts, and address mismatches. It is a key supporting core for building an intelligent, predictive migration management and control system for hydropower scheduling tasks.

[0095] The instantaneous task loss risk coefficient generated by a deep learning model pre-trained based on historical data for multi-dimensional dynamic evaluation of the task migration link status is compared and analyzed with a pre-set reference threshold for the instantaneous task loss risk coefficient to intelligently determine whether the current task has an instantaneous loss risk. The judgment logic is as follows:

[0096] If the instantaneous loss risk coefficient of a task is greater than the preset reference threshold for instantaneous loss risk coefficient of a task, then the current task is determined to have an instantaneous loss risk; if the instantaneous loss risk coefficient of a task is less than or equal to the preset reference threshold for instantaneous loss risk coefficient of a task, then the current task is determined not to have an instantaneous loss risk.

[0097] The dynamic risk control and takt management module automatically triggers a dynamic risk control mechanism for migration operations when it identifies a risk of instantaneous task loss in the migration link. This reduces the time sensitivity of migration tasks, extends the allowed execution time of migration operations, and dynamically adjusts the takt of task scheduling and resource allocation priority. It proactively reduces the frequency of high-frequency migration operations, thereby achieving "frequency reduction and risk avoidance." That is, before the migration risk is eliminated, a more robust and low-impact scheduling strategy is used to avoid the risk of task loss and interruption caused by resource contention, link congestion, or node status fluctuations, ensuring the continuity of critical computing tasks and the overall operational security of the system.

[0098] The dynamic risk control and cycle time management module intelligently adjusts key parameters of the migration operation to respond promptly and mitigate risks in the migration chain during task migration. In particular, when the system identifies a potential risk of instantaneous task loss in the migration chain, it can automatically trigger a series of control measures. These measures mainly include reducing the time sensitivity of migration tasks, extending the allowable execution time of migration operations, and dynamically adjusting the task scheduling cycle time and resource allocation priority, thereby avoiding task loss or interruption caused by factors such as link quality fluctuations, uneven node load, or resource contention.

[0099] Specifically, reducing the time sensitivity of migration tasks means allowing migrations to be completed within a more relaxed time window, rather than following a fixed and overly tight migration cycle. This avoids problems such as incomplete state synchronization and memory mapping misalignment caused by overly rushed migration tasks, thereby improving the success rate of task migration. Extending the allowed execution time of migration operations means that when a migration task is high-risk, the system will not rashly execute the migration. Instead, by extending the migration time window, it ensures that all states and data are transferred and received more comprehensively and stably during the migration process, reducing the probability of data loss during the task migration. At the same time, the core function of dynamically adjusting the task scheduling takt time and resource allocation priority is to reduce the frequency of high-frequency migration operations, avoid over-scheduling when the migration link is under high load, reduce the risks caused by resource shortages, link congestion, or node contention, and ensure that the system can perform migrations during periods of low load and low risk, thereby maintaining the efficient and stable operation of the system.

[0100] These control strategies proactively mitigate adverse effects such as resource contention, link fluctuations, and excessive node load caused by high-frequency migration by implementing a "frequency reduction and risk avoidance" approach. This ensures the safe migration and execution of tasks and avoids system turbulence caused by over-scheduling. Ultimately, this module ensures the continuity and operational safety of critical computing tasks, improves the adaptability and fault tolerance of the hydropower station dispatching system in complex environments, and, in particular, dynamically responds to and adjusts scheduling strategies when facing potential network latency, abnormal node loads, or link instability. This effectively reduces the probability of failure during migration and increases the success rate of task migration.

[0101] Once a risk of momentary task loss is detected in the migration process, a dynamic risk control mechanism for the migration operation is automatically triggered. This mechanism reduces the time sensitivity of the migration task, extends the allowed execution time of the migration operation, and dynamically adjusts the task scheduling takt time and resource allocation priority. It proactively reduces the frequency of high-frequency migration operations, thereby achieving "frequency reduction and risk avoidance." The specific steps are as follows:

[0102] Once a risk of momentary task loss is identified during the migration process, the time sensitivity attribute of the scheduled task is first adjusted. This relaxes the constraint that migration actions must be completed within a limited time frame, preventing state loading failure due to an excessively narrow migration window. The task migration time window is then dynamically adjusted by calculating the extended execution time factor for the expanded migration operation. The expression for calculating the extended execution time factor for the migration operation is as follows:

[0103] ,

[0104] In the formula, λ1 is the time window expansion sensitivity coefficient, representing the system's response strength to the "time tolerance strategy" adopted after the risk exceeds the limit, i.e., the time expansion multiple coefficient corresponding to each unit of risk exceeding the limit, with a value range of [0.5, 2.0]. The more important the task and the more conservative the system, the larger the value. d This is the instantaneous loss risk coefficient, with a value ranging from [0,1]. A higher value indicates a higher risk. R th It is the reference threshold for the instantaneous loss risk coefficient, α w This is the allowed execution duration extension factor for migration operations. It represents the time tolerance factor dynamically extended based on the risk level from the original migration time limit. It is used to control the maximum migration duration allocated to the task by the scheduling engine. A typical value is α. w =1: No expansion, risk is within acceptable limits; α w =1.5: Expanded by 50%, suitable for high-risk migration tasks; α w >2: Forced slowdown when highly unstable links or systems are severely congested;

[0105] This step calculates the expansion coefficient using a normalized measure of the risk exceeding the limit. If R d >R th This allows for a larger tolerance time window during task migration, ensuring sufficient time to complete state synchronization and mapping, and reducing the probability of data loss.

[0106] The allowed execution time extension factor α based on the acquired migration operation. w The clock speed reduction rate of the scheduled task in the scheduling system is calculated, and the scheduling trigger frequency of the task is dynamically adjusted to avoid high-frequency migration commands increasing the system link pressure or causing link conflicts. The calculation expression is as follows:

[0107] ,

[0108] In the formula, λ2 is the clock decay gain coefficient, representing the time tolerance spread factor (i.e., α) in the risk response. w The nonlinear amplification effect on the beat adjustment amplitude, with a value range of 1–10, δ max This is the upper limit of the clock speed reduction ratio, representing the maximum allowable value of the scheduling clock speed reduction ratio. It is used to limit the slowdown of the system scheduling rhythm to prevent the system response from being too slow or the scheduling frequency from being too low. It is usually set to a fixed value, such as 3, 5, or 10. δ p This is the scheduling tempo reduction factor, which represents the "deceleration factor" that the scheduling rhythm needs to be applied to the task in the scheduling system after identifying a risk of instantaneous task loss. In other words, the original scheduling frequency will be reduced by this factor. The value range is: δ p >1;

[0109] By squaring the time tolerance, the amplification effect of the "high-risk signal" of unstable migration tasks on the system scheduling rhythm is reflected, while an upper limit is set to avoid affecting the global scheduling efficiency under extreme conditions.

[0110] Based on the instantaneous loss risk coefficient R d Migration operation allows for a multiplier α to extend execution time. w With the deceleration ratio δ of the scheduling cycle p The scheduling resource compression priority coefficient of the target task is jointly calculated and used to adjust its sorting priority in the global resource contention queue, avoiding strong binding between high-risk tasks and critical path resources. The calculation expression is as follows:

[0111] ,

[0112] In the formula, λ3 is the nonlinear curvature factor of the compression mapping, used to control the steepness of the sigmoid compression function response, affecting the sensitivity and magnitude of task-priority compression; e is the natural base; and π is the coefficient of performance. r This is a resource compression priority coefficient, representing the priority weight of the current task in the task scheduling resource allocation queue. It is used to dynamically adjust the sorting position or scoring factor of tasks when allocating resources in the scheduler, preventing high-risk tasks from crowding out scheduling resources. The value range is 0 < π. r <1, the smaller the value, the stronger the system's priority compression of the task.

[0113] This step uses non-linear compression with a sigmoid function to proactively lower the priority of tasks with high risk, slow scheduling pace, and large time requirements. From a resource scheduling perspective, this avoids their high sensitivity from interfering with critical migration channels, thus achieving avoidance and off-peak scheduling strategies.

[0114] This invention achieves intelligent full-process control over the potential instantaneous loss risk of hydropower scheduling tasks during hot migration by constructing a dynamic and elastic resource expansion mechanism with multi-module collaborative capabilities. This effectively improves the system's scheduling stability and resource adaptability in high-concurrency, heterogeneous node environments. The system acquires the real-time operating status and link quality parameters of key nodes through a distributed state awareness module. Combined with data preprocessing and structured construction operations, it ensures the temporal consistency and analytical adaptability of multi-source data. Based on this, it extracts and quantifies key indicators that accurately reflect migration risks, such as the degree of cache paging interference and the trend of mapping space misalignment. A deep learning model is used to intelligently determine the migration link status of the current task, identifying potential loss risks during migration in advance. Finally, based on the risk identification results, the scheduling system dynamically adjusts the migration duration, scheduling rhythm, and task priority, proactively avoiding unstable resource states and achieving flexible control of the scheduling rhythm. This significantly improves the safety and continuity of the task migration process, ensuring the real-time performance of critical hydropower scheduling tasks and the overall reliability of the system operation.

[0115] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0116] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0117] It should be noted that, in this document, the use of relational terms such as "first" and "second" is merely for distinguishing one entity or operation from another, and does not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0118] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0120] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0122] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0123] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0124] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A dynamically elastic resource-expanding computing server for hydropower stations, characterized in that, It includes a distributed state awareness acquisition module, a data preprocessing and structuring module, a feature extraction and risk quantification analysis module, a deep learning intelligent assessment module, and a dynamic risk control and cycle management module. The distributed state awareness acquisition module continuously and in real time collects the operating status and link quality parameter data of each key node in the migration link through the distributed state awareness system deployed on the central dispatch control platform of the hydropower station computing server during the hot migration of hydropower dispatching tasks. The data preprocessing and structuring module preprocesses the multi-dimensional raw data collected in real time. After preprocessing, it constructs a data set with unified field definitions and time index markers, and stores it in a structured format according to time sequence. The feature extraction and risk quantification analysis module, based on the preprocessed structured data set, uses feature engineering methods to extract key indicators that sensitively reflect the potential instantaneous loss risk of the task in the migration link, and performs comprehensive analysis on the extracted key indicators to quantify the severity of the instantaneous loss risk during the task migration process. The deep learning intelligent evaluation module inputs key feature indicators that have undergone comprehensive analysis and risk quantification into a deep learning model that has been pre-trained based on historical data. The model performs multi-dimensional dynamic evaluation of the task migration link status and intelligently determines whether there is a risk of instantaneous loss of the current task. The dynamic risk control and takt management module automatically triggers a dynamic risk control mechanism for migration operations when it detects that there is a risk of instantaneous loss of tasks in the migration link. This reduces the time sensitivity of migration tasks, extends the allowed execution time of migration operations, and dynamically adjusts the takt of task scheduling and the priority of resource allocation. It proactively reduces the frequency of high-frequency migration operations, thereby achieving "frequency reduction and risk avoidance". When a risk of momentary task loss is detected in the migration chain, a dynamic risk control mechanism for the migration operation is automatically triggered. This reduces the time sensitivity of the migration task, extends the allowed execution time of the migration operation, and dynamically adjusts the task scheduling takt time and resource allocation priority. This proactively reduces the frequency of high-frequency migration operations, thus achieving "frequency reduction and risk avoidance." The specific steps are as follows: When a risk of momentary task loss is detected in the migration chain, firstly, the time sensitivity attribute of the scheduled task is adjusted. By calculating the extended execution time extension factor for the migration operation, the task migration time window is dynamically adjusted. The expression for calculating the extended execution time extension factor for the migration operation is as follows: In the formula, It is the sensitivity coefficient for time window expansion. It is a reference threshold for the risk factor of instantaneous loss. This is the multiple by which the migration operation is allowed to extend its execution time. It is the instantaneous loss risk coefficient; the allowable execution time extension multiple based on the acquired migration operation. The clock speed reduction rate of the scheduled task in the scheduling system is calculated, and the scheduling trigger frequency of the task is dynamically adjusted. The calculation expression is as follows: In the formula, It is the beat attenuation gain coefficient. This is the upper limit of the beat deceleration ratio. It is the scheduling cycle deceleration ratio; based on the instantaneous loss risk coefficient. Migration operations are allowed to extend execution time by a certain multiple. With the deceleration ratio of the scheduling cycle time The scheduling resource compression priority coefficient of the target task is jointly calculated and used to adjust the sorting priority in the global resource contention queue. The calculation expression is as follows: In the formula, It is the nonlinear curvature factor of the compression mapping. It is the natural base. It is the resource compression priority coefficient.

2. The dynamic elastic resource expansion hydropower station computing server according to claim 1, characterized in that, During the hot migration of hydropower dispatching tasks, the link status is collected through a distributed status awareness system deployed on the central dispatching and control platform of the hydropower station's computing power server. The specific steps are as follows: resource monitoring probes and communication quality acquisition components are deployed at key locations of each computing node and network link involved in the hot migration to ensure independent collection of local resource status and link information; each sensing unit continuously collects node data parameters at a set frequency; all collected operating status and link quality data are transmitted back to the central dispatching and control platform in real time through the internal data bus or monitoring channel, and are uniformly received and partitioned by the dispatching engine.

3. The dynamic elastic resource expansion hydropower station computing server according to claim 1, characterized in that, Based on the preprocessed structured dataset, feature engineering methods are used to extract key indicators that sensitively reflect the potential instantaneous loss risk of tasks during the migration process. The extracted indicators include the degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node, and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node. The degree of access conflict between the migrating task and the locally running task in the memory page table cache area of ​​the hot migration node, and the fluctuation trend of the address offset of the task's virtual memory space mapped to the target node are comprehensively analyzed under the detection window to generate cache paging interference indicators and mapping space misalignment indicators. The severity of the instantaneous loss risk during the task migration process is quantified by the cache paging interference indicators and mapping space misalignment indicators.

4. The dynamic elastic resource expansion hydropower station computing server according to claim 3, characterized in that, The specific steps for generating a cache paging interference index by comprehensively analyzing the access conflict level between migration tasks and locally running tasks in the memory page table cache area of ​​a hot-migrating node under the detection window are as follows: Conflict detection is performed on the access behavior of migration tasks and locally running tasks in the memory page table cache area of ​​the hot-migrating node. The access trajectories of migration tasks and local tasks to page table cache resources are recorded, and the number of overlapping accesses within the same memory page mapping area is calculated. The frequency of conflicts occurring between migration tasks and local tasks during cache access is defined as the conflict frequency factor. The expression for calculating the conflict frequency factor is as follows: In the formula, This represents the number of conflicts when migration tasks and local tasks access the same cache page table entry within the detection window. It is the sum of the cumulative number of accesses to the current memory page mapping region in the page table cache by the migration task and the local task. This is the conflict frequency factor; weighting factors are set according to the impact of access to different memory page mapping regions on the migration task, and then the local conflict degree is calculated. The calculation formula is as follows: In the formula, It is a regional weighting factor. It is the local conflict level; based on the local conflict level of each memory region. Calculate the overall cache paging interference index. The cache paging interference index is used to characterize the intensity of resource interference experienced by the migration task in the entire page table cache space. The calculation formula is as follows: In the formula, It is a cache paging interference indicator. It is the first Local conflict level of each memory page mapping region It is the first The number of accesses to the memory page mapped region. It represents the total number of memory page mapped regions. It is a logarithmic function.

5. A dynamic elastic resource expansion hydropower station computing server according to claim 3, characterized in that, The specific steps for generating a mapping space misalignment index by comprehensively analyzing the address offset fluctuation trend of the task's virtual memory space mapped to the target node under the detection window are as follows: During the task hot migration process, obtain the sequence of virtual address segments of the current task on the source node. and the address segment sequence after remapping in the target node ,in, , It is the first in the source node A virtual address, It is the total number of virtual addresses. , It is the first in the target node For each pair of virtual addresses, calculate the structure offset using the following expression: In the formula, It is an address hash mapping function. It is the address mapping offset; based on the obtained address mapping offset sequence, a mapping space misalignment index is generated, and the generation formula is as follows: In the formula, It is the address mapping offset of the previous virtual address unit. It is the first The virtual address unit and the first The minimum value of the address mapping offset of each virtual address unit. It is a mapping space misalignment indicator.

6. A dynamic elastic resource expansion hydropower station computing server according to claim 3, characterized in that, The cache paging interference index and mapping space misalignment index, which have undergone comprehensive analysis and risk quantification, are input into a deep learning model that has been pre-trained based on historical data. The model generates a task instantaneous loss risk coefficient, and the task migration link status is dynamically evaluated in multiple dimensions based on the task instantaneous loss risk coefficient, so as to intelligently determine whether there is an instantaneous loss risk for the current task.

7. A dynamic elastic resource expansion hydropower station computing server according to claim 6, characterized in that, The system compares the instantaneous task loss risk coefficient generated by a deep learning model trained on historical data to perform multi-dimensional dynamic evaluation of the task migration link status with a pre-set reference threshold for instantaneous task loss risk coefficient. This comparison intelligently determines whether the current task has an instantaneous loss risk. The judgment logic is as follows: if the instantaneous task loss risk coefficient is greater than the pre-set reference threshold, then the current task is judged to have an instantaneous loss risk; if the instantaneous task loss risk coefficient is less than or equal to the pre-set reference threshold, then the current task is judged not to have an instantaneous loss risk.

Citation Information

Patent Citations

  • Big data platform scheduling task and data collaborative smooth migration method and system

    CN119576506A