Heterogeneous computing power resource dynamic allocation method oriented to ocean supercomputing environment
By monitoring the calculation amount and real-time level of marine data processing tasks in real time, dynamically calibrate task priorities and redistribute resources, the problems of uneven resource utilization and task delay in marine supercomputing environments are solved, and load balancing and efficient resource utilization are achieved.
Patent Information
- Application Number
- CN202510805137.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing technology lacks fine-grained identification of dynamic computing requirements and real-time level of task in marine supercomputing environments, resulting in uneven resource utilization and delayed execution of critical tasks, making it difficult to cope with resource overload problems in marine task environments with high frequency changes.
By monitoring the calculation amount and real-time level of marine data processing tasks in real time, dynamically calibrate task priorities, filter suitable heterogeneous computing resources, and trigger resource redistribution when load rate is overloaded, and priority migration of low-priority tasks is given.
It realizes efficient utilization of heterogeneous computing resources, reduces critical tasks latency, ensures load balancing, and improves system operation efficiency and task completion quality.
Smart Images

Figure CN120336033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of marine information processing, and particularly to a method for dynamically allocating heterogeneous computing power resources for a marine supercomputing environment. Background Art
[0002] With the continuous improvement of the requirements for real-time performance and computational accuracy of marine observation, simulation, early warning, and analysis tasks, marine scientific research and engineering decision-making increasingly rely on supercomputing platforms for rapid processing and high-frequency scheduling of massive data. Currently, supercomputing environments for marine applications mostly adopt heterogeneous computing power architectures, covering multiple types of computing nodes such as CPUs, GPUs, and FPGAs, to adapt to the concurrent execution requirements of different task models. However, due to the significant differences in computational intensity and response time limits among various tasks, if a static resource scheduling strategy is adopted, it often leads to uneven resource utilization and delayed execution of critical tasks, severely restricting the overall operating efficiency of the system and the quality of task completion.
[0003] Existing technologies often lack fine-grained identification of the dynamic computational requirements and real-time levels of tasks during the resource scheduling process, cannot accurately match heterogeneous computing power resources according to task priorities, and have not formed a complete resource reallocation mechanism to cope with resource overload situations. Especially in the marine task environment with high-frequency changes, traditional scheduling methods are difficult to perceive resource load fluctuations in real time and difficult to dynamically adjust the task deployment location, resulting in high-priority tasks being easily delayed and the system being prone to local bottlenecks. Therefore, there is an urgent need for a method for dynamically allocating heterogeneous computing power resources for a marine supercomputing environment to solve the above problems. Summary of the Invention
[0004] Based on the above objectives, the present invention provides a method for dynamically allocating heterogeneous computing power resources for a marine supercomputing environment.
[0005] The method for dynamically allocating heterogeneous computing power resources for a marine supercomputing environment includes the following steps: S1: Real-time monitor the dynamic demand characteristics of marine data processing tasks, including task computational volume and task real-time level; S2: Dynamically calibrate the task priority according to the task computational volume and real-time level in S1; S3: Based on the task priority, scan the current heterogeneous computing power resource pool, and filter out target resources that simultaneously meet the task computational volume requirements and whose response speed matches the real-time level; S4: Allocate the task to the target resource, and real-time monitor the change in the load rate of the target resource; S5: If the load rate of the target resource exceeds the preset threshold, trigger resource reallocation, and preferentially migrate low-priority tasks to other available computing power resources; S6: Output the final resource allocation plan, and record the completion time of the corresponding tasks and the utilization rate of each computing resource.
[0006] Optionally, the S1 specifically includes: S11: Receive the description information of the task to be executed through the ocean supercomputer task scheduling interface, and extract the source code call graph, input data scale identifier, and the user-set deadline; S12: Invoke the static instruction count analyzer to perform per-node instruction statistics on the source code call graph, and combine the input data scale identifier to obtain the task computation amount value using a linear upscaling model; S13: Calculate the difference Δt between the user-set deadline and the current time, and determine the task real-time level Level in the preset real-time level mapping table according to the difference Δt. Specifically, when Δt ≤ 1s, the corresponding level is L1; when 1s < Δt ≤ 5s, the corresponding level is L2; when Δt > 5s, the corresponding level is L3; S14: Form a dynamic demand feature data pair with the task computation amount and the real-time level Level, write it into the demand feature cache queue, and refresh and update it at a fixed period of 100ms.
[0007] Optionally, the S12 specifically includes: S121: Based on the task source code call graph extracted in S11, perform static instruction counting on each function node to obtain the basic instruction count corresponding to each node, denoted as , where i is the node number; S122: Obtain the total input data volume D according to the input data scale identifier, and find the corresponding scale growth coefficient ; S123: Construct a linear upscaling model based on the total node instruction volume and the scale factor, and calculate the task computation amount value. The formula is: , where C represents the task computation amount value; represents the data scale growth coefficient; N represents the total number of function nodes in the call graph.
[0008] Optionally, the S2 specifically includes: S21: Receive the task computation amount value C and the real-time level L output by S1; S22: According to the real-time level L, look up the corresponding weight coefficient from the preset real-time weight mapping table, where L1 corresponds to = 0.6, L2 corresponds to = 0.3, L3 corresponds to = 0.1; S23: Perform normalization processing on the task computation amount C to obtain the normalized computation amount value ; S24: Combine with the weight coefficient to calculate the task priority value P, and its calculation formula is: , where the value range of P is [0, 1]; S25: Store the calculated priority value P in the task scheduling table, and sort it in descending order of the priority value. The scheduling table is updated every 500 ms.
[0009] Optionally, the specific steps of S3 are as follows: S31: Receive the task priority value P, the task computation amount value C, and the task real-time level L output by S2, initialize the resource scanning queue, and traverse all computing nodes in the current heterogeneous computing power resource pool; S32: Obtain the available computing power parameters of each computing node, including the current idle instruction throughput rate R and the current response delay time T; S33: Perform a preliminary screening on each computing node, and eliminate the nodes whose current instruction throughput rate R is lower than the minimum processing capacity required by the task computation amount C; S34: Among the nodes after the preliminary screening, look up the corresponding maximum allowable response delay according to the task real-time level L, and screen out all nodes with a response delay time T > ; S35: Add the remaining nodes that meet the requirements of instruction throughput rate and response delay to the candidate resource set, sort them in descending order of the instruction throughput rate, and select the node ranked first as the target resource node to carry the current task.
[0010] Optionally, the specific steps of S32 are as follows: S321: Read the hardware performance counter of the target computing node within the monitoring period to obtain the current main frequency F and the instruction-level parallelism I of the node; S322: Calculate the theoretical peak instruction throughput rate of the node according to F and I, and the calculation formula is: ; S323: Call the operating system resource monitoring interface to obtain the current CPU utilization rate U of the node, and then calculate the current idle instruction throughput rate R of the node. The formula is: ; S324: Read the task queue length Q of the node and the task service rate of the node, and calculate the queue waiting time , and the formula is: ; S325: Measure the average context switching time S of the node through the kernel timer; S326: According to Obtain the current response delay time T of the node from S, and the formula is: .
[0011] Optionally, the S4 specifically includes: S41: The task scheduling manager sends a task start instruction packet to the target resource node selected by S3 through the Internet. The instruction packet includes a task identifier, an input data path, the required memory capacity, and an execution image fingerprint; S42: After parsing the instruction packet, the target resource node calls the local container orchestration service, pulls the corresponding computing image based on the execution image fingerprint and allocates an independent namespace, and then mounts the input data path to complete the instantiation of the task running environment; S43: The scheduling manager issues the computing entry function and parameters through the zero-copy message queue. The target resource node immediately starts the task main thread and records the task start timestamp after receiving the message; S44: During the task execution, the built-in resource monitoring agent of the node calls the hardware performance counter at a sampling period of 200 ms to obtain the current instruction throughput rate used by the node and the peak instruction throughput rate , and calculate the node load rate according to the following formula , and the formula is: ; S45: The node reports the load rate and the task execution progress percentage to the scheduling manager in real time through the gRPC streaming channel to continuously monitor the load change of the target resource node.
[0012] Optionally, the S5 specifically includes: S51: After receiving the current load rate reported by the target resource node, the scheduling manager compares it with the preset load rate threshold . If the condition > is met, the resource reallocation process is immediately entered; S52: The scheduling manager queries all the tasks currently being executed in the target node, sorts them in ascending order according to the corresponding priority value P of each task, and constructs a low-priority task list; S53: Select tasks from the low-priority task list in turn, scan all candidate idle nodes in the heterogeneous computing power resource pool that are not the current node, and filter out the nodes that meet both the task computing volume requirements and the response delay limit; S54: Sort the nodes that meet the conditions in descending order according to their idle instruction throughput rate, select the optimal node as the target migration node, and send a task migration instruction packet. The task migration instruction packet includes a task context snapshot, input data path mapping information, and recovery image information; S55: After receiving the instruction packet, the target migration node completes data mapping and computing image preparation, loads the original task context, and starts migration without interrupting the task process; S56: The scheduling manager updates the task scheduling table and the resource mapping table, marking the completion of an effective task migration.
[0013] Optionally, the preset load rate threshold has the following calculation formula: , where represents the preset load rate threshold of the target resource node; represents the average load rate of the node in the most recent P scheduling cycles; represents the standard deviation of the load rate of the node in the most recent P scheduling cycles; is the margin coefficient.
[0014] Optionally, S6 specifically includes: S61: After completing the task assignment of the current batch, the scheduling manager forms a one-to-one mapping between each task and its assigned target resource node, constructs a resource allocation mapping table, and the mapping table includes task identification, resource node number, priority value, scheduling timestamp, and estimated computing time consumption; S62: After the task is completed, the target resource node sends back the task completion status to the scheduling manager through the node management interface, including task identification, actual start timestamp, and actual end timestamp, and the scheduling manager calculates the task completion time based on this; S63: During the task execution process, the resource monitoring agent of the target resource node periodically collects local CPU utilization data, and the scheduling manager performs numerical integration and averaging processing on the CPU utilization data during this period after the task is completed to calculate the average computing power utilization during the task execution; S64: The scheduling manager organizes the data of the target resource number, task completion time, and average computing power utilization corresponding to each task to generate a structured resource allocation result report and outputs it in JSON format.
[0015] Advantages of the present invention: In the present invention, by quantifying the task computing amount and response time limit in real time during the task submission phase, mapping the two into computable priorities, and combining the linear upscaling model and the weight mapping table to drive resource screening, high-computing amount-high-real-time tasks can be accurately matched to heterogeneous nodes that meet the throughput and latency requirements in the first scheduling, avoiding queuing waiting of critical tasks due to resource mismatch and effectively shortening the task start delay.
[0016] In the present invention, node load is monitored through millisecond-level performance counting sampling and dynamic threshold method. When node overload is detected, seamless migration of low-priority tasks is automatically triggered, and a structured resource allocation report is output after the tasks are completed, providing data support for subsequent scheduling strategy optimization and performance modeling. This mechanism ensures balanced load on each node and efficient utilization of global resources, and reduces the risk of task delay caused by local overload. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 Schematic diagram of the dynamic allocation method of heterogeneous computing power resources in an embodiment of the present invention; Figure 2 Schematic diagram of the real-time monitoring dynamic demand feature process in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The present invention will be described in detail below with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0020] As Figure 1 - Figure 2 shown, the dynamic allocation method of heterogeneous computing power resources for the ocean supercomputing environment includes the following steps: S1: Real-time monitor the dynamic demand characteristics of ocean data processing tasks, including task computing volume and task real-time level; S2: Dynamically calibrate the task priority according to the task computing volume and real-time level in S1, where the task with high computing volume and high real-time level has the highest priority; S3: Based on the task priority, scan the current heterogeneous computing power resource pool, and screen out the target resources that simultaneously meet the task computing volume requirements and whose response speed matches the real-time level; S4: Allocate the task to the target resource, and real-time monitor the change of the load rate of the target resource; S5: If the load rate of the target resource exceeds the preset threshold, trigger resource reallocation, and preferentially migrate low-priority tasks to other available computing power resources; S6: Output the final resource allocation plan, and record the completion time of the corresponding task and the utilization rate of each computing power resource.
[0021] S1 specifically includes: S11: receiving description information of the task to be executed through the Ocean Supercomputing Task Scheduling Interface, extracting the source code call graph, input data scale identifier, and the deadline set by the user; S12: calling the static instruction count analyzer to perform node-by-node instruction statistics on the source code call graph, and combining the input data scale identifier, using the linear upscaling model to obtain the task computing value; S13: Calculate the difference Δt between the user-set deadline and the current time, and determine the task real-time level Level in the preset real-time level mapping table according to the difference Δt. Specifically, when Δt≤1s, it corresponds to level L1; 1s<Δt≤5s corresponds to level L2; Δt>5s corresponds to level L3; S14: The task computing amount and the real-time level Level are combined into a dynamic demand feature data pair, written into the demand feature cache queue, and refreshed and updated at a fixed period of 100ms; the above steps achieve real-time and fine-grained monitoring of the task computing amount and real-time level by accurately quantifying the difference between the number of task source code instructions and the deadline, providing an accurate and reliable basis for subsequent priority setting and resource screening.
[0022] S12 specifically includes: S121: Based on the task source code call graph extracted in S11, static instruction counting is performed on each function node to obtain the basic instruction number corresponding to each node, recorded as , where i is the node number; S122: Obtain the total amount of input data D according to the input data scale identifier, and search for the corresponding scale growth coefficient from the preset scale weight table , this coefficient represents the linear expansion relationship between data scale and task computing amount; Table 1 Preset scale weights Input data scale range (MB) Corresponding scale growth coefficient α Corresponding description 0 < D ≤ 10 1 Base scale, default instruction volume ratio 10 < D ≤ 50 1.2 Slight growth, mildly extended model 50 < D ≤ 100 1.5 Medium extension, cache hit needs to be considered 100 < D ≤ 500 1.9 Significant growth, increased I / O load D > 500 2.5 Ultra-large scale, accompanied by a large amount of memory scheduling and thread scheduling overhead S123: Construct a linear up-scaling model based on the total node instruction amount and the scale factor, and calculate the task calculation amount value. The formula is: , where C represents the task computation value; Represents the data scale growth coefficient; N represents the total number of function nodes in the call graph; the above steps can efficiently quantify the task computing intensity, avoid the estimation deviation of computing requirements in the task scheduling process, and thus improve the accuracy and stability of resource allocation.
[0023] S2 specifically includes: S21: Receive the task computation volume value C and the real-time level L output by S1, where C represents the number of millions of instructions required for task computation, and L represents the real-time level, with a value range of {L1, L2, L3}, corresponding to high, medium, and low real-time tasks respectively; S22: According to the real-time level L, look up the corresponding weight coefficient from the preset real-time weight mapping table , where L1 corresponds to = 0.6, L2 corresponds to = 0.3, L3 corresponds to = 0.1; S23: Perform normalization processing on the task computation volume C to obtain the normalized computation volume value , and its calculation formula is: , where C is the current task computation volume, with the unit of millions of instructions (MIPS); are respectively the minimum and maximum computation volumes of all tasks to be processed within the current scheduling period, with the unit of millions of instructions; has a value range of [0, 1]; S24: Combine with the weight coefficient to calculate the task priority value P, and its calculation formula is: , where P has a value range of [0, 1], and the larger the value, the higher the priority; S25: Store the calculated priority value P in the task scheduling table, and sort it in descending order of the priority value for subsequent resource allocation process calls. The scheduling table is updated every 500 ms; The above steps, through the combination of real-time weights and computation volume normalization indicators, can achieve a unified priority evaluation mechanism for multi-tasks under different data scales and time sensitivities, ensuring that the scheduling strategy always gives priority to critical and urgent tasks under limited resources, and effectively improving the utilization efficiency of supercomputer resources and the rationality of task scheduling.
[0024] S3 specifically includes: S31: Receive the task priority value P, the task computation volume value C, and the task real-time level L output by S2, and initialize the resource scanning queue, and traverse all computing nodes in the current heterogeneous computing power resource pool; S32: Obtain the available computing power parameters of each computing node, including the current idle instruction throughput rate R and the current response delay time T, where the instruction throughput rate represents the maximum number of instructions that can be processed per unit time, and the response delay time represents the average waiting time for the node to start computing after receiving the task instruction; S33: Perform a preliminary screening on each computing node, and eliminate the nodes whose current instruction throughput rate R is lower than the minimum processing capacity required by the task computation volume C; S34: Among the preliminarily screened nodes, retrieve the corresponding maximum allowable response latency according to the task real-time level L , and filter out all nodes with response latency time T > ; S35: Add the remaining nodes that meet the requirements of instruction throughput rate and response latency to the candidate resource set, sort them in descending order of instruction throughput rate, and select the node ranked first as the target resource node to carry the current task; The above steps perform a dual screening on the nodes in the heterogeneous resource pool from two dimensions by combining the task computing volume and the real-time level, ensuring that the selected target resource not only has sufficient computing power but also can complete the response processing within the task time limit, guaranteeing the response reliability and scheduling adaptability of the computing task in the high-concurrency scenario of ocean supercomputing.
[0025] S32 specifically includes: S321: Read the hardware performance counters of the target computing node within the monitoring period to obtain the current main frequency F and instruction-level parallelism I of the node; S322: Calculate the theoretical peak instruction throughput rate of the node according to F and I , and the calculation formula is: ; S323: Call the operating system resource monitoring interface to obtain the current CPU utilization rate U of the node, and then calculate the current idle instruction throughput rate R of the node. The formula is: ; S324: Read the node task queue length Q and the node task service rate , and calculate the queue waiting time , and the formula is: ; S325: Measure the average context switching time S of the node through the kernel timer; S326: Obtain the current response latency time T of the node according to and S. The formula is: ; The above steps can accurately obtain the remaining processing power and response latency of each computing node within a millisecond-level cycle by jointly quantifying the hardware counters and the real-time monitoring data of the operating system, providing a reliable data basis for the subsequent accurate matching of tasks and resources, thereby improving the real-time performance and efficiency of heterogeneous computing power resource allocation.
[0026] S4 specifically includes: S41: The task scheduling manager sends a task start instruction packet to the target resource node selected in S3 through the interconnection network. The instruction packet includes the task identifier, input data path, required memory capacity, and execution image fingerprint; S42: After the target resource node parses the instruction packet, it calls the local container orchestration service, pulls the corresponding computing image based on the execution image fingerprint and allocates an independent namespace, and then mounts the input data path to complete the instantiation of the task running environment; S43: The scheduling manager issues the computing entry function and parameters through the zero-copy message queue. After receiving the message, the target resource node immediately starts the task main thread and records the task start timestamp; S44: During the task execution, the built-in resource monitoring agent of the node calls the hardware performance counter at a sampling period of 200 ms to obtain the current instruction throughput rate used by the node and the peak instruction throughput rate , and calculates the node load rate according to the following formula , the formula is: ; S45: The node reports the load rate and the task execution progress percentage to the scheduling manager in real time through the gRPC streaming channel, realizing continuous monitoring of the load change of the target resource node; The above steps use container orchestration to quickly start tasks and combine high-frequency performance counting sampling to calculate the dimensionless load rate, realizing millisecond-level task distribution and real-time load transparent monitoring, enhancing the dynamic response ability of computing power resource scheduling and the accuracy of resource utilization.
[0027] S5 specifically includes: S51: After the scheduling manager receives the current load rate reported by the target resource node , it compares it with the preset load rate threshold . If the condition > is met, it immediately enters the resource reallocation process; S52: The scheduling manager queries all the tasks currently being executed in the target node, sorts them in ascending order according to the corresponding priority value P of each task, and constructs a low-priority task list; S53: Select tasks from the low-priority task list in turn, scan all candidate idle nodes in the heterogeneous computing power resource pool that are not the current node, and filter out the nodes that meet both the task computing volume requirements and the response delay limit; S54: Sort the nodes that meet the conditions in descending order according to their idle instruction throughput rate, select the optimal node as the target migration node, and send a task migration instruction packet, which includes the task context snapshot, input data path mapping information, and recovery image information; S55: After receiving the instruction packet, the target migration node completes data mapping and computing image preparation, and loads the original task context to start the migration without interrupting the task process; S56: The scheduling manager updates the task scheduling table and the resource mapping table, marking the completion of an effective task migration. The above steps achieve a distributed computing power adaptive scheduling mechanism under the condition of overloaded target resources by preferentially identifying and migrating low-priority tasks and intelligently screening target nodes based on the current resource pool status, effectively alleviating the node computing bottleneck and ensuring the running quality of high-priority tasks.
[0028] Preset load rate threshold The calculation formula is: , where represents the preset load rate threshold of the target resource node; represents the average load rate of the node in the most recent P scheduling cycles; represents the standard deviation of the load rate of the node in the most recent P scheduling cycles; is the margin coefficient, used to reflect the tolerance of the system to load fluctuations, and its value range is [1.5, 2.5], which is specifically determined by the task real-time level. Specifically, when the task real-time level is L1, the margin coefficient is 1.5; when the task real-time level is L2, the margin coefficient is 2.0; when the task real-time level is L3, the margin coefficient is 2.5.
[0029] S6 specifically includes: S61: After the scheduling manager completes the task allocation of the current batch, it forms a one-to-one mapping between each task and its allocated target resource node, constructs a resource allocation mapping table, and the mapping table includes task identification, resource node number, priority value, scheduling timestamp, and estimated calculation time-consuming; S62: After the task is completed, the target resource node sends back the task completion status to the scheduling manager through the node management interface, including task identification, actual start timestamp, and actual end timestamp. The scheduling manager calculates the task completion time based on this, and the calculation method is the task end time minus the task start time; S63: During the task execution process, the resource monitoring agent of the target resource node periodically collects local CPU utilization data. After the task is completed, the scheduling manager performs numerical integration and averaging on the CPU utilization data during this period to calculate the average computing power utilization during the task execution; S64: The scheduling manager organizes the data of the target resource number, task completion time, and average computing power utilization corresponding to each task to generate a structured resource allocation result report and outputs it in JSON format. The above steps ensure the traceability and quantifiable analysis ability of the scheduling process by standardizing the recording of task completion time and computing power usage, and combining structured output and historical archiving.
[0030] The present invention covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. Additionally, well-known methods, processes, procedures, components, and circuits, etc. are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0031] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for dynamically allocating heterogeneous computing power resources for an ocean supercomputing environment, characterized in that It includes the following steps: S1: Real-time monitor the dynamic demand characteristics of ocean data processing tasks, including task computing volume and task real-time level; S2: Dynamically calibrate the task priority according to the task computing volume and real-time level in S1; S3: Based on the task priority, scan the current heterogeneous computing power resource pool, and filter out the target resources that simultaneously meet the task computing volume requirements and whose response speed matches the real-time level; S4: Allocate the task to the target resource and real-time monitor the change of the load rate of the target resource; S5: If the load rate of the target resource exceeds the preset threshold, trigger resource reallocation, and preferentially migrate low-priority tasks to other available computing power resources; S6: Output the final resource allocation plan, and record the completion time of the corresponding task and the utilization rate of each computing power resource.
2. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 1, wherein, The specific content of S1 includes: S11: Receive the description information of the task to be executed through the ocean supercomputer task scheduling interface, and extract the source code call graph, input data scale identifier, and the deadline set by the user; S12: Call the static instruction count analyzer to perform node-by-node instruction statistics on the source code call graph, and combine the input data scale identifier to obtain the task computing volume value using a linear upscaling model; S13: Calculate the difference Δt between the deadline set by the user and the current time, and determine the task real-time level Level in the preset real-time level mapping table according to the difference Δt. Specifically, when Δt ≤ 1s, it corresponds to level L1; when 1s < Δt ≤ 5s, it corresponds to level L2; when Δt > 5s, it corresponds to level L3; S14: Form a dynamic demand characteristic data pair with the task computing volume and the real-time level Level, write it into the demand characteristic cache queue, and refresh and update it at a fixed period of 100 ms.
3. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 2, characterized in that, The specific content of S12 includes: S121: Based on the task source code call graph extracted in S11, perform static instruction counting on each function node, and obtain the basic instruction count corresponding to each node, denoted as , where i is the node number; S122: Obtain the total amount of input data D according to the input data scale identifier, and look up the corresponding scale growth coefficient from the preset scale weight table ; S123: Construct a linear upscaling model based on the total node instruction volume and the scaling factor, and calculate the task computation volume value. The formula is: , where C represents the task computation volume value; represents the data scale growth coefficient; N represents the total number of function nodes in the call graph.
4. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 3, wherein The specific content of S2 includes: S21: Receive the task computing volume value C and the real-time level L output by S1; S22: Obtain the corresponding weight coefficient from the preset real-time weight mapping table according to the real-time level L , where L1 corresponds to = 0.6, L2 corresponds to = 0.3, L3 corresponds to = 0.1; S23: Normalize the task computation amount C to obtain the normalized computation amount value ; S24: Combine with the weight coefficient to calculate the task priority value P, and its calculation formula is: , where the value range of P is [0, 1]; S25: Store the calculated priority value P in the task scheduling table, and sort it in descending order according to the priority value. The scheduling table is updated every 500 ms.
5. The heterogeneous computing power resource dynamic allocation method for the ocean supercomputing environment according to claim 4, wherein The specific content of S3 includes: S31: Receive the task priority value P, the task computing volume value C, and the task real-time level L output by S2, and initialize the resource scanning queue, and traverse all computing nodes in the current heterogeneous computing power resource pool; S32: Obtain the available computing power parameters of each computing node, including the current idle instruction throughput rate R and the current response delay time T; S33: Perform a preliminary screening on each computing node, and eliminate the nodes whose current instruction throughput rate R is lower than the minimum processing capacity required by the task computing volume C; S34: Among the initially screened nodes, retrieve the corresponding maximum allowable response delay according to the task real-time level L , and filter out all nodes with response delay time T > ; S35: Add the remaining nodes that meet the instruction throughput rate and response delay requirements to the candidate resource set, sort them in descending order according to the instruction throughput rate, and select the node ranked first as the target resource node to carry the current task.
6. The heterogeneous computing power resource dynamic allocation method for the ocean supercomputing environment according to claim 5, wherein The specific content of S32 includes: S321: Read the hardware performance counter of the target computing node within the monitoring period to obtain the current main frequency F and the instruction-level parallelism I of the node; S322: Calculate the theoretical peak instruction throughput rate of the node based on F and I , and the calculation formula is as follows: ; S323: Call the operating system resource monitoring interface to obtain the current CPU utilization U of the node, and then calculate the current idle instruction throughput rate R of the node. The formula is: ; S324: Read the node task queue length Q and the node task service rate , calculate the queue waiting time , the formula is: ; S325: Measure the average context switching time S of the node through the kernel timer; S326: According to and S, obtain the current response delay time T of the node. The formula is: .
7. The heterogeneous computing power resource dynamic allocation method for the ocean supercomputing environment according to claim 1, wherein The specific content of S4 includes: S41: The task scheduling manager sends a task start instruction packet to the target resource node selected by S3 through the Internet. The instruction packet includes a task identifier, an input data path, the required memory capacity, and an execution image fingerprint. S42: After parsing the instruction packet, the target resource node calls the local container orchestration service, pulls the corresponding computing image based on the execution image fingerprint, and allocates an independent namespace. Subsequently, it mounts the input data path to complete the instantiation of the task running environment. S43: The scheduling manager issues the computing entry function and parameters through a zero-copy message queue. The target resource node immediately starts the task main thread and records the task start timestamp after receiving the message. S44: During task execution, the resource monitoring agent built into the node calls the hardware performance counter at a sampling period of 200 ms to obtain the current instruction throughput rate used by the node and the peak instruction throughput rate , and calculates the node load rate according to the following formula , the formula is: ; S45: The node reports the load ratio and the task execution progress percentage to the scheduling manager in real time through the gRPC streaming channel, so as to continuously monitor the load changes of the target resource node.
8. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 1, characterized in that The specific steps of S5 are as follows: S51: After the scheduling manager receives the currently reported load rate from the target resource node it compares it with a preset load rate threshold and if the condition > is met, it immediately enters the resource reallocation process; S52: The scheduling manager queries all the tasks currently being executed in the target node, sorts them in ascending order according to the corresponding priority value P of each task, and constructs a low-priority task list. S53: Select tasks from the low-priority task list in sequence, scan all candidate idle nodes in the heterogeneous computing power resource pool that are not the current node, and filter out the nodes that simultaneously meet the task computing volume requirements and response latency limits. S54: Sort the qualified nodes in descending order according to their idle instruction throughput rate, select the optimal node as the target migration node, and send a task migration instruction packet. The task migration instruction packet includes a task context snapshot, input data path mapping information, and recovery image information. S55: After receiving the instruction packet, the target migration node completes data mapping and computing image preparation, loads the original task context, and starts the migration without interrupting the task process. S56: The scheduling manager updates the task scheduling table and the resource mapping table, marking the completion of an effective task migration.
9. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 8, characterized in that The preset load rate threshold has the following calculation formula: , where represents the preset load rate threshold of the target resource node; represents the average load rate of the node in the most recent P scheduling cycles; represents the standard deviation of the load rate of the node in the most recent P scheduling cycles; is the margin coefficient.
10. The heterogeneous computing power resource dynamic allocation method for ocean supercomputing environment according to claim 1, characterized in that The specific steps of S6 are as follows: S61: After completing the task allocation for the current batch, the scheduling manager forms a one-to-one mapping between each task and its allocated target resource node, constructs a resource allocation mapping table. The mapping table includes a task identifier, a resource node number, a priority value, a scheduling timestamp, and an estimated computing time consumption. S62: After the task is completed, the target resource node sends back the task completion status to the scheduling manager through the node management interface, including the task identifier, the actual start timestamp, and the actual end timestamp. The scheduling manager calculates the task completion time based on this. S63: During the task execution process, the resource monitoring agent of the target resource node periodically collects the local CPU utilization data. After the task is completed, the scheduling manager performs numerical integration and averaging on the CPU utilization data during this period to calculate the average computing power utilization during the task execution. S64: The scheduling manager organizes the data of the target resource number, task completion time, and average computing power utilization corresponding to each task to generate a structured resource allocation result report and outputs it in JSON format.
Citation Information
Patent Citations
Edge computing task processing method, edge server and storage medium
CN115061800A
Real-time quality monitoring method and system in production process of low-voltage power distribution cabinet
CN118966881A
Load-aware scheduling method based on deep learning
CN119065835A
Scheduling automation system application state management method
CN119292745A
Cluster load balancing processing method based on cloud computing
CN119718688A
Cited By
High-speed rail multi-band communication dynamic computing power optimization method and system
CN120568400A
Adaptive computing network integrated arranging and scheduling method based on load and SLA (Service Level Agreement)
CN121603396A