Weather-based HPC resource scheduling method and system
Through the accurate classification of meteorological data, scientific division of task priorities, intelligent allocation of resources and dynamic monitoring, the problem of uneven resource allocation in HPC resource scheduling is solved, efficient and timely resource scheduling is achieved, and the smooth progress of meteorological business and scientific research is ensured.
Patent Information
- Application Number
- CN202510316614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
The existing HPC resource scheduling methods do not fully consider the spatio-temporal characteristics of meteorological data, resulting in uneven resource allocation and low scheduling efficiency, which cannot meet the needs of meteorological services and scientific research for high precision and timeliness.
Through clustering analysis, high-frequency accessed meteorological data blocks are identified and stored in the cache area. A task priority list is generated in combination with the hierarchical analysis method, resource allocation is used to use the first adaptation algorithm, and dynamic load balancing is achieved through distributed probes, meteorological spatiotemporal distribution coordinate system is constructed, and resource allocation is adjusted using convolutional neural network.
It realizes efficient and dynamic allocation of HPC system resources, improves scheduling efficiency, ensures timely processing of meteorological tasks and stable operation of the system, and meets the high-precision needs of meteorological business and scientific research.
Smart Images

Figure CN120256100A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource scheduling, and specifically to a weather-based HPC resource scheduling method and system. Background Art
[0002] With the deep integration of meteorological science and high-performance computing (HPC), weather-based HPC resource scheduling has gradually become the mainstream technology. Different from general big data, meteorological data has strong spatio-temporal characteristics. From the perspective of spatial characteristics, the meteorological data collected from all over the world covers diversified information. At the same time, from the perspective of temporal characteristics, meteorological data is continuously generated at different frequencies, and meteorological tasks with different time scales have different requirements for processing time and timeliness.
[0003] Therefore, based on the spatial and temporal characteristics of meteorological data, the HPC system faces challenges in resource allocation. In data management, the difference in access frequency of meteorological data is not taken seriously, and the speed of retrieving high-frequency meteorological data is slow; the division of task priorities lacks systematicness, and data dependencies and resource characteristics between tasks are often ignored; resource allocation is not flexible, and node load imbalance is likely to occur in the face of sudden tasks, and it is difficult to adapt to changes in resource requirements. The resource scheduling is out of touch with the spatio-temporal laws of meteorological data, and the overall scheduling efficiency is low and the effect is poor, making it difficult to meet the advanced needs of meteorological business and scientific research.
[0004] Therefore, there is an urgent need to design a method and system for efficiently and reasonably allocating resources when executing meteorological tasks according to the spatio-temporal characteristics of meteorological data. Summary of the Invention
[0005] In order to at least overcome the above deficiencies in the prior art, the purpose of the present application is to provide a weather-based HPC resource scheduling method and system.
[0006] In the first aspect, an embodiment of the present application provides a weather-based HPC resource scheduling method, which includes the following steps:
[0007] Based on meteorological data and its corresponding access frequency, using a clustering analysis algorithm, classify the meteorological data, identify meteorological data blocks with high-frequency access, and store the meteorological data blocks in the cache area close to the computing node to form a high-frequency data cache set;
[0008] Use the analytic hierarchy process to analyze the timeliness requirements, data dependency relationships, and resource demand characteristics of each meteorological task in the high-frequency data cache set to generate a task priority list;
[0009] Use the first-fit algorithm to allocate resources to the task priority list, thereby generating a resource scheduling plan;
[0010] Based on the resource scheduling scheme, distributed probes are deployed to the computing nodes at all levels of the HPC system to monitor the CPU usage rate, memory occupancy, and network bandwidth consumption of the HPC system in real time. According to the preset load balancing threshold, the task allocation is dynamically adjusted to the computing nodes with lighter loads, thereby generating a dynamic load balancing mechanism;
[0011] The meteorological data blocks concentrated in the high-frequency data cache are divided into three-dimensional grids according to region, altitude, and time to form a meteorological spatio-temporal distribution coordinate system;
[0012] The resource occupancy ratio, the number of idle resources, and the coordinate values corresponding to each meteorological block in the meteorological spatio-temporal distribution coordinate system during the execution of meteorological tasks in the task priority list are sent to the trained convolutional neural network model for prediction, and the resource configuration is adjusted secondly according to the prediction results to complete the resource scheduling of meteorological tasks.
[0013] The purpose of the present invention is to overcome the problems that the traditional weather-based HPC resource scheduling method does not fully consider the spatio-temporal characteristics of meteorological data, resulting in uneven resource allocation and low scheduling efficiency, and being unable to meet the requirements of meteorological services and scientific research for high-precision and timely resource scheduling. Based on this, the inventors of this application provide a weather-based HPC resource scheduling method and system. Through the precise classification of meteorological data, the scientific division of task priorities, the intelligent allocation of resources, and the secondary optimization based on dynamic monitoring, it is ensured that the resources of the HPC system can be allocated in real time and efficiently according to the needs of meteorological tasks, greatly improving the overall scheduling efficiency and the effectiveness of meteorological task processing, and ensuring the stable operation of meteorological services and the smooth progress of scientific research work.
[0014] To address the above-mentioned drawbacks, this application realizes the optimal scheduling of resources through a series of innovative algorithm combinations and refined process designs. First, according to the access frequency of meteorological data, the clustering analysis algorithm is used to identify high-frequency data and store it in the high-speed cache area. In this embodiment, the high-speed cache area is a high-speed storage area built based on a solid-state drive (SSD), which has the advantages of fast read and write speeds and fast access speeds, can greatly reduce data read latency, make the high-frequency access data more rapid during IO data transmission, and greatly improve the access efficiency of high-frequency data.
[0015] Next, the Analytic Hierarchy Process is used to comprehensively analyze the characteristics of meteorological tasks and construct a reasonable task priority list. For the analysis of the time - efficiency requirements of meteorological tasks, not only the deadlines set by the tasks themselves are considered, but also the development trends of meteorological phenomena are taken into account. For example, tasks related to predicting rapidly developing severe convective weather are given higher priorities to ensure that urgent and critical tasks are processed first. When considering data - dependency relationships, a detailed task - dependency graph is constructed to clearly show the data flow and mutual constraints between tasks, avoiding task delays caused by insufficient data preparation. For instance, before conducting regional climate simulation tasks, the prerequisite tasks such as basic meteorological data collection and pre - processing required are accurately identified, and resources are reasonably allocated according to the priority order. Regarding the characteristics of resource requirements, the CPU core count, memory capacity, and network bandwidth requirements of each task are carefully evaluated.
[0016] Then, the First - Fit algorithm is used to intelligently allocate resources according to the priorities, generating a preliminary resource - scheduling plan. The First - Fit algorithm starts allocating resources from the task with the highest priority, first meeting the full requirements of high - priority meteorological tasks for CPU core count, memory, and network bandwidth, and then allocating to other meteorological tasks in turn.
[0017] Then, by introducing distributed probes and a dynamic load - balancing mechanism, the system status is monitored in real - time. When a node overload is detected, the task - resource allocation is immediately adjusted to maintain system stability. The distributed probes are evenly deployed at each key node of the HPC system, including computing nodes, storage nodes, network switching nodes, etc. They collect core metrics such as CPU usage, memory occupancy, and network - bandwidth consumption of the nodes at 5 - second intervals and transmit the data back to the monitoring center in real - time. The monitoring center makes accurate judgments based on preset load - balancing thresholds. When it is detected that the CPU usage of a certain computing node exceeds 80%, memory occupancy exceeds 70%, or network - bandwidth consumption exceeds 90% and other preset overload thresholds, the node is immediately marked as an overloaded node. At the same time, the resource - preemption mechanism and the listening mechanism are quickly started. When the resource - preemption mechanism is executed, according to the task - priority list, low - priority meteorological tasks are quickly located, the difference between the resources required by the current meteorological task and the remaining resources of the overloaded node is accurately calculated, and part of the resources already allocated to the low - priority meteorological tasks are skillfully recovered according to the preset 30% resource - recovery ratio and preferentially supplied to the current meteorological task. The listening mechanism runs synchronously, continuously monitoring the execution status of the previous and / or multiple previous meteorological tasks. When it is detected that a task releases resources, all or part of the released resources are allocated to the current meteorological task according to the dynamic resource - adaptation formula until the load returns to the normal range, effectively maintaining the stable and efficient operation of the system.
[0018] Finally, combined with the resource information during the execution of meteorological tasks in the task priority list, the coordinate values of the meteorological spatio-temporal distribution coordinate system, and the convolutional neural network model, predict the trend of resource demand, and perform a second fine adjustment of resource allocation to make the resource scheduling closely fit the dynamic changes of meteorological tasks. When constructing the meteorological spatio-temporal distribution coordinate system, obtain the geographical information of the meteorological data blocks in the high-frequency data cache set. Using a 0.5-degree longitude and latitude difference as the division granularity, divide the globe or a specified area into multiple sub-regions, and each sub-region corresponds to a unique geographical code; according to the collection height of each sub-region, divide the height interval at intervals of 1500 meters, and classify the meteorological data blocks in each sub-region into the corresponding intervals according to the height interval, and at the same time label the corresponding height identifier for each height interval, where the height identifier includes the ground layer, the low-altitude layer, the mid-altitude layer, and the high-altitude layer; at a preset time interval of 1 hour, sort out the timestamp information of the meteorological data blocks in each sub-region, group the meteorological data blocks belonging to the same time period into one group, and calibrate the time serial number for each group of data; perform three-dimensional coordinate construction on the meteorological data blocks after being divided by geographical location, height, and time dimensions to form a meteorological spatio-temporal distribution coordinate system; transmit the resource occupancy ratio, the number of idle resources, and the coordinate values in the meteorological spatio-temporal distribution coordinate system of the resource allocation model during the execution of meteorological tasks to the trained convolutional neural network model for prediction. The convolutional layer of the convolutional neural network model uses a 3×3 small convolutional kernel to extract input features, calculates the extracted features through the ReLU activation function, the pooling layer uses a 2×2 maximum pooling operation to reduce the dimension of the input features after activation processing, and the fully connected layer integrates the input features after convolutional and pooling processing, and predicts the storage change rate of the cache area within 12 hours, and predicts the change ratio of the computing resource requirements of each meteorological task within 3 hours. Compare the predicted storage change rate of the cache area within 12 hours with the preset storage change rate. If it is less than the preset storage change rate, start the data cleaning mechanism, traverse the high-frequency data cache set, and remove the 20% of the meteorological data blocks closest to the expiration time from the cache area to free up space for the meteorological data blocks to be newly stored in the cache area; compare the predicted change ratio of the computing resource requirements of each meteorological task within 3 hours with the preset change threshold. If the change ratio of the computing resource requirements of the current meteorological task is greater than the preset change threshold, start the resource allocation mechanism at this time, and recycle part of the resources from other low-priority meteorological tasks and resource pools in a ratio of 1:3 and allocate them to the current meteorological task to overcome the disadvantages of traditional resource scheduling for big data in all aspects.
[0019] Further, the specific steps for forming a high-frequency data cache set based on meteorological data and its corresponding access frequency include:
[0020] Read the storage file of meteorological data, and through data cleaning, remove the invalid values and duplicate values in the meteorological data to generate a meteorological data set;
[0021] Feature extraction is performed on the meteorological data set to form meteorological data features. The KMeans clustering algorithm is used to classify the meteorological data according to the types of meteorological data features, and meteorological data clustering subsets are generated.
[0022] The access frequency of data points in each clustering subset is statistically counted. Combining with the timestamp information corresponding to the data points, the access times of the data points are statistically counted. Sorted according to the access times, the data points with an access frequency higher than the set access frequency threshold are screened out and labeled as high-frequency accessed meteorological data blocks.
[0023] The high-speed cache area close to the computing node is determined through the system configuration file, and a cache partition for storing high-frequency meteorological data is divided. Using the mounting instruction of the system configuration file, the meteorological data blocks are stored in the high-speed cache area to form a high-frequency data cache set.
[0024] According to the access frequency of data points in each clustering subset, a time limit is set for the meteorological data blocks stored in the high-speed cache area, and the meteorological data blocks are removed from the high-frequency data cache set after the expiration.
[0025] Furthermore, the specific steps for generating a task priority list using the analytic hierarchy process include:
[0026] Parse each meteorological task in the high-frequency data cache set, extract the key information of the meteorological task to construct a task information list, where the key information includes the task name, geographical area range, timeliness requirement, and task execution duration.
[0027] Each meteorological task is divided into four levels: urgent, high, medium, and low according to the timeliness requirement, and a preset reference time range is set for each level in combination with the task execution duration.
[0028] Traverse the high-frequency data cache set, determine the data dependency relationship of each meteorological task, and build a task dependency graph according to the dependency relationship.
[0029] According to the data dependency relationship, timeliness requirement, and task execution duration in the task dependency graph, a standardized scoring mechanism is used to evaluate the resource requirements of each meteorological task, and a corresponding resource configuration weight is configured for each meteorological task as its resource requirement characteristic.
[0030] Taking the timeliness requirement, data dependency relationship, and resource requirement characteristic as the analysis indicators of the analytic hierarchy process, the analytic hierarchy process is used to sort the task priorities of each meteorological task, thereby generating a task priority list.
[0031] Furthermore, the specific steps for generating a resource scheduling scheme using the first fit algorithm include:
[0032] Initialize the current resource allocation status of the HPC system, obtain the current available resource information, and transfer the available resource information to the resource pool. Among them, the available resource information includes the number of idle CPU cores, available memory space, and unoccupied network bandwidth;
[0033] Traverse the task priority list in descending order of task priority. According to the resource requirement characteristics of each meteorological task, use the first-fit algorithm to first allocate the resources in the resource pool to the meteorological tasks with higher priority, and then allocate them to other meteorological tasks in turn.
[0034] Furthermore, the specific steps for generating a dynamic load balancing mechanism based on the resource scheduling scheme include:
[0035] Based on the resource scheduling scheme, start the resource allocation process. Use distributed probes to collect data on CPU usage, memory occupancy, and network bandwidth consumption of each level of computing nodes in the HPC system in real time, and transmit the data back to the monitoring center in real time;
[0036] The monitoring center analyzes and judges the backhaul data based on the preset load balancing threshold. When it is detected that the computing node corresponding to the current meteorological task exceeds the load threshold, mark the node as an overloaded node, and at the same time continuously start the resource preemption mechanism and the listening mechanism;
[0037] Execute the resource preemption mechanism. According to the task priority list, quickly locate the meteorological tasks with low priority, calculate the difference between the resources required by the current meteorological task and the remaining resources of the overloaded node, and accurately recover some resources from the resources already allocated to the low-priority meteorological tasks according to the preset resource recovery ratio, and give priority to supplementing the current meteorological task;
[0038] Synchronously run the listening mechanism, continuously monitor the execution status of the previous and / or previous multiple meteorological tasks. When it is detected that a task releases resources, allocate all or part of the released resources to the current meteorological task according to the dynamic resource adaptation formula until the load returns to the normal range, and finally generate a dynamic load balancing mechanism.
[0039] Furthermore, the expression of the dynamic resource adaptation formula is:
[0040]
[0041] In the formula, R allocated - represents the amount of resources allocated to the current meteorological task, R r - represents the amount of resources that have been recovered, R g - represents the resource gap of the current meteorological task, R n - represents the amount of newly released resources, ω cpu 、ω mem and ω bwrespectively represent the resource allocation weights of the current task on meteorological data CPU, memory, and network bandwidth. and respectively represent the available resource amounts of the newly released resources on CPU, memory, and network bandwidth.
[0042] Furthermore, the specific steps for forming the meteorological spatio-temporal distribution coordinate system include:
[0043] Obtain the regional information of the meteorological data blocks in the high-frequency data cache set, use the longitude and latitude difference as the division granularity, divide the globe or a specified area into multiple sub-regions, and each sub-region corresponds to a unique regional code;
[0044] Divide the height intervals according to the collection height of each sub-region, classify the meteorological data blocks in each sub-region into the corresponding intervals according to the height intervals, and at the same time label the corresponding height identifiers for each height interval, where the height identifiers include the ground layer, low-altitude layer, mid-altitude layer, and high-altitude layer;
[0045] At a preset time interval, sort out the timestamp information of the meteorological data blocks in each sub-region, group the meteorological data blocks belonging to the same time period into one group, and calibrate a time serial number for each group of data;
[0046] Construct a three-dimensional coordinate for the meteorological data blocks divided by region, height, and time dimensions, so as to form a meteorological spatio-temporal distribution coordinate system. Among them, the X-axis represents the regional code value, the Y-axis represents the height identifier, and the Z-axis represents the time serial number.
[0047] Furthermore, the specific steps for the secondary adjustment of resource allocation include:
[0048] Extract the computing resource ratio information and storage resource ratio information corresponding to each meteorological data block of the meteorological tasks being executed in the resource allocation model, and at the same time obtain the idle resource quantity information in the resource allocation model, determine the coordinate values corresponding to each meteorological data block in the meteorological spatio-temporal distribution coordinate system, and summarize and integrate the resource ratio of each meteorological data block, its corresponding coordinate value information, and the idle resource quantity to form a complete meteorological data set;
[0049] Assign the integrated meteorological data set as input features to the trained convolutional neural network model;
[0050] The convolutional layer of the convolutional neural network model uses a small 3×3 convolutional kernel to extract input features, calculates the extracted features through the ReLU activation function. The pooling layer uses a 2×2 max pooling operation to reduce the dimension of the input features after activation processing. The fully connected layer integrates the input features after convolution and pooling processing, and predicts the storage change rate of the cache area within 12 hours, as well as predicts the change ratio of the computing resource requirements of each meteorological task within 3 hours;
[0051] Compare the predicted storage change rate of the cache area within 12 hours with the preset storage change rate. If it is less than the preset storage change rate, start the data cleaning mechanism, traverse the high-frequency data cache set, and remove the 20% of the meteorological data blocks closest to the expiration time from the cache area;
[0052] Compare the predicted change ratio of the computing resource requirements of each meteorological task within 3 hours with the preset change threshold. If the change ratio of the computing resource requirements of the current meteorological task is greater than the preset change threshold, start the resource allocation mechanism at this time, and recycle part of the resources from other low-priority meteorological tasks and resource pools according to a ratio of 1:3 and allocate them to the current meteorological task.
[0053] In a second aspect, the application example provides a system using any of the above weather-based HPC resource scheduling methods, including:
[0054] A high-frequency data storage unit, configured to store meteorological data blocks with an access frequency higher than the set access frequency threshold in the cache area;
[0055] A task priority list generation unit, configured to generate a task priority list for each meteorological data block in the high-frequency data cache set according to the time limit requirements, data dependency relationships, and resource requirement characteristics of its meteorological tasks;
[0056] A resource allocation unit, configured to perform the first resource allocation on the task priority list to generate a resource scheduling plan;
[0057] A load balancing monitoring unit, configured to dynamically adjust the resource allocation of overloaded nodes based on the resource scheduling plan to generate a dynamic load balancing mechanism;
[0058] A spatio-temporal coordinate system construction unit, configured to perform three-dimensional grid division on the meteorological data blocks in the high-frequency data cache set according to region, altitude, and time to form a meteorological spatio-temporal distribution coordinate system;
[0059] A resource secondary adjustment unit, configured to send the resource occupancy ratio, the number of idle resources, and the coordinate values in the meteorological spatio-temporal distribution coordinate system of the resource configuration model when executing meteorological tasks to the trained convolutional neural network model for prediction, and perform secondary adjustment of resource configuration according to the prediction results.
[0060] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0061] A weather-based HPC resource scheduling method and system according to the present invention, through the above technical solutions, on the one hand, uses clustering analysis to accurately locate high-frequency access meteorological data and quickly store it, greatly improving the data reading efficiency; uses the analytic hierarchy process to deeply analyze the task characteristics, scientifically plan the task priorities, ensure that urgent and critical meteorological tasks are processed first, and effectively avoid task jams; then combines the first-fit algorithm to reasonably allocate initial resources to ensure the maximization of resource utilization. On the other hand, it uses distributed probes to monitor the system load in real time, dynamically adjusts the task allocation, and maintains the stable and efficient operation of the system; and by constructing a meteorological spatio-temporal distribution coordinate system and combining it with a convolutional neural network model to predict the resource trend, it can anticipate the changes in the storage and computing resource requirements in advance, clean the cache and allocate resources in a timely manner, so that the resource scheduling closely conforms to the dynamic changes of meteorological services. This all-round and refined scheduling strategy avoids the problems of uneven resource allocation and low scheduling efficiency in traditional methods, and effectively meets the advanced requirements of meteorological services and scientific research for high-precision and timely resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not limit the embodiments of the present invention. In the drawings:
[0063] Figure 1 is a flowchart of the scheduling method of the present invention;
[0064] Figure 2 is a flowchart of forming a high-frequency data cache set of the present invention;
[0065] Figure 3 is a flowchart of generating a task priority list of the present invention;
[0066] Figure 4 is a flowchart of generating a dynamic load balancing mechanism of the present invention;
[0067] Figure 5 is a flowchart of forming a meteorological spatio-temporal distribution coordinate system of the present invention;
[0068] Figure 6 is a flowchart of performing secondary adjustment of resource configuration of the present invention;
[0069] Figure 7 is a schematic structural diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the embodiments and the accompanying drawings. The illustrative embodiments and descriptions thereof of the present invention are only used to explain the present invention and shall not be construed as limiting the present invention.
[0071] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as limiting the present invention.
[0072] In addition, the terms "first" and "second" are only used for descriptive purposes and shall not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0073] In the present invention, unless otherwise clearly specified and defined, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0074] In the present invention, unless otherwise clearly specified and defined, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely indicates that the horizontal height of the first feature is lower than that of the second feature.
[0075] In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of the disclosure of the present invention.
[0076] In the following description, suffixes such as "module", "component", "assembly" or "unit" are only used to facilitate the description of the present invention, and they have no specific meaning by themselves. Therefore, they can be used interchangeably.
[0077] The present invention will be further described in detail below in conjunction with specific embodiments and the accompanying drawings.
[0078] Please refer to Figure 1 , which is a schematic flowchart of a weather-based HPC resource scheduling method provided by an embodiment of the present invention. Further, the specific steps of the weather-based HPC resource scheduling method include:
[0079] Based on meteorological data and its corresponding access frequencies, using a clustering analysis algorithm, classify the meteorological data, identify meteorological data blocks with high-frequency access, and store the meteorological data blocks in a cache area close to the computing nodes to form a high-frequency data cache set;
[0080] Use the analytic hierarchy process to analyze the timeliness requirements, data dependency relationships, and resource demand characteristics of each meteorological task in the high-frequency data cache set to generate a task priority list;
[0081] Use the first-fit algorithm to allocate resources to the task priority list to generate a resource scheduling plan;
[0082] Based on the resource scheduling plan, deploy distributed probes to the computing nodes at all levels of the HPC system, monitor the CPU usage rate, memory occupancy, and network bandwidth consumption of the HPC system in real time, and dynamically adjust the task allocation to the computing nodes with lighter loads according to a preset load balancing threshold to generate a dynamic load balancing mechanism;
[0083] Divide the meteorological data blocks in the high-frequency data cache set into three-dimensional grids according to region, altitude, and time to form a meteorological spatio-temporal distribution coordinate system;
[0084] Send the resource occupancy ratio, the number of idle resources, and the coordinate values corresponding to each meteorological block in the meteorological spatio-temporal distribution coordinate system of each meteorological task in the task priority list to a trained convolutional neural network model for prediction, and perform a secondary adjustment of resource allocation according to the prediction results to complete the resource scheduling of meteorological tasks.
[0085] In the implementation of this embodiment, according to the access frequency of meteorological data, the clustering analysis algorithm is used to identify high-frequency data and store it in the cache area. In this embodiment, the cache area is a high-speed storage area built based on a solid-state drive (SSD), which has the advantages of fast read and write speed and fast access speed, can greatly reduce data read latency, make high-frequency access data more rapid during IO data transmission, and significantly improve the access efficiency of high-frequency data.
[0086] Next, the analytic hierarchy process is used to comprehensively analyze the characteristics of meteorological tasks and construct a reasonable task priority list. For the analysis of the time-effect requirements of meteorological tasks, not only the deadline set by the task itself is considered, but also the development trend of meteorological phenomena is comprehensively considered; for example, tasks related to predicting the upcoming rapidly developing severe convective weather are given higher priorities to ensure that urgent and critical tasks are processed first; when considering data dependencies, by constructing a detailed task dependency graph, the data flow and mutual restriction relationships between each task are clearly presented to avoid task jams caused by insufficient data preparation; for example, before performing the regional climate simulation task, the prerequisite tasks such as the collection and preprocessing of the basic meteorological data required are accurately identified, and resources are reasonably arranged according to the priority order; for the resource demand characteristics, the requirements of each task for the number of CPU cores, memory capacity, and network bandwidth are carefully evaluated.
[0087] Then, the first-fit algorithm is used to intelligently allocate resources according to the priority to generate a preliminary resource scheduling plan. The first-fit algorithm starts allocating resources from the task with the highest priority, first meeting the full requirements of high-priority meteorological tasks for the number of CPU cores, memory, and network bandwidth, and then allocating them to other meteorological tasks in turn.
[0088] Then, by introducing a distributed probe and a dynamic load balancing mechanism, the system status is monitored in real time. When a node overload is detected, the task resource allocation is immediately adjusted to maintain system stability. The distributed probes are evenly deployed on each key node of the HPC system, including computing nodes, storage nodes, network switching nodes, etc. They collect core metrics such as CPU usage, memory occupancy, and network bandwidth consumption of the nodes at an interval of every 5 seconds and transmit the data back to the monitoring center in real time. The monitoring center makes accurate judgments based on the preset load balancing thresholds. When it is detected that the CPU usage of a certain computing node exceeds 80%, the memory occupancy exceeds 70%, or the network bandwidth consumption exceeds 90% and other predefined overload thresholds, the node is immediately marked as an overloaded node. At the same time, the resource preemption mechanism and the listening mechanism are quickly started. When the resource preemption mechanism is executed, according to the task priority list, the low-priority meteorological tasks are quickly located, and the difference between the resources required for the current meteorological task and the remaining resources of the overloaded node is accurately calculated. According to the preset 30% resource recovery ratio, some resources are skillfully recovered from the resources already allocated to the low-priority meteorological tasks and preferentially supplied to the current meteorological task. The listening mechanism runs synchronously, continuously monitoring the execution status of the previous and / or previous multiple meteorological tasks. When it is detected that a task releases resources, all or part of the released resources are allocated to the current meteorological task according to the dynamic resource adaptation formula until the load returns to the normal range, effectively maintaining the stable and efficient operation of the system.
[0089] Finally, combined with the resource information during the execution of meteorological tasks, the coordinate values of the meteorological spatio-temporal distribution coordinate system, and the convolutional neural network model in the task priority list, predict the trend of resource demand, and perform a second fine-tuning of resource allocation to make the resource scheduling closely fit the dynamic changes of meteorological tasks. When constructing the meteorological spatio-temporal distribution coordinate system, obtain the geographical information of the meteorological data blocks in the high-frequency data cache set. Using a 0.5-degree longitude and latitude difference as the division granularity, divide the globe or a specified area into multiple sub-regions, and each sub-region corresponds to a unique geographical code; according to the acquisition height of each sub-region, divide the height intervals at intervals of 1500 meters, and classify the meteorological data blocks in each sub-region into the corresponding intervals according to the height intervals, and at the same time label the corresponding height identifiers for each height interval, where the height identifiers include the ground layer, the low-altitude layer, the middle-altitude layer, and the high-altitude layer; at a preset time interval of 1 hour, sort out the timestamp information of the meteorological data blocks in each sub-region, group the meteorological data blocks belonging to the same time period into one group, and label a time sequence number for each group of data; perform three-dimensional coordinate construction on the meteorological data blocks divided by geographical location, height, and time dimensions to form a meteorological spatio-temporal distribution coordinate system; send the resource occupancy ratio, the number of idle resources of the resource allocation model during the execution of meteorological tasks, and the coordinate values in the meteorological spatio-temporal distribution coordinate system to the trained convolutional neural network model for prediction. The convolutional layer of the convolutional neural network model uses a 3×3 small convolutional kernel to extract input features, calculates the extracted features through the ReLU activation function, the pooling layer uses a 2×2 maximum pooling operation to reduce the dimension of the input features after activation processing, and the fully connected layer integrates the input features after convolution and pooling processing, and predicts the storage change rate of the cache area within 12 hours, and predicts the change ratio of the computing resource requirements of each meteorological task within 3 hours. Compare the predicted storage change rate of the cache area within 12 hours with the preset storage change rate. If it is less than the preset storage change rate, start the data cleaning mechanism, traverse the high-frequency data cache set, and remove 20% of the meteorological data blocks closest to the expiration time from the cache area to make room for the meteorological data blocks to be newly stored in the cache area; compare the predicted change ratio of the computing resource requirements of each meteorological task within 3 hours with the preset change threshold. If the change ratio of the computing resource requirements of the current meteorological task is greater than the preset change threshold, start the resource allocation mechanism at this time, and recycle part of the resources from other low-priority meteorological tasks and resource pools in a ratio of 1:3 and allocate them to the current meteorological task to overcome the disadvantages of traditional resource scheduling for big data in all aspects.
[0090] In one possible implementation, please refer to Figure 2 , the specific steps for forming the high-frequency data cache set based on meteorological data and its corresponding access frequency include:
[0091] Read the storage file of the meteorological data, and through data cleaning, remove the invalid values and duplicate values in the meteorological data to generate a meteorological data set;
[0092] Extract features from the meteorological data set to form meteorological data features. Use the KMeans clustering algorithm to classify the meteorological data according to the types of the meteorological data features, and generate meteorological data clustering subsets;
[0093] Statistically analyze the access frequency of data points in each clustering subset. Combine the timestamp information corresponding to the data points to count the access times of the data points. Sort according to the access times, filter out the data points with an access frequency higher than the set access frequency threshold, and label them as high-frequency accessed meteorological data blocks;
[0094] Determine the high-speed cache area close to the computing node through the system configuration file, divide a cache partition for storing high-frequency meteorological data, and use the mounting instruction of the system configuration file to store the meteorological data blocks in the high-speed cache area to form a high-frequency data cache set.
[0095] Calibrate the expiration date of the meteorological data blocks stored in the high-speed cache area according to the access frequency of data points in each clustering subset. After expiration, remove the meteorological data blocks from the high-frequency data cache set.
[0096] In the implementation of this embodiment, the high-speed storage area is a hierarchical storage structure, which is specifically divided into four storage structures: the upper layer, the upper middle layer, the lower middle layer, and the lower layer from top to bottom. The division basis is the time interval of the meteorological data block from the expiration date. Among them, the lower layer stores the set of meteorological data blocks within 5 days from the expiration date, the lower middle layer stores the set of meteorological data blocks within 10 days from the expiration date, the upper middle layer stores the set of meteorological data blocks within 15 days from the expiration date, and the upper layer stores the set of meteorological data blocks within 20 days from the expiration date; when some meteorological data blocks in the lower layer reach the expiration date, the system automatically clears them from the bottom layer. As time goes by, in the set of meteorological data blocks stored in the lower middle layer, if the expiration date of some data blocks changes from within 10 days to within 5 days, the set of meteorological data blocks that meet this condition will migrate from the lower middle layer to the lower layer; similarly, in the set of meteorological data blocks stored in the upper middle layer, if the expiration date of some data blocks changes from within 15 days to within 10 days, these sets of meteorological data blocks will migrate from the upper middle layer to the lower middle layer, and so on. The set of meteorological data blocks in the upper layer also migrates to the upper middle layer according to the same time limit rule. The cache area vacated in the upper cache area due to data block migration will be used to store new meteorological data blocks. Through this way of hierarchical storage and dynamic migration, the storage management of the high-speed cache area can be further optimized, and the storage and access efficiency of meteorological data can be improved.
[0097] In a possible implementation, please refer to Figure 3 for details. The specific steps for generating the task priority list using the analytic hierarchy process include:
[0098] Parse each meteorological task in the high-frequency data cache set, extract the key information of the meteorological task to construct a task information list, where the key information includes the task name, geographical area range, time limit requirement, and task execution duration;
[0099] Divide each meteorological task into four levels: urgent, high, medium, and low according to the time limit requirement, and set a preset reference time range for each level in combination with the task execution duration;
[0100] Traverse the high-frequency data cache set, determine the data dependency relationship of each meteorological task, and build a task dependency graph according to the dependency relationship;
[0101] According to the data dependency relationship, time limit requirement, and task execution duration in the task dependency graph, use a standardized scoring mechanism to evaluate the resource requirements of each meteorological task, and configure a corresponding resource allocation weight for each meteorological task as its resource requirement characteristic;
[0102] Take the time limit requirement, data dependency relationship, and resource requirement characteristic as the analysis indicators of the analytic hierarchy process, and use the analytic hierarchy process to sort the task priorities of each meteorological task, so as to generate a task priority list.
[0103] When this embodiment is implemented, the execution process of the standardized scoring mechanism includes: for a meteorological task that is at the core computing node of the critical data processing link and has an urgent time limit, such as real-time typhoon path prediction, since it needs to quickly process a large amount of complex data, the basic score for the CPU core number requirement is set to 8-10 points; for a meteorological task at the edge branch, such as historical meteorological data archiving, the basic score for the memory capacity is set to 2-4 points; for a meteorological task such as global climate simulation with a large amount of data processing and high memory requirements, the basic score for the memory capacity is set to 9-10 points; for tasks such as local urban short-term air quality monitoring, the basic score for the memory capacity is set to 3-5 points; in terms of network bandwidth requirements, for a meteorological task that receives and distributes satellite meteorological observation data streams in real time, due to the high requirement for data timeliness, the basic score is 8-10 points; for a meteorological task mainly based on local data, such as statistical analysis based on local historical data, the basic score is 2-4 points; then, through a preset weighting algorithm, combine the basic scores of each meteorological task to calculate the resource allocation weight of each meteorological task to adapt to its resource requirements and ensure the scheduling efficiency.
[0104] In the implementation of this embodiment, the task priority list can be a serial linked list or a binary tree linked list structure. When the amount of meteorological task data is within 100, and the basis for determining task priorities is relatively simple, mainly arranged in the chronological order of task initial entry into the system, without complex data dependencies and resource competition relationships, a serial linked list is adopted. The nodes of the serial linked list are connected in series in sequence; when the amount of meteorological task data exceeds 100, and task priorities need to comprehensively consider timeliness urgency, there are complex nested dependency relationships between tasks, and operations such as frequent high-priority task queue-jumping and dynamic adjustment of associated task priorities are involved, a binary tree linked list is selected. The binary tree linked list constructed in the form of a binary search tree can perform efficient binary sorting based on task key attribute values, and the average time complexity of searching, inserting, and deleting task nodes is O(logn). It can quickly locate tasks in large-data-volume and high-complexity task scenarios, accurately allocate resources, and effectively improve the system's ability to handle complex task scheduling.
[0105] In one possible implementation manner, the specific steps of generating the resource scheduling plan by using the first-fit algorithm include:
[0106] Initialize the current resource allocation status of the HPC system, obtain the current available resource information, and send the available resource information to the resource pool, where the available resource information includes the number of idle CPU cores, available memory space, and unoccupied network bandwidth;
[0107] Traverse the task priority list in descending order of task priority. According to the resource requirement characteristics of each meteorological task, use the first-fit algorithm to preferentially allocate the resources in the resource pool to high-priority meteorological tasks, and then allocate them to other meteorological tasks in sequence.
[0108] In one possible implementation manner, please refer to Figure 4 for reference. The specific steps of generating the dynamic load balancing mechanism based on the resource scheduling plan include:
[0109] Based on the resource scheduling plan, start the resource allocation process, use distributed probes to collect data on CPU usage, memory occupancy, and network bandwidth consumption of each level of computing nodes in the HPC system in real time, and transmit the data back to the monitoring center in real time;
[0110] The monitoring center analyzes and judges the transmitted data based on a preset load balancing threshold. When it is detected that the computing node corresponding to the current meteorological task exceeds the load threshold, mark the node as an overloaded node, and at the same time continuously start the resource preemption mechanism and the listening mechanism;
[0111] Execute the resource preemption mechanism. According to the task priority list, quickly locate the low-priority meteorological tasks, calculate the difference between the resources required by the current meteorological task and the remaining resources of the overloaded nodes, and accurately reclaim a part of the resources from the resources already allocated to the low-priority meteorological tasks according to the preset resource recovery ratio, and preferentially supply the current meteorological task;
[0112] Synchronously run the monitoring mechanism, continuously monitor the execution status of the previous and / or previous multiple meteorological tasks. When it is detected that a task releases resources, all or part of the released resources are allocated to the current meteorological task according to the dynamic resource adaptation formula until the load returns to the normal range, and finally a dynamic load balancing mechanism is generated.
[0113] In a possible implementation manner, the expression of the dynamic resource adaptation formula is:
[0114]
[0115] In the formula, R allocated - represents the amount of resources allocated to the current meteorological task, R r - represents the amount of resources that have been reclaimed, R g - represents the resource gap of the current meteorological task, R n - represents the amount of newly released resources, ω cpu 、ω mem and ω bw respectively represent the resource configuration weights of the current task on the CPU, memory, and network bandwidth of meteorological data, and respectively represent the available resource amounts of the newly released resources on the CPU, memory, and network bandwidth.
[0116] In a possible implementation manner, please refer to Figure 5 for reference. The specific steps for forming the meteorological spatio-temporal distribution coordinate system include:
[0117] Obtain the geographical information of the meteorological data blocks in the high-frequency data cache set. Using the longitude and latitude difference as the division granularity, divide the globe or a specified area into multiple sub-regions, and each of the sub-regions corresponds to a unique geographical code;
[0118] Divide the height intervals according to the collection height of each sub-region, classify the meteorological data blocks in each sub-region into the corresponding intervals according to the height intervals, and at the same time label the corresponding height identifiers for each height interval, where the height identifiers include the ground layer, the low-altitude layer, the middle-altitude layer, and the high-altitude layer;
[0119] At a preset time interval, organize the timestamp information of the meteorological data blocks in each sub-region, group the meteorological data blocks belonging to the same time period into one group, and label a time serial number for each group of data;
[0120] Construct a three-dimensional coordinate system for the meteorological data blocks after being divided by region, altitude, and time dimensions, thereby forming a meteorological spatio-temporal distribution coordinate system. Among them, the X-axis represents the region coding value, the Y-axis represents the altitude identifier, and the Z-axis represents the time serial number.
[0121] In a possible implementation manner, please refer to Figure 6 , the specific steps for the secondary adjustment of the resource configuration include:
[0122] Extract the computing resource occupancy information and storage resource occupancy information corresponding to each meteorological data block that is executing a meteorological task in the resource configuration model. At the same time, obtain the free resource quantity information in the resource configuration model, determine the coordinate values corresponding to each meteorological data block in the meteorological spatio-temporal distribution coordinate system, and summarize and integrate the resource occupancy of each meteorological data block, its corresponding coordinate value information, and the free resource quantity to form a complete meteorological data set;
[0123] Assign the integrated meteorological data set as input features to a trained convolutional neural network model;
[0124] The convolutional layer of the convolutional neural network model uses a 3×3 small convolutional kernel to extract the input features, calculates the extracted features through the ReLU activation function, the pooling layer uses a 2×2 max pooling operation to reduce the dimension of the input features after activation processing, and the fully connected layer integrates the input features after convolution and pooling processing, and predicts the storage change rate of the cache area within 12 hours, and predicts the change ratio of the computing resource requirements of each meteorological task within 3 hours;
[0125] Compare the predicted storage change rate of the cache area within 12 hours with the preset storage change rate. If it is less than the preset storage change rate, start the data cleaning mechanism, traverse the high-frequency data cache set, and remove the 20% of the meteorological data blocks that are closest to the expiration time from the cache area;
[0126] Compare the predicted change ratio of the computing resource requirements of each meteorological task within 3 hours with the preset change threshold. If the change ratio of the computing resource requirements of the current meteorological task is greater than the preset change threshold, at this time, start the resource allocation mechanism, and recycle part of the resources from other low-priority meteorological tasks and resource pools according to a ratio of 1:3 and allocate them to the current meteorological task.
[0127] In this embodiment, the training process of the convolutional neural network model is as follows:
[0128] Use a large amount of relevant data during the execution of historical meteorological tasks as training samples. This training sample covers resource allocation information and corresponding meteorological spatio-temporal distribution characteristics under different time periods, different regions, and different meteorological task types. Specifically, it includes the resource occupancy of each meteorological task at different times during execution, including CPU usage rate, memory occupancy, and network bandwidth consumption data. It also includes the detailed coordinate information of each meteorological data block in the meteorological spatio-temporal distribution coordinate system, as well as the storage change situation of the data in the high-frequency data cache set during the task execution process.
[0129] Preprocess the training samples. For the resource occupancy information, perform normalization processing to make its numerical range between [0,1], so that the model can converge faster and better. For the coordinate values of the meteorological spatio-temporal distribution coordinate system, perform standard encoding at the same time, and convert the height identifier into a digital vector form that is easy for the model to understand, ensuring the consistency and comparability of the data.
[0130] Randomly divide the preprocessed training samples into a training set and a test set according to the ratio of 80% and 20%. Use the Stochastic Gradient Descent (SGD) algorithm to adjust the weight parameters of the model to gradually reduce the value of the loss function. During the training process, set the initial learning rate to 0.001. As the number of training epochs increases, every 10 epochs, the learning rate decays to 0.9 times the original value to ensure that the model can converge more precisely in the later stage of training. At the same time, introduce the Early Stopping mechanism to monitor the loss change on the test set. When the loss of the test set no longer decreases for 5 consecutive epochs, stop training to prevent the model from overfitting.
[0131] Use the test set to evaluate the performance of the trained model. Calculate the accuracy, recall rate, and F1 value between the predicted value and the true value. If the accuracy of the model is lower than 85%, further adjust and optimize the hyperparameters of the training algorithm and retrain until the model shows good prediction accuracy and stability on the test set, and can accurately predict the storage change in the cache area and the change in meteorological task resource requirements, providing strong decision support for the entire weather-based HPC resource scheduling system.
[0132] In a second aspect, please refer to Figure 7 , this application example provides a system adopting any one of the above weather-based HPC resource scheduling methods, including:
[0133] A high-frequency data storage unit, configured to store meteorological data blocks with an access frequency higher than a set access frequency threshold into the cache area;
[0134] A task priority list generation unit, configured to generate a task priority list for each meteorological data block in the high-frequency data cache set according to the timeliness requirements, data dependency relationships, and resource requirement characteristics of its meteorological tasks;
[0135] A resource allocation unit, configured to perform a first resource allocation on the task priority list to generate a resource scheduling plan;
[0136] A load balancing monitoring unit, configured to dynamically adjust the resource allocation of overloaded nodes based on the resource scheduling plan to generate a dynamic load balancing mechanism;
[0137] A spatio-temporal coordinate system construction unit, configured to perform three-dimensional grid division on the meteorological data blocks in the high-frequency data cache set according to region, altitude, and time to form a meteorological spatio-temporal distribution coordinate system;
[0138] A resource secondary adjustment unit, configured to send the resource occupancy ratio, the number of idle resources, and the coordinate values in the meteorological spatio-temporal distribution coordinate system when the resource configuration model executes meteorological tasks to a trained convolutional neural network model for prediction, and perform secondary adjustment of resource configuration according to the prediction results.
[0139] The above specific embodiments further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A weather-based HPC resource scheduling method according to claim 1, characterized in that, Including the following steps: Based on meteorological data and its corresponding access frequency, using a clustering analysis algorithm, classify the meteorological data, identify meteorological data blocks with high access frequency, and store the meteorological data blocks in a cache area close to the computing nodes to form a high-frequency data cache set; Use the analytic hierarchy process to analyze and process the time-sensitive requirements, data dependency relationships, and resource demand characteristics of each meteorological task in the high-frequency data cache set to generate a task priority list; Use the first-fit algorithm to allocate resources for the task priority list to generate a resource scheduling plan; Based on the resource scheduling plan, deploy distributed probes to the computing nodes at all levels of the HPC system, monitor the CPU usage rate, memory occupancy, and network bandwidth consumption of the HPC system in real time, and dynamically adjust the task allocation to the computing nodes with lighter loads according to the preset load balancing threshold to generate a dynamic load balancing mechanism; Perform three-dimensional grid division on the meteorological data blocks in the high-frequency data cache set according to region, altitude, and time to form a meteorological spatio-temporal distribution coordinate system; Send the resource occupancy ratio, idle resource quantity of each meteorological task when executing the meteorological task, and the coordinate values corresponding to each meteorological block in the meteorological spatio-temporal distribution coordinate system in the task priority list to the trained convolutional neural network model for prediction, and perform secondary adjustment of resource configuration according to the prediction results to complete the resource scheduling of meteorological tasks.
2. The method for scheduling HPC resources based on weather according to claim 1, wherein The specific steps for forming the high-frequency data cache set based on meteorological data and its corresponding access frequency include: Read the storage file of the meteorological data, and through data cleaning, remove the invalid values and duplicate values in the meteorological data to generate a meteorological data set; Extract features from the meteorological data set to form meteorological data features, use the KMeans clustering algorithm, and classify the meteorological data according to the types of the meteorological data features to generate meteorological data clustering subsets; Count the access frequency of data points in each clustering subset, combine the timestamp information corresponding to the data points, count the access times of the data points, sort according to the access times, filter out the data points with access frequency higher than the set access frequency threshold, and label them as meteorological data blocks with high access frequency; Determine the cache area close to the computing nodes through the system configuration file, divide the cache partition for storing high-frequency meteorological data, and use the mounting instruction of the system configuration file to store the meteorological data blocks in the high-speed cache area to form a high-frequency data cache set. Calibrate the expiration period of the meteorological data blocks stored in the high-speed cache area according to the access frequency of data points in each clustering subset, and remove the meteorological data blocks from the high-frequency data cache set after expiration.
3. The weather-based HPC resource scheduling method according to claim 1, wherein The specific steps for generating the task priority list using the analytic hierarchy process include: Parse each meteorological task in the high-frequency data cache set, extract the key information of the meteorological task to construct a task information list, where the key information includes task name, geographical area range, time-sensitive requirement, and task execution duration; Divide each of the meteorological tasks into four levels: urgent, high, medium, and low according to the aging requirement, and set a preset reference time range for each level in combination with the task execution duration; Traverse the high-frequency data cache set, determine the data dependencies of each meteorological task, and build a task dependency graph according to the dependencies; According to the data dependencies in the task dependency graph, the aging requirement, and the task execution duration, use a standardized scoring mechanism to evaluate the resource requirements of each meteorological task, and configure a corresponding resource allocation weight for each meteorological task as its resource requirement characteristic; Use the aging requirement, data dependencies, and resource requirement characteristics as the analysis indicators of the analytic hierarchy process, and use the analytic hierarchy process to sort the priorities of each meteorological task, so as to generate a task priority list.
4. A weather-based HPC resource scheduling method according to claim 3, characterized in that, The specific steps for generating the resource scheduling plan using the first-fit algorithm include: Initialize the current resource allocation status of the HPC system, obtain the current available resource information, and send the available resource information to the resource pool, where the available resource information includes the number of idle CPU cores, available memory space, and unoccupied network bandwidth; Traverse the task priority list in descending order of task priority. According to the resource requirement characteristics of each meteorological task, use the first-fit algorithm to first allocate the resources in the resource pool to the meteorological tasks with high priority, and then allocate them to other meteorological tasks in turn.
5. The weather-based HPC resource scheduling method according to claim 4, wherein The specific steps for generating the dynamic load balancing mechanism based on the resource scheduling plan include: Based on the resource scheduling plan, start the resource allocation process, and use distributed probes to collect the CPU usage rate, memory occupancy, and network bandwidth consumption data of each level of computing nodes in the HPC system in real time, and send the data back to the monitoring center in real time; The monitoring center analyzes and judges the backhaul data according to the preset load balancing threshold. When it is detected that the computing node corresponding to the current meteorological task exceeds the load threshold, mark the node as an overloaded node, and at the same time continuously start the resource preemption mechanism and the listening mechanism; Execute the resource preemption mechanism, quickly locate the meteorological tasks with low priority according to the task priority list, calculate the difference between the resources required by the current meteorological task and the remaining resources of the overloaded node, and accurately recover some resources from the resources already allocated to the low-priority meteorological tasks according to the preset resource recovery ratio, and give priority to supplementing the current meteorological task; Synchronously run the listening mechanism, continuously monitor the execution status of the previous and / or previous meteorological tasks. When it is detected that a task releases resources, allocate all or part of the released resources to the current meteorological task according to the dynamic resource adaptation formula until the load returns to the normal range, and finally generate a dynamic load balancing mechanism.
6. The weather-based HPC resource scheduling method according to claim 5, wherein The expression of the dynamic resource adaptation formula is: Wherein, R allocated - represents the amount of resources allocated to the current meteorological task, R r - represents the amount of resources that have been recycled, R g - represents the resource gap of the current meteorological task, R n - represents the amount of newly released resources, ω cpu 、ω mem and ω bw respectively represent the resource configuration weights of the current task on meteorological data CPU, memory, and network bandwidth, and respectively represent the available resource amounts of the newly released resources on CPU, memory, and network bandwidth.
7. A weather-based HPC resource scheduling method according to claim 5, characterized in that The specific steps for forming the meteorological spatio-temporal distribution coordinate system include: Obtain the geographical information of the meteorological data blocks in the high-frequency data cache set, use the longitude and latitude difference as the division granularity, divide the globe or a specified area into multiple sub-regions, and each sub-region corresponds to a unique geographical code; Divide height intervals according to the acquisition height of each of the sub-regions, classify the meteorological data blocks in each of the sub-regions into corresponding intervals according to the height intervals, and at the same time label corresponding height identifiers for each of the height intervals, where the height identifiers include ground layer, low altitude layer, middle altitude layer, and high altitude layer; At a preset time interval, organize the timestamp information of the meteorological data blocks in each of the sub-regions, group the meteorological data blocks belonging to the same time period into a group, and assign a time serial number to each group of data; Construct a three-dimensional coordinate system for the meteorological data blocks divided by region, height, and time dimensions, so as to form a meteorological spatio-temporal distribution coordinate system, where the X-axis represents the region coding value, the Y-axis represents the height identifier, and the Z-axis represents the time serial number.
8. A weather-based HPC resource scheduling method according to claim 7, characterized in that The specific steps for performing the secondary adjustment of the resource allocation include: Extract the computing resource occupancy information and storage resource occupancy information corresponding to each meteorological data block performing a meteorological task in the resource allocation model, and at the same time obtain the idle resource quantity information in the resource allocation model, determine the coordinate values corresponding to each meteorological data block in the meteorological spatio-temporal distribution coordinate system, and summarize and integrate the resource occupancy of each meteorological data block, its corresponding coordinate value information, and the idle resource quantity to form a complete meteorological data set; Use the integrated meteorological data set as input features and assign them to a trained convolutional neural network model; The convolutional layer of the convolutional neural network model uses a 3×3 small convolutional kernel to extract the input features, calculates the extracted features through a ReLU activation function, the pooling layer uses a 2×2 max pooling operation to reduce the dimension of the input features after activation processing, and the fully connected layer integrates the input features after convolution and pooling processing, and predicts the storage change rate of the cache area within 12 hours, and predicts the calculation resource demand change ratio of each meteorological task within 3 hours; Compare the predicted storage change rate of the cache area within 12 hours with a preset storage change rate. If it is less than the preset storage change rate, start the data cleaning mechanism, traverse the high-frequency data cache set, and remove the 20% of the meteorological data blocks closest to the expiration time from the cache area; Compare the predicted calculation resource demand change ratio of each meteorological task within 3 hours with a preset change threshold. If the calculation resource demand change ratio of the current meteorological task is greater than the preset change threshold, at this time start the resource allocation mechanism, and recycle part of the resources from other low-priority meteorological tasks and resource pools in a ratio of 1:3 and allocate them to the current meteorological task.
9. The HPC resource scheduling system based on weather according to claims 1 to 8, characterized in that, Including: A high-frequency data storage unit configured to store meteorological data blocks with an access frequency higher than a set access frequency threshold in the cache area; A task priority list generation unit configured to generate a task priority list for each meteorological data block in the high-frequency data cache set according to the timeliness requirements, data dependency relationships, and resource demand characteristics of its meteorological task; A resource allocation unit configured to perform a first resource allocation on the task priority list to generate a resource scheduling plan; The load balancing monitoring unit is configured to dynamically adjust the resource allocation of overloaded nodes based on the resource scheduling scheme, and generate a dynamic load balancing mechanism; The spatio-temporal coordinate system construction unit is configured to perform three-dimensional grid division on the meteorological data blocks in the high-frequency data cache according to regions, heights, and time to form a meteorological spatio-temporal distribution coordinate system; The resource secondary adjustment unit is configured to send the resource occupancy ratio, the number of idle resources, and the coordinate values in the meteorological spatio-temporal distribution coordinate system of the resource configuration model when executing meteorological tasks to the trained convolutional neural network model for prediction, and perform secondary adjustment of resource configuration according to the prediction results.
Citation Information
Cited By
Carbon satellite ground system resource dynamic allocation method based on service awareness
CN121508613A
Business-aware-based carbon satellite ground system resource dynamic allocation method
CN121508613B