Data processing method and device for high-performance computing power environment, medium and product

Through real-time performance monitoring and dynamic resource scheduling, combined with task feature analysis, and a gradual adjustment strategy is adopted, the problem of frequent resource evaluation and severe adjustment in a high-performance computing environment is solved, and the efficient and stable operation of the system and the precise optimization of computing resources are achieved.

CN120336038APending Publication Date: 2025-07-18BEIJING YIYONG TIMES TECH CO LTD

Patent Information

Application Number
CN202510836239.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the existing high-performance computing power environment, resource evaluation is frequent and vigorously adjusted, resulting in low system processing efficiency and inability to effectively adapt to dynamic and changing computing needs.

Method used

Through real-time performance monitoring and dynamic resource scheduling, resource configuration is optimized in real time based on system performance indicators, and a gradual adjustment strategy is adopted to avoid mutations in resource configuration, and tasks are merged or split in combination with task feature analysis to achieve smooth transition of resource configuration and system stability.

Benefits of technology

It improves the overall processing efficiency of the system, ensures efficient utilization and stable operation of resources, adapts to the dynamic needs of different computing scenarios, and avoids the overhead and instability caused by resource evaluation and adjustment in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336038A_ABST
    Figure CN120336038A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device for a high-performance computing power environment, a medium and a product, and relates to the field of electrical digital data processing.The method comprises the steps that input computing data is received, data fragmentation processing is conducted on the computing data, and a plurality of parallel computing tasks are generated; starting an executive program of the parallel computing task based on default resource configuration, and collecting system performance indexes in the running process of the executive program; calculating an execution efficiency value of the execution program based on the system performance index; when the execution efficiency value is lower than a preset efficiency threshold value, generating optimized resource configuration of the parallel computing task based on the execution efficiency value; and dynamically adjusting the resource configuration parameters based on the real-time resource scheduling interface, so that the default resource configuration of the executive program is smoothly transited to the optimized resource configuration in the running process. By implementing the application, the data processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic digital data processing, and in particular to a data processing method, apparatus, medium and product for a high-performance computing power environment. Background Art

[0002] With the continuous growth of the scale and complexity of data processing, the demand for data processing in high-performance computing power environments is increasing day by day. In fields such as financial transactions, scientific computing, and real-time analysis, higher requirements are put forward for the timeliness and resource utilization efficiency of data processing. Especially in a mixed workload scenario, it is necessary to process small-scale data with high real-time requirements while also taking into account resource-intensive large-scale computing tasks.

[0003] In related technologies, high-performance computing power environments usually adopt a priority scheduling scheme to process data. This scheme first divides the input data into priorities and distributes the data to different processing queues. The processing system processes the data in each queue in turn according to the priority order, and adopts an independent resource allocation strategy for each queue. During the processing, the system splits large-scale data into multiple subtasks through task splitting and distributes them to different computing nodes for parallel processing.

[0004] However, due to the need for resource evaluation and task planning for each processing, the overall processing efficiency of the system is low. Summary of the Invention

[0005] This application provides a data processing method, apparatus, medium and product for a high-performance computing power environment to improve data processing efficiency.

[0006] In a first aspect, this application provides a data processing method for a high-performance computing power environment, which is applied to a data processing device. The method includes: receiving input calculation data, performing data sharding processing on the calculation data to generate multiple parallel computing tasks; starting an execution program of the parallel computing tasks based on a default resource configuration, and collecting system performance metrics during the running of the execution program; calculating an execution efficiency value of the execution program based on the system performance metrics; when the execution efficiency value is lower than a preset efficiency threshold, generating an optimized resource configuration for the parallel computing tasks based on the execution efficiency value; and dynamically adjusting resource configuration parameters based on a real-time resource scheduling interface, so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation.

[0007] In the above embodiments, the data processing device generates parallel tasks by fragmenting the calculation data, and monitors the system performance metrics in real time during the execution process. Only when the execution efficiency is lower than the threshold, the resource configuration is dynamically optimized, realizing a smooth transition of the resource configuration. This dynamic resource scheduling mechanism based on real-time performance feedback avoids the overhead caused by repeated resource evaluations in the traditional solutions, improves the overall processing efficiency of the system, and ensures the stability of the system operation.

[0008] Combined with some embodiments of the first aspect, in some embodiments, before the step of starting the execution program of the parallel computing task based on the default resource configuration and collecting the system performance metrics during the running process of the execution program, the method further includes: performing clustering analysis on the resource usage data in the historical execution records to obtain multiple resource usage patterns; calculating the benchmark usage amounts of various resources based on the distribution characteristics of the resource usage patterns; setting the default resource configuration according to the benchmark usage amounts; the default resource configuration includes the initial allocation values of the number of CPU cores, the memory capacity, and the I / O bandwidth.

[0009] In the above embodiments, the data processing device performs clustering analysis on the historical execution records, identifies typical resource usage patterns, and sets the default resource configuration including CPU, memory, and I / O configurations accordingly, avoiding the problem of blindly allocating resources, making the initial resource allocation more reasonable, reducing the range and frequency of subsequent resource adjustments, and improving the startup efficiency of the system.

[0010] Combined with some embodiments of the first aspect, in some embodiments, the step of calculating the execution efficiency value of the execution program based on the system performance metrics specifically includes: collecting the CPU utilization rate, the memory usage rate, and the I / O waiting time according to a preset sampling period; performing normalization processing and weighted calculation on the CPU utilization rate, the memory usage rate, and the I / O waiting time to obtain the system resource utilization degree; obtaining the number of calculation tasks completed within a unit time, and calculating the task completion rate; calculating the execution efficiency value based on the system resource utilization degree and the task completion rate.

[0011] In the above embodiments, the data processing device uses multi-dimensional performance metrics to evaluate the system operation state, including the CPU utilization rate, the memory usage rate, and the I / O waiting time, etc. Obtaining the system resource utilization degree through normalization processing and weighted calculation, and then calculating the execution efficiency in combination with the task completion rate, can comprehensively reflect the system performance state and provide an accurate basis for resource optimization.

[0012] In some embodiments in combination with some embodiments of the first aspect, the step of dynamically adjusting resource configuration parameters based on the real-time resource scheduling interface to smoothly transition the default resource configuration of the execution program to the optimized resource configuration during operation specifically includes: obtaining the CPU quota difference, memory limit difference, and I / O bandwidth difference between the default resource configuration and the optimized resource configuration; calculating the CPU adjustment amount, memory adjustment amount, and I / O bandwidth adjustment amount for each adjustment cycle according to the CPU quota difference, memory limit difference, and I / O bandwidth difference; adjusting the CPU quota of the current adjustment cycle through the real-time resource scheduling interface and collecting the real-time CPU usage rate; when the fluctuation range of the real-time CPU usage rate is less than the first preset threshold, adjusting the memory limit of the current adjustment cycle through the real-time resource scheduling interface and collecting the real-time memory usage rate; when the fluctuation range of the real-time memory usage rate is less than the second preset threshold, adjusting the I / O bandwidth of the current adjustment cycle through the real-time resource scheduling interface and collecting the real-time I / O waiting time.

[0013] In the above embodiments, the data processing device adopts a step-by-step and progressive resource adjustment strategy, first adjusting the CPU quota, then adjusting the memory limit after the system stabilizes, and finally adjusting the I / O bandwidth, avoiding the impact on the system caused by drastic changes in resources and ensuring the stable operation of the system during the resource adjustment process.

[0014] In some embodiments in combination with some embodiments of the first aspect, after the step of adjusting the I / O bandwidth of the current adjustment cycle through the real-time resource scheduling interface and collecting the real-time I / O waiting time when the fluctuation range of the real-time memory usage rate is less than the second preset threshold, the method further includes: calculating the weighted sum of the real-time CPU usage rate, real-time memory usage rate, and real-time I / O waiting time to obtain the system stability score of the current adjustment cycle; when the system stability score is greater than the preset stability threshold, using the CPU adjustment amount, memory adjustment amount, and I / O bandwidth adjustment amount as the adjustment parameters for the next adjustment cycle.

[0015] In the above embodiments, the data processing device evaluates the effect of the current adjustment by calculating the system stability score, and only continues the next round of adjustment when the system stability meets the requirements, being able to promptly detect and correct improper resource adjustments and avoid the system falling into an unstable state.

[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of dynamically adjusting resource configuration parameters based on a real-time resource scheduling interface so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation, the method further includes: when the execution efficiency value under the optimized resource configuration is still lower than a preset efficiency threshold, obtaining the load characteristics of the parallel computing tasks currently being executed; based on the load characteristics, merging or splitting the parallel computing tasks to generate a new set of parallel computing tasks; and rescheduling the new set of parallel computing tasks based on the optimized resource configuration.

[0017] In the above embodiments, when the data processing device still cannot achieve the expected efficiency after resource optimization, it will reorganize the tasks based on the load characteristics, optimize the task granularity by merging or splitting, and improve the processing efficiency of the system.

[0018] In combination with some embodiments of the first aspect, in some embodiments, before the step of merging or splitting the parallel computing tasks based on the load characteristics to generate a new set of parallel computing tasks, the method further includes: obtaining the computational complexity and data dependency relationships of each parallel computing task, and establishing a task association graph; calculating the data transfer overhead between tasks according to the task association graph, and constructing a cost function for task reorganization based on the data transfer overhead; when it is determined based on the cost function that the data transfer overhead between adjacent tasks is greater than a preset transmission threshold, marking the adjacent tasks as tasks to be merged; and when it is determined based on the cost function that the computational complexity of a single task is greater than a preset computational threshold, marking the single task as a task to be split.

[0019] In the above embodiments, the data processing device analyzes the data dependency relationships and data transfer overhead between tasks by establishing a task association graph, constructs a cost function for task reorganization, and determines the task merging or splitting strategy accordingly, which can minimize the data transfer overhead between tasks and improve the overall processing efficiency of the system.

[0020] In a second aspect, an embodiment of the present application provides a data processing device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the data processing device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0021] In a third aspect, an embodiment of the present application provides a computer program product containing instructions, and when the computer program product runs on a data processing device, it causes the data processing device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0022] Fourthly, an embodiment of the present application provides a computer-readable storage medium, including instructions, which, when running on a data processing device, cause the data processing device to execute the method described in the first aspect and any possible implementation manner in the first aspect.

[0023] It can be understood that the data processing device provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiments of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, and will not be elaborated here.

[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Since a data processing method based on real-time performance monitoring and dynamic resource scheduling is adopted, the system can optimize resource allocation in real time according to execution efficiency and achieve smooth transition, effectively solving the problems of frequent resource evaluation and drastic adjustment in the prior art, and then realizing the efficient utilization and stable operation of system resources; avoiding the frequent resource evaluation and planning processes in the traditional scheme, reducing system overhead, and ensuring the stability of the system through progressive adjustment, and finally improving the overall efficiency of data processing.

[0025] 2. Since a resource configuration initialization method based on historical data analysis is adopted, the system can identify typical resource usage patterns and set reasonable default configurations accordingly, effectively solving the problems of blind and unreasonable initial resource allocation in the prior art, and then realizing the efficient utilization of resources in the system startup stage; making the initial resource allocation more accurate, reducing the need for subsequent adjustment, avoiding resource waste, and improving the startup efficiency of the system.

[0026] 3. Since a task reorganization and optimization method based on load characteristics is adopted, the system can reasonably reorganize tasks when the expected efficiency cannot be achieved after resource optimization, effectively solving the problem that only relying on resource adjustment cannot further improve performance in the prior art, and then realizing the in-depth optimization of system performance; improving the processing efficiency of the system and enabling the system to better adapt to different computing scenarios. Description of the Drawings

[0027] Figure 1 is a flowchart of a data processing method in a high-performance computing power environment in an embodiment of the present application; Figure 2 is another flowchart of a data processing method in a high-performance computing power environment in an embodiment of the present application; Figure 3 is a schematic structural diagram of an entity device of a data processing device in an embodiment of the present application. Detailed Embodiments

[0028] The terms used in the following embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application, the singular forms "a", "an", "above-mentioned", "the", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations of one or more of the listed items.

[0029] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0030] For ease of understanding, the application scenarios of the embodiments of this application are introduced below.

[0031] During the Double Eleven period, a large e-commerce platform needs to process a huge amount of order transaction data. The platform uses a distributed computing framework to process data and decomposes the order processing tasks into multiple parallel computing tasks. However, due to the volatility of the order volume, which surges from zero o'clock in the morning and gradually stabilizes in the morning, and there is a small peak at noon, this results in a drastic change in the computing resource requirements accordingly. If a fixed resource allocation scheme is adopted, either there will be insufficient resources during the peak period, leading to delays in order processing, or there will be a waste of resources during the low period. At the same time, different types of order processing tasks (such as ordinary orders, flash sale orders, cross-border orders, etc.) have different requirements for computing resources. Such a dynamically changing business scenario poses a severe challenge to the intelligent scheduling and optimization of computing resources.

[0032] In the related art, the dynamic allocation of computing resources can be achieved by adopting a resource adjustment strategy with a fixed threshold. For example, when the system detects that the CPU usage rate exceeds the preset threshold, the number of CPU cores is increased, or when the memory usage rate exceeds a certain fixed value, the memory capacity is expanded; this method cannot accurately identify the dependency relationship between tasks and the data flow characteristics, and the sudden resource adjustment is likely to cause drastic fluctuations in system performance. The scenario of the data processing method using the high-performance computing power environment in the related art is introduced below.

[0033] A financial institution uses existing resource scheduling technologies to process high-frequency trading data. This technology adopts a preset fixed resource allocation threshold. When the CPU usage rate is detected to exceed 90%, the number of CPU cores is increased. When the memory usage rate exceeds 85%, the memory capacity is expanded. However, in actual operation, due to the failure to consider the dependencies between tasks and the characteristics of data flow, the effect of resource adjustment is not good. For example, although the CPU usage rate of some computing tasks is not high, there is frequent data interaction, and the fixed-threshold adjustment scheme cannot recognize this situation. At the same time, the resource adjustment adopts a mutation-style adjustment, directly jumping the resource configuration from one state to another, resulting in severe fluctuations in system performance. Especially at critical moments such as the opening of the market and the release of major news, this unstable adjustment method may lead to response delays in the trading system and affect the trading execution efficiency.

[0034] However, by using the data processing method for a high-performance computing power environment in the embodiments of the present application, through in-depth analysis of task characteristics and a multi-step smooth adjustment strategy, intelligent optimization of computing resources is achieved. It can not only accurately identify the data dependencies between tasks but also ensure the stability of system performance. The following introduces the scenarios using the data processing method for a high-performance computing power environment in the present application.

[0035] A cloud computing service provider adopts the solution of the present application to optimize computing resource scheduling. When a user submits a machine learning training task, the system first analyzes the computing characteristics and data dependencies of the task and constructs a task association graph. For model parameter update tasks with frequent data interaction, the system identifies the high data flow overhead among them and automatically combines these tasks for processing. For feature engineering tasks with high computing complexity, the system splits them into multiple parallel subtasks. During the resource adjustment process, the system adopts a progressive adjustment strategy to achieve a smooth transition of resource configuration through multiple small steps. For example, when it is necessary to increase the GPU computing power, the system will gradually improve the computing ability within multiple adjustment cycles while monitoring the stability of training performance; this method not only ensures the training efficiency of machine learning tasks but also avoids performance fluctuations caused by resource adjustment.

[0036] It can be seen that by using the data processing method for a high-performance computing power environment in the embodiments of the present application, while achieving dynamic adjustment of computing resources, it can also effectively solve the mutation adjustment problem in traditional fixed-threshold schemes, thereby achieving precise optimization and smooth transition of resource configuration.

[0037] The following describes the process of the method provided in this embodiment. Please refer to Figure 1 , which is a schematic flowchart of the data processing method for a high-performance computing power environment in the embodiments of the present application.

[0038] S101. Receive the input calculation data, perform data sharding on the calculation data, and generate multiple parallel calculation tasks.

[0039] Among them, the calculation data represents the set of original data that needs to be processed, which can include various types such as structured data and unstructured data; data sharding is the process of dividing large-scale data into multiple smaller data subsets according to a predetermined rule; parallel calculation tasks represent independent calculation units that can be executed simultaneously, and each calculation unit is responsible for processing a data subset; the data processing device refers to a hardware device with data receiving, processing, and calculation capabilities, which can be a computing device such as a server or a workstation.

[0040] After receiving the calculation request submitted by the user, the data processing device needs to process the input large-scale data to support parallel calculation. Specifically, the data processing device first receives the input calculation data and analyzes it according to the format, size, and calculation characteristics of the data. Then, based on factors such as data correlation and calculation complexity, it selects an appropriate sharding strategy to divide the data into multiple data subsets with appropriate sizes and balanced loads. Finally, the data processing device creates corresponding calculation tasks for each data subset, configures attributes such as input and output parameters and execution priorities of the tasks, and generates a set of parallelizable calculation tasks.

[0041] In some embodiments, data sharding and parallel task generation can be achieved in various ways: Optionally, the data processing device can adopt a sharding method based on data blocks. First, the data is divided into multiple data blocks according to a fixed size, and then the dependency relationships between the data blocks are analyzed. The independent data blocks are assigned to different calculation tasks, and the data blocks with dependency relationships are assigned to the same calculation task; Optionally, the data processing device can also adopt a sharding method based on calculation characteristics. According to the calculation complexity, memory occupancy and other characteristics of the data, the data with similar calculation characteristics is divided into the same subset to achieve balanced distribution of the calculation load. It can be understood that other data sharding methods can also be adopted, such as sharding based on data type, sharding based on time window, etc. The specific sharding method to be adopted can be selected according to the actual application scenario. In addition, when performing data sharding, the integrity and consistency of the data also need to be considered to ensure that the sharded data can be correctly processed.

[0042] In practical applications, data sharding may encounter the problem of data skew, that is, the computational workload of some data subsets is greater than that of other subsets, resulting in a decrease in the efficiency of parallel computing. To address this, the data processing device can adopt a dynamic load balancing strategy: First, sample and analyze the data to identify the characteristics that may cause data skew; then preprocess the data based on these characteristics, such as scattering or further subdividing hot data; finally, when generating parallel tasks, ensure that the computational load of each task is relatively balanced by merging or splitting tasks. For example, in the scenario of user behavior log analysis, the data volume of some popular users or popular products may be much larger than other data. At this time, the hot data can be randomly assigned to multiple computing tasks to avoid a single task from processing too much data.

[0043] S102. Start the execution program of the parallel computing task based on the default resource configuration, and collect system performance metrics during the running of the execution program.

[0044] Among them, the default resource configuration refers to the set of computational resource parameters initially allocated by the data processing device to the execution program, including the number of CPU cores, memory capacity, I / O bandwidth, etc.; the execution program refers to the software entity responsible for scheduling and executing parallel computing tasks; the system performance metrics refer to the various quantitative metrics reflecting the running state of the data processing device, including CPU usage, memory occupancy, I / O waiting time, network throughput, etc.; collection refers to the process of regularly obtaining and recording these metric data through the system monitoring module.

[0045] After the data processing device completes the generation of parallel computing tasks, it needs to allocate computational resources for these tasks and start the execution. Specifically, the data processing device first allocates corresponding computational resources for each parallel computing task according to the pre-set default resource configuration, including CPU time slices, memory space, I / O bandwidth, etc. Then it starts the execution program, and the execution program is responsible for scheduling these parallel tasks to run within the allocated resources. During the task execution process, the data processing device continuously monitors the system running state, collects various performance metrics according to the pre-set sampling period (such as every second or every minute), and stores these metric data in the performance log to provide a basis for subsequent performance analysis and optimization.

[0046] In some embodiments, the acquisition and processing of performance metrics can be achieved in various ways: Optionally, the data processing device can adopt a system call-based approach to obtain underlying performance data such as CPU usage, memory allocation status, I / O operation statistics, etc. by calling the performance counter interface provided by the operating system, and then aggregate and standardize this raw data to generate performance metrics for analysis; Optionally, the data processing device can also adopt an application-layer monitoring approach, embedding performance probes in the executing program to record application-level metrics such as task execution time, resource occupancy, queue length, etc., and performing comprehensive analysis in combination with system-level metrics. It can be understood that other performance monitoring methods can also be adopted, such as distributed tracing, log analysis, etc., and the specific monitoring method can be determined according to actual requirements. In addition, when performing performance monitoring, the system overhead brought by the monitoring itself also needs to be considered to ensure that the monitoring behavior does not affect the system performance.

[0047] In practical applications, the acquisition of performance metrics will encounter data noise problems, that is, due to factors such as system load fluctuations and network jitters, the collected performance data shows abnormal fluctuations, affecting the accuracy of subsequent analysis. In this regard, the data processing device can adopt a multi-level filtering strategy: First, perform outlier detection on the raw performance data to eliminate data points that deviate significantly from the normal range; then use methods such as moving window averaging to smooth the data and reduce the impact of short-term fluctuations; finally, through statistical analysis methods, such as median smoothing, exponentially weighted moving average, etc., extract the metric values that reflect the true performance state of the system. For example, when collecting the CPU usage rate, there may be outliers where the instantaneous usage rate exceeds 100%. At this time, historical data and context information need to be combined for correction to ensure the accuracy and reliability of the performance metrics.

[0048] S103. Calculate the execution efficiency value of the execution program based on the system performance metrics.

[0049] Among them, the execution efficiency value represents a comprehensive evaluation index for measuring the running performance of the execution program, reflecting the overall effect of system resource utilization and task processing; the system resource utilization degree refers to the ratio of the actual usage level of each system resource to its capacity upper limit; the task completion rate represents the ratio of the number of successfully processed computing tasks to the number of planned processing tasks per unit time; the standardization process refers to the mathematical processing process of converting performance metrics with different dimensions into a comparable unified scale.

[0050] The data processing device needs to comprehensively analyze the collected multi-dimensional performance metrics to evaluate the running efficiency of the execution program. Specifically, the data processing device first preprocesses the collected original performance metrics, including data cleaning, outlier handling, and time series alignment. Then, through normalization processing, metrics with different dimensions such as CPU utilization rate, memory usage rate, and I / O waiting time are converted to a unified evaluation scale. Next, based on the preset weight coefficients, the normalized metrics are weighted and calculated to obtain a comprehensive metric reflecting the overall resource utilization status of the system. At the same time, the number of tasks completed within the current time window is counted, and the task completion rate is calculated. Finally, the system resource utilization degree and the task completion rate are combined according to a specific calculation formula to obtain the final execution efficiency value.

[0051] In practical applications, problems with unreasonable metric weights may occur in the execution efficiency calculation, that is, the preset metric weights cannot accurately reflect the actual impact degree of different performance metrics on the overall efficiency, resulting in distorted efficiency evaluation results. In response to this, the data processing device can adopt an adaptive weight adjustment strategy: first, through correlation analysis, study the relationship between each performance metric and the actual processing effect; then, based on historical operation data, use machine learning algorithms (such as random forest, gradient boosting tree, etc.) to establish a metric importance evaluation model; finally, according to the evaluation results of the model, dynamically adjust the weight coefficients of each metric. For example, in I / O-intensive tasks, the impact of I / O waiting time on the execution efficiency may be greater than that of CPU utilization rate. At this time, the weights of I / O-related metrics should be appropriately increased; while in compute-intensive tasks, more attention should be paid to the impact of CPU-related metrics. In this way, the efficiency calculation can better adapt to different computing scenarios and load characteristics.

[0052] S104. When the execution efficiency value is lower than the preset efficiency threshold, generate an optimized resource configuration for the parallel computing task based on the execution efficiency value.

[0053] Among them, the preset efficiency threshold represents the lowest execution efficiency level that the system expects to maintain and is the benchmark value for judging whether the current execution efficiency needs to be optimized; the optimized resource configuration represents the improved resource allocation plan obtained through calculation and analysis, including the reallocation parameters of resources such as CPU, memory, and I / O; the execution efficiency deviation refers to the gap between the current execution efficiency value and the preset efficiency threshold; the resource elasticity range represents the maximum and minimum limits that various resources can be adjusted.

[0054] When the data processing device detects a decrease in execution efficiency, it needs to promptly adjust the resource allocation to improve system performance. Specifically, the data processing device first compares the calculated execution efficiency value with a preset efficiency threshold. When the efficiency value is lower than the threshold, it triggers an optimization process. Then, it analyzes the specific reasons for the efficiency decrease, including resource bottleneck identification, load characteristic analysis, etc. Next, based on the degree of efficiency deviation, combined with the current usage status and historical change trends of various resources, it calculates the amount of resources that need to be adjusted. Finally, within the resource elasticity range, it generates a new resource allocation plan, which needs to balance resource utilization efficiency and system stability to avoid performance fluctuations caused by excessive adjustment.

[0055] In some embodiments, the generation of optimized resource allocation can be achieved in various ways: Optionally, the data processing device can adopt a rule-based optimization method, pre-defining a series of resource adjustment rules, such as increasing the CPU quota when the CPU utilization rate continuously exceeds 90%, and expanding the memory limit when the memory usage rate approaches the upper limit, and triggering the corresponding adjustment rules according to the current state; Optionally, the data processing device can also adopt a performance model-based optimization method, establishing a mathematical model between execution efficiency and resource allocation, and obtaining the optimal resource allocation plan by solving the optimization problem. This model will consider multiple factors such as resource constraints and performance goals. It can be understood that other optimization methods can also be adopted, such as reinforcement learning, genetic algorithms, etc. The specific optimization method to be adopted can be determined according to the system scale and performance requirements. In addition, when generating the optimized configuration, the cost and risk of resource adjustment also need to be considered to ensure the feasibility of the optimization plan.

[0056] In practical applications, resource optimization will encounter resource competition problems, that is, multiple parallel tasks simultaneously request an increase in resource quotas, while the available resources of the system are limited and cannot meet the needs of all tasks. In response, the data processing device can adopt a multi-level priority strategy: First, set priority weights for different types of tasks to reflect the importance and urgency of the tasks; then establish a utility function for resource allocation to evaluate the improvement effect of increasing resources on the execution efficiency of different tasks; finally, based on task priorities and resource utilities, use methods such as priority queues or proportional allocation to reasonably allocate the limited resources. For example, when multiple computing tasks all need to increase the memory quota, the needs of critical business tasks can be preferentially met, and other tasks can be allocated in a progressive or proportional manner. In this way, the optimal overall performance improvement can be achieved under resource constraints.

[0057] S105. Dynamically adjust the resource configuration parameters based on the real-time resource scheduling interface, so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation.

[0058] Among them, the real-time resource scheduling interface represents a set of program interfaces provided by the data processing device for dynamically adjusting system resource allocation, supporting online modification of parameters such as CPU quota, memory limit, and I / O bandwidth; the resource configuration parameters represent specific values describing the resource allocation status, including the allocation amount and limit value of various resources; smooth transition means avoiding system instability caused by sudden changes in resource configuration through progressive adjustment; the adjustment period represents the time interval between two resource adjustment operations.

[0059] After obtaining the optimized resource configuration, the data processing device needs to achieve dynamic changes in resource configuration through a reasonable adjustment strategy. Specifically, the data processing device first calculates the difference between the default configuration and the optimized configuration, and determines the adjustment strategy based on the size of the difference and the system state. Then, the resource adjustment amount is divided into multiple smaller adjustment steps, and only a small part of the adjustment is performed within each adjustment period. After each adjustment, the adjustment effect is evaluated by monitoring system metrics, and the adjustment amplitude of the next step is dynamically adjusted according to the system response. The entire adjustment process needs to ensure the continuous and stable operation of the system, avoiding a decline in service quality caused by too fast or too drastic resource adjustment.

[0060] In some embodiments, smooth adjustment of resource configuration can be achieved in various ways: Optionally, the data processing device can adopt an adjustment method based on feedback control, monitoring the response characteristics of the system in each adjustment period, such as the degree of performance fluctuation and the change in resource utilization rate, and dynamically adjusting the step size and rate of resource allocation according to this feedback information to achieve a stable transition of the system state; Optionally, the data processing device can also adopt an adjustment method based on prediction, by establishing a system performance model, predicting the system response that may be caused by different adjustment strategies, and selecting the adjustment path that can minimize performance fluctuations. It can be understood that other adjustment methods can also be adopted, such as hierarchical adjustment, priority adjustment, etc. The specific adjustment method to be adopted can be determined according to the stability requirements and performance goals of the system. In addition, when performing resource adjustment, the dependency relationship between multiple resources also needs to be considered to ensure the coordinated change of various resources.

[0061] In practical applications, resource adjustment may encounter the problem of performance oscillation, that is, due to improper adjustment strategies, the system performance fluctuates repeatedly between different states, and a stable optimization effect cannot be achieved. In response to this, the data processing device can adopt an adaptive adjustment strategy: first, set multiple adjustment sensitivity levels to reflect the response characteristics of the system to resource changes; then observe the stability indicators of the system after each adjustment, including the amplitude of performance fluctuation, the change rate of resource utilization, etc.; when it is detected that the performance fluctuation exceeds the expectation, automatically reduce the adjustment sensitivity, reduce the adjustment step size or extend the adjustment period; on the contrary, when the system is stable, the adjustment sensitivity can be appropriately increased to accelerate the optimization process. For example, when adjusting the CPU quota, if it is observed that the task execution time fluctuates greatly, the adjustment amount each time should be reduced, and the observation period should be increased; when the system gradually stabilizes, the normal adjustment rhythm can be gradually restored. Through this dynamic adaptation mechanism, the optimal adjustment of resource allocation can be achieved on the premise of ensuring system stability.

[0062] The following further describes the more specific process of the method provided in this embodiment. Please refer to Figure 2 , which is another process schematic diagram of the data processing method for the high-performance computing power environment in the embodiment of the present application.

[0063] S201. Receive the input calculation data, perform data sharding processing on the calculation data, and generate multiple parallel calculation tasks.

[0064] Referring to step S101, the data processing device will shard the input data and create parallel tasks.

[0065] S202. Perform clustering analysis on the resource usage data in the historical execution records to obtain multiple resource usage patterns.

[0066] Among them, the historical execution record represents the set of resource usage data of the calculation tasks previously run by the data processing device; the resource usage data includes runtime statistical information such as CPU usage time series, memory occupancy changes, I / O operation records, etc.; clustering analysis refers to the process of classifying similar resource usage behaviors through data mining methods; the resource usage pattern represents the task type with similar resource consumption characteristics.

[0067] The data processing device needs to analyze historical data to identify typical resource usage patterns. Specifically, the data processing device first extracts resource usage data from the historical record database, and performs preprocessing such as cleaning and standardization on the data. Then select a suitable feature representation method to convert the time series data into feature vectors that can be used for clustering. Then use a clustering algorithm to group the feature vectors to identify the task sets with similar resource usage characteristics. Finally, evaluate and screen the clustering results to determine the most representative resource usage patterns.

[0068] In some embodiments, the identification of resource usage patterns can be achieved in various ways: Optionally, the data processing device can adopt distance-based clustering methods, such as algorithms like K-means, DBSCAN, etc., to cluster tasks with similar resource usage characteristics into one category, and evaluate the clustering effect through metrics such as the silhouette coefficient; Optionally, the data processing device can also adopt density-based clustering methods, and automatically discover high-density regions as typical usage patterns by analyzing the distribution density of resource usage data. It can be understood that other clustering methods can also be used, such as hierarchical clustering, spectral clustering, etc., and the specific clustering method to be adopted can be determined according to data characteristics and analysis requirements. In addition, when performing clustering analysis, the temporal correlation of the data also needs to be considered, and a time window mechanism can be introduced to capture the dynamic characteristics of resource usage.

[0069] In practical applications, clustering analysis may encounter the problem of unstable patterns, that is, over time, the resource usage patterns in historical data may change, resulting in the clustering results not being able to accurately reflect the current resource usage characteristics. In response to this, the data processing device can adopt an incremental learning strategy: First, set the validity period of historical data, eliminate expired data or assign it a lower weight; then, through an online clustering algorithm, continuously update the resource usage patterns; finally, establish a pattern evolution tracking mechanism to monitor the change trend of resource usage patterns and adjust the clustering model in a timely manner. For example, a sliding time window can be used to maintain the resource usage data in the most recent period, and clustering analysis is performed regularly to ensure the timeliness of the patterns.

[0070] S203. Calculate the baseline usage amounts of various resources based on the distribution characteristics of the resource usage patterns.

[0071] Among them, the distribution characteristics represent the statistical characteristics of resource usage amounts across different tasks and time periods, including mean, variance, quantiles, etc.; the baseline usage amount refers to the reference configuration value of various resources, which serves as the basic basis for resource allocation; statistical feature extraction refers to the process of calculating representative statistics from resource usage data.

[0072] The data processing device needs to determine reasonable resource baseline values based on the identified resource usage patterns. Specifically, the data processing device first conducts statistical analysis on the historical data under each resource usage pattern to calculate the probability distribution characteristics of the resource usage amounts. Then, according to the system's performance objectives and resource constraints, an appropriate statistic (such as the 75th percentile) is selected as the basis for calculating the baseline value. Next, the correlation relationship between different resources is considered to coordinate and balance the initially calculated baseline values. Finally, the adjusted values are used as the baseline usage amounts of various resources.

[0073] In some embodiments, the calculation of the baseline usage can be achieved in various ways: Optionally, the data processing device can adopt a method based on statistical analysis to determine a reasonable baseline value by calculating statistics such as the mean, standard deviation, and quantile of the resource usage, in combination with the service quality requirements of the system; Optionally, the data processing device can also adopt a method based on performance modeling to establish a relationship model between the resource usage and the system performance, and obtain the optimal baseline value by solving an optimization problem. It can be understood that other calculation methods can also be adopted, such as heuristic algorithms, expert systems, etc. The specific calculation method to be adopted can be determined according to the system characteristics and performance requirements. In addition, when determining the baseline usage, the dynamic characteristics of the system also need to be considered, and an elastic adjustment range can be set.

[0074] In practical applications, resource dependency problems will be encountered in the calculation of the baseline value, that is, there are usage correlations between different types of resources. Optimizing the baseline value of each resource separately may lead to poor overall performance. In this regard, the data processing device can adopt a multi-dimensional optimization strategy: First, establish an association model of resource usage to describe the dependency relationships between different resources; Then, transform the calculation of the baseline value into a multi-objective optimization problem, considering the usage efficiency of multiple resources at the same time; Finally, through methods such as Pareto optimality, find the optimal combination of the baseline values of various resources. For example, when determining the baseline usage of the CPU and memory, the impact of memory access on the CPU performance needs to be considered, and the configuration combination that can achieve the overall optimum is selected.

[0075] S204. Set the default resource configuration according to the baseline usage.

[0076] Among them, the default resource configuration represents the set of initial resource parameters allocated to tasks when the system starts; the configuration parameters include specific values such as the number of CPU cores, the memory capacity limit, and the I / O bandwidth upper limit; the configuration generation rule defines the method of converting the baseline usage into actual configuration parameters.

[0077] The data processing device needs to convert the calculated baseline usage into an executable resource configuration. Specifically, the data processing device first determines the configuration granularity and value range of various resources according to the hardware specifications and management strategies of the system. Then, map the baseline usage to the actual configuration parameter space, considering the minimum allocation unit and step value of the resources. Next, fine-tune and optimize the configuration parameters according to the scheduling strategy and performance goals of the system. Finally, generate a complete default resource configuration plan, including all necessary resource limit parameters.

[0078] In some embodiments, the default configuration can be set in various ways: Optionally, the data processing device can adopt a rule-based method, predefined a series of configuration conversion rules, and automatically generate configuration parameters according to the benchmark usage and system status; Optionally, the data processing device can also adopt a template-based method, maintain multiple preset configuration templates, and select the template that best matches the benchmark usage for parameter adjustment. It can be understood that other configuration methods can also be adopted, such as dynamic calculation, progressive adjustment, etc. The specific configuration method can be determined according to the management requirements and operation and maintenance strategies of the system. In addition, when setting the default configuration, appropriate resource margins need to be reserved to cope with load fluctuations.

[0079] In practical applications, resource fragmentation problems will occur in the default configuration setting, that is, due to the granularity limitations of the configuration parameters, the benchmark usage cannot be fully matched, resulting in a decrease in resource utilization efficiency. In this regard, the data processing device can adopt an intelligent configuration strategy: First, analyze the resource allocation limitations of the system, such as the binding rules of CPU cores, the memory page size, etc.; then design an alignment algorithm for the configuration parameters to convert the benchmark usage into a configuration value that meets the system limitations; finally, through resource pool management, rationally organize and allocate fragmented resources. For example, when the CPU benchmark usage requires 2.5 cores, a finer-grained allocation can be achieved through the CPU shares mechanism to avoid resource waste.

[0080] S205. Start the execution program of the parallel computing task based on the default resource configuration, and collect system performance metrics during the running of the execution program.

[0081] Referring to step S102, the data processing device will start the task execution and monitor the system performance.

[0082] S206. Calculate the execution efficiency value of the execution program based on the system performance metrics.

[0083] Referring to step S103, the data processing device will calculate the overall execution efficiency of the task.

[0084] In some embodiments, the data processing device will quantitatively evaluate the system performance, that is, the data processing device will collect the CPU utilization rate, memory usage rate, and I / O waiting time according to a preset sampling period; perform normalization processing and weighted calculation on the CPU utilization rate, memory usage rate, and I / O waiting time to obtain the system resource utilization; obtain the number of computing tasks completed per unit time, and calculate the task completion rate; based on the system resource utilization and the task completion rate, calculate the execution efficiency value.

[0085] Among them, the sampling period represents the time interval for the data processing device to collect performance data; the CPU utilization rate refers to the proportion of the actual working time of the CPU core in the total time; the memory usage rate represents the proportion of the allocated memory in the total available memory; the I / O waiting time refers to the time when the task is blocked due to waiting for the I / O operation to complete; the normalization process refers to the mathematical processing process of converting indicators with different dimensions into a unified scale; the weighted calculation means performing a comprehensive calculation by assigning weight coefficients according to the importance of different indicators; the system resource utilization degree is a comprehensive indicator reflecting the overall resource usage status of the data processing device; the task completion rate represents the ratio of the actual number of completed tasks to the planned number of tasks; the execution efficiency value is the final evaluation indicator for measuring the running performance of the data processing device.

[0086] The data processing device needs to establish a scientific performance evaluation mechanism to evaluate the execution efficiency through the collection and analysis of multi-dimensional indicators. Specifically, the data processing device first sets a reasonable sampling period (such as 1 second or 5 seconds) and regularly collects performance data in three dimensions: CPU, memory, and I / O. Then, preprocess the collected raw data, including outlier filtering and data smoothing. Next, normalize the indicators with different dimensions through methods such as min-max or z-score to unify their numerical ranges. After that, set weight coefficients according to the impact degree of each type of resource on performance, and calculate the weighted average value to obtain the resource utilization degree. At the same time, count the number of tasks successfully completed within a unit time, divide it by the planned number of completed tasks to get the task completion rate. Finally, combine the resource utilization degree and the task completion rate according to the preset calculation formula to generate an evaluation value reflecting the overall execution efficiency.

[0087] In some embodiments, the calculation of performance evaluation indicators can be achieved in various ways: Optionally, the data processing device can store historical performance data by setting up multiple-level buffers. First, perform real-time filtering on the newly collected data to eliminate obviously abnormal sampling points, then use methods such as moving window averaging to smooth data fluctuations, and finally calculate various indicators based on the processed data; Optionally, the data processing device can also adopt an adaptive weight scheme. By analyzing the correlation between different resource indicators and the actual performance, dynamically adjust the weight coefficients of each indicator in the calculation process to make the evaluation result more accurately reflect the system state. It can be understood that other indicator calculation methods can also be adopted, such as fuzzy comprehensive evaluation, machine learning, etc. The specific calculation method to be adopted can be determined according to actual needs. In addition, the timeliness of indicators needs to be considered during the calculation process, and a time decay factor can be introduced to highlight the importance of recent data.

[0088] In practical applications, data instability problems may occur in performance metric calculations, that is, due to external interference or load fluctuations, the sampled data fluctuates violently, affecting the accuracy of the evaluation results. In response, the data processing device can adopt a hierarchical filtering strategy: first, use methods such as median filtering at the raw data level to remove burst noise; then, stabilize the trend of index changes through techniques such as exponential smoothing at the feature level; finally, introduce a reference mechanism for historical evaluation results at the decision-making level to avoid violent fluctuations in evaluation results. For example, when it is detected that the CPU utilization rate of a certain sample suddenly jumps from 50% to 90%, the historical data can be combined to determine whether this change is reasonable and adjust the data processing method accordingly. Through this multi-level data processing mechanism, the accuracy and stability of performance evaluation can be improved.

[0089] S207. When the execution efficiency value is lower than the preset efficiency threshold, generate an optimized resource configuration for the parallel computing task based on the execution efficiency value.

[0090] Referring to step S104, the data processing device generates an optimized resource configuration.

[0091] S208. Dynamically adjust the resource configuration parameters based on the real-time resource scheduling interface, so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation.

[0092] Referring to step S105, the data processing device smoothly adjusts the resource configuration parameters.

[0093] In some embodiments, the data processing device adopts a multi-step resource adjustment strategy, that is, the data processing device obtains the CPU quota difference, memory limit difference, and I / O bandwidth difference between the default resource configuration and the optimized resource configuration; calculates the CPU adjustment amount, memory adjustment amount, and I / O bandwidth adjustment amount for each adjustment cycle according to the CPU quota difference, memory limit difference, and I / O bandwidth difference; adjusts the CPU quota of the current adjustment cycle through the real-time resource scheduling interface and collects the real-time CPU utilization rate; when the fluctuation range of the real-time CPU utilization rate is less than the first preset threshold, adjusts the memory limit of the current adjustment cycle through the real-time resource scheduling interface and collects the real-time memory utilization rate; when the fluctuation range of the real-time memory utilization rate is less than the second preset threshold, adjusts the I / O bandwidth of the current adjustment cycle through the real-time resource scheduling interface and collects the real-time I / O waiting time.

[0094] Among them, the resource configuration difference represents the change amount of the optimized configuration relative to the default configuration; the adjustment period refers to the time interval for the data processing device to perform a resource adjustment once; the adjustment amount represents the specific value of modifying the resource parameters in each period; the real-time resource scheduling interface refers to the set of program interfaces for dynamically adjusting system resources; the real-time utilization rate represents the immediate usage status after the resource adjustment; the fluctuation range refers to the change degree of the resource utilization rate before and after the adjustment; the preset threshold represents the benchmark value for judging the system stability.

[0095] It should be noted that the first preset threshold and the second preset threshold are respectively used to judge whether the fluctuations of the CPU utilization rate and the memory utilization rate are within an acceptable range. Among them, the first preset threshold is usually set to a fluctuation range of 5% - 10% to ensure that the system runs stably after the CPU resource adjustment before proceeding to the next memory adjustment; the second preset threshold is usually set to a fluctuation range of 3% - 8% to ensure that the system runs stably after the memory resource adjustment before proceeding to the next I / O bandwidth adjustment. The settings of these two thresholds are based on the system's requirements for resource stability. A smaller threshold means more stringent stability requirements but may slow down the resource adjustment process, while a larger threshold may bring a faster adjustment speed but there is a certain risk of instability.

[0096] When the data processing device determines that the resource configuration needs to be adjusted, it must adopt a progressive adjustment strategy to ensure the stable operation of the system. Specifically, the data processing device first calculates the differences between the optimized configuration and the default configuration in the three dimensions of CPU, memory, and I / O. Then, according to the total difference and the preset adjustment period, it calculates the specific adjustment amount for each period to ensure that a single adjustment will not cause too much impact on the system. Next, it adjusts in the order of CPU, memory, and I / O, and monitors the usage of the corresponding resources after each adjustment. Only when the fluctuation of the CPU utilization rate after adjustment is within an acceptable range, will it continue to adjust the memory configuration; similarly, only when the memory usage is stable, will it perform the I / O bandwidth adjustment. This serial adjustment mechanism can effectively prevent system instability during the resource adjustment process.

[0097] In some embodiments, the dynamic adjustment of resource allocation can be achieved in various ways: Optionally, the data processing device can adopt an adjustment strategy based on feedback control. First, set an initial adjustment step size, and then monitor the system response after each adjustment. When a large fluctuation is detected, automatically reduce the step size, and appropriately increase the step size when the system is stable. Through this adaptive mechanism, the smoothness of the adjustment is ensured; Optionally, the data processing device can also adopt an adjustment strategy based on a prediction model. By establishing a mathematical model between resource adjustment and system response, predict in advance the system fluctuations that may be caused by different adjustment schemes, and select the adjustment path with the least risk. It can be understood that other resource adjustment methods can also be adopted, such as hierarchical adjustment, priority adjustment, etc. The specific adjustment method to be adopted can be determined according to the system characteristics and performance requirements. In addition, when performing resource adjustment, the dependency relationship between different resources also needs to be considered to ensure the rationality of the adjustment order.

[0098] In practical applications, resource adjustment will encounter oscillation problems, that is, due to improper setting of adjustment parameters, the system fluctuates repeatedly between different states and cannot reach a stable state. To this end, the data processing device can adopt a progressive adjustment strategy: First, divide the resource adjustment into multiple smaller adjustment steps, and set a sufficient observation period after each adjustment; Then, by monitoring the response characteristics of the system in real time, including the change trend and fluctuation amplitude of resource utilization rate; When it is found that the system shows an oscillation trend, automatically reduce the adjustment rate or suspend the adjustment. For example, if the CPU utilization rate fluctuates back and forth on both sides of the target value for three consecutive cycles, it means that the current adjustment step size may be too large, and the adjustment amount needs to be reduced or the adjustment period needs to be extended. Through this dynamic adaptation mechanism, system instability during the resource adjustment process can be effectively avoided.

[0099] In some embodiments, the data processing device will evaluate the stability of resource adjustment, that is, the data processing device will calculate the weighted sum of the real-time CPU utilization rate, real-time memory utilization rate, and real-time I / O waiting time to obtain the system stability score for the current adjustment cycle; When the system stability score is greater than the preset stability threshold, the CPU adjustment amount, memory adjustment amount, and I / O bandwidth adjustment amount are used as the adjustment parameters for the next adjustment cycle.

[0100] Among them, the weighted sum represents the comprehensive value obtained by weighted calculation of multiple performance indicators according to their importance; The system stability score refers to the quantitative indicator reflecting the current operating stability of the data processing device; The preset stability threshold represents the reference value for judging whether the system reaches a stable state; The adjustment parameters refer to the specific numerical set used for resource adjustment in the next cycle; Performance fluctuation represents the change amplitude of the system operation index.

[0101] The data processing device needs to evaluate the adjustment effect at the end of each adjustment cycle and provide a decision-making basis for the adjustment of the next cycle. Specifically, the data processing device first collects the real-time performance data of the CPU, memory, and I / O in three dimensions during the current cycle. Then, according to the degree of influence of different indicators on system stability, weight coefficients are set, and these performance indicators are weighted and calculated to obtain a quantitative score reflecting the overall stability. Next, the calculated stability score is compared with a preset stability threshold. When the score exceeds the threshold, it indicates that the current adjustment strategy is effective, and the adjustment amount of this cycle can be used as the adjustment parameter for the next cycle; otherwise, the adjustment strategy needs to be re-evaluated. This feedback-based adjustment mechanism can ensure the continuous effectiveness of resource adjustment.

[0102] In some embodiments, the evaluation of system stability can be achieved in various ways: Optionally, the data processing device can adopt an evaluation method based on statistical analysis. First, perform variance analysis on the time series data of performance indicators to calculate the degree of fluctuation, then combine the mean level of performance indicators for comprehensive scoring, and finally set multiple stability levels to guide subsequent adjustments; Optionally, the data processing device can also adopt an evaluation method based on trend prediction. By establishing a time series model of performance indicators, analyze the change trend and periodic characteristics of the indicators, predict the development direction of system stability, and thus make forward-looking adjustment decisions. It can be understood that other stability evaluation methods can also be adopted, such as fuzzy evaluation, machine learning, etc. The specific evaluation method to be adopted can be determined according to actual needs. In addition, when performing stability evaluation, it is also necessary to consider the performance fluctuations at different time scales, and short-term fluctuations and long-term trends can be analyzed simultaneously.

[0103] In practical applications, false stability problems may occur in stability evaluation, that is, the system seemingly reaches a stable state, but actually is in an unstable critical state and is prone to losing balance due to small perturbations. In this regard, the data processing device can adopt a multi-dimensional stability verification strategy: First, analyze the stability characteristics of the system from multiple time scales, including instantaneous stability, short-term stability, and long-term stability; then, by applying small-amplitude test perturbations, observe the system's recovery ability; finally, combine historical data to analyze the stability performance of the system in similar states. For example, when the system reaches a seemingly stable state, the anti-perturbation ability of the system can be tested by slightly adjusting the load. If the system can quickly return to the stable state, it indicates that the current state is a true stable state. Through this comprehensive verification mechanism, false stable states can be effectively identified and avoided.

[0104] S209. When the execution efficiency value under optimized resource allocation is still lower than the preset efficiency threshold, obtain the load characteristics of the currently executing parallel computing tasks.

[0105] Among them, the load characteristics represent the resource consumption patterns and computing characteristics shown during the operation of parallel computing tasks, including computing density, memory access pattern, I / O operation frequency, etc.; computing density refers to the number of computing operations per unit time; the memory access pattern describes the read and write behavior characteristics of the task to the memory; the I / O operation frequency reflects the frequency of external data exchange of the task.

[0106] When the data processing device finds that the efficiency still fails to meet the standard after resource optimization, it needs to conduct in-depth analysis at the task level. Specifically, the data processing device first collects the runtime statistical information of each parallel task, including CPU usage distribution, memory access statistics, I / O operation records, etc. Then, it extracts features from these statistical data to identify the computing characteristics, data access characteristics, and resource dependency characteristics of the tasks. Finally, it normalizes these feature data to generate a standardized load characteristic description, providing a basis for subsequent task optimization.

[0107] In some embodiments, the acquisition and analysis of load characteristics can be achieved in multiple ways: Optionally, the data processing device can adopt a method based on performance counters to collect microarchitecture event data during task execution through hardware performance counters, such as cache hit rate, branch prediction accuracy, memory bandwidth usage, etc., to characterize the execution characteristics of tasks from the bottom layer; Optionally, the data processing device can also adopt a method based on application profiling to insert performance probes into the task code to record high-level information such as function call relationships, execution time distribution, and resource access patterns. It can be understood that other feature analysis methods can also be adopted, such as behavior modeling, statistical sampling, etc. The specific analysis method to be adopted can be determined according to the monitoring capabilities and analysis requirements of the system. In addition, when conducting feature analysis, the timeliness and representativeness of the data also need to be considered to ensure that the obtained features can truly reflect the execution status of the tasks.

[0108] In practical applications, data sparsity problems will be encountered in load characteristic acquisition, that is, for some tasks, due to short execution time or unstable behavior, it is difficult to collect sufficient feature data. In this regard, the data processing device can adopt a feature enhancement strategy: First, expand the original feature data set through multiple samplings and data accumulation; then, use similarity analysis to find tasks with similar behavior patterns from historical data and draw on their feature information; finally, through a feature fusion algorithm, integrate the directly collected features and reference features to generate a more complete feature description. For example, for short-term running batch tasks, a more reliable load characteristic model can be constructed through data accumulation and statistical analysis of multiple executions.

[0109] S210. Merge or split the parallel computing tasks based on the load characteristics to generate a new set of parallel computing tasks.

[0110] Among them, task merging means combining multiple tasks with relatively low computational loads or strong data dependencies into one task; task splitting refers to dividing a task with an overly heavy computational load into multiple smaller subtasks; task set restructuring is the process of adjusting the task granularity through merging or splitting operations; data dependency relationships represent the input-output associations between tasks.

[0111] The data processing device needs to perform reasonable task restructuring according to the load characteristics of the tasks. Specifically, the data processing device first analyzes the computational complexity and resource requirements of each task to identify tasks with overly heavy or light loads. Then it constructs a data flow relationship graph between the tasks to evaluate the data interaction overhead between the tasks. Next, according to the task characteristics and system status, it formulates a task restructuring strategy to determine the tasks that need to be merged or split. Finally, it performs the task restructuring operation to adjust the computational boundaries and data flow directions of the tasks, generating a new task set.

[0112] In some embodiments, task restructuring can be achieved in multiple ways: Optionally, the data processing device can adopt a graph partitioning-based method, model the task dependency relationships as a directed graph, and determine the task merging and splitting schemes through the minimum cut algorithm to minimize the data transmission overhead between tasks; Optionally, the data processing device can also adopt a load balancing-based method, evaluate the effects of different restructuring schemes through a task load prediction model, and select a restructuring strategy that can achieve the most balanced computational load. It can be understood that other restructuring methods can also be adopted, such as greedy algorithms, dynamic programming, etc. The specific restructuring method to be adopted can be determined according to the task characteristics and system requirements. In addition, when performing task restructuring, the overhead of the restructuring operation itself also needs to be considered to avoid overly frequent adjustments.

[0113] In practical applications, task restructuring will encounter task dependency conflict problems, that is, there are complex dependency relationships between some tasks, and improper restructuring may lead to deadlocks or performance degradation. In response to this, the data processing device can adopt a dependency-aware restructuring strategy: First, establish a complete dependency chain of the tasks, including direct dependencies and indirect dependencies; then through critical path analysis, identify the critical tasks that affect the overall execution time; finally, on the premise of ensuring that the dependency relationships are not damaged, optimize the organization method of the tasks. For example, for a task sequence with pipeline characteristics, the overall performance can be optimized by adjusting the granularity of the pipeline stages.

[0114] In some embodiments, the data processing device optimizes the task structure based on dependencies. That is, the data processing device obtains the computational complexity and data dependencies of each parallel computing task, and builds a task association graph; calculates the data transfer overhead between tasks according to the task association graph, and constructs a cost function for task reorganization based on the data transfer overhead; when it is determined based on the cost function that the data transfer overhead between adjacent tasks is greater than a preset transmission threshold, marks the adjacent tasks as tasks to be merged; when it is determined based on the cost function that the computational complexity of a single task is greater than a preset computational threshold, marks the single task as a task to be split.

[0115] Among them, the computational complexity represents the amount of computing resources and time overhead required for task execution; the data dependency refers to the input-output association between tasks; the task association graph represents a directed graph structure describing the dependencies between tasks; the data transfer overhead represents the resource consumption of data transmission and communication between tasks; the cost function is a mathematical model used to evaluate the pros and cons of a task reorganization plan; the preset transmission threshold represents the data transfer overhead benchmark value for judging whether tasks need to be merged; the preset computational threshold represents the computational complexity benchmark value for judging whether tasks need to be split.

[0116] The data processing device needs to reasonably reorganize the parallel computing tasks to improve the execution efficiency. Specifically, the data processing device first analyzes the computational characteristics of each task, including the amount of computation, algorithmic complexity, resource requirements, etc., and at the same time identifies the data transfer relationships between tasks. Then it constructs a directed graph reflecting the task dependencies, where nodes represent tasks and edges represent data flow directions. Next, a cost function is constructed based on factors such as the data transfer volume and communication frequency between tasks, which is used to quantitatively evaluate the overhead of task reorganization. When the data transfer overhead between adjacent tasks exceeds the preset threshold, it indicates that there is frequent data interaction between these tasks, and they should be considered for merging to reduce the communication overhead. Similarly, when the computational complexity of a single task exceeds the preset threshold, it indicates that this task will become a performance bottleneck, and it should be considered to be split into smaller-grained subtasks.

[0117] In some embodiments, the decision-making for task reorganization can be achieved in various ways: Optionally, the data processing device can adopt a reorganization method based on graph analysis. First, it uses a graph clustering algorithm to identify the tightly connected subgraphs in the task association graph, then evaluates the data flow characteristics within each subgraph, and finally determines the set of tasks that can be merged according to the preset merging conditions, while determining the large tasks that need to be split through load analysis. Optionally, the data processing device can also adopt a reorganization method based on performance modeling. By establishing a prediction model for task execution time and resource consumption, it simulates the system performance under different reorganization scenarios and selects the reorganization strategy that can maximize the execution efficiency. It can be understood that other task reorganization methods can also be adopted, such as heuristic search, dynamic programming, etc. The specific reorganization method to be adopted can be determined according to the task characteristics and system requirements. In addition, when performing task reorganization, the overhead of the reorganization operation itself also needs to be considered to ensure that the performance improvement brought by the reorganization is greater than the reorganization overhead.

[0118] In practical applications, task reorganization will encounter the problem of optimization conflicts, that is, task merging can reduce communication overhead but increase load imbalance, and task splitting can improve parallelism but increase scheduling overhead. There is a trade-off relationship between these optimization goals. In response, the data processing device can adopt a multi-objective optimization strategy: First, establish a comprehensive evaluation model that includes multiple optimization goals such as communication overhead, load balance, and scheduling overhead; then balance the importance of each optimization goal by setting the target weights in different scenarios; finally, use a heuristic algorithm to search for the optimal reorganization plan that meets multiple constraint conditions. For example, in an environment with limited network bandwidth, the weight of communication overhead in the optimization goal can be appropriately increased to prioritize reducing data transmission between tasks; while in the case of sufficient computing resources, more attention can be paid to the optimization goal of load balance. Through this dynamic balance mechanism, the most suitable task reorganization plan can be found in different scenarios.

[0119] S211. Reschedule the parallel computing task set based on the optimized resource allocation.

[0120] Among them, rescheduling means making a new round of resource allocation and execution arrangement for the reorganized task set; the scheduling policy defines the execution order and resource allocation scheme of tasks; and the execution queue management refers to organizing and controlling the tasks waiting to be executed and the tasks being executed.

[0121] The data processing device needs to formulate a new execution plan for the reorganized task set. Specifically, the data processing device first evaluates the resource requirements of the new task set, including the amount of computation, memory occupancy, and I / O requirements, etc. Then, based on the optimized resource allocation, appropriate execution resources are allocated to each task. Next, according to the priority and dependency relationship of the tasks, the execution order and parallelism of the tasks are determined. Finally, the tasks are submitted to the execution queue to start a new round of task execution.

[0122] In some embodiments, task rescheduling can be achieved in various ways: Optionally, the data processing device can adopt a priority-based scheduling method, set the execution priority according to the importance and urgency of tasks, and give priority to scheduling the execution of critical tasks; Optionally, the data processing device can also adopt a resource reservation-based scheduling method, reserve dedicated resources for different types of tasks, and avoid performance fluctuations caused by resource competition. It can be understood that other scheduling methods can also be adopted, such as fair scheduling, preemptive scheduling, etc. The specific scheduling method to be adopted can be determined according to the scheduling requirements of the system and the service quality requirements. In addition, when performing scheduling, the load balancing of the system also needs to be considered to avoid hot spot problems in resource usage.

[0123] In practical applications, task rescheduling will encounter the problem of scheduling jitter, that is, due to inaccurate estimation of task execution time or system load fluctuations, the actual execution effect deviates greatly from the expectation. In this regard, the data processing device can adopt an adaptive scheduling strategy: First, establish a prediction model for task execution time, and continuously optimize the prediction accuracy through historical data and real-time feedback; Then set up a scheduling feedback mechanism to monitor the actual execution situation of tasks; Finally, dynamically adjust the scheduling strategy according to the execution deviation, such as adjusting task priorities or reallocating resources. For example, when it is found that the actual execution time of a certain type of task far exceeds the expectation, the parallelism of this type of task can be appropriately reduced to avoid system overload.

[0124] In the embodiments of the present application, due to the adoption of a multi-dimensional analysis method based on task characteristics and system performance indicators, combined with a progressive resource adjustment strategy and a task dynamic reorganization mechanism, it is possible to accurately grasp the resource demand characteristics of parallel computing tasks, achieve precise optimization and smooth transition of resource allocation. At the same time, by establishing a task association graph and data flow analysis, intelligently identifying the dependencies between tasks, and adopting a reasonable task merging and splitting strategy, the problems of inaccurate resource allocation, unsmooth adjustment, and fixed task granularity in the traditional fixed threshold scheme are effectively solved. Through the comprehensive analysis and feedback adjustment of multi-dimensional performance indicators, the reliability of the resource optimization process is ensured.

[0125] The data processing device in the embodiments of the present invention application will be described from the perspective of hardware processing below. Please refer to Figure 3 , which is a schematic structural diagram of an entity device of the data processing device in the embodiments of the present application.

[0126] It should be noted that Figure 3 The structure of the data processing device shown is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.

[0127] Such as Figure 3As shown, the data processing device includes a CPU 301 which can perform various appropriate actions and processes according to the program stored in the ROM 302 or the program loaded into the RAM 303 from the storage section 308, such as executing the method described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.

[0128] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, button switches, etc.; an output section 307 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required so that a computer program read from it can be installed into the storage section 308 as required.

[0129] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the CPU 301, various functions defined in the present invention are executed.

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings.

[0131] Specifically, the data processing device of this embodiment includes a processor and a memory. A computer program is stored on the memory. When the computer program is executed by the processor, the data processing method for the high-performance computing power environment provided in the above embodiment is implemented.

[0132] As another aspect, the present invention also provides a computer-readable storage medium. This storage medium can be included in the data processing device described in the above embodiment; or it can exist alone without being assembled into the data processing device. The above storage medium carries one or more computer programs. When the one or more computer programs are executed by a processor of the data processing device, the data processing device implements the data processing method for the high-performance computing power environment provided in the above embodiment.

[0133] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

[0134] As used in the above embodiments, depending on the context, the term "when..." can be interpreted to mean "if...", or "after...", or "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if detecting (the stated condition or event)" can be interpreted to mean "if determining...", or "in response to determining...", or "when detecting (the stated condition or event)", or "in response to detecting (the stated condition or event)".

Claims

1. A data processing method for a high-performance computing power environment, characterized in that Applied to a data processing device, the method includes: Receiving the input calculation data, performing data sharding processing on the calculation data, and generating multiple parallel computing tasks; Based on the default resource configuration, starting the execution program of the parallel computing tasks, and collecting system performance metrics during the running of the execution program; Calculating the execution efficiency value of the execution program based on the system performance metrics; When the execution efficiency value is lower than the preset efficiency threshold, generating an optimized resource configuration for the parallel computing tasks based on the execution efficiency value; Dynamically adjusting the resource configuration parameters based on the real-time resource scheduling interface, so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation.

2. The method according to claim 1, wherein Before the step of starting the execution program of the parallel computing tasks based on the default resource configuration and collecting system performance metrics during the running of the execution program, the method further includes: Performing clustering analysis on the resource usage data in the historical execution records to obtain multiple resource usage patterns; Calculating the benchmark usage amounts of various resources based on the distribution characteristics of the resource usage patterns; Setting the default resource configuration according to the benchmark usage amounts; the default resource configuration includes the initial allocation values of the number of CPU cores, the memory capacity, and the I / O bandwidth.

3. The method according to claim 1, wherein The step of calculating the execution efficiency value of the execution program based on the system performance metrics specifically includes: Collecting the CPU utilization rate, the memory usage rate, and the I / O waiting time according to a preset sampling period; Performing normalization processing and weighted calculation on the CPU utilization rate, the memory usage rate, and the I / O waiting time to obtain the system resource utilization degree; Obtaining the number of computing tasks completed within a unit time, and calculating the task completion rate; Calculating the execution efficiency value based on the system resource utilization degree and the task completion rate.

4. The method according to claim 1, wherein The step of dynamically adjusting the resource configuration parameters based on the real-time resource scheduling interface, so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation specifically includes: Obtaining the CPU quota difference, the memory limit difference, and the I / O bandwidth difference between the default resource configuration and the optimized resource configuration; Calculating the CPU adjustment amount, the memory adjustment amount, and the I / O bandwidth adjustment amount for each adjustment period according to the CPU quota difference, the memory limit difference, and the I / O bandwidth difference; Adjusting the CPU quota of the current adjustment period through the real-time resource scheduling interface, and collecting the real-time CPU utilization rate; When the fluctuation range of the real-time CPU utilization rate is less than the first preset threshold, adjusting the memory limit of the current adjustment period through the real-time resource scheduling interface, and collecting the real-time memory usage rate; When the fluctuation range of the real-time memory usage rate is less than the second preset threshold, adjusting the I / O bandwidth of the current adjustment period through the real-time resource scheduling interface, and collecting the real-time I / O waiting time.

5. The method according to claim 4, wherein After the step of when the fluctuation range of the real-time memory usage rate is less than the second preset threshold, adjusting the I / O bandwidth of the current adjustment period through the real-time resource scheduling interface, and collecting the real-time I / O waiting time, the method further includes: Calculate the weighted sum of the real-time CPU usage rate, the real-time memory usage rate, and the real-time I / O waiting time to obtain the system stability score for the current adjustment period; When the system stability score is greater than the preset stability threshold, use the CPU adjustment amount, the memory adjustment amount, and the I / O bandwidth adjustment amount as the adjustment parameters for the next adjustment period.

6. The method according to claim 1, wherein After the step of dynamically adjusting the resource configuration parameters based on the real-time resource scheduling interface so that the default resource configuration of the execution program smoothly transitions to the optimized resource configuration during operation, the method further includes: When the execution efficiency value under the optimized resource configuration is still lower than the preset efficiency threshold, obtain the load characteristics of the parallel computing tasks currently being executed; Based on the load characteristics, merge or split the parallel computing tasks to generate a new set of parallel computing tasks; Reschedule the new set of parallel computing tasks based on the optimized resource configuration.

7. The method according to claim 6, wherein Before the step of merging or splitting the parallel computing tasks based on the load characteristics to generate a new set of parallel computing tasks, the method further includes: Obtain the computational complexity and data dependency relationships of each parallel computing task, and establish a task association graph; Calculate the data transfer overhead between tasks according to the task association graph, and construct a cost function for task reorganization based on the data transfer overhead; When it is determined based on the cost function that the data transfer overhead between adjacent tasks is greater than the preset transmission threshold, mark the adjacent tasks as tasks to be merged; When it is determined based on the cost function that the computational complexity of a single task is greater than the preset computational threshold, mark the single task as a task to be split.

8. A data processing device, characterized in that, The data processing device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the data processing device to execute the method according to any one of claims 1-7.

9. A computer-readable storage medium, comprising instructions, characterized in that, When the instruction runs on the data processing device, cause the data processing device to execute the method according to any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product runs on the data processing device, cause the data processing device to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Computing power data management system and method based on distributed computing

    CN119025283A

  • Chip platform resource dynamic scheduling method

    CN119322684A

  • Python monitoring task resource use method and system

    CN119512883A

  • Big data platform scheduling task and data collaborative smooth migration method and system

    CN119576506A

  • Resource scheduling method, device and equipment

    CN120179414A

Cited By

  • Server data acquisition method and device

    CN120821645A