Heterogeneous computing power dynamic allocation method based on big data analysis
Patent Information
- Application Number
- CN202510870222.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120803705A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of heterogeneous computing power allocation, in particular to a heterogeneous computing power dynamic allocation method based on big data analysis. BACKGROUND
[0002] Heterogeneous computing power allocation involves unified scheduling and allocation of multiple types of computing resources, including coordinated management of multiple different types of computing units such as central processors, graphics processors, and other customized acceleration hardware, aiming to improve overall resource utilization efficiency and the scientific nature of computing task allocation, and meet dynamic computing needs in different business scenarios. Among them, the traditional heterogeneous computing power dynamic allocation method refers to the use of a preset static allocation strategy or a rule-based scheduling method based on current system load to assign specific type tasks to designated computing units in a large-scale data processing and multi-type computing resource environment. Common implementation methods include static resource binding based on task type or adjusting the allocation ratio according to the real-time utilization rate of the computing unit, but fail to conduct in-depth analysis on the dynamic changes of computing power demand in a big data analysis environment, resulting in allocation decision lag and low resource utilization.
[0003] The existing technology mainly depends on task type and current node utilization rate, and only makes allocation decisions after the arrival of tasks based on static rules or simple load information, lacking deep analysis of features based on multi-source data tags. In the face of varying task characteristics, sudden demand increases or resource heterogeneity environment, it is difficult to effectively determine the matching degree between tasks and computing resources, and there is a lack of quick response to node state changes and task structure adjustments during the allocation process, resulting in long-term idle resources or high local node load. The allocation strategy is difficult to adapt to dynamic task distribution and real-time resource fluctuations, affecting the balance and agility of overall scheduling. SUMMARY
[0004] The purpose of the present application is to solve the problems existing in the prior art and to provide a heterogeneous computing power dynamic allocation method based on big data analysis.
[0005] In order to achieve the above purpose, the present application adopts the following technical scheme: a heterogeneous computing power dynamic allocation method based on big data analysis, comprising the following steps:
[0006] S1: Based on the initial stage of data task, analyze the data task source and target, map the source type and target path to feature parameters according to a unified standard, compare the task type and dependency attribute, classify the emergency level, combine the performance of heterogeneous computing power resources, and calculate the maximum item, minimum item and balance index of the parameter to obtain normalized feature data;
[0007] S2: Based on the normalized feature data, judge the change of the balanced index and the business scene reference data, analyze the difference between the two, identify the key task according to the adaptability of the computing resource, adjust the task scheduling state, and obtain the adaptive task index set;
[0008] S3: Based on the adaptive task index set, calculate the task queue number of the computing resource, compare the switching efficiency combined with the historical scheduling switching situation, judge the task processing capacity and analyze the energy consumption, optimize the idle window and abnormal state discrimination, and obtain the computing power running distribution identifier;
[0009] S4: Based on the computing power running distribution identifier, screen the computing power resource running normally, analyze the frequency and temperature data, compare the bandwidth and temperature control range, calculate the response performance, judge the relationship between the task concurrency number and the running state, adjust the running level, and obtain the resource state classification mark.
[0010] The application improves that the normalized feature data includes task label dimension, parameter normalization information and feature aggregation result, the adaptive task index set includes task priority index, distribution mark information and scheduling adaptation type, the computing power running distribution identifier includes resource distribution information, running label feature and node exception identifier, and the resource state classification mark includes performance interval label, load state type and stability grouping number.
[0011] The application improves that the acquisition step of the normalized feature data is specifically:
[0012] S111: Based on the data task initial stage, analyze the source type information and target path data of the data task, calculate the structural difference between the source and the target, compare the task type and the dependency relationship, judge the consistency of the task parameter mapping, and obtain the source target structure deviation;
[0013] S112: According to the source target structure deviation, calculate the emergency level label of each task, compare the matching relationship between the emergency level and the computing resource performance label, screen the tasks meeting the feature requirements of each group, and generate the emergency level resource adaptation group;
[0014] S113: Based on the emergency level resource adaptation group, calculate the normalized performance of each group of tasks, judge the balanced relationship between the tasks, screen the key performance of each feature dimension, and obtain the normalized feature data.
[0015] The application improves that the acquisition step of the adaptive task index set is specifically:
[0016] S211: Based on the normalized feature data, analyze its balance performance with the business scenario reference data, combine the task labels, parameter normalization content and feature aggregation content, compare the type and dependent parameters of each task, determine the offset of the task in the feature distribution, and generate the task feature offset;
[0017] S212: Based on the task feature offset, the performance labels of each node in the heterogeneous computing resource pool are compared to analyze the computing power, response capability, and energy consumption characteristics of the node. By matching the computing requirements of the task, the adaptation between each task and the resource node is screened to obtain the task resource adaptation difference;
[0018] S213: Based on the task resource adaptation difference, adjust the scheduling relationship between the task and the heterogeneous computing resources, determine the matching performance of each task in the current resource scheduling, analyze the tasks that need to be adjusted first in the scheduling state, and obtain the adaptation task index set.
[0019] The present invention is improved in that the steps for obtaining the computing power operation distribution identifier are specifically as follows:
[0020] S311: Based on the adapted task index set, calculate the change in the number of tasks queued for each heterogeneous computing power node in a continuous scheduling cycle, analyze the task type distribution and scheduling completion status, determine the task processing status of the computing power node in the same cycle, identify the behavior pattern of inconsistent processing speed, and obtain a processing capacity fluctuation sequence;
[0021] S312: Based on the processing capacity fluctuation sequence, select computing nodes with concentrated processing fluctuation resources, calculate the changing trends of task waiting time and energy consumption records of the nodes, determine whether task response delays occur repeatedly, determine stability characteristics, and obtain an abnormal node sequence;
[0022] S313: Based on the abnormal node sequence, analyze the combined characteristics of the load quantity, task response interval, energy consumption change and main frequency fluctuation of each type of node, calculate the computing power response fluctuation intensity, determine whether the response capability presents a periodic unstable characteristic, and obtain the computing power operation distribution identifier.
[0023] The present invention is improved in that the steps of obtaining the resource status classification label are specifically as follows:
[0024] S411: Based on the computing power operation distribution identifier, analyze the main frequency, temperature, and queue length parameters, determine whether each parameter meets the normal operation standard, select computing power nodes that meet the standard, and obtain a screening set of operating nodes;
[0025] S412: Based on the running node screening set, the frequency and temperature of each computing power node are compared, the frequency bandwidth adaptation difference and the temperature control adaptation difference are calculated in combination with the standard running bandwidth and the temperature control range, and a node response performance sequence is obtained;
[0026] S413: According to the node response performance sequence, the number of concurrent tasks of the computing power node is adjusted, the synchronous change of the frequency and the temperature is analyzed, the stress and the response state of the node in the task running process are calculated, the belonging category of the computing power node is judged, and a resource state classification mark is obtained.
[0027] The application improves that the step further comprises:
[0028] S5: Based on the resource state classification mark, the frequency change trend and the temperature rise rate of the computing power resource are analyzed, the task emergency level priority is judged, the computing power resource with stable frequency and gentle temperature change is identified, the allocation relationship is optimized, and a resource dynamic allocation sequence is obtained.
[0029] The resource dynamic allocation sequence comprises a task allocation structure, a computing power mapping relationship and a priority allocation sequence.
[0030] The application improves that the acquisition step of the resource dynamic allocation sequence is specifically:
[0031] S511: Based on the resource state classification mark, the frequency change and the temperature rise process of each computing power node in a continuous task scheduling period are analyzed, whether the frequency change of each computing power node is stable and whether the temperature rise is uniform are judged according to the frequency and temperature monitoring information at each time point, and a frequency and temperature change feature is generated.
[0032] S512: Based on the frequency and temperature change feature, the computing power node with stable frequency and uniform temperature change is screened, the high-priority task is compared with the screened node in combination, the scheduling matching between the node and the task is judged, and a task node adaptation combination is obtained.
[0033] S513: Based on the task node adaptation combination, the allocation relationship between the task and the computing power node is adjusted, the running state grouping and the task distribution of each node are analyzed, the task scheduling sequence and the computing power node bearing distribution are optimized, and a resource dynamic allocation sequence is obtained.
[0034] Compared with the prior art, the application has the advantages and positive effects that:
[0035] In the present application, by standardizing the multi-dimensional label of data tasks, normalizing mapping and feature aggregation, the depth mining and classification of task features can be realized, the dynamic linkage of task demand and computing resource state is completed according to big data analysis, the correlation between tasks and heterogeneous computing resources is finely quantized, the resource response performance and grading mechanism are combined, the task allocation strategy is automatically adjusted, the resource adaptation degree and allocation sequence matching during task processing are guaranteed, the real-time response to dynamic computing pressure and complex task structure is strengthened, the sensitivity of resource cooperation and scheduling is improved, and the adaptability of the allocation mechanism to sudden tasks and fluctuating scenes is effectively enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 The main step flowchart of the present application is shown in the figure;
[0037] Figure 2 The acquisition flowchart of normalized feature data in the present application is shown in the figure;
[0038] Figure 3 The acquisition flowchart of the adaptive task index set in the present application is shown in the figure;
[0039] Figure 4 The acquisition flowchart of the computing power running distribution identifier in the present application is shown in the figure;
[0040] Figure 5 The acquisition flowchart of the resource state grading mark in the present application is shown in the figure;
[0041] Figure 6 The acquisition flowchart of the resource dynamic allocation sequence in the present application is shown in the figure. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0043] In the description of the present application, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0044] EMBODIMENT
[0045] Please refer to Figure 1The application provides a technical scheme: a heterogeneous computing power dynamic allocation method based on big data analysis, comprising the following steps:
[0046] S1: based on the initial stage of the data task, analyzing the data task source and target information, mapping the source type and target path by using a unified standard, comparing the task type and dependent parameters, classifying the emergency level in a structured manner, and combining the performance label of the heterogeneous computing power resource to calculate the maximum item, minimum item and balance index of each parameter, obtaining normalized feature data;
[0047] S2: based on the normalized feature data, judging the amplitude change of the balance index and the business scenario reference data, analyzing the difference between the two, and identifying the key task of the change amplitude according to the adaptability of the heterogeneous computing power resource, adjusting the state of the task in the computing power resource scheduling as a to-be-allocated object, obtaining an adaptive task index set;
[0048] S3: based on the adaptive task index set, calculating the task queue number of the computing power resource, comparing the computing power resource switching efficiency through the historical scheduling switching situation, judging the task processing capacity of each computing power resource in the difference period, and combining the energy consumption condition to optimize the discrimination logic of the idle window and the computing power abnormal state, obtaining a computing power running distribution identifier;
[0049] S4: based on the computing power running distribution identifier, screening the computing power resources running normally, analyzing the main frequency and temperature of each computing power resource, comparing the standard running bandwidth and temperature control range, calculating the response performance of the computing power resource, judging the relationship between the task concurrency number and the computing power resource running state, and adjusting the computing power resource running level, obtaining a resource state classification label;
[0050] S5: based on the resource state classification label, analyzing the main frequency change trend and temperature rise rate of the computing power resource, judging the priority of the task emergency level, identifying the computing power resource with stable main frequency and gentle temperature change, adjusting the allocation relationship between the computing power resource and the high-priority task, and obtaining a resource dynamic allocation sequence.
[0051] The normalized feature data includes task label dimensions, parameter normalization information and feature aggregation results, the adaptive task index set includes task priority index, allocation mark information and scheduling adaptation type, the computing power running distribution identifier includes resource distribution information, running label features and node exception identifier, the resource state classification label includes performance interval label, load state type and stability grouping number, and the resource dynamic allocation sequence includes task deployment structure, computing power mapping relationship and priority allocation sequence.
[0052] In S1, the unified standard refers to the same set of data formatting, normalization and numbering rules for all input data task sources, targets, types, dependencies, urgency levels and other parameters. Different sources of diversified data parameters are standardized to the same dimension and data structure, which is convenient for subsequent comparison and calculation. Mapping is to convert the original parameters (such as source type, target path) into numbers, labels or standardized features according to the "unified standard" through a corresponding relationship table, eliminating the differences in expression methods and data structures of various task parameters. Dependency parameters refer to the attribute information of other data, processes, pre-task resources that a task needs to depend on in each data task, such as data dependency, dependency level, dependency category, etc., which are commonly used to determine task scheduling priority and resource allocation relationship. Structured mode refers to qualitative grouping and labeling of non-quantitative parameters such as task urgency level through pre-set data structures (such as hierarchical table, enumeration type classification), which are standardized for operation and logical judgment. Heterogeneous computing resources refer to diversified computing nodes and computing units composed of different types (such as CPU, GPU, FPGA, special-purpose acceleration chips, etc.) on data centers or computing platforms, which have different computing performance, energy consumption, bandwidth, response characteristics, etc. The maximum item is the feature item with the largest value in a group of parameters. The minimum item is the feature item with the smallest value. The balanced index is the mean value of all standardized parameters or other statistical quantities reflecting the balance of overall characteristics, which is used to measure the overall performance and distribution trend of a single task in multiple dimensions.
[0053] In S2, the amplitude change refers to the numerical difference between the balanced index and the business scenario reference parameter (i.e. the preset target data or industry experience data), which is usually used to measure the distance between the actual task characteristics and the ideal business demand. The difference between the two is the difference between the normalized feature data and the business scenario reference data after comparison and calculation, which can be expressed as positive or negative deviation or similarity size, and is used for task screening. Adaptability describes the matching degree of a task in the current heterogeneous computing resource pool to achieve efficient scheduling and allocation, including performance, energy consumption, structure and other dimensions. The change amplitude is a task with a significant difference (significant deviation) or a high match (high match) between the normalized feature data and the reference data, which is the object that needs to be focused on or prioritized in resource allocation strategy. The to-be-allocated object is the task or task number set selected after the above comparison and judgment to participate in the subsequent resource allocation process.
[0054] In S3, the number of task queues refers to the number of data tasks currently waiting for processing or being processed by each computing resource node (such as a GPU or an FPGA), which reflects the resource load and the degree of busyness; the computing resource switching efficiency indicates the efficiency required for the computing resource to switch from processing one task to another during task allocation and scheduling, which can be represented by task switching frequency, switching time consumption, etc., and measures the agility of node scheduling; the task processing capability refers to the number of data tasks that can be completed by the computing resource node within a certain period of time or the data processing speed, which is usually related to the computing performance, bandwidth, memory, etc. of the node hardware resources; the idle window refers to the idle time segment or available time interval of the computing resource node during no task processing or waiting for a new task, which is used to dynamically determine the schedulability of the resource; the computing abnormal state refers to abnormal conditions that occur during the operation of the resource node, such as sudden increase in computing delay, abnormal increase in energy consumption, queue congestion, etc., which affect task scheduling and system stability.
[0055] In S4, running normally means that all indicators (frequency, temperature, queue length, energy consumption, etc.) of the computing resource node at the current time are within the preset normal range, and there is no abnormal mark; the response performance refers to the actual response of the computing resource node when receiving and processing task requests, including response speed, processing time, stability, etc.; the number of concurrent tasks indicates the number of data tasks processed in parallel by a single computing resource node at the same time, reflecting the concurrent processing capability and bandwidth utilization of the node; the running level is the level or category divided according to the current frequency, temperature, energy consumption, concurrency, etc. of the node, which is used for resource priority, allocation strategy adjustment, etc.
[0056] In S5, the frequency change trend refers to the fluctuation trajectory and change direction of the frequency (i.e. the clock frequency of CPU / GPU, etc.) of the computing resource node within a period of time, reflecting the dynamic changes of computing capability scheduling; the high-priority task is a task that is considered to be more in need of priority scheduling among a group of tasks according to the standards of emergency level, dependency degree, business importance, etc.; the allocation relationship refers to the mapping and binding state between the final established task and the computing resource node, i.e. which heterogeneous computing node actually carries each task.
[0057] Please refer to Figure 2 The step of obtaining normalized feature data is specifically:
[0058] S111: Based on the initial stage of data tasks, analyze the source type information and target path data of the data tasks, calculate the structural difference between the source and the target, compare the task types and dependency relationships, judge the consistency of task parameter mapping, and obtain the source-target structure deviation;
[0059] Extract the "source type" field from the task configuration file, which is generally defined by the data source system, such as "log system", "business database", "IoT edge device", etc., and convert it to a unified label format such as "source type 1", "source type 2", etc., then parse the "target path" field and unify it to a standard format path expression such as " / data / warehouse / zone1 / tableA", and extract the total number of fields, nested levels, data type structure, etc. after classifying the original path. In the analysis of source and target structure, extract the field number, field type sequence, and hierarchical structure expression one by one, and then compare the structural characteristics between the two, for example, there are 10 fields in the source and 12 fields in the target, the source contains 6 numerical types and 4 string types, while the target contains 5 numerical types and 7 string types, the difference in field number is 2, the difference in string field is 3, and the difference in nested level is 1. After the difference is quantified, the structural deviation between the source and the target is evaluated, and the task type field is further analyzed, such as "batch processing", "real-time streaming", "periodic task", etc. After classification and numbering, it is mapped, and the task dependency information is extracted, including the number of dependent tasks, whether it depends on external interfaces, etc. The dependency relationship is converted into dependency depth, dependency path length, etc. Parameters are matched with the complexity of task type and dependency, for example, a task type is "batch processing", and the number of dependent tasks is 3, one of which comes from an external system, marked as high dependency strength. Then compare the task with the target task type and dependency parameters at the field level, check if the task type is the same, the number of dependencies is equal, and the dependency source is consistent, etc. If any parameter is inconsistent, it is recorded as a deviation item. Finally, all comparison results are summarized to get the overall difference between the source and the target in structure, type, and dependency path. By combining the structural field difference, type mismatch, and dependency parameter inconsistency, the mapping consistency between the source and target structures is judged, and the source-target structure deviation is formed.
[0060] S112: According to the source-target structure deviation, calculate the emergency level label of each task, compare the matching relationship between the emergency level and the performance label of the computing resource, and select the tasks that meet the requirements of each group of characteristics to generate an emergency level resource adaptation group.
[0061] First, set the five-level segmentation threshold of structural deviation, divide the tasks into emergency levels 1 to 5 from low to high according to the deviation value, for example, the deviation value less than 0.3 is level 1, and more than 1.2 is level 5, the deviation value of each task is located in the interval, and the corresponding emergency level label is given, then the performance label information of each resource node in the heterogeneous computing resource pool is extracted, the resource type (such as CPU, GPU, FPGA, etc.) is read one by one, the frequency value, the number of supported threads, the system available memory and the average response time, the resources with the frequency higher than 3.5GHz, the thread number not less than 16, the memory capacity more than 64GB and the average response time less than 20ms are marked as high-performance nodes, and the resources with the frequency lower than 2GHz, the thread less than 8 and the memory less than 16GB are marked as low-performance nodes, then the task emergency level and each performance interval node are matched one by one, for the tasks with emergency level 1 or 2, only the adaptability of low-performance nodes is considered, if at least two of the node frequency, memory and response speed meet the requirements, it is recorded as a suitable task, for the tasks with level 4 or 5, only in the high-performance node is screened, and the same three performance requirements are met as the adaptation condition, each task is matched with the resource node list one by one, the corresponding relationship between the task and the resource node is constructed, the task number and its emergency level, the number of allocable resource nodes are combined to form a structured record, for example, the task "data_stream_07" is with emergency level 5, and the adaptation node number is "GPU04" and "FPGA02", then the adaptation entry {data_stream_07, 5, [GPU04, FPGA02]} is established, after all the tasks are matched in this way, the adaptation resource nodes of each group of tasks are arranged according to the task emergency level, and the emergency level resource adaptation group is formed.
[0062] S113: Based on the emergency level resource adaptation group, the formula is used:
[0063]
[0064] The normalized performance of each group of tasks is calculated, the balance relationship between tasks is judged, the key performance of each feature dimension is screened, and the normalized feature data is obtained, wherein, FS ij represents the normalized difference performance of task i under resource j, PS i represents the path type structure deviation amplitude of task i, TS j represents the performance label mapping amount of resource j, RS j represents the paired feature amplitude of resource j, LS jk represents the normalized amplitude of resource j in the kth feature dimension, and n is the combined dimension of resource label.
[0065] The normalized difference performance is used to measure the comprehensive adaptation degree and feature difference between the data task and the heterogeneous computing resource under the standardized parameters, and is an important parameter for priority discrimination and grouping scheduling in resource dynamic allocation.
[0066] FS ij represents the normalized difference performance of task i on resource j, PS i is the task structure offset, which is directly derived from the difference between the source structure code and the target path code in S111; TS j is the resource performance label quantization item, which is mapped after the dimension unification of multiple indicators such as the frequency of the hardware device, the operation bandwidth, and the number of concurrent threads; RS j is the feature offset amplitude generated by the pairing of the resource and the task label, considering factors such as scheduling period difference, task load offset, and bandwidth fitting degree; LS jk is the standard feature performance of the resource in the kth feature dimension (such as the normalized energy consumption index, the normalized temperature distribution index, and the normalized bandwidth index), which is a normalized coefficient after dimension standardization, and the total number of dimensions is set to n = 3. Taking the pairing of task T003 and resource j as an example for illustration:
[0067] The source structure code of task T003 is 11, and the target path code is 17, so PS i = 6;
[0068] The performance label TS j of the resource paired with this task is 8.1, the resource pairing feature amplitude RS j is 5.5, and the normalized features of the resource in the energy consumption, temperature, and bandwidth dimensions are LS j1 = 0.7, LS j2 = 0.6, and LS j3 = 0.8.
[0069] Substitute the data into the formula for step-by-step calculation as follows:
[0070] First, calculate the square root term:
[0071]
[0072] Then calculate the denominator term:
[0073]
[0074] Substitute the calculation of the overall expression:
[0075]
[0076] Therefore, the normalized difference of task T003 under the current resource combination is-1.22, which is used to represent the difference between the task and the resource in the structure adaptation and performance matching dimension. All such values of all tasks will be used in subsequent stages to form sorting or integration processing as data input for dynamic scheduling strategies. The formula calculation logic is: the numerator part represents the difference between the structure deviation of the task itself and the overall performance composition strength of the resource, the denominator part uses the normalized sum of the resource characteristics in each dimension as an adjustment factor to set a punishment balance mechanism for tasks with large differences, avoid judgment deviation caused by a single characteristic in task resource adaptation, and overall reflect the comprehensive adaptation degree of the task in the current resource characteristic mapping structure.
[0077] Please refer to Figure 3 The acquisition step of the adaptive task index set is specifically:
[0078] S211: Based on the normalized feature data, analyze the balanced performance between the normalized feature data and the business scenario reference data, compare the types and dependent parameters of each task, judge the deviation of the task in the feature distribution, and generate the task feature deviation;
[0079] The normalized parameter vector of each task is divided into three components: task label information, parameter normalization value, and feature aggregation vector. The task label information is derived from the fields extracted from the task initialization metadata, including task type label, emergency level marker, and dependency depth identifier. The parameter normalization value is derived from the performance-related values extracted after structured analysis of the task, such as data volume, processing period, dependency path length, and response time target. These values are converted to standardized decimal values between 0 and 1 through uniform normalization operation. The feature aggregation vector is a multi-dimensional feature measurement vector formed by weighted aggregation of multiple normalized parameters. Then, the business scenario reference data is read, which is a collection of standard feature samples induced from historical typical tasks. After classification by task type, the reference equilibrium vector is established, which corresponds to the ideal distribution median of each parameter in different task categories. Then, the normalized parameters of the current task are compared with the reference vector item by item. For each dimension of the feature, the absolute difference between the normalized parameter and the reference value is calculated. After averaging or median processing of all dimension differences, the equilibrium deviation degree of the task is formed. Then, the task type field is reclassified, such as task T01 for "periodic processing" class and task T02 for "real-time response" class. The dependency depth, cross-system dependency marker, and dependency task value in the task dependency parameter are called respectively. The offset degree of the task and the reference scene in the dependency structure is compared. Each task's dependency parameter is compared with the reference value item by item. For example, the dependency depth of T01 is 3, the reference task is 1, the deviation is 2, the cross-system dependency marker is yes, the reference task is no, the deviation is 1, the number of dependent tasks is 4, the reference value is 2, and the deviation is 2. Finally, the dependency parameter deviation values are added up and weighted to construct the dependency feature offset of the task. The parameter normalization value and the dependency offset value are merged. The total offset of the task is the fusion result of all dimension deviations. For example, the parameter offset score of T01 is 0.26, the dependency offset is 5, and the final feature offset is 0.53 after weighting. The task feature offset is generated.
[0080] S212: Based on the task feature offset, compare with the performance label of each node in the heterogeneous computing resource pool, analyze the operation ability, response ability and energy consumption characteristics of the node, match the computing demand of the task, filter the adaptation between each task and resource node, and obtain the task resource adaptation difference.
[0081] The performance tags of all computing nodes in the heterogeneous computing resource pool are uniformly extracted and structured, including computing node types (such as GPU, CPU, FPGA, etc.), running frequency, response time, task throughput per unit time, and performance output value per unit power consumption. Then the computing resource demand indicators of each task are read, and the computing scale, task concurrency, required response time, etc. are extracted from the task initialization information and normalization parameters, and are compared with the corresponding performance items of the resource nodes one by one to determine whether each performance parameter can meet the task demand threshold. The standard is set as follows: if the task response time target is less than 20ms, the resource node response capability needs to be less than or equal to 15ms to be considered to meet the requirement; if the task computing scale is high, the resource throughput capability needs to be greater than 1000ops / sec to be considered to match; if the task concurrency is higher than 10 threads, the minimum concurrent thread number of the node needs to be greater than or equal to 12 threads or more; if any performance dimension does not meet the above constraints, it is recorded as a mismatch item. The matching condition of each task and each computing node is recorded by the number of matching items, for example, task T05 and GPU node G01, which meet the response and concurrency but the power consumption is higher than the tolerance value, the number of matching items is 2, and the difference is 1. The matching difference value between the task and all nodes is then calculated and merged to obtain the average adaptation difference value between each task and all resource nodes. The higher the value, the stronger the inadaptability between the current task and the resource. Further, the resource matching results of all tasks are sorted to form an adaptation difference list between the tasks and the resources. In combination with the tasks with a task offset greater than 0.5 as the key analysis object, the node with the lowest adaptation degree is selected as the poorly adapted node in the resource matching, and the node number and task number are recorded to form an adaptation difference record set, and the task resource adaptation difference is obtained.
[0082] S213: Based on the task resource adaptation difference, the scheduling relationship between the task and the heterogeneous computing resource is adjusted, the matching performance of each task in the current resource scheduling is determined, the task that needs to be adjusted in priority is analyzed, and an adaptation task index set is obtained.
[0083] First, read the task allocation state record in the current scheduling system, check whether the resource node corresponding to each task is located in the high difference interval, set the matching evaluation benchmark: if the adaptation difference value between the task and the allocated node exceeds 0.6, it is considered to be not adapted, and marked as needing adjustment state, then traverse all task scheduling allocation relations, according to the "task number-resource number-difference value" structure in the scheduling record, the scheduling adjustment judgment of each task state is carried out in turn, if the difference value of a task matching node is lower than 0.3, the task characteristic offset is lower than 0.4, the allocation state is reserved, if the matching difference value of a task exceeds 0.7, and the task offset is higher than 0.6, the task is marked as "urgent adjustment task", and added to the priority adjustment queue, combined with the available resource state table in the scheduling system, the current idle or low load resource node is screened, and it is judged whether the node meets the minimum three performance index requirements of the high difference task, if it meets, the resource is taken as the candidate replacement node, the task number, current node number and recommended replacement node number are recorded, after the scheduling relationship of all replacement recommended tasks is summarized, the task scheduling adjustment suggestion table is formed, the scheduling adjustment mark is generated for each task, the task set in the need to reschedule or priority replacement state is arranged according to the task number, and the adaptive task index set is obtained.
[0084] Please refer to Figure 4 , the acquisition step of the computing power running distribution identifier is specifically:
[0085] S311: Based on the adaptive task index set, calculate the task queuing number change of each heterogeneous computing power node in the continuous scheduling period, analyze the task type distribution and scheduling completion state, judge the task processing situation of the computing power node in the same period, identify the behavior mode of inconsistent processing speed, and obtain the processing capacity fluctuation sequence;
[0086] The task queue records of each heterogeneous computing power node in the continuous scheduling period are called, the number of tasks to be processed of each node in each period is counted according to the scheduling time axis in the period unit (such as every 5 minutes as a period), a sequence structure of the scheduling period and the queue number is formed, for example, the queue number of GPU node G02 in 10 periods is [3, 4, 5, 8, 9, 11, 7, 5, 3, 2] in turn, then the type distribution of each node is counted respectively, the type tags of the tasks entering the queue are extracted from each period, such as "periodic task", "stream computing task", "delay tolerant task", etc., the types are classified and counted, whether the proportion of different task types in the same node in different periods fluctuates is compared, then the completion time of each task is read according to the scheduling completion record, the number of tasks completed in each period is calculated, the number of completions and the number of queues are compared, the task backlog or release trend is obtained, the number of tasks processed by the node in the same period is classified and counted, if the number of tasks queued in the adjacent period continues to increase but the number of completions does not increase synchronously, it is judged that the processing speed of the node does not keep up with the task arrival speed, and it is marked as a period section with slowed processing speed, if the number of task completions of a node fluctuates greatly in different periods, for example, 5 in the 3rd period, 1 in the 4th period, and 6 in the 5th period, the difference value of the processing speed fluctuation item is recorded as 5, the change difference value of the completion quantity between periods is calculated to form a fluctuation value sequence, a processing speed fluctuation reference value is further set, when the fluctuation difference value of the node is greater than 3 for three consecutive periods, it is judged that there is inconsistent processing behavior, the fluctuation behavior is recorded in the task behavior log, and the cross analysis result of the task type change, the processing quantity change and the queue number of the node in the whole observation period is extracted, which is marked as "high fluctuation node", the fluctuation difference value sequence of all nodes is numbered and classified to form a processing capacity fluctuation sequence.
[0087] S312: Based on the processing capacity fluctuation sequence, the resources in the processing fluctuation set of the computing power node are screened, the change trend of the task waiting time and the energy consumption record of the node is calculated, whether the task response delay occurs repeatedly is judged, the stability feature is determined, and the abnormal node sequence is obtained;
[0088] The fluctuation records of each computing power node are screened, the number of periods in which the fluctuation amplitude exceeds the specified threshold in multiple periods is read and accumulated, and if a node has 9 periods of fluctuation difference exceeding the set threshold of 3 in 20 periods of the observation period, the fluctuation period ratio is 45%, the set fluctuation concentration judgment threshold is 40%, and if it exceeds this threshold, it is judged that the node is a fluctuation concentration node, and it is added to the resource list of fluctuation concentration. Subsequently, the task waiting time record of such nodes in the above period is extracted, the average value of the waiting time of all tasks in the waiting queue in each period is calculated, and the period average waiting time sequence is obtained. If the average waiting time repeatedly rises and then falls in multiple periods, for example, period 1 is 12 ms, period 2 rises to 25 ms, period 3 falls to 13 ms, and period 4 rises to 27 ms again, it is recorded as a repeated fluctuation behavior of waiting time. The energy consumption records of the node in the corresponding period are read synchronously, such as power consumption, temperature change, etc., the fluctuation rate of the energy consumption between periods is calculated, and if there is a short-time peak in the energy consumption between periods, for example, the power consumption of period 5 is 75 W, the power consumption of period 6 rises to 110 W, and the power consumption of period 7 falls to 80 W, forming a high-frequency mutation, it is judged that the energy consumption is unstable. The node state of the period is cross-checked according to the consistency of the fluctuation of the waiting time and the mutation of the energy consumption, and if the period overlap rate of the two types of indicators is higher than 60%, it is recorded as an abnormal index resonance. Further, whether the number of times the node is recorded as "high fluctuation processing state" in the entire observation period is greater than 30% of the total number of periods is judged to determine whether the node has processing stability deviation. If the above conditions are met, the state label of the node is updated to "abnormal", and the number of the node is added to the abnormal node sequence. Finally, an abnormal node sequence containing the numbers of all unstable resources is generated.
[0089] S313: Based on the abnormal node sequence, analyze the combined characteristics of the load quantity, task response interval, energy consumption change and frequency fluctuation of each type of node, and use the formula:
[0090]
[0091] Calculate the computing power response fluctuation intensity to determine whether it presents a phased instability feature of response capability, and obtain the computing power running distribution identifier, wherein, GH z represents the computing power response fluctuation intensity of the zth node, n GH represents the total number of computing power nodes, LG z,t represents the task load quantity of the zth node at time t, FG z,t represents the frequency value of the zth node at time t, CG zi,t represents the energy consumption change of the zth node at time t, Gτ z,t represents the task response interval of the zth node at time t, LG z,t-1denotes the number of task loads of the z-th node at the time t-1, FG z,t-1 denotes the frequency value of the z-th node at the time t-1, CG z,t-1 denotes the energy consumption change amount of the z-th node at the time t-1, Gτ z,t-1 denotes the task response interval of the z-th node at the time t-1.
[0092] The computing power running distribution identifier is a numerical index for measuring the dynamic change range of the response capability (such as task processing speed, load condition, frequency change, energy consumption, response interval, etc.) of a heterogeneous computing power resource node to an assigned task over a continuous running period, which can intuitively compare the scheduling stability of different computing power nodes.
[0093] The running behavior of each identified node in a specified scheduling period is structured and arranged, and four types of characteristics, i.e., the number of task loads, the frequency value, the energy consumption change amount, and the task response interval, are merged and processed in sequence according to the period. Taking the Z2 node as an example, assuming that the parameters of the Z2 node in the scheduling period t are the number of task loads 80, the frequency value 2.0 GHz, the energy consumption change amount 150 kJ, and the task response interval 130 ms; the parameters of the Z2 node in the period t-1 are the number of task loads 75, the frequency value 2.2 GHz, the energy consumption change amount 145 kJ, and the task response interval 125 ms, after normalization processing based on the standard interval of the node running index, the normalization results are as follows:
[0094] Period t: LG z,t = 0.80, FG z,t = 0.65, CG zi,t = 0.72, Gτ z,t = 0.81;
[0095] Period t-1: LG z,t-1 = 0.75, FG z,t-1 = 0.71, CG z,t-1 = 0.70, Gτ z,t-1 = 0.78, the above parameters are substituted into the formula, where n GH = 1, and the substitution gives:
[0096]
[0097] The results show that the response behavior of Z2 node in the adjacent two scheduling periods has small fluctuation range, the computing power response ability remains relatively stable under the multi-dimensional characteristics of task load, frequency change, energy consumption fluctuation and response interval, and the numerical result 0.000038 is lower than the identification limit 0.0001 set for the stability of response behavior, so it can be judged that the node is included in the abnormal node sequence, but it does not show persistent or sudden abnormal behavior, which shows that its running state is within the acceptable range. The computing power response fluctuation intensity is a core indicator for representing the stability of node behavior, and the state labels of different computing nodes in the resource pool can be constructed by classifying and summarizing the fluctuation intensities of multiple nodes, and the task mapping strategy and node priority selection logic can be optimized based on the results at the task scheduling level.
[0098] Referring to Figure 5 , the obtaining step of the resource state classification label is specifically:
[0099] S411: Based on the computing power running distribution identifier, analyze its frequency, temperature and queue length parameters, judge whether each parameter meets the normal running standard, select the computing power nodes meeting the standard, and obtain the running node screening set;
[0100] Three types of key running indicator parameters are extracted from each heterogeneous computing power node, including frequency, temperature and current task queue length. The judgment of the frequency value is based on the rated frequency range of the computing node, for example, the standard running frequency range of a certain type of GPU is 1.5GHz to 2.1GHz, and if the measured value of the frequency of a certain node is 1.87GHz, it is marked as within the running standard interval, and if it is lower than 1.5GHz or higher than 2.1GHz, it is marked as abnormal. The judgment of temperature is based on the working temperature control range set by the chip manufacturer of each node, for example, the safe running temperature range of FPGA is set to 40℃ to 80℃, and if the temperature of a certain node is 78.4℃ in the sampling period, it is judged to be normal, and if it is higher than 85℃ or lower than 35℃, it is in an abnormal state. The judgment of the queue length is based on the maximum number of concurrent tasks that the node processing capacity can bear, for example, the theoretical maximum concurrent processing capacity of a certain CPU node is 12 threads, and if the current queue length exceeds 16 tasks, it is judged to be in an overload queuing state. The judgment criteria are as follows: the frequency is within the rated range, the temperature is within the safe interval, and the queue length is lower than the concurrency threshold. The node marked as "running normally" is evaluated in three dimensions. The counter is used to record whether each parameter meets the running standard. If two or more of the three parameters are abnormal, the node is not included in the screening set. If all parameters meet the standard or only one parameter deviates slightly from the set boundary (such as temperature rising to 82℃ but stable), it is still recorded as a running normal node. The node number, frequency value, temperature value and current queue length are packaged into a node state table item. All computing nodes meeting the running standard are screened out according to the screening rule to construct the running node screening set.
[0101] S412: Based on the running node screening set, compare the frequency and temperature of each computing power node, combine the standard running bandwidth and temperature control range, calculate the frequency bandwidth adaptation difference and temperature control adaptation difference, and obtain the node response performance sequence;
[0102] Read the frequency and current running temperature of each node, call the device state interface according to the node type to obtain the frequency stability interval value and temperature floating interval value within every 5 seconds, for example, the frequency fluctuation range of GPU node G07 in the current cycle is 1.84GHz to 1.89GHz, the average value is 1.86GHz, and the temperature range is 73℃ to 75℃, the average value is 74.1℃, then read the standard running bandwidth requirement value and the corresponding temperature control tolerance interval of the current task, for example, a task needs to be kept at a bandwidth lower limit of 1.8GHz, and the temperature should not exceed 76℃, the difference between the node frequency average value and the task lower limit value is 0.06GHz, and the judgment is that the frequency adaptation meets the requirements, and the temperature floating value and the task temperature control upper limit are subtracted, and the temperature control adaptation difference is 1.9℃, the frequency adaptation difference threshold is set to ±0.2GHz, and the temperature control adaptation difference threshold is set to ±3℃, if the difference value is within the threshold range, the node is considered to have good response performance, the two difference values are recorded and combined into a response deviation vector, for example, the frequency difference of node G07 is 0.06, and the temperature difference is 1.9, which is constructed as [0.06, 1.9], execute the above steps on all screened nodes, arrange all node response deviation vectors into a sequence, mark the nodes with response deviation whose frequency difference exceeds ±0.2GHz or temperature difference exceeds ±3℃, remove the remaining node response index value and sort the nodes according to the frequency difference from small to large and the temperature difference from small to large, and sort the node overall response performance, and form the node response performance sequence.
[0103] S413: According to the node response performance sequence, adjust the task concurrency number of the computing power node, analyze the synchronous change of the frequency and temperature, calculate the stress and response state of the node in the task running process, and use the formula:
[0104]
[0105] Determine the belonging category of the computing power node, and obtain the resource state classification label, wherein, RY s represents the computing power response state index, n RY represents the total number of running nodes participating in statistics, MY v represents the task concurrency number of the vth node, TY v represents the current temperature of the vth node, TY b represents the center temperature of the temperature control range, and ΔFY orepresents the change of the main frequency of the oth node, Y o represents the bandwidth matching parameter of the oth node, CY o represents the number of task processing per unit time of the oth node.
[0106] The computing power response state index refers to a comprehensive quantitative result for reflecting the matching degree between the current running state of a group of heterogeneous computing power resource nodes and the task demand after a normalized analysis and calculation of multiple parameters such as task concurrency, temperature, main frequency, bandwidth and processing capacity of the nodes in the process of processing data tasks.
[0107] According to the node response performance sequence, the task concurrency quantity parameter of the screened node is adjusted, and the current task processing queue of each node is sorted and counted. The number of concurrent tasks of node A is 5, the number of concurrent tasks of node B is 6, and the number of concurrent tasks of node C is 4. At the same time, the temperature monitoring records of each node in the latest scheduling period are detected, which are node A: 75℃, node B: 80℃, and node C: 72℃. Based on the center temperature 70℃ of the set temperature control range, the temperature deviation degrees of each node are calculated, which are 5℃, 10℃ and 2℃ respectively. The main frequency change parameters are detected, which are 0.4GHz for node A, 0.3GHz for node B, and 0.5GHz for node C. The bandwidth matching coefficient parameters are collected, which are 0.8, 0.9 and 0.85 respectively. The number of processing tasks per unit time is collected, which is node A: 15, node B: 14, and node C: 13. The normalized operation is performed on the above parameters, and the normalized input variables are as follows:
[0108] MY v : 0.83 for node A, 1.00 for node B, and 0.67 for node C;
[0109] TY v : 0.71 for node A, 0.86 for node B, and 0.57 for node C;
[0110] ΔFY o : 0.80 for node A, 0.60 for node B, and 1.00 for node C;
[0111] Yλ o : 0.70 for node A, 0.90 for node B, and 0.80 for node C;
[0112] CY o : 1.00 for node A, 0.93 for node B, and 0.87 for node C;
[0113] The corresponding variables are substituted into the following formula for calculation:
[0114] wherein the normalized temperature control center temperature TY b = 0.50, and the numerator part is calculated as follows:
[0115] 0.83 · |0.71 - 0.50| + 1.00 · |0.86 - 0.50| + 0.67 · |0.57 - 0.50|;
[0116] = 0.83 · 0.21 + 1.00 · 0.36 + 0.67 · 0.07;
[0117] = 0.1743 + 0.36 + 0.0469 = 0.5812;
[0118] The denominator part is calculated as follows:
[0119]
[0120] The calculation shows that the computing power response state index is:
[0121]
[0122] The result shows that the computing power response state index RY s is 0.2825, which is in the lowest response section of the set classification standard, indicating that the overall temperature control of the current participating node deviates slightly, the task load concurrent adaptation is good, the frequency fluctuation is relatively slow, the bandwidth utilization and the task processing capability are matched well, and the running stability and scheduling adaptability are strong. The smaller the value is, the more stable the node response is, and the closer the running state is to the preset scheduling ideal state. On the contrary, it indicates that the deviation degree rises. Combined with the segmented attribution standard, the numerical interval is divided into grades to generate a differentiated resource state classification label.
[0123] Please refer to Figure 6 , the acquisition steps of the resource dynamic allocation sequence are as follows:
[0124] S511: Based on the resource state classification label, analyze the frequency variation and temperature rise process of each computing power node in the continuous task scheduling period, and according to the frequency and temperature monitoring information at each time point, judge whether the frequency variation of each computing power node is stable and the temperature rise is uniform, and generate the frequency and temperature variation characteristics;
[0125] First, the running monitoring data of each computing node in the continuous scheduling period is extracted, the frequency sampling value and temperature monitoring value of each period are read, and the sequence structure is formed in time sequence, for example, the frequency record of node C07 in 10 scheduling periods is [2.10, 2.11, 2.09, 2.10, 2.10, 2.10, 2.11, 2.10, 2.09, 2.10] GHz, and the temperature record is [64.1, 64.3, 64.6, 64.8, 65.0, 65.3, 65.4, 65.7, 65.9, 66.2]℃, then the inter-period difference value of the frequency sequence is calculated, and it is judged whether the frequency fluctuation between adjacent two periods exceeds the set frequency change threshold, the set frequency stability judgment threshold is 0.05GHz, that is, if the frequency difference value of any adjacent two periods does not exceed 0.05GHz, it is judged to be stable, for example, the maximum fluctuation amplitude of the above frequency sequence is 0.02GHz, which meets the stability standard, then the temperature rise value between any two periods in the temperature sequence is read, and whether the change gradient is within the set temperature rise uniformity standard range is compared, the set temperature rise uniformity judgment threshold is that the single period does not exceed 0.4℃, if it exceeds, it is judged that the temperature rise is not uniform, in the above temperature sequence, the maximum single period increase is 0.3℃, and the overall fluctuation is within the threshold, it is judged that the temperature change is uniform, the same judgment process is performed on all nodes, whether the frequency is stable and whether the temperature is uniform is marked respectively, and the two judgment results are recorded and integrated into node feature label, for example, node C07 is marked as "frequency stable, temperature rise uniform", nodes that do not meet any one of the two conditions are marked as "unstable" or "temperature rise too fast", the state evaluation records of all nodes are summarized, and the frequency temperature change feature is generated.
[0126] S512: Based on the frequency temperature change feature, the computing nodes with stable frequency and uniform temperature change are screened, the high priority tasks are compared with the screened nodes in combination according to the task emergency level, the scheduling matching between the nodes and the tasks is judged, and the task node adaptive combination is obtained.
[0127] The node number list that meets the conditions of stable frequency and uniform temperature rise is screened out, and is recorded as a stable node set. The latest frequency and temperature average value of each node are read, and are compared with the task set marked as high priority in the task scheduling system. The tasks with emergency level of 4 or 5 are selected as high priority objects from the task set. The required frequency lower limit value and temperature tolerance upper limit value in the task metadata of the tasks are extracted and compared. For example, the frequency requirement of task A07 is ≥2.0 GHz, and the temperature control tolerance upper limit is 70℃. If the frequency of node D09 is 2.08 GHz and the temperature is 65.5℃, it is judged that the node meets the scheduling adaptation conditions of the task. If the frequency of the node is lower than the required frequency lower limit of the task or the temperature is higher than the allowed temperature upper limit of the task, it is judged that the adaptation is not suitable. The judgment process is set as: task frequency requirement-node frequency value ≤0.1 GHz and node temperature ≤task temperature control upper limit. Only when the above conditions are met, the adaptation is recorded as successful. All tasks and nodes that meet the adaptation conditions are paired to form a structured combination item. The recorded content includes task number, emergency level, node number, frequency comparison value, and temperature comparison value, such as {A07, level 5, D09, difference 0.08, difference-4.5}. For each high priority task in the stable node set, find the adaptation node. If there are multiple nodes that meet the conditions, select the node with the smallest frequency difference and the smallest temperature difference as the priority matching node. After completing the task node adaptation record, all matching pairs are summarized and arranged to generate a task node adaptation combination.
[0128] S513: Based on the task node adaptation combination, the allocation relationship between the tasks and the computing power nodes is adjusted. The running state grouping of each node and the task distribution condition are analyzed. The task scheduling sequence and the computing power node bearing distribution are optimized. The resource dynamic allocation sequence is obtained.
[0129] First, read the current task scheduling record table all the task allocation node state, the new matching combination and the original scheduling relationship comparison, judge whether there is a task need to be rescheduled or node need to be reconfigured situation, for example, task B02 original binding on node E11 but the new adaptive combination is E15, then the task state is updated to "need to adjust", then read the node running state grouping number and the existing task quantity, build the current task distribution atlas, calculate the total number of tasks carried by each node, scheduling interval time, average response time and other indicators, then select all "need to be redistributed" tasks from the task adaptive combination, preferentially match high-level tasks to the current empty or low load nodes, define the task scheduling order priority: task emergency level is higher than 4, the current unbound node is preferred, the current load task quantity is less than the average value of the node, read the update time point of each task bound resource when executing the scheduling rearrangement process, preferentially select the task whose allocation lag time exceeds 3 times the scheduling period as the scheduling adjustment candidate, complete the final mapping according to the node response performance and task priority, update the scheduling table to generate the new allocation structure record of the current task and node, finally integrate all the completed matching adjustment task and node binding relationship, form the resource dynamic allocation sequence optimized according to the task order and node load distribution.
[0130] The above is only the preferred embodiment of the present application, not other forms of the present application, any skilled in the art may use the above disclosed technical content to change or modify the equivalent embodiment applied to other fields, but any simple modification, equivalent change and modification of the above embodiment without departing from the technical solution content of the present application, according to the technical essence of the present application, still belongs to the protection scope of the technical solution of the present application.
Claims
1. A method for dynamic allocation of heterogeneous computing power based on big data analysis, characterized in that: The following steps are involved: S1: Based on the initial stage of data tasks, analyze the source and target of data tasks, map the source type and target path into characteristic parameters according to unified standards, compare task types and dependency attributes, classify urgency levels, combine the performance of heterogeneous computing resources, and calculate the maximum and minimum terms and balance indicators of the parameters to obtain normalized characteristic data; S2: Based on the normalized feature data, determine the changes in the balance index and the business scenario reference data, analyze the differences between the two, identify key tasks based on the adaptability of computing resources, adjust the task scheduling status, and obtain an adapted task index set; S3: Based on the adapted task index set, the number of task queues of computing resources is calculated, the switching efficiency is compared with the historical scheduling switching situation, the task processing capacity is determined and the energy consumption is analyzed, the idle window and abnormal state identification are optimized, and the computing power operation distribution identifier is obtained; S4: Based on the computing power operation distribution identifier, filter the computing power resources that are operating normally, analyze the main frequency and temperature data, compare the bandwidth and temperature control range, calculate the response performance, determine the relationship between the number of concurrent tasks and the operation status, adjust the operation level, and obtain the resource status classification label.
2. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The normalized feature data includes task label dimensions, parameter normalization information, and feature aggregation results; the adapted task index set includes task priority index, allocation mark information, and scheduling adaptation type; the computing power operation distribution identifier includes resource distribution information, operation label characteristics, and node abnormality identifier; the resource status classification label includes performance interval label, load status type, and stability group number.
3. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the normalized feature data are specifically as follows: S111: Based on the initial stage of the data task, the source type information and target path data of the data task are analyzed, the structural difference between the source and the target is calculated, and the consistency of the task parameter mapping is determined by comparing the task type and dependency relationship to obtain the source-target structural deviation; S112: Calculate the urgency level label of each task based on the source target structure deviation, compare the matching relationship between the urgency level and the computing resource performance label, select tasks that meet the characteristic requirements of each group, and generate an urgency level resource adaptation group; S113: Based on the emergency level resource adaptation group, calculate the normalized performance of each group of tasks, determine the balance relationship between the tasks, screen the key performance of each feature dimension, and obtain normalized feature data.
4. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the adaptation task index set are specifically as follows: S211: Based on the normalized feature data, analyze its balance performance with the business scenario reference data, combine the task labels, parameter normalization content and feature aggregation content, compare the type and dependent parameters of each task, determine the offset of the task in the feature distribution, and generate the task feature offset; S212: Based on the task feature offset, the performance labels of each node in the heterogeneous computing resource pool are compared to analyze the computing power, response capability, and energy consumption characteristics of the node. By matching the computing requirements of the task, the adaptation between each task and the resource node is screened to obtain the task resource adaptation difference; S213: Based on the task resource adaptation difference, adjust the scheduling relationship between the task and the heterogeneous computing resources, determine the matching performance of each task in the current resource scheduling, analyze the tasks that need to be adjusted first in the scheduling state, and obtain the adaptation task index set.
5. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the computing power operation distribution identifier are as follows: S311: Based on the adapted task index set, calculate the change in the number of tasks queued for each heterogeneous computing power node in a continuous scheduling cycle, analyze the task type distribution and scheduling completion status, determine the task processing status of the computing power node in the same cycle, identify the behavior pattern of inconsistent processing speed, and obtain a processing capacity fluctuation sequence; S312: Based on the processing capacity fluctuation sequence, select computing nodes with concentrated processing fluctuation resources, calculate the changing trends of task waiting time and energy consumption records of the nodes, determine whether task response delays occur repeatedly, determine stability characteristics, and obtain an abnormal node sequence; S313: Based on the abnormal node sequence, analyze the combined characteristics of the load quantity, task response interval, energy consumption change and main frequency fluctuation of each type of node, calculate the computing power response fluctuation intensity, determine whether the response capability presents a periodic unstable characteristic, and obtain the computing power operation distribution identifier.
6. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The steps for obtaining the resource status classification label are specifically as follows: S411: Based on the computing power operation distribution identifier, analyze the main frequency, temperature, and queue length parameters, determine whether each parameter meets the normal operation standard, select computing power nodes that meet the standard, and obtain a screening set of operating nodes; S412: Based on the screening set of running nodes, compare the main frequency and temperature of each computing power node, combine the standard operating bandwidth and temperature control range, calculate the main frequency-bandwidth adaptation difference and the temperature control adaptation difference, and obtain a node response performance sequence; S413: According to the node response performance sequence, adjust the number of concurrent tasks of the computing power node, analyze the synchronous changes of the main frequency and temperature, calculate the pressure and response status of the node during the task operation, determine the category of the computing power node, and obtain the resource status classification label.
7. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 1 is characterized in that: The steps also include: S5: Based on the resource status classification labels, analyze the frequency change trend and temperature increase rate of the computing resources, determine the task urgency level and priority, identify computing resources with stable frequency and gentle temperature change, optimize the allocation relationship, and obtain a dynamic resource allocation sequence; The dynamic resource allocation sequence includes a task allocation structure, a computing power mapping relationship, and a priority allocation sequence.
8. The method for dynamic allocation of heterogeneous computing power based on big data analysis according to claim 7 is characterized in that: The steps for obtaining the resource dynamic allocation sequence are specifically as follows: S511: Based on the resource status classification label, analyze the main frequency change and temperature rise process of each computing power node during the continuous task scheduling cycle, and determine whether the main frequency change of each computing power node is stable and whether the temperature rise is uniform based on the main frequency and temperature monitoring information at each time point, and generate a main frequency and temperature change feature; S512: Based on the main frequency and temperature variation characteristics, computing nodes with stable main frequencies and uniform temperature variations are selected. Based on the task urgency level, high-priority tasks are compared with the selected nodes for adaptation, scheduling compatibility between the nodes and tasks is determined, and an adaptation combination of task nodes is obtained. S513: Based on the task node adaptation combination, adjust the distribution relationship between tasks and computing power nodes, analyze the operation status grouping and task distribution status of each node, optimize the task scheduling sequence and computing power node load distribution, and obtain a dynamic resource allocation sequence.
Citation Information
Cited By
Computing power server-oriented full-stack resource big data analysis dynamic tuning method
CN121411952A