A GPU computing power dynamic scheduling system

By monitoring the thermal expansion and path disturbance of the GPU core in real time, and optimizing the resource scheduling strategy, the problem of unbalanced resource allocation in the existing technology is solved, and load balancing between GPU computing units and system execution efficiency is improved.

CN120179045BActive Publication Date: 2025-07-18CHENGDU WANDA ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510646898.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-07-18
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing resource scheduling technologies cannot monitor the thermal expansion and path disturbance of GPU computing resources in a high concurrency environment in real time, resulting in task delays and calculation performance degradation, unbalanced resource allocation, and affecting the overall performance of the system.

Method used

The thermal expansion monitoring module, path disturbance identification module, load migration module, scheduling label revision module and node priority module are adopted to monitor the thermal expansion and path disturbance of the GPU core in real time, and generate thermally coupled structures of thermal path coverage marks and instruction disturbances, optimize resource scheduling strategies, and achieve dynamic load balancing.

Benefits of technology

It improves the GPU resource utilization rate and task execution stability, enhances the resource scheduling accuracy and overall system execution efficiency in multi-task concurrency environment, and avoids excessive competition and idle resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179045B_ABST
    Figure CN120179045B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of resource scheduling, and specifically to a dynamic GPU computing power scheduling system. The system includes a thermal expansion monitoring module, a path perturbation identification module, a load migration module, a scheduling label revision module, and a node priority ranking module. In the present invention, through the real-time monitoring of the thermal expansion response of the GPU core and the instruction path behavior, it effectively avoids the performance bottleneck caused by overheating in the task execution area, ensures better coordination of heat conduction and computing load during resource scheduling, and combines the dynamic adjustment of the thermal response time and the path compression scheduling label. It can perform dynamic optimization scheduling according to the changes in task load and resource utilization rate, enhance the load balance between GPU computing units, avoid excessive competition and idle of resources, improve the overall execution efficiency and response speed of the system, and through the technology of optimizing resource scheduling by analyzing thermal behavior and path perturbation, it effectively improves the accuracy and stability of GPU computing power scheduling in a multi-task concurrent execution environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource scheduling, and in particular, to a dynamic scheduling system for GPU computing power. Background Art

[0002] The technical field of resource scheduling includes the dynamic management and reasonable allocation mechanisms of computing resources in various hardware and software environments. The core content of this technical field is to face the multi-task concurrent execution environment for multi-type computing resources such as general-purpose processors, graphics processors, application-specific integrated circuits, field-programmable gate arrays and other hardware resources. By perceiving the resource status, analyzing the task load and formulating a scheduling strategy, the resources can be efficiently matched and distributed in multiple dimensions such as computing density, task priority, and load balancing. This technical field also includes cross-platform and cross-architecture resource integration scheduling algorithm design, scheduling rule formulation, and task mapping mechanisms in heterogeneous resource environments, systematically building a connection bridge between computing resources and business requirements, and realizing the dynamic overall planning and refined management of resources in the space-time dimension.

[0003] Among them, the dynamic scheduling system for GPU computing power refers to the resource management problem of the graphics processor in a multi-task scenario. By constructing a task recognition method to monitor the load characteristics of the current GPU task in real time and comparing the task characteristics with the GPU resource status, using a scheduling priority determination method to judge the scheduling order of each GPU core, and then combining the task type classification method to map different types of tasks to the appropriate GPU computing units, the orderly distribution of GPU resources and the coordination of computing power load between nodes are realized. This system uses task analysis technology, resource status perception method, and scheduling rule matching mechanism to complete the distribution process of computing power tasks.

[0004] The existing resource scheduling technologies mainly rely on the static comparison of task load and resource status for scheduling decisions, ignoring the dynamic changes in GPU computing resources caused by thermal expansion and path disturbances during the actual execution process. This results in the phenomenon of task delay or calculation performance degradation due to excessive temperature or large path execution fluctuations in a high-concurrency environment. For example, when multiple tasks are executed on the same GPU core simultaneously, if the heat conduction coverage and path disturbances are not monitored and dynamically adjusted in real time, some tasks will not be able to be completed in time due to insufficient computing resources, affecting the execution efficiency of the entire system. The task scheduling mechanism in the prior art cannot accurately perceive the specific requirements of different task types for GPU resources, resulting in uneven resource allocation and reduced overall system performance in complex task combinations and multi-task environments. This scheduling mechanism based on fixed rules and load balancing has limitations in dealing with complex resource allocation requirements under heterogeneous computing platforms and cannot fully adapt to the changing workload and dynamic changes in computing resources. Summary of the Invention

[0005] The object of the present invention is to solve the deficiencies existing in the prior art, and a GPU computing power dynamic scheduling system is proposed.

[0006] To achieve the above object, the present invention adopts the following technical solutions. A GPU computing power dynamic scheduling system includes:

[0007] The thermal expansion monitoring module obtains the time series record of the GPU core thermal sensor and the motherboard thermal expansion response trajectory, calls the thermal expansion speed, temperature rise duration and diffusion direction path number in the core running state, identifies whether there is information covering the task instruction execution area in the thermal conduction during the scheduling period, and generates thermal path coverage mark information;

[0008] Based on the thermal path coverage mark information, the path perturbation identification module retrieves whether the path numbers are continuous and extracts the positions where path jump behaviors occur. If the path numbers are not continuous in two adjacent cycles and the memory access order is interrupted, the marked segment is a perturbed behavior path, and an instruction perturbation thermal coupling linkage structure is generated;

[0009] The load migration module calls the cache continuous read and write cycles between the GPU core located by the instruction perturbation thermal coupling linkage structure and the adjacent cores, retrieves whether there are path jump stable sequences and cache hit cycle stable nodes in the adjacent cores, and generates a core migration path mapping trajectory set;

[0010] The scheduling label revision module uses the core migration path mapping trajectory set, compares the thermal response time change behavior and the execution order change trajectory, and screens the continuous path structure and the record of shortened cache after migration to generate a path compression scheduling label.

[0011] As a further solution of the present invention, the thermal path coverage mark information includes a thermal expansion path sequence number, a starting boundary of the thermal conduction coverage area, and a task path overlap index point. The instruction perturbation thermal coupling linkage structure includes a perturbed path number set, a thermal area coupling mapping index, and an in-cycle jump synchronization mark. The core migration path mapping trajectory set includes a target core path jump sequence, a cross-core path delay trajectory line, and a path mapping section. The path compression scheduling label includes a path compression identification number, a cache usage fallback section, and a sequential behavior continuity mark.

[0012] As a further solution of the present invention, the thermal expansion monitoring module includes:

[0013] The time series extraction sub-module obtains the time series record of the GPU core thermal sensor and the motherboard thermal expansion response trajectory, detects the core running state parameters and the motherboard thermal expansion response values at multiple moments, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal sensor value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value;

[0014] The diffusion path recognition sub-module, based on the temperature rise duration sequence values, calls the thermal diffusion response trajectory numbers and the core working state paths within the corresponding time period, calculates the unit time thermal sensitivity change value and the diffusion range value within each trajectory, and filters according to the difference between the diffusion value and the path number to obtain the set of thermal diffusion direction path numbers;

[0015] The hot path coverage marking sub-module matches the corresponding execution area of the task instruction within the task scheduling period according to the set of thermal diffusion direction path numbers, determines whether there is a spatial overlap between the thermal diffusion propagation area in the path number set and the instruction execution position, calculates the hot path conduction matching degree, analyzes the spatial position index and the instruction number corresponding to the marked coverage area, and generates the hot path coverage marking information.

[0016] As a further solution of the present invention, the path perturbation recognition module includes:

[0017] The path number extraction sub-module, based on the hot path coverage marking information, obtains the path numbers of the GPU core within the task running period, the start and end timestamps of the instruction segments, the access memory instruction order, and the cache hit time record, extracts the path numbers within two adjacent periods, sorts the path numbers, and retrieves the continuity of the number sequence, filters the path number segments that do not meet the number continuity determination condition, and generates the path number interruption interval value;

[0018] The access memory order retrieval sub-module, according to the path number interruption interval value, calls the access memory instruction order and the cache hit time record, extracts the access memory segments corresponding to the path interruption, compares the access memory order in the segment with the access memory order of the same period path, calculates the access memory order interruption intensity value, and obtains the set of perturbation behavior path segments;

[0019] The thermal coupling linkage structure construction sub-module calls the set of perturbation behavior path segments, extracts the path jump positions and the hot path segment numbers, performs position mapping and coincidence interval judgment on the jump positions and the hot path numbers, and generates the instruction perturbation thermal coupling linkage structure.

[0020] As a further solution of the present invention, the load migration module includes:

[0021] The perturbation detection sub-module, based on the instruction perturbation thermal coupling linkage structure, extracts the cache continuous read and write cycles, the cache latency behavior, and the set of path jump points between the GPU core and the adjacent cores, monitors whether the cache continuous read and write cycles in the set of GPU core path jump points are in the cache latency behavior change interval, filters the path jump points contained in the adjacent cores associated with the change interval, and generates the cache latency perturbation intensity measure;

[0022] The path screening sub-module calls the cache latency perturbation metric to determine whether there are path jump stable sequences and cache hit cycle stable nodes in the neighboring cores. If the node hit cycle fluctuation is within the hit cycle jump constraint interval, the neighboring cores are extracted as candidate path cores, and a set of path stability values is generated.

[0023] The core mapping sub-module constructs the migration paths between cores and maps the node jump trajectory path sequences according to the set of path stability values, calculates the migration trajectory offset intensity value, and evaluates the trajectory mapping continuity of the candidate cores to obtain the core migration path mapping trajectory set.

[0024] As a further solution of the present invention, the scheduling label revision module includes:

[0025] The path trajectory extraction sub-module extracts the start and end nodes of cross-core migration, the number of path jumps, and the migration frequency value of the execution path based on the execution instruction records and the corresponding core numbers in the core migration path mapping trajectory set. Through joint comparison of the migration frequency value and the number of jumps, a cross-core path jump trajectory is generated.

[0026] The cache structure screening sub-module calls the consecutive path segment number sequence and path interruption positions in the cross-core path jump trajectory, and generates a cache hit fluctuation interval by statistically analyzing the cache read times and cache hit records of the instruction sets before and after the path segments and comparing the change rate of the read times with the hit record fluctuation interval.

[0027] The execution sequence comparison sub-module extracts the sequence structure numbers of the instructions before and after scheduling under the same path based on the path set corresponding to the cache hit fluctuation interval, compares the execution cycle values and start time offset values of the execution instructions, and screens out the paths with abnormal offset trends to obtain the execution sequence offset fluctuation degree.

[0028] The path compression sub-module calls the abnormal path numbers and offset direction attributes in the execution sequence offset fluctuation degree, adjusts the sorting order of the path segments according to the offset direction attributes, and compresses the length values of the consecutive jump path segments to obtain the path compression scheduling label.

[0029] As a further solution of the present invention, the system further includes a node priority sorting module:

[0030] The node priority sorting module calls the path compression scheduling label, extracts the path compression node numbers, cache behavior records, and jump sequence stable times, identifies the number of stable path formations and the behavior trend intervals, sorts and schedules the nodes according to the path structure coherence, and generates a GPU stable execution path sorted list.

[0031] The GPU stable execution path sorted list includes sorted node numbers, behavior stability sequence values, and scheduling priority index codes.

[0032] As a further solution of the present invention, the node priority sorting module includes:

[0033] Based on the path compression scheduling label, the path compression extraction sub-module extracts the path compression node number, the cache behavior record corresponding to the node, and the stable time of the jump sequence, establishes a pairing relationship between the stable time of the jump sequence and the path compression node number according to the node order, determines whether the stable time of the jump sequence shows a continuous increasing trend and marks the corresponding numbered paragraph, and generates a compressed node number interval sequence;

[0034] The behavior trend division sub-module calls the compressed node number interval sequence, analyzes the cache behavior record corresponding to the node number, divides according to the synchronous fluctuation characteristics of the hit frequency, write ratio, and replacement rate in the time dimension in the cache behavior record, evaluates the numbered attribution relationship of multiple types of fluctuation trends in the path segment, and generates a behavior change trend interval;

[0035] The node sorting sub-module screens the node numbers with stable jump sequence characteristics according to the behavior change trend interval, judges the difference direction of the hit frequency and the replacement rate in the cache behavior record corresponding to the node number, selects the node numbers with the hit frequency exceeding the replacement rate, and arranges them in order according to the original path position of the node to generate a sorted coherent node sequence;

[0036] The execution path adjustment sub-module calls the sorted coherent node sequence, combines the adjacent structure between the corresponding stable time of the jump sequence and the original node number sequence, judges whether there is a node coherence gap in the path, retains the path segment numbers with a complete coherent structure, eliminates the paragraph numbers of the nodes with abnormal jumps, and integrates the consecutive numbers to form a stable path node combination to obtain the GPU stable execution path sorting list.

[0037] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0038] In the present invention, through real-time monitoring of the thermal expansion response of the GPU core and the behavior of the instruction path, during the task execution process, it is possible to accurately identify the thermal conduction coverage and path perturbation behaviors, effectively avoid performance bottlenecks caused by overheating in the task execution area, and ensure better coordination of thermal conduction and computing load during resource scheduling. By capturing the path jump points and stable nodes of the cache behavior, the generated path migration mapping trajectory provides a more refined resource scheduling reference for task allocation, effectively improving the utilization rate of GPU resources and the stability of task execution. Combining the thermal response time with the dynamic adjustment of the path compression scheduling label, it is possible to perform dynamic optimization scheduling according to changes in task load and resource utilization rate, enhance the load balancing between GPU computing units, avoid excessive competition and idleness of resources, and improve the overall execution efficiency and response speed of the system. The technology of optimizing resource scheduling through thermal behavior and path perturbation analysis effectively improves the accuracy and stability of GPU computing power scheduling in a multi-task concurrent execution environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the system flow chart of the present invention;

[0040] Figure 2 is the flow chart of the thermal expansion monitoring module in the present invention;

[0041] Figure 3 is the flow chart of the path perturbation identification module in the present invention;

[0042] Figure 4 is the flow chart of the load migration module in the present invention;

[0043] Figure 5 is the flow chart of the scheduling label revision module in the present invention;

[0044] Figure 6 is the flow chart of the node priority sorting module in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0045] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more, unless otherwise specifically defined.

[0047] Please refer to Figure 1 , a GPU computing power dynamic scheduling system includes:

[0048] The thermal expansion monitoring module obtains the time series records of the GPU core thermal sensors and the motherboard thermal expansion response trajectories, calls the thermal expansion speed, temperature rise duration, and diffusion direction path number in the core running state, identifies whether there is information covering the task instruction execution area in the heat conduction during the scheduling period, and generates heat path coverage marking information;

[0049] Based on the heat path coverage marking information, the path perturbation identification module obtains the path number of the GPU core, the start and end timestamps of the instruction segments, the memory access instruction order, and the cache hit time records during the task running cycle, retrieves whether the path numbers are continuous and extracts the positions where path jump behaviors occur. If the path numbers are discontinuous in two adjacent cycles and the memory access order is interrupted, the segment is marked as a perturbed behavior path, and it is judged whether there is an overlap between the path perturbation behavior and the heat path area, generating an instruction perturbation thermal coupling linkage structure;

[0050] The load migration module calls the cache continuous read and write cycles, cache latency behaviors, and path jump point sets between the GPU core located by the instruction perturbation thermal coupling linkage structure and the adjacent cores, retrieves whether there are stable path jump sequences and stable cache hit cycle nodes in the adjacent cores. If it is determined to be a stable core, it is set as the path migration target, generating a core migration path mapping trajectory set;

[0051] The scheduling label revision module uses the cross-core execution path behavior sequence formed by the core migration path mapping trajectory set, compares the heat response time change behavior and the execution order change trajectory, filters the continuous path structure and the cache shortening record after migration, and generates a path compression scheduling label;

[0052] The node priority sorting module calls the path compression scheduling label, extracts the path compression node numbers, cache behavior records, and jump sequence stable times, identifies the number of stable path formations and the behavior trend intervals, sorts the scheduling nodes according to the path structure coherence, and generates a GPU stable execution path sorted list;

[0053] The hot path coverage marking information includes the hot expansion path sequence number, the starting boundary of the heat conduction coverage area, and the task path overlap index point. The instruction perturbation thermal coupling linkage structure includes the perturbation path number set, the thermal zone coupling mapping index, and the in-cycle jump synchronization mark. The core migration path mapping trajectory set includes the target core path jump sequence, the cross-core path delay trajectory line, and the path mapping section. The path compression scheduling label includes the path compression identification number, the cache usage fallback section, and the sequential behavior continuity identification. The GPU stable execution path sorted list includes the sorting node number, the behavior stability sequence value, and the scheduling priority index code.

[0054] Please refer to Figure 2 , the hot expansion monitoring module includes:

[0055] The time series extraction sub-module obtains the GPU core thermal sensor time series record and the motherboard hot expansion response trajectory, detects the core operation state parameters and the motherboard hot expansion response values at multiple moments, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal sensitive value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value;

[0056] Multiple highly sensitive thermal sensor nodes need to be arranged on the GPU chip structure. The sensor nodes continuously sample at intervals of 100 ms, record the temperature readings of each sensor at the current moment, and set four thermal points P1 to P4 within a certain detection period. The data they read are 72.3 °C, 74.1 °C, 75.5 °C, and 73.0 °C in sequence. After the acquisition is completed, the four groups of data are numbered according to the time sequence and classified into the time series structure. At the same time, the core operating state parameters are synchronously collected among the obtained thermal data points, such as the current load value, voltage change rate, instruction throughput rate, etc. It is set that the core current is 3.6 A, the voltage change rate is 0.15 V / s, and the instruction throughput rate is 1800 instructions / s during this period. The parameters and the thermal data are combined into a complete state data vector for this period. After the data set is constructed, it is necessary to take the continuous temperature rise stage within the detection period as the screening target, extract the data segment with continuous thermal change and rise, and set it between 100 ms and 300 ms. The temperature of the P2 sensor continuously rises from 74.1 °C to 78.0 °C, then this time period is identified as the temperature rise stage, and it is necessary to calculate the temperature rise value increase and the time period length. The temperature rise amplitude of this segment is 3.9 °C, and the time span is 200 ms. This data is numbered as the "temperature rise duration data point" and bound to the core operating state intensity of this period to form the "temperature rise duration sequence value". During this process, it is also necessary to judge whether the temperature rise change rate exceeds the set thermal change threshold. The set thermal change threshold is 15 °C / s, and the temperature rise change rate of this segment is 3.9 °C / 0.2 s = 19.5 °C / s, so it is determined that it meets the trigger requirement and is recorded as a valid temperature rise duration data point, completing the integration process of the time series data and obtaining the temperature rise duration sequence value.

[0057] Based on the temperature rise duration sequence value, the diffusion path identification sub-module calls the thermal diffusion response trajectory number and the core working state path within the corresponding period, calculates the thermal change value per unit time and the diffusion range value within each trajectory, and screens according to the difference between the diffusion value and the path number to obtain the set of thermal diffusion direction path numbers;

[0058] Locate the spatial diffusion path number it is in, call the set of recorded main board thermal expansion response trajectory numbers, and set that there are three trajectories numbered R1, R2, and R3 in the previous cycle. Among them, R1 corresponds to the path from the GPU core to the main control chip, R2 corresponds to the path from the GPU core to the south bridge chip, and R3 corresponds to the path from the GPU core to the memory module. For each trajectory path, extract the thermal change value per unit time and the diffusion range value between points on the path. Set that the temperature rises by 1.8 °C per unit time in path R1, and the diffusion range value is 7.5 mm². The corresponding values in path R2 are 2.1 °C and 8.3 mm², and in path R3 are 1.2 °C and 5.9 mm². Calculate the degree of difference between the thermal expansion value and the path number. Through the calculation of the normalized response rate of the path number, classify the path numbers with a normalized value exceeding 0.6 into the "effective diffusion path" set. Set the normalized response rate values of R2 and R3 to 0.71 and 0.68, both of which exceed the threshold, while R1 is 0.53 and is excluded. Only retain paths R2 and R3 as the set of effective diffusion path sequence numbers. On this basis, establish a one-to-one correspondence between the temperature rise duration sequence value and the path number, and combine the core operation state intensity to evaluate the intensity level of the thermal expansion response on each path. Set the operation state intensity bound to path R2 to 0.88 and that on path R3 to 0.92. Then both paths meet the requirement that the thermal expansion response rate reaches the path thermal expansion response rate threshold of 0.85, and are marked as the propagable path numbers, completing the identification process of the thermal expansion path and obtaining the set of thermal expansion direction path numbers.

[0059] According to the set of thermal expansion direction path numbers, the thermal path coverage marking sub-module matches the execution area corresponding to the task instruction within the task scheduling period, and judges whether the thermal expansion propagation area in the path number set overlaps with the instruction execution position in space, using the formula:

[0060] ;

[0061] Calculate the thermal path conduction matching degree, analyze the spatial position index and instruction number corresponding to the marked coverage area, and generate thermal path coverage marking information;

[0062] Among them, is the thermal path conduction matching degree, represents the duration of the path number in the task instruction, represents the spatial displacement difference corresponding to the thermal expansion path number , represents the path number corresponding thermal expansion response rate, represents the path number core load gain value within the detection period, is the number of path numbers within the detection period;

[0063] Table 1 Task Thermal Expansion Parameter Table:

[0064] ;

[0065] As shown in Table 1, the path parameters collected for each task number are the basic participation items of the thermal path conduction matching degree value. By quantifying the basic thermophysical characteristic values, a quantitative index reflecting the coincidence degree of the thermal path and the instruction path can be calculated, and it can guide whether to execute the generation of marker coverage information;

[0066] It is necessary to clarify the specific numbers and task durations of each task in the scheduling period, and set the numbers , and The corresponding durations are 5.2 seconds, 7.1 seconds, and 4.8 seconds respectively. It is necessary to quantify the path displacement difference corresponding to each task . The spatial displacement differences obtained through the thermal expansion path detection results are 1.5 mm, 2.1 mm, and 1.2 mm. At the same time, it is necessary to collect the thermal expansion response rate of this path during the monitoring period and the core load gain . The obtained rate values are 0.8 °C / s, 1.0 °C / s, 0.7 °C / s, and the load gains are 0.3, 0.4, 0.2. Calculate the task propagation error term under each path number, and use the absolute value form of each item to express the difference between the product of the task duration and the thermal path displacement difference and the square root of the sum of the thermal expansion rate squared and the load gain, that is:

[0067] ;

[0068] Calculate separately as: ;

[0069] as: ;

[0070] as: ;

[0071] Sum up the above difference results to obtain the numerator term:

[0072] ;

[0073] Directly sum up the time, path thermal expansion rate, and load gain of each task, and the denominator term of each item is obtained as : ;

[0074] : ;

[0075] : ;

[0076] The denominator after summation is:

[0077] ;

[0078] Substitute into the formula:

[0079] ;

[0080] This result indicates that the average matching error intensity of the heat path conduction in the task instruction path is 1.243, which is used to evaluate the spatial coincidence degree between the execution position of the current scheduling task and the predicted heat propagation path, reflect whether the task is in the heat risk area, guide the task to avoid the hot spot area, improve the heat balance and operation reliability. If the determination threshold of the heat path conduction matching degree is set to 1.1, since 1.243 > 1.1, it shows that there is a coverage relationship between the task path and the heat expansion path, and spatial position indexes and identifications should be added to the instruction path to generate heat path coverage marking information.

[0081] Please refer to Figure 3 , the path perturbation identification module includes:

[0082] Based on the heat path coverage marking information, the path number extraction sub-module obtains the path numbers of the GPU cores during the task operation cycle, the start and end timestamps of the instruction segments, the memory access instruction order, and the cache hit time records, extracts the path numbers in two adjacent cycles, sorts the path numbers, and retrieves the continuity of the number sequence, filters out the path number segments that do not meet the number continuity determination conditions, and generates path number interruption interval values;

[0083] Retrieve the GPU scheduling log, parse the allocation records of the thread blocks in the scheduling information, extract the path number information of each instruction executed in the GPU core in each cycle, and set the task transfer path numbers in cycle as P123 and P124, and in cycle as P125 and P126. Each path number corresponds to a certain number of instruction segments, and their start and end timestamps can be marked by accessing the start and end time information of the instruction scheduling in the access log. For example, if an instruction segment starts at 10.02ms and ends at 10.31ms, it can be recorded as [10.02, 10.31]. Extract the memory access instruction execution order from the memory access stack record. For example, the access order in this cycle is instruction 10, instruction 12, instruction 15, and read the cache hit time together with the cache status record. Set the cache hit time of access instruction 12 to 1.3ms, and sort the extracted path number sequence in time order, such as During the period, it is recorded as P123 → P124. If the starting number of the period is P126, then the broken number P125 in the middle is the interruption number, indicating that there is a path number jump behavior during the period switch, providing a data basis for subsequent judgment of whether the path is continuous. On this basis, all path number sequences are compared pairwise to determine whether they are continuous in number, that is, whether there are situations such as missing numbers and number jumps. For example, if the number sequence should be P123 → P124 → P125 → P126, and the actual measurement result is P123 → P124 → P126, then it is judged that the P125 path is missing, regarded as a path discontinuity event. In this case, the corresponding path segment is marked as the path number interruption interval, the set of path segments with path number breaks is marked, and recorded as the path number interruption interval value.

[0084] The memory access order retrieval sub-module calls the memory access instruction order and cache hit time record according to the path number interruption interval value, extracts the memory access fragment corresponding to the path interruption, and compares the memory access order in the fragment with the memory access order of the same-period path, using the formula:

[0085] ;

[0086] Calculate the memory access order interruption intensity value to obtain the path fragment set of disturbance behaviors;

[0087] Among them, represents the memory access order interruption intensity value, represents the position number of the th memory access instruction in the disturbance fragment, represents the position number of the th instruction in the same-period comparison fragment, represents the corresponding th cache hit latency value, represents the th path cache hit count, represents the number of memory access instruction entries in the disturbance fragment, represents the number of cache hit record segments in the corresponding path;

[0088] Identify whether there is a break in the sequence of path numbers within two consecutive cycles. If the ending path number in cycle C1 is 25 and the starting path number in cycle C2 is 28, it is regarded as an interruption interval of [25, 28]. Further locate the memory access segments between path numbers 25 and 28, arrange the memory access instructions in them in chronological order, and respectively extract the actual execution position numbers of the instructions in their perturbed path numbers and the comparison position numbers in the normal path numbers. For example, if the position of the memory access instruction numbered 2 in the perturbed path is 25 and the position of this type of instruction in the comparison path is 22, then the sequential offset between the two is 3. After that, extract the offset values of the memory access instructions in the interruption segment, and calculate the total sum of the instruction position differences through absolute value calculation, which provides a reference for calculating the memory access interruption intensity. Obtain the latency and hit count of each memory access instruction hitting the cache. In the table "Example Table of Memory Access Interruption Behavior", the number of the second memory access instruction in the perturbed path is 25, the cache hit latency is 0.9 ms, and the hit count is 2. Summarize all instructions in turn and substitute them into the formula:

[0089] Let the position of the perturbed path of memory access instruction 1 be 12, the position of the comparison path be 10, the memory access time be 1.2 milliseconds, and the cache hit count be 3;

[0090] Let the position of the perturbed path of memory access instruction 2 be 25, the position of the comparison path be 22, the memory access time be 0.9 milliseconds, and the cache hit count be 2;

[0091] Let the position of the perturbed path of memory access instruction 3 be 37, the position of the comparison path be 36, the memory access time be 1.5 milliseconds, and the cache hit count be 4;

[0092] Let the position of the perturbed path of memory access instruction 4 be 49, the position of the comparison path be 52, the memory access time be 1.1 milliseconds, and the cache hit count be 1;

[0093] ;

[0094] ;

[0095] ;

[0096] ;

[0097] The result shows that there are obvious interruption characteristics in the current memory access segment in the instruction execution order. The memory access order interruption intensity value is a numerical index used to measure the severity of the disturbance of the order during the execution of memory access instructions. By comparing the difference between the original path order and the actual order of the current interruption segment, it helps to identify problems such as abnormal paths, cache behavior misalignment, and improper scheduling, which belong to the manifestations of disturbance behavior. By integrating the path interruption position difference and the cache latency hit behavior, the quantitative evaluation of the impact degree of path disturbance is realized, which helps to distinguish the boundary between normal fluctuations and abnormal disturbances and can be classified into the set of path segments of disturbance behavior.

[0098] The thermocouple linkage structure construction sub-module calls the set of path segments of disturbance behavior, extracts the path jump position and the hot path segment number, judges the position mapping and coincidence interval between the jump position and the hot path number, and generates an instruction disturbance thermocouple linkage structure;

[0099] Import the set of path segment numbers marked as disturbance behavior into the cache mapping module, scan and extract the path jump positions marked in each segment. If the jump point occurs between instruction 35 and instruction 47 where the path number is from P124 to P126, the jump position is recorded as the instruction number range [35, 47]. Then import the hot path area marking information, match the position range for each hot path segment number, and set that the hot path segment numbers H003 and H004 cover the instruction sections [30, 45] and [46, 60] in the path respectively. Then there is an overlap between the jump position range and the hot path number, and the overlapping part needs to be mapped. Set that instructions 35 to 45 belong to H003, and instructions 46 to 47 belong to H004, that is, this jump behavior partially coincides with both the hot path segments H003 and H004. Extract the memory access interruption intensity values of each overlapping position point, call the sequence of interruption intensity values calculated above, set the intensity corresponding to the position of instruction 35 to be 1.3 and instruction 46 to be 1.6, map each point in this sequence to the hot path, form a quantitative record of the disturbance impact under the hot path segment, integrate the overlapping point position numbers, path jump relationships and the corresponding memory access interruption intensity values, classify them according to the hot path number, and further organize them into an instruction disturbance thermocouple linkage structure.

[0100] Please refer to Figure 4 ,the load migration module includes:

[0101] Based on the instruction disturbance thermocouple linkage structure, the disturbance detection sub-module extracts the cache continuous read-write cycles, cache latency behavior and the set of path jump points between the GPU core and adjacent cores, monitors whether the cache continuous read-write cycles in the set of path jump points of the GPU core are in the cache latency behavior change interval, screens the path jump points contained in adjacent cores associated with the change interval, and generates a cache latency disturbance intensity measure;

[0102] Invoke the GPU scheduling control instruction to force the generation of a thermal coupling response on a specific thread scheduling node. At this time, locate the target GPU core and obtain the data exchange status between it and adjacent cores. In this state, by reading the cache access logs of each core, extract the number of read and write operations of the path jump point within a unit time, and obtain the cache continuous read and write cycle. For the case where the single jump cycle ranges from 110 to 140 nanoseconds, if the number of read and write operations exceeds 3000 times per second, it is marked as a high-frequency interval of the continuous cycle. In the read latency behavior log, use the cache access time difference as an evaluation index, extract the access latency value corresponding to the path jump point as the cache latency behavior. If the latency value is in the range of 4.5 to 7.5 milliseconds, it is regarded as a medium-high perturbation range. At the same time, combine the access paths that have mutated continuously more than 3 times in the jump point set and mark them as mutant jump points to identify the high-frequency response nodes of the current core under scheduling perturbation. Analyze the above-extracted read and write cycles and latency behaviors corresponding to the jump points to determine whether they form periodic perturbation characteristics. If a certain jump point has significant latency fluctuations (such as the jump amplitude exceeds 1.5 milliseconds) in multiple perturbation cycles, it is determined as a perturbation key node. Aggregate this node to identify whether there is a synchronous perturbation phenomenon in adjacent cores, that is, if the adjacent cores have a response fluctuation with a latency amplitude within ±1 millisecond in the same time window, it indicates that this perturbation is a cross-core correlation perturbation, which is classified as a candidate interference path. Correspondingly, extract its cache access frequency and cycle offset data, and further count the perturbation frequency distribution. If the perturbation frequency of this path exceeds 20% of the total number of jumps within a unit time, it is recorded as a strong cache perturbation path. Through the screening of the full path set, output a set of path nodes with clear cache perturbation response characteristics, and generate a cache delay perturbation intensity metric.

[0103] The path screening sub-module invokes the cache delay perturbation intensity metric to determine whether there are stable path jump sequences and cache hit cycle stable nodes in adjacent cores. If the node hit cycle fluctuation is within the hit cycle jump constraint interval, extract the adjacent cores as candidate path cores and generate a set of path stability values;

[0104] The synchronous path structure of the involved adjacent cores is identified. By analyzing the jump sequence changes in each adjacent core, the periodic jump frequency of the node in the continuous time slice is recorded. If the fluctuation amplitude of the frequency in three consecutive cycles is less than 5%, it is preliminarily determined to be a stable path jump sequence. The path cycle records of a certain adjacent core are 60, 62, and 61 nanoseconds respectively. The jump amplitude of the sequence is only 2 nanoseconds, which meets the requirements of stable path characteristics. The volatility of the node hit cycle is evaluated. After reading the hit cycle log, the hit count and interval in the corresponding path are extracted, and the hit cycle variation range is calculated. If the range is within ±1.5 milliseconds, it is regarded as a hit cycle stable node. If a certain jump If the hit cycle range in the path is between 59.5 and 61 milliseconds, the stable node setting standard is met. At this time, the adjacent core is recorded as a candidate path core, and the path mapping parameters under the core are extracted, that is, the time sequence and corresponding address mapping information in its node sequence, as well as the cycle parameter set, that is, the hopping frequency distribution corresponding to the stable cycle sequence. The candidate path core information that meets the path stability and hit stability conditions is summarized. In the actual instance, if there is a core A with a hopping frequency of 2%, a hit cycle fluctuation of 1.2 milliseconds, a continuous path length greater than 5 jump units, and all node cycles between 60±2 nanoseconds, it is regarded as a stable path core candidate, and a path stability value set is constructed.

[0105] The core mapping submodule constructs the inter-core migration path and maps the node jump trajectory path sequence according to the path stability value set, using the formula:

[0106] ;

[0107] Calculate the migration trajectory offset strength value, evaluate the trajectory mapping continuity of the candidate core, and obtain the core migration path mapping trajectory set;

[0108] in, represents the migration trajectory offset intensity value, Indicates the jump path The period value of Indicates the path The path continuity, Indicates the path The cache disturbance amplitude is Indicates the path The average jump span, Indicates the path The hit cycle offset, Indicates the path The stable period value of Indicates the hit cycle offset adjustment value. represents the path stability trade-off term, Indicates the path cycle judgment limit;

[0109] The set of call path stability values contains 4 candidate paths, numbered P1, P2, P3, and P4 respectively. Each path has 7 key parameters: the jump path period , path continuity , cache perturbation amplitude , average jump span , hit cycle offset , stable cycle value . The data is as follows: The period of path P1 is 120 ns, the continuity is 0.85, the cache perturbation is 5.2 ms, the span is 30, the hit offset is 2.1, and the stable cycle value is 60; The period of path P2 is 115 ns, the continuity is 0.78, the cache perturbation is 6.1 ms, the span is 28, the hit offset is 2.3, and the stable cycle value is 62; The period of path P3 is 134 ns, the continuity is 0.91, the cache perturbation is 4.8 ms, the span is 33, the hit offset is 1.9, and the stable period is 59; The period of path P4 is 112 ns, the continuity is 0.69, the cache perturbation is 7.0 ms, the span is 25, the hit offset is 2.6, and the stable period is 63;

[0110] By comparing the cache perturbation amplitudes of each path , initially screen out the path nodes with cache perturbation values higher than the set offset determination limit of 10 ms. Then, based on the square value of the path time continuity and the cache perturbation amplitude , take the absolute value of the difference between the square root of their sum and the jump path period to measure the cycle offset rate of each path under perturbation. Add the jump spans of each node and the hit cycle offset after weighting, and calculate the total sum to construct the following denominator expression. At the same time, take the absolute value of the difference between the sum of the stable cycle values and the stable determination threshold and then weight it;

[0111] is the jump cycle of each path node, and the data is recorded by the monitoring point;

[0112] is the path continuity, which can be obtained by normalizing the reciprocal of the standard deviation of the jump time;

[0113] is the cache perturbation amplitude, which is calculated from the cache cycle fluctuation;

[0114] is the average jump span, which can be obtained by the average jump distance between the previous and next jump nodes of the path;

[0115] is the hit cycle offset, calculated from the deviation between the node hit interval and the ideal hit cycle;

[0116] is the path stability period value, obtained by measuring the jump sequence period;

[0117] and and are the set weights and thresholds;

[0118] Based on the path jump data, numerical substitution calculations are performed as follows:

[0119] For P1: ;

[0120] For P2: Similarly ;

[0121] For P3: ;

[0122] For P4: ;

[0123] Summing up the above results, the numerator part is 471.68, and the denominator part is:

[0124] ;

[0125] The total stable period is: ;

[0126] Therefore, the offset term is: ;

[0127] Substitute into the formula for calculation: ;

[0128] This result indicates that the offset intensity value under path P4 is the lowest. The offset intensity value of the migration trajectory represents the degree of deviation in terms of time, space, cache, jump, etc. between the execution path and the target core mapping trajectory when the task migrates between multiple cores. Therefore, select its node jump path to establish the core migration path mapping trajectory set.

[0129] Please refer to Figure 5 , the scheduling label revision module includes:

[0130] The path trajectory extraction sub-module extracts the start and end nodes of cross-core migration, the number of path jumps, and the migration frequency value of the execution path based on the execution instruction records and the corresponding core numbers in the core migration path mapping trajectory set, and generates the cross-core path jump trajectory by jointly comparing the migration frequency value and the number of jumps;

[0131] Unfold the instruction execution sequence in the scheduling log one by one, determine the core number to which each record belongs, and build a task migration path list based on this number. Each path corresponds to a continuous migration process from the source core to the target core. Suppose in a multi-core server environment, task M001 runs on core 1 for half of the time and then is transferred to core 4. Then the starting node of this migration path is core 1, and the ending node is core 4. By traversing the execution instruction segments involved in the path, mark the positions where migration actions occur, record the node numbers of the actions as jump nodes, count the total number of jump nodes in each path, and combine their occurrence frequencies within a unit time to build a data list of the frequency of cross-core jumps. In a certain high-load computing task, path P78 is migrated 5 times within one minute and contains scheduling segments of 3 different cores. This path can be determined as a cross-core migration path. To avoid misidentifying low-frequency noise paths, set an identification criterion for cross-core jumps, and only retain records with more than 3 jumps and containing 2 or more core numbers in the path. The system filters out the paths that meet the conditions from all scheduling paths and outputs their starting and ending nodes, the number of jumps, and the set of core numbers to generate the cross-core path jump trajectory.

[0132] The cache structure screening sub-module calls the continuous path segment number sequence and path interruption positions in the cross-core path jump trajectory, and generates a cache hit fluctuation range by counting the cache read times and cache hit records of the instruction sets before and after the path segment, and comparing the change rate of the read times and the fluctuation range of the hit records.

[0133] For each paragraph marked as a migration path, read the instruction sets executed before and after the interruption of this path segment, and count the cache read times and hit times generated respectively according to the cache interaction records during the scheduling process of the instructions. Suppose in path segment D41, the system reads 35 cache commands, and the hit record is 28 times. After the path interruption, in segment D42 on the migration target core, 40 cache commands are read, and the hit record is only 21 times. By comparing the cache call behaviors of the front and back segments, the impact of migration on the cache hit situation can be intuitively judged. The system takes adjacent path segments as units, compares the difference in read times and the floating range of the hit records respectively, and introduces paragraphs with a read difference exceeding a certain proportion and a significant hit difference as the discrimination basis during the judgment. Set path segments with a read increase of more than 30% and a hit rate decrease of more than 20% to be marked as significantly cache fluctuating by the system. In an example application, the hit times of path segment D58 decrease from 25 to 12 after migration, and the read times increase from 30 to 47. Then this segment is classified as a hit fluctuation path segment. All path segments that meet the above conditions and their interruption positions are recorded and uniformly output to form a cache hit fluctuation range.

[0134] The execution sequence comparison sub-module extracts the sequence structure numbers of the instructions before and after scheduling under the same path according to the path set corresponding to the cache hit fluctuation range, compares the execution cycle value and the start time offset value of the executed instructions, filters out the paths with abnormal offset trends, and obtains the execution sequence offset fluctuation degree;

[0135] For each path number, extract the execution instruction sequence structure in the two stages before and after scheduling, and mark its start time point and corresponding duration on the time axis. Taking path P91 as an example, the start time of the task before migration is 240 milliseconds and it takes 15 milliseconds to complete; after migration, due to the core scheduling delay, the new execution start time is 270 milliseconds and the duration is 22 milliseconds. Through this time comparison, it is possible to calculate whether there is an elongation of the execution cycle before and after migration of this path, as well as the offset amplitude of the scheduling start time. The system further statistically analyzes similar changes in a batch of paths, and forms an offset trend index by accumulating multiple instruction cycle changes and start time differences. If multiple instructions in a certain path show continuous delay in start-up, increase in execution cycle, etc., this path will be judged as an abnormal migration path by the system. In an actual calculation scenario, during the parallel processing of multiple frames of images in an image processing task, if the start time of processing a single frame of image after migration of path P105 is delayed by more than 30 milliseconds and the cycle increase exceeds 10 milliseconds three times in a row, it can be included in the abnormal list. The system summarizes the numbers of all abnormal paths and their corresponding delay directions to obtain the execution sequence offset fluctuation degree.

[0136] The path compression sub-module calls the abnormal path number and offset direction attribute in the execution sequence offset fluctuation degree, adjusts the sorting order of the path segments according to the offset direction attribute, and compresses the length value of the continuous jump path segments to obtain the path compression scheduling label;

[0137] According to the marked delay or advance direction in each path, rearrange the order of its path segment numbers. Suppose a path contains paragraph numbers P300, P302, and P305, where the offset of P300 is 25 milliseconds, P302 is 15 milliseconds, and P305 is 40 milliseconds. Then the new path sorting is P302, P300, P305, in order to compress the redundancy of the execution chain caused by scheduling offsets. Based on the new sorting, compress the length of the continuous jump structure in each path segment, that is, merge the paragraphs that are logically jump but highly continuous in execution, reducing the total number of path segment numbers. Suppose the original path contains 8 paragraphs, and after compression, they are merged into 5 paragraphs. The system records the change in the number of paragraphs before and after compression and constructs path index information, including the path number of each segment, the compressed length, the compression direction, and the compression amplitude. In one instance, path P410 is compressed from 9 segments to 6 segments, and the compression amplitude exceeds 30%. The system marks this segment as a valid compression segment and writes it into the index table. Summarize the index content of all compressed path segments uniformly, and generate a callable path recognition tag set according to multi-dimensional fields such as path number sorting and compression rate sorting, forming a path compression scheduling tag.

[0138] Please refer to Figure 6 , the node priority sorting module includes:

[0139] Based on the path compression scheduling tag, the path compression extraction sub-module extracts the path compression node numbers, the cache behavior records corresponding to the nodes, and the stable time of the jump sequence. Establish a pairing relationship between the stable time of the jump sequence and the path compression node numbers according to the node order, judge whether the stable time of the jump sequence shows a continuous increasing trend and mark the corresponding numbered paragraphs, and generate a compressed node number interval sequence;

[0140] Extract and process the path compression node numbers, the cache behavior records corresponding to the nodes, and the stable time of the jump sequence in sequence to obtain the node call records in the path compression scheduling label. Store the jump association paths between the node numbers in the form of a graph structure, extract the set of numbers of the nodes connected by each jump edge, and parse the time tag field thereof to extract the stable time parameter value when the jump action occurs. This stable time field is in milliseconds and records the duration of the state of a certain node during adjacent jumps. Set the stable time of the jump between node numbers 301 and 302 to 420 ms, then mark it as the stable duration of the 301→302 path. Based on the depth-first traversal method of the graph structure, pair the stable time fields of the jump sequence with the corresponding node numbers in topological order of the node numbers to form a set of stable sequence pairs. At the same time, record the cache behavior records before each node jump, including the hit status (hit / miss), the number of write requests, and the replacement operation count, and store them in the corresponding node fields in the form of an array. Set the hit status of node number 305 to 22 hits, 9 write requests, and 5 replacement times. After constructing the pairing relationship between the stable time of the jump sequence and the node number, it is necessary to judge whether the pairing sequence shows an increasing trend. The direction is judged by the difference between adjacent stable times. Let the adjacent two jump times be and , if , then mark it as increasing. If several consecutive jump sequences satisfy the monotonically increasing stable time, that is, they are classified into a continuously growing paragraph, perform index marking processing on the number sequence to form a numbered section group, and generate a set of node numbers in the continuously growing interval as the compressed path segment output. Set the stable times in the jump sequence 301→302→303→304 to 420 ms, 440 ms, 470 ms, and 490 ms respectively, which can be classified into a continuous interval, and the corresponding number set is {301, 302, 303, 304}, generating a compressed node number interval sequence.

[0141] The sub-module for dividing the behavior trend calls the compressed node number interval sequence, analyzes the cache behavior records corresponding to the node numbers, and divides them according to the synchronous fluctuation characteristics of the hit frequency, write ratio, and replacement rate in the time dimension in the cache behavior records, evaluates the numbered belonging relationship of various fluctuation trends in the path segment, and generates the interval of the behavior change trend.

[0142] Analyze the cache behavior records corresponding to the node numbers, extract the data sets of the hit frequency, write ratio, and replacement rate for each node, construct a behavior record matrix in a multi-dimensional time axis synchronization manner, set the node number as the row index, and the behavior record fields as column vectors. Set the behavior record of node number 305 as 22 hits, 9 write requests, and 5 replacement operations. Then the write ratio is 9 / (22 + 9) ≈ 0.29, and the replacement rate is 5 / (22 + 9 + 5) ≈ 0.14, forming a behavior record vector [0.29, 0.14]. After forming the behavior record matrix, use the time series change trend between the behavior vectors as the classification basis, adopt the sliding time window technology to calculate the co-fluctuation of the behavior changes, define the behavior vector change rate as the difference value of the Euclidean distance between the vectors of adjacent two nodes. If the change rate between consecutive multiple nodes remains within a fixed range Δr, it is classified into the same behavior trend section. Let Δr = 0.1. Then if the behavior vector change rate between node i and i + 1 is 0.05, and between i + 1 and i + 2 is 0.07, the three nodes are grouped together. Evaluate the set of numbers in the path segment in turn, delimit the number intervals belonging to various fluctuation trend segments, and generate the behavior change trend intervals.

[0143] The node sorting sub-module filters the node numbers with stable jump sequence characteristics according to the behavior change trend intervals, judges the difference direction of the hit frequency and replacement rate in the cache behavior records corresponding to the node numbers, selects the node numbers with the hit frequency exceeding the replacement rate, and arranges them in order according to the original path position of the nodes to generate a sorted coherent node sequence;

[0144] Perform screening and judgment operations on the nodes, call the node number and its corresponding jump sequence stable time and include them in the screening range, and calculate the node stable feature coefficient:

[0145] S = T×(H - R);

[0146] Where, T is the jump sequence stable time, H is the hit frequency, and R is the corresponding number of replacement rates;

[0147] If the value of S is positive, it means that the node has a hit-dominated behavior in the stable jump stage. Set T = 420ms, H = 22, R = 5, then:

[0148] S = 420×(22 - 5) = 7140;

[0149] Set the node stable feature reference value As the screening threshold, set according to the median value of the average stable features of the node behavior samples. If the current node , it is considered to have the characteristics of stable jump behavior. Further compare the difference between the hit frequency and the replacement rate value, and judge whether the hit frequency exceeds the replacement rate. Those that meet the conditions are included in the set of reserved node numbers, and then sorted according to the original number order of the nodes in the path compression structure. Set the original structure order as 301→303→305→306, and if the selected nodes are 303, 305, and 306, then keep their order unchanged and arrange them as 303→305→306 to generate a sorted and coherent node sequence.

[0150] The execution path adjustment sub-module calls the sorted and coherent node sequence, combines the adjacent structure between the stable time of the corresponding jump sequence and the original node number sequence, judges whether there is a node coherence gap in the path, retains the path segment numbers with a complete coherent structure, eliminates the paragraph numbers where the jump abnormal nodes are located, and integrates the consecutive numbers to form a stable path node combination to obtain the GPU stable execution path sorted list;

[0151] Combined with the paired structure of the stable time field of the jump sequence and the original node number sequence, judge the path coherence gap. Let the node and The stable time between them is , number difference:

[0152] ;

[0153] If and ≥ ;

[0154] is the average value of the jump stable time, then it is recognized as a coherent path segment;

[0155] If >1 or < , then it is a jump abnormal segment, and such node pairs are eliminated;

[0156] Set between nodes 301→304 =3, =360ms, =400ms, then judge that this segment is an abnormal paragraph and do not include it in the path structure. Retain the node segment group that meets the continuous number and stable jump time, construct a stable node combination path set, such as 303→304→305→306, which is a structure continuous and jump stable segment, and integrate them in turn to form a complete path segment number set and output it as the GPU stable execution path sorted list.

[0157] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A GPU computing power dynamic scheduling system, characterized in that, The system includes: The thermal expansion monitoring module obtains the time series records of the GPU core thermal sensor and the motherboard thermal expansion response trajectory, calls the thermal expansion speed, temperature rise duration, and diffusion direction path number in the core running state, identifies whether there is information covering the task instruction execution area in the heat conduction during the scheduling period, and generates heat path coverage marking information; Based on the heat path coverage marking information, the path perturbation identification module retrieves whether the path numbers are continuous and extracts the positions where path jump behaviors occur. If the path numbers are discontinuous in two adjacent cycles and the memory access order is interrupted, the marked segment is a perturbed behavior path, and an instruction perturbation thermal coupling linkage structure is generated; The load migration module calls the cache continuous read and write cycles between the GPU core located by the instruction perturbation thermal coupling linkage structure and the adjacent cores, retrieves whether there are path jump stable sequences and cache hit cycle stable nodes in the adjacent cores, and generates a core migration path mapping trajectory set; The scheduling label revision module uses the core migration path mapping trajectory set, compares the heat response time change behavior and the execution order change trajectory, filters the continuous path structure and the record of shortened cache after migration, and generates a path compression scheduling label.

2. The GPU computing power dynamic scheduling system according to claim 1, wherein The heat path coverage marking information includes the thermal expansion path sequence number, the starting boundary of the heat conduction coverage area, and the task path overlap index point. The instruction perturbation thermal coupling linkage structure includes a set of perturbed path numbers, a thermal area coupling mapping index, and an in-cycle jump synchronization mark. The core migration path mapping trajectory set includes the target core path jump sequence, the cross-core path delay trajectory line, and the path mapping section. The path compression scheduling label includes a path compression identification number, a cache usage fallback section, and a sequential behavior continuity mark.

3. The GPU computing power dynamic scheduling system according to claim 1, wherein The thermal expansion monitoring module includes: The time series extraction sub-module obtains the time series records of the GPU core thermal sensor and the motherboard thermal expansion response trajectory, detects the core running state parameters and the motherboard thermal expansion response values at multiple moments, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal sensor value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value; Based on the temperature rise duration sequence value, the diffusion path identification sub-module calls the thermal expansion response trajectory number and the core working state path in the corresponding time period, calculates the unit time thermal change value and the diffusion range value in each trajectory, and filters according to the difference between the diffusion value and the path number to obtain a set of thermal expansion direction path numbers; According to the set of thermal expansion direction path numbers, the heat path coverage marking sub-module matches the corresponding execution area of the task instruction in the task scheduling period, judges whether the thermal expansion propagation area in the path number set overlaps with the instruction execution position in space, calculates the heat path conduction matching degree, analyzes the spatial position index and the instruction number corresponding to the marked coverage area, and generates heat path coverage marking information.

4. The GPU computing power dynamic scheduling system according to claim 3, wherein The path perturbation identification module includes: Based on the hot path coverage marking information, the path number extraction sub-module obtains the path numbers of the GPU cores, the start and end timestamps of instruction segments, the memory access instruction order, and the cache hit time records during the task running cycle, extracts the path numbers within two adjacent cycles, sorts the path numbers, and retrieves the continuity of the number sequence, filters out the path number segments that do not meet the number continuity determination conditions, and generates path number interruption interval values; According to the path number interruption interval values, the memory access order retrieval sub-module calls the memory access instruction order and cache hit time records, extracts the memory access segments corresponding to the path interruptions, compares the memory access order in the segments with the memory access order of the paths in the same cycle, calculates the memory access order interruption intensity values, and obtains the set of perturbed behavior path segments; The thermal coupling linkage structure construction sub-module calls the set of perturbed behavior path segments, extracts the path jump positions and hot path segment numbers, performs position mapping and coincidence interval judgment on the jump positions and hot path numbers, and generates an instruction perturbation thermal coupling linkage structure.

5. The GPU computing power dynamic scheduling system according to claim 4, wherein, The load migration module includes: Based on the instruction perturbation thermal coupling linkage structure, the perturbation detection sub-module extracts the cache continuous read / write cycles, cache latency behaviors, and path jump point sets between the GPU core and neighboring cores, monitors whether the cache continuous read / write cycles in the GPU core path jump point set are in the cache latency behavior change interval, filters out the path jump points contained in the neighboring cores associated with the change interval, and generates the cache latency perturbation intensity measure; The path screening sub-module calls the cache latency perturbation intensity measure, determines whether there are path jump stable sequences and cache hit cycle stable nodes in the neighboring cores. If the node hit cycle fluctuates within the hit cycle jump constraint interval, it extracts the neighboring cores as candidate path cores and generates a set of path stability values; According to the set of path stability values, the core mapping sub-module constructs the migration path between cores and maps the node jump trajectory path sequence, calculates the migration trajectory offset intensity value, evaluates the trajectory mapping continuity of the candidate cores, and obtains the core migration path mapping trajectory set.

6. The GPU computing power dynamic scheduling system according to claim 5, wherein The scheduling label revision module includes: Based on the execution instruction records and corresponding core numbers in the core migration path mapping trajectory set, the path trajectory extraction sub-module extracts the start and end nodes of cross-core migration, the number of path jumps, and the migration frequency values of the execution path, performs a joint comparison according to the migration frequency values and the number of jumps, and generates the cross-core path jump trajectory; The cache structure screening sub-module calls the continuous path segment number sequence and path interruption positions in the cross-core path jump trajectory, counts the cache read times and cache hit records of the instruction sets before and after the path segments, compares the change rate of the read times and the fluctuation interval of the hit records, and generates the cache hit fluctuation interval; According to the path sets corresponding to the cache hit fluctuation interval, the execution sequence comparison sub-module extracts the sequence structure numbers of the instructions before and after scheduling under the same path, compares the execution cycle values and start time offset values of the execution instructions, filters out the paths with abnormal offset trends, and obtains the execution sequence offset fluctuation degree; The path compression sub-module calls the abnormal path number and offset direction attribute in the execution sequence offset fluctuation degree, adjusts the path segment sorting order according to the offset direction attribute, compresses the length value of the continuous jump path segment, and obtains the path compression scheduling label.

7. The GPU computing power dynamic scheduling system according to claim 1, wherein The system further includes a node priority sorting module: The node priority sorting module calls the path compression scheduling label, extracts the path compression node number, cache behavior record, and jump sequence stability time, identifies the number of stable path formations and the behavior trend interval, sorts and schedules the nodes according to the path structure coherence, and generates a GPU stable execution path sorting list; The GPU stable execution path sorting list includes the sorted node number, behavior stability sequence value, and scheduling priority index code.

8. The GPU computing power dynamic scheduling system according to claim 7, characterized in that The node priority sorting module includes: The path compression extraction sub-module, based on the path compression scheduling label, extracts the path compression node number, the cache behavior record corresponding to the node, and the jump sequence stability time, establishes a pairing relationship between the jump sequence stability time and the path compression node number according to the node order, determines whether the jump sequence stability time shows a continuous increasing trend, and marks the corresponding numbered paragraph to generate a compressed node number interval sequence; The behavior trend division sub-module calls the compressed node number interval sequence, analyzes the cache behavior record corresponding to the node number, divides according to the synchronous fluctuation characteristics of the hit frequency, write ratio, and replacement rate in the cache behavior record in the time dimension, evaluates the numbered attribution relationship of multiple types of fluctuation trends in the path segment, and generates a behavior change trend interval; The node sorting sub-module, according to the behavior change trend interval, filters the node numbers with stable jump sequence characteristics, judges the difference direction of the hit frequency and replacement rate in the cache behavior record corresponding to the node number, selects the node numbers with the hit frequency exceeding the replacement rate, and arranges them in order according to the original path position of the node to generate a sorted coherent node sequence; The execution path adjustment sub-module calls the sorted coherent node sequence, combines the adjacent structure between the corresponding jump sequence stability time and the original node number sequence, judges whether there is a node coherence gap in the path, retains the path segment numbers with a complete coherent structure, eliminates the paragraph numbers of the jump abnormal nodes, and integrates the continuous numbers to form a stable path node combination to obtain a GPU stable execution path sorting list.

Citation Information

Patent Citations

  • Host power consumption and temperature balance control method and system, medium and program product

    CN119536958A

  • HBase client main and standby switching method and system based on fault perception

    CN119537484A