GPU computing power dynamic scheduling system

By monitoring the thermal expansion response and instruction path behavior of the GPU core in real time, identifying thermal conduction coverage and path disturbance behavior, and generating path migration mapping trajectory and path compression scheduling tags, the problem of unbalanced resource allocation in the existing technology is solved, and the utilization rate of GPU resources and the overall performance of the system is improved.

CN120179045AActive Publication Date: 2025-06-20CHENGDU WANDA ELECTRONIC TECH CO LTD

Patent Information

Application Number
CN202510646898.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The prior art cannot accurately perceive the specific needs of GPU resources of different task types in high concurrency environments, resulting in uneven resource allocation and reducing overall system performance.

Method used

Through the thermal expansion monitoring module, path disturbance identification module, load migration module and scheduling tag revision module, the thermal expansion response and instruction path behavior of the GPU core are monitored in real time, the thermal conduction coverage and path disturbance behavior are identified, and the path migration mapping trajectory and path compression scheduling tags are generated to optimize the allocation of GPU resources.

Benefits of technology

It effectively avoids performance bottlenecks caused by overheating of the task execution area, improves the utilization rate of GPU resources and the stability of task execution, and enhances the overall execution efficiency and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179045A_ABST
    Figure CN120179045A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of resource scheduling, in particular to a GPU computing power dynamic scheduling system which comprises a thermal expansion monitoring module, a path disturbance identification module, a load migration module, a scheduling label revision module and a node priority ordering module. According to the method, the GPU core thermal expansion response and the instruction path behavior are monitored in real time, so that the performance bottleneck caused by overheating of a task execution area is effectively avoided, the heat conduction and the calculation load during resource scheduling are ensured to be better coordinated, the dynamic adjustment of the thermal response time and the path compression scheduling label is combined, and the scheduling efficiency is improved. Dynamic optimization scheduling can be carried out according to changes of task loads and resource utilization rates, load balance among GPU computing units is enhanced, excessive competition and idleness of resources are avoided, the overall execution efficiency and response speed of a system are improved, and the resource scheduling technology is optimized through thermal behavior and path disturbance analysis. And the accuracy and the stability of GPU computing power scheduling in a multi-task concurrent execution environment are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource scheduling, and in particular, to a dynamic scheduling system for GPU computing power. Background Art

[0002] The technical field of resource scheduling includes the dynamic management and reasonable allocation mechanisms of computing resources in various hardware and software environments. The core content of this technical field is to face the multi-task concurrent execution environment for multi-type computing resources such as general-purpose processors, graphics processors, application-specific integrated circuits, field-programmable gate arrays and other hardware resources. By perceiving the resource status, analyzing the task load and formulating a scheduling strategy, the resources can be efficiently matched and distributed in multiple dimensions such as computing density, task priority and load balancing. This technical field also includes cross-platform and cross-architecture resource integration scheduling algorithm design, scheduling rule formulation, and task mapping mechanisms in heterogeneous resource environments, systematically building a connection bridge between computing resources and business requirements, and realizing the dynamic overall planning and refined management of resources in the time and space dimensions.

[0003] Among them, the dynamic scheduling system for GPU computing power refers to the resource management problem of the graphics processor in a multi-task scenario. By constructing a task recognition method to monitor the load characteristics of the current GPU task in real time and comparing the task characteristics with the GPU resource status, using a scheduling priority determination method to judge the scheduling order of each GPU core, and then combining a task type classification method to map different types of tasks to the appropriate GPU computing units, the orderly distribution of GPU resources and the coordination of computing power load between nodes are realized. This system uses task analysis technology, resource status perception methods, and scheduling rule matching mechanisms to complete the distribution process of computing power tasks.

[0004] The existing resource scheduling technologies mainly rely on static comparison of task load and resource status for scheduling decisions, ignoring the dynamic changes of GPU computing resources caused by thermal expansion and path disturbances during the actual execution process. This leads to phenomena such as task delays or decreased computing performance due to excessive temperature or large path execution fluctuations in a high-concurrency environment. For example, when multiple tasks are executed on the same GPU core simultaneously, if real-time monitoring and dynamic adjustment of heat conduction coverage and path disturbances are not performed, some tasks may not be completed in time due to insufficient computing resources, affecting the execution efficiency of the entire system. The task scheduling mechanism in the prior art cannot accurately perceive the specific requirements of different task types for GPU resources, resulting in unbalanced resource allocation and reduced overall system performance in complex task combinations and multi-task environments. This scheduling mechanism based on fixed rules and load balancing has limitations in dealing with complex resource allocation requirements under heterogeneous computing platforms and cannot fully adapt to changing workloads and dynamic changes of computing resources. Summary of the Invention

[0005] The object of the present invention is to solve the disadvantages existing in the prior art, and a GPU computing power dynamic scheduling system is proposed.

[0006] To achieve the above object, the present invention adopts the following technical solutions. A GPU computing power dynamic scheduling system includes: The thermal expansion monitoring module obtains the time series records of the GPU core thermal sensors and the motherboard thermal expansion response trajectory, calls the thermal expansion speed, temperature rise duration, and diffusion direction path number in the core running state, identifies whether there is information covering the task instruction execution area in the heat conduction during the scheduling period, and generates heat path coverage marking information; Based on the heat path coverage marking information, the path perturbation identification module retrieves whether the path numbers are continuous and extracts the positions where path jump behaviors occur. If the path numbers are not continuous in two adjacent periods and the memory access order is interrupted, the marked segment is a perturbed behavior path, and an instruction perturbation thermal coupling linkage structure is generated; The load migration module calls the cache continuous read and write cycles between the GPU core located by the instruction perturbation thermal coupling linkage structure and the adjacent cores, retrieves whether there are path jump stable sequences and cache hit cycle stable nodes in the adjacent cores, and generates a core migration path mapping trajectory set; The scheduling label revision module uses the core migration path mapping trajectory set, compares the heat response time change behavior and the execution order change trajectory, and screens the continuous path structure and the cache shortening record after migration to generate a path compression scheduling label.

[0007] As a further solution of the present invention, the heat path coverage marking information includes a thermal expansion path sequence number, a starting boundary of the heat conduction coverage area, and a task path overlap index point. The instruction perturbation thermal coupling linkage structure includes a perturbed path number set, a thermal area coupling mapping index, and an in-cycle jump synchronization mark. The core migration path mapping trajectory set includes a target core path jump sequence, a cross-core path delay trajectory line, and a path mapping section. The path compression scheduling label includes a path compression identification number, a cache usage fallback section, and a sequential behavior continuity mark.

[0008] As a further solution of the present invention, the thermal expansion monitoring module includes: The time series extraction sub-module obtains the time series records of the GPU core thermal sensors and the motherboard thermal expansion response trajectory, detects the core running state parameters and the motherboard thermal expansion response values at multiple moments, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal sensor value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value; The diffusion path recognition sub-module, based on the temperature rise duration sequence values, calls the thermal diffusion response trajectory numbers and the core working state paths within the corresponding time period, calculates the unit time thermal sensitivity change values and the diffusion range values within each trajectory, and filters according to the difference between the diffusion values and the path numbers to obtain the set of thermal diffusion direction path numbers; The hot path coverage marking sub-module matches the corresponding execution area of the task instruction within the task scheduling period according to the set of thermal diffusion direction path numbers, determines whether there is a spatial overlap between the thermal diffusion propagation area in the path number set and the instruction execution position, calculates the hot path conduction matching degree, analyzes the spatial position index and the instruction number corresponding to the marked coverage area, and generates the hot path coverage marking information.

[0009] As a further solution of the present invention, the path perturbation recognition module includes: The path number extraction sub-module, based on the hot path coverage marking information, obtains the path numbers of the GPU core, the start and end timestamps of the instruction segments, the memory access instruction order, and the cache hit time records during the task operation period, extracts the path numbers within two adjacent periods, sorts the path numbers and retrieves the continuity of the number sequence, filters the path number segments that do not meet the number continuity determination conditions, and generates the path number interruption interval value; The memory access order retrieval sub-module, according to the path number interruption interval value, calls the memory access instruction order and the cache hit time records, extracts the memory access segments corresponding to the path interruption, compares the memory access order in the segment with the memory access order of the same period path, calculates the memory access order interruption intensity value, and obtains the set of perturbation behavior path segments; The thermal coupling linkage structure construction sub-module calls the set of perturbation behavior path segments, extracts the path jump positions and the hot path segment numbers, performs position mapping and coincidence interval judgment on the jump positions and the hot path numbers, and generates the instruction perturbation thermal coupling linkage structure.

[0010] As a further solution of the present invention, the load migration module includes: The perturbation detection sub-module, based on the instruction perturbation thermal coupling linkage structure, extracts the cache continuous read and write cycles, the cache latency behavior, and the path jump point set between the GPU core and the adjacent cores, monitors whether the cache continuous read and write cycles in the GPU core path jump point set are in the cache latency behavior change interval, filters the path jump points contained in the adjacent cores associated with the change interval, and generates the cache latency perturbation intensity measure; The path screening sub-module calls the cache latency perturbation intensity measure to determine whether there are path jump stable sequences and cache hit cycle stable nodes in the adjacent cores. If the node hit cycle fluctuates within the hit cycle jump constraint interval, the adjacent cores are extracted as candidate path cores, and the set of path stability values is generated; The core mapping sub-module constructs the inter-core migration paths and maps the node jump trajectory path sequences according to the set of path stability values, calculates the migration trajectory offset intensity value, evaluates the trajectory mapping continuity of the candidate cores, and obtains the core migration path mapping trajectory set.

[0011] As a further solution of the present invention, the scheduling label revision module includes: The path trajectory extraction sub-module extracts the start and end nodes of cross-core migration, the number of path jumps and the migration frequency value of the execution path based on the execution instruction records and the corresponding core numbers in the core migration path mapping trajectory set, and generates a cross-core path jump trajectory by jointly comparing the migration frequency value and the number of jumps. The cache structure screening sub-module calls the continuous path segment number sequence and the path interruption position in the cross-core path jump trajectory, and generates a cache hit fluctuation range by statistically analyzing the cache read times and cache hit records of the instruction sets before and after the path segment, and comparing the change rate of the read times and the fluctuation range of the hit records. The execution sequence comparison sub-module extracts the sequence structure numbers of the instructions before and after scheduling under the same path according to the path set corresponding to the cache hit fluctuation range, compares the execution cycle value and the start time offset value of the execution instructions, screens the paths with abnormal offset trends, and obtains the execution sequence offset fluctuation degree. The path compression sub-module calls the abnormal path numbers and offset direction attributes in the execution sequence offset fluctuation degree, adjusts the sorting order of the path segments according to the offset direction attributes and compresses the length value of the continuous jump path segments, and obtains the path compression scheduling label.

[0012] As a further solution of the present invention, the system further includes a node priority sorting module: The node priority sorting module calls the path compression scheduling label, extracts the path compression node numbers, cache behavior records and jump sequence stable times, identifies the number of stable path formations and the behavior trend intervals, sorts and schedules the nodes according to the path structure coherence, and generates a GPU stable execution path sorting list. The GPU stable execution path sorting list includes sorted node numbers, behavior stability sequence values, and scheduling priority index codes.

[0013] As a further solution of the present invention, the node priority sorting module includes: The path compression extraction sub-module extracts the path compression node numbers, the cache behavior records corresponding to the nodes and the jump sequence stable times based on the path compression scheduling label, establishes a pairing relationship between the jump sequence stable times and the path compression node numbers in the node order, determines whether the jump sequence stable times show a continuous increasing trend and marks the corresponding numbered paragraphs, and generates a compressed node number interval sequence. The behavior trend division sub-module calls the compressed node number interval sequence, analyzes the cache behavior records corresponding to the node numbers, and divides them according to the synchronous fluctuation characteristics of the hit frequency, write ratio, and replacement rate in the time dimension in the cache behavior records. It evaluates the number attribution relationship of multiple types of fluctuation trends in the path segment and generates a behavior change trend interval. The node sorting sub-module filters the node numbers with stable jump sequence characteristics according to the behavior change trend interval, judges the difference direction of the hit frequency and replacement rate in the cache behavior records corresponding to the node numbers, selects the node numbers with the hit frequency exceeding the replacement rate, and arranges them in order according to the original path position of the nodes to generate a sorted coherent node sequence. The execution path adjustment sub-module calls the sorted coherent node sequence, combines the adjacent structure between the stable time of the corresponding jump sequence and the original node number sequence, judges whether there are node coherence gaps in the path, retains the path segment numbers with complete coherent structures, eliminates the paragraph numbers where the jump abnormal nodes are located, and integrates the continuous numbers to form a stable path node combination to obtain the GPU stable execution path sorted list.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, through the real-time monitoring of the GPU core thermal expansion response and the instruction path behavior, the heat conduction coverage and path perturbation behaviors can be accurately identified during the task execution process, effectively avoiding the performance bottleneck caused by overheating in the task execution area, and ensuring better coordination of heat conduction and computing load during resource scheduling. By capturing the stable nodes of the path jump points and cache behaviors, the generated path migration mapping trajectory provides a more refined resource scheduling reference for task allocation, effectively improving the utilization rate of GPU resources and the stability of task execution. Combining the dynamic adjustment of the heat response time and the path compression scheduling label, it can perform dynamic optimization scheduling according to the changes in task load and resource utilization rate, enhance the load balance between GPU computing units, avoid excessive competition and idle of resources, and improve the overall execution efficiency and response speed of the system. The technology of optimizing resource scheduling through heat behavior and path perturbation analysis effectively improves the accuracy and stability of GPU computing power scheduling in a multi-task concurrent execution environment. Brief Description of the Drawings

[0015] Figure 1 is the system flow chart of the present invention; Figure 2 is the flow chart of the thermal expansion monitoring module in the present invention; Figure 3 is the flow chart of the path perturbation identification module in the present invention; Figure 4 is the flow chart of the load migration module in the present invention; Figure 5It is the flowchart of the scheduling label revision module in the present invention; Figure 6 It is the flowchart of the node priority sorting module in the present invention. Specific embodiments

[0016] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0017] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, in the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0018] Please refer to Figure 1 , a GPU computing power dynamic scheduling system includes: The thermal expansion monitoring module obtains the time series records of the GPU core thermal sensors and the motherboard thermal expansion response trajectories, calls the thermal expansion speed, temperature rise duration and diffusion direction path number under the core running state, identifies whether there is information covering the task instruction execution area in the heat conduction during the scheduling period, and generates heat path coverage marking information; Based on the heat path coverage marking information, the path perturbation identification module obtains the path number of the GPU core, the start and end timestamps of the instruction fragments, the memory access instruction sequence and the cache hit time record during the task running period, retrieves whether the path numbers are continuous and extracts the positions where path jump behaviors occur. If the path numbers are not continuous in two adjacent periods and the memory access order is interrupted, the fragment is marked as a perturbed behavior path, and it is judged whether there is an overlap between the path perturbation behavior and the heat path area, and an instruction perturbation thermal coupling linkage structure is generated; The load migration module calls the cache continuous read and write cycles, cache latency behaviors and path jump point sets between the GPU core located by the instruction perturbation thermal coupling linkage structure and the adjacent cores, retrieves whether there are stable path jump sequences and cache hit cycle stable nodes in the adjacent cores. If it is determined as a stable core, it is set as the path migration target, and a core migration path mapping trajectory set is generated; The scheduling label revision module adopts the cross-core execution path behavior sequence composed of the core migration path mapping trajectory set, compares it with the hot response time change behavior and the execution order change trajectory, filters the continuous path structure and the cache shortening record after migration, and generates a path compression scheduling label; The node priority sorting module calls the path compression scheduling label, extracts the path compression node number, cache behavior record and jump sequence stable time, identifies the number of stable path formations and the behavior trend interval, sorts the scheduling nodes according to the path structure coherence, and generates a GPU stable execution path sorting list; The hot path coverage marking information includes the hot expansion path sequence number, the starting boundary of the heat conduction coverage area, and the task path overlap index point. The instruction perturbation thermal coupling linkage structure includes the perturbation path number set, the thermal area coupling mapping index, and the in-cycle jump synchronization mark. The core migration path mapping trajectory set includes the target core path jump sequence, the cross-core path delay trajectory line, and the path mapping section. The path compression scheduling label includes the path compression identification number, the cache usage fallback section, and the sequential behavior continuity mark. The GPU stable execution path sorting list includes the sorted node number, the behavior stability sequence value, and the scheduling priority index code.

[0019] Please refer to Figure 2 , the hot expansion monitoring module includes: The time series extraction sub-module obtains the GPU core thermal sensor time series record and the motherboard hot expansion response trajectory, detects the core operation state parameters and the motherboard hot expansion response values at multiple moments, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal sensitivity value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value; It is necessary to deploy multiple highly sensitive thermal sensor nodes on the GPU chip structure. The sensor nodes continuously sample at intervals of 100 ms, record the temperature readings of each sensor at the current moment, and set four thermal points P1 to P4 within a certain detection period. The data they read are 72.3 °C, 74.1 °C, 75.5 °C, and 73.0 °C in sequence. After the acquisition is completed, the four groups of data are numbered according to the time sequence and classified into the time series structure. At the same time, the core operating state parameters are synchronously collected from the obtained thermal data points, such as the current load value, voltage change rate, instruction throughput rate, etc. It is set that the core current is 3.6 A, the voltage change rate is 0.15 V / s, and the instruction throughput rate is 1800 instructions / s during this period. The parameters and the thermal data are combined into a complete state data vector for this period. After the data set is constructed, it is necessary to take the continuous temperature rise stage within the detection period as the screening target, extract the data segment with continuously rising thermal sensitivity changes, and set it between 100 ms and 300 ms. The temperature of the P2 sensor continuously rises from 74.1 °C to 78.0 °C, then this time period is identified as the temperature rise stage, and it is necessary to calculate the temperature rise value increase and the time period length. The temperature rise amplitude of this segment is 3.9 °C, and the time span is 200 ms. This data is numbered as a "temperature rise duration data point" and bound to the core operating state intensity of this period to form a "temperature rise duration sequence value". During this process, it is also necessary to judge whether the temperature rise change rate exceeds the set thermal sensitivity change threshold. The set thermal sensitivity change threshold is 15 °C / s, and the temperature rise change rate of this segment is 3.9 °C / 0.2 s = 19.5 °C / s, so it is determined that it meets the trigger requirement and is recorded as a valid temperature rise duration data point, completing the integration process of the time series data and obtaining the temperature rise duration sequence value.

[0020] Based on the temperature rise duration sequence value, the diffusion path identification sub-module calls the thermal diffusion response trajectory number and the core working state path within the corresponding period, calculates the unit-time thermal sensitivity change value and the diffusion range value within each trajectory, and screens according to the difference between the diffusion value and the path number to obtain the set of thermal diffusion direction path numbers; Locate the spatial diffusion path number it is in, call the set of recorded main board thermal expansion response trajectory numbers, and set that there are three trajectories numbered R1, R2, and R3 in the previous cycle. Among them, R1 corresponds to the path from the GPU core to the main control chip, R2 corresponds to the path from the GPU core to the south bridge chip, and R3 corresponds to the path from the GPU core to the memory module. For each trajectory path, extract the thermal change value per unit time and the diffusion range value between each point on the path. Set that the temperature rises by 1.8 °C per unit time in path R1, and the diffusion range value is 7.5 mm². The corresponding values in path R2 are 2.1 °C and 8.3 mm², and in path R3 are 1.2 °C and 5.9 mm². Calculate the degree of difference between the thermal expansion value and the path number. Through the calculation of the normalized response rate of the path number, classify the path numbers with a normalized value exceeding 0.6 into the "effective diffusion path" set. Set the normalized response rate values of R2 and R3 to 0.71 and 0.68, both of which exceed the threshold, while R1 is 0.53 and is excluded. Only retain paths R2 and R3 as the set of effective diffusion path sequence numbers. On this basis, establish a one-to-one correspondence between the temperature rise duration sequence value and the path number, and combine the core operation state intensity to evaluate the intensity level of the thermal expansion response on each path. Set the operation state intensity bound to path R2 to 0.88 and that on path R3 to 0.92. Then both paths meet the requirement that the thermal expansion response rate reaches the path thermal expansion response rate threshold of 0.85, and are marked as the propagable path numbers, completing the identification process of the thermal expansion path and obtaining the set of thermal expansion direction path numbers.

[0021] According to the set of thermal expansion direction path numbers, the thermal path coverage marking sub-module matches the execution area corresponding to the task instruction within the task scheduling period, and judges whether the thermal expansion propagation area in the path number set overlaps spatially with the instruction execution position. Use the formula: ; Calculate the thermal path conduction matching degree, analyze the spatial position index and instruction number corresponding to the marked coverage area, and generate thermal path coverage marking information; Among them, is the thermal path conduction matching degree, represents the path number in the task instruction 's duration, represents the spatial displacement difference corresponding to the thermal expansion path number , represents the path number 's corresponding thermal expansion response rate, represents the path number 's core load gain value within the detection period, is the number of path numbers within the detection period; Table 1 Task thermal expansion parameter table: ; As shown in Table 1, the path parameters collected for each task number are the basic participation items of the thermal path conduction matching degree value. By quantifying the basic thermophysical characteristic values, a quantitative index reflecting the coincidence degree of the thermal path and the instruction path can be calculated, and it can guide whether to execute the marker coverage information generation; It is necessary to clarify the specific numbers and task durations of each task in the scheduling period, and set the numbers , and The corresponding durations are 5.2 seconds, 7.1 seconds, and 4.8 seconds respectively. It is necessary to quantify the path displacement difference corresponding to each task . Through the thermal expansion path detection results, the spatial displacement differences are 1.5 mm, 2.1 mm, and 1.2 mm. At the same time, it is necessary to collect the thermal expansion response rate and the core load gain . The obtained rate values are 0.8 °C / s, 1.0 °C / s, 0.7 °C / s, and the load gains are 0.3, 0.4, 0.2. Calculate the task propagation error term under each path number, and use the absolute value form of each item to express the difference between the product of the task duration and the thermal path displacement difference and the square root of the sum of the square of the thermal expansion rate and the load gain, that is: ; Calculate separately as: ; as: ; as: ; Sum the above difference results to obtain the numerator term: ; Directly sum the time, path thermal expansion rate, and load gain of each task to obtain the denominator term for each item as : ; : ; : ; The denominator after summation is: ; Substitute into the formula: ; The result shows that the average matching error intensity of the thermal path conduction in the task instruction path is 1.243, which is used to evaluate the spatial coincidence degree between the execution position of the current scheduling task and the predicted thermal propagation path, reflect whether the task is in the thermal risk area, guide the task to avoid the hot spot area, improve the thermal balance and operation reliability. If the determination threshold of the thermal path conduction matching degree is set to 1.1, since 1.243 > 1.1, it indicates that there is an overlay relationship between the task path and the thermal expansion path, and the spatial position index and identification should be added to the instruction path to generate the thermal path coverage mark information.

[0022] Please refer to Figure 3 , the path perturbation identification module includes: Based on the thermal path coverage mark information, the path number extraction sub-module obtains the path number of the GPU core, the start and end timestamps of the instruction segments, the memory access instruction order and the cache hit time record during the task running period, extracts the path numbers in two adjacent periods, sorts the path numbers and retrieves the continuity of the number sequence, and filters out the path number segments that do not meet the number continuity determination condition to generate the path number interruption interval value; Retrieve the GPU scheduling log, parse the allocation record of the thread blocks in the scheduling information, extract the path number information of each instruction executed in the GPU core in each period, and set the task transfer path numbers in the period as P123 and P124, and in the period as P125 and P126. Each path number corresponds to a certain number of instruction segments, and its start and end timestamps can be marked by accessing the time information of the start and end of the instruction scheduling in the access log. For example, if an instruction segment starts at 10.02ms and ends at 10.31ms, it can be recorded as [10.02, 10.31]. Extract the memory access instruction execution order from the memory access stack record. For example, the access order in this period is instruction 10, instruction 12, instruction 15, and read the cache hit time together with the cache status record. Set the cache hit time of access instruction 12 to 1.3ms, and sort the extracted path number sequence in time order, such as recorded as P123 → P124 in the period. If If the cycle start number is P126, then the number P125 of the middle break is the interrupt number, indicating that there is a path number jump behavior during cycle switching, providing a data basis for subsequent judgment of whether the path is continuous. On this basis, pairwise comparison is performed on all path number sequences to determine whether they are consecutive in number, that is, whether there are situations such as missing numbers and number jumps. For example, if the number sequence should be P123→P124→P125→P126 and the measured result is P123→P124→P126, then it is determined that the P125 path is missing, regarded as a path discontinuity event. In this case, the corresponding path segment is marked as the path number interruption interval, the set of path segments with path number breaks is marked, and recorded as the path number interruption interval value.

[0023] According to the path number interruption interval value, the memory access order retrieval sub-module calls the memory access instruction order and cache hit time record, extracts the memory access fragment corresponding to the path interruption, and compares the memory access order in the fragment with the memory access order of the same-cycle path, using the formula: ; Calculate the memory access order interruption intensity value to obtain the set of path fragments of the disturbance behavior; Among them, represents the memory access order interruption intensity value, represents the position number of the th memory access instruction in the disturbance fragment, represents the position number of the th instruction in the same-cycle comparison fragment, represents the corresponding th cache hit latency value, represents the th path cache hit count, represents the number of memory access instruction entries in the disturbance fragment, represents the number of cache hit record segments in the corresponding path; Identify whether there is a break in the sequence of path numbers within two consecutive cycles. If the ending path number in cycle C1 is 25 and the starting path number in cycle C2 is 28, the interruption interval is considered to be [25, 28]. Further locate the memory access segments between path numbers 25 and 28, arrange the memory access instructions in chronological order, and extract the actual execution position numbers of the instructions in the perturbed path numbers and the comparison position numbers in the normal path numbers respectively. For example, if the position of the memory access instruction numbered 2 in the perturbed path is 25 and the position of this type of instruction in the comparison path is 22, the order offset between the two is 3. Then, extract the offset values of the memory access instructions in the interrupted segment, and calculate the total sum of the instruction position differences through absolute value calculation, which provides a reference for calculating the memory access interruption intensity. Obtain the latency and hit count of each memory access instruction hitting the cache. In the table "Example Table of Memory Access Interruption Behavior", the number of the second memory access instruction in the perturbed path is 25, the cache hit latency is 0.9 ms, and the hit count is 2. Summarize all instructions in turn and substitute them into the formula: Let the position of the perturbed path of memory access instruction 1 be 12, the position of the comparison path be 10, the memory access time be 1.2 milliseconds, and the cache hit count be 3; The position of the perturbed path of memory access instruction 2 is 25, the position of the comparison path is 22, the memory access time is 0.9 milliseconds, and the cache hit count is 2; The position of the perturbed path of memory access instruction 3 is 37, the position of the comparison path is 36, the memory access time is 1.5 milliseconds, and the cache hit count is 4; The position of the perturbed path of memory access instruction 4 is 49, the position of the comparison path is 52, the memory access time is 1.1 milliseconds, and the cache hit count is 1; ; ; ; ; The result shows that there are obvious interruption characteristics in the instruction execution order of the current memory access segment. The memory access order interruption intensity value is a numerical index used to measure the severity of the perturbation of the order during the execution of memory access instructions. By comparing the difference between the original path order and the actual order of the current interrupted segment, it helps to identify problems such as abnormal paths, cache behavior misalignment, and improper scheduling, which belong to the manifestations of perturbation behavior. By integrating the path interruption position difference and the cache latency hit behavior, a quantitative evaluation of the impact degree of path perturbation is realized, which helps to distinguish the boundary between normal fluctuations and abnormal perturbations and can be classified into the set of path segments of perturbation behavior.

[0024] The thermal coupling linkage structure construction sub-module calls the disturbance behavior path segment set, extracts the path jump positions and the thermal path segment numbers, performs position mapping and coincidence interval judgment on the jump positions and the thermal path numbers, and generates an instruction disturbance thermal coupling linkage structure; Import the set of path segment numbers marked as disturbance behaviors into the cache mapping module, scan and extract the path jump positions marked in each segment. If the jump point occurs between instructions 35 and 47 where the path numbers are P124→P126, the jump position is recorded as the instruction number range [35, 47]. Then import the thermal path region marking information, perform position range matching for each thermal path segment number. Set the thermal path segment numbers H003 and H004 to cover the instruction sections [30, 45] and [46, 60] in the path respectively. Then there is an overlap between the jump position range and the thermal path number, and the overlapping part needs to be mapped. Set that instructions 35 to 45 belong to H003, and instructions 46 to 47 belong to H004. That is, this jump behavior partially coincides with both the thermal path segments H003 and H004. Extract the memory access interruption intensity values at each overlapping position point, call the sequence of interruption intensity values calculated previously. Set the intensity corresponding to the position of instruction 35 to be 1.3 and that of instruction 46 to be 1.6. Map each point in this sequence to the thermal path to form a quantitative record of the disturbance influence under the thermal path segment. Integrate the overlapping point position numbers, path jump relationships, and the corresponding memory access interruption intensity values, classify them according to the thermal path number, and further organize them into an instruction disturbance thermal coupling linkage structure.

[0025] Please refer to Figure 4 , the load migration module includes: Based on the instruction disturbance thermal coupling linkage structure, the disturbance detection sub-module extracts the cache continuous read / write cycles, cache latency behaviors, and the set of path jump points between the GPU core and adjacent cores, monitors whether the cache continuous read / write cycles in the set of path jump points of the GPU core are within the cache latency behavior change interval, filters the path jump points contained in the adjacent cores associated with the change interval, and generates a cache latency disturbance intensity measure; Invoke the GPU scheduling control instruction to force the generation of a thermal coupling response on a specific thread scheduling node. At this time, locate the target GPU core and obtain the data exchange status between it and adjacent cores. In this state, by reading the cache access logs of each core, extract the number of read and write operations of the path jump point per unit time, and obtain the cache continuous read and write cycle. For the case where the single jump cycle ranges from 110 to 140 nanoseconds, if the number of read and write operations exceeds 3000 times per second, it is marked as a high-frequency interval of the continuous cycle. In the read latency behavior log, use the cache access time difference as the evaluation index, extract the access latency value corresponding to the path jump point as the cache latency behavior. If the latency value is in the range of 4.5 to 7.5 milliseconds, it is regarded as a medium-high perturbation range. At the same time, combine the access paths that have mutated continuously more than 3 times in the jump point set and mark them as mutation jump points to identify the high-frequency response nodes of the current core under scheduling perturbation. Analyze the above-extracted read and write cycles and latency behaviors corresponding to the jump points to determine whether they constitute periodic perturbation characteristics. If a certain jump point has significant latency fluctuations (such as the jump amplitude exceeds 1.5 milliseconds) in multiple perturbation cycles, it is determined as a perturbation key node, aggregate this node, and identify whether there is a synchronous perturbation phenomenon in adjacent cores, that is, if the adjacent cores have a response fluctuation with a latency amplitude within ±1 millisecond in the same time window, it indicates that this perturbation is a cross-core correlation perturbation, which is classified as a candidate interference path, extract the cache access frequency and cycle offset data corresponding to it, further count the perturbation frequency distribution. If the perturbation frequency of this path exceeds 20% of the total number of jumps per unit time, it is recorded as a strong cache perturbation path. Through the screening of the full path set, output a set of path nodes with clear cache perturbation response characteristics and generate the cache delay perturbation intensity measure.

[0026] The path screening sub-module invokes the cache delay perturbation intensity measure to determine whether there are path jump stable sequences and cache hit cycle stable nodes in adjacent cores. If the hit cycle fluctuation of the node is within the hit cycle jump constraint interval, extract the adjacent cores as candidate path cores and generate a set of path stability values; The synchronous path structure of the involved adjacent cores is identified. By analyzing the jump sequence changes in each adjacent core, the periodic jump frequency of the node in the continuous time slice is recorded. If the fluctuation amplitude of the frequency in three consecutive cycles is less than 5%, it is preliminarily determined to be a stable path jump sequence. The path cycle records of a certain adjacent core are 60, 62, and 61 nanoseconds respectively. The jump amplitude of the sequence is only 2 nanoseconds, which meets the requirements of stable path characteristics. The volatility of the node hit cycle is evaluated. After reading the hit cycle log, the hit count and interval in the corresponding path are extracted, and the hit cycle variation range is calculated. If the range is within ±1.5 milliseconds, it is regarded as a hit cycle stable node. If a certain jump If the hit cycle range in the path is between 59.5 and 61 milliseconds, the stable node setting standard is met. At this time, the adjacent core is recorded as a candidate path core, and the path mapping parameters under the core are extracted, that is, the time sequence and corresponding address mapping information in its node sequence, as well as the cycle parameter set, that is, the hopping frequency distribution corresponding to the stable cycle sequence. The candidate path core information that meets the path stability and hit stability conditions is summarized. In the actual instance, if there is a core A with a hopping frequency of 2%, a hit cycle fluctuation of 1.2 milliseconds, a continuous path length greater than 5 jump units, and all node cycles between 60±2 nanoseconds, it is regarded as a stable path core candidate, and a path stability value set is constructed.

[0027] The core mapping submodule constructs the inter-core migration path and maps the node jump trajectory path sequence according to the path stability value set, using the formula: ; Calculate the migration trajectory offset strength value, evaluate the trajectory mapping continuity of the candidate core, and obtain the core migration path mapping trajectory set; in, represents the migration trajectory offset intensity value, Indicates the jump path The period value of Indicates the path The path continuity, Indicates the path The cache disturbance amplitude is Indicates the path The average jump span, Indicates the path The hit cycle offset, Indicates the path The stable period value of Indicates the hit cycle offset adjustment value. represents the path stability trade-off term, Indicates the path cycle judgment limit; Call path stability value set, which contains 4 candidate paths, numbered P1, P2, P3, and P4 respectively. Each path has 7 key parameters: jump path period , path continuity , cache perturbation amplitude , average jump span , hit cycle offset , stable cycle value , and the data is as follows: The period of path P1 is 120ns, the continuity is 0.85, the cache perturbation is 5.2ms, the span is 30, the hit offset is 2.1, and the stable cycle value is 60; The period of path P2 is 115ns, the continuity is 0.78, the cache perturbation is 6.1ms, the span is 28, the hit offset is 2.3, and the stable cycle value is 62; The period of path P3 is 134ns, the continuity is 0.91, the cache perturbation is 4.8ms, the span is 33, the hit offset is 1.9, and the stable period is 59; The period of path P4 is 112ns, the continuity is 0.69, the cache perturbation is 7.0ms, the span is 25, the hit offset is 2.6, and the stable period is 63; By comparing the cache perturbation amplitudes of each path , initially screen out the path nodes with cache perturbation values higher than the set offset determination limit value of 10ms. Then, based on the square value of the path time continuity and the square root of the sum of the cache perturbation amplitude , take the absolute value of the difference from the jump path period to measure the cycle offset rate of each path under perturbation. Weightedly add the jump span of each node and the hit cycle offset , and sum them up to construct the following denominator expression. At the same time, take the absolute value of the difference between the sum of the stable cycle values and the stable determination threshold and weight it; is the jump cycle of each path node, and the data is recorded by the monitoring point; is the path continuity, which can be obtained by normalizing the reciprocal of the standard deviation of the jump time; is the cache perturbation amplitude, which is calculated from the cache cycle fluctuation; is the average jump span, which can be obtained by the average jump distance between the jump nodes before and after the path; is the hit cycle offset, which is calculated from the deviation between the node hit interval and the ideal hit cycle; is the path stable cycle value, which is measured from the jump sequence cycle; and 、 to set weights and thresholds; Based on the path jump data, numerical substitution calculations are performed as follows: For P1: ; For P2: Similarly ; For P3: ; For P4: ; The sum of the above results gives the numerator part as 471.68, and the denominator part is: ; The total stable period is: ; Therefore, the offset term is: ; Substitute into the formula for calculation: ; This result shows that the offset intensity value under path P4 is the lowest. The offset intensity value of the migration trajectory represents the degree of deviation in aspects such as time, space, cache, and jump between the execution path and the target core mapping trajectory when the task migrates between multiple cores. Therefore, select its node jump path to establish a core migration path mapping trajectory set.

[0028] Please refer to Figure 5 , the scheduling label revision module includes: The path trajectory extraction sub-module extracts the start and end nodes of cross-core migration, the number of path jumps, and the migration frequency value of the execution path based on the execution instruction records and corresponding core numbers in the core migration path mapping trajectory set. Through joint comparison of the migration frequency value and the number of jumps, a cross-core path jump trajectory is generated; Expand the instruction execution sequence in the scheduling log one by one, determine the core number to which each record belongs, and build a task migration path list based on this number. Each path corresponds to a continuous migration process from the source core to the target core. Suppose in a multi-core server environment, task M001 runs on core 1 for half and then is transferred to core 4. Then the starting node of this migration path is core 1, and the ending node is core 4. By traversing the execution instruction segments involved in the path, mark the positions where migration actions occur, record the node numbers of the actions as jump nodes, count the total number of jump nodes in each path, and combine their occurrence frequencies within a unit time to build a data list of the frequency of cross-core jumps. In a certain high-load computing task, path P78 is migrated 5 times within one minute and contains scheduling segments of 3 different cores. This path can be determined as a cross-core migration path. To avoid misidentifying low-frequency noise paths, set an identification criterion for cross-core jumps, and only retain records with more than 3 jumps and containing 2 or more core numbers in the path. The system filters out the paths that meet the conditions from all scheduling paths and outputs their starting and ending nodes, the number of jumps, and the set of core numbers to generate the cross-core path jump trajectory.

[0029] The cache structure screening sub-module calls the continuous path segment number sequence and path interruption positions in the cross-core path jump trajectory. By counting the cache read times and cache hit records of the instruction sets before and after the path segment, comparing the change rate of the read times and the fluctuation range of the hit records, a cache hit fluctuation range is generated. For each paragraph marked as a migration path, read the instruction sets executed before and after the interruption of this path segment. According to the cache interaction records of the instructions during the scheduling process, count the cache read times and hit times respectively generated. Suppose in path segment D41, the system reads 35 cache commands, and the hit record is 28 times. After the path interruption, segment D42 reads 40 cache commands on the migration target core, and the hit record is only 21 times. By comparing the cache call behaviors of the front and back segments, the impact of migration on the cache hit situation can be intuitively judged. The system takes adjacent path segments as units, compares the difference in read times and the floating range of the hit records respectively. During the judgment process, introduce paragraphs with a read difference exceeding a certain proportion and a significant hit difference as the judgment basis. Set that path segments with a read increase of more than 30% and a hit rate decrease of more than 20% will be marked by the system as having significant cache fluctuations. In an example application, the hit times of path segment D58 decrease from 25 to 12 after migration, and the read times increase from 30 to 47. Then this segment is classified as a hit fluctuation path segment. All path segments that meet the above conditions and their interruption positions are recorded and uniformly output to form the cache hit fluctuation range.

[0030] The execution sequence comparison sub-module extracts the sequence structure numbers of the instructions before and after scheduling under the same path according to the path set corresponding to the cache hit fluctuation range, compares the execution cycle value and the start time offset value of the executed instructions, filters out the paths with abnormal offset trends, and obtains the execution sequence offset fluctuation degree; For each path number, extract the execution instruction sequence structure in the two stages before and after scheduling, and mark its start time point and corresponding duration on the time axis. Taking path P91 as an example, the start time of the task before migration is 240 milliseconds and it takes 15 milliseconds to complete; after migration, due to the core scheduling delay, the new execution start time is 270 milliseconds and the duration is 22 milliseconds. Through this time comparison, it can be calculated whether there is an elongation of the execution cycle before and after the migration of this path, as well as the offset amplitude of the scheduling start time. The system further statistically analyzes similar changes in a batch of paths, and forms an offset trend index by accumulating multiple instruction cycle changes and start time differences. If multiple instructions in a certain path show continuous delayed start, increased execution cycle, etc., this path will be judged as an abnormal migration path by the system. In the actual calculation scenario, during the parallel processing of multiple frames of images in the image processing task, if the start time of processing a single frame of image after migration of path P105 is delayed by more than 30 milliseconds and the cycle increase exceeds 10 milliseconds three times in a row, it can be included in the abnormal list. The system summarizes the numbers of all abnormal paths and their corresponding delay directions to obtain the execution sequence offset fluctuation degree.

[0031] The path compression sub-module calls the abnormal path numbers and offset direction attributes in the execution sequence offset fluctuation degree, adjusts the sorting order of the path segments according to the offset direction attribute, and compresses the length value of the continuous jump path segments to obtain the path compression scheduling label; According to the marked delay or advance direction in each path, rearrange the order of its path segment numbers. Suppose a path contains paragraph numbers P300, P302, and P305, where the offset of P300 is 25 milliseconds, P302 is 15 milliseconds, and P305 is 40 milliseconds. Then the new path sorting is P302, P300, P305, so as to compress the redundancy of the execution chain caused by scheduling offsets. On the basis of the new sorting, perform length compression on the continuous jump structures in each path segment, that is, merge the paragraphs that logically have jumps but are highly continuous in execution, reducing the total number of path segment numbers. Suppose the original path contains 8 paragraphs, and after compression, they are merged into 5 paragraphs. The system records the change in the number of paragraphs before and after compression and constructs path index information, including the path number of each segment, the compressed length, the compression direction, and the compression amplitude. In one instance, path P410 is compressed from 9 segments to 6 segments, and the compression amplitude exceeds 30%. The system marks this segment as a valid compression segment and writes it into the index table. Summarize the index contents of all compressed path segments uniformly, and generate a callable path recognition tag set according to multi-dimensional fields such as path number sorting and compression rate sorting, forming a path compression scheduling tag.

[0032] Please refer to Figure 6 , the node priority sorting module includes: Based on the path compression scheduling tag, the path compression extraction sub-module extracts the path compression node numbers, the cache behavior records corresponding to the nodes, and the stable time of the jump sequence. Establish a pairing relationship between the stable time of the jump sequence and the path compression node numbers according to the node order, judge whether the stable time of the jump sequence shows a continuous increasing trend and mark the corresponding numbered paragraphs, and generate a compressed node number interval sequence; Extract and process the path compression node numbers, the cache behavior records corresponding to the nodes, and the stable time of the jump sequence in sequence to obtain the node call records in the path compression scheduling tag. Store the jump association paths between the node numbers in the form of a graph structure. Extract the set of numbers of the nodes connected by each jump edge, and parse the time tag field thereof to extract the stable time parameter value when the jump action occurs. This stable time field is in milliseconds and records the duration of the state of a certain node during adjacent jumps. Set the stable time of the jump between node numbers 301 and 302 to 420 ms, then mark it as the stable duration of the 301→302 path. Based on the depth-first traversal method of the graph structure, pair the stable time fields of the jump sequence with the corresponding node numbers in the topological order of the node numbers to form a set of stable sequence pairs. At the same time, record the cache behavior records before each node jump, including the hit status (hit / miss), the number of write requests, and the replacement operation count, and store them in the corresponding node fields in the form of an array. Set the hit status of node number 305 to 22 hits, 9 write requests, and 5 replacement times. After constructing the pairing relationship between the stable time of the jump sequence and the node number, it is necessary to judge whether the pairing sequence shows an increasing trend. The direction is judged by the difference between adjacent stable times. Let the adjacent two jump times be and , if , then it is marked as increasing. If several consecutive jump sequences satisfy the monotonically increasing stable time, that is, they are classified into a continuously growing paragraph, the number sequence is indexed and marked to form a number section group, and the set of node numbers in the continuously growing interval is generated as the compressed path segment output. Set the stable times in the jump sequence 301→302→303→304 to 420 ms, 440 ms, 470 ms, and 490 ms in sequence, which can be classified into a continuous interval, and the corresponding number set is {301, 302, 303, 304}, generating a compressed node number interval sequence.

[0033] The behavior trend sub-module calls the compressed node number interval sequence, analyzes the cache behavior records corresponding to the node numbers, and divides them according to the synchronous fluctuation characteristics of the hit frequency, write ratio, and replacement rate in the cache behavior records in the time dimension, evaluates the number attribution relationship of various fluctuation trends in the path segment, and generates the behavior change trend interval; Analyze the cache behavior records corresponding to the node numbers, extract the data sets of the hit frequency, write ratio, and replacement rate for each node, construct a behavior record matrix in a multi-dimensional time axis synchronization manner, set the node number as the row index, and the behavior record fields as the column vectors. Suppose the behavior record of node number 305 is 22 hits, 9 write requests, and 5 replacement operations. Then the write ratio is 9 / (22 + 9) ≈ 0.29, and the replacement rate is 5 / (22 + 9 + 5) ≈ 0.14, forming a behavior record vector [0.29, 0.14]. After forming the behavior record matrix, use the time series change trend between the behavior vectors as the classification basis, adopt the sliding time window technology to calculate the co-fluctuation of the behavior changes, define the behavior vector change rate as the difference value of the Euclidean distance between the vectors of adjacent two nodes. If the change rate between consecutive multiple nodes remains within a fixed range Δr, they are classified into the same behavior trend section. Suppose Δr = 0.1. Then if the behavior vector change rate between node i and i + 1 is 0.05, and between i + 1 and i + 2 is 0.07, the three nodes are grouped together. Evaluate the set of numbers in the path segment in turn, delimit the number intervals belonging to various fluctuation trend segments, and generate the behavior change trend intervals.

[0034] The node sorting sub-module filters the node numbers with stable jump sequence features according to the behavior change trend intervals, judges the difference direction of the hit frequency and replacement rate in the cache behavior records corresponding to the node numbers, selects the node numbers with the hit frequency exceeding the replacement rate, and arranges them in order according to the original path position of the nodes to generate a sorted coherent node sequence; Perform screening and judgment operations on the nodes, call the node number and its corresponding jump sequence stable time and include them in the screening range, and calculate the node stable feature coefficient: S = T×(H - R); Where, T is the jump sequence stable time, H is the hit frequency, and R is the corresponding number of replacement rates; If the value of S is positive, it means that the node has a hit-dominant behavior in the stable jump stage. Set T = 420ms, H = 22, R = 5, then: S = 420×(22 - 5) = 7140; Set the node stable feature reference value As the screening threshold, it is set according to the median value of the average stable features of the node behavior samples. If the current node , it is considered to have stable jump behavior characteristics, and the difference between the hit frequency and the replacement rate is further compared to determine whether the hit frequency exceeds the replacement rate. Those that meet the conditions are included in the reserved node number set, and then sorted according to the original numbering order of the nodes in the path compression structure. The original structure order is set to 301→303→305→306, and the selected nodes are 303, 305, and 306, then their order is kept unchanged and arranged as 303→305→306 to generate a sorted coherent node sequence.

[0035] The execution path adjustment submodule calls the sorted coherent node sequence, combines the adjacent structure between the corresponding jump sequence stabilization time and the original node number sequence, determines whether there is a node coherence gap in the path, retains the path segment number with a complete coherent structure, removes the segment number of the abnormal jump node, and integrates the continuous numbering to form a stable path node combination to obtain the GPU stable execution path sorting list; Combined with the pairing structure of the jump sequence stable time field and the node original number sequence, the path continuity gap is judged and the node and The stabilization time is , number difference: ; like and ≥ ; If it is the average value of the jump stability time, it is considered to be a coherent path segment; like >1 or < , then it is a jump abnormal segment, and such node pairs are removed; Set between nodes 301→304 =3, =360ms, =400ms, the segment is judged as an abnormal segment and is not included in the path structure. The node segment group that meets the requirements of continuous numbering and stable jump time is retained to build a stable node combination path set, such as 303→304→305→306, which is a continuous structure and stable jump segment. The complete path segment number set is sequentially integrated and output as a GPU stable execution path sorting list.

[0036] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A GPU computing power dynamic scheduling system, characterized in that: The system comprises: The thermal expansion monitoring module obtains the GPU core thermal sensor time series record and the mainboard thermal expansion response trajectory, calls the thermal expansion speed, temperature rise duration and diffusion direction path number under the core running state, identifies whether there is information about whether heat conduction covers the task instruction execution area within the scheduling cycle, and generates thermal path coverage mark information; The path disturbance identification module searches whether the path number is continuous and extracts the position where the path jump behavior occurs based on the hot path coverage mark information. If the path number is discontinuous in two adjacent cycles and the memory access sequence is interrupted, the marked segment is a disturbance behavior path, and an instruction disturbance thermal coupling linkage structure is generated; The load migration module calls the cache continuous read and write cycle between the GPU core located by the instruction perturbation thermal coupling structure and the adjacent core, searches whether there is a path jump stable sequence and a cache hit cycle stable node in the adjacent core, and generates a core migration path mapping trajectory set; The scheduling label revision module adopts the core migration path mapping trajectory set, compares the thermal response time change behavior and the execution order change trajectory, screens the continuous path structure and the post-migration cache shortening record, and generates a path compression scheduling label.

2. The GPU computing power dynamic scheduling system according to claim 1, characterized in that: The thermal path coverage mark information includes the thermal expansion path sequence number, the starting boundary of the heat conduction coverage area, and the task path overlap index point. The instruction disturbance thermal coupling linkage structure includes a disturbance path number set, a thermal zone coupling mapping index, and an intra-cycle jump synchronization mark. The core migration path mapping trajectory set includes a target core path jump sequence, a cross-core path delay trajectory line, and a path mapping segment. The path compression scheduling label includes a path compression identification number, a cache usage fallback segment, and a sequential behavior continuity identification.

3. The GPU computing power dynamic scheduling system according to claim 1, characterized in that: The thermal expansion monitoring module comprises: The time series extraction submodule obtains the GPU core thermal sensor time series record and the mainboard thermal expansion response trajectory, detects the core operating status parameters and the mainboard thermal expansion response values ​​at multiple times, combines and numbers the read data sequences in chronological order, calculates the increase in the thermal value and the corresponding time period length in the continuous temperature rise stage, and obtains the temperature rise duration sequence value; The diffusion path identification submodule calls the thermal expansion response trajectory number and the core working state path in the corresponding period based on the temperature rise time sequence value, calculates the thermal sensitivity change value and the diffusion range value per unit time in each trajectory, and screens according to the difference between the diffusion value and the path number to obtain the thermal expansion direction path number set; The thermal path coverage marking submodule matches the execution area corresponding to the task instruction in the task scheduling cycle according to the thermal expansion direction path number set, determines whether the thermal expansion propagation area in the path number set overlaps with the instruction execution position in space, calculates the thermal path conduction matching degree, analyzes the spatial position index and instruction number corresponding to the mark coverage area, and generates thermal path coverage marking information.

4. The GPU computing power dynamic scheduling system according to claim 3, characterized in that: The path disturbance identification module comprises: The path number extraction submodule obtains the path number of the GPU core in the task running cycle, the start and end timestamps of the instruction fragment, the memory access instruction sequence and the cache hit time record based on the hot path coverage mark information, extracts the path number in two adjacent cycles, sorts the path number and searches for the continuity of the number sequence, filters the path number fragments that do not meet the number continuity judgment condition, and generates the path number interruption interval value; The memory access sequence retrieval submodule calls the memory access instruction sequence and cache hit time record according to the path number interruption interval value, extracts the memory access fragment corresponding to the path interruption, compares the memory access sequence in the fragment with the memory access sequence of the same cycle path, calculates the memory access sequence interruption intensity value, and obtains the disturbance behavior path fragment set; The thermal coupling linkage structure construction submodule calls the disturbance behavior path segment set, extracts the path jump position and the thermal path segment number, performs position mapping and overlap interval judgment on the jump position and the thermal path number, and generates an instruction disturbance thermal coupling linkage structure.

5. The GPU computing power dynamic scheduling system according to claim 4, characterized in that: The load migration module includes: The disturbance detection submodule extracts the cache continuous read and write cycle, cache delay behavior and path jump point set between the GPU core and the adjacent core based on the instruction disturbance thermal coupling structure, monitors whether the cache continuous read and write cycle in the GPU core path jump point set is in the cache delay behavior change interval, screens the path jump points contained in the adjacent core associated with the change interval, and generates a cache delay disturbance intensity value; The path screening submodule calls the cache delay disturbance strength quantity to determine whether there is a path jump stable sequence and a cache hit cycle stable node in the adjacent core. If the node hit cycle fluctuation is within the hit cycle jump constraint interval, the adjacent core is extracted as a candidate path core to generate a path stability value set. The core mapping submodule constructs the inter-core migration path and maps the node jump trajectory path sequence according to the path stability value set, calculates the migration trajectory offset strength value, evaluates the trajectory mapping continuity of the candidate core, and obtains the core migration path mapping trajectory set.

6. The GPU computing power dynamic scheduling system according to claim 5, characterized in that: The scheduling label revision module includes: The path trajectory extraction submodule extracts the cross-core migration start and end nodes, the path jump number and the migration frequency value of the execution path based on the execution instruction record and the corresponding core number in the core migration path mapping trajectory set, and generates a cross-core path jump trajectory according to a joint comparison between the migration frequency value and the jump number; The cache structure screening submodule calls the continuous path segment number sequence and the path interruption position in the cross-core path jump trajectory, and generates a cache hit fluctuation interval by counting the cache read times and cache hit records of the instruction set before and after the path segment, and comparing the read times change rate with the hit record fluctuation interval; The execution sequence comparison submodule extracts the sequence structure numbers of the instructions before and after the scheduling under the same path according to the path set corresponding to the cache hit fluctuation range, compares the execution cycle value and the start time offset value of the execution instruction, screens the path with abnormal offset trend, and obtains the execution sequence offset fluctuation; The path compression submodule calls the abnormal path number and the offset direction attribute in the execution sequence offset fluctuation, adjusts the path segment sorting order according to the offset direction attribute and compresses the continuous jump path segment length value to obtain the path compression scheduling label.

7. The GPU computing power dynamic scheduling system according to claim 1, characterized in that: The system also includes a node prioritization module: The node priority sorting module calls the path compression scheduling tag, extracts the path compression node number, cache behavior record and jump sequence stability time, identifies the number of stable path formation and behavior trend interval, sorts the scheduling nodes according to the path structure coherence, and generates a GPU stable execution path sorting list; The GPU stable execution path sorting list includes a sorting node number, a behavior stability sequence value, and a scheduling priority index code.

8. The GPU computing power dynamic scheduling system according to claim 7, characterized in that: The node prioritization module includes: The path compression extraction submodule extracts the path compression node number, the node corresponding cache behavior record and the jump sequence stability time based on the path compression scheduling label, establishes a pairing relationship between the jump sequence stability time and the path compression node number according to the node order, determines whether the jump sequence stability time presents a continuous increasing trend and marks the corresponding numbered segment, and generates a compression node number interval sequence; The behavior trend division submodule calls the compressed node number interval sequence, analyzes the cache behavior records corresponding to the node numbers, divides them according to the synchronous fluctuation characteristics between the hit frequency, write ratio and replacement rate in the cache behavior records in the time dimension, evaluates the number attribution relationship of multiple types of fluctuation trends in the path segment, and generates a behavior change trend interval; The node sorting submodule selects the node numbers with stable jump sequence characteristics according to the behavior change trend interval, performs difference direction judgment on the hit frequency and replacement rate in the cache behavior record corresponding to the node number, selects the node numbers whose hit frequency exceeds the replacement rate, and arranges them in order according to the original path position of the node to generate a sorted and coherent node sequence; The execution path adjustment submodule calls the sorted coherent node sequence, combines the adjacent structure between the corresponding jump sequence stabilization time and the original node number sequence, determines whether there is a node coherence gap in the path, retains the path segment number with a complete coherent structure, removes the segment number where the jump abnormal node is located, and integrates the continuous numbering to form a stable path node combination to obtain the GPU stable execution path sorting list.

Citation Information

Patent Citations

  • Host power consumption and temperature balance control method and system, medium and program product

    CN119536958A

  • HBase client main and standby switching method and system based on fault perception

    CN119537484A

  • Monitoring device mesh network systems and methods

    US20080186871A1

  • Devices and methods for thermal management

    WO2018121937A1

Cited By

  • GPU (Graphics Processing Unit) server resource dynamic allocation management method and device, equipment and medium

    CN120723472A

  • GPU server resource dynamic allocation management method, device, equipment and media

    CN120723472B

  • Intelligent network computing power resource management system

    CN120762906A

  • A wisdom network computing power resource management system

    CN120762906B

  • Efficient vehicle monitoring system based on vehicle-road-cloud cooperation

    CN120935534A