Computing power resource cross-platform migration scheduling method and device, equipment and medium
By obtaining user demand instructions in the computing power scheduling, matching computing power resource pools, and adopting GPU passthrough virtualization technology and GPU pooling scheduling algorithm, the problem of difficult cross-platform resource scheduling is solved, and efficient utilization of computing power resources is achieved.
Patent Information
- Application Number
- CN202511112552.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-09
- Publication Date
- 2025-11-18
AI Technical Summary
In existing computing power scheduling technologies, the different programming languages and instructions of various platforms make cross-platform resource scheduling difficult, resulting in some chip clusters being idle or overloaded, causing serious waste of resources and affecting scheduling efficiency.
By acquiring user demand instructions and using task computing power and latency requirements as constraints, the computing power resource pool is matched, resources with small architectural differences and high performance are selected, and dynamic allocation is performed using GPU passthrough virtualization technology and GPU pooling scheduling algorithm to achieve cross-platform migration scheduling.
It improves the efficiency of cross-platform scheduling of computing resources, makes full use of heterogeneous hardware resources, reduces resource waste, and optimizes the dynamic allocation and use of computing resources.
Smart Images

Figure CN120973536A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of computing power scheduling, and particularly relates to a computing power resource cross-platform migration scheduling method and device, equipment and a medium. BACKGROUND
[0002] With the development of computing power scheduling technology, computing power pooling and virtualization technology appears. The core of the technology is to split physical resources into virtual units that can be dynamically allocated through a hardware abstraction layer, supporting multi-task parallel computing.
[0003] In traditional computing power scheduling technology, various types of heterogeneous hardware resources adopt an independent cluster management mode. An enterprise deploys a chip cluster to run training for a server, and builds another chip cluster on other architecture devices to process compatibility verification. Each chip cluster sets up an independent scheduling system, and after manually estimating the task demand, the task demand is allocated to a specified hardware partition.
[0004] However, due to the differences in programming languages and instructions of various platforms, when the instruction set type of a sudden task request does not match the current idle resources, the system cannot schedule resources across platforms, resulting in long-term idling of some chip clusters and continuous overloading of some chip clusters. This phenomenon causes waste of tens of millions of computing power assets, and for resource allocation such as GPU, regardless of the size of the demand, only the whole GPU card can be allocated and used, causing more serious performance idling and restricting the work efficiency of computing power resource scheduling. SUMMARY
[0005] Therefore, it is necessary to provide a computing power resource cross-platform migration scheduling method, device, equipment and medium capable of improving the work efficiency of computing power resource scheduling.
[0006] In a first aspect, the application provides a computing power resource cross-platform migration scheduling method, comprising:
[0007] obtaining a user demand instruction, the user demand instruction including a task computing power demand instruction, a task time delay demand instruction and a task instruction set;
[0008] matching the task instruction set with a preset computing power resource pool as a constraint condition of the task computing power demand instruction and the task time delay demand instruction to obtain a first computing power resource pool, the first computing power resource pool including computing power resources meeting the task computing power demand instruction and the task time delay demand instruction;
[0009] when the first computing power resource pool meets a preset instruction conversion requirement, selecting a task computing power resource from the first computing power resource pool, and converting the task instruction set according to the task computing power resource to obtain a standard instruction set; the resource architecture type of the task computing power resource is different from the resource architecture type of the task instruction set;
[0010] According to the standard instruction set, the task computing power resource is dynamically allocated to obtain a computing power allocation resource result, and the standard instruction set is processed based on the computing power allocation resource result to obtain task feedback information.
[0011] Further, the task instruction set is matched with the preset computing power resource pool under the constraint conditions of the task computing power demand instruction and the task time delay demand instruction to obtain a first computing power resource pool, including:
[0012] For each computing power resource in the computing power resource pool, the following formula is used to calculate the first load score corresponding to the computing power resource:
[0013]
[0014] Wherein, S(v i ) is the first load score of the computing power resource v i , α is the performance weight, C i is the peak computing power of the computing power resource v i , R c is the task computing power demand instruction, δ v is the virtualization overhead coefficient, β is the time delay weight, T c is the task time delay demand instruction, L i is the task queuing time delay, τ is the startup time consumption, Δ is the transmission time delay, γ is the storage weight, S f is the idle video memory of the computing power resource v i , S t is the total video memory of the computing power resource v i , δ is the architecture weight, is the architecture difference score of the computing power resource v i ;
[0015] Wherein, the architecture difference score is calculated by the following formula:
[0016]
[0017] Wherein, is the architecture difference score of the computing power resource v i , ω1 is the bit width weight, W d(x,i) is the data bit width difference value between the computing power resource v i and the task instruction set, ω2 is the vector unit weight, U d(x,i) is the vector processing unit difference value between the computing power resource v i and the task instruction set, ω3 is the hardware weight, H d(x,i) is the hardware compatibility difference value between the computing power resource v i and the task instruction set, ω4 is the conversion weight, C (x,i) is the conversion index of the computing power resource vi and the task instruction set.
[0018] selecting computing power resources in the first load scoring pool greater than a preset first load threshold to obtain a first computing power resource pool.
[0019] Further, when the first computing power resource pool meets a preset instruction conversion requirement, a task computing power resource is selected from the first computing power resource pool, and the task instruction set is converted according to the task computing power resource to obtain a standard instruction set, including:
[0020] According to the order of increasing architecture difference scores, computing power resources are selected from the first computing power resource pool until the task computing power requirement instruction is met, and a second computing power resource pool is constructed based on the selected computing power resources;
[0021] According to the second computing power resource pool, the task instruction set is converted to obtain a corresponding standard instruction set;
[0022] The conversion performance loss of the task instruction set and each standard instruction set is calculated to obtain the performance loss rate corresponding to each computing power resource;
[0023] According to the order of increasing performance loss rate, computing power resources are selected from the second computing power resource pool, and the selected computing power resources are configured as task computing power resources;
[0024] The instruction conversion requirement is that the first computing power resource pool includes computing power resources whose resource architecture types are different from the resource architecture types of the task instruction set.
[0025] Further, according to the second computing power resource pool, the task instruction set is converted to obtain a corresponding standard instruction set, including:
[0026] According to a preset instruction division rule, the task instruction set is split to obtain a basic instruction set; the basic instruction set is used to represent a set of basic instructions;
[0027] According to the resource architecture type of the computing power resource in the second computing power resource pool and the task instruction set, a corresponding instruction conversion rule is selected from a preset mapping rule set;
[0028] Each basic instruction in the basic instruction set is traversed, and a corresponding conversion instruction is determined from the instruction conversion rule;
[0029] According to the conversion instruction, the basic instruction is converted to obtain a corresponding conversion atomic instruction;
[0030] All conversion atomic instructions are combined in the traversal order to obtain the standard instruction set.
[0031] Further, the conversion performance loss of the task instruction set and each standard instruction set is calculated to obtain the performance loss rate corresponding to each computing power resource, including:
[0032] According to the preset historical task data, instructions matched with the historical task data are selected from the task instruction set to obtain an initial instruction type set;
[0033] Based on the initial instruction type set, the occurrence frequency of each instruction in the initial instruction type set is calculated to obtain a weighted weight of each instruction;
[0034] Based on the task instruction set, the original cycle instruction number of each instruction is calculated;
[0035] Based on each standard instruction set, the conversion cycle instruction number of each instruction in each standard instruction set is calculated;
[0036] Based on the original cycle instruction number of each instruction and the conversion cycle instruction number in each standard instruction set, the performance loss rate corresponding to each computing resource is calculated.
[0037] Further, the task computing resources are dynamically allocated according to the standard instruction set to obtain a computing resource allocation result, and the standard instruction set is processed based on the computing resource allocation result to obtain task feedback information, including:
[0038] According to the task computing demand instruction and the virtualization overhead coefficient, the computing compensation computing data is calculated;
[0039] According to the task computing resource, the computing compensation computing data is adjusted to obtain virtual computing compensation computing data;
[0040] The GPU pass virtual algorithm and the GPU pooling scheduling algorithm are adopted, and the task computing resources are dynamically allocated according to the virtual computing compensation computing data to obtain a computing resource allocation result;
[0041] The standard instruction set is processed by using the computing resource allocation result to obtain the task feedback information.
[0042] Further, the method further comprises:
[0043] The task feedback information is migrated to a preset storage space, and the data of the computing resource allocation result is cleared to obtain a cleared computing resource allocation;
[0044] The cleared computing resource allocation is added to the computing resource pool to obtain an updated computing resource pool.
[0045] In a second aspect, the application also provides a computing resource cross-platform migration scheduling device, comprising:
[0046] An instruction acquisition module is configured to acquire user demand instructions, the user demand instructions including task computing demand instructions, task time delay demand instructions, and a task instruction set;
[0047] The first matching module is configured to match the task instruction set with a preset computing power resource pool under the constraint conditions of the task computing power requirement instruction and the task time delay requirement instruction, and obtain a first computing power resource pool, wherein the first computing power resource pool comprises computing power resources meeting the task computing power requirement instruction and the task time delay requirement instruction.
[0048] The second matching module is configured to select a task computing power resource from the first computing power resource pool when the first computing power resource pool meets preset instruction conversion requirements, and convert the task instruction set according to the task computing power resource, to obtain a standard instruction set; the resource architecture type of the task computing power resource is different from the resource architecture type of the task instruction set.
[0049] The task completion module is configured to dynamically allocate the task computing power resource according to the standard instruction set, to obtain a computing power allocation resource result, process the standard instruction set based on the computing power allocation resource result, and obtain task feedback information.
[0050] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any of the computing power resource cross-platform migration scheduling methods according to the first aspect of the present application when executing the computer program.
[0051] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement any of the computing power resource cross-platform migration scheduling methods according to the first aspect of the present application.
[0052] The above computing power resource cross-platform migration scheduling method, device, computer device and storage medium can obtain a user requirement instruction, wherein the user requirement instruction comprises a task computing power requirement instruction, a task time delay requirement instruction and a task instruction set; match the task instruction set with a preset computing power resource pool under the constraint conditions of the task computing power requirement instruction and the task time delay requirement instruction, to obtain a first computing power resource pool, wherein the first computing power resource pool comprises computing power resources meeting the task computing power requirement instruction and the task time delay requirement instruction; select a task computing power resource from the first computing power resource pool when the first computing power resource pool meets preset instruction conversion requirements, and convert the task instruction set according to the task computing power resource, to obtain a standard instruction set; the resource architecture type of the task computing power resource is different from the resource architecture type of the task instruction set; dynamically allocate the task computing power resource according to the standard instruction set, to obtain a computing power allocation resource result, process the standard instruction set based on the computing power allocation resource result, and obtain task feedback information, thereby realizing cross-platform dynamic allocation of computing power resources and improving the work efficiency of computing power resource scheduling. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0054] Figure 1 A flowchart of a computing resource cross-platform migration scheduling method provided by an embodiment of the present application is shown in the figure.
[0055] Figure 2 A structural diagram of a computing resource cross-platform migration scheduling device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will further describe the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain the present application, and are not used to limit the present application.
[0057] In one embodiment, as shown in Figure 1 , a computing resource cross-platform migration scheduling method is provided, and the present embodiment takes the method applied to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is realized through the interaction of the terminal and the server. In the present embodiment, the method includes the following S101-S104, wherein:
[0058] S101, obtaining a user demand instruction, the user demand instruction including a task computing power demand instruction, a task time delay demand instruction and a task instruction set.
[0059] Specifically, the terminal obtains a user demand instruction, the user demand instruction including a task computing power demand instruction, a task time delay demand instruction and a task instruction set. The task computing power demand instruction is used to indicate how much standardized TFLOPS computing power is needed to complete the task; the task time delay demand instruction is used to indicate the maximum response time for completing the task, including computing resource matching time, transmission delay and processing time of the computing resource processing the task; the task instruction set refers to the task program instruction set and the type of hardware architecture written. The terminal obtains the above three instructions, and identifies the data indicated by the instructions, so as to schedule computing resources satisfying the above three instructions to process the task.
[0060] S102, matching the task instruction set with a preset computing resource pool as a constraint condition of the task computing power demand instruction and the task time delay demand instruction, to obtain a first computing resource pool, the first computing resource pool including computing resources satisfying the task computing power demand instruction and the task time delay demand instruction.
[0061] Specifically, the terminal selects, as a constraint condition, the task computing power requirement instruction and the task time delay requirement instruction, selects computing power resources meeting the task computing power requirement instruction and the task time delay requirement instruction from a preset computing power resource pool, and forms a first computing power resource pool. Illustratively, the preset computing power resource pool can be constructed according to computing power resource hardware that is completed in actual work registration connection, and can include computing power resources of multiple different resource architecture types. Illustratively, the computing power resource is a set of hardware computing units that can be dynamically scheduled and allocated, is a standardized and measurable service unit that abstracts parallel processing capability of a physical chip, includes bound memory, storage and network bandwidth, and can be allocated to complete task required computing.
[0062] S103, when the first computing power resource pool meets preset instruction conversion requirements, selecting a task computing power resource from the first computing power resource pool, and converting a task instruction set according to the task computing power resource to obtain a standard instruction set; the resource architecture type of the task computing power resource is different from the resource architecture type of the task instruction set.
[0063] Specifically, when the first computing power resource pool includes computing power resources of different resource architecture types from the task instruction set, a task computing power resource is selected from the first computing power resource pool as a selection standard of smaller architecture difference and better performance consumption, and a task instruction set is converted according to the resource architecture type of the task computing power resource to obtain a standard instruction set, so that the task can be processed under the resource architecture type of the task computing power resource.
[0064] S104, dynamically allocating the task computing power resource according to the standard instruction set to obtain a computing power allocation resource result, processing the standard instruction set based on the computing power allocation resource result to obtain task feedback information.
[0065] Specifically, the terminal dynamically allocates the task computing power resource according to the standard instruction set to obtain a computing power allocation resource result, and transmits the standard instruction set to the corresponding computing power allocation resource result for processing to obtain task feedback information. Wherein, the GPU pass-through virtual technology is used to virtualize the hardware corresponding to the task computing power resource, and the GPU pooling scheduling algorithm is used to dynamically cut and allocate the task computing power resource, so that the computing power resource is fully used. Optionally, the task feedback information includes but is not limited to a processing result obtained by the computing power allocation resource result according to the task of the task instruction set, such as a model parameter training result or a data calculation result of data training.
[0066] The power resource cross-platform migration scheduling method provided by the embodiment, by obtaining the task power demand instruction, the task time delay demand instruction and the task instruction set of the user, selecting the power resource based on the task power demand instruction and the task time delay demand instruction, for the power resource with different resource architecture types and task instruction sets, taking the smaller architecture difference and the better performance consumption as the selection standard, selecting the better heterogeneous power resource, determining as the task power resource, and converting the task instruction set to obtain the standard instruction set; based on the standard instruction set, using the GPU pass-through virtual technology to dynamically allocate the task power resource, obtaining the power allocation resource result, and processing the standard instruction set based on the power allocation resource result, obtaining the task feedback information, realizing the cross-platform scheduling power resource, improving the scheduling efficiency between different architecture power hardware, and realizing the dynamic power resource cutting allocation, improving the work efficiency of the power resource scheduling.
[0067] In one of the embodiments, the task instruction set is matched with the preset power resource pool with the task power demand instruction and the task time delay demand instruction as the constraint condition, to obtain the first power resource pool, including:
[0068] S201, for each power resource in the power resource pool, the following formula is used to calculate the first load score corresponding to the power resource:
[0069]
[0070] Wherein, S(v i ) is the first load score of the power resource v i , α is the performance weight, C i is the peak power of the power resource v i , R c is the task power demand instruction, δ v is the virtualization overhead coefficient, β is the time delay weight, T c is the task time delay demand instruction, L i is the task queuing time delay, τ i is the startup time, Δ is the transmission time delay, γ is the storage weight, S f is the idle video memory of the power resource v i , S t is the total video memory of the power resource v i , δ is the architecture weight, is the architecture difference score of the power resource v i ;
[0071] Wherein, the architecture difference score is calculated by the following formula:
[0072]
[0073] Wherein, It is computing power resources v i The architectural difference score, ω1 is the bit width weight, W d(x,i) It is computing power resources v i The difference in data bit width between the task instruction set and the data bit width, ω2 is the vector unit weight, U d(x,i) It is computing power resources v i The difference between the vector processing unit and the task instruction set, ω3 is the hardware weight, H d(x,i) It is computing power resources v i Hardware compatibility difference with the task instruction set, ω4 is the transformation weight, C (x,i) It is the conversion index between computing power resources (vi) and task instruction sets.
[0074] Specifically, for each computing resource in the computing resource pool, the aforementioned formula is used to quantify the matching degree between the computing resource and the task through four dimensions: computing performance, latency, storage, and architectural differences, thus calculating the first load score corresponding to the computing resource. Illustratively, S(v i ) is computing power resource v i The first load score is used to determine the matching degree between computing resources and tasks. A higher first load score indicates a higher degree of matching between computing resources and tasks; conversely, a lower first load score indicates a lower degree of matching between computing resources and tasks. Optionally, computing resources v i Peak computing power C i Standardized TFLOPS computing power is used to characterize computing resources. Schematic, the virtualization overhead coefficient δ... v It is computing power resources v i Performance overhead of virtualization. Optionally, task queuing latency L. i It is computing power resources v i Current task queue waiting time. Indicatively, startup time τ. i It is computing power resources v i The initialization delay. Optionally, the transmission delay Δ is the network transmission time. Illustratively, computing resources v i Free video memory S f It is computing power resources v i Real-time available video memory. Optionally, computing resources v i Total video memory S t It is computing power resources v i Total video memory. Indicatively, computing resources v i Architectural Difference Score This is used to characterize the differences in task instruction sets and computing resource hardware architecture. Optionally, performance weight α, latency weight β, storage weight γ, and architecture weight δ represent the focus on four dimensions: computing performance, latency, storage, and architecture differences, respectively. These weights can be set according to the specific needs of the work, where the sum of α, β, γ, and δ is 1. Illustratively, computing resources vi and the data bit width difference value W of the task instruction set d(x,i) is the absolute value of the instruction set bit width difference of the task instruction set and the computing resource, which can be extracted according to the architecture manual. Optionally, the computing resource v i and the vector processing unit difference value U of the task instruction set d(x,i) is used to represent the resource architecture vector processing hardware capability difference of the task instruction set and the computing resource. Illustratively, the computing resource v i and the hardware compatibility difference value H of the task instruction set d(x,i) is used to represent the hardware underlying compatibility of the computing resource and the task instruction set. Optionally, the conversion index C of the computing resource vi and the task instruction set (x,i) is used to represent the instruction conversion degree of the computing resource v i and the software level of the task instruction set. If the instructions of the task instruction set can all be converted to the hardware architecture of the computing resource v i , C (x,i) is 0, otherwise C (x,i) is positive infinity. Illustratively, the bit width weight ω1, the vector unit weight ω2, the hardware weight ω3 and the conversion weight ω4 respectively represent the attention tendency of the data bit width difference value, the vector unit difference value, the hardware compatibility difference value and the conversion coefficient, which can be set according to the attention to the four dimensions in actual work, and the sum of ω1, ω2, ω3 and ω4 is 1.
[0075] S202, selecting a computing resource with a first load score greater than a preset first load threshold from a computing resource pool to obtain a first computing resource pool.
[0076] Specifically, based on the obtained first load scores of the computing resources, the computing resources with the first load scores greater than the preset first load threshold are selected from the computing resource pool to form the first computing resource pool. The first load threshold can be set according to the requirements of the four dimensions of computing performance, latency, storage and architecture difference in actual work.
[0077] The computing resource cross-platform migration scheduling method provided in this embodiment takes the task computing demand instruction and the task latency demand instruction of a task as constraint conditions, evaluates the load of each computing resource in the preset computing resource pool from the four dimensions of computing performance, latency, storage and architecture difference, obtains a first load score, selects a computing resource with a first load score greater than a first load threshold, and forms a first computing resource pool, thereby reducing the cases of resource over-provisioning and architecture blind selection, avoiding the scheduling range of inefficient resources, and effectively improving the work efficiency of computing resource scheduling.
[0078] In one of the embodiments, when the first computing resource pool meets the preset instruction conversion requirement, the task computing resource is selected from the first computing resource pool, and the task instruction set is converted according to the task computing resource, to obtain a standard instruction set, including:
[0079] S301, selecting the computing resource from the first computing resource pool in the order of increasing architecture difference score until the task computing resource requirement instruction is met, and constructing the second computing resource pool based on the selected computing resource.
[0080] Specifically, according to the obtained first computing resource pool, the computing resources in the first computing resource pool are sorted in the order of increasing architecture difference score, and the computing resource whose performance meets the task computing resource requirement instruction is selected to construct the second computing resource pool.
[0081] S302, converting the task instruction set according to the computing resources in the second computing resource pool to obtain the corresponding standard instruction set.
[0082] Specifically, based on the computing resources contained in the second computing resource pool, the conversion rule is obtained according to the resource architecture of each computing resource and the task instruction set, the task instruction set is converted to obtain the standard instruction set consistent with the architecture of the computing resource. Illustratively, the performance loss caused by the conversion of the standard instruction set is used to select the computing resource for processing the task instruction set.
[0083] S303, calculating the conversion performance loss of the task instruction set and each standard instruction set to obtain the performance loss rate corresponding to each computing resource.
[0084] Specifically, the terminal selects the high-frequency instruction from the task instruction set, and calculates the number of instructions per cycle of the high-frequency instruction in the preset resource architecture of the task instruction set and the corresponding resource architecture of each standard instruction set through a formula, and quantifies the conversion performance loss of the task instruction set converted to each standard instruction set according to the ratio of the number of instructions per cycle, to obtain the performance loss rate corresponding to each computing resource architecture. Illustratively, the performance loss rate is used to evaluate the performance attenuation of the converted task instruction set. Optionally, the number of instructions per cycle is used to represent the number of instructions completed per clock cycle, and the larger the value, the higher the performance, which can be obtained from the chip manual or measured in the laboratory.
[0085] S304, selecting the computing resource from the second computing resource pool in the order of increasing performance loss rate, and configuring the selected computing resource as the task computing resource.
[0086] Specifically, according to the order of performance loss rate increment, the computing power resources in the second computing power resource pool are arranged, and the computing power resource corresponding to the minimum performance loss rate is preferentially selected as the task computing power resource for processing the task instruction set. If there are computing power resources with the same performance loss rate and the same minimum value, the computing power resource corresponding to the more optimal architecture difference score and the first load score is selected.
[0087] S305, wherein the instruction conversion requirement is that the computing power resource in the first computing power resource pool includes a resource architecture type different from the resource architecture type of the task instruction set.
[0088] Specifically, the instruction conversion requirement is that the computing power resource in the first computing power resource pool includes a resource architecture type different from the resource architecture type of the task instruction set, so as to convert the instruction set and match the heterogeneous computing power resources.
[0089] The method provided by the embodiment of the application comprises the following steps: selecting, in the order of increment of the architecture difference score, computing power resources satisfying the task computing power requirement instruction and having a more optimal architecture difference score from a first computing power resource pool to form a second computing power resource pool; converting the task instruction set according to the resource architecture of each computing power resource in the second computing power resource pool and calculating the performance loss rate caused by the conversion of each computing power resource; selecting the computing power resource based on the loss rate and the task computing power requirement instruction to obtain a task computing power resource for processing the task instruction set; and taking the more optimal performance loss across resource architectures as the selection standard to provide a more suitable computing power resource for processing the task of cross-platform computing power scheduling, thereby improving the work efficiency of computing power resource scheduling.
[0090] In one of the embodiments, the task instruction set is converted according to each computing power resource in the second computing power resource pool to obtain a corresponding standard instruction set, which comprises:
[0091] S401, the task instruction set is split according to a preset instruction division rule to obtain a basic instruction set; the basic instruction set is used to represent a set of basic instructions.
[0092] Specifically, the terminal splits the task instruction set according to a preset instruction division rule to obtain each basic instruction included in the task instruction set, and the basic instructions are combined according to the execution order to obtain a basic instruction set; illustratively, the instruction division rule can be set according to the function unit or the operation code type in actual work, such as splitting the “matrix multiplication CUDA kernel function” into “loading register data”, “parallel computing”, and “writing back to the memory” basic instructions.
[0093] S402, the corresponding instruction conversion rule is selected from a preset mapping rule set according to the resource architecture type of the computing power resource in the second computing power resource pool and the task instruction set.
[0094] Specifically, the terminal selects a corresponding instruction conversion rule from the preset mapping rule set according to the resource architecture type of the computing resource in the second computing resource pool and the resource architecture of the task instruction set. Illustratively, the mapping rule set is obtained by counting each computing hardware architecture actually included in the work and designing the instruction conversion rule between each resource architecture, such as the vector permutation instruction in the common CUDA and ARM architecture, which constitutes one of the conversion rules of shfl sync()-NEON. When used, the conversion rule between the two resource architectures can be obtained by searching the resource architecture type of the computing resource and the resource architecture of the task instruction set.
[0095] S403, traversing each basic instruction in the basic instruction set, determining the corresponding conversion instruction from the instruction conversion rule.
[0096] Specifically, the terminal traverses each basic instruction in the basic instruction set, and confirms the corresponding conversion instruction according to the instruction conversion rule. Illustratively, the conversion instruction is used to indicate the conversion of the basic instruction into the conversion atomic instruction of the corresponding resource architecture.
[0097] S404, converting the basic instruction according to the conversion instruction to obtain the corresponding conversion atomic instruction.
[0098] Specifically, the terminal converts each basic instruction based on the conversion instruction, and adjusts according to the resource architecture of the computing resource to obtain the corresponding conversion atomic instruction. Illustratively, for the vector multiplication instruction, if the resource architecture of the computing resource does not satisfy the bit number of the basic instruction, the basic instruction is disassembled to meet the requirements of the computing resource. Optionally, the operation target register is checked, and if it does not exist, a return mechanism is triggered to reselect the computing resource to convert the task instruction set.
[0099] S405, combining all the conversion atomic instructions in the traversal order to obtain the standard instruction set.
[0100] Specifically, the terminal combines all the conversion atomic instructions in the traversal order, preserves the dependency relationship of the basic instructions, and optimizes the sequence conflict through the instruction scheduler to obtain the standard instruction set corresponding to any computing resource.
[0101] The computing resource cross-platform migration scheduling method provided in this embodiment disassembles the task instruction set to obtain the basic instruction set, selects the instruction conversion rule from the preset mapping rule set, converts each basic instruction in the basic instruction set based on the instruction conversion rule, and performs reasonable processing of instruction execution to obtain the standard instruction set corresponding to the computing resource, realizes the task instruction set conversion across the resource architecture platform, fully utilizes the computing resources of each resource architecture, and effectively improves the work efficiency of the computing resource scheduling.
[0102] In one of the embodiments, the translation performance loss of the task instruction set and each standard instruction set is calculated to obtain the performance loss rate corresponding to each computing resource, including:
[0103] S501, according to the preset historical task data, selecting instructions matched with the historical task data from the task instruction set to obtain an initial instruction type set.
[0104] Specifically, the terminal selects instructions matched with the historical task data from the task instruction set based on the preset historical task data to obtain an initial instruction type set. Illustratively, the historical task data can be recorded according to high-frequency instructions used in actual work, such as floating-point multiply-add (FMA) and convolution (CONV), so as to measure the performance loss rate of the overall translation of the task instruction set by calculating the translation efficiency of the high-frequency instructions.
[0105] S502, based on the initial instruction type set, calculating the occurrence frequency of each instruction in the initial instruction type set to obtain the weighted weight of each instruction.
[0106] Specifically, according to the initial instruction type set, the occurrence frequency of each instruction in the initial instruction type set is calculated, and the total number of occurrences of all instructions is calculated. For each instruction, the occurrence frequency of the instruction is divided by the total number of occurrences of all instructions to obtain the weighted weight of the instruction.
[0107] S503, based on the task instruction set, calculating the original cycle instruction number of each instruction.
[0108] Specifically, for each instruction in the task instruction set, the number of instructions completed per cycle of the instruction on the resource architecture corresponding to the task instruction set is obtained to obtain the original cycle instruction number. Illustratively, the original cycle instruction number can be calculated by querying a technical manual or actual measurement. Alternatively, the original cycle instruction number is used to represent the number of instructions completed per clock cycle of the instruction under the preset resource architecture type of the task instruction set, and the larger the value, the higher the performance.
[0109] S504, based on each standard instruction set, calculating the translation cycle instruction number of each instruction in each standard instruction set.
[0110] Specifically, for each instruction in the task instruction set, the number of instructions completed per cycle of the instruction on the resource architecture corresponding to the standard instruction set is obtained to obtain the translation cycle instruction number. Illustratively, the translation cycle instruction number can be calculated by querying a technical manual or actual measurement. Alternatively, the translation cycle instruction number is used to represent the number of instructions completed per clock cycle of the instruction under the resource architecture type of the computing resource corresponding to each standard instruction set, and the larger the value, the higher the performance.
[0111] S505, calculate the performance loss rate corresponding to each computing resource based on the original cycle instruction number of each instruction and the converted cycle instruction number of each standard instruction set.
[0112] Specifically, based on the obtained weighted weight, original cycle instruction number and converted cycle instruction number of each instruction, the conversion loss of each instruction is obtained by dividing the original cycle instruction number by the converted cycle instruction number, and multiplied by the corresponding weighted weight, and the conversion loss of each instruction is weighted and combined to obtain the performance loss rate corresponding to the conversion of the task instruction set on each computing resource.
[0113] The computing resource cross-platform migration scheduling method provided in this embodiment confirms the initial instructions appearing in the task instruction set through the preset historical task data, determines the weighted weight of each high-frequency instruction according to the number of occurrences of each initial instruction and the total number of instruction occurrences, and obtains the performance loss rate corresponding to the conversion of the task instruction set on the computing resource by calculating the cycle occurrence number of each initial instruction on the preset resource architecture and the cycle occurrence number after conversion, and by weighting and combining, and by calculating the performance loss rate through the cycle occurrence number of the high-frequency initial instruction set, the calculation redundancy caused by the calculation of low-frequency instructions is reduced.
[0114] In one of the embodiments, the task computing resource is dynamically allocated according to the standard instruction set to obtain a computing resource allocation result, and the standard instruction set is processed based on the computing resource allocation result to obtain task feedback information, including:
[0115] S601, calculate the computing power compensation computing power data according to the task computing power demand instruction and the virtualization overhead coefficient.
[0116] Specifically, the terminal calculates the performance loss caused by the hardware virtualization of the computing resource according to the task computing power demand instruction and the virtualization overhead coefficient, and obtains the compensated computing power compensation computing power data to be sufficient to process the task.
[0117] S602, adjust the computing power compensation computing power data according to the task computing resource to obtain the virtual computing power compensation computing power data.
[0118] Specifically, the terminal obtains the idle computing power of the computing resource according to the actual load of the computing resource hardware, and if the idle computing power is not enough to process the task corresponding to the computing power compensation computing power data, the terminal dynamically fine-tunes the computing power compensation computing power data, sets the compensation coefficient through the linear interpolation method, and multiplies the compensation coefficient and the computing power compensation computing power data to obtain the virtual computing power compensation computing power data. Illustratively, the adjustment process refers to the historical resource fluctuation data to ensure that the compensation data is highly consistent with the actual availability.
[0119] S603, the GPU pass virtualization algorithm and the GPU pooling scheduling algorithm are adopted, the task computing power resource is dynamically allocated according to the virtual computing power compensation computing power data, and a computing power allocation resource result is obtained.
[0120] Specifically, the terminal adopts the GPU pass virtualization technology and the GPU pooling scheduling algorithm, dynamically allocates the task computing power resource according to the virtual computing power compensation computing power data, and directly binds the hardware architecture of the computing power resource to obtain the computing power allocation resource result. Illustratively, the GPU pass virtualization technology is a physical GPU mapping technology, which directly maps the physical GPU device to the virtual machine by bypassing the virtualization intermediate layer, realizes zero-intermediate-layer access of the virtual machine to the GPU hardware, and the specific process can be that a privileged IOMMU mapping domain is created in the host server, the PCIe device space (register, video memory, DMA engine) of the physical GPU is directly projected to the virtual machine, and the virtual machine directly controls the GPU hardware through the native driver, eliminating the virtualization software overhead and ensuring that the cloud host obtains not less than 95% of the computing performance of the physical GPU. Alternatively, the GPU pooling scheduling algorithm is a computing power unit splitting and merging algorithm, which abstracts the GPU into standardized computing power units and can select a plurality of standardized computing power units to combine into an arbitrary computing power resource to process tasks by using the resource grid and dynamic roaming mechanism.
[0121] S604, the standard instruction set is processed by using the computing power allocation resource result, and task feedback information is obtained.
[0122] Specifically, the standard instruction set is transmitted to the computing power allocation resource result for processing, and task feedback information is obtained. Illustratively, the task feedback information includes but is not limited to the task processing result.
[0123] The computing power resource cross-platform migration scheduling method provided in this embodiment compensates the computing power requirement of the computing power resource according to the task computing power demand instruction and the virtualization overhead coefficient, obtains the compensated computing power compensation computing power data, adjusts the compensated computing power compensation computing power data according to the actual idle computing power of the computing power resource, obtains the virtual computing power compensation computing power data, adopts the GPU pass virtualization algorithm and the GPU pooling scheduling algorithm, dynamically cuts and allocates the hardware architecture of the computing power resource according to the virtual computing power compensation computing power data, and directly connects to obtain the computing power allocation resource result. The task feedback information is obtained by processing the representation instruction set based on the computing power allocation resource result, the cross-platform computing power resource is fully utilized, the idle computing power of the computing power resource is used for dynamic allocation of computing power, the task is processed based on the dynamically allocated computing power resource, and the working efficiency of the computing power resource scheduling is improved.
[0124] In one of the embodiments, the method further comprises:
[0125] S701, migrate the task feedback information to the preset storage space, and clear the data of the computing power allocation resource result to obtain the cleared computing power allocation resource.
[0126] Specifically, the terminal migrates the task feedback information to the preset military-grade encrypted storage space, and performs data clearing on the computing power allocation resource result allocated by dynamic cutting to obtain the cleared computing power allocation resource.
[0127] S702, add the cleared computing power allocation resource to the computing power resource pool to obtain an updated computing power resource pool.
[0128] Specifically, the cleared computing power allocation resource is re-added to the computing power resource pool, and the idle state of the adjacent resource block under the same computing power resource is detected. If it is idle, a fragment consolidation algorithm is used to merge the cleared computing power allocation resource with the adjacent resource block to obtain a computing power resource with stronger computing power performance, thereby obtaining an updated computing power resource pool.
[0129] To further illustrate the scheme of the application example in this embodiment, a specific example is used for illustration:
[0130] Based on the computing power resource cross-platform migration scheduling method, a computing power scheduling cloud platform can be developed for users with computing power needs, and the specific functions include:
[0131] (I) Physical infrastructure layer.
[0132] Provide a physical computing power base for domestic and imported mixed deployment, support unified management of storage network and security equipment, and support mixed scheduling of domestic and imported chips.
[0133] (II) Computing power resource pooling layer.
[0134] Abstract the physical computing power resource into a standardized computing power unit to support dynamic allocation and elastic expansion. Computing power dynamic allocation supports vGPU creation of domestic and imported chips, and supports hyper-converged storage in terms of storage.
[0135] (III) Computing power resource management layer.
[0136] Pool and schedule physical hardware to realize full-life-cycle management and control of cross-architecture resources. This includes allocating computing power through vGPU virtualization and pass-through technology, and automatically matching domestic / imported chips according to task requirements and chip idle load capacity.
[0137] (IV) Computing power resource service layer.
[0138] As a unified portal, it provides standardized computing power services for AI frameworks and industry applications. Users register task requirements through open interfaces, and the system automatically orchestrates resource pooling capabilities and dynamically generates resource lists, supporting filtering available resources by instruction set.
[0139] The functions implemented by the computing power scheduling cloud platform include:
[0140] (1) GPU pool dynamic allocation.
[0141] Each time a user requests resources, available GPUs / vGPUs are dynamically selected from the GPU pool and provided to the user. During this period, data disks are mounted in real time, business data is saved, and GPU resource usage is audited. GPU resource pooling scheduling and dynamic roaming allocation of computing power are used. When the user completes the use of AI computing power resources, the resources are automatically recycled, and the resources are periodically recycled to achieve dynamic use of GPUs - time-sharing multiplexing of GPU computing power, dynamic recycling of GPUs - automatic recycling of heterogeneous computing power.
[0142] (2) Unified scheduling of heterogeneous resources.
[0143] Build a resource pool, which is a unit of organization and division of computing resources, network resources, storage resources, and GPU resources, providing computing power support for different business scenarios.
[0144] Multiple resource pools for different purposes and SLA service levels can be created and managed simultaneously.
[0145] X86 and ARM architecture server heterogeneous resource pool unified scheduling management can be built.
[0146] (3) Heterogeneous GPU pooling scheduling.
[0147] The computing power platform provides GPU passthrough virtualization, GPU virtualization (vGPU), and GPU redirection technology, using a GPU passthrough virtualization device driver model and resource grid pooling scheduling algorithm to achieve virtualization of general-purpose standard GPU devices in host servers, as well as dynamic scheduling and allocation management of AI cloud hosts and GPU virtualization resources. The cloud host has a three-dimensional graphics image computing processing performance not less than 95% of the same configuration physical graphics PC workstation, and inference training capability not only supports AI inference training, AI development testing, but also supports cloud design rendering, AI video processing, and other business scenarios.
[0148] (4) Super-converged computing power resources.
[0149] Adopting a centerless symmetrical architecture, there is no metadata expansion bottleneck, supporting PB-level storage capacity.
[0150] Using elastic Hash algorithm, it has close to linear scalability.
[0151] A multi-copy redundancy mechanism is provided to ensure data fault tolerance and high availability.
[0152] A variety of copy mechanisms such as single copy, 2+1 copy, 3 copy and erasure code are provided.
[0153] (5) Computing power resource migration.
[0154] Cross-storage computing power services, online migration of computing power resources between heterogeneous storage pools such as SAN storage, hyper-converged, and local storage can be supported simultaneously.
[0155] (6) High reliability of computing power resources.
[0156] When the host server fails, the computing power platform can schedule the normal operation of the cloud host affected by the failed host server within a few minutes, ensuring the high availability of all cloud hosts in the resource pool and ensuring the high reliability of the computing power service.
[0157] (7) Protection of residual information of computing power resources.
[0158] A residual information protection management mechanism based on memory and disk is provided to prevent security threats through residual information attacks on host server memory, disk, and storage.
[0159] The computing power resource cross-platform migration scheduling method provided in the embodiment realizes high-security, zero-residue, and low-fragmentation cross-platform computing power resource scheduling and migration, effectively improving the work efficiency of computing power resource scheduling.
[0160] In the above computing power resource cross-platform migration scheduling method, the user demand instruction includes a task computing power demand instruction, a task time delay demand instruction, and a task instruction set; the task instruction set is matched with a preset computing power resource pool as a constraint condition of the task computing power demand instruction and the task time delay demand instruction, to obtain a first computing power resource pool including computing power resources meeting the task computing power demand instruction and the task time delay demand instruction; when the first computing power resource pool meets a preset instruction conversion requirement, a task computing power resource is selected from the first computing power resource pool, and the task instruction set is converted according to the task computing power resource to obtain a standard instruction set; the resource architecture type of the task computing power resource is different from that of the task instruction set; the task computing power resource is dynamically allocated according to the standard instruction set to obtain a computing power allocation resource result, the standard instruction set is processed based on the computing power allocation resource result to obtain task feedback information, and cross-platform dynamic allocation of computing power resources is realized, improving the work efficiency of computing power resource scheduling.
[0161] It should be understood that, although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise explicitly stated herein, there is no strict order limitation on the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowchart involved in each embodiment as described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0162] Based on the same inventive concept, the embodiments of the present application also provide a computing resource cross-platform migration scheduling device for implementing the computing resource cross-platform migration scheduling method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more computing resource cross-platform migration scheduling device embodiments provided below can refer to the limitations of the computing resource cross-platform migration scheduling method described above, which will not be repeated here.
[0163] In one exemplary embodiment, as shown in Figure 2 A computing resource cross-platform migration scheduling device 200 is provided, comprising:
[0164] An instruction acquisition module 201 is configured to acquire a user demand instruction, wherein the user demand instruction includes a task computing resource demand instruction, a task time delay demand instruction, and a task instruction set;
[0165] A first matching module 202 is configured to match the task instruction set with a preset computing resource pool by taking the task computing resource demand instruction and the task time delay demand instruction as constraint conditions, to obtain a first computing resource pool, wherein the first computing resource pool includes computing resources that meet the task computing resource demand instruction and the task time delay demand instruction;
[0166] A second matching module 203 is configured to select a task computing resource from the first computing resource pool when the first computing resource pool meets a preset instruction conversion requirement, and convert the task instruction set according to the task computing resource, to obtain a standard instruction set; the resource architecture type of the task computing resource is different from the resource architecture type of the task instruction set;
[0167] A task completion module 204 is configured to dynamically allocate the task computing resource according to the standard instruction set, to obtain a computing resource allocation result, process the standard instruction set based on the computing resource allocation result, and obtain task feedback information.
[0168] Further, the first matching module is further configured to:
[0169] For each computing resource in the computing resource pool, the first load score corresponding to the computing resource is calculated using the following formula:
[0170]
[0171] wherein S(v i ) is the first load score of the computing resource v i , α is the performance weight, C i is the peak computing power of the computing resource v i , R c is the task computing power requirement instruction, δ v is the virtualization overhead coefficient, β is the time delay weight, T c is the task time delay requirement instruction, L i is the task queuing time delay, τ i is the startup time consumption, Δ is the transmission time delay, γ is the storage weight, S f is the free video memory of the computing resource v i , S t is the total video memory of the computing resource v i , δ is the architecture weight, is the architecture difference score of the computing resource v i ;
[0172] wherein the architecture difference score is calculated using the following formula:
[0173]
[0174] wherein, is the architecture difference score of the computing resource v i , ω1 is the bit width weight, W d(x,i) is the data bit width difference value between the computing resource v i and the task instruction set, ω2 is the vector unit weight, U d(x,i) is the vector processing unit difference value between the computing resource v i and the task instruction set, ω3 is the hardware weight, H d(x,i) is the hardware compatibility difference value between the computing resource v i and the task instruction set, ω4 is the conversion weight, C (x,i) is the conversion index of the computing resource vi and the task instruction set.
[0175] The computing resource with the first load score greater than the preset first load threshold in the computing resource pool is selected to obtain a first computing resource pool.
[0176] Further, the second matching module includes:
[0177] The second computing resource pool construction unit is configured to select computing resources from the first computing resource pool in an order of increasing architecture difference scores until the task computing resource demand instruction is met, and construct a second computing resource pool based on the selected computing resources.
[0178] The instruction set conversion unit is configured to convert the task instruction set according to each computing resource in the second computing resource pool to obtain a corresponding standard instruction set.
[0179] The loss rate calculation unit is configured to calculate the performance loss of the conversion of the task instruction set and each standard instruction set to obtain a performance loss rate corresponding to each computing resource.
[0180] The computing resource confirmation unit is configured to select computing resources from the second computing resource pool in an order of increasing performance loss rates, and configure the selected computing resources as task computing resources.
[0181] The instruction conversion requirement is that the computing resources in the first computing resource pool include computing resources whose resource architecture types are different from the resource architecture type of the task instruction set.
[0182] Further, the instruction set conversion unit is further configured to:
[0183] Split the task instruction set according to a preset instruction division rule to obtain a basic instruction set; the basic instruction set is used to represent a set of basic instructions.
[0184] According to the resource architecture type of the computing resource in the second computing resource pool and the task instruction set, select a corresponding instruction conversion rule from a preset mapping rule set.
[0185] Iterate through each basic instruction in the basic instruction set, and determine a corresponding conversion instruction from the instruction conversion rule.
[0186] Convert the basic instruction according to the conversion instruction to obtain a corresponding conversion atomic instruction.
[0187] Combine all conversion atomic instructions in the traversal order to obtain the standard instruction set.
[0188] Further, the loss rate calculation unit is further configured to:
[0189] According to the preset historical task data, select instructions matching the historical task data from the task instruction set to obtain an initial instruction type set.
[0190] Based on the initial instruction type set, calculate the occurrence frequency of each instruction in the initial instruction type set to obtain a weighted weight of each instruction.
[0191] Based on the task instruction set, calculate the original cycle instruction number of each instruction.
[0192] based on each standard instruction set, the number of conversion cycles of each instruction in each standard instruction set is calculated;
[0193] based on the original cycle instruction number of each instruction and the conversion cycle instruction number in each standard instruction set, the performance loss rate corresponding to each computing resource is calculated.
[0194] Further, the task completion module is also used for:
[0195] According to the task computing power demand instruction and the virtualization overhead coefficient, the computing power compensation computing power data is calculated;
[0196] According to the task computing power resource, the computing power compensation computing power data is adjusted to obtain virtual computing power compensation computing power data;
[0197] GPU pass virtual algorithm and GPU pooling scheduling algorithm are used, and the virtual computing power compensation computing power data is used to dynamically allocate the task computing power resource to obtain the computing power allocation resource result;
[0198] The standard instruction set is processed by using the computing power allocation resource result to obtain the task feedback information.
[0199] Further, the device further includes a storage migration module for:
[0200] The task feedback information is migrated to the preset storage space, and the data of the computing power allocation resource result is cleared to obtain the cleared computing power allocation resource;
[0201] The cleared computing power allocation resource is added to the computing power resource pool to obtain the updated computing power resource pool.
[0202] In one embodiment, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and the processor implements the steps in each method embodiment described above when executing the computer program.
[0203] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in each method embodiment described above.
[0204] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The device embodiment described above is only schematic, wherein the components shown as separate components can or can not be physically separate, and the components shown as a unit can or can not be a physical unit, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure. Those skilled in the art can understand and implement it without creative labor.
[0205] The above-described embodiments only express several implementation manners of the present application, which are described in detail, but cannot be understood as a limitation on the patent scope of the application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.
Claims
1. A method for cross-platform migration and scheduling of computing resources, characterized in that, The method includes: Obtain user requirement instructions, which include task computing power requirement instructions, task latency requirement instructions, and task instruction sets; Using the task computing power requirement instruction and the task latency requirement instruction as constraints, the task instruction set is matched with a preset computing power resource pool to obtain a first computing power resource pool. The first computing power resource pool includes computing power resources that satisfy the task computing power requirement instruction and the task latency requirement instruction. When the first computing power resource pool meets the preset instruction conversion requirements, task computing power resources are selected from the first computing power resource pool, and the task instruction set is converted according to the task computing power resources to obtain a standard instruction set; the resource architecture type of the task computing power resources is different from the resource architecture type of the task instruction set. The computing resources for the task are dynamically allocated according to the standard instruction set to obtain the computing resource allocation result. The standard instruction set is then processed based on the computing resource allocation result to obtain task feedback information.
2. The method according to claim 1, characterized in that, The first computing resource pool is obtained by matching the task instruction set with a preset computing resource pool, using the task computing power requirement instruction and the task latency requirement instruction as constraints, including: For each computing resource in the computing resource pool, the first load score corresponding to the computing resource is calculated using the following formula: Wherein, S(v i ) is computing power resource v i The first load score, α is the performance weight, C i It is computing power resources v i peak computing power, R c It is a task computing power requirement instruction, δ v β is the virtualization overhead coefficient, β is the latency weight, and T is the latency weight. c It is a task latency requirement instruction, L i It is the task queuing delay, τ i Δ is the startup time, γ is the transmission latency, and S is the storage weight. f Computing resources v i Free video memory, S t It is computing power resources v i Total video memory, δ is the architecture weight. It is computing power resources v i Architectural differences score; The architecture difference score is calculated using the following formula: in, It is computing power resources v i The architectural difference score, ω1 is the bit width weight, W d(x,i) It is computing power resources v i The difference in data bit width between the task instruction set and the data bit width, ω2 is the vector unit weight, U d(x,i) It is computing power resources v i The difference between the vector processing unit and the task instruction set, ω3 is the hardware weight, H d(x,i) It is computing power resources v i Hardware compatibility difference with the task instruction set, ω4 is the transformation weight, C (x,i) It is the conversion index between computing power resources (VI) and task instruction sets; The computing resources in the computing resource pool whose first load score is greater than a preset first load threshold are selected to obtain the first computing resource pool.
3. The method according to claim 2, characterized in that, When the first computing power resource pool meets the preset instruction conversion requirements, task computing power resources are selected from the first computing power resource pool, and the task instruction set is converted according to the task computing power resources to obtain a standard instruction set, including: The computing resources are selected from the first computing resource pool in ascending order of the architecture difference score until the task computing power requirement instruction is met, and a second computing resource pool is constructed based on the selected computing resources. Based on the computing resources in the second computing resource pool, the task instruction set is converted to obtain the corresponding standard instruction set; Calculate the conversion performance loss between the task instruction set and each of the standard instruction sets to obtain the performance loss rate corresponding to each of the computing resources; The computing resources are selected from the second computing resource pool in ascending order of performance loss rate, and the selected computing resources are configured as the task computing resources. The instruction conversion requirement is that the first computing resource pool includes computing resources whose resource architecture type is different from the resource architecture type of the task instruction set.
4. The method according to claim 3, characterized in that, The step of converting the task instruction set according to each computing resource in the second computing resource pool to obtain the corresponding standard instruction set includes: According to the preset instruction division rules, the task instruction set is split into basic instruction sets; the basic instruction set is used to represent the set of basic instructions. Based on the resource architecture type and task instruction set of the computing resources in the second computing resource pool, select the corresponding instruction conversion rule from the preset mapping rule set; Iterate through each basic instruction in the basic instruction set and determine the corresponding conversion instruction from the instruction conversion rules; The basic instruction is converted according to the conversion instruction to obtain the corresponding conversion atomic instruction; All the aforementioned conversion atomic instructions are combined in traversal order to obtain the standard instruction set.
5. The method according to claim 3, characterized in that, The calculation of the conversion performance loss between the task instruction set and each of the standard instruction sets, to obtain the performance loss rate corresponding to each of the computing resources, includes: Based on preset historical task data, select the instruction that matches the historical task data from the task instruction set to obtain an initial instruction type set; Based on the initial instruction type set, the occurrence frequency of each instruction in the initial instruction type set is calculated to obtain the weighted weight of each instruction; Based on the task instruction set, the original cycle instruction number of each instruction is calculated; Based on each of the aforementioned standard instruction sets, the number of conversion cycle instructions for each instruction in each of the aforementioned standard instruction sets is calculated; Based on the original cycle instruction count of each instruction and the conversion cycle instruction count in each standard instruction set, the performance loss rate corresponding to each computing resource is calculated.
6. The method according to claim 2, characterized in that, The process of dynamically allocating computing resources for the task according to the standard instruction set to obtain computing resource allocation results, and processing the standard instruction set based on the computing resource allocation results to obtain task feedback information includes: The computing power compensation data is calculated based on the task computing power requirement instruction and the virtualization overhead coefficient. Based on the computing power resources of the task, the computing power compensation data is adjusted to obtain virtual computing power compensation data; The GPU passthrough virtualization algorithm and GPU pooling scheduling algorithm are used to dynamically allocate the computing power resources of the task based on the virtual computing power compensation computing power data, so as to obtain the computing power allocation resource result; The standard instruction set is processed using the results of the computing power allocation to obtain the task feedback information.
7. The method according to claim 6, characterized in that, The method further includes: The task feedback information is migrated to a preset storage space, and the data of the computing power allocation result is cleared to obtain cleared computing power allocation resources; The cleared computing power allocation resources are added to the computing power resource pool to obtain the updated computing power resource pool.
8. A cross-platform migration and scheduling device for computing resources, characterized in that, The device includes: The instruction acquisition module is used to acquire user demand instructions, which include task computing power demand instructions, task latency demand instructions, and task instruction sets. The first matching module is used to match the task instruction set with a preset computing resource pool based on the task computing power requirement instruction and the task latency requirement instruction as constraints, to obtain a first computing resource pool. The first computing resource pool includes computing resources that satisfy the task computing power requirement instruction and the task latency requirement instruction. The second matching module is used to select task computing resources from the first computing resource pool when the first computing resource pool meets the preset instruction conversion requirements, and convert the task instruction set according to the task computing resources to obtain a standard instruction set; the resource architecture type of the task computing resources is different from the resource architecture type of the task instruction set. The task completion module is used to dynamically allocate computing resources for the task according to the standard instruction set, obtain the computing resource allocation result, process the standard instruction set based on the computing resource allocation result, and obtain task feedback information.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Cross-platform resource scheduling method and system for hybrid cloud environment
CN121842131A
Cross-platform resource scheduling method and system for hybrid cloud environment
CN121842131B