Task allocation method, electronic device, and computer-readable storage medium
By mapping target tasks to affine connection manifolds on the GPU and allocating them according to the manifold curvature, the problem of uneven task allocation is solved, improving processor efficiency and resource utilization.
Patent Information
- Application Number
- CN202511232504.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-29
AI Technical Summary
In existing technologies, uneven task allocation among different computing cores of a GPU leads to decreased processing efficiency, with one CUDA core and one Tensor core being in operation while the other may be relatively idle.
By mapping the target task onto an affine connection manifold, the task is allocated to the first and second computational cores according to the manifold curvature. The manifold curvature of the first region is lower than that of the second region, thus achieving uniform task allocation.
It improves the processor's processing efficiency, achieves even task allocation between the first and second computing cores, and makes full use of the computing resources of both.
Smart Images

Figure CN120743556B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to a task allocation method, electronic device, and computer-readable storage medium. Background Technology
[0002] With the rapid development of artificial intelligence and deep learning technologies, GPUs (Graphics Processing Units), as core hardware for high-performance computing, have a significant impact on AI (Artificial Intelligence) training and inference through their models and computational methods. In AI applications using GPUs, different computing cores within the GPU handle different computational tasks. For example, CUDA cores (Compute Unified Device Architecture cores) are primarily responsible for handling simple tasks that do not involve matrix computations, while Tensor cores are mainly responsible for accelerating complex tasks such as matrix operations.
[0003] In related technologies, when allocating tasks to a GPU, tasks are usually allocated according to the proportion of matrix operations. Tasks with a higher proportion of non-matrix operations are assigned to CUDA cores, while tasks with a higher proportion of matrix operations are assigned to Tensor cores. This results in one CUDA core or Tensor core being in operation while the other may be relatively idle. Therefore, there is an uneven distribution of tasks, which leads to a decrease in the processing efficiency of the GPU. Summary of the Invention
[0004] This application provides a task allocation method, an electronic device, and a computer-readable storage medium to at least solve the problem of reduced processor processing efficiency caused by uneven task allocation to different computing cores in related technologies.
[0005] This application provides a task allocation method applied to a processor. The processor includes a first computing core and a second computing core. The first computing core is used to perform general-purpose parallel operations, and the second computing core is used to perform matrix operations. The task allocation method includes: when the processor receives a target task, determining the affine connection manifold of the target task, the affine connection manifold being used to characterize the distribution characteristics of the computational load of the target task; and allocating the target task to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold; wherein, the first computational load of the target task allocated to the first computing core corresponds to a first region in the affine connection manifold, the second computational load of the target task allocated to the second computing core corresponds to a second region in the affine connection manifold, and the manifold curvature of the first region is lower than that of the second region.
[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing any of the above-described task allocation methods when executing the computer program.
[0007] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described task allocation methods.
[0008] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described task allocation methods.
[0009] This application maps the target task onto an affine connection manifold, determines the manifold curvature of the affine connection manifold, and divides the affine connection manifold into a first region and a second region according to the curvature. The first computational load corresponding to the first region in the target task is allocated to the first computing core, and the second computational load corresponding to the second region in the target task is allocated to the second computing core. This enables rapid invocation of the first and second computing cores, improves the uniformity of task allocation between the first and second computing cores, thereby improving processor efficiency, and allows for fine-grained allocation of target tasks, ensuring full utilization of the computing resources of both the first and second computing cores. Attached Figure Description
[0010] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 The following are schematic diagrams of server architectures according to some embodiments of this application;
[0012] Figure 2 The present application shows a module topology diagram of a server processor according to some embodiments;
[0013] Figure 3 Flowcharts of task allocation methods according to some embodiments of this application are shown;
[0014] Figure 4 Structural block diagrams of a task allocation apparatus according to some embodiments of this application are shown;
[0015] Figure 5 Structural block diagrams of electronic devices according to some embodiments of this application are shown. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0017] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0018] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] In order to solve the technical problems in the above-mentioned related technologies, this application proposes a task allocation method, which will be further described in detail below with reference to the accompanying drawings.
[0020] The specific application environment architecture or specific hardware architecture on which the task allocation method depends is described here.
[0021] In some embodiments of this application, exemplarily, Figure 1 The following are schematic diagrams of server architectures according to some embodiments of this application. Figure 2 The following is a module topology diagram of a server processor according to some embodiments of this application, such as... Figure 1 and Figure 2As shown, the server architecture supports the Hopper8-GPU module. This architecture features eight Graphics Processing Units (GPUs) and two Central Processing Units (CPUs). The CPUs and GPUs are interconnected via a PCIe Switch for rapid expansion of peripheral components. High-speed communication between the GPUs is achieved through high-speed interconnect technology. The two CPUs and the platform controller are centrally located on the motherboard, while the eight GPUs are mounted on the GPU substrate.
[0022] like Figure 2 As shown, the server architecture module inputs 128 threads of PCIe 5.0 signals, which are then connected to the graphics processor after passing through 8 retimer chips, thereby enhancing the link driving capability.
[0023] For example, the task assignment method is used to assign tasks to objects such as... Figure 1 and Figure 2 In the server architecture, the graphics processor is used for task allocation; that is, the task allocation method is applied to the processing of the graphics processor.
[0024] Embodiments of this application provide a task allocation method applied to a processor. The processor includes a first computing core and a second computing core. The first computing core is used to perform general-purpose parallel computing, and the second computing core is used to perform matrix operations. Figure 3 Flowcharts of task allocation methods according to some embodiments of this application are shown, such as Figure 3 As shown, the task allocation methods include:
[0025] Step 302: When the processor receives the target task, determine the affine connection manifold of the target task. The affine connection manifold is used to characterize the distribution characteristics of the computational load of the target task.
[0026] In this embodiment, the processor includes a first computing core and a second computing core. The first computing core is responsible for performing simple computational loads such as general parallel operations, and the second computing core is responsible for performing complex computational loads such as accelerated matrix operations.
[0027] For example, the processor is a graphics processing unit (GPU), the first computing core is the CUDA core in the CPU, and the second computing core is the Tensor core in the GPU. The CUDA core is a general-purpose parallel processing core in the GPU used to perform floating-point and integer operations, capable of processing large-scale data streams in parallel. One CUDA core is responsible for executing one instruction but can operate on multiple data units, making it particularly suitable for repetitive computational tasks, primarily used for traditional floating-point and integer operations. The Tensor core is used for matrix multiplication and accumulation operations, the core computational tasks of deep learning, capable of performing matrix operations with extremely high throughput. In AI (Artificial Intelligence) computing, the CUDA core and Tensor core form a complementary and synergistic relationship, and their collaborative mode has a significant impact on the GPU's computing power.
[0028] In this embodiment, when the processor receives a target task, it determines the affine connection manifold of the target task. The affine connection manifold can map the target task into the affine manifold and describe the distribution characteristics of the computational load of the target task.
[0029] Step 304: Based on the manifold curvature of the affine connection manifold, assign the target task to at least one of the first computing core and the second computing core;
[0030] In this context, the first computational load allocated to the first computing core in the target task corresponds to the first region in the affine connected manifold, and the second computational load allocated to the second computing core in the target task corresponds to the second region in the affine connected manifold. The manifold curvature of the first region is lower than that of the second region.
[0031] In this embodiment, the curvature of the affine connection manifold reflects the uniformity of the target task's distribution on the manifold; that is, different computational loads of varying complexity correspond to different curvatures in the affine connection manifold. Specifically, a first region with lower curvature in the affine connection manifold corresponds to a first computational load of lower computational complexity in the target task, and a second region with higher curvature in the affine connection manifold corresponds to a second computational load of higher computational complexity in the target task.
[0032] The curvature of the affine connected manifold allows us to determine its first and second regions. Since the curvature of the first region is lower than that of the second region, the computational complexity of the first computational load corresponding to the first region is lower than that of the second computational load corresponding to the second region. It's understandable that the first computational core performs general-purpose parallel computations, while the second computational core performs matrix operations. Since the computational complexity of general-purpose parallel computations is lower than that of matrix operations, allocating the first computational load corresponding to the first region of the affine connected manifold to the first computational core and the second computational load corresponding to the second region to the second computational core enables rapid invocation of both the first and second computational cores.
[0033] The following explanation uses an example where the first computing core is a CUDA core and the second computing core is a Tensor core:
[0034] In related technologies, when allocating tasks to the processor, the allocation is usually based on the proportion of matrix operation load. Tasks with a high proportion of matrix operation load are allocated to Tensor cores, and tasks with a low proportion of matrix operation load are allocated to CUDA cores. This allocation method may result in one of the Tensor cores or CUDA cores being in a saturated state while the other is in a relatively idle state, which means there is a problem of uneven task allocation.
[0035] In this embodiment, the target task is mapped to an affine connected manifold, and the first computational load corresponding to the first region with lower manifold curvature in the affine connected manifold is allocated to the CUDA core, and the second computational load corresponding to the second region with higher manifold curvature in the affine connected manifold is allocated to the Tensor core.
[0036] For example, the target task is a Transformer (attention-based deep learning model) inference task, which includes self-attention layer computational load and residual connection computational load. The Transformer inference task is mapped to an affine connection manifold, where the first region of the affine connection manifold with lower manifold curvature corresponds to the residual connection computational load. Therefore, the residual connection computational load is assigned as the first computational load to the CUDA core, and the second region of the affine connection manifold with higher manifold curvature corresponds to the self-attention layer computational load. Therefore, the self-attention layer computational load is assigned to the Tensor core.
[0037] In this embodiment, by mapping the target task onto an affine connection manifold and determining the manifold curvature, the affine connection manifold is divided into a first region and a second region according to the curvature. The first computational load corresponding to the first region in the target task is allocated to the first computing core, and the second computational load corresponding to the second region in the target task is allocated to the second computing core. This enables rapid invocation of the first and second computing cores, improves the uniformity of task allocation between the first and second computing cores, thereby improving the processor's processing efficiency. Furthermore, it allows for fine-grained allocation of target tasks, ensuring that the computing resources of the first and second computing cores are fully utilized.
[0038] In some embodiments of this application, determining the affine connection manifold of the target task includes: constructing a target geometric association between a first computing core and a second computing core; determining a target mapping relationship between the computational load and manifold curvature in the target task based on the target geometric association; and determining the affine connection manifold of the target task corresponding to the target task based on the target mapping relationship.
[0039] In this embodiment, the target geometric association between the first and second computational cores is constructed using Christopher symbols. The essence of Christopher symbols is a differential coordinate converter describing the manifold structure. Christopher symbols provide a continuously differentiable geometric description, thereby improving the accuracy of describing the distribution characteristics of the computational load of the target task.
[0040] For example, the Christopher notation is defined by the first derivative of the basis vectors as shown in equation (1):
[0041] (1)
[0042] in, e j Coordinates x j Covariant basis vectors in direction e j It is the corresponding contravariant basis vector. Indicates along direction e j covariant derivative, It is the Christopher symbol, representing the basis vector. e j In direction ei The rate of change on.
[0043] It should be noted that the computing unit of the processor is regarded as a manifold, and each point represents a computing unit. The computing unit can be a thread block or a stream processor. The computing unit includes, but is not limited to, the first computing core and the second computing core. A local coordinate system (x^1, x^2, ..., x^n) is defined on the affine connection manifold, where (n) is the dimension of the manifold. With the local coordinates, the covariant basis vector and the contravariant basis vector can be calculated accordingly.
[0044] In this embodiment, the uneven distribution of the target task on the manifold is calculated by using Christopher notation, and a target mapping relationship between each computational load in the target task and the curvature of the manifold is established. After the target mapping relationship is constructed, the corresponding affine connected manifold can be constructed based on the target mapping relationship and each computational load in the target task, which improves the accuracy of constructing the affine connected manifold representation and enables the affine connected manifold to accurately represent the distribution characteristics of the computational load in the target task.
[0045] In some embodiments of this application, allocating a target task to at least one of a first computing core and a second computing core based on the manifold curvature of an affine connection manifold includes: dividing the affine connection manifold into a first region and a second region, wherein the curvature of the first region is less than the curvature of the second region; allocating a first computational load corresponding to the first region of the target task to the first computing core; and allocating a second computational load corresponding to the second region of the target task to the second computing core; wherein the computational complexity of the first computational load is less than the computational complexity of the second computational load.
[0046] In this embodiment, after obtaining the affine connection manifold, the affine connection manifold is divided into regions to obtain a first region with low curvature and a second region with high curvature. The first computational load corresponding to the first region with low curvature is usually processing rule data transfer, and the second computational load corresponding to the second region with high curvature is usually processing load matrix operations. Therefore, the first computational load corresponding to the first region is allocated to the first computing core, and the second computational load corresponding to the second region is allocated to the second computing core.
[0047] For example, the curvature of the affine connection manifold is related to the computational load distribution characteristics in the target task. The geometric characteristics of the affine connection manifolds constructed by different target tasks are also different. Therefore, the proportions of the first region and the second region in the entire affine connection manifold are also different.
[0048] For example, consider a Transformer (a deep learning model based on attention mechanisms) inference task. If the target task only includes self-attention layer computational load, the entire affine connection manifold corresponding to the target task is a high-curvature region, meaning the affine connection manifold only includes the second region and excludes the first region. Therefore, all computational load in the target task is allocated to the second computational core. If the target task only includes residual computational load, the entire affine connection manifold corresponding to the target task is a low-curvature region, meaning the affine connection manifold only includes the first region and excludes the second region. Therefore, all computational load in the target task is allocated to the second computational core.
[0049] In this embodiment, after the affine connection manifold is constructed, the affine connection manifold is divided into a first region and a second region according to the curvature of the manifold. That is, the first region represents the first computational load with lower computational complexity, and the second region represents the second computational load with higher computational complexity. This realizes that the first computational load and the second computational load are respectively allocated to the corresponding first computing core and second computing core according to the computational complexity of the computational load in the target task, thereby improving the speed of calling the first computing core and the second computing core in the processor.
[0050] In some embodiments of this application, dividing an affine connection manifold into a first region and a second region includes: determining a first region and a second region in the affine connection manifold based on a first curvature threshold, wherein the curvature of the first region is less than or equal to the first curvature threshold, and the curvature of the second region is greater than the first curvature threshold.
[0051] In this embodiment, a first curvature threshold is used to divide a first region and a second region in an affine connected manifold. Regions in the affine connected manifold with manifold curvature less than or equal to the first curvature threshold are defined as the first region, and regions with manifold curvature greater than the first curvature threshold are defined as the second region. The first curvature threshold is determined through extensive simulations and field measurements, ensuring that the first region, divided by the first curvature threshold, matches a first computational load with lower computational complexity, and that the second region, divided by the first curvature threshold, matches a second computational load with higher computational complexity. This further improves the accuracy of allocating target tasks to the first and second computing cores.
[0052] In some embodiments of this application, after allocating the target task to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, the task allocation method further includes: determining the Riemannian manifold corresponding to the computing space of the processor, wherein the Riemannian manifold is used to characterize the degree of concentration of computational load in the computing space; performing load reconstruction processing on the target task allocated to the processor according to the Gaussian curvature of the Riemannian manifold; and allocating the target task after load reconstruction processing to the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, so as to reduce the load difference between the computational load in the first computing core and the computational load in the second computing core.
[0053] In this embodiment, the Riemannian manifold is matched with the computational space of the processor after the target task is assigned, and the Riemannian manifold can characterize the degree of aggregation of computational load in the computational space of the processor after the target task is assigned.
[0054] The Gaussian curvature of the Riemannian manifold accurately characterizes the degree of computational load concentration in the computation space. If the concentration of computational load is high, at least one of the first and second computing cores may exceed its predetermined task allocation capacity, thus affecting the processor's overall task processing efficiency. Conversely, if the concentration is low, at least one of the first and second computing cores may experience wasted computational resources. In this case, load refactoring of the allocated target tasks based on the Gaussian curvature of the Riemannian manifold can split highly concentrated computational loads to reduce their concentration, and merge less concentrated computational loads to increase their concentration.
[0055] For example, the load refactoring process includes splitting the computational load and merging the computational load.
[0056] Specifically, the target task within the current computation space consists of a first computational load and a second computational load allocated based on an affine connected manifold. Therefore, either the first or second computational load undergoes load reconfiguration. After load reconfiguration, the target task is redistributed to the first and second computational cores based on the manifold curvature of the affine connected manifold. Since the manifold curvature of the first and second computational loads may change after load reconfiguration, it is necessary to redistribute them based on the manifold curvature to ensure the accuracy of the load allocation. The redistributed first and second computational loads have a lower degree of computational load aggregation, thereby further balancing the computational load between the first and second computational loads and reducing the load difference between the computational load in the first and second computational cores.
[0057] In this embodiment, after the target task is assigned to one of the first computing core and the second computing core, the processor's computing space is mapped to a Riemannian manifold. The Gaussian curvature of the Riemannian manifold can be used to assess the current degree of computational load aggregation of the processor. Based on this, the target task is reconfigured. This reduces the situation where at least one of the first computing core and the second computing core exceeds the predetermined task allocation capacity due to a high degree of computational load aggregation in the computing space, and also reduces the situation where at least one of the first computing core and the second computing core wastes computational resources due to a low degree of computational load aggregation in the computing space.
[0058] In some embodiments of this application, determining the Riemann manifold corresponding to the processor's computation space includes: obtaining at least two computation units in the processor, wherein the at least two computation units include a first computation core and a second computation core; determining at least two manifold points based on the at least two computation units; and constructing a Riemann manifold based on the at least two manifold points.
[0059] In this embodiment, the processor's computation space is abstracted as a Riemannian manifold, where each computational unit in the computation space corresponds to a manifold point in the Riemannian manifold. During the construction of the Riemannian manifold, at least two computational units in the processor are obtained, including but not limited to a first computational core and a second computational core. At least two corresponding manifold points are determined based on these at least two computational units, and each manifold point corresponds one-to-one with one of the at least two computational units. After determining the at least two manifold points, the Riemannian manifold can be constructed using these at least two manifold points.
[0060] For example, the processor is a GPU, and the computing unit includes, but is not limited to, thread blocks and streaming multiprocessors (SMs). The computing unit includes CUDA cores and Tensor cores, namely the first computing core and the second computing core.
[0061] In this embodiment, by mapping each computing unit in the processor's computing space to a manifold point of a Riemannian manifold, and then constructing a Riemannian manifold based on at least two manifold points corresponding to at least two computing units, the matching between the constructed Riemannian manifold and the processor's computing space is improved, thereby improving the accuracy of the load aggregation degree in the computing space determined based on the curvature of the Riemannian manifold.
[0062] In some embodiments of this application, before performing load refactoring on the target tasks assigned to the first computing core and the second computing core based on the Gaussian curvature of the Riemann manifold, the task allocation method further includes: obtaining the metric tensor component of the target task on the Riemann manifold; and obtaining the Riemann curvature tensor component of the Riemann manifold, wherein the Riemann curvature tensor component is related to the hardware topology of the computing space; and calculating the ratio between the Riemann curvature tensor component and the metric tensor component to determine the Gaussian curvature.
[0063] In this embodiment, the Gaussian curvature of the Riemann manifold can be calculated by the ratio of the Riemann curvature tensor component to the metric tensor component. Therefore, after constructing the Riemann manifold, the Riemann curvature tensor component of the Riemann manifold and the metric tensor component of the target task on the Riemann manifold are obtained.
[0064] For example, the expression (2) for the Gaussian curvature of a Riemannian manifold is as follows:
[0065] (2)
[0066] in, K For Gaussian curvature, R 1212 For the Riemann curvature tensor components, det ( g () is a metric tensor component.
[0067] In this embodiment, the Riemann curvature tensor component of the Riemann curve is determined based on the hardware topology of the computing space, and the metric tensor component of the target task on the Riemann manifold is obtained. Gaussian curvature can be obtained by calculating the ratio of the Riemann curvature tensor component and the metric tensor component. By modeling the computational load distribution as an energy function optimization problem on the Riemann manifold, load balancing on the first and second computing cores is achieved through task allocation driven by Gaussian curvature.
[0068] In some embodiments of this application, obtaining the metric tensor components of the target task on the Riemannian manifold includes: obtaining the communication overhead and computational intensity required by the target task; and determining the metric tensor components of the target task based on the communication overhead and computational intensity.
[0069] In this embodiment, the load concentration of the target task in the computing space is related to the communication overhead and computational intensity required when the target task is processed. That is, the higher the communication overhead and computational intensity required by the target task, the greater the likelihood of a high load concentration of the target task in the computing space. Conversely, the lower the communication overhead and computational intensity required by the target task, the lower the likelihood of a high load concentration of the target task in the computing space. Therefore, by constructing the components of the metric tensor based on the communication overhead and computational intensity required by the target task, the accuracy of the Gaussian curvature of the calculated Riemannian manifold can be further improved.
[0070] In some embodiments of this application, the target tasks allocated to the first computing core and the second computing core are refactored according to the Gaussian curvature of the Riemannian manifold, including: splitting the computational load of the target tasks when the Gaussian curvature is greater than a second curvature threshold; and merging the computational load of the target tasks when the Gaussian curvature is less than a third curvature threshold.
[0071] In this embodiment, a higher Gaussian curvature of the Riemannian manifold indicates a higher degree of load concentration of the assigned target task in the computational space; conversely, a lower Gaussian curvature indicates a lower degree of load concentration of the assigned target task in the computational space. Excessive load concentration of the target task in the computational space leads to decreased task processing efficiency, while excessively low load concentration leads to wasted runtime resources. Therefore, load reconfiguration of the target task is necessary based on the degree of load concentration.
[0072] For example, the first computational load with lower computational complexity in the target task can be merged into a second computational load with higher computational complexity. Specifically, the first computational load is a general parallel computational load, which is merged into a matrix operation load, i.e., the second computational load.
[0073] For example, the second computational load with higher computational complexity in the target task can be split into the first computational load with lower computational complexity. Specifically, the second computational load is a matrix load, which can be split into a general parallel computational load, i.e., the first computational load.
[0074] In this embodiment, if the Gaussian curvature is greater than the second curvature threshold, it is determined that the target task has a high degree of load concentration in the computing space. At this time, it is necessary to split the target task, thereby splitting the computational load executed in one processing unit into multiple processing units, thereby reducing the degree of load concentration of the target task in the computing space.
[0075] If the Gaussian curvature is less than the third curvature threshold, it is determined that the load concentration of the target task in the computing space is low. At this time, it is necessary to merge the target tasks, thereby merging the computational load executed in multiple processing units into one processing unit, thereby improving the load concentration of the target task in the computing space.
[0076] For example, the second curvature threshold ranges from 0 to 0, and the third curvature threshold ranges from 0 to 0. Specifically, for instance, if both the second and third curvature thresholds are 0, and the Gaussian curvature of the Riemannian manifold is greater than 0, then the computational load in the target task is split; if the Gaussian curvature of the Riemannian manifold is less than 0, then the computational load in the target task is merged.
[0077] In this embodiment, by setting a second curvature threshold and a third curvature threshold, and comparing the calculated Gaussian curvature of the Riemannian manifold with the second curvature threshold and the third curvature threshold respectively, it is possible to determine whether the Riemannian manifold has an excessively high or low load concentration. When the load concentration is too high, the computational load in the allocated target tasks is split; when the load concentration is too low, the allocated target tasks are merged, thereby completing the load reconfiguration process. This allows the target tasks after load reconfiguration to be redistributed to the first computing core and / or the second computing core, further improving the load balance between the first and second computing cores. This achieves dynamic adjustment of the computational load in the first and second computing cores, reducing the occurrence of slow task scheduling response or failures caused by the inability to quickly adapt to changes when there is a high load or sudden computing demand exceeding the predetermined resource allocation capacity.
[0078] In some embodiments of this application, after the target task after load reconstruction is assigned to the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, the task assignment method further includes: returning to the step of obtaining the Riemannian manifold corresponding to the computing space of the processor until the Gaussian curvature is less than or equal to the second curvature threshold and greater than or equal to the third curvature threshold.
[0079] In this embodiment, after the target task completes the load refactoring process and the target task after load refactoring is redistributed to the first computing core and the second computing core, the process returns to the step of obtaining the Riemann manifold corresponding to the processor's computing space and performing load refactoring on the target task based on the Gaussian curvature of the Riemann manifold, until the Gaussian curvature is between the second curvature threshold and the third curvature threshold. When the Gaussian curvature is between the second curvature threshold and the third curvature threshold, it is determined that the load aggregation degree in the current processor's computing space is in a balanced state, and there is no need to continue the load refactoring process.
[0080] For example, a Markov decision process is constructed based on Gaussian curvature. The Markov decision process can reinforce learning and dynamic programming and is used to model sequential decision problems in stochastic environments. The core idea is to iterate through states, actions, and rewards to find the optimal policy, thereby determining the second curvature threshold and the third curvature threshold.
[0081] In this embodiment, after the target task undergoes computational load reconstruction, and the processed target task is allocated to the first computing core and / or the second computing core based on the affine connection manifold, there may still be an imbalance in the degree of load aggregation. Therefore, the steps of constructing the Riemannian manifold and performing load reconstruction on the target task based on the Riemannian manifold are returned until the Gaussian curvature of the Riemannian manifold converges. This improves the task load balance within the first and second computing cores and further reduces the possibility that one of the first and second computing cores is running under high load while the other is idle.
[0082] In some embodiments of this application, assigning a target task to at least one of a first computing core and a second computing core based on the manifold curvature of an affine connection manifold includes: performing matrix rotation processing on the second computational load assigned to the second computing core through a three-dimensional rotation group, so that the alignment direction of the rotated second computational load matches the target alignment direction; and inputting the rotated second computational load into the second computing core.
[0083] In this embodiment, the three-dimensional rotation group is a Lie group SO(3), which represents the rotational symmetry of matrix operations. Since the second computational core is used to perform matrix operations, rotating the second computational load allocated to the second computational core to the target alignment direction by using the three-dimensional rotation group of the Lie group can improve the processing efficiency of the second computational load in the second computational core.
[0084] It should be noted that in three-dimensional space, the set of all linear transformations that preserve the vector length and rotation direction constitutes a continuous Lie group. Tensors with rotational symmetry, such as convolution kernels and attention matrices, also have a parameter space that is essentially an SO(3) manifold.
[0085] The following explanation uses matrix multiplication, where the second computational core is a Tensor core, as an example:
[0086] In matrix multiplication, if the data distribution does not meet the hardware alignment requirements of the Tensor Core, such as a 16×16×16 WMMA (Warp-level Matrix Multiply and Accumulate) block that does not match the target alignment direction, it will cause the Tensor Core's performance to degrade. The Tensor Core's processing performance can be improved by rotating the input matrix of the second computational load to the target alignment direction through group action transformation.
[0087] In this embodiment of the application, the processing efficiency of the second computing core is improved by performing matrix rotation on the second computing core through a three-dimensional rotation group and then transferring the matrix-rotated second computing core to the second computing core.
[0088] In some embodiments of this application, a matrix rotation process is performed on the second computational load allocated to the second computing core using a three-dimensional rotation group. This includes: parameterizing the input matrix of the second computational load to obtain the Lie algebra elements of the second computational load; mapping the Lie algebra elements of the load to a three-dimensional rotation group using an exponential mapping method to obtain a rotation matrix; and rotating the rotation matrix according to the target alignment direction to obtain the rotated second computational load.
[0089] In this embodiment, during the matrix rotation of the second computational load through a three-dimensional rotation group, the input matrix of the second computational load needs to be parameterized into Lie algebra elements first, and then the Lie algebra elements are mapped to the three-dimensional rotation group to determine the rotation matrix. Finally, the rotation matrix is rotated according to the target alignment direction to obtain the second computational load after matrix rotation. Lie algebra is an algebraic structure related to Lie groups. The Lie algebra corresponding to the three-dimensional rotation group is so(3), where the elements are 3×3 antisymmetric matrices ω. Exponential mapping is a mathematical operation that maps Lie algebra elements to Lie group elements. It can map the antisymmetric matrix ω to the rotation matrix in the Lie group SO(3). Since the calculation of Lie algebra elements is relatively more stable, it reduces the numerical error caused by accuracy issues in FP16 calculation and can avoid geometric distortion to a certain extent.
[0090] For example, the second computational core is a Tensor core, which can accelerate the exponential mapping process using the atomic multiply-accumulate instructions of the Tensor core.
[0091] For example, the exponential mapping is calculated using the following algorithm (3):
[0092] (3)
[0093] in, ω For antisymmetric matrices, R =exp( ω ) is a rotation matrix. I It is a 3×3 identity matrix. I Specifically .
[0094] In some embodiments of this application, the heterogeneous computing collaborative architecture constructed through the task allocation method is tested on an AI server. The focus is on parameters such as speedup and memory usage in the training of large AI models. Further optimization can be carried out in the future through methods such as fractional Fourier transform for computational flow analysis and establishing a Wasserstein distance-driven memory allocation strategy.
[0095] Embodiments of this application also provide a task allocation device applied to a processor. The processor includes a first computing core and a second computing core. The first computing core is used to perform general-purpose parallel operations, and the second computing core is used to perform matrix operations. Figure 4 Structural block diagrams of task allocation apparatuses according to some embodiments of this application are shown, such as... Figure 4 As shown, the task allocation device 400 includes:
[0096] The determination module 402 is used to determine the affine connection manifold of the target task when the processor receives the target task. The affine connection manifold is used to characterize the distribution characteristics of the computational load of the target task.
[0097] The allocation module 404 is used to allocate the target task to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold.
[0098] In this context, the first computational load allocated to the first computing core in the target task corresponds to the first region in the affine connected manifold, and the second computational load allocated to the second computing core in the target task corresponds to the second region in the affine connected manifold. The manifold curvature of the first region is lower than that of the second region.
[0099] In this embodiment, by mapping the target task onto an affine connection manifold and determining the manifold curvature, the affine connection manifold is divided into a first region and a second region according to the curvature. The first computational load corresponding to the first region in the target task is allocated to the first computing core, and the second computational load corresponding to the second region in the target task is allocated to the second computing core. This enables rapid invocation of the first and second computing cores, improves the uniformity of task allocation between the first and second computing cores, thereby improving the processor's processing efficiency. Furthermore, it allows for fine-grained allocation of target tasks, ensuring that the computing resources of the first and second computing cores are fully utilized.
[0100] Embodiments of this application also provide an electronic device. Figure 5 Structural block diagrams of electronic devices according to some embodiments of this application are shown, such as Figure 5 As shown, the electronic device 500 includes a memory 502 and a processor 504. The memory 502 stores a computer program, and the processor 504 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0101] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described task allocation method embodiments at runtime.
[0102] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0103] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described task allocation method embodiments.
[0104] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described task allocation method embodiments.
[0105] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] The foregoing has provided a detailed description of a task allocation method, apparatus, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method of task allocation, characterized by, The application is applied to a processor, the processor comprises a first computing core and a second computing core, the first computing core is used for executing general parallel operation, the second computing core is used for executing matrix operation, and the task allocation method comprises: In the case that the processor receives a target task, an affine connection manifold of the target task is determined, the affine connection manifold is used for characterizing the distribution characteristics of operation load of the target task; Based on the manifold curvature of the affine connection manifold, the target task is allocated to at least one of the first computing core and the second computing core; Wherein, the first operation load allocated to the first computing core in the target task corresponds to a first region in the affine connection manifold, the second operation load allocated to the second computing core in the target task corresponds to a second region in the affine connection manifold, and the manifold curvature of the first region is lower than that of the second region; A target geometric association of the first computing core and the second computing core is constructed; Based on the target geometric association, a target mapping relationship between operation load and manifold curvature in the target task is determined; According to the target mapping relationship, the affine connection manifold corresponding to the target task of the target task is determined; After the target task is allocated to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, the task allocation method further comprises: Determine the Riemannian manifold corresponding to the computing space of the processor, wherein the Riemannian manifold is used to characterize the aggregation degree of operation load in the computing space; According to the Gaussian curvature of the Riemannian manifold, the load reconstruction processing is performed on the target task allocated to the processor; After the target task is reconstructed, the target task is allocated to the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, so as to reduce the load difference between the operation load in the first computing core and the operation load in the second computing core.
2. The task allocation method according to claim 1, characterized in that, The target task is allocated to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, comprising: The affine connection manifold is divided into the first region and the second region, wherein the curvature of the first region is less than that of the second region; The first operation load corresponding to the first region in the target task is allocated to the first computing core; The second operation load corresponding to the second region in the target task is allocated to the second computing core; Wherein, the operation complexity of the first operation load is less than that of the second operation load.
3. The task allocation method according to claim 2, wherein, The affine connection manifold is divided into the first region and the second region, comprising: According to the first curvature threshold, the first region and the second region in the affine connection manifold are determined, wherein the curvature of the first region is less than or equal to the first curvature threshold, and the curvature of the second region is greater than the first curvature threshold.
4. The task allocation method according to claim 1, characterized by, The Riemannian manifold corresponding to the computing space of the processor is determined, comprising: acquire at least two computing units in the processor, wherein the at least two computing units comprise a computing unit in the first computing core and a computing unit in the second computing core; determine at least two manifold points according to the at least two computing units; construct the Riemannian manifold based on the at least two manifold points.
5. The task allocation method according to claim 1, wherein, Before the load reconstruction processing of the target task distributed to the first computing core and the second computing core according to the Gaussian curvature of the Riemannian manifold, the task distribution method further comprises: acquire the metric tensor component of the target task on the Riemannian manifold; and acquire the Riemann curvature tensor component of the Riemannian manifold, wherein the Riemann curvature tensor component is related to the hardware topology of the computing space; determine the Gaussian curvature by ratio calculation of the Riemann curvature tensor component and the metric tensor component.
6. The task allocation method according to claim 5, wherein, The acquisition of the metric tensor component of the target task on the Riemannian manifold comprises: acquire the communication overhead and the computing intensity required by the target task; determine the metric tensor component of the target task according to the communication overhead and the computing intensity.
7. The task allocation method of claim 1, wherein, The load reconstruction processing of the target task distributed to the first computing core and the second computing core according to the Gaussian curvature of the Riemannian manifold comprises: in the case that the Gaussian curvature is greater than a second curvature threshold, split the operation load of the target task; in the case that the Gaussian curvature is less than a third curvature threshold, combine the operation load of the target task.
8. The task allocation method of claim 7, wherein, After the target task after the load reconstruction processing is distributed to the first computing core and the second computing core based on the manifold curvature of the affine connection manifold, the task distribution method further comprises: return to execute the step of acquiring the Riemannian manifold corresponding to the computing space of the processor until the Gaussian curvature is less than or equal to the second curvature threshold and greater than or equal to the third curvature threshold.
9. The task allocation method according to any one of claims 1 to 3, characterized in that, The distribution of the target task to at least one of the first computing core and the second computing core based on the manifold curvature of the affine connection manifold comprises: perform matrix rotation processing on the second operation load distributed to the second computing core by a three-dimensional rotation group, so that the alignment direction of the second operation load after rotation processing matches the target alignment direction; input the second operation load after rotation processing to the second computing core.
10. The task allocation method according to claim 9, wherein, The matrix rotation processing on the second operation load distributed to the second computing core by a three-dimensional rotation group comprises: parameterize the input matrix of the second operation load to obtain Lie algebra elements of the second operation load; map the Lie algebra elements of the load to a three-dimensional rotation group by exponential mapping to obtain a rotation matrix; perform rotation processing on the rotation matrix according to the target alignment direction to obtain the second operation load after rotation processing.
11. An electronic device, comprising: comprise: a memory for storing a computer program; a processor for executing the computer program to realize the steps of the task distribution method according to any one of claims 1 to 10.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the task allocation method in any one of claims 1 to 10.
13. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the task allocation method in any one of claims 1 to 10.
Citation Information
Patent Citations
Battery assembly process defect real-time detection and classification method and system
CN119904704A
Environment monitoring method and system based on 5G
CN119989241A