Computing power distribution method and related product
By acquiring the task description information and resource pool status of the target task, computing resources are dynamically allocated, solving the problem of limited computing resources and realizing computing resource sharing and efficient utilization.
Patent Information
- Application Number
- CN202511544784.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-16
AI Technical Summary
Given limited computing resources, how can we make full use of these resources and achieve computing power sharing to alleviate the problem of resource scarcity?
By obtaining the task description information of the target task, multiple subtasks and their urgency weights and resource requirements are determined. Combined with the status information of the computing power resource pool, computing power resources are dynamically allocated to achieve multi-dimensional resource utilization.
It has enabled the full utilization and sharing of computing resources, improved computing efficiency, and reduced resource fragmentation and waste.
Smart Images

Figure CN121349692A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a method for allocating computing power and related products. BACKGROUND
[0002] With the rapid development of artificial intelligence (AI), large language models (LLM) have gradually become a hotspot of research and application due to their strong language understanding and generation capabilities. More and more enterprises begin to apply them to practical scenarios, such as intelligent customer service, virtual assistants, content creation, content review, machine translation, etc. In view of the above background, it can be predicted that the demand and scale of computing power for artificial intelligence will be higher and higher in the future, and the lack of computing power resources is inevitable. In the case of limited computing power resources, how to make full use of computing power resources and realize computing power sharing to alleviate the lack of computing power resources is a problem to be solved at present. SUMMARY
[0003] Therefore, the present disclosure aims to provide a method for allocating computing power and related products.
[0004] To achieve the above purpose, the present disclosure provides a method for allocating computing power, comprising: obtaining a target task and first task description information corresponding to the target task; determining a plurality of subtasks and second task description information corresponding to each of the subtasks according to the target task and the first task description information; determining an urgency weight and a resource demand combination corresponding to each of the subtasks according to each of the second task description information; wherein the resource demand combination comprises a demand amount of at least one computing power resource; obtaining state information of each computing power resource in a preset computing power resource pool; determining computing power allocation parameter information of each of the subtasks according to the urgency weight, the resource demand combination, and the state information of each of the computing power resources.
[0005] Based on the same inventive concept, the present disclosure also provides a device for allocating computing power, comprising: a first obtaining module configured to obtain a target task and first task description information corresponding to the target task; a decomposition module configured to determine a plurality of subtasks and second task description information corresponding to each of the subtasks according to the target task and the first task description information; a demand module configured to determine an urgency weight and a resource demand combination corresponding to each of the subtasks according to each of the second task description information; wherein the resource demand combination comprises a demand amount of at least one computing power resource; a second obtaining module, configured to obtain state information of each computing resource in a preset computing resource pool; an allocating module, configured to determine computing resource allocation parameter information of each subtask according to the urgency weight, the resource requirement combination and the state information of each computing resource.
[0006] Based on the same inventive concept, the embodiments of the present disclosure further provide an electronic device, comprising a memory, a processor and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the allocation method according to any one of the above embodiments when executing the program.
[0007] Based on the same inventive concept, the embodiments of the present disclosure further provide a non-transitory computer readable storage medium, which stores computer instructions for causing a computer to execute the allocation method according to any one of the above embodiments.
[0008] Based on the same inventive concept, the embodiments of the present disclosure further provide a computer program product, which comprises computer program instructions, and when the computer program instructions run on a computer, the computer program instructions cause the computer to execute the allocation method according to any one of the above embodiments.
[0009] As can be seen from the above, the allocation method of computing resource and related products provided by the present disclosure specifically include obtaining a target task and corresponding first task description information; determining a plurality of subtasks and second task description information corresponding to each subtask according to the target task and the first task description information; determining an urgency weight and a resource requirement combination corresponding to each subtask according to each second task description information; wherein the resource requirement combination includes a demand amount of at least one computing resource; obtaining state information of each computing resource in a preset computing resource pool; and determining computing resource allocation parameter information of each subtask according to the urgency weight, the resource requirement combination and the state information of each computing resource. Such a technical solution determines the computing resource allocation of the target task through the urgency weight, the resource requirement combination of the plurality of subtasks included in the target task and the state information of each computing resource in the computing resource pool, so that the dynamic multi-dimensional computing resource is fully considered in the computing resource allocation, the technical effect of fully utilizing the computing resource is achieved, the computing resource sharing is realized and the lack of computing resource is alleviated. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, brief introductions will be given to the drawings needed to be used in the embodiments or the related art descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without any creative effort.
[0011] Figure 1 A framework structure schematic diagram of a software-defined heterogeneous AI computing resource pool is shown. Figure 2 A structure schematic diagram of a scenario framework provided by the embodiment of the present disclosure is shown. Figure 3 A schematic diagram of an optical character recognition splitting process in a picture provided by the embodiment of the present disclosure is shown. Figure 4 A schematic diagram of multi-file and multi-end task decomposition provided by the embodiment of the present disclosure is shown. Figure 5 A schematic diagram of multi-end computing resource scheduling provided by the embodiment of the present disclosure is shown. Figure 6 A schematic diagram of an end-side computing resource monitoring and management architecture provided by the embodiment of the present disclosure is shown. Figure 7 A cosine similarity algorithm formula provided by the embodiment of the present disclosure is shown. Figure 8 A flowchart of a cross-end heterogeneous computing power allocation method provided by the embodiment of the present disclosure is shown. Figure 9 A structure schematic diagram of a cross-end heterogeneous computing power allocation device provided by the embodiment of the present disclosure is shown. Figure 10 A structure schematic diagram of an electronic device provided by the embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0012] To make the objectives, technical solutions, and advantages of the present disclosure clearer, further detailed descriptions will be given below with reference to the embodiments and the accompanying drawings.
[0013] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood as the general meanings understood by those skilled in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not represent any order, number, or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar terms mean that the elements or objects before the terms cover the elements or objects listed after the terms and their equivalents, without excluding other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right", and the like only represent relative positional relationships, which can change accordingly when the absolute positions of the described objects change.
[0014] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0015] For example, in response to receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0016] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0017] It can be understood that the above notification and obtaining of the authorization of the user are only illustrative, and do not limit the implementation manners of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0018] As described in the background section, in the case of limited computing power resources, how to fully utilize the computing power resources to realize computing power sharing to alleviate the lack of computing power resources is a problem to be solved at present. However, most of the related technologies regard the graphic processing unit (GPU) hardware resources of the computing power cluster as the main factor affecting the performance of the task, ignoring the influence of other dimensional resources such as the center processing unit (CPU), field programmable gate array (FPGA), application specific integrated circuit chip (ASIC), memory, network, etc. in the computing power cluster, especially the emerging heterogeneous computing power processors such as neural network processing unit (NPU) and tensor processing unit (TPU) on the end side, which provide a better solution for tensor processing and neural network processing of large models.
[0019] Therefore, the disclosure provides a cross-end heterogeneous computing power allocation method and related products. The allocation method specifically includes obtaining a target task and corresponding first task description information of the target task; determining a plurality of sub-tasks and second task description information corresponding to each sub-task according to the target task and the first task description information; determining an urgency weight and a resource requirement combination corresponding to each sub-task according to each second task description information; wherein the resource requirement combination includes a demand amount of at least one computing power resource; obtaining state information of each computing power resource in a preset computing power resource pool; and determining computing power allocation parameter information of each sub-task according to each urgency weight, the resource requirement combination, and the state information of each computing power resource. Such a technical solution determines the computing power allocation of the target task through the urgency weight, the resource requirement combination of the plurality of sub-tasks included in the target task, and the state information of each computing power resource in the computing power resource pool, so that the dynamic multi-dimensional computing power resources are fully considered in the computing power allocation, the technical effect of fully utilizing the computing power resources is achieved, the computing power sharing is realized, and the lack of computing power resources is alleviated.
[0020] In order to facilitate understanding of the technical solutions of the disclosure, some related technologies related to the disclosure are introduced below.
[0021] Computing power fusion technology With the popularization of AI applications and the growth of computing demand, the computing power structure presents the characteristics of diversification and fragmentation. In order to effectively integrate various resources and achieve efficient utilization, the fusion of virtualization, containerization and heterogeneous resource pooling technology becomes the key. Among them, through the virtualization technology, the physical server can be divided into multiple independent virtual servers, each virtual server can independently run different operating systems and application programs, so as to realize flexible allocation of resources; the containerization technology allows multiple isolated application instances to run on the same operating system, greatly reducing resource consumption and improving resource utilization, especially suitable for rapid deployment and expansion of AI applications; and the heterogeneous resource pooling technology uniformly manages computing resources of different types and different architectures (such as CPU, GPU, FPGA, ASIC, etc.), and provides a unified interface for upper-layer applications, realizing efficient deployment of resources. The fusion of these technologies not only can build a unified resource management platform, realize cross-platform and cross-device resource scheduling, and significantly improve the performance and efficiency of the overall system.
[0022] Computing power scheduling solution How to make full use of scarce and expensive computing power resources, maximize the sharing of computing power, and reduce the non-distributable fragment probability, the optional idea is to first cut the GPU, AI chip and the like on different nodes through virtualization software, then report the cut resources to the cluster scheduling framework plug-in, and finally the cluster scheduling framework plug-in flexibly schedules and allocates resources according to the task requirements, so that resources can be orderly supplied according to the actual needs of the task.
[0023] In the related art, the solution includes two parts of the cluster scheduling framework plug-in and the node virtualization software.
[0024] In some embodiments, the cluster scheduling framework plug-in: through the high-performance computing power network to open the inter-server channel, so that the computing power resources such as CPU, GPU, AI chip distributed in each server can realize interconnection, transparent sharing through high-speed lossless network. According to the task requirements for resources, through advanced scheduling strategy, the task is scheduled to different nodes. If the task requires the resources of the whole card, it will be scheduled to the node with idle whole card, and if the task requires the resources of the fine-grained card, it will be scheduled to the node with idle virtual card, so as to realize efficient allocation of resources.
[0025] However, whether it is scheduled to the whole card or the virtual card, the cluster scheduling framework plug-in will only allocate the appropriate computing power resources according to the task requirements, and cannot determine whether the task is really using the computing power resources. Exemplarily, the cluster scheduling framework plug-in includes gpushare and elastic-gpu based on K8S open source, GPU Sharing based on Volcano open source, etc.
[0026] In some embodiments, the node virtualization software: through the user mode or kernel mode, the computing power resources are virtualized, which can realize simple or arbitrary ratio cutting of computing power and video memory dimensions, and can realize aggregation of single machine and multiple cards. For multi-machine and multi-card cross-node resource requirements, the task management module is relied on to cut the job into multiple distributed tasks, and then the task is scheduled to the appropriate node by the cluster scheduling framework plug-in, so the node virtualization software and the cluster scheduling framework plug-in must be used at the same time to solve the above problems to some extent. Exemplarily, the node virtualization software includes cGPU, qGPU, GPU virtualization solutions of cloud manufacturers, and open source GPU Manager.
[0027] Computing power pooling solution Compared with the computing power scheduling solution composed of cluster scheduling framework plug-ins and node virtualization software, the computing power resource pooling solution is a resource pooling based on software-defined technology on top of hardware AI computing power, which can realize the decoupling of AI application and computing power. Specifically, AI applications can be deployed arbitrarily, and running AI computing tasks can be hot migrated without being bound to AI servers. AI applications can be deployed on CPU servers, access AI server computing power through remote calls, and use AI computing power on demand throughout the data center. AI applications are truly running when they are allocated appropriate computing power resources from the entire data center AI resource pool. When the AI application execution is completed, the computing power resources are released back to the data center AI resource pool, allowing other AI applications to use them. For AI application cross-node resource requirements, instead of relying on the upper task management module to split the job into multiple tasks, multiple node AI computing resources can be directly aggregated for use.
[0028] Exemplarily, as shown in Figure 1 The heterogeneous AI computing resource pool refers to integrating different types and capabilities of computing resources (such as CPU, GPU, FPGA, ASIC, etc.) together, providing efficient, flexible, and scalable AI computing power services through intelligent scheduling and management. Through software-defined means, the management and configuration of computing power resources are abstracted, providing a flexible programmable computing environment. This allows developers to focus on AI application development and optimization without worrying about underlying hardware details. AI businesses use resource pool computing power resources on demand without the need for restart to adjust the required resources. After the computing power resource pooling, dynamic mounting and dynamic release realize efficient rotation of computing power resources, solving the problems of static allocation, exclusive use, and difficulty in recycling. The computing power pooling scheduling platform provides a variety of scheduling strategies, including but not limited to global resource scheduling, resource location scheduling, heterogeneous resource scheduling, cross-end scheduling, priority scheduling, and label scheduling.
[0029] Intelligent task splitting and merging technology The intelligent task splitting and merging technology is a technology for optimizing the use of computing resources, which can split a large computing task into several small tasks, execute them on different computing nodes, and then merge the results of each node to achieve the final computing goal. By breaking down complex tasks into manageable small tasks and executing them in parallel on multiple computing nodes, computing efficiency is greatly improved. At the same time, through intelligent scheduling and result merging, the consistency and fault tolerance of the computation are guaranteed. The key components of this technology include large task splitting, sub-task scheduling, sub-task execution, and multi-task result merging.
[0030] Among them, large task splitting is to decompose a large computing task into several independent small tasks, and each small task can be executed independently. By splitting the task, the computing resources of multiple devices can be fully utilized to speed up the calculation. According to the nature and characteristics of the task, select the appropriate division strategy. For example, for data parallel tasks, you can divide by data volume; for task parallel tasks, you can divide by the computational complexity of subtasks. Ensure that the computational load of each subtask is roughly the same, and avoid overloading some nodes to affect overall efficiency. Design a reasonable fault tolerance mechanism that can automatically redistribute tasks or resume calculations when a node fails.
[0031] Subtask scheduling is to allocate the split subtasks to different computing nodes for execution, ensuring that subtasks can be efficiently allocated to the most suitable computing nodes. According to the current load, computing power and other factors of the node, dynamically adjust the task allocation strategy. Manage and monitor the resource usage of the computing nodes to ensure that resources are allocated reasonably. According to the urgency and importance of the task, set different priorities and execute high-priority tasks first.
[0032] Subtask execution is the execution of subtasks allocated to each computing node, ensuring that subtasks can be executed correctly and produce the expected results. For different types of computing tasks, optimize local computing processes to improve execution efficiency, save intermediate results generated during subtask execution, and facilitate subsequent merging operations. Optimize the communication mechanism between nodes to reduce data transmission delay and bandwidth consumption.
[0033] Multi-task result merging is to merge the intermediate results generated by each computing node to generate the final computing result, ensuring that the computing results of each subtask can be correctly integrated to achieve the expected computing goal. During the merging process, ensure that the results of all subtasks are consistent to avoid conflicts or omissions. Only merge the parts that have changed since the last merge to reduce the complexity of the merging operation. When an error occurs during the merging process, it can be detected and recovered in time to ensure the correctness of the final result.
[0034] In order to make the technical solutions of the present disclosure clearer and easier to understand, the scene architecture of the cross-end heterogeneous computing power allocation method provided by the embodiments of the present disclosure is introduced below in conjunction with the drawings.
[0035] The cross-end heterogeneous computing power allocation method provided by the embodiments of the present disclosure includes but is not limited to being applied in the application scenarios as shown in Figure 2 , such as Figure 2As shown, the application scenario includes a cross-terminal server 201, a dispatcher 202, a running log database, a resource information database, an original task queue 203, a decomposed task queue 204, and a hardware resource 205. The cross-terminal server 201, the running log database, and the resource information database can be connected through a wired or wireless communication network. The cross-terminal server 201, the running log database, and the resource information database can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. The cross-terminal server 201, the running log database, and the resource information database can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms.
[0036] In some embodiments, the cross-terminal server 201 can dynamically submit original tasks to the original task queue 203 to update the original task queue 203. Here, the cross-terminal server 201 can submit attribute information such as task type, task size, and task priority at the same time as submitting the original task.
[0037] In some embodiments, the cross-terminal server 201 can also obtain the computing power state list maintained by the terminal-side computing power manager.
[0038] It should be noted that the object of terminal-side computing power management is terminal-side computing power. The parameters of terminal-side computing power can be different for different hardware computing power. For example, for CPU, the specific parameters are CPU core number, thread number, architecture platform, and clock frequency (i.e., the number of instructions executed per second), such as an 8-core 16-thread ARM platform at 2.5 GHz; for GPU, the specific parameters are GPU core number, core frequency, and video memory capacity / type; for NPU, the performance data is measured in Tera Operations Per Second (TOPS) per second, which can be obtained by running a benchmark test; for memory, the specific parameters are capacity size and frequency; and for local disk, the specific parameters are capacity size and interface type.
[0039] Figure 6 An architecture schematic diagram of terminal-side computing power resource monitoring and management provided by an embodiment of the present disclosure is shown. As shown in the figure, Figure 6 As shown, the architecture of terminal-side computing power resource monitoring and management includes a computing power index monitor, a terminal-side computing power manager, a computing power state list, and a cross-terminal server. With such an architecture, the underlying heterogeneous computing power chips can be abstracted into a computing power resource pool that can be allocated and used.
[0040] In some embodiments, the computing power index monitor can monitor the computing power resources of each computing power chip in real time, that is, the performance indicators and computing power usage status of each computing power chip can be monitored in real time. Specifically, the computing power index monitor can monitor the usage of computing power chips such as CPU, GPU, NPU, memory, disk, network input / output (Input / Output, IO), and the tasks being executed by the computing power chips. As shown in Figure 6 The computing power index monitor can collect performance indicator data of the computing power, including CPU usage, GPU usage, memory usage, NPU usage, and network IO indicator data.
[0041] Further, the computing power index monitor reports the collected data to the end-side computing power manager over the network. The end-side computing power manager maintains a list of computing power states in the region and finally publishes the list of computing power states to the cross-end server, providing a basis for task computing power scheduling. In addition, by monitoring the performance indicators of the computing nodes in real time, such as CPU utilization, GPU utilization, and memory usage, it is helpful to discover performance bottlenecks and optimize them.
[0042] Optionally, the computing power state list includes at least one computing power resource and state information of the computing power resource. Here, the state information of the computing power resource is used to describe the utilization state of the computing power resource, which can include usage state, idle state, etc. The idle state is convenient to know which computing power resources can participate in distribution and scheduling. For schedulable computing power resources, parameters can be used to describe them, such as the computing power chip to which they belong, the number of idle CPU cores, the number of idle CPU threads, the idle memory capacity, the idle disk capacity, the idle disk type, whether the NPU is idle, whether the GPU is idle, etc.
[0043] As can be seen, the computing power state list provided by the end-side computing power manager provides an important basis for the distribution of cross-end heterogeneous computing power.
[0044] In some embodiments, the state information can include a predicted load rate of the computing power resource; the determination step of the predicted load rate specifically includes: Real-time acquisition of the usage rate of the computing power resource; here, the usage rate can be acquired by means of the computing power index monitor, which will not be described in detail; Sorting and preprocessing the usage rate of the computing power resource in time sequence to obtain a time data set of the usage rate; here, the preprocessing can include removing outliers, filling missing values, etc.
[0045] Based on the time data set, trend analysis is performed to determine the predicted load rate. Here, trend analysis can use naive extrapolation, state space / Kalman filtering, machine learning regression, deep learning, etc., which are not limited by the present disclosure.
[0046] The technical solution can effectively allocate future idle resources by predicting the load rate, reduce resource idling and overuse, and enhance the stability and reliability of the computing power allocation.
[0047] By monitoring and predicting the usage rate of various types of computing power resources, it is beneficial to subsequently combine resource demand data in the task queue to dynamically predict resource demand and analyze state changes, adjust computing resource allocation parameters, enhance resource utilization flexibility, and improve task processing response speed and adaptability.
[0048] In some embodiments, the running log database can be responsible for collecting multi-dimensional resource usage, execution time, task type, and other task log data of the task. It should be noted that the task log data can assist in task scheduling as historical data, and train resource demand prediction models, etc. Here, Figure 2 The task model in the task model database can include a resource demand prediction model for predicting the impact of multi-dimensional resource demand on performance and guiding the implementation of adaptive resource adjustment of tasks in the resource allocation process.
[0049] The resource information database can store the computing power state list obtained by the cross-end server 201. It should be noted that the resource model can be based on the computing power state list to determine the computing power resources available to the task, providing a reference for task scheduling.
[0050] In some embodiments, in the current allocation round, first, the scheduler 202 can send a task request to the original task queue 203 to obtain a target task. Next, the scheduler 202 can determine the resource demand of the target task using the resource demand prediction model, which can adaptively adjust the multi-dimensional resource demand based on the attribute information of the target task. Then, the scheduler 202 can deploy the target task to a computing power device (such as hardware resource 205) that can meet the resource demand based on the resource demand of the target task and the state information of the computing power resource provided by the resource model, realizing dynamic allocation of cross-end computing power resources. Finally, update the state information of the hardware resource 205 and the execution state information of the target task for the next round of allocation. Here, the updated state information of the hardware resource 205 can be recorded in the resource information database, and the execution state information of the target task can be recorded in the running log database.
[0051] It should be noted that the hardware resource 205 can include CPU, GPU, NPU, TPU, FPGA, etc. Here, CPU, GPU, NPU, TPU, FPGA are not limited to a single device, but can come from multiple devices (such as local devices, cloud devices, etc.).
[0052] The technical solution considers the resource requirement characteristics of different attribute tasks and full utilization of cross-end multi-dimensional resources, and can solve the mismatching problem of tasks among multiple dimensional resources, and improve the deployment efficiency of cross-end tasks and the utilization rate of multi-dimensional resources.
[0053] In some embodiments, in order to facilitate the allocation of computing power resources, the target task can be decomposed into multiple subtasks, and the scheduler 202 can allocate computing power resources for each subtask and update to the decomposed task queue 204. By decomposing the target task into at least one subtask, it is helpful to balance the distribution of computing tasks among different nodes in the heterogeneous AI computing resource pool, avoid the situation that some nodes are overloaded while other nodes are idle, and provide overall computing efficiency.
[0054] The application scenarios of the resource requirement prediction model and the cross-end heterogeneous computing power allocation method according to the example embodiments of the present disclosure will be described below. Figure 2 It should be noted that the above application scenarios are only shown for the purpose of facilitating the understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0055] For intelligent computing power allocation, the task resource requirement specifically includes one or more of GPU application amount, CPU application amount, NPU application amount, TPU application amount, memory application amount, network bandwidth, IO occupation, etc. Therefore, the task resource requirement involved in the allocation process focuses on multiple dimensions such as GPU, CPU, NPU, TPU, memory, network, and IO. It should be understood that the utilization efficiency of a task for a single resource is affected by other dimensional resources, and therefore, the embodiments of the present disclosure provide a resource requirement prediction model for a task. The resource requirement prediction model can predict the resource requirement of a single complete task, or predict the resource requirement of each subtask belonging to a task, which is not limited by the present disclosure.
[0056] In some embodiments, the resource requirement prediction model can be constructed by the following steps: First, a plurality of sample tasks and their corresponding description information are obtained. Here, the description information can be used to describe the attributes of the sample tasks, and can include task urgency weight, task size, computing complexity, data throughput, environment data such as network condition, geographic location, timestamp, etc.
[0057] Next, run the sample task on the CPU, GPU, NPU, TPU, memory, etc. power resource alone, determine the running time of the sample task and the resource demand test value occupied. Based on the selected multiple sample tasks and their multi-dimensional resource demand test values, a resource regression model is constructed. Here, if the historical running data of the task with the same task description information as the historical data record, such as CPU usage, memory occupancy, disk I / O, etc., can directly use the historical data, and the present disclosure does not limit this.
[0058] It should be noted that the characteristic value of the resource regression model can be the number of CPU cores, GPU percentage occupancy, NPU power value, memory capacity demand, network bandwidth, IO occupancy, etc. The target value Y of the resource regression model j may be the execution time of the jth sample.
[0059] Then, based on the resource regression model and the preset range of the execution time of the sample task, the sample resource demand characteristics corresponding to the sample task are determined. The sample task, sample task description information, and sample resource demand characteristics can form a first sample, and multiple first samples can construct a power data set. It should be noted that the preset range of the execution time of the sample task can be the minimum execution time, or the minimum execution time + 10% of the minimum execution time. In other words, the preset range of the execution time can be set based on the demand of power processing, and the present disclosure does not limit this. It should be understood that the larger the preset range of the execution time, the smaller the limit on resource demand, that is, there is more possibility of resource allocation.
[0060] Finally, the initial model is trained based on the power data set to obtain a resource demand prediction model. Here, the initial model can be a neural network model, a random forest, a support vector machine, etc., and the present disclosure does not limit this.
[0061] Exemplarily, the step of training the initial model can include: data division, dividing the power data set into a training set, a validation set, and a test set; hyperparameter tuning, adjusting the model parameters by cross-validation method to obtain the best performance.
[0062] Using the resource demand prediction model, the resource demand can be predicted for different tasks, improving the efficiency of processing multiple tasks, optimizing the calculation speed and energy consumption, and reducing network bandwidth and traffic loss.
[0063] Next, the allocation method of cross-end heterogeneous computing power is described in detail taking the picture recognition scene as an example.
[0064] First, for the local large picture, the number of rows is large, for example, the picture contains more than 20 lines of text, build an OCR recognition model, use a single CPU single-thread serial processing, it takes about 11 seconds to process on the RK3588 board. If the recognition process is decomposed and the computing power is allocated, the recognition time can be effectively shortened.
[0065] Figure 3 A picture optical character recognition splitting process schematic diagram provided by the embodiments of the present disclosure is shown. As shown in Figure 3 The picture area is divided into three areas, area 301, area 302, and area 303, and the data of the three areas is input into three threads, and the text line area is recognized at the same time. After recognition, each area will draw the area of the text line recognized by the area, for example, area 301 recognizes 4 lines of text, which are line 3011, line 3012, line 3013, and line 3014. The four lines of text are input into four threads, and single-character recognition is performed at the same time. For single-line text, it can be further divided, for example, line 3012 is divided into three sentences, which are sentence 301211, sentence 30122, and sentence 30123. The three sentences are input into three threads, and single-character recognition is performed at the same time.
[0066] From Figure 3 It can be seen that by dividing a task into small tasks, parallel recognition can be used instead of serial recognition. Multiple small tasks can run on different threads. Different threads can run on different computing resources, such as CPU+GPU.
[0067] For multi-end multi-file scenarios, assuming that there are many pictures to be recognized, batch processing is required. In addition to local multi-thread processing optimization, cross-end collaborative processing can also be performed according to the networking device state. Suppose 100 jpg pictures, each 300 kB, are transmitted through a local area network, 20 pictures are 6M, and it takes about 500 milliseconds to transmit under Wifi direct connection. Processing a single OCR recognition including more than 20 lines takes more than 6 seconds, and 20 pictures take at least 100 seconds. The transmission time can be completely ignored. Figure 4 A multi-file, multi-end task splitting schematic diagram provided by the embodiments of the present disclosure is shown. As shown in Figure 4 The file group 4011 can be divided into multiple threads, which can run on multiple ends and different computing resources, such as running on the master device 401, the slave devices 402, 403, 404, and 405. Here, Figure 5 A multi-end computing resource scheduling schematic diagram provided by the embodiments of the present disclosure is shown. As shown in Figure 5 The multiple ends can be local ends or cloud ends, which are not limited by the present disclosure.
[0068] Furthermore, tasks can be further allocated based on the computing resources (CPU, GPU, memory, etc.) of each device. After the operation is completed, the results are summarized and returned to the original device for final filtering and processing.
[0069] It should be understood that breaking down large tasks into relatively independent smaller tasks (corresponding to subtasks) allows for advance processing based on business workflows. However, the selection of computing power equipment and the allocation of subtasks require dynamic allocation based on the resource requirements of the subtasks and the real-time operational status of computing resources. Here, when breaking down large tasks, the descriptive information of the subtasks is simultaneously determined.
[0070] Next, using the aforementioned resource demand prediction model, the resource demand combination for each subtask can be determined. This resource demand combination includes the demand for at least one computing resource (e.g., GPU, CPU), such as GPU requests, CPU requests, NPU requests, TPU requests, memory requests, network bandwidth, and I / O usage.
[0071] Then, the status information of each computing resource is obtained, and the computing power allocation parameter information of each subtask is determined based on the urgency weight, resource demand combination and status information of each computing resource.
[0072] In some embodiments, the step of determining the alternative computing resources for the sub-task based on the resource requirement combination and the status information of each computing resource may specifically include: Based on the urgency weights of multiple subtasks, the combination of resource requirements, and alternative computing resources, a preset algorithm is used to determine the matching degree between the target task and the computing resource pool.
[0073] Optionally, the preset algorithm includes a cosine similarity algorithm. For example... Figure 7 As shown, S is the weighted cosine similarity score, used to quantify the matching degree between the target task and the resource. Ai is the resource usage of subtask i, representing the amount of target resources consumed by the subtask during execution. Bi is the available amount of resource type i, representing the total amount of resource type i currently available for allocation. Pi is the priority weight of resource i, used to adjust the importance of resource type i in the matching process. Ui is the urgency weight of subtask i, used to adjust the urgency of subtask i in resource allocation. n is the number of subtasks and resource types involved, used to define the range of the summation operation. i is an index variable used to represent different task or resource types.
[0074] Exemplarily, the specific execution process can include the following steps: for each subtask i, collecting resource usage amount Ai (which can be determined based on resource requirement combination), resource available amount Bi (which can be determined based on alternative computing resource), priority Pi of corresponding weight factor resource (which can be determined based on alternative computing resource, and state information can include priority weight), and urgency Ui of subtask (which can correspond to urgency weight), substituting the parameters into the formula to calculate the matching degree S, which is used to evaluate the matching degree of the target task and the plurality of computing resources. Here, Pi and Ui are used to adjust the calculation importance of resources and subtasks, and resources with high priority and subtasks with high urgency obtain greater consideration in the allocation process. Considering the number, importance and urgency of resources and subtasks, it is ensured that the plurality of resources are allocated to match the subtask demand.
[0075] In response to determining that the matching degree meets the preset threshold, the computing power allocation parameter information of each subtask is determined based on the alternative computing resource. Here, when the matching degree meets the preset threshold, it indicates that the alternative computing power resource can meet the computing power resource demand of each subtask, and thus the computing power allocation parameter information of each subtask can be determined.
[0076] In response to determining that the matching degree does not meet the preset threshold, the alternative computing power resource of the subtask is re-determined. Here, when the matching degree does not meet the preset threshold, it indicates that the alternative computing power resource cannot meet the computing power resource demand of each subtask, and therefore the alternative computing power resource needs to be re-determined, and the matching degree needs to be re-calculated until the matching degree meets the preset threshold.
[0077] It should be understood that, in the case of alternative computing power resource being unique, the division of subtasks can be re-performed (for example, further split), or the supply of computing power resource can be prompted to be increased, for example, cloud devices are introduced.
[0078] Finally, based on the computing power allocation parameter information of each subtask, each subtask is allocated to the matched computing power resource.
[0079] Optionally, the allocation of subtasks can be implemented by means of OpenCL computing power framework. Here, in the OpenCL computing power framework, dynamic scheduling includes three main runtime modules: device analyzer, which is used to collect and analyze the performance (such as memory, computing power and I / O) of the device; kernel analyzer, which analyzes and predicts the execution time of the kernel on different devices; task scheduler, which schedules tasks (such as subtasks) in the command queue marked with a scheduling strategy to the device. The dynamic scheduling of OpenCL tasks is implemented based on the above three runtime modules. The device analyzer and the kernel analyzer analyze the OpenCL device and the kernel respectively, judge the degree of fit between the OpenCL device and the kernel, and provide data basis for the task scheduler.
[0080] Therefore, the technical solution provided by the embodiments of the present disclosure considers not only GPU but also CPU, network, I / O, etc. during the allocation of computing power, fully utilizes the interconnection and intercommunication capability between cross-terminal devices, and improves the processing speed of tasks.
[0081] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of the present embodiment can also be applied to a distributed scenario, and be completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the present disclosure, and the multiple devices can interact with each other to complete the method.
[0082] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order described above and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0083] Based on the same inventive concept, the present disclosure also provides a computing power allocation method corresponding to any of the above-mentioned embodiment methods. As shown in Figure 8 the allocation method specifically includes: S801: obtaining a target task and first task description information corresponding to the target task; as Figure 2 shown, the target task can come from the original task queue 203, and the first task description information can include attribute information such as task urgency weight, task size, calculation complexity, data throughput, and environment data, which is not limited by the present disclosure; S803: determining a plurality of subtasks and second task description information corresponding to each of the subtasks according to the target task and the first task description information; the method of splitting the target task into subtasks can be referred to as described above, and will not be described in detail; based on the situation of the target task, an appropriate task splitting method can also be selected, which is not limited by the present disclosure; S805: determining an urgency weight and a resource demand combination corresponding to each of the subtasks according to each of the second task description information; wherein the resource demand combination includes a demand amount of at least one computing power resource; here, the second task description information can include the urgency weight, or the urgency weight can be determined based on the second task description information, for example, some description words (such as basic and bottom layer) correspond to a high urgency weight, and some description words (such as extension) correspond to a low urgency weight; the resource demand combination can be determined based on the resource demand prediction model described above; S807: Obtain state information of each computing resource in the preset computing resource pool; here, the state information can be derived from the computing state list, which can refer to the foregoing related description, and will not be repeated here; Figure 6 S809: Determine the computing allocation parameter information of each subtask according to the emergency weight, the resource requirement combination, and the state information of each computing resource.
[0084] In some embodiments, S803 specifically includes: comparing the second task description information with preset historical data; wherein the historical data includes at least one historical description information and its corresponding historical resource requirement combination; here, as shown in the figure, the historical data can be recorded in a running log database; Figure 2 in response to determining that the second task description information and any of the historical description information match successfully, determining the resource requirement combination of the corresponding subtask based on the corresponding historical resource requirement combination; in response to determining that the second task description information and each of the historical description information all fail to match, determining the resource requirement combination of the subtask based on the second task description information by using a resource requirement prediction model.
[0085] With such a technical solution, for the matched subtask, the resource requirement combination can be determined without going through the resource requirement prediction model, which helps to save allocation time and improve allocation efficiency.
[0086] In some embodiments, the resource requirement preset model is obtained by training an initial model using a computing data set; wherein the computing data set includes a plurality of first samples; wherein the first sample includes sample task description information and sample resource requirement features corresponding to the execution time of the sample task within a preset range.
[0087] In some embodiments, the sample resource requirement features are determined based on a resource regression model; the resource regression model is determined based on the sample task and its corresponding multi-dimensional resource requirement test value.
[0088] Here, the multi-dimensional resource can include at least one of CPU, GPU, FPGA, ASIC, NPU, TPU, network, and IO; the test value can be the usage; it should be noted that the process of training the resource requirement preset model can refer to the foregoing, and will not be repeated here.
[0089] In some embodiments, the state information includes the usage rate of the computing resource; The state information of each computing resource in the preset computing resource pool is acquired, and specifically includes: Referring to Figure 6 , an algorithm state list issued by an end-side computing power manager is acquired; the algorithm state list includes at least one computing resource and state information of the computing resource; the state information is acquired by an algorithm index monitor.
[0090] In some embodiments, the state information includes a predicted load rate of the computing resource; and the determination of the predicted load rate specifically includes: The usage rate of the computing resource is acquired in real time; The usage rate of the computing resource is sorted in time sequence and preprocessed to obtain a time data set of the usage rate; Based on the time data set, trend analysis is performed to determine the predicted load rate.
[0091] In some embodiments, S809 specifically includes: According to the resource requirement combination and the state information of each computing resource, the alternative computing resource of the subtask is determined; According to the urgency weight of the plurality of subtasks, the resource requirement combination and the alternative computing resource, a preset algorithm is used to determine the matching degree of the target task and the computing resource pool; optionally, the preset algorithm includes a cosine similarity algorithm; In response to determining that the matching degree meets a preset threshold, the algorithm allocation parameter information of each subtask is determined based on the alternative computing resource; In response to determining that the matching degree does not meet the preset threshold, the alternative computing resource of the subtask is re-determined.
[0092] In some embodiments, the hardware corresponding to the computing resource pool includes at least one of a CPU, a GPU, an FPGA, an ASIC, an NPU, a TPU, a network and an IO; and the computing resource includes at least one of a computing resource provided by a CPU, a GPU, an FPGA, an ASIC, an NPU, a TPU, a network and an IO.
[0093] In some embodiments, the computing resource of the computing resource pool includes resources of a virtual node and resources of a physical node; and / or Referring to Figure 5 As shown, the computing resource of the computing resource pool includes a computing resource provided by a local device and a computing resource provided by a cloud device.
[0094] In some embodiments, the target task corresponds to at least one target object; and the subtasks include minimum parallel task units divided based on the target object. Here, each subtask is divided into minimum parallel task units, and the resource requirement of each subtask is small, facilitating the allocation of computing power resources.
[0095] Based on the same inventive concept, the disclosure also provides an allocation device of computing power corresponding to any of the above-mentioned embodiment methods.
[0096] Reference Figure 9 The allocation device comprises: A first acquisition module 901 configured to acquire a target task and first task description information corresponding to the target task; A decomposition module 903 configured to determine a plurality of subtasks and second task description information corresponding to each subtask according to the target task and the first task description information; A requirement module 905 configured to determine an urgency weight and a resource requirement combination corresponding to each subtask according to each second task description information; wherein the resource requirement combination includes a requirement amount of at least one computing power resource; A second acquisition module 907 configured to acquire state information of each computing power resource in a preset computing power resource pool; An allocation module 909 configured to determine computing power allocation parameter information of each subtask according to each urgency weight, the resource requirement combination, and the state information of each computing power resource.
[0097] For the convenience of description, the above device is described as various modules respectively described in terms of functions. Of course, when implementing the disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0098] The device of the above-mentioned embodiments is used to implement the corresponding allocation method in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.
[0099] Based on the same inventive concept, the disclosure also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the allocation method of any of the above-mentioned embodiments when executing the program.
[0100] Figure 10A more specific electronic device hardware structure schematic diagram provided by the embodiment is shown. The device can include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for internal communication.
[0101] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present specification.
[0102] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0103] The input / output interface 1030 is configured to connect an input / output module to realize information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0104] The communication interface 1040 is configured to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0105] The bus 1050 includes a channel to transmit information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.
[0106] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain components necessary to implement the embodiments of the present application, and does not necessarily contain all the components shown in the figure.
[0107] The electronic device of the above embodiment is used to implement the corresponding allocation method in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0108] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the allocation method as claimed in any of the above embodiments.
[0109] The computer-readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0110] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to perform the allocation method as claimed in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0111] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a computer program product including computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the allocation method. Corresponding to the execution subject of each step in each embodiment of the allocation method, the processor performing the corresponding step can belong to the corresponding execution subject.
[0112] The computer program product of the above embodiments is used to make the computer and / or the processor execute the allocation method as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0113] It should be understood by those of ordinary skill in the art that the discussion of any of the above embodiments is merely exemplary and is not intended to suggest that the scope of the disclosure (including the claims) is limited to these examples; the above embodiments or technical features among different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes to the aspects of the embodiments of the disclosure as described above, which are not provided in detail for the sake of brevity. It is intended that each of the individual aspects of the embodiments of the disclosure, as well as any combination of the aspects, be considered individually and in combination.
[0114] In addition, in order to simplify the description and discussion, and so as not to make the embodiments of the disclosure difficult to understand, the well-known power / ground connections of integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. In addition, the devices can be shown in the form of block diagrams in order to avoid making the embodiments of the disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented to implement the embodiments of the disclosure (i.e., these details should be fully within the understanding of those skilled in the art). Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the disclosure, it will be apparent to those skilled in the art that the embodiments of the disclosure can be practiced without these specific details or with variations on these specific details. Therefore, these descriptions should be considered as illustrative rather than limiting.
[0115] Although the disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0116] The embodiments of the disclosure are intended to cover all such alternatives, modifications and variations as falling within the broad scope of the appended claims. Accordingly, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the disclosure shall be included in the protection scope of the disclosure.
Claims
1. A method of distributing computational effort, characterized by, The method comprises: obtaining a target task and corresponding first task description information thereof; determining a plurality of subtasks and corresponding second task description information of each subtask according to the target task and the first task description information; determining an urgency weight and a resource requirement combination corresponding to each subtask according to each second task description information; wherein the resource requirement combination comprises a demand amount of at least one computing resource; obtaining state information of each computing resource in a preset computing resource pool; determining computing resource allocation parameter information of each subtask according to the urgency weight, the resource requirement combination, and the state information of each computing resource.
2. The dispensing method of claim 1, wherein, The determination of the urgency weight and the resource requirement combination corresponding to each subtask according to each second task description information comprises: comparing the second task description information with preset historical data; wherein the historical data comprises at least one historical description information and a corresponding historical resource requirement combination; in response to determining that the second task description information and any historical description information match successfully, determining the resource requirement combination of the corresponding subtask based on the corresponding historical resource requirement combination; in response to determining that the second task description information and each historical description information all fail to match, determining the resource requirement combination of the subtask based on the second task description information using a resource requirement prediction model.
3. The dispensing method of claim 2, wherein, The resource requirement prediction model is obtained by training an initial model using a computing resource data set; wherein the computing resource data set comprises a plurality of first samples; wherein the first sample comprises sample task description information and sample resource requirement features corresponding to the execution time of the sample task within a preset range.
4. The dispensing method of claim 3, wherein, The sample resource requirement features are determined based on a resource regression model; the resource regression model is determined based on the sample task and corresponding multi-dimensional resource requirement test values.
5. The dispensing method of claim 1, wherein, The state information comprises a usage rate of the computing resource; The obtaining of the state information of each computing resource in the preset computing resource pool comprises: obtaining a computing resource state list published by an end-side computing resource manager; wherein the computing resource state list comprises at least one computing resource and state information of the computing resource; wherein the state information is obtained by a computing resource index monitor.
6. The dispensing method of claim 5, wherein, The state information comprises a predicted load rate of the computing resource; the determination of the predicted load rate comprises: real-time obtaining of the usage rate of the computing resource; sorting and preprocessing the usage rate of the computing resource in time sequence to obtain a time data set of the usage rate; performing trend analysis based on the time data set to determine the predicted load rate.
7. The method of distributing according to claim 1, wherein, The determination of the computing resource allocation parameter information of each subtask according to the urgency weight, the resource requirement combination, and the state information of each computing resource comprises: determining alternative computing resources of the subtask according to the resource requirement combination and the state information of each computing resource; determining a matching degree of the target task and the computing resource pool using a preset algorithm according to the urgency weight, the resource requirement combination, and the alternative computing resources of a plurality of subtasks. In response to determining that the matching degree meets a preset threshold, determine, based on the candidate computing resource, computing resource allocation parameter information of each of the sub-tasks; In response to determining that the matching degree does not meet the preset threshold, redetermine the candidate computing resource of the sub-tasks.
8. The dispensing method of claim 7, wherein, The preset algorithm includes a cosine similarity algorithm.
9. The dispensing method of claim 1, wherein, The hardware corresponding to the computing resource pool includes at least one of a CPU, a GPU, an FPGA, an ASIC, an NPU, a TPU, a network, and an IO; and the computing resource includes at least one of a CPU, a GPU, an FPGA, an ASIC, an NPU, a TPU, a network, and an IO.
10. The dispensing method of claim 1, wherein, The computing resource of the computing resource pool includes resources of a virtual node and resources of a physical node; and / or The computing resource of the computing resource pool includes computing resources provided by a local device and computing resources provided by a cloud device.
11. The dispensing method of claim 1, wherein, The target task corresponds to at least one target object; and the sub-task includes a minimum parallel task unit divided based on the target object.
12. An apparatus for allocating computational effort, characterized by The method comprises: The first obtaining module is configured to obtain a target task and first task description information corresponding to the target task; The decomposition module is configured to determine, according to the target task and the first task description information, a plurality of sub-tasks and second task description information corresponding to each of the sub-tasks; The demand module is configured to determine, according to each of the second task description information, an urgency weight and a resource demand combination corresponding to each of the sub-tasks; wherein the resource demand combination includes a demand amount of at least one computing resource; The second obtaining module is configured to obtain state information of each computing resource in a preset computing resource pool; The allocation module is configured to determine, according to each of the urgency weight, the resource demand combination, and the state information of each of the computing resources, computing resource allocation parameter information of each of the sub-tasks.
13. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, wherein, The processor, when executing the computer program, implements the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium, comprising: The non-transitory computer readable storage medium stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 11.
15. A computer program product, characterised in that, The computer program instructions, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 11.