Constraint-based cloud resource scheduling optimization method, computer device and medium
By obtaining the latency requirements of user terminals and the target ratio of the number of terminals, cloud resources are allocated reasonably, solving the waste problem caused by unreasonable cloud resource configuration, and realizing the reduction of cloud resource costs and the improvement of utilization.
Patent Information
- Application Number
- CN202511270435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-08
AI Technical Summary
In existing technologies, unreasonable cloud resource allocation leads to waste and increases unnecessary cloud resource costs.
By obtaining the latency requirements of user terminals and the target proportion of the number of terminals, the cloud resource allocation ratio is determined, and cloud resources are allocated reasonably with the total cost of the cloud resource pool as the target.
It effectively reduced cloud resource costs and improved cloud resource utilization.
Smart Images

Figure CN120785896B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of end-to-cloud fusion processing, and particularly relates to a cloud resource scheduling optimization method based on constraint conditions, a computer device and a medium. BACKGROUND
[0002] In related technologies, when computing tasks of user terminals are processed by using end-to-cloud fusion technology, data processing is usually performed according to a cloud resource configuration specified in advance. However, the cloud resource configuration specified in advance is often unreasonable, causing some allocated cloud resources to be wasted and increasing unnecessary cloud resource costs.
[0003] Therefore, how to reasonably configure cloud resources to reduce cloud resource costs is a problem that needs to be solved at present. SUMMARY
[0004] Embodiments of the present application provide a cloud resource scheduling optimization method based on constraint conditions, a computer device and a medium, aiming to make cloud resource allocation more reasonable and thus reduce cloud resource costs.
[0005] In a first aspect, embodiments of the present application provide a cloud resource scheduling optimization method based on constraint conditions, which comprises:
[0006] obtaining delay requirement gears of a plurality of user terminals of a cloud resource pool for computing tasks respectively;
[0007] obtaining a terminal quantity target proportion of the plurality of user terminals that need to meet the delay requirement gears;
[0008] determining cloud resource allocation proportions among the plurality of user terminals based on the delay requirement gears and the terminal quantity target proportion, with the goal of constraining a total cloud resource cost of the cloud resource pool;
[0009] allocating cloud resources to the plurality of user terminals according to the cloud resource allocation proportions.
[0010] In some embodiments, the obtaining of the delay requirement gears of the plurality of user terminals of the cloud resource pool for computing tasks respectively comprises:
[0011] obtaining historical computing tasks of each of the user terminals in a historical period;
[0012] confirming a task type of each of the historical computing tasks, wherein the task type comprises at least one of real-time video rendering, interactive data analysis and offline batch processing;
[0013] determining the delay requirement gear of each of the user terminals based on the task type.
[0014] In some embodiments, the determining the latency requirement level of each user terminal based on the task type comprises:
[0015] For each user terminal, determining a target task type with the largest number of historical computing tasks in the task types of the historical computing tasks of the user terminal;
[0016] Determining the latency requirement level of the user terminal based on the target task type.
[0017] In some embodiments, the determining the cloud resource allocation ratio between the user terminals based on the latency requirement level and the target terminal quantity ratio, with the objective of constraining the total cloud resource cost of the cloud resource pool, comprises:
[0018] Determining a target minimum latency pre-associated with each latency requirement level;
[0019] Determining a computing resource cost of the corresponding user terminal based on the target minimum latency and the target task type of the corresponding user terminal, wherein the total cloud resource cost is equal to the sum of the computing resource costs of the user terminals of the cloud resource pool;
[0020] Determining the product of the total number of user terminals of the cloud resource pool and the target terminal quantity ratio as a target user terminal quantity;
[0021] Determining the cloud resource allocation ratio based on the computing resource cost and the target user terminal quantity, with the objective of constraining the total cloud resource cost of the cloud resource pool.
[0022] In some embodiments, the determining the computing resource cost of the corresponding user terminal based on the target minimum latency and the target task type of the corresponding user terminal comprises:
[0023] Obtaining a cost prediction model pre-associated with the target task type of the corresponding user terminal;
[0024] Inputting the target minimum latency and the target task type of the corresponding user terminal into the cost prediction model;
[0025] Receiving a cost value output by the cost prediction model as the computing resource cost of the corresponding user terminal.
[0026] In some embodiments, the cost prediction model is generated by:
[0027] obtain a plurality of historical computing tasks belonging to the target task type, a historical processing delay and a historical resource cost of each of the historical computing tasks;
[0028] perform model training based on the historical processing delay and the historical resource cost to obtain the cost prediction model.
[0029] In some embodiments, the cloud resource allocation ratio is determined based on the computing resource cost and the target number of user terminals, with the objective of constraining the total cost of cloud resources of the cloud resource pool.
[0030] Among the plurality of user terminals, a plurality of first user terminals matching the target number of user terminals are determined based on the computing resource cost.
[0031] Among the plurality of user terminals, user terminals other than the first user terminals are all regarded as second user terminals.
[0032] The computing resource cost of each of the second user terminals is reduced.
[0033] The cloud resource allocation ratio is determined based on the computing resource cost of each of the first user terminals and the reduced computing resource cost of each of the second user terminals.
[0034] In some embodiments, the target proportion of terminal quantity required to meet the corresponding delay requirement level among the plurality of user terminals is obtained as follows:
[0035] Among a plurality of preset time periods, a target time period in which a current time point is located is determined.
[0036] A preset proportion associated with the target time period is taken as the target proportion of terminal quantity.
[0037] In a second aspect, embodiments of the present application provide a cloud resource scheduling optimization apparatus based on constraint conditions, which comprises:
[0038] An obtaining module is configured to obtain delay requirement levels of a plurality of user terminals of a cloud resource pool for computing tasks, and obtain a target proportion of terminal quantity required to meet the corresponding delay requirement level among the plurality of user terminals.
[0039] A determining module is configured to determine a cloud resource allocation ratio among the plurality of user terminals based on the delay requirement level and the target proportion of terminal quantity, with the objective of constraining the total cost of cloud resources of the cloud resource pool.
[0040] An allocating module is configured to allocate cloud resources to the plurality of user terminals according to the cloud resource allocation ratio.
[0041] In a third aspect, embodiments of the present application provide a computer device, comprising a processor and a memory, wherein the memory stores a computer program configured to be executed by the processor to implement the constraint-based cloud resource scheduling optimization method according to any one of the above.
[0042] In a fourth aspect, embodiments of the present application provide a computer readable storage medium, which stores a computer program configured to be executed by a processor to implement the constraint-based cloud resource scheduling optimization method according to any one of the above.
[0043] In a fifth aspect, embodiments of the present application provide a computer program product comprising computer programs or instructions, which are executed by a processor to implement the constraint-based cloud resource scheduling optimization method according to any one of the above.
[0044] The beneficial effects of embodiments of the present application are as follows:
[0045] In embodiments of the present application, the cloud resource allocation ratio among the plurality of user terminals is determined based on the delay requirement gear of each user terminal of the cloud resource pool for the computing task, and the target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement gear, so as to constrain the total cost of the cloud resources of the cloud resource pool, thereby the cloud resources can be reasonably allocated by using the constraint of the total cost of the cloud resources to reduce the cost of the cloud resources. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0047] Figure 1 is an embodiment flowchart of the constraint-based cloud resource scheduling optimization method provided by the embodiments of the present application;
[0048] Figure 2 is another embodiment flowchart of the constraint-based cloud resource scheduling optimization method provided by the embodiments of the present application;
[0049] Figure 3 is still another embodiment flowchart of the constraint-based cloud resource scheduling optimization method provided by the embodiments of the present application;
[0050] Figure 4is an embodiment structure schematic diagram of the cloud resource scheduling optimization device based on a constraint condition provided by an embodiment of the present application.
[0051] Figure 5 is an embodiment structure schematic diagram of the computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0052] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person skilled in the art without creative work fall within the scope of protection of the present application.
[0053] In the description of the present application, the meaning of "multiple" is two or more than two, unless otherwise explicitly and specifically limited. In addition, in the description of the present application, the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include one or more of the features.
[0054] In a first aspect, the embodiments of the present application provide a cloud resource scheduling optimization method based on a constraint condition. Referring to Figure 1 , Figure 1 is an embodiment flow schematic diagram of the cloud resource scheduling optimization method based on a constraint condition. In Figure 1 , the cloud resource scheduling optimization method based on a constraint condition can include:
[0055] 101, obtaining delay requirement gears of multiple user terminals of a cloud resource pool for a computing task, respectively.
[0056] In the embodiments of the present application, end-cloud fusion is implemented based on a cloud resource pool. The cloud resource pool refers to a resource pool that aggregates computing resources of cloud servers. The computing resources may, for example, include CPU (Central Processing Unit) resources, GPU (Graphics Processing Unit) resources, and other cloud resources of the cloud servers. The multiple user terminals of the cloud resource pool refer to multiple user terminals connected to the cloud resource pool, so as to implement data processing of the computing task by using end-cloud fusion technology. The user terminal refers to a terminal device of a user, and each user can trigger the computing task through the corresponding terminal device. Taking the interior design industry as an example, the computing task can include the generation task of a virtual reality three-dimensional picture related to interior design. The computing task can implement data processing by using end-cloud fusion technology to obtain the computing result of the computing task.
[0057] In the embodiments of the present application, each user terminal has a corresponding delay requirement for a computing task. The delay requirement level refers to the level of the delay requirement. For example, different levels of delay requirements can be classified in advance to obtain a plurality of preset delay requirement levels. The plurality of delay requirement levels can include high, medium, low, and the like.
[0058] 102. Obtain a target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement level.
[0059] In the embodiments of the present application, in an ideal case, the data processing delay of each user terminal for a computing task needs to meet the delay requirement of the corresponding delay requirement level to ensure the high performance of the user terminal. However, for the purpose of reducing the cost of cloud resources, at least part of the user terminals need to meet the corresponding delay requirement level, and a corresponding target proportion of the number of terminals is set.
[0060] In some embodiments of the present application, obtaining a target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement level can include: determining a target time period in which a current time point is located in a plurality of preset time periods, wherein the plurality of preset time periods can include a cloud resource use peak time period, a cloud resource use valley time period, and the like; and a preset proportion associated with the target time period is used as the target proportion of the number of terminals, for example, a preset proportion associated with the cloud resource use peak time period can be greater than a preset proportion associated with the cloud resource use valley time period, so as to meet the cloud resource use demand of more user terminals at the same time.
[0061] 103. Based on the delay requirement level and the target proportion of the number of terminals, determine a cloud resource allocation proportion between the plurality of user terminals, with the constraint of the total cost of cloud resources of the cloud resource pool as the target.
[0062] In the embodiments of the present application, the total cost of cloud resources of the cloud resource pool refers to the cost of the cloud resources allocated to the plurality of user terminals in the cloud resource pool. The total cost of cloud resources can be positively correlated with the total amount of cloud resources allocated to the plurality of user terminals in the cloud resource pool. The specific calculation method of the total cost of cloud resources can be set based on actual needs, which is not limited herein, for example, the total amount of cloud resources allocated to the plurality of user terminals in the cloud resource pool can be directly used as the total cost of cloud resources of the cloud resource pool.
[0063] In the embodiments of the present application, the target of constraining the total cost of cloud resources of the cloud resource pool is to reduce or minimize the total cost of cloud resources of the cloud resource pool. Since each user terminal corresponds to a delay requirement level, the cloud resource allocation ratio among the multiple user terminals can be determined based on the delay requirement level requirement and the target proportion of the number of terminals, so as to reduce or minimize the total cost of cloud resources of the cloud resource pool.
[0064] 104. Perform cloud resource allocation to the multiple user terminals according to the cloud resource allocation ratio.
[0065] In the embodiments of the present application, the computing resources of the cloud servers are aggregated in the cloud resource pool, and the amount of the computing resources of the cloud servers is the total amount of cloud resources. The cloud resource allocation is performed to the multiple user terminals, so that the ratio between the amounts of cloud resources allocated to the multiple user terminals is equal to the cloud resource allocation ratio. For example, when the total amount of allocable cloud resources has been limited or the number of user terminals of the cloud resource pool is large, the embodiments of the present application can try to meet the delay requirements of the multiple user terminals matching the target proportion of the number of terminals, improve the utilization rate of cloud resources, and reduce the cost of cloud resources.
[0066] It can be seen that in the above embodiments of the present application, the cloud resource allocation ratio among the multiple user terminals is determined based on the delay requirement levels of the multiple user terminals of the cloud resource pool for computing tasks, and the target proportion of the number of terminals in the multiple user terminals that need to meet the corresponding delay requirement levels, so as to constrain the total cost of cloud resources of the cloud resource pool. Therefore, the cloud resources can be reasonably allocated by using the constraint on the total cost of cloud resources, so as to reduce the cost of cloud resources.
[0067] In some embodiments of the present application, as shown in Figure 2 based on the embodiments shown in Figure 1 The obtaining of the delay requirement levels of the multiple user terminals of the cloud resource pool for computing tasks can include:
[0068] 201. Obtain historical computing tasks of each user terminal in a historical period.
[0069] In the embodiments of the present application, the historical computing task refers to a computing task for data processing of the corresponding user terminal in a historical period. Each user terminal has performed data processing of at least one historical computing task in the historical period, that is, each user terminal has at least one historical computing task in the historical period.
[0070] 202. Confirm the task type of each historical computing task, wherein the task type includes at least one of real-time video rendering, interactive data analysis, and offline batch processing.
[0071] In the embodiments of the present application, the computing tasks are classified in advance, so that a plurality of task types are obtained, and thus the task type of each historical computing task can be determined. Since the computing tasks of real-time video rendering, interactive data analysis, and offline batch processing have different requirements for data processing delay, for example, real-time video rendering requires lower data processing delay, and offline batch processing does not require lower data processing delay, the task type includes at least one of real-time video rendering, interactive data analysis, and offline batch processing.
[0072] 203. Determine the delay requirement gear of each user terminal based on the task type.
[0073] In the embodiments of the present application, for each user terminal, the delay requirement gear of the user terminal is different when the task types of the historical computing tasks of the user terminal are different.
[0074] In some embodiments of the present application, determining the delay requirement gear of each user terminal based on the task type can include: for each user terminal, determining a target task type with the largest number of historical computing tasks among the task types of the plurality of historical computing tasks of the user terminal, for example, when the number of historical computing tasks with the task type of real-time video rendering is the largest among the plurality of historical computing tasks of the user terminal, the real-time video rendering is taken as the target task type; and determining the delay requirement gear of the user terminal based on the target task type, for example, when the target task type is real-time video rendering, the delay requirement gear of the user terminal is high, and for example, when the target task type is interactive data analysis, the delay requirement gear of the user terminal is medium, and for example, when the target task type is offline batch processing, the delay requirement gear of the user terminal is low. In this way, the delay requirement gear of the user terminal can be determined based on the target task type with the largest number of historical computing tasks, and thus the cloud resources allocated to the user terminal are more likely to be fully utilized, so as to improve the utilization rate of the cloud resources.
[0075] As can be seen, in the above-mentioned embodiments of the present application, the historical computing tasks of each user terminal in the historical period are obtained, the task type of each historical computing task is determined, and then the delay requirement gear of each user terminal is determined based on the task type, so that the determination of the delay requirement gear is more accurate.
[0076] In some embodiments of the present application, as shown in Figure 3 based on the above-mentioned embodiments shown in Figure 2 to determine the cloud resource allocation ratio among the plurality of user terminals based on the delay requirement gear and the target ratio of the number of terminals, which can include:
[0077] 301. Determine the target minimum delay pre-associated with each delay requirement gear.
[0078] In the embodiments of the present application, each latency requirement gear is pre-associated with a target minimum latency, and the target minimum latencies pre-associated with different latency requirement gears are different. The target minimum latency pre-associated with the high latency requirement gear is less than the target minimum latency pre-associated with the medium latency requirement gear. The target minimum latency pre-associated with the medium latency requirement gear is less than the target minimum latency pre-associated with the low latency requirement gear. The specific value of the target minimum latency can be set based on actual needs, which is not limited here.
[0079] 302. Determine the computing resource cost of the corresponding user terminal based on the target minimum latency and the target task type of the corresponding user terminal, wherein the total cost of cloud resources is equal to the sum of the computing resource costs of the multiple user terminals of the cloud resource pool.
[0080] In the embodiments of the present application, since the amount of computing resources required to achieve the target minimum latency when the user terminal performs data processing on the computing task is usually different when the target minimum latency is different, the corresponding computing resource cost is also usually different.
[0081] In some embodiments of the present application, the determination of the computing resource cost can be determined based on a neural network model. Specifically, determining the computing resource cost of the corresponding user terminal based on the target minimum latency and the target task type of the corresponding user terminal can include: obtaining a cost prediction model pre-associated with the target task type of the corresponding user terminal, wherein the cost prediction model is a neural network model, and different target task types are pre-associated with different cost prediction models; input the target minimum latency and the target task type of the corresponding user terminal into the cost prediction model; receive the cost value output by the cost prediction model as the computing resource cost of the corresponding user terminal. In this way, by distinguishing different target task types, the neural network model can be used to accurately predict the computing resource cost of the user terminal.
[0082] In some embodiments of the present application, an example description is given to the generation process of the cost prediction model. Specifically, the cost prediction model can be generated by: obtaining a plurality of historical computing tasks belonging to a target task type, a historical processing delay of each historical computing task, and a historical resource cost of each historical computing task, wherein the historical resource cost can be determined based on a total amount of historical cloud resources allocated to the corresponding historical computing task when the corresponding user terminal performs data processing on the corresponding historical computing task, for example, the historical resource cost can be positively correlated with the total amount of historical cloud resources, and the specific calculation rule of the historical resource cost can be set based on actual needs, which is not limited herein; based on the historical processing delay and the historical resource cost, model training is performed to obtain the cost prediction model. For example, the historical processing delay can be used as a sample for model training, and the corresponding historical resource cost can be used as a label of the sample, and the cost prediction model can be obtained through model training.
[0083] 303. Determine the product of the total number of the plurality of user terminals of the cloud resource pool and the terminal quantity target ratio as the target user terminal quantity.
[0084] In embodiments of the present application, the target user terminal quantity, i.e., the number of user terminals in the plurality of user terminals of the cloud resource pool that need to meet the corresponding delay requirement level.
[0085] 304. Based on the computing resource cost and the target user terminal quantity, determine the cloud resource allocation ratio with the constraint of the total cloud resource cost of the cloud resource pool as the target.
[0086] In embodiments of the present application, since each user terminal corresponds to a computing resource cost, the cloud resource allocation ratio among the plurality of user terminals can be determined based on the constraint of the computing resource cost and the target user terminal quantity, so as to reduce or minimize the total cloud resource cost of the cloud resource pool.
[0087] In some embodiments of the present application, based on the total cost of the cloud resources of the constrained cloud resource pool and the target number of user terminals, the cloud resource allocation ratio is determined, which can include: in the plurality of user terminals, the computing resource cost is used to determine a plurality of first user terminals matched with the target number of user terminals, for example, the plurality of user terminals can be ranked from low to high according to the computing resource cost, and the top ranked plurality of user terminals of the target number of user terminals are taken as the plurality of first user terminals; in the plurality of user terminals, the user terminals other than the first user terminals are all taken as second user terminals; the computing resource cost of each second user terminal is reduced, for example, the product of the computing resource cost of each second user terminal and a preset cost correction coefficient can be taken as the reduced computing resource cost of the corresponding second user terminal, the preset cost correction coefficient is greater than zero and less than 1; based on the computing resource cost of each first user terminal and the reduced computing resource cost of each second user terminal, the cloud resource allocation ratio is determined, for example, the ratio between the computing resource cost of each first user terminal and the reduced computing resource cost of each second user terminal can be directly taken as the cloud resource allocation ratio.
[0088] In some embodiments of the present application, since the user terminal is based on the end-to-cloud integration technology to process the computing task, the determination method of the cloud resource allocation ratio can also be optimized. Specifically, based on the computing resource cost of each first user terminal and the reduced computing resource cost of each second user terminal, the cloud resource allocation ratio is determined, which can include: obtaining a first local resource cost corresponding to the local resource amount of each first user terminal and a second local resource cost corresponding to the local resource amount of each second user terminal; for each first user terminal, the difference between the computing resource cost of the first user terminal and the corresponding first local resource cost is determined and taken as the first cloud resource cost; for each second user terminal, the difference between the reduced computing resource cost of the second user terminal and the corresponding second local resource cost is determined and taken as the second cloud resource cost; the ratio between the first cloud resource cost of each first user terminal and the second cloud resource cost of each second user terminal is taken as the cloud resource allocation ratio, so that the determined cloud resource allocation ratio is more accurate and the cloud resources are more reasonable.
[0089] Among the local resource amount of the first user terminal and the local resource amount of the second user terminal, the local resource amount refers to the amount of computing resources that can be used for data processing of a computing task at the corresponding user terminal in a local cloud fusion. The local resource amount can be obtained according to the local resource configuration information of the corresponding user terminal, wherein the local resource configuration information records the specific circumstances of the local resources of the user terminal, such as the local CPU resource amount and GPU resource amount of the user terminal. Therefore, the local resource amount may, for example, include at least one of the local CPU resource amount and the GPU resource amount of the user terminal.
[0090] It can be seen that in the above embodiments of the present application, the cloud resource allocation ratio is determined based on the computing resource cost of the corresponding user terminal and the target user terminal quantity, so as to constrain the total cloud resource cost of the cloud resource pool, so that the cloud resource can be reasonably allocated by using the constraint of the total cloud resource cost, thereby further reducing the cloud resource cost.
[0091] In a second aspect, based on the cloud resource scheduling optimization method based on the constraint condition in the above embodiments, the embodiments of the present application provide a cloud resource scheduling optimization device based on a constraint condition. The cloud resource scheduling optimization device based on the constraint condition is used to execute the steps in any embodiment of the above cloud resource scheduling optimization method based on the constraint condition. Specifically, referring to Figure 4 , the cloud resource scheduling optimization device 400 can include:
[0092] The acquisition module 401 is configured to acquire the delay requirement gear of each of the plurality of user terminals of the cloud resource pool for the computing task, and acquire the target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement gear.
[0093] The determination module 402 is configured to determine the cloud resource allocation ratio among the plurality of user terminals based on the delay requirement gear and the target proportion of the number of terminals, with the target of constraining the total cloud resource cost of the cloud resource pool.
[0094] The allocation module 403 is configured to allocate cloud resources to the plurality of user terminals according to the cloud resource allocation ratio.
[0095] In a third aspect, the embodiments of the present application provide a computer device integrating any cloud resource scheduling optimization device based on a constraint condition provided by the embodiments of the present application. The computer device includes a processor and a memory, and the memory stores a computer program configured to be executed by the processor to implement the cloud resource scheduling optimization method based on the constraint condition as described in any of the above embodiments, for example:
[0096] Obtain the latency requirement levels for computing tasks from multiple user terminals in the cloud resource pool; obtain the target proportion of the number of terminals that need to meet the corresponding latency requirement levels; with the total cloud resource cost of the cloud resource pool as the objective, determine the cloud resource allocation ratio among the multiple user terminals based on the latency requirement levels and the target proportion of the number of terminals; allocate cloud resources to the multiple user terminals according to the cloud resource allocation ratio.
[0097] Fourthly, embodiments of this application provide a computer device that integrates any of the constraint-based cloud resource scheduling optimization devices provided in embodiments of this application. For example... Figure 5 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:
[0098] The computer device may include components such as a processor 501 with one or more processing cores, a storage unit 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 5 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0099] The processor 501 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the storage unit 502, and by calling data stored in the storage unit 502, thereby providing overall monitoring of the computer device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 501.
[0100] The storage unit 502 can be used to store software programs and modules, and the processor 501 executes various function applications and data processing by running the software programs and modules stored in the storage unit 502. The storage unit 502 can mainly include a storage program area and a storage data area, wherein the storage program area can store operating systems, application programs required by at least one function (such as a sound playing function, an image playing function, etc.), and the like; and the storage data area can store data created according to the use of the computer device, etc. In addition, the storage unit 502 can include a high-speed random access memory, and can also include a nonvolatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the storage unit 502 can also include a memory controller to provide the processor 501 with access to the storage unit 502.
[0101] The computer device further includes a power supply 503 for supplying power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 503 can also include one or more than one direct current or alternating current power supply, a recharging system, a power failure detection circuit, a power converter or inverter, a power state indicator, and the like.
[0102] The computer device can further include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0103] Although not shown, the computer device can also include a display unit and the like, which will not be described here. Specifically in the embodiments of the present application, the processor 501 in the computer device will load the executable file corresponding to the process of one or more than one application program into the storage unit 502 according to the following instructions, and run the application program stored in the storage unit 502 by the processor 501, so as to realize various functions, for example:
[0104] Obtain delay requirement gears of a plurality of user terminals of a cloud resource pool for a computing task respectively; obtain a terminal quantity target proportion of the plurality of user terminals that need to meet corresponding delay requirement gears; determine a cloud resource allocation proportion between the plurality of user terminals based on the delay requirement gears and the terminal quantity target proportion, with a constraint of a total cost of cloud resources of the cloud resource pool as a target; and allocate cloud resources to the plurality of user terminals according to the cloud resource allocation proportion.
[0105] In a fifth aspect, embodiments of the present application provide a computer readable storage medium, which can include a Read Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc. The computer readable storage medium stores a computer program configured to be executed by a processor to implement the constraint-based cloud resource scheduling optimization method according to any one of the preceding aspects, for example:
[0106] obtaining a delay requirement level of each of a plurality of user terminals in a cloud resource pool for a computing task; obtaining a target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement level; determining a cloud resource allocation proportion between the plurality of user terminals based on the delay requirement level and the target proportion of the number of terminals, with a total cost of cloud resources of the cloud resource pool as a constraint; and allocating cloud resources to the plurality of user terminals according to the cloud resource allocation proportion.
[0107] In a sixth aspect, embodiments of the present application provide a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium and executes the computer instructions, so that the computer device implements the constraint-based cloud resource scheduling optimization method according to any one of the preceding aspects, for example:
[0108] obtaining a delay requirement level of each of a plurality of user terminals in a cloud resource pool for a computing task; obtaining a target proportion of the number of terminals in the plurality of user terminals that need to meet the corresponding delay requirement level; determining a cloud resource allocation proportion between the plurality of user terminals based on the delay requirement level and the target proportion of the number of terminals, with a total cost of cloud resources of the cloud resource pool as a constraint; and allocating cloud resources to the plurality of user terminals according to the cloud resource allocation proportion.
[0109] The above describes the embodiments of the present application in detail, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application; meanwhile, for those skilled in the art, the specific implementation manner and application range can be changed according to the idea of the present application, and the above description of the present application should not be understood as a limitation.
Claims
1. A method for cloud resource scheduling optimization based on constraints, characterized in that, The constraint-based cloud resource scheduling optimization method comprises: obtaining delay requirement levels of multiple user terminals of a cloud resource pool for a computing task respectively; obtaining a target proportion of terminal quantity of multiple user terminals that need to meet the delay requirement levels; determining a cloud resource allocation proportion among the multiple user terminals based on the delay requirement levels and the target proportion of terminal quantity, with a constraint of a total cloud resource cost of the cloud resource pool; allocating cloud resources to the multiple user terminals according to the cloud resource allocation proportion; wherein the obtaining of the delay requirement levels of the multiple user terminals of the cloud resource pool for the computing task respectively comprises: obtaining historical computing tasks of each user terminal in a historical period; confirming a task type of each historical computing task, wherein the task type comprises at least one of real-time video rendering, interactive data analysis, and offline batch processing; determining the delay requirement level of each user terminal based on the task type; the determining of the delay requirement level of each user terminal based on the task type comprises: for each user terminal, determining a target task type with the largest number of historical computing tasks among the task types of the multiple historical computing tasks of the user terminal; determining the delay requirement level of the user terminal based on the target task type; the determining of the cloud resource allocation proportion among the multiple user terminals based on the delay requirement levels and the target proportion of terminal quantity, with the constraint of the total cloud resource cost of the cloud resource pool comprises: determining a target minimum delay pre-associated with each delay requirement level; determining a computing resource cost of a corresponding user terminal based on the target minimum delay and the target task type of the corresponding user terminal, wherein the total cloud resource cost is equal to a sum of the computing resource costs of the multiple user terminals of the cloud resource pool; determining a product of a total quantity of the multiple user terminals of the cloud resource pool and the target proportion of terminal quantity as a target user terminal quantity; determining the cloud resource allocation proportion based on the computing resource cost and the target user terminal quantity, with the constraint of the total cloud resource cost of the cloud resource pool. 2.The constraint-based cloud resource scheduling optimization method of claim 1, wherein, the determining of the computing resource cost of the corresponding user terminal based on the target minimum delay and the target task type of the corresponding user terminal comprises: obtaining a cost prediction model pre-associated with the target task type of the corresponding user terminal; inputting the target minimum delay and the target task type of the corresponding user terminal into the cost prediction model; receiving a cost value output by the cost prediction model as the computing resource cost of the corresponding user terminal. 3.The constraint-based cloud resource scheduling optimization method of claim 2, wherein, the cost prediction model is generated by: obtaining multiple historical computing tasks belonging to the target task type, historical processing delay, and historical resource cost of each historical computing task; performing model training based on the historical processing delay and the historical resource cost to obtain the cost prediction model. 4.The constraint-based cloud resource scheduling optimization method of claim 1, wherein, The cloud resource total cost of the constraint of the cloud resource pool is taken as a target, the cloud resource allocation ratio is determined based on the computing resource cost and the target user terminal quantity, and the method comprises the steps of: In the plurality of user terminals, the computing resource cost is used to determine a plurality of first user terminals matched with the target user terminal quantity; In the plurality of user terminals, the user terminals except the first user terminals are all regarded as second user terminals; The computing resource cost of each second user terminal is reduced; The cloud resource allocation ratio is determined based on the computing resource cost of each first user terminal and the reduced computing resource cost of each second user terminal. 5.The constraint-based cloud resource scheduling optimization method of claim 1, wherein, The terminal quantity target ratio of the plurality of user terminals needing to meet the corresponding delay requirement gear is obtained, and the method comprises the steps of: In a plurality of preset time periods, a target time period where a current time point is located is determined; A preset ratio associated with the target time period is taken as the terminal quantity target ratio.
6. A computer device, comprising: The computer device comprises a processor and a memory, the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the constraint-based cloud resource scheduling optimization method in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is configured to be executed by the processor to implement the constraint-based cloud resource scheduling optimization method in any one of claims 1 to 5.
Citation Information
Patent Citations
Request scheduling method for multi-task edge video analysis
CN119723299A