Dynamic allocation method and device for computing power resources

By dynamically allocating computing resources in the natural language interaction platform, high-priority tasks are allocated to fast computing devices according to task priority and device status, and low-priority tasks are allocated to slow computing devices, the problem of unreasonable resource allocation is solved and user experience and task processing efficiency is improved.

CN120371545AActive Publication Date: 2025-07-25国家超级计算天津中心

Patent Information

Application Number
CN202510873462.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

When the existing natural language interaction platform handles different tasks, the allocation of computing resources is unreasonable, resulting in tasks with high computing resources being unable to be processed in time, response delays increase, user experience becomes worse, and resource waste is serious, and overall task processing efficiency is inefficient.

Method used

By parsing the text input by user, obtaining task priority scores, distinguishing high-priority and low-priority tasks, and assigning high-priority tasks to fast computing devices, and assigning low-priority tasks to slow computing devices. Use the lightweight user intention identification model to perform task priority evaluation, and dynamic resource allocation is performed in combination with the device call status.

Benefits of technology

It improves the response speed of high-priority tasks, reduces user waiting time, improves user satisfaction, makes full use of resources, avoids waste, and maintains the overall task processing capabilities of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371545A_ABST
    Figure CN120371545A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a computing power resource dynamic allocation method and device, and relates to the technical field of computing power allocation, and the method specifically comprises the steps: analyzing a text input by a user, so as to obtain a task priority score of a target task corresponding to the text input by the user; if the task priority score is greater than or equal to a first preset threshold value, determining that the target task is a high-priority task; obtaining a calling state of each first device; the calling state comprises an idle state and a use state; when a plurality of to-be-used first devices in an idle state exist, determining a target first device from the plurality of to-be-used first devices, and allocating the target task to the target first device; and when the plurality of first devices are all in the use state, distributing the target task to the second device in the idle state. The satisfaction degree and the experience feeling of the user in the interaction process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of computing power allocation, and particularly to a method and device for dynamically allocating computing power resources. Background Art

[0002] With the development of intelligent interaction technology, the user demands of natural language interaction platforms such as ChatGPT have started to show high frequency and diversification. Currently, natural language interaction platforms usually allocate all tasks to the same computing power resource for processing, and it is often difficult to take into account tasks with different computing power requirements. As a result, when the platform processes different tasks, tasks with high computing power requirements and tasks with low computing power requirements both occupy the same computing power resource. Therefore, for some tasks with high computing power requirements, due to competing for the same computing power resource with a large number of other tasks with low computing power requirements, the tasks with high computing power requirements cannot be processed in a timely manner, the response delay increases, and the user experience deteriorates.

[0003] At the same time, due to the unreasonable allocation of computing power resources, it is easy to cause resource waste. The above problems make the overall task processing efficiency of the entire platform low and difficult to meet the growing user demands. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a method and device for dynamically allocating computing power resources, which can improve the satisfaction and experience of users during the interaction process.

[0005] In a first aspect, the embodiments of the present application provide a method for dynamically allocating computing power resources, including: Parsing the user input text to obtain the task priority score of the target task corresponding to the user input text; If the task priority score is greater than or equal to the first preset threshold, determining that the target task is a high-priority task; the high-priority task is a high-priority task that needs to be processed in real time; Obtaining the call status of each first device; the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state; When there are multiple first devices in the idle state, determining a target first device from the multiple first devices and allocating the target task to the target first device; When all of the multiple first devices are in the used state, then allocating the target task to a second device in the idle state; the second device is a slow computing device for processing low-priority tasks.

[0006] As an optional implementation manner of the embodiments of the present application, when there are multiple first devices in the idle state, determining a target first device from the multiple first devices includes: Obtain the computing power resources corresponding to multiple first devices to be used respectively, and the computing power requirements corresponding to the target task, so as to calculate the predicted response times corresponding to multiple first devices to be used respectively; Determine the target first device based on the size relationship between the predicted response time and the preset response time threshold.

[0007] As an optional implementation manner of an embodiment of the present application, after parsing the user input text to obtain the task priority score of the target task corresponding to the user input text, the method further includes: If the task priority score is less than or equal to the second preset threshold, determine that the target task is a low-priority task; the second preset threshold is less than the first preset threshold; Obtain the call status of each second device; Determine the second devices to be used that are in an idle state from multiple second devices, determine the target second device from multiple second devices to be used, and assign the target task to the target second device.

[0008] As an optional implementation manner of an embodiment of the present application, determining the second devices to be used that are in an idle state from multiple second devices, and determining the target second device from multiple second devices to be used includes: Obtain the computing power resources corresponding to multiple second devices to be used respectively, and the computing power requirements corresponding to the target task, so as to determine the maximum task throughput efficiency; Determine the target second device based on the maximum task throughput efficiency.

[0009] As an optional implementation manner of an embodiment of the present application, parsing the user input text to obtain the task priority of the target task corresponding to the user input text includes: Based on the lightweight user intention recognition model, parse the urgency and real-time requirements of the target task corresponding to the user input text to obtain the corresponding task priority score.

[0010] As an optional implementation manner of an embodiment of the present application, the method further includes: Obtain the status information of the target task, and predict the expected task delay tolerance in combination with the user's historical behavior; Monitor the current response duration corresponding to the target task; Adjust the computing power resources of the target first device according to the size relationship between the current response duration and the expected task delay tolerance.

[0011] As an optional implementation manner of an embodiment of the present application, the method further includes: After the target task is executed, collect the complete response time and computing power utilization rate corresponding to the target task; Update the first preset threshold and the second preset threshold according to the complete response time and the computing power utilization rate.

[0012] In a second aspect, an embodiment of the present application provides a dynamic allocation device for computing power resources, including: A parsing unit, configured to parse the user input text to obtain the task priority score of the target task corresponding to the user input text; A determining unit, configured to determine that the target task is a high-priority task if the task priority score is greater than or equal to the first preset threshold; the high-priority task is a high-priority task that needs to be processed in real time; An obtaining unit, configured to obtain the call status of each first device; the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state; A first allocation unit, when there are multiple first devices in the idle state, determines a target first device from the multiple first devices and allocates the target task to the target first device; A second allocation unit, configured to, when all of the multiple first devices are in the used state, allocate the target task to a second device in the idle state; the second device is a slow computing device for processing low-priority tasks.

[0013] As an optional implementation manner of the embodiment of the present application, the first allocation unit is specifically configured to obtain the computing power resources corresponding to the multiple first devices to be used respectively, and the computing power requirements corresponding to the target task, so as to calculate the predicted response times corresponding to the multiple first devices to be used respectively; determine the target first device based on the size relationship between the predicted response time and the preset response time threshold.

[0014] As an optional implementation manner of the embodiment of the present application, the parsing unit is further configured to determine that the target task is a low-priority task if the task priority score is less than or equal to the second preset threshold; the low-priority task is a low-priority task that needs to be processed in batches; obtain the call status of each second device; determine a second device to be used in the idle state from the multiple second devices, and determine a target second device from the multiple second devices to be used, and allocate the target task to the target second device.

[0015] As an optional implementation manner of the embodiment of the present application, the second allocation unit is specifically configured to obtain the computing power resources corresponding to the multiple second devices to be used respectively, and the computing power requirements corresponding to the target task, so as to determine the maximum task throughput efficiency; determine the target second device based on the maximum task throughput efficiency.

[0016] As an optional implementation manner of an embodiment of the present application, the parsing unit is specifically configured to parse the urgency and real-time requirements of the target task corresponding to the user input text based on a lightweight user intention recognition model, and obtain the corresponding task priority score.

[0017] As an optional implementation manner of an embodiment of the present application, the obtaining unit is further configured to obtain the status information of the target task, and combine the user's historical behavior to predict the expected task delay tolerance; monitor the current response duration corresponding to the target task; and adjust the computing power resources of the target first device according to the magnitude relationship between the current response duration and the expected task delay tolerance.

[0018] As an optional implementation manner of an embodiment of the present application, the computing power resource dynamic allocation device further includes an updating unit, which is specifically configured to collect the complete response time and computing power utilization rate corresponding to the target task after the execution of the target task; and update the first preset threshold and the second preset threshold according to the complete response time and the computing power utilization rate.

[0019] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program; and the processor is used to cause the electronic device to implement the computing power resource dynamic allocation method of any one of the above embodiments when executing the computer program.

[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is caused to implement the computing power resource dynamic allocation method of any one of the above embodiments.

[0021] The specific method for dynamically allocating computing power resources provided by the embodiments of this application is as follows: Parse the user input text to obtain the task priority score of the target task corresponding to the user input text; if the task priority score is greater than or equal to the first preset threshold, determine that the target task is a high-priority task; a high-priority task is a high-priority task that needs to be processed in real time; obtain the call status of each first device; the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state; when there are multiple first devices in the idle state, determine the target first device from the multiple first devices and allocate the target task to the target first device; when all the first devices are in the used state, allocate the target task to the second device in the idle state; the second device is a slow computing device for processing low-priority tasks. The above method can accurately distinguish high-priority tasks and low-priority tasks by parsing the user input text to obtain the task priority score. Furthermore, high-priority tasks can be promptly allocated to appropriate computing devices, that is, high-speed computing devices, avoiding resource contention with low-priority tasks, ensuring that tasks with high computing power resource requirements can be processed first, and effectively reducing the response latency of high-priority tasks; furthermore, in a natural language interaction platform, urgent questions or high-complexity requirements raised by users can be promptly responded to and processed, reducing the user's waiting time and enhancing the user's satisfaction and experience during the interaction process.

[0022] At the same time, when all fast computing devices are in the used state, allocate high-priority tasks to the idle slow computing device, that is, the second device; thereby making full use of idle computing power resources, avoiding resource waste, and also ensuring that tasks can still be processed when resources are scarce, maintaining the overall task processing ability of the platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required to be referred to in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is one of the step flowcharts of the method for dynamically allocating computing power resources provided by the embodiments of this application; Figure 2 is the second of the step flowcharts of the method for dynamically allocating computing power resources provided by the embodiments of this application; Figure 3It is the third step flowchart of the computing power resource dynamic allocation method provided by the embodiment of the present application; Figure 4 It is the structural schematic diagram of the computing power resource dynamic allocation device provided by the embodiment of the present application; Figure 5 It is the hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0026] In order to more clearly understand the above-mentioned objects, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0027] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all the embodiments.

[0028] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In addition, in the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more.

[0029] It should be noted that in this article, the term "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0030] The embodiment of the present application provides a computing power resource dynamic allocation method. Referring to Figure 1 as shown, the computing power resource dynamic allocation method includes the following S101-S105: S101. Parse the user input text to obtain the task priority score of the target task corresponding to the user input text.

[0031] In the intelligent interaction platform, the types of tasks input by users are diverse, and their computing requirements and urgency levels vary. Embodiments of this application parse the text input by users to further explore the key information behind the tasks corresponding to the text input by users, thereby determining the priority scores of the tasks, so as to provide a basis for reasonably allocating computing resources subsequently.

[0032] Specifically, the specific implementation manner of parsing the text input by users to obtain the task priority scores of the target tasks corresponding to the text input by users can be: based on a lightweight user intention recognition model, infer and analyze the urgency and real-time requirements of the target tasks corresponding to the text input by users, and obtain the corresponding task priority scores.

[0033] In some embodiments, the lightweight user intention recognition model can be constructed using a miniaturized language model with model parameters between 1.5B and 3B, where B is an order of magnitude unit representing one billion (Billion). Through technical means such as parameter pruning and model quantization, the model scale and computational complexity are significantly reduced, the computing power consumption during inference is reduced, enabling it to complete inference within milliseconds and meet the real-time processing requirements.

[0034] Furthermore, the lightweight user intention recognition model is used to parse the text input by users to quickly and accurately evaluate the urgency and real-time requirements of the tasks corresponding to the text input by users. Then, based on this, task priority scoring is performed for them. The priority level of the task corresponding to the current text input by the user is obtained through the task priority scores, which is convenient for allocating the corresponding computing power resources for the task subsequently.

[0035] Specifically, after receiving the text input by users, the lightweight user intention recognition model will first parse it and convert it into corresponding tasks. Exemplarily, the text input by user A is parsed and converted into a target task, and then an embedding operation is performed on the target task: ; where represents the embedding operation, which is used to map the text-form task to a fixed-dimensional vector space, convert the text into a numerical vector representation that can be processed by a computer, and extract the key semantic features of the text for subsequent analysis and calculation.

[0036] Then, with reference to the following formula, inference is performed based on the embedding features using the lightweight user intention recognition model and the model parameter weights, and then the task priority scores are calculated: .

[0037] Where is the lightweight user intention recognition model, is the model parameter weight; P is the task priority score, and the value range of P is [0, 1]. The numerical value of P is the main basis for measuring the task priority.

[0038] Through this step, the urgency and real-time requirements of the target task corresponding to the user input text can be analyzed to obtain the task priority score, providing a basis for subsequent task scheduling and resource allocation.

[0039] S102. If the task priority score is greater than or equal to the first preset threshold, determine that the target task is a high-priority task.

[0040] In the embodiments of the present application, the task priority level can be determined by the task priority score corresponding to each user input text; specifically, for the priority degree of the task, the tasks are divided into two priority levels, including high-priority tasks and low-priority tasks. Among them, high-priority tasks refer to tasks that require immediate response. Such tasks usually need to be processed in real time and will then be allocated to fast computing devices in the subsequent process to ensure that a response can be obtained within milliseconds; low-priority tasks refer to tasks that allow delayed processing, and are often tasks that can be batch-processed. Such tasks have low requirements for real-time performance and will be allocated to slow computing devices for processing.

[0041] It should be noted that in the embodiments of the present application, after determining the priority level corresponding to each task, the corresponding priority label can be marked for each task in the way of tagging. For example, high priority is represented by and low priority is represented by . Then, the tasks with the corresponding identifiers are allocated to the resource pool of the corresponding fast computing device or the resource pool of the slow computing device, which is convenient for allocating the corresponding computing device for the task subsequently.

[0042] Combined with the above S101, if P > (such as set to 0.7), it indicates that the urgency and real-time requirements of the task are high, and it is thus determined as a high-priority task. If P < (such as set to 0.3), it means that the urgency and real-time requirements of the task are low, belonging to low-priority tasks. It should be noted that and can be set by developers according to actual needs, and the embodiments of the present application do not make any limitations.

[0043] It should also be noted that when the priority score of a task is in the middle range, that is, when it does not meet the determination ranges of high priority and low priority, the system will further evaluate the real-time requirements of the task. By analyzing factors such as the nature of the task, its association with other tasks, and the degree of impact on the business process, it is determined whether it requires a faster response or can tolerate a certain delay. If it is a task involving the connection of critical business processes and has a direct impact on other tasks, its priority may be increased; if the task is not time-sensitive and does not affect the overall business progress, the priority can be appropriately reduced.

[0044] It is also possible to dynamically adjust the task allocation strategy in combination with the current resource utilization situation of the system. If the resources of the fast computing device (the first device) are relatively abundant, even if the task does not clearly meet the high-priority standard, it can be allocated to the fast computing device for processing to speed up the task execution; if the load of the slow computing device (the second device) is low and the task has no strict real-time requirements, even if the task priority is slightly higher than the low-priority threshold, it can be allocated to the slow computing device to make full use of the batch processing ability of the slow computing device and improve the overall resource utilization.

[0045] S103. Obtain the call status of each first device.

[0046] Among them, the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state.

[0047] Specifically, after determining that the target task corresponding to the user input text is a high-priority task, that is, a high-priority level task, it is necessary to immediately allocate the corresponding computing device from the computing resource pool corresponding to the fast computing device to process the target task in a timely manner; before determining the target first device, it is necessary to obtain the call status of multiple fast computing devices (the first device) to accurately grasp the resource situation of the fast computing device and ensure that high-priority tasks can be timely allocated to available devices for processing, avoiding task waiting and backlog, and ensuring the efficiency and timeliness of system response.

[0048] S104. When there are multiple first devices in the idle state, determine the target first device from the multiple first devices and allocate the target task to the target first device.

[0049] In some embodiments, determining the target first device from multiple first devices may be a comprehensive judgment based on the computing power of the device and the specific requirements of the current task to determine the target first device, which is the best choice for processing the current target task; then allocate the target task to the target first device, and the target first device completes the target task. Through clarifying the device status in the embodiments of the present application, the target first device suitable for processing the target task is reasonably determined.

[0050] It should be noted that in many scenarios, the target task can be divided into multiple subtasks, and these subtasks can be respectively assigned to the corresponding target first devices in the idle state to complete. Therefore, the number of target first devices in the embodiments of the present application is not limited.

[0051] Assume the task has a computing requirement of (the computing power required for a unit task), then the scheduling matching relationship for high-priority tasks is as follows: .

[0052] Among them, represents the current usage of the fast computing device i (which can take values 0 or 1, 1 means available, that is, in the idle state, 0 means unavailable, in the usage state); represents the computing power of the fast computing device i; is the number of slow computing devices; represents that the task belongs to a high-priority task.

[0053] S105. When all of the multiple first devices are in the usage state, the target task is assigned to the second device in the idle state.

[0054] Among them, the second device is a slow computing device, which is used to process low-priority tasks.

[0055] In some embodiments, after obtaining the call status of each first device, it may be encountered that all first devices are in the usage state, that is, there is no idle fast computing device available to process the current high-priority target task; at this time, the slow computing device can be flexibly allocated to ensure the progress of the task, and then find the second device in the idle state to replace the first device to process the target task; although the processing speed of the second device is relatively slow, it can also be used to process high-priority tasks to avoid the task from waiting or backlogging for a long time due to lack of available resources and ensure the continuous progress of the task.

[0056] Similarly, before calling the second device, it is also necessary to check whether it is in the idle state. If there is an idle second device, the task scheduling module will assign the target task to it. During the processing, the second device will gradually complete the task processing according to its own processing ability and task scheduling rules. At the same time, the system will continuously monitor the status of the first device. Once a first device is idle, some high-priority tasks processed on the second device may be transferred back to the first device to speed up the processing speed.

[0057] It should be noted that when multiple first devices are all in use, when the target task is assigned to a second device in an idle state, if there are a relatively large number of second devices in an idle state, the second device that can meet the current target task can also be preferentially selected to process the target task. Specifically, the computing power margin of each second device in an idle state can be compared with the computing power required for the current target task to determine the second device used to process the target task. The above process can be expressed by the following formula: .

[0058] Among them, represents the computing power margin of a certain second device, represents the computing power required for the current target task, represents the current target task, that is represents that the current target task is a high-priority task.

[0059] Furthermore, by using the above formula, it is judged whether the computing power margin of a certain second device can meet the computing requirements of the high-priority task . If it is satisfied , the high-priority task will be temporarily assigned to the slow computing device for execution.

[0060] It should also be noted that in the actual computing power resource allocation process, it may also be encountered that after the target task is assigned to the target first device, due to other reasons, during the process of the target first device processing the target task, its remaining computing power cannot support the completion of the target task. At this time, part of the target task can also be transferred and assigned to an idle second device, and the second device will share the current target task to ensure the smooth completion of the current target task.

[0061] The dynamic allocation method of computing power resources provided by the embodiments of this application is specifically as follows: Parse the user input text to obtain the task priority score of the target task corresponding to the user input text; if the task priority score is greater than or equal to the first preset threshold, determine that the target task is a high-priority task; the high-priority task is a high-priority task that needs to be processed in real time; obtain the call status of each first device; the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state; when there are multiple first devices in the idle state, determine the target first device from the multiple first devices and allocate the target task to the target first device; when all the first devices are in the used state, allocate the target task to the second device in the idle state; the second device is a slow computing device for processing low-priority tasks. The above method can accurately distinguish high-priority tasks and low-priority tasks by parsing the user input text to obtain the task priority score. Furthermore, high-priority tasks can be timely allocated to appropriate computing devices, that is, high-speed computing devices, to avoid competing for resources with low-priority tasks, ensure that tasks with high computing power resource requirements can be processed first, and effectively reduce the response delay of high-priority tasks; furthermore, in a natural language interaction platform, urgent questions or high-complexity requirements raised by users can be timely responded to and processed, reducing the user waiting time and enhancing the user's satisfaction and experience during the interaction process.

[0062] At the same time, when all fast computing devices are in the used state, allocate high-priority tasks to the idle slow computing device, that is, the second device; thereby making full use of idle computing power resources, avoiding resource waste, and also ensuring that tasks can still be processed when resources are scarce, maintaining the overall task processing ability of the platform.

[0063] As an extension and refinement of the above embodiment, refer to Figure 2 As shown, the embodiments of this application also provide a dynamic allocation method of computing power resources, including the following S201 to S204: S201. Parse the user input text to obtain the task priority score of the target task corresponding to the user input text.

[0064] The description of this step refers to the description of S101 above and will not be elaborated here.

[0065] S202. If the task priority score is less than or equal to the second preset threshold, determine that the target task is a low-priority task.

[0066] Among them, the second preset threshold is less than the first preset threshold.

[0067] Specifically, when the task priority score is less than or equal to the second preset threshold, it is determined that the target task is a low-priority task, that is, a low-priority task. The characteristic of this type of task is that a certain delay is allowed and it is suitable for batch processing.

[0068] S203. Obtain the call status of each second device.

[0069] Specifically, before invoking the second device to process the target task, by obtaining the call status of the second device, it is determined whether each device is currently processing a task. Then, based on the obtained status information of the second device, the devices in the idle state are screened out from multiple second devices, providing accurate information for subsequent task allocation, ensuring that the task can be allocated to a suitable device, and avoiding low task processing efficiency or waste of device resources caused by inappropriate device selection.

[0070] S204. Determine the second devices to be used that are in the idle state from multiple second devices, determine the target second device from multiple second devices to be used, and allocate the target task to the target second device.

[0071] In the embodiment of the present application, for the screened idle second devices, the system will further evaluate their performance metrics. According to the characteristics and requirements of the target task, multiple devices that are most suitable for processing the target task are selected from the idle second devices as the target second devices. Then, the target task is sent to the corresponding target second device to process the target task.

[0072] Assume that the computing requirement of task Tk is (the required computing power of a unit task), then the scheduling relationship of the low-priority task is as follows: .

[0073] Among them, represents the usage status of the slow computing device (the value can be 0 or 1, 1 means available, that is, in the idle state, 0 means unavailable, in the usage state); represents the computing power of the slow computing device ; represents the total computing requirement of the low-priority task, usually a batch processing task; is the number of slow computing devices; represents the task belongs to the low-priority task.

[0074] In the embodiment of the present application, the task priority score is obtained by parsing the user input text and compared with the second preset threshold to accurately determine the low-priority tasks. Then, the call status of the second device, i.e., the slow computing device, is obtained, and the idle second device to be used is determined from multiple second devices, and the target device is further selected to allocate tasks; thus, the resources of the slow computing device can be utilized to enable low-priority tasks to be processed on appropriate devices, which helps to balance the computing resource load within the system. Avoid competing with high-priority tasks for computing power resources. High-priority tasks are processed by fast computing devices, and low-priority tasks are processed by slow computing devices, ensuring the stable operation of the system and improving the overall task processing ability.

[0075] Specifically, the specific implementation method of determining the second device in the idle state from multiple second devices and determining multiple target second devices in step S204 above can refer to the following steps 1 and 2: Step 1: Obtain the computing power resources corresponding to multiple second devices to be used respectively, and the computing power requirements corresponding to the target task to determine the maximum task throughput efficiency.

[0076] In the embodiment of the present application, the slow computing device often works in a batch processing manner to process low-priority tasks. Since the throughput efficiency refers to the total amount of tasks processed within a certain period of time, it can effectively measure the task processing ability of the slow computing device. The throughput efficiency can be calculated by the following formula:

[0077] Wherein, represents the throughput efficiency, represents the total computing requirements of the low-priority tasks, represents the total computing ability of the slow computing device, is the number of slow computing devices, represents the slow computing device 's computing ability.

[0078] Specifically, the current multiple idle second devices to be used are combined through the above formula, the throughput efficiency under multiple combinations is calculated, and it is determined which device combinations have high throughput efficiency and are thus more suitable for processing the current target task. For example, when multiple slow computing devices are available for selection, compare the throughput efficiency under different device combinations, and preferentially select the combination with high efficiency as the target second device to allocate tasks to improve the overall processing efficiency and maximize the resource utilization rate.

[0079] Step 2: Determine the target second device based on the maximum task throughput efficiency.

[0080] Furthermore, based on the throughput efficiency calculated in S2041 above, multiple target second devices are further determined, which facilitates the subsequent allocation of target tasks to these multiple second devices.

[0081] Similarly, in many scenarios, a target task can be divided into multiple subtasks, and these subtasks can be respectively assigned to corresponding target first devices in an idle state for completion. Therefore, the number of target first devices is not limited in the embodiments of the present application.

[0082] In the embodiments of the present application, the slow computing device can also be further optimized in combination with its actual usage. Specifically, the throughput of the slow computing device can be obtained to further measure whether resource allocation for low-priority tasks needs to be optimized currently. The calculation method of throughput specifically refers to the following formula: .

[0083] Wherein, represents throughput; is the total computing amount of low-priority tasks; is the total task execution time of all slow computing devices. During the optimization process, the throughput is gradually iteratively obtained to maximize the value of the throughput, so as to guide the slow computing device to process low-priority tasks, allocate resources and schedule tasks more efficiently, so that more low-priority tasks can be processed per unit time and the overall processing efficiency can be improved.

[0084] As an extension and refinement of the above embodiments, in combination with the embodiments shown in S101 to S105 and S201 to S204 above, the embodiments of the present application further include the following steps A and B: Step A: After the target task is executed, collect the complete response time and computing power utilization rate corresponding to the target task.

[0085] In the embodiments of the present application, after the target task is executed, by collecting the complete response time and computing power utilization rate corresponding to the target task, key feedback is provided for the system to optimize resource allocation and scheduling strategies. If it is found that the complete response time of a certain type of task always exceeds expectations while the computing power utilization rate is not high, the system can adjust the task allocation rules accordingly, such as trying to allocate this type of task to different devices or adjusting the device combination to improve the processing efficiency. Similarly, if the computing power utilization rate is at an unreasonable level for a long time, whether it is too high or too low, the scheduling strategy can be adjusted accordingly. For example, when the computing power utilization rate is too high, increase device resources or optimize the task queuing mechanism to avoid device overload; when the computing power utilization rate is too low, re-evaluate the task allocation method to make the device make more full use of the computing power.

[0086] Specifically, the complete response time corresponding to the target task is the entire time span from the generation of the target task to the completion of task processing and the feedback of results. It covers the time consumed in various links such as task queuing, scheduling and allocation, and device processing. The computing power utilization rate is the actual utilization ratio of the computing power resources used during the task execution. By monitoring the working status of computing cores such as the CPU and GPU of the device, and the occupancy of resources such as memory and storage, the average utilization percentage of computing power resources during the task execution is calculated.

[0087] Step B: Update the first preset threshold and the second preset threshold according to the complete response time and the computing power utilization rate.

[0088] If the collected complete response times are generally short, it indicates that the current system has a high task processing efficiency, and the first preset threshold can be appropriately reduced to further improve the real-time requirement of task processing; if the complete response time is long and exceeds the expectation, it may be necessary to increase the first preset threshold to avoid unreasonable task allocation due to excessive requirements. For example, if the first preset threshold for high-priority tasks was originally 10 milliseconds (ms), but it is found that the average response time is only 5 milliseconds (ms) during actual processing, the threshold can be considered to be lowered to 8 milliseconds (ms).

[0089] When the computing power utilization rate remains at a low level for a long time, it indicates that the resources are not fully utilized, and the second threshold can be adjusted to allow the device to run at a higher load, so as to allocate more tasks and improve the resource utilization rate; if the computing power utilization rate is often too high and the device is on the verge of overload, it may be necessary to reduce the second threshold to reduce the task allocation volume and ensure the stable operation of the device. For example, if the computing power utilization rate of the device has been maintained at 40%, which is far lower than the device's carrying capacity, the second threshold can be appropriately increased to enable the device to undertake more tasks.

[0090] Furthermore, by updating the first preset threshold and the second threshold with the actual data (complete response time and computing power utilization rate) after each task execution, the task scheduling and resource allocation strategies can be made more in line with the actual operating conditions, achieving dynamic optimization.

[0091] As an extension and refinement of the above embodiments, referring to Figure 3 As shown, the embodiment of the present application further provides a method for dynamically allocating computing power resources, including the following S301 to S306: S301: Parse the user input text to obtain the task priority score of the target task corresponding to the user input text.

[0092] For the description of this step, refer to the description of S101 above, and details will not be elaborated here.

[0093] S302: If the task priority score is greater than or equal to the first preset threshold, determine that the target task is a high-priority task.

[0094] For the description of this step, please refer to the description of S102 above, which will not be elaborated here.

[0095] S303. Obtain the call status of each first device.

[0096] Among them, the first device is a fast computing device for processing high-priority tasks; the call status includes an idle state and a used state.

[0097] For the description of this step, please refer to the description of S103 above, which will not be elaborated here.

[0098] S304. When at least one of the multiple first devices is in an idle state, obtain the computing power resources corresponding to each of the multiple first devices to be used and the computing power requirements corresponding to the target task, so as to calculate the predicted response time corresponding to each of the multiple first devices to be used.

[0099] In this step, when at least one of the multiple first devices is in an idle state, the specific implementation method for determining the target first device from the multiple first devices can be to first obtain the computing power resources corresponding to each of the multiple first devices to be used and the computing power requirements corresponding to the target task, so as to calculate the predicted response time corresponding to each of the multiple first devices to be used.

[0100] Specifically, it is necessary to obtain the computing power requirements of the target task, which include the computing complexity and data volume of the task, etc. Then, according to the computing power resources of each first device and the computing power requirements of the current target task, calculate the predicted response time for each idle first device to process the target task. The predicted response time refers to the estimated time required for the device to start processing the task until it completes the task. The calculation formula for the predicted response time can be referred to as follows: .

[0101] Among them, represents the predicted response time, W is the computing power requirement of the task; represents the computing power capacity of the fast computing device; is the maximum response time set for high-priority tasks, that is, the preset response time threshold, such as 10 milliseconds (ms). Then, through the above formula, calculate the predicted response time of each first device in the idle state. It is convenient to compare with the preset response time threshold subsequently. This threshold is determined according to the real-time requirements of the task and the overall performance goals of the system.

[0102] S305. Based on the size relationship between the predicted response time and the preset response time threshold, determine the target first device and allocate the target task to the target first device.

[0103] Compare the predicted response times corresponding to multiple first devices calculated in S304 with a preset response time threshold, which is a time limit set for high-priority tasks and represents the maximum acceptable time range for tasks to be processed on the devices.

[0104] Based on the calculated predicted response times, screen out the first devices whose predicted response times meet the preset response time threshold as the target first devices. If the predicted response times of multiple devices all meet the requirements, the target first devices can be further determined according to other factors, such as the remaining resources of the devices and the historical processing performance of the devices. After determining the target first devices, the target tasks can be assigned to these devices for processing. Through the above steps, determining the target first devices according to the response time can effectively ensure the timely processing of high-priority tasks and improve the overall performance and response speed of the system.

[0105] In the embodiments of the present application, the actual usage of the fast computing devices can also be combined to further optimize the fast computing devices, that is, optimize according to the average response delay size feedback by the fast computing devices. The specific calculation formula for the average response delay size is as follows:

[0106] where L represents the average response delay; represents the sum of the complete response times of each high-priority task, that is, the total difference between the completion time and the start time; is the number of high-priority tasks. Furthermore, by obtaining the above ratio L and continuously minimizing it, the system can optimize the scheduling and execution of high-priority tasks on the fast computing devices, enabling tasks to be completed faster.

[0107] S306. When all of the multiple first devices are in use, assign the target tasks to the second devices in the idle state.

[0108] Among them, the second devices are slow computing devices for processing low-priority tasks.

[0109] The embodiments of the present application further illustrate how to determine the target first devices, that is, by calculating the predicted response times, preferentially select the devices that can process tasks quickly, avoid assigning tasks to devices that are idle but have low processing efficiency, make the resources of the fast computing devices be more reasonably utilized, prevent resource idleness or inefficient use, and improve the overall resource utilization efficiency. At the same time, selecting devices and assigning tasks according to the predicted response times can avoid the situation where some devices are overloaded while some devices are idle, balance the load among the fast computing devices, extend the service life of the devices, and ensure the long-term stable operation of the system.

[0110] As an extension and refinement of the above embodiments, after each task execution, in order to improve the user interaction experience, the dynamic allocation method of computing power resources will be continuously optimized through user task data and the results of user interaction feedback.

[0111] Specifically, it is necessary to collect the response time, task latency tolerance, and interaction satisfaction for each task execution within a period of time; among them, the interaction satisfaction is obtained by user experience scoring and is used to evaluate whether the task results meet the user's needs. Then, high-priority tasks and low-priority tasks are divided, and the corresponding average interaction satisfaction is calculated for high-priority tasks and low-priority tasks respectively; to adjust the strategy when allocating fast computing devices and slow computing devices according to the feedback average interaction satisfaction, the calculation formula of the average interaction satisfaction is as follows:

[0112] Among them, is the interaction satisfaction corresponding to each task; is the actual response time of the i-th task; N is the total number of tasks; is the latency tolerance of the i-th task; S represents the average interaction satisfaction; , this condition is for all tasks i, and it is required that the actual response time of the i-th task is less than or equal to its latency tolerance .

[0113] Furthermore, through the average interaction satisfaction, it is further decided whether to optimize the way of allocating computing power resources for high-priority tasks; then for the optimization method of fast computing devices, that is, when the average interaction satisfaction of high-priority tasks is lower than the preset threshold, the computing power resources allocated to high-priority tasks by fast computing devices can be increased to further meet the real-time requirements of high-priority tasks; specifically, the scheduling parameters used when allocating computing power resources for high-priority tasks can be optimized : .

[0114] Among them, is the new scheduling parameter obtained through optimization; is the learning factor (determining the amplitude of each update of the scheduling parameter), is the gradient update of the current scheduling strategy for the interaction satisfaction, reflecting the change direction and degree of the influence of the current scheduling method on the interaction satisfaction; Gradient update of the resource allocation method for the current scheduling on the interactive satisfaction, making the resource allocation of the fast computing device more in line with the needs of high-priority tasks; through the above update formula, the resource allocation of the fast computing device is more in line with the needs of high-priority tasks. At the same time, according to the average interactive satisfaction feedback from task execution, continuously iteratively update the scheduling parameters, enabling the system to dynamically adjust the strategy during resource allocation to adapt to different task loads and user demand changes. It should be noted that the scheduling parameters are multiple parameters related to task resource allocation, which are not limited here and can be set according to the actual situation.

[0115] It is also possible to optimize the computing power resource ratio adopted when allocating computing power resources for high-priority tasks : .

[0116] Among them, represents the computing power resource ratio of high-priority tasks; is the learning rate ( determining the amplitude of each update of the computing power resource ratio); and then through the average interactive satisfaction feedback in real time during the execution of high-priority tasks, continuously iteratively update to obtain the latest , to further dynamically adapt to the changes in user needs.

[0117] Similarly, through the average interactive satisfaction, further decide whether to optimize the way of allocating computing power resources for low-priority tasks; then for the optimization method of slow computing devices, that is, when the average interactive satisfaction of low-priority tasks is lower than the preset threshold, the optimization method can be to arrange task batches reasonably, enabling the slow computing device to perform batch processing more efficiently and increasing the task processing volume per unit time, that is, improving the throughput. Similarly, based on the feedback data after the execution of low-priority tasks, such as whether the response time is within the acceptable range and the interactive satisfaction situation, gradually adjust the task allocation strategy of the slow computing device. For example, if it is found that the interactive satisfaction of a certain type of low-priority task is low when processed on a slow computing device, the task allocation can be adjusted to other more suitable slow computing devices, or the device configuration can be optimized to improve the overall interactive satisfaction.

[0118] Furthermore, obtain the average interactive satisfaction for all tasks over a period of time through the above formula. By adjusting the task allocation strategy and then observing the average interactive satisfaction feedback in real time in the subsequent period of time, determine whether the previous adjustment is effective, and gradually iteratively adjust the above method to make the average interactive satisfaction reach the optimal.

[0119] In the embodiments of the present application, the above-mentioned optimization of the computing power resource scheduling methods for high-priority tasks and low-priority tasks respectively through the average interaction satisfaction can also be optimized overall by minimizing the objective function shown in the following formula , to balance the overall global computing and interaction performance, enabling the system to have the ability to optimize dynamic load. Considering both user interaction satisfaction and computing resource utilization efficiency, while meeting the user experience, computing power resources are efficiently utilized; specifically, the objective function is expressed as:

[0120] wherein, is the average response time of the task, is the average delay tolerance of the task. In the formula, , this part measures the degree of difference between the actual average response time and the average delay tolerance. The smaller the difference, the more the task response conforms to the acceptable range of the user, and the higher the interaction satisfaction. is the weight coefficient, used to adjust the importance of this part in the objective function, the larger it is, the more the system pays attention to interaction satisfaction.

[0121] M represents the total number of computing power devices, is the total computing capacity of the computing power devices in use ( is the device usage status, is the device computing power capacity); in the formula, this part reflects the utilization efficiency of computing resources, and the smaller the value, the higher the utilization rate of computing power resources. is the weight coefficient, used to adjust the degree of attention to computing efficiency, the larger it is, the more the system attaches importance to the efficient utilization of computing resources.

[0122] In actual operation, according to different application scenarios and requirements, reasonably set the values of and . For example, in an online game scenario with extremely high real-time requirements, can be appropriately increased to give priority to ensuring interaction satisfaction; in a scenario of large-scale data processing with relatively low real-time requirements, can be increased to improve the utilization efficiency of computing resources. By continuously optimizing this objective function, the system can dynamically adjust the task allocation and resource scheduling strategies to adapt to different task loads and user requirements, and achieve the balance of computing and interaction performance.

[0123] As an extension and refinement of the above embodiments, a method for dynamically allocating computing power resources shown in the present application further includes the following steps 1 to 3: Step 1: Obtain the status information of the target task and predict the expected task delay tolerance in combination with the user's historical behavior.

[0124] In some embodiments, the status information of the target task may include task priority, task type (data search, data analysis, etc.), task scale, resource utilization, etc. The user's historical behavior includes the degree of acceptance of delay by the user when performing similar tasks in the past (i.e., the user's historical delay tolerance data), and the expected task delay tolerance is the maximum delay time allowed by the user for the task.

[0125] Furthermore, in combination with various relevant information of the target task, according to the user's historical behavior, referring to the task delay tolerance prediction function shown in the following formula, calculate the ideal delay tolerance of the user for the target task, that is, the expected task delay tolerance : .

[0126] Where g(*) is the tolerance prediction function; is the context information of the current task (such as priority, task type, etc.); H is the user's historical delay tolerance data; U is the hardware resource utilization of the current task.

[0127] It should be noted that to ensure the smoothness of the interaction with the user, the predicted expected task delay tolerance needs to fall within a reasonable range, that is, set reasonable upper and lower limits for it, and the upper and lower limits are expressed as: . Where and respectively represent the upper and lower limits of the predicted expected task delay tolerance.

[0128] Step 2: Monitor the current response duration corresponding to the target task.

[0129] In some embodiments, by monitoring the progress of the target task, that is, recording in real time the time spent by the target task from the start of processing to the current moment, to obtain the current response duration Rk.

[0130] Step 3: Adjust the computing power resources of the target first device according to the size relationship between the current response duration and the expected task delay tolerance.

[0131] Furthermore, compare the current response duration Rk with the expected task delay tolerance . When (δ is a constant, such as 0.9), it means that the current task response time is within the acceptable range of the user, and continue to maintain the computing power resource allocation of the current target first device (fast computing device) without adjustment. When , which means that the task response time is too long, approaching or exceeding the user-acceptable range. At this time, computing resources can be added to the target first device, and the operating frequency of the target first device can be increased, etc., to reduce the current task response time and ensure that the task can be completed within the user-acceptable latency tolerance.

[0132] According to the actual execution situation of the task and the user's tolerance to latency, dynamically adjust the computing resources of the fast computing device to ensure the completion efficiency of high-priority tasks. In terms of resource allocation, a balance can be found between throughput and response speed, avoiding waste caused by over-allocation of resources and preventing excessive task latency caused by insufficient resources, thereby improving the user experience and the overall performance of the system.

[0133] As an extension and refinement of the above embodiments, the embodiments of the present application also set up a corresponding computing resource dynamic allocation system for the computing resource dynamic allocation method. The system includes: an analysis module, a calculation module, an interface module, and a scheduling module. Specifically: The analysis module is used to parse the user input text to determine the device for processing the task corresponding to the user input text in combination with the method of the above embodiments.

[0134] The calculation module is composed of multiple sub-calculation units, including a fast computing device sub-unit, a slow computing device calculation sub-unit, etc.; different sub-units are respectively equipped with their own computing resource pools.

[0135] The interface module realizes on-demand connection and dynamic scheduling of hardware modules through standardized interface design. The above design can be completed through the PCIe interface and the NVLink interface; the PCIe interface is a common high-speed serial computer expansion bus standard for connecting the motherboard to various hardware devices; the NVLink interface can achieve high-speed interconnection between GPUs and improve data transmission bandwidth.

[0136] The scheduling module can dynamically select and call appropriate modules from numerous calculation sub-units according to the current task requirements, and at the same time optimize the allocation of computing power to ensure the efficient execution of tasks.

[0137] When constructing the computing resource dynamic allocation system, in order to be able to real-time count the total computing power of the currently connected and available devices and reasonably allocate tasks in the future, it is necessary to obtain the total computing power of the entire system. The calculation formula is as follows: .

[0138] Among them, is the total computing power of the system; M is the total number of connected sub-calculation units; is the computing power of the i-th device (such as trillion floating-point operations per second TFLOPS); represents the device connection status (1 means connected, 0 means not connected).

[0139] Through the above formula, the overall computing power of the system can be calculated. When the system is expanded, for example, when a new computing subunit is connected, only the size of M needs to be increased. If its is 1, then it is further updated according to the formula.

[0140] It should be noted that the computing power resource dynamic allocation system also includes the following three parts of functions: The first part: Inspecting new devices when the system is expanded When a new computing device is connected to the system, with the help of a standardized interface, the device registration module checks the device according to the compatibility parameters The exemplary check expression is as follows: .

[0141] Among them, is the check result (1 means passed, 0 means not passed), is the device compatibility parameter (involving hardware architecture, communication protocol, etc.), is the minimum compatibility requirement of the current system. If the result of the above check expression is 1, it indicates that the device meets the minimum requirements. When the device compatibility parameter meets the minimum requirements, it is determined to pass the check. The devices that pass the check will be added to the computing power resource pool, and can become the computing resources that can be called by the system and dynamically join the scheduling process to prepare for subsequent task allocation.

[0142] Furthermore, after the new device passes the inspection, it can be put into subsequent normal use.

[0143] The second part: Distributed task execution and migration of the system For a certain task , it can be broken down into a set of subtasks executed by multiple threads or multiple computing subunits , and each subtask is a subset of, and the sum of the computing requirements of all subtasks is equal to the global task computing requirement , specifically referring to the following formula: .

[0144] For example, in big data distributed computing, a large-scale data processing task can be split into multiple subtasks according to data partitioning. According to the computing power resource status of the entire system, the subtasks are allocated to available sub-computing units for execution respectively. The scheduling module coordinates the data flow between modules through the bus or network interface to ensure the collaborative work of each subtask.

[0145] When the system executes the above-mentioned multiple subtasks, it is also necessary to continuously monitor the resource utilization rate corresponding to each sub-computation unit in the system; when the resource utilization rate of a certain sub-computation unit module is higher than the specified threshold, it is determined that the sub-computation unit is overloaded. For example, the server CPU utilization rate continuously exceeds 80% (assuming is 80%). If there is in the sub-computation unit, then the system migrates tasks according to the following formula: .

[0146] Among them, is the amount of task calculation to be migrated, is the migration ratio factor; is the resource utilization rate of the current sub-computation unit. Furthermore, by migrating tasks, the load between modules can be made more balanced, preventing the performance bottleneck of a single device from affecting the overall system efficiency. For example, in a distributed file storage system, some file transfer tasks of overloaded storage nodes are migrated to idle nodes.

[0147] Part Two: Targeted Resource Expansion and Compatibility When facing multiple types of computing tasks, for example, when the target task is a high-priority task, there may be a scenario where it needs to be jointly completed by fast computing devices and slow computing devices. Therefore, when deciding on the computing devices corresponding to the target task, it is necessary to consider whether the current overall computing power resources can support the target task, which is convenient for subsequent allocation and expansion of computing power resources. Since the computing power resources in this application are divided into two main categories, fast computing devices and slow computing devices; assuming there are fast computing devices and slow computing devices in the current entire system, the total computing power resources can be described by the following formula: .

[0148] Among them, represents the computing power of the i-th fast computing device; represents the computing power of the j-th slow computing device; is the number of fast computing devices; is the number of slow computing devices.

[0149] Through the above formula, the total computing power resources of fast computing devices and slow computing devices can be calculated, providing a basis for subsequent scheduling and allocation. Furthermore, based on the total computing power and task requirements, tasks can be reasonably arranged to different devices to achieve efficient resource utilization, which helps to determine whether there is sufficient computing power to handle newly submitted high-priority tasks.

[0150] Therefore, when it is determined according to the above formula that the system needs to be expanded, new devices can be introduced in a planned manner to address the problem of insufficient computing power. The computing power of the devices to be introduced needs to be calculated according to the following formula: 。

[0151] Among them, is the total computing power requirement of the current task, is the available computing power of the current system. For example, if the total computing power requirement is 100 TFLOPS and the available computing power of the current system is 80 TFLOPS, then , that is, new devices with at least 20 TFLOPS of computing power need to be introduced.

[0152] At the same time, it should be noted that before expanding new sub-computing units, it is also necessary to automatically adapt the hardware protocol through the interface layer design to ensure the unity of the global environment. The formula corresponding to the above conditions is: 。

[0153] Among them, is the hardware compatibility parameter, is the interface protocol compatibility parameter, is the node device compatibility parameter. For example, in the interface protocol of the new hardware device and the compatibility of the node device itself, the lower compatibility parameter value is taken as the overall hardware compatibility standard.

[0154] Furthermore, through the above method, it is ensured that the new sub-computing unit can seamlessly integrate into the existing architecture, realize dynamic performance expansion in complex scenarios, and avoid system failures or performance degradation caused by hardware compatibility problems.

[0155] Therefore, the computing power resource dynamic allocation system of the embodiment of the present application can dynamically propose expansion requirements according to the change of the current task load, ensuring that the computing power resources fit the actual load. During the peak period of tasks, the computing power gap is evaluated in a timely manner and expansion is planned to avoid affecting the task processing efficiency due to insufficient computing power.

[0156] Based on the same inventive concept, as an implementation of the above method, the embodiment of the present application also provides a computing power resource dynamic allocation device. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details of the foregoing method embodiment will not be described one by one in this embodiment. However, it should be clear that a computing power resource dynamic allocation device in this embodiment can correspondingly implement all the contents of the foregoing method embodiment.

[0157] The embodiment of the present application provides a computing power resource dynamic allocation device, Figure 4 which is a schematic structural diagram of the computing power resource dynamic allocation device. As shown in Figure 4 shown, the computing power resource dynamic allocation device 400 includes: A parsing unit 401, configured to parse the user input text to obtain the task priority score of the target task corresponding to the user input text; A determination unit 402, configured to determine that the target task is a high-priority task if the task priority score is greater than or equal to a first preset threshold; the high-priority task is a high-priority task that needs to be processed in real time; An acquisition unit 403, configured to acquire the call status of each first device; the first device is a fast computing device for processing high-priority tasks; the call status includes an idle status and a used status; A first allocation unit 404, when there are multiple first devices in the idle state, determines a target first device from the multiple first devices, and allocates the target task to the target first device; A second allocation unit 405, configured to, when all of the multiple first devices are in the used state, allocate the target task to a second device in the idle state; the second device is a slow computing device for processing low-priority tasks.

[0158] As an optional implementation manner of an embodiment of the present application, the first allocation unit 404 is specifically configured to acquire the computing power resources corresponding to each of the multiple first devices to be used, and the computing power requirement corresponding to the target task, so as to calculate the predicted response time corresponding to each of the multiple first devices to be used; based on the magnitude relationship between the predicted response time and a preset response time threshold, determine the target first device.

[0159] As an optional implementation manner of an embodiment of the present application, the parsing unit 401 is further configured to determine that the target task is a low-priority task if the task priority score is less than or equal to a second preset threshold; the second preset threshold is less than the first preset threshold; acquire the call status of each second device; determine a second device to be used in the idle state from the multiple second devices, and determine a target second device from the multiple second devices to be used, and allocate the target task to the target second device.

[0160] As an optional implementation manner of an embodiment of the present application, the second allocation unit 405 is specifically configured to acquire the computing power resources corresponding to each of the multiple second devices to be used, and the computing power requirement corresponding to the target task, so as to determine the maximum task throughput efficiency; based on the maximum task throughput efficiency, determine the target second device.

[0161] As an optional implementation manner of an embodiment of the present application, the parsing unit 401 is specifically configured to parse the urgency and real-time requirements of the target task corresponding to the user input text based on a lightweight user intention recognition model, and acquire the corresponding task priority score.

[0162] As an optional implementation manner of the embodiment of the present application, the obtaining unit 403 is further configured to obtain the status information of the target task, and predict the expected task delay tolerance in combination with the user's historical behavior; monitor the current response duration corresponding to the target task; and adjust the computing power resources of the target first device according to the magnitude relationship between the current response duration and the expected task delay tolerance.

[0163] As an optional implementation manner of the embodiment of the present application, the computing power resource dynamic allocation device further includes an updating unit, which is specifically configured to collect the complete response time and computing power usage rate corresponding to the target task after the execution of the target task ends; and update the first preset threshold and the second preset threshold according to the complete response time and the computing power usage rate.

[0164] Based on the same inventive concept, the embodiments of the present disclosure further provide an electronic device. Figure 5 The structural schematic diagram of the electronic device provided by the embodiment of the present disclosure is shown in Figure 5 As shown, the electronic device provided in this embodiment includes: a memory 501 and a processor 502. The memory 501 is used to store a computer program; the processor 502 is configured to execute the computing power resource dynamic allocation method provided in the above embodiment when executing the computer program.

[0165] Based on the same inventive concept, the embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, a computing device is enabled to implement the computing power resource dynamic allocation method provided in the above embodiment.

[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0167] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0168] The memory may include non - permanent memory in the form of computer - readable media, random access memory (RAM) and / or non - volatile memory such as read - only memory (ROM) or flash RAM. The memory is an example of computer - readable media.

[0169] Computer - readable media includes permanent and non - permanent, removable and non - removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer - readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices or any other non - transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer - readable media does not include transitory computer - readable media such as modulated data signals and carrier waves.

[0170] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for dynamically allocating computing power resources, characterized in that, Including: Parsing the user input text to obtain the task priority score of the target task corresponding to the user input text; If the task priority score is greater than or equal to a first preset threshold, determining that the target task is a high-priority task; Obtaining the call status of each first device; The first device is a fast computing device for processing the high-priority task; the call status includes an idle state and a used state; When there are multiple first devices to be used in the idle state, determining a target first device from the multiple first devices to be used, and allocating the target task to the target first device; When all of the multiple first devices are in the used state, allocating the target task to a second device in the idle state; the second device is a slow computing device for processing low-priority tasks.

2. The method according to claim 1, characterized in that, The step of, when there are multiple first devices in the idle state, determining a target first device from the multiple first devices includes: Obtaining the computing power resources corresponding to the multiple first devices to be used respectively, and the computing power requirements corresponding to the target task, to calculate the predicted response times corresponding to the multiple first devices to be used respectively; Determining the target first device based on the magnitude relationship between the predicted response time and a preset response time threshold.

3. The method according to claim 1, characterized in that, After parsing the user input text to obtain the task priority score of the target task corresponding to the user input text, the method further includes: If the task priority score is less than or equal to a second preset threshold, determining that the target task is a low-priority task; the second preset threshold is less than the first preset threshold; Obtaining the call status of each second device; Determining a second device to be used in the idle state from the multiple second devices, and determining a target second device from the multiple second devices to be used, and allocating the target task to the target second device.

4. The method according to claim 3, characterized in that, The step of determining a second device to be used in the idle state from the multiple second devices, and determining a target second device from the multiple second devices to be used includes: Obtaining the computing power resources corresponding to the multiple second devices to be used respectively, and the computing power requirements corresponding to the target task, to determine the maximum task throughput efficiency; Determining the target second device based on the maximum task throughput efficiency.

5. The method according to claim 1, characterized in that The step of parsing the user input text to obtain the task priority of the target task corresponding to the user input text includes: Based on a lightweight user intention recognition model, parsing the urgency and real-time requirements of the target task corresponding to the user input text, to obtain the corresponding task priority score.

6. The method according to claim 1, characterized in that, The method further includes: Obtaining the status information of the target task, and predicting the expected task delay tolerance in combination with the user's historical behavior; Monitoring the current response duration corresponding to the target task; Adjusting the computing power resources of the target first device according to the magnitude relationship between the current response duration and the expected task delay tolerance.

7. The method according to claim 3, wherein The method further includes: After the target task is executed, collecting the complete response time and computing power utilization rate corresponding to the target task; Update the first preset threshold and the second preset threshold according to the complete response time and the computing power utilization rate.

8. A computing power resource dynamic allocation device, characterized in that, Comprising: A parsing unit configured to parse the user input text to obtain a task priority score of a target task corresponding to the user input text; A determining unit configured to determine that the target task is a high-priority task if the task priority score is greater than or equal to a first preset threshold; An obtaining unit configured to obtain the call status of each first device; The first device is a fast computing device for processing the high-priority task; the call status includes an idle state and a used state; A first allocation unit, when there are multiple first devices in the idle state, determines a target first device from the multiple first devices and allocates the target task to the target first device; A second allocation unit configured to, when all of the multiple first devices are in the used state, allocate the target task to a second device in the idle state; the second device is a slow computing device for processing low-priority tasks.

9. An electronic device, characterized in that, Comprising: A memory and a processor, the memory is configured to store a computer program; the processor is configured to, when executing the computer program, enable the electronic device to implement the computing power resource dynamic allocation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a computing device, the computing device is enabled to implement the computing power resource dynamic allocation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Task optimization method and device

    CN114741075A

  • Intelligent scheduling system and method based on calculation power demand prediction

    CN120066720A

  • Scheduler, multi-core processor system, and scheduling method

    US20130138886A1

  • Method and apparatus for processing task, and device and storage medium

    WO2024041401A1

Cited By

  • Camera computing power dynamic allocation method and system based on heterogeneous computing power cutting

    CN121792873A

  • A camera computing power dynamic allocation method and system based on heterogeneous computing power pruning

    CN121792873B