Dynamic allocation method and device for AI reasoning tasks
By real-time monitoring and dynamic allocation of AI inference tasks to computing nodes in the cloud-edge collaborative system, and adjusting the model complexity, the computing limitation problem in the cloud and edge-side is solved, and efficient resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202510748989.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-06
AI Technical Summary
In AI applications, traditional cloud processing introduces latency and bandwidth limitations, while edge processing is limited by computing power, memory and energy consumption. How to dynamically schedule AI inference workloads between edge computing and cloud computing for more efficient performance and response speed is a key issue.
By monitoring the load, memory, network latency and energy consumption of computing nodes in the cloud-edge collaborative system in real time, dynamically allocate AI inference tasks to different computing nodes, and adjust the complexity of the AI model according to the node status, and optimize the task allocation strategy using a hierarchical computing architecture and reinforcement learning algorithm.
It realizes efficient resource utilization among different computing nodes, improves the execution efficiency and overall performance of AI inference tasks, and adapts to different task requirements and environmental changes.
Smart Images

Figure CN120343029A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence computing, and more particularly, to a method and apparatus for dynamically allocating AI inference tasks. Background Art
[0002] In many AI applications, especially in autonomous driving, smart cities, industrial automation, and Internet of Things devices, real-time response is crucial. Although traditional cloud AI processing has powerful computing capabilities, due to its dependence on network transmission, it introduces significant latency and bandwidth limitations that cannot be ignored. Edge processing, on the other hand, can reduce latency but is typically limited by computing power, memory, and energy consumption. Therefore, how to dynamically schedule AI inference workloads between edge computing and cloud computing to achieve more efficient performance and response speed has become the key to improving the effectiveness of AI applications. Summary of the Invention
[0003] To optimize the performance of AI applications, the embodiments described herein provide a method and apparatus for dynamically allocating AI inference tasks, as well as a computer-readable storage medium storing a computer program. By considering multiple factors such as real-time network conditions, load, energy availability, and required response time, AI inference tasks are dynamically allocated among edge devices, edge servers, and the cloud to achieve flexible workload scheduling and meet the real-time and computing requirements of different application scenarios.
[0004] According to a first aspect of the present disclosure, there is provided a method for dynamically allocating AI inference tasks, including: real-time monitoring of the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in a cloud-edge collaboration system, where the computing nodes include local edge devices, edge servers, and cloud platforms; dynamically allocating tasks to different computing nodes according to the latency requirements, computing requirements, and data privacy requirements of the tasks, as well as the real-time status of each computing node; adjusting the complexity of the AI model used to execute the tasks according to the resource usage of the computing nodes; and optimizing the task allocation strategy according to the historical performance data and real-time network conditions during the task execution process.
[0005] In some embodiments of the present disclosure, the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the real-time monitoring cloud-edge collaboration system are monitored. The computing nodes include local edge devices, edge servers, and cloud platforms, including: using system monitoring tools to track the usage rates of CPUs, GPUs, and threads, the number of processing tasks, and the execution time of the computing nodes, and evaluating the load levels of each node; checking the memory usage of each computing node and monitoring the total memory and remaining memory of each node; running network performance tests to evaluate the latency time, upload and download bandwidth, and packet loss rate between different computing nodes; monitoring the energy consumption of each computing node in real time; and transmitting the monitoring data from each computing node to a centralized data platform, storing the historical monitoring data using a distributed storage system, and displaying the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in real time.
[0006] In some embodiments of the present disclosure, tasks are dynamically allocated to different computing nodes according to the latency requirements, computing requirements, and data privacy requirements of the tasks, as well as the real-time status of each computing node, including: collecting data of the computing nodes through sensors, analyzing the task requirements based on the collected data, determining whether the task has strict latency requirements, clarifying the computing resources required by the task, and whether it involves sensitive data; classifying AI inference tasks and allocating AI inference tasks to different computing nodes through a hierarchical computing architecture.
[0007] In some embodiments of the present disclosure, classifying AI inference tasks and allocating AI inference tasks to different computing nodes through a hierarchical computing architecture includes: preferentially allocating low-latency tasks, lightweight inference tasks, preliminary data processing, and AI tasks involving sensitive data to local edge devices; preferentially allocating tasks with high latency tolerance and high computing requirements to cloud platforms; and dynamically selecting the optimal node according to the current load conditions of each computing node.
[0008] In some embodiments of the present disclosure, adjusting the complexity of the AI model used to execute tasks according to the resource usage of the computing nodes includes: for computing nodes with limited resources, reducing the number of layers or neurons of the AI model by means of quantization, pruning, or model segmentation; for computing nodes with sufficient resources, increasing the number of layers or neurons of the AI model.
[0009] In some embodiments of the present disclosure, optimizing the task allocation strategy according to historical performance data and real-time network conditions during task execution includes: real-time tracking and recording the processing capacity, latency, load, and computing resource usage of each computing node; real-time monitoring of network bandwidth, latency, and stability to determine whether the current network environment is suitable for task execution or whether task allocation needs to be adjusted; continuously monitoring the execution time, resource consumption, and task completion quality of inference tasks; and dynamically adjusting the task allocation strategy through a reinforcement learning algorithm.
[0010] In some embodiments of the present disclosure, dynamically adjusting the task allocation strategy through a reinforcement learning algorithm includes: defining a state space and an action space, where the state space includes historical performance data, the current network condition, and task complexity information, and the action space includes decisions to allocate tasks to different computing nodes; calculating a reward based on the effect of task execution, and giving a positive reward if the task is successfully completed within a predetermined time and the computing resources are used properly; and giving a negative reward if the task execution fails or the resources are used improperly.
[0011] In some embodiments of the present disclosure, the method further includes: during the data transmission process between different computing nodes, compressing and encrypting the transmitted data across nodes.
[0012] According to a second aspect of the present disclosure, there is provided a dynamic allocation device for AI inference tasks. The device includes at least one processor; and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the device is caused to: real-time monitor the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the cloud-edge collaboration system, where the computing nodes include local edge devices, edge servers, and cloud platforms; dynamically allocate tasks to different computing nodes according to the latency requirements, computing requirements, and data privacy requirements of the tasks and the real-time status of each computing node; adjust the complexity of the AI model used to execute the tasks according to the resource usage of the computing nodes; and optimize the task allocation strategy according to the historical performance data and real-time network conditions during task execution.
[0013] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the steps of the method according to the first aspect of the present disclosure.
[0014] The method and device for dynamically allocating AI inference tasks according to embodiments of the present disclosure can effectively cope with differences in computing capabilities of different nodes and avoid overconsumption of resources by monitoring the resource usage of computing nodes in real time and dynamically adjusting the complexity of the AI models for task execution. In a multi-level computing environment such as local edge devices, edge servers, and cloud platforms, the execution of tasks can adjust the complexity of the computing model according to the current state of the nodes, thereby achieving more efficient resource utilization, and can automatically adjust the task allocation strategy according to the requirements of different tasks, environmental changes, and resource status. With the accumulation of task execution, continuous learning and optimization are carried out, and finally a more intelligent and efficient task allocation model is formed. Therefore, this solution can significantly improve the execution efficiency and overall performance of AI inference tasks in the cloud-edge collaboration system. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To briefly describe the technical solutions of the embodiments of the present disclosure more clearly, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the following-described drawings only relate to some embodiments of the present disclosure and do not limit the present disclosure, where: Figure 1 is an exemplary flowchart of the method for dynamically allocating AI inference tasks according to embodiments of the present disclosure; Figure 2 is a schematic block diagram of the device for dynamically allocating AI inference tasks according to embodiments of the present disclosure.
[0016] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts shall also fall within the scope of protection of the present disclosure.
[0018] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the subject matter of the present disclosure belongs. Further, it will be understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless expressly defined herein otherwise. Additionally, terms such as "first" and "second" are only used to distinguish one component (or a part of a component) from another component (or another part of a component).
[0019] To improve the processing efficiency and response speed of AI inference tasks, embodiments of the present disclosure provide a dynamic allocation method for AI inference tasks, which can optimize the allocation and execution of AI inference tasks in a cloud-edge collaborative environment. Combining real-time resource monitoring, intelligent task scheduling, and AI model adaptive adjustment can significantly improve the overall performance of the system and optimize resource utilization efficiency.
[0020] Figure 1 FIG. shows an exemplary flowchart of a dynamic allocation method for AI inference tasks according to an embodiment of the present disclosure. Refer to Figure 1 As shown, at block S102 of the dynamic allocation method 100 for AI inference tasks, the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the cloud-edge collaborative system are monitored in real time. The computing nodes include local edge devices, edge servers, and cloud platforms.
[0021] According to an embodiment of the present disclosure, the cloud-edge collaborative system includes multi-level computing nodes, consisting of local edge devices, edge servers, and cloud platforms. The local edge devices (such as smartphones, cameras, embedded AI chips) are responsible for the preliminary processing of tasks, especially suitable for low-latency and high-frequency processing requirements, and can execute some relatively simple AI models or tasks, such as image preprocessing, preliminary analysis of sensor data, etc. The edge servers are located closer to users, usually in facilities such as 5G base stations and micro data centers, and can perform more in-depth analysis and calculations. The cloud platform is suitable for tasks that require a large amount of computing resources, data storage, and large-scale model training. Monitoring agents are deployed on each computing node (local edge device, edge server, and cloud platform). Each agent is responsible for collecting the usage of local resources and transmitting the data to a centralized monitoring platform.
[0022] System monitoring tools (such as top, htop, Prometheus) can be used to track the CPU, GPU, and thread usage rates, the number of processed tasks, and the execution time of the computing nodes, and evaluate the load level of each node. Tools (such as free, vmstat, Prometheus) are used to collect and analyze the memory usage, check the memory usage of each computing node, and monitor the total memory and remaining memory of each node. Network performance tests (such as ping or traceroute) are run to evaluate the latency time, upload and download bandwidth usage, and packet loss rate between different computing nodes. Energy monitoring hardware is used to monitor the energy consumption of each computing node.
[0023] Finally, the monitoring data is transmitted from each computing node to a centralized data platform, and a distributed storage system is used to store the historical monitoring data and display the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in real time.
[0024] Subsequently, in block S104, according to the latency requirements, computing requirements, data privacy requirements of the task, and the real-time status of each computing node, the task is dynamically allocated to different computing nodes.
[0025] Edge computing executes preliminary tasks with low computing resource consumption at a location closer to the data source, which can reduce data transmission time and bandwidth requirements. Cloud computing can utilize a large amount of computing resources, storage space, and powerful computing capabilities for high-precision analysis and task completion, but may be affected by network latency and thus impact the response speed. Allocating computing tasks to the cloud can handle large-scale training tasks, but for inference tasks that require low latency, edge devices are more suitable. The system can dynamically select edge devices or the cloud for inference processing according to different task requirements and resource availability.
[0026] According to an embodiment of the present disclosure, the hierarchical computing architecture of the multi-level AI processing node can efficiently schedule according to the requirements of different tasks and the availability of computing resources, thereby realizing the efficient allocation and execution of AI inference tasks. Edge devices such as smartphones, cameras, and embedded AI chips are mainly responsible for executing lightweight inference tasks or preliminary data processing. For example, performing preliminary analysis and detection of images or basic preprocessing of voice signals. In addition, AI tasks involving sensitive data are processed at the local node to ensure data privacy and security. Edge servers such as 5G base stations and micro data centers can handle tasks that require more computing resources, or medium-complexity inference tasks transmitted from edge devices. For example, partial inference of deep learning models, data aggregation, and preliminary data analysis. Cloud computing is suitable for AI tasks that require a large amount of computing resources, such as deep learning training, large-scale data processing, and high-precision inference.
[0027] Data of the computing node can be collected through sensors, and the task requirements are analyzed according to the collected data. For example, cameras, microphones, IoT devices, etc. are used to collect data of the computing node in real time, including sensor information such as images, sounds, temperature, and humidity, for evaluating the current environment and task requirements. After data preprocessing, the task requirements are analyzed to determine whether the task has strict latency requirements, clarify the computing resources required for the task (such as CPU, GPU, memory, etc.), and whether it involves sensitive data.
[0028] Classify AI inference tasks and allocate AI inference tasks to different computing nodes through a hierarchical computing architecture. For example, low-latency tasks, lightweight inference tasks, preliminary data processing, and AI tasks involving sensitive data are preferentially allocated to local edge devices. Among them, low-latency tasks are tasks with high real-time requirements, such as speech recognition, real-time video processing, etc. Lightweight inference tasks are tasks with low computational complexity and are suitable for fast processing and real-time feedback. Preliminary data processing tasks are preliminary operations such as preprocessing data and feature extraction, and these operations usually do not require large-scale computing resources. Tasks with high latency tolerance and high computational requirements are preferentially allocated to the cloud platform. For example, deep learning model training, large-scale data analysis, etc. are preferentially allocated to the cloud platform for processing. Dynamically select the optimal node according to the current load situation of each computing node. For example, if a certain node has a high load and tight resources, the task can be allocated to a node with a lower load. During the task execution process, if the resource usage of a certain node changes (such as too high load, increased latency, etc.), the task can be transferred to other nodes through dynamic migration to ensure the efficiency and stability of the system. Through comprehensive analysis of the latency requirements, computational requirements, data privacy requirements of the tasks and the real-time status of each computing node, optimal allocation of tasks can be achieved.
[0029] Since data exchange between computing nodes is inevitable in AI inference tasks. In this environment, the security and transmission efficiency of data are particularly important. In an embodiment of the present disclosure, during the data transmission process between different computing nodes, data compression and encryption are performed on the transmission data across nodes, which can ensure the security and efficiency of data transmission to the greatest extent.
[0030] Then in block S106, according to the resource usage of the computing node, adjust the complexity of the AI model used to execute the task.
[0031] According to an embodiment of the present disclosure, for resource-constrained computing nodes, the number of layers or neurons of an AI model can be reduced by means of quantization, pruning, or model splitting. Quantization is to reduce the floating-point precision of the AI model to a lower precision (such as using integers with low bit widths instead of floating-point numbers), thereby reducing the storage and computing requirements of the model. For example, quantizing a 32-bit floating model to an 8-bit integer model can significantly reduce memory usage and computing load. Pruning is to reduce the computational complexity of the model by removing redundant neurons or connections in the neural network. For example, structured pruning or unstructured pruning techniques can be adopted to delete parameters that have less impact on the model performance. Model splitting is to split a large-scale AI model into multiple sub-models, and each sub-model is only responsible for a part of the model's functions. In this way, the model can be distributed and executed among multiple computing nodes, reducing the computational burden on a single node. The split model can run the simpler parts on local devices, while the more complex parts are pushed to edge servers or the cloud with stronger computing capabilities for processing.
[0032] For computing nodes with sufficient resources, increase the number of layers or neurons of the AI model. By increasing the number of layers of the neural network, the learning ability and expression ability of the model can be improved. More layers usually mean that the model can learn more complex features and are suitable for processing high-precision or complex tasks. Increasing the number of neurons in each layer can increase the computing power of the model, enabling it to process more complex features and tasks. This usually increases the computing requirements, but in an environment with sufficient resources, the performance of the model can be effectively improved.
[0033] During the task execution process, continuously monitor the resource usage of each computing node (such as memory, computing load, bandwidth, etc.), and evaluate in real time whether the complexity of the model needs to be adjusted. If resource tension is detected (for example, insufficient memory, high CPU load, etc.), the size and precision of the model can be automatically adjusted to ensure the continuous execution of the task without resource overload.
[0034] Finally, in block S108, optimize the task allocation strategy according to the historical performance data and real-time network conditions during the task execution process.
[0035] By tracking and recording in real time metrics such as the processing capacity, latency, load, and computing resource usage of each node, understand the performance of different nodes when processing different tasks. Monitor the network bandwidth, latency, and stability in real time to determine whether the current network environment is suitable for task execution or whether the task allocation needs to be adjusted. Continuously monitor the execution time, resource consumption, and task completion quality of the inference task. Dynamically adjust the task allocation strategy through reinforcement learning algorithms.
[0036] In some embodiments of the present disclosure, the specific process of the feedback-based optimization mechanism is as follows: perform preliminary workload allocation according to the current network conditions, historical performance data, and the basic characteristics of the tasks (such as task type, priority, resource requirements). During the task execution, continuously monitor the execution status of the tasks, including the load of nodes, the progress of tasks, network stability, etc. Dynamically adjust the workload allocation strategy through a reinforcement learning algorithm to optimize task execution. The policy update may include adjusting the task allocation priority, selecting a more suitable computing node, or adjusting the parameters of the allocation strategy. Among them, the reinforcement learning algorithm includes: defining the state space and the action space. The state space includes information such as historical performance data, current network conditions, and task complexity. The action space includes decisions to allocate tasks to different computing nodes, such as edge devices, edge servers, or cloud computing nodes. Calculate the reward according to the effect of task execution. For example, if the task is successfully completed within the predetermined time and the computing resources are used properly, the system will give a positive reward; if the task execution fails or the resources are used improperly, a negative reward will be given. Through the feedback mechanism, the system can determine which tasks should be preferentially allocated to edge devices, which need to be processed by edge servers, and which need to be pushed to the cloud. This not only ensures the fast processing of low-latency tasks but also does not waste the computing resources of edge nodes on high-complexity tasks.
[0037] Figure 2 is a schematic block diagram of a dynamic allocation device for AI inference tasks according to an embodiment of the present disclosure. As Figure 2 shown, the device 200 may include a processor 210 and a memory 220 storing a computer program. When the computer program is executed by the processor 210, the device 200 can execute the steps of the method 100 as Figure 1 shown. In one example, the device 200 may be a computer device or a cloud computing node. The device 200 can monitor the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the cloud-edge collaboration system in real time. The computing nodes include local edge devices, edge servers, and cloud platforms; dynamically allocate tasks to different computing nodes according to the latency requirements, computing requirements, and data privacy requirements of the tasks and the real-time status of each computing node; adjust the complexity of the AI model used to execute the tasks according to the resource usage of the computing nodes; and optimize the task allocation strategy according to the historical performance data and real-time network conditions during the task execution process.
[0038] In some embodiments of the present disclosure, the device 200 can use system monitoring tools to track the CPU, GPU, and thread utilization rates, the number of processing tasks, and the execution time of computing nodes, evaluate the load level of each node; check the memory usage of each computing node, and monitor the total memory and remaining memory of each node; run network performance tests to evaluate the latency time, upload and download bandwidth, and packet loss rate between different computing nodes; monitor the energy consumption of each computing node in real time; and transmit the monitoring data from each computing node to a centralized data platform, store the historical monitoring data using a distributed storage system, and display the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in real time.
[0039] In some embodiments of the present disclosure, the device 200 can collect data of computing nodes through sensors, analyze the task requirements based on the collected data, determine whether the task has strict requirements for latency, clarify the computing resources required by the task, and whether it involves sensitive data; classify AI inference tasks, and allocate AI inference tasks to different computing nodes through a hierarchical computing architecture.
[0040] In some embodiments of the present disclosure, the device 200 can preferentially allocate low-latency tasks, lightweight inference tasks, preliminary data processing, and AI tasks involving sensitive data to local edge devices; preferentially allocate tasks with high latency tolerance and high computing requirements to cloud platforms; dynamically select the optimal node according to the current load conditions of each computing node.
[0041] In some embodiments of the present disclosure, for computing nodes with limited resources, the device 200 can reduce the number of layers or neurons of the AI model through quantization, pruning, or model segmentation; for computing nodes with sufficient resources, increase the number of layers or neurons of the AI model. In some embodiments of the present disclosure, the device 200 can track and record the processing capabilities, latency, load, and computing resource usage of each computing node in real time; monitor the network bandwidth, latency, and stability in real time, and determine whether the current network environment is suitable for task execution or whether task allocation needs to be adjusted; continuously monitor the execution time, resource consumption, and task completion quality of inference tasks; dynamically adjust the task allocation strategy through a reinforcement learning algorithm. In some embodiments of the present disclosure, the device 200 can define a state space and an action space. The state space includes historical performance data, the current network condition, and task complexity information. The action space includes decisions to allocate tasks to different computing nodes; calculate rewards based on the execution effect of the task. If the task is successfully completed within a predetermined time and the computing resources are used properly, a positive reward is given; if the task execution fails or the resources are used improperly, a negative reward is given.
[0042] In some embodiments of the present disclosure, the device 200 may perform data compression and encryption on the transmission data across nodes during the data transmission process between different computing nodes.
[0043] In an embodiment of the present disclosure, the processor 210 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. The memory 220 may be any type of memory implemented using data storage technologies, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk memory, etc.
[0044] In addition, in an embodiment of the present disclosure, the device 200 may also include an input device 230, such as a keyboard, a mouse, etc. Additionally, the device 200 may further include an output device 240, such as a display, etc.
[0045] In other embodiments of the present disclosure, a computer-readable storage medium storing a computer program is further provided, wherein the computer program, when executed by a processor, is capable of implementing the steps of the method as Figure 1 shown.
[0046] In summary, according to the method and device for dynamically allocating AI inference tasks in embodiments of the present disclosure, by monitoring the resource usage of computing nodes in real time and dynamically adjusting the complexity of the AI model for task execution, it is possible to effectively cope with the differences in computing capabilities of different nodes and avoid excessive resource consumption. In a multi-level computing environment such as local edge devices, edge servers, and cloud platforms, the execution of tasks can adjust the complexity of the computing model according to the current state of the nodes, thereby achieving more efficient resource utilization, and can automatically adjust the task allocation strategy according to the requirements of different tasks, environmental changes, and resource status. With the accumulation of task execution, continuous learning and optimization are carried out, and finally a more intelligent and efficient task allocation model is formed. Therefore, this solution can significantly improve the execution efficiency and overall performance of AI inference tasks in the cloud-edge collaboration system.
[0047] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0048] Unless the context clearly dictates otherwise herein, the singular forms of words used in this specification and the appended claims also include the plural, and vice versa. Thus, when reference is made to the singular, the corresponding plural is generally included. Similarly, the terms "comprising" and "including" are to be construed as inclusive rather than exclusive. Likewise, the term "or" should be interpreted as inclusive, unless expressly prohibited by the context herein. Where the term "exemplary" is used herein, particularly when it is after a list of terms, the "exemplary" is merely illustrative and explanatory and should not be considered exclusive or extensive.
[0049] Further aspects and scope of adaptability become apparent from the description provided herein. It should be understood that the various aspects of the present application may be implemented individually or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present application.
[0050] The above has described in detail several embodiments of the present disclosure. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The protection scope of the present disclosure is defined by the appended claims.
Claims
1. A dynamic allocation method for AI inference tasks, characterized in that Including: Real-time monitoring of the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the cloud-edge collaboration system, where the computing nodes include local edge devices, edge servers, and cloud platforms; Dynamically allocating tasks to different computing nodes according to the latency requirements, computing demands, data privacy requirements of the tasks, and the real-time status of each computing node; Adjusting the complexity of the AI model used to execute the task according to the resource usage of the computing node; And Optimizing the task allocation strategy according to the historical performance data and real-time network conditions during the task execution process.
2. The dynamic allocation method for AI inference tasks according to claim 1, wherein The real-time monitoring of the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in the cloud-edge collaboration system, where the computing nodes include local edge devices, edge servers, and cloud platforms includes: Using system monitoring tools to track the usage rates of the CPU, GPU, and threads, the number of tasks processed, and the execution time of the computing node, and evaluating the load level of each node; Checking the memory usage of each computing node and monitoring the total memory and remaining memory of each node; Running network performance tests to evaluate the latency time, upload and download bandwidth, and packet loss rate between different computing nodes; Real-time monitoring of the energy consumption of each computing node; and Transmitting the monitoring data from each computing node to a centralized data platform, storing the historical monitoring data using a distributed storage system, and displaying the workload, memory availability, network latency, bandwidth, and energy consumption of each computing node in real time.
3. The dynamic allocation method for AI inference tasks according to claim 1, wherein The dynamically allocating tasks to different computing nodes according to the latency requirements, computing demands, data privacy requirements of the tasks, and the real-time status of each computing node includes: Collecting data of the computing node through sensors, analyzing the task requirements based on the collected data, determining whether the task has strict latency requirements, clarifying the computing resources required by the task, and whether it involves sensitive data; and Classifying AI inference tasks and allocating AI inference tasks to different computing nodes through a hierarchical computing architecture.
4. The dynamic allocation method for AI inference tasks according to claim 3, wherein The classifying AI inference tasks and allocating AI inference tasks to different computing nodes through a hierarchical computing architecture includes: Prioritizing the allocation of low-latency tasks, lightweight inference tasks, preliminary data processing, and AI tasks involving sensitive data to local edge devices; Prioritizing the allocation of tasks with high latency tolerance and high computing demands to the cloud platform; and Dynamically selecting the optimal node according to the current load situation of each computing node.
5. The dynamic allocation method for AI inference tasks according to claim 1, characterized in that The adjusting the complexity of the AI model used to execute the task according to the resource usage of the computing node includes: For computing nodes with resource constraints, reducing the number of layers or neurons of the AI model by means of quantization, pruning, or model splitting; For computing nodes with sufficient resources, increasing the number of layers or neurons of the AI model.
6. The dynamic allocation method for AI inference tasks according to claim 1, characterized in that, The optimizing the task allocation strategy according to the historical performance data and real-time network conditions during the task execution process includes: Real-time tracking and recording the processing capacity, latency, load, and computing resource usage of each computing node; Monitor network bandwidth, latency, and stability in real time to determine whether the current network environment is suitable for task execution or whether task allocation needs to be adjusted; Continuously monitor the execution time, resource consumption, and task completion quality of the inference task; and Dynamically adjust the task allocation strategy through a reinforcement learning algorithm.
7. The dynamic allocation method for AI inference tasks according to claim 6, characterized in that The dynamically adjusting the task allocation strategy through the reinforcement learning algorithm includes: Define the state space and action space. The state space includes historical performance data, the current network condition, and task complexity information. The action space includes decisions to allocate tasks to different computing nodes; Calculate the reward based on the effect of task execution. If the task is successfully completed within the predetermined time and the computing resources are used properly, a positive reward is given; if the task execution fails or the resources are used improperly, a negative reward is given.
8. The dynamic allocation method for AI inference tasks according to claim 1, characterized in that, The method further includes: During the data transmission process between different computing nodes, compress and encrypt the transmitted data across nodes.
9. A dynamic allocation device for AI inference tasks, characterized in that, The device includes: At least one processor; and At least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the device is caused to perform the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Edge computing task unloading optimization method and system
CN118567851A
Elastic configuration and two-way collaborative adaptation system for application computing task
CN119668849A
AI algorithm process and service all-in-one machine based on edge computing
CN119690673A
Cloud-side collaborative real-time task scheduling method for intelligent network connection automobile
CN119835335A
Resource allocation for tasks of unknown complexity
US20180107508A1
Cited By
Edge node collaborative query method and system based on cloud edge collaboration
CN120524043A
A cloud-edge collaborative query method and system for edge nodes
CN120524043B
High-performance parallel computing dynamic scheduling method and system based on eBPF
CN120540821A
An eBPF-based high-performance parallel computing dynamic scheduling method and system
CN120540821B
Multi-station resource scheduling decomposition method and device based on cloud edge collaboration
CN120750925A