Task request processing method and device, and storage medium
Patent Information
- Application Number
- CN202611139274.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-22
AI Technical Summary
[0005]本申请的主要目的在于提供一种任务请求的处理方法、装置及存储介质,以解决相关技术中处理任务请求的算力调度效率较低的问题
[0017]在本申请实施例中,通过接收任务请求,通过中心管控平台基于全局资源状态视图生成调度策略,并将调度策略发送至与调度策略对应的区域协同代理,其中,任务请求携带有算力需求信息与模型标识信息,全局资源状态视图用于表示每个边缘设备的资源状态信息,中心管控平台对应多个区域协同代理,每个区域协同代理对应多个边缘设备;通过区域协同代理接收调度策略,并基于调度策略确定目标边缘设备和调度指令,并将调度指令发送至目标边缘设备中的边缘代理;通过边缘代理接收调度指令,并依据调度指令在目标边缘设备中确定用于处理任务请求的图形处理器资源和模型快照,并基于图形处理器资源和模型快照处理任务请求,解决了现有技术中处理任务请求的算力调度效率较低的技术问题。
Smart Images

Figure CN122802595A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to a method, apparatus, and storage medium for processing task requests. Background Technology
[0002] Edge computing is a distributed open platform technology that integrates computing, storage, networking, and core application capabilities at the network edge, close to the data source or user side. With the explosive growth of applications such as the Internet of Things, the Industrial Internet, and smart security, massive amounts of heterogeneous edge devices (such as smart gateways, edge boxes, and graphics processing units) are being deployed widely.
[0003] However, existing edge computing scheduling technologies suffer from low efficiency in handling task requests, hindering the full potential of edge computing performance. Firstly, at the heterogeneous device access level, existing technologies lack a unified resource modeling mechanism. The computing power metrics of edge devices with different architectures and operating systems are difficult to standardize and abstract, leading to inaccurate assessment of the actual available computing power and often necessitating resource reservation strategies, resulting in idle computing power. Secondly, at the large-scale node management level, traditional solutions often employ a centralized scheduling model, where all task requests and status feedback must converge to a central node for processing. As the node scale reaches millions, control plane traffic congestion becomes severe, and scheduling decision latency is high, making it difficult to meet the real-time requirements of edge scenarios. Furthermore, in terms of computing power scheduling, existing mechanisms often use dedicated allocation of entire GPUs, and model switching requires service restarts or weight reloading, resulting in task switching latency of seconds or even tens of seconds, which cannot adapt to the multi-task, high-concurrency computing power demands of high-density edge inference scenarios.
[0004] There is currently no effective solution to the problem of low computing power scheduling efficiency in processing task requests in related technologies. Summary of the Invention
[0005] The main objective of this application is to provide a method, apparatus, and storage medium for processing task requests, so as to solve the problem of low computing power scheduling efficiency in processing task requests in related technologies.
[0006] To achieve the above objectives, according to one aspect of this application, a method for processing task requests is provided. The method includes: receiving a task request; generating a scheduling policy based on a global resource status view through a central management platform; and sending the scheduling policy to a regional collaborative agent corresponding to the scheduling policy. The task request carries computing power requirement information and model identification information. The global resource status view represents the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The method also includes: receiving the scheduling policy through the regional collaborative agents; determining a target edge device and a scheduling instruction based on the scheduling policy; and sending the scheduling instruction to an edge agent in the target edge device. Finally, the method includes: receiving the scheduling instruction through the edge agents; determining, according to the scheduling instruction, the graphics processing unit (GPU) resources and model snapshots used to process the task request in the target edge device; and processing the task request based on the GPU resources and model snapshots.
[0007] Furthermore, before receiving a task request, the method further includes: determining a capability descriptor, wherein the type of the capability descriptor includes device information, hardware parameters, software parameters, and interface specifications; mapping the hardware and software environments of multiple heterogeneous edge devices to capability descriptors based on plug-in adapters, and transmitting the capability descriptors to the edge agents corresponding to the edge devices; determining resource status information based on the capability descriptors through the edge agents, sending the resource status information to the corresponding regional collaboration agents, and sending the resource status information to the central management platform through the regional collaboration agents; and updating the global resource status view based on the resource status information through the central management platform.
[0008] Furthermore, before sending the resource status information to the central management platform through the regional collaborative agent, the method also includes: integrating the resource status information uploaded by different edge devices in the corresponding region according to the regional collaborative agent to obtain multiple resource status information; and using a conflict-free replication data type algorithm to perform data convergence processing on the multiple resource status information to obtain multiple processed resource status information.
[0009] Furthermore, the scheduling strategy generated by the central management platform based on the global resource status view includes: parsing the computing power requirement information in the task request through the central management platform to obtain the target computing power value and the target video memory value; selecting edge devices whose resource status information meets the target computing power value and the target video memory value from the global resource status view to obtain multiple candidate edge devices; calculating the adaptation weight of each candidate edge device based on the network connectivity, current load priority, and model identification information of each candidate edge device, and determining the target edge device from multiple candidate edge devices based on the adaptation weight; determining the graphics processor resource allocation ratio, task priority, and model context snapshot path information based on the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, and integrating the graphics processor resource allocation ratio, task priority, and model context snapshot path information to obtain the scheduling strategy.
[0010] Furthermore, determining the target edge device and scheduling instructions based on the scheduling strategy includes: receiving the scheduling strategy through the regional collaborative agent, determining the target edge device according to the scheduling strategy, and determining the regional range to which the target edge device belongs based on the pre-stored topology mapping relationship; if the regional range is located within the corresponding region of the regional collaborative agent, extracting the graphics processor resource allocation ratio, task priority, and model context snapshot path information from the scheduling strategy, and determining the scheduling instructions for the target edge device based on the graphics processor resource allocation ratio, task priority, and model context snapshot path information.
[0011] Furthermore, according to the scheduling instructions, determining the graphics processor resources and model snapshots for processing task requests in the target edge device includes: using the virtualization device plugin deployed in the target edge device to partition the graphics processor resources of the target edge device according to the graphics processor resource partitioning ratio in the scheduling instructions, and determining the graphics processor resources for processing task requests based on the partitioning results and task priorities; using the scheduler extension script in the target edge device to load the model snapshot from the local cache according to the model context snapshot path information in the scheduling instructions.
[0012] Furthermore, in the process of processing task requests based on graphics processor resources and model snapshots, the method also includes: if the network connection between the edge agent and the regional collaborative agent is interrupted, the first state information is stored according to the regional collaborative agent, the task request is continued to be processed according to the scheduling instructions cached locally by the target edge device, and the second state information is stored according to the processing result; after the network connection between the edge agent and the regional collaborative agent is restored, the second state information is sent to the regional collaborative agent through incremental synchronization according to the edge agent, the first state information and the second state information are processed by the regional collaborative agent using a conflict-free copy data type algorithm, and the data convergence processing result is sent to the central management platform.
[0013] To achieve the above objectives, according to another aspect of this application, a task request processing apparatus is provided. The apparatus includes: a receiving unit, configured to receive a task request, generate a scheduling policy based on a global resource status view through a central management platform, and send the scheduling policy to a regional collaborative agent corresponding to the scheduling policy; wherein the task request carries computing power requirement information and model identification information, the global resource status view represents the resource status information of each edge device, the central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices; a determining unit, configured to receive the scheduling policy through the regional collaborative agents, determine the target edge device and scheduling instructions based on the scheduling policy, and send the scheduling instructions to the edge agent in the target edge device; and a processing unit, configured to receive the scheduling instructions through the edge agents, determine the graphics processing unit resources and model snapshots for processing the task request in the target edge device according to the scheduling instructions, and process the task request based on the graphics processing unit resources and model snapshots.
[0014] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein, when the program is running, a method for processing any task request is provided to control the device where the computer-readable storage medium is located to perform such processing.
[0015] According to another aspect of this application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a processing method for performing any kind of task request.
[0016] According to another aspect of this application, a computer program product is provided, including computer instructions, which, when executed by a processor, implement the steps of a task request processing method for any of the above.
[0017] In this embodiment, by receiving a task request, a scheduling strategy is generated by the central management platform based on a global resource status view, and the scheduling strategy is sent to the regional collaborative agent corresponding to the scheduling strategy. The task request carries computing power requirement information and model identification information. The global resource status view represents the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The regional collaborative agent receives the scheduling strategy, determines the target edge device and scheduling instructions based on the scheduling strategy, and sends the scheduling instructions to the edge agent in the target edge device. The edge agent receives the scheduling instructions, determines the graphics processor resources and model snapshots for processing the task request in the target edge device according to the scheduling instructions, and processes the task request based on the graphics processor resources and model snapshots. This solves the technical problem of low computing power scheduling efficiency in processing task requests in the prior art.
[0018] The three-level collaborative control architecture generates scheduling strategies based on task requests through a central control platform, determines target edge devices and scheduling instructions through regional collaborative agents, and determines the graphics processor resources and model snapshots in the target edge devices to process task requests based on the scheduling instructions through edge agents in the edge devices. The task requests are then processed based on the graphics processor resources and model snapshots. This architecture achieves efficient computing power scheduling for task requests, avoids control congestion, and improves the efficiency of computing power scheduling. Attached Figure Description
[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 A hardware block diagram of a computer terminal for implementing a task request processing method is shown.
[0021] Figure 2 This is a flowchart of a task request processing method provided according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of a three-level collaborative control architecture in the task request processing method provided in the embodiments of this application;
[0023] Figure 4 This is a schematic diagram of fine-grained resource scheduling in the task request processing method provided in the embodiments of this application;
[0024] Figure 5 This is a schematic diagram of a task request processing apparatus provided according to an embodiment of this application;
[0025] Figure 6This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, do not violate public order and good morals, and provide corresponding access points for users to choose to authorize or refuse. For example, this system has interfaces with relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0029] Example 1
[0030] According to an embodiment of this application, a method embodiment for processing task requests is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a task request processing method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0032] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0033] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the task request processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned task request processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0034] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0035] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0036] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for handling task requests is shown. Figure 2 This is a flowchart of a task request processing method according to Embodiment 1 of this application.
[0037] Step S201: Receive task request, generate scheduling policy based on global resource status view through central management platform, and send scheduling policy to regional collaborative agent corresponding to scheduling policy.
[0038] Optionally, refer to Figure 3 As shown, the three-tier collaborative control architecture of this application embodiment includes a central management and control platform, regional collaborative agents, and edge agents. The central management and control platform is deployed in the cloud and can be a microservice or script program used to receive task requests from users. The central management and control platform is also used to read global policies through a preset configuration file, which may include device access policies, resource scheduling policies, and fault handling policies. The central management and control platform is also used to make resource scheduling decisions, allocating global computing resources and coordinating cross-regional business scheduling based on the status and task requirements of edge nodes in each region. The central management and control platform is also used to receive and display resource status information reported by regional collaborative agents in real time, such as displaying device online status, resource utilization, and business operation status, and issuing alarms when anomalies occur.
[0039] Optionally, the task request carries computing power requirement information and model identification information. The central management platform calculates and generates a scheduling policy based on the global resource status view and sends the scheduling policy to the corresponding regional collaborative agent. The computing power requirement information is the processing capacity requirement declared in the task request, which may include a target computing power value (e.g., in TOPS, i.e., trillions of operations per second) and a target video memory value (e.g., the number of GB required for model loading, i.e., gigabytes), used to filter edge devices with insufficient hardware resources. The model identification information is the unique identifier or version number of the artificial intelligence model to be executed, used to locate the model's context snapshot in the edge device for fast loading. The global resource status view represents the resource status information of each edge device and can be a logical view displaying the real-time resource status information of all edge devices, thus providing a global view for generating the scheduling policy. The central management platform corresponds to multiple regional collaborative agents, each regional collaborative agent corresponds to a region and manages multiple edge devices within that region. The scheduling policy may include the identification information of the target edge device and the target region to which the target edge device belongs. Each regional collaborative agent corresponds to a managed region and can send the scheduling policy to the regional collaborative agent corresponding to the target region in the scheduling policy.
[0040] Step S202: Receive the scheduling policy through the regional collaborative agent, determine the target edge device and scheduling instructions based on the scheduling policy, and send the scheduling instructions to the edge agent in the target edge device.
[0041] Optionally, refer to Figure 3 Regional collaborative agents are deployed on the edge gateways of each region. These agents can be lightweight processes, and each region can have one primary collaborative agent and one backup collaborative agent, enabling primary / backup failover and improving operational stability. The regional collaborative agent acts as an intermediary layer between the edge agents and the central management platform, aggregating regional status, forwarding the central management platform's scheduling policies, and synchronizing status information between edge devices and the central management platform, thus reducing the load on the central management platform.
[0042] Optionally, after the regional collaborative agent receives the scheduling policy, it can parse the identification information of the target edge device in the scheduling policy and determine whether the target edge device belongs to the management area of the regional collaborative agent. If so, it can extract the execution parameters for the target edge device from the scheduling policy to convert the scheduling policy into a command that the target edge device can understand and execute, thereby obtaining the scheduling instruction and sending it to the edge agent in the target edge device.
[0043] Step S203: Receive scheduling instructions through the edge agent, determine the graphics processor resources and model snapshots for processing task requests in the target edge device according to the scheduling instructions, and process the task requests based on the graphics processor resources and model snapshots.
[0044] Optionally, the edge agent is deployed on edge devices. The edge agent can be a lightweight process. A model snapshot refers to a memory image of the running state of an artificial intelligence model, including weights, biases, etc., used to quickly restore the model state and avoid the high latency caused by repeatedly loading the model. The edge agent can also be used to collect the local state of the edge device (e.g., CPU / GPU utilization, memory usage, network status, business running status, local cache information, etc.), execute scheduling instructions issued by the regional collaborative agent, perform autonomous management in the event of a network outage (e.g., independently execute local tasks based on locally cached policies and task data, record device status and task operation logs to ensure continuous task operation), and synchronize data through incremental synchronization after reconnection.
[0045] In summary, this three-level collaborative control architecture achieves efficient computing power scheduling for task requests by generating scheduling strategies based on task requests through a central control platform, determining target edge devices and scheduling instructions through regional collaborative agents, and identifying graphics processor resources and model snapshots for processing task requests in the target edge devices based on scheduling instructions through edge agents in the edge devices. The system then processes task requests based on the graphics processor resources and model snapshots.
[0046] To improve the efficiency of computing power scheduling, optionally, before receiving a task request, the method further includes: determining a capability descriptor, wherein the type of the capability descriptor includes device information, hardware parameters, software parameters, and interface specifications; mapping the hardware and software environments of multiple heterogeneous edge devices to capability descriptors based on plug-in adapters, and transmitting the capability descriptors to the edge agents corresponding to the edge devices; determining resource status information through the edge agents based on the capability descriptors, sending the resource status information to the corresponding regional collaborative agents, and sending the resource status information to the central management platform through the regional collaborative agents; and updating the global resource status view through the central management platform based on the resource status information.
[0047] Optionally, multiple heterogeneous edge devices may include GPU servers, smart gateways, edge boxes, etc. A capability descriptor is a standardized resource abstraction model used to uniformly represent the capabilities of edge devices with different architectures to eliminate heterogeneity, enabling the generation of scheduling policies to evaluate the available computing power of all edge devices using a unified standard. The capability descriptor can be in JSON (a lightweight data format) format, including four types of information: device information including static attributes such as device ID, name, deployment location, and installation time; hardware parameters including computing power (e.g., CPU computing power), memory capacity, storage capacity, heterogeneous acceleration unit type, and peripheral interface type (e.g., Ethernet); software parameters including operating system type, supported communication protocols, and pre-installed services; and interface specifications, i.e., a unified device access interface and data interaction format.
[0048] Optionally, plug-in adapters are lightweight software modules deployed on edge devices, responsible for protocol conversion and data standardization. This allows them to abstract the capabilities of existing devices into a standard format without modifying the original device firmware, reducing access costs. Plug-in adapters are developed for specific hardware architectures and operating system combinations; different combinations have different plug-in adapters. Deployed on edge devices, these adapters collect native hardware sensor data and software configuration information in real time and convert them into the aforementioned standard capability descriptor format.
[0049] Optionally, the edge proxy receives capability descriptors from the pluggable adapter and combines these descriptors with the local state of the edge device collected by the edge proxy to form resource status information. The edge proxy can then send this resource status information to the corresponding regional collaborative proxy in the region where the edge device belongs via heartbeat packets or incremental synchronization mechanisms. For example, the edge proxy can act as a resident process, periodically (perhaps every 5 seconds) collecting local capability descriptors and generating data packets containing resource status information. These data packets may include the edge device ID, timestamp, and updated parameters. The edge proxy sends these data packets to the corresponding regional collaborative proxy via a long connection. If the edge device goes offline and then comes back online, the edge proxy only sends the parameters that changed during the offline period to the regional collaborative proxy via incremental synchronization, reducing bandwidth consumption. After the central management platform receives the resource status information from the regional collaborative proxy, it can update the global resource status view based on this information.
[0050] In summary, by abstracting the resources of heterogeneous edge devices into unified and standardized data through capability descriptors, and then using regional edge agents as an intermediate layer to report resource status information, the global resource status view is updated, laying a data foundation for efficient computing power scheduling. At the same time, the edge agent only needs to communicate with the regional collaborative agent, and the regional collaborative agent then communicates with the central management platform. This hierarchical reporting mechanism avoids network storms and bandwidth congestion caused by large-scale edge devices directly connecting to the center, thus improving the efficiency of computing power scheduling.
[0051] To improve the efficiency of computing power scheduling, the method may optionally include the following steps before sending resource status information to the central management platform through the regional collaborative agent: integrating resource status information uploaded by different edge devices in the corresponding region according to the regional collaborative agent to obtain multiple resource status information; and using a conflict-free copy data type algorithm to perform data convergence processing on the multiple resource status information to obtain multiple processed resource status information.
[0052] Optionally, to ensure data consistency, a data convergence process is performed before the regional collaborative agent sends resource status information to the central management platform. The regional collaborative agent collects resource status information from all edge devices within its region by listening to message queues or receiving reporting requests from various edge agents. Each resource status information packet may include the edge device ID, collection timestamp, and capability descriptor. An improved CRDT algorithm (i.e., a conflict-free replication data type algorithm) can be used to perform data convergence processing on the aforementioned multiple original resource status information packets. This process transforms potentially conflicting or duplicated resource status information into unique and accurate resource status information, eliminating data redundancy and conflicts, and providing high-quality resource status information to the central management platform. For example, each resource status information packet can carry a globally monotonically increasing timestamp. When the regional collaborative agent receives multiple resource status information packets with the same edge device ID, it compares the corresponding timestamps and retains the resource status information packet with the largest timestamp, ignoring other older or conflicting versions. This converges multiple conflicting or redundant resource status information packets, which may be caused by network latency or retransmission, into a unique and up-to-date state, resulting in processed resource status information.
[0053] In summary, by integrating resource status information uploaded by different edge devices within the corresponding region based on the regional collaborative agent, multiple resource status information is obtained. A conflict-free copy data type algorithm is used to perform data convergence processing on the multiple resource status information to obtain processed multiple resource status information. This ensures the accuracy of the global resource status view, provides a foundation for efficient computing power scheduling, and improves the efficiency of computing power scheduling.
[0054] To improve the efficiency of computing power scheduling, optionally, the scheduling strategy generated by the central management platform based on the global resource status view includes: parsing the computing power demand information in the task request through the central management platform to obtain the target computing power value and the target video memory value; filtering edge devices whose resource status information meets the target computing power value and the target video memory value from the global resource status view to obtain multiple candidate edge devices; calculating the adaptation weight of each candidate edge device based on the network connectivity, current load priority, and model identification information of each candidate edge device, and determining the target edge device from the multiple candidate edge devices based on the adaptation weight; determining the graphics processor resource allocation ratio, task priority, and model context snapshot path information based on the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, and integrating the graphics processor resource allocation ratio, task priority, and model context snapshot path information to obtain the scheduling strategy.
[0055] Optionally, the task request can be in JSON format. The target computing power (in TOPS) and target GPU memory (in GB) are obtained by parsing the task request. For example, if the task request is to run an object detection model for real-time video analysis, the parsed target computing power might be 4.0 TOPS and the target GPU memory might be 2.0 GB. The central management platform iterates through all edge devices in the global resource status view to check if their available computing power and available GPU memory reach the target values. Devices that meet both conditions are marked as candidate edge devices. Available GPU memory can be determined by subtracting the GPU memory utilization rate in the edge device's local status collected by the edge agent from the maximum GPU memory capacity in the capability descriptor. Available computing power is determined by the computing power in the capability descriptor (as the theoretical maximum computing power) and the GPU utilization rate in the edge device's local status collected by the edge agent in the resource status information. Edge devices that meet both conditions can be considered candidate edge devices.
[0056] Optionally, the local status of the edge device collected by the edge agent in the resource status information may also include network connectivity, current load priority, and model identification information. The adaptation weight is obtained by calculating a weighted sum based on the candidate edge device's network connectivity (A), current load priority (B), and the hit result of the edge device's local cache on the model identification information (C). For example, network connectivity can be based on a score of network latency; the higher the network latency, the lower the score of A. The current load priority reflects load information; the higher the load, the lower the value of the current load priority B (i.e., the worse the stability). If the edge device's local cache already contains a snapshot of the model context corresponding to the model identification information, i.e., the model identification information has been hit, then C is 1; otherwise, C is 0. The adaptation weight comprehensively considers network quality, system stability, and the hit result of the local cache on the model identification information, and can determine the edge device with the highest adaptation weight as the target edge device.
[0057] Optionally, after identifying the target edge device, the central management platform can query the current remaining resource information of the target edge device (such as the number of remaining GPU computing units and the number of remaining video memory GB) based on the global resource status view. Combined with the priority of the task requests (high priority, medium priority, low priority), the graphics processor resource allocation ratio can be determined (for example, a high-priority task can exclusively occupy a logical GPU block, while a low-priority task shares 0.1 TOPS of computing power). Different priority task requests have different graphics processor resource allocation ratios; the higher the task priority, the higher the proportion of graphics processor resources it occupies. Simultaneously, based on the model identification information, the storage path of the model snapshot is found in the global resource status view to obtain the model context snapshot path information. This information is then integrated into a scheduling strategy in JSON format.
[0058] In summary, by analyzing the computing power requirement information in task requests through the central management platform, the target computing power value and target video memory value are obtained. Edge devices whose resource status information meets the target computing power value and target video memory value are selected from the global resource status view, resulting in multiple candidate edge devices. Based on the network connectivity, current load priority, and model identification information of each candidate edge device, the adaptation weight of each candidate edge device is calculated, and the target edge device is determined from the multiple candidate edge devices based on the adaptation weight. According to the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, the graphics processor resource allocation ratio, task priority, and model context snapshot path information are determined. Finally, the graphics processor resource allocation ratio, task priority, and model context snapshot path information are integrated to obtain the scheduling strategy, thus improving the efficiency of computing power scheduling.
[0059] To improve the efficiency of computing power scheduling, optionally, determining the target edge device and scheduling instructions based on the scheduling policy includes: receiving the scheduling policy through a regional collaborative agent, determining the target edge device according to the scheduling policy, and determining the regional range to which the target edge device belongs based on the pre-stored topology mapping relationship; if the regional range is located within the corresponding region of the regional collaborative agent, extracting the graphics processor resource allocation ratio, task priority, and model context snapshot path information from the scheduling policy, and determining the scheduling instructions for the target edge device based on the graphics processor resource allocation ratio, task priority, and model context snapshot path information.
[0060] Optionally, after receiving the scheduling policy from the central management platform, the regional collaborative agent receives the target edge device identification information (e.g., edge device ID). The regional collaborative agent determines whether the target edge device corresponding to the target edge device identification information belongs to the managed region by querying the pre-stored topology mapping relationship maintained locally. The pre-stored topology mapping relationship includes the identification information of each edge device and its corresponding region. If the target edge device is located within the corresponding region of the regional collaborative agent, execution parameters are extracted from the scheduling policy. Execution parameters may include the graphics processor resource allocation ratio (e.g., {"gpu_split_ratio":"1 / 8","memory_split_gb":2}, indicating that the GPU's computing cores are divided into 8 equal parts, with the current task requesting one part, requiring 2GB of VRAM), task priority (e.g., high priority HIGH), and model context snapshot path information (e.g., / gpu_cache / model_v1_snapshot.bin). Encapsulating these execution parameters into JSON format data yields the scheduling instruction.
[0061] In summary, by receiving scheduling policies through a regional collaborative agent, determining the target edge device based on the scheduling policies, and identifying the regional range to which the target edge device belongs based on the pre-stored topology mapping relationship, if the regional range is located within the corresponding region of the regional collaborative agent, the GPU resource allocation ratio, task priority, and model context snapshot path information are extracted from the scheduling policy. Based on the GPU resource allocation ratio, task priority, and model context snapshot path information, scheduling instructions for the target edge device are determined, thus avoiding bandwidth congestion and latency issues caused by centralized scheduling and improving the efficiency of computing power scheduling.
[0062] To improve the efficiency of computing power scheduling, optionally, determining the graphics processor resources and model snapshots for processing task requests in the target edge device according to the scheduling instructions includes: using the virtualization device plugin deployed in the target edge device to partition the graphics processor resources of the target edge device according to the graphics processor resource partitioning ratio in the scheduling instructions, and determining the graphics processor resources for processing task requests based on the partitioning results and task priorities; using the scheduler extension script in the target edge device to load the model snapshot from the local cache according to the model context snapshot path information in the scheduling instructions.
[0063] Optionally, refer to Figure 4 As shown, a virtualization device plugin can be deployed in the target edge device. This plugin, developed based on containerization technology, parses scheduling instructions through an edge agent and extracts the graphics processor (GPU) resource allocation ratio. Based on this ratio, the virtualization device plugin logically allocates the physical GPU resources of the target edge device and ultimately determines the GPU resources allocated to the current task according to task priority. For example, the virtualization device plugin reads the GPU resource allocation ratio (1 / 8 card or a specific computing power value of 0.125 TOPS), calls the underlying program interface, and divides the physical GPU's computing power units into multiple logical computing power units and the video memory into multiple logical video memory units. The virtualization device plugin queries the task priority; if it is a high-priority task, it prioritizes allocating the maximum available resources from idle logical computing power units; if it is a medium-priority or low-priority task, it allocates resources from the remaining resources, following the order of high priority, medium priority, and low priority. Higher-priority tasks have the right to preempt resources from lower-priority tasks. By binding allocated logical computing units and logical memory units into an independent container or process space using virtualization device plugins, graphics processing unit resources are obtained for processing task requests. This fine-grained allocation of GPU resources through virtualization technology enables super-resolution multiplexing of GPU resources. Furthermore, the virtualization device plugins can monitor GPU load; if the load exceeds a threshold of 85%, the allocation ratio is automatically adjusted to avoid overload, thus ensuring fine-grained allocation of graphics processing unit resources and maximizing hardware utilization.
[0064] Optionally, a scheduler extension script runs on the target edge device. The scheduling instructions are parsed by the edge agent, and the model context snapshot path information is extracted, referencing... Figure 4Based on this path information, the scheduler extension script loads model snapshots (e.g., snapshots of models A, B, and C) from the local cache to prepare the task execution environment. For example, the scheduler extension script can be a lightweight Python script that reads the model context snapshot path information from the scheduling instructions and calls the GPU inference engine's loading interface to map the binary snapshot data under that path to the GPU memory. This efficiently restores the running state of the AI model, enabling hot-switching of the model and reducing model loading time to milliseconds. During task execution, the scheduler extension script periodically saves snapshots of the inference model's running context and dynamically allocates tasks based on task requirements and GPU resource status, thereby achieving load balancing and improved computing power utilization. When a model switch is needed, the scheduler extension script directly calls the snapshot in the local cache to quickly load the model context without restarting the service or reloading model weights, further improving the response speed and efficiency of computing power scheduling.
[0065] In summary, by utilizing the virtualization device plugin deployed in the target edge device, the graphics processor resources of the target edge device are allocated according to the graphics processor resource allocation ratio in the scheduling instruction, and the graphics processor resources used to process task requests are determined based on the allocation results and task priorities; by utilizing the scheduler extension script in the target edge device, model snapshots are loaded from the local cache according to the model context snapshot path information in the scheduling instruction, the overall efficiency of computing power scheduling is improved.
[0066] To improve the efficiency of computing power scheduling, optionally, in the process of processing task requests based on graphics processor resources and model snapshots, the method further includes: if the network connection between the edge agent and the regional collaborative agent is interrupted, the first state information is stored according to the regional collaborative agent, the task request is continued to be processed according to the scheduling instructions cached locally by the target edge device, and the second state information is stored according to the processing result; after the network connection between the edge agent and the regional collaborative agent is restored, the second state information is sent to the regional collaborative agent through incremental synchronization according to the edge agent, the first state information and the second state information are processed by the regional collaborative agent using a conflict-free copy data type algorithm, and the data convergence processing result is sent to the central management platform.
[0067] Optionally, a heartbeat mechanism can be used to detect the network connection between the edge agent and the regional collaborative agent. If the edge agent does not receive a heartbeat response from the regional collaborative agent three consecutive times, it is determined that the network connection between the two is interrupted. At this time, the edge agent switches to local autonomous mode and continues to execute the issued scheduling instructions based on the first status information sent locally by the regional collaborative agent before the network disconnection. The first status information may include the task corresponding to the current task request, the proportion of allocated GPU resources, model context snapshot path information, etc. During execution, the edge agent continuously monitors the processing status of the task corresponding to the task request, such as GPU utilization, task completion progress, error logs, etc., and can record the task processing status as second status information and store it in the local persistent storage of the target edge device.
[0068] Optionally, if the edge agent detects the restoration of network connectivity with the regional collaborative agent via heartbeat response, it reads the second state information through the edge agent and compares it with the first state information to generate an incremental data list. The incremental data list only includes fields that changed during the network outage (e.g., a subtask in the task corresponding to the task request has been completed, and the GPU load has changed from a to b). This incremental data list is then sent to the regional collaborative agent to avoid congestion caused by a large amount of data upon network recovery. The regional collaborative agent can use an improved CRDT algorithm based on the first state information before the network outage to merge the second state information with the first state information. For example, the second state information with the newer timestamp can be used to overwrite older values in the first state information, eliminating data conflicts and obtaining the data convergence processing result. The data convergence processing result is then uploaded to the central management platform through the regional collaborative agent to ensure that the central management platform's scheduling strategy decisions are based on the latest data.
[0069] In summary, if the network connection between the edge agent and the regional collaborative agent is interrupted, the regional collaborative agent stores the first state information, continues to process task requests based on the scheduling instructions cached locally on the target edge device, and stores the second state information based on the processing result. After the network connection between the edge agent and the regional collaborative agent is restored, the edge agent sends the second state information to the regional collaborative agent through incremental synchronization. The regional collaborative agent uses a conflict-free replication data type algorithm to perform data convergence processing on the first and second state information and sends the data convergence processing result to the central management platform. This avoids data transmission timeouts and scheduling failures caused by conflicting data due to network instability, thus improving the efficiency and accuracy of computing power scheduling.
[0070] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0071] Example 2
[0072] This application also provides a task request processing apparatus. It should be noted that the task request processing apparatus of this application can be used to execute the task request processing method provided in this application. The task request processing apparatus provided in this application will be described below.
[0073] According to embodiments of this application, an apparatus for implementing the above-described task request processing method is also provided, such as... Figure 5 As shown, the device includes:
[0074] The receiving unit 501 is used to receive task requests, generate scheduling strategies based on the global resource status view through the central management platform, and send the scheduling strategies to the regional collaborative agents corresponding to the scheduling strategies. The task requests carry computing power requirement information and model identification information. The global resource status view is used to represent the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices.
[0075] The determining unit 502 is used to receive the scheduling policy through the regional collaborative agent, determine the target edge device and the scheduling instruction based on the scheduling policy, and send the scheduling instruction to the edge agent in the target edge device.
[0076] The processing unit 503 is configured to receive scheduling instructions through the edge agent, determine the graphics processor resources and model snapshots for processing the task request in the target edge device according to the scheduling instructions, and process the task request based on the graphics processor resources and model snapshots.
[0077] The task request processing apparatus provided in this application embodiment receives task requests through a receiving unit 501, generates a scheduling strategy based on a global resource status view through a central management platform, and sends the scheduling strategy to the regional collaborative agent corresponding to the scheduling strategy. The task request carries computing power requirement information and model identification information. The global resource status view represents the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. A determining unit 502 receives the scheduling strategy through the regional collaborative agents, determines the target edge device and scheduling instructions based on the scheduling strategy, and sends the scheduling instructions to the edge agent in the target edge device. A processing unit 503 receives the scheduling instructions through the edge agents, determines the graphics processor resources and model snapshots for processing the task request in the target edge device according to the scheduling instructions, and processes the task request based on the graphics processor resources and model snapshots. This solves the problem of low computing power scheduling efficiency in processing task requests in related technologies, thereby improving the computing power scheduling efficiency for processing task requests.
[0078] Optionally, in the task request processing apparatus provided in this application embodiment, the apparatus further includes: a descriptor determination unit, configured to determine a capability descriptor before receiving a task request, wherein the type of the capability descriptor includes device information, hardware parameters, software parameters, and interface specifications; a mapping unit, configured to map the hardware and software environments of multiple heterogeneous edge devices into capability descriptors according to a plug-in adapter, and transmit the capability descriptors to the edge agent corresponding to the edge device; a sending unit, configured to determine resource status information based on the capability descriptor through the edge agent, send the resource status information to the corresponding regional collaborative agent, and send the resource status information to the central management platform through the regional collaborative agent; and an update unit, configured to update the global resource status view based on the resource status information through the central management platform.
[0079] Optionally, in the task request processing apparatus provided in this application embodiment, the apparatus further includes: an integration unit, configured to integrate resource status information uploaded by different edge devices in the corresponding region according to the regional collaborative agent before sending resource status information to the central management platform through the regional collaborative agent, to obtain multiple resource status information; and a convergence processing unit, configured to perform data convergence processing on the multiple resource status information using a conflict-free replication data type algorithm, to obtain multiple processed resource status information.
[0080] Optionally, in the task request processing apparatus provided in this application embodiment, the receiving unit 501 includes: a parsing module, used to parse the computing power requirement information in the task request through a central management platform to obtain the target computing power value and the target video memory value; a filtering module, used to filter edge devices whose resource status information meets the target computing power value and the target video memory value from the global resource status view to obtain multiple candidate edge devices; a calculation module, used to calculate the adaptation weight of each candidate edge device based on the network connectivity, current load priority, and model identification information of each candidate edge device, and determine the target edge device from multiple candidate edge devices based on the adaptation weight; and a strategy determination module, used to determine the graphics processor resource allocation ratio, task priority, and model context snapshot path information based on the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, and integrate the graphics processor resource allocation ratio, task priority, and model context snapshot path information to obtain a scheduling strategy.
[0081] Optionally, in the task request processing apparatus provided in this application embodiment, the determining unit 502 includes: a device determining module, configured to receive a scheduling policy through a regional collaborative agent, determine a target edge device according to the scheduling policy, and determine the regional range to which the target edge device belongs according to a pre-stored topology mapping relationship; and an instruction determining module, configured to extract the graphics processor resource allocation ratio, task priority, and model context snapshot path information from the scheduling policy if the regional range is located within the corresponding region of the regional collaborative agent, and determine a scheduling instruction for the target edge device according to the graphics processor resource allocation ratio, task priority, and model context snapshot path information.
[0082] Optionally, in the task request processing apparatus provided in this application embodiment, the processing unit 501 includes: a segmentation module, used to segment the graphics processor resources of the target edge device according to the graphics processor resource segmentation ratio in the scheduling instruction by utilizing the virtualization device plugin deployed in the target edge device, and to determine the graphics processor resources used to process the task request according to the segmentation result and task priority; and a loading module, used to load the model snapshot from the local cache according to the model context snapshot path information in the scheduling instruction by utilizing the scheduler extension script in the target edge device.
[0083] Optionally, in the task request processing apparatus provided in this application embodiment, the apparatus further includes: an edge autonomous unit, configured to, during the process of processing task requests based on graphics processor resources and model snapshots, if the network connection between the edge agent and the regional collaborative agent is interrupted, store first state information according to the regional collaborative agent, continue processing the task request according to the scheduling instructions cached locally by the target edge device, and store second state information according to the processing result; and an information synchronization unit, configured to, after the network connection between the edge agent and the regional collaborative agent is restored, send the second state information to the regional collaborative agent through incremental synchronization according to the edge agent, perform data convergence processing on the first state information and the second state information through the conflict-free copy data type algorithm of the regional collaborative agent, and send the data convergence processing result to the central management platform.
[0084] It should be noted that the receiving unit 501, determining unit 502, and processing unit 503 mentioned above correspond to steps S201 to S203 in Embodiment 1. The instances and application scenarios implemented by the units and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules can also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.
[0085] Example 3
[0086] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) Processor 602, memory 604, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0087] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0088] The processor can access information and applications stored in memory via a transmission device to execute the following steps: receiving a task request; generating a scheduling policy based on a global resource status view through a central management platform; and sending the scheduling policy to the corresponding regional collaborative agent. The task request carries computing power requirement information and model identification information. The global resource status view represents the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The processor receives the scheduling policy through the regional collaborative agent, determines the target edge device and scheduling instructions based on the policy, and sends the scheduling instructions to the edge agent in the target edge device. The edge agent receives the scheduling instructions, determines the graphics processing unit (GPU) resources and model snapshots for processing the task request in the target edge device based on the instructions, and processes the task request based on the GPU resources and model snapshots.
[0089] The processor can also invoke information and applications stored in memory via a transmission device to perform the following steps: determining capability descriptors, wherein the types of capability descriptors include device information, hardware parameters, software parameters, and interface specifications; mapping the hardware and software environments of multiple heterogeneous edge devices to capability descriptors based on plug-in adapters, and transmitting the capability descriptors to the edge agents corresponding to the edge devices; determining resource status information based on the capability descriptors through the edge agents, sending the resource status information to the corresponding regional collaboration agents, and sending the resource status information to the central management platform through the regional collaboration agents; and updating the global resource status view based on the resource status information through the central management platform.
[0090] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: integrate the resource status information uploaded by different edge devices in the corresponding region according to the regional collaborative agent to obtain multiple resource status information; use a conflict-free copy data type algorithm to perform data convergence processing on the multiple resource status information to obtain multiple processed resource status information.
[0091] The processor can also access information and applications stored in memory via a transmission device to perform the following steps: parse the computing power requirement information in the task request through the central management platform to obtain the target computing power value and target video memory value; filter edge devices whose resource status information meets the target computing power value and target video memory value from the global resource status view to obtain multiple candidate edge devices; calculate the adaptation weight of each candidate edge device based on its network connectivity, current load priority, and model identification information, and determine the target edge device from the multiple candidate edge devices based on the adaptation weight; determine the graphics processor resource allocation ratio, task priority, and model context snapshot path information based on the target edge device's current remaining resource information, task request priority, and model identification information, and integrate the graphics processor resource allocation ratio, task priority, and model context snapshot path information to obtain a scheduling strategy.
[0092] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: receive the scheduling policy through the regional cooperation agent, determine the target edge device according to the scheduling policy, and determine the regional range to which the target edge device belongs according to the pre-stored topology mapping relationship; if the regional range is located in the corresponding region of the regional cooperation agent, extract the graphics processor resource allocation ratio, task priority and model context snapshot path information from the scheduling policy, and determine the scheduling instruction for the target edge device according to the graphics processor resource allocation ratio, task priority and model context snapshot path information.
[0093] The processor can also access information and applications stored in memory via a transmission device to perform the following steps: using the virtualization device plugin deployed in the target edge device, the graphics processor resources of the target edge device are allocated according to the graphics processor resource allocation ratio in the scheduling instruction, and the graphics processor resources used to process task requests are determined according to the allocation results and task priorities; using the scheduler extension script in the target edge device, the model snapshot is loaded from the local cache according to the model context snapshot path information in the scheduling instruction.
[0094] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: If the network connection between the edge agent and the regional collaborative agent is interrupted, the first state information is stored according to the regional collaborative agent, the task request is continued to be processed according to the scheduling instructions cached locally by the target edge device, and the second state information is stored according to the processing result; after the network connection between the edge agent and the regional collaborative agent is restored, the second state information is sent to the regional collaborative agent through incremental synchronization according to the edge agent, and the first and second state information are processed by the regional collaborative agent using a conflict-free copy data type algorithm to perform data convergence processing on the first and second state information, and the data convergence processing result is sent to the central management platform.
[0095] This application provides a solution for processing task requests. By receiving a task request, a central management platform generates a scheduling policy based on a global resource status view and sends the policy to the corresponding regional collaborative agent. The task request carries computing power requirement information and model identification information. The global resource status view represents the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The regional collaborative agent receives the scheduling policy, determines the target edge device and scheduling instructions based on the policy, and sends the scheduling instructions to the edge agent in the target edge device. The edge agent receives the scheduling instructions and, based on the instructions, determines the graphics processing unit (GPU) resources and model snapshots for processing the task request in the target edge device. The task request is then processed based on the GPU resources and model snapshots. This solution addresses the technical problem of low computing power scheduling efficiency in processing task requests in existing technologies.
[0096] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.
[0097] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0098] Example 4
[0099] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the task request processing method provided in Embodiment 1.
[0100] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0101] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: receiving a task request, generating a scheduling policy based on a global resource status view through a central management platform, and sending the scheduling policy to the regional collaborative agent corresponding to the scheduling policy, wherein the task request carries computing power requirement information and model identification information, the global resource status view is used to represent the resource status information of each edge device, the central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices; receiving the scheduling policy through the regional collaborative agent, determining the target edge device and scheduling instructions based on the scheduling policy, and sending the scheduling instructions to the edge agent in the target edge device; receiving the scheduling instructions through the edge agent, determining the graphics processor resources and model snapshots for processing the task request in the target edge device according to the scheduling instructions, and processing the task request based on the graphics processor resources and model snapshots.
[0102] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: determining a capability descriptor, wherein the type of the capability descriptor includes device information, hardware parameters, software parameters, and interface specifications; mapping the hardware and software environments of multiple heterogeneous edge devices to capability descriptors according to the plug-in adapter, and transmitting the capability descriptors to the edge agents corresponding to the edge devices; determining resource status information based on the capability descriptors through the edge agents, sending the resource status information to the corresponding regional collaboration agents, and sending the resource status information to the central management platform through the regional collaboration agents; and updating the global resource status view based on the resource status information through the central management platform.
[0103] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: integrating resource status information uploaded by different edge devices within the corresponding region according to the regional collaborative agent to obtain multiple resource status information; and using a conflict-free copy data type algorithm to perform data convergence processing on the multiple resource status information to obtain processed multiple resource status information.
[0104] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: parsing the computing power requirement information in the task request through the central management platform to obtain the target computing power value and the target video memory value; filtering edge devices whose resource status information meets the target computing power value and the target video memory value from the global resource status view to obtain multiple candidate edge devices; calculating the adaptation weight of each candidate edge device based on the network connectivity, current load priority, and model identification information of each candidate edge device, and determining the target edge device from multiple candidate edge devices based on the adaptation weight; determining the graphics processor resource allocation ratio, task priority, and model context snapshot path information based on the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, and integrating the graphics processor resource allocation ratio, task priority, and model context snapshot path information to obtain a scheduling strategy.
[0105] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: receiving a scheduling policy through a regional collaborative agent, determining a target edge device based on the scheduling policy, and determining the regional range to which the target edge device belongs based on a pre-stored topology mapping relationship; if the regional range is located within the corresponding region of the regional collaborative agent, extracting the graphics processor resource allocation ratio, task priority, and model context snapshot path information from the scheduling policy, and determining a scheduling instruction for the target edge device based on the graphics processor resource allocation ratio, task priority, and model context snapshot path information.
[0106] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: using a virtualization device plugin deployed in the target edge device, dividing the graphics processor resources of the target edge device according to the graphics processor resource allocation ratio in the scheduling instruction, and determining the graphics processor resources used to process task requests based on the allocation result and task priority; using a scheduler extension script in the target edge device, loading a model snapshot from the local cache according to the model context snapshot path information in the scheduling instruction.
[0107] Optionally, in this embodiment, the computer-readable storage medium is further configured to store program code for performing the following steps: if the network connection between the edge agent and the regional collaborative agent is interrupted, the first state information is stored according to the regional collaborative agent, the task request is continued to be processed according to the scheduling instructions cached locally by the target edge device, and the second state information is stored according to the processing result; after the network connection between the edge agent and the regional collaborative agent is restored, the second state information is sent to the regional collaborative agent by the edge agent through incremental synchronization, the first state information and the second state information are processed by the regional collaborative agent using a conflict-free copy data type algorithm, and the data convergence processing result is sent to the central management platform.
[0108] This application also provides a computer program product, which, when executed on a data processing device, is adapted to perform the processing method steps of a task request.
[0109] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0110] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0111] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0112] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0113] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0114] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0115] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for processing task requests, characterized in that, include: Upon receiving a task request, a scheduling strategy is generated based on a global resource status view through a central management platform, and the scheduling strategy is sent to the regional collaborative agent corresponding to the scheduling strategy. The task request carries computing power requirement information and model identification information. The global resource status view is used to represent the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The scheduling policy is received through the regional collaborative agent, and the target edge device and scheduling instructions are determined based on the scheduling policy. The scheduling instructions are then sent to the edge agent in the target edge device. The edge agent receives the scheduling instruction and determines the graphics processor resources and model snapshots in the target edge device to process the task request based on the scheduling instruction, and processes the task request based on the graphics processor resources and model snapshots.
2. The method according to claim 1, characterized in that, Before receiving a task request, the method further includes: Determine the capability descriptor, wherein the type of the capability descriptor includes device information, hardware parameters, software parameters, and interface specifications; Based on the plug-in adapter, the hardware and software environments of multiple heterogeneous edge devices are mapped to the capability descriptors, and the capability descriptors are transmitted to the edge agent corresponding to the edge device; The edge agent determines the resource status information based on the capability descriptor, sends the resource status information to the corresponding regional collaboration agent, and then sends the resource status information to the central management platform through the regional collaboration agent. The central management platform updates the global resource status view based on the resource status information.
3. The method according to claim 2, characterized in that, Before sending the resource status information to the central management platform through the regional collaborative agent, the method further includes: Based on the regional collaborative agent, the resource status information uploaded by different edge devices in the corresponding region is integrated to obtain multiple resource status information; A conflict-free copy data type algorithm is used to perform data convergence processing on the multiple resource status information to obtain the processed multiple resource status information.
4. The method according to claim 1, characterized in that, The scheduling strategy generated by the central management platform based on the global resource status view includes: The central control platform parses the computing power requirement information in the task request to obtain the target computing power value and the target video memory value. From the global resource status view, edge devices whose resource status information meets the target computing power value and the target memory value are selected to obtain multiple candidate edge devices; Based on the network connectivity, current load priority, and model identification information of each candidate edge device, the adaptation weight of each candidate edge device is calculated, and the target edge device is determined from the plurality of candidate edge devices based on the adaptation weight; Based on the current remaining resource information of the target edge device, the priority of the task request, and the model identification information, the graphics processor resource allocation ratio, task priority, and model context snapshot path information are determined, and the scheduling strategy is obtained by integrating the graphics processor resource allocation ratio, the task priority, and the model context snapshot path information.
5. The method according to claim 1, characterized in that, Determining the target edge device and scheduling instructions based on the aforementioned scheduling strategy includes: The scheduling policy is received through the regional collaborative agent, the target edge device is determined according to the scheduling policy, and the regional range to which the target edge device belongs is determined according to the pre-stored topology mapping relationship. If the area is located within the corresponding area of the regional collaborative agent, the graphics processor resource allocation ratio, task priority, and model context snapshot path information are extracted from the scheduling strategy, and a scheduling instruction for the target edge device is determined based on the graphics processor resource allocation ratio, the task priority, and the model context snapshot path information.
6. The method according to claim 5, characterized in that, The graphics processor resources and model snapshots determined in the target edge device according to the scheduling instructions for processing the task request include: Using the virtualization device plugin deployed in the target edge device, the graphics processor resources of the target edge device are divided according to the graphics processor resource allocation ratio in the scheduling instruction, and the graphics processor resources used to process the task request are determined according to the allocation result and the task priority. The scheduler extension script in the target edge device is used to load the model snapshot from the local cache according to the model context snapshot path information in the scheduling instruction.
7. The method according to claim 1, characterized in that, In processing the task request based on the graphics processor resources and model snapshot, the method further includes: If the network connection between the edge agent and the regional collaborative agent is interrupted, the first state information is stored according to the regional collaborative agent, the task request is processed according to the scheduling instruction cached locally by the target edge device, and the second state information is stored according to the processing result. After the network connection between the edge agent and the regional collaborative agent is restored, the edge agent sends the second status information to the regional collaborative agent through incremental synchronization. The regional collaborative agent then uses a conflict-free replication data type algorithm to perform data convergence processing on the first and second status information, and sends the data convergence processing result to the central management platform.
8. A task request processing apparatus, characterized in that, include: The receiving unit is used to receive task requests, generate scheduling strategies based on the global resource status view through the central management platform, and send the scheduling strategies to the regional collaborative agents corresponding to the scheduling strategies. The task requests carry computing power requirement information and model identification information. The global resource status view is used to represent the resource status information of each edge device. The central management platform corresponds to multiple regional collaborative agents, and each regional collaborative agent corresponds to multiple edge devices. The determining unit is configured to receive the scheduling policy through the regional collaborative agent, determine the target edge device and scheduling instructions based on the scheduling policy, and send the scheduling instructions to the edge agent in the target edge device; The processing unit is configured to receive the scheduling instruction through the edge agent, determine the graphics processor resources and model snapshots for processing the task request in the target edge device according to the scheduling instruction, and process the task request based on the graphics processor resources and model snapshots.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the task request processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the task request processing method according to any one of claims 1 to 7.
11. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the task request processing method according to any one of claims 1 to 7.