Processing method and device of reasoning task and electronic equipment
By using the target algorithm and bidirectional graph link structure in wireless edge networks to optimize resource allocation and model splitting points, the delay and energy consumption problems of multi-user DNN inference tasks are solved, and efficient resource management and task processing are achieved.
Patent Information
- Application Number
- CN202510360106.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
In the wireless edge network environment, multi-user DNN-based inference tasks have problems such as increasing latency and increasing energy consumption in dealing with model splitting and resource competition among users.
In each decision cycle, the first server receives the status information of the user equipment, combines the target algorithm and the two-way graph link structure, determines the model splitting points and resource allocation strategies of each user equipment, optimizes resource allocation and model splitting point selection, and realizes intelligent resource scheduling and model splitting.
It effectively reduces inference delay and energy consumption, improves resource utilization efficiency, and ensures efficient dynamic management of DNN inference tasks in a multi-user environment.
Smart Images

Figure CN120335949A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of wireless edge computing, and in particular, to a method, an apparatus, and an electronic device for processing inference tasks. Background Art
[0002] In a wireless edge network environment, the application of deep neural network (DNN) technology is becoming increasingly widespread, especially showing great potential in AI tasks such as image recognition and natural language processing. However, with the rapid growth of the scale and computational complexity of DNN models, as well as the limitations of the computing resources of edge devices, it becomes impractical to locally execute DNN inference tasks. For this reason, leveraging the rich computing resources of edge servers, collaborative offloading of DNN inference has become a typical solution. However, when facing the scenario of multi-user sharing resources, how to effectively perform model splitting, resource scheduling, and interference coordination among users to simultaneously reduce the inference latency and energy consumption has become a key problem to be solved urgently.
[0003] Currently, wireless edge DNN collaborative inference mainly adopts methods based on mathematical programming and deep reinforcement learning schemes. Mathematical programming methods such as linear programming and greedy algorithms can perform resource allocation, but they have poor adaptability in a dynamic environment and are difficult to adjust in real time to cope with network state changes. Although the deep reinforcement learning scheme can adaptively optimize resource allocation, when facing problems such as a large number of user devices, complex model splitting, and coupled resources, its training difficulty and actual application effect are limited. Especially in a multi-user scenario, how to balance the task latency and energy consumption of users while ensuring fair resource allocation has become a technical challenge.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The present application provides a method, an apparatus, and an electronic device for processing inference tasks, so as to at least solve the technical problems in the prior art that in a wireless edge network environment, the multi-user DNN-based inference tasks have deficiencies in aspects such as model splitting and resource competition among users, resulting in increased inference latency and energy consumption.
[0006] According to one aspect of the present application, a method for processing an inference task is provided, including: in each decision cycle, receiving status information corresponding to N user devices through a first server, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user devices; determining, according to a target algorithm, the status information of each user device, and the target model of each user device, a target policy for each user device, where the target policy at least includes a model split point of each user device, and the model split point is used to divide the computational part processed locally by the user device and the computational part remotely processed by a second server in the inference task, the target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information, the structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm for bidirectional traversal to select and adjust the model split point, and the target model is used to calculate the inference task; processing the inference tasks on each user device respectively according to the target policy of each user device.
[0007] Optionally, the target policy of each user device further includes a task window, and processing the inference tasks on each user device respectively according to the target policy of each user device, where the task window is used to characterize the number of inference tasks processed in each decision cycle, including: taking the inference tasks of each user device determined according to the task window in the target policy of each user device as the target tasks of each user device; according to the model split point of each user device, splitting the target model corresponding to each user device into a first part and a second part, where the first part is used to characterize the front-end computational layer of the target model determined based on the model split point, and the second part is used to characterize the back-end computational layer of the target model determined based on the model split point; processing the target tasks of each user device according to the first part and the second part of the target model corresponding to each user device.
[0008] Optionally, the target policy of each user device further includes a transmission channel and a transmission power. The target tasks of each user device are processed according to the first part and the second part of the target model corresponding to each user device. The transmission channel includes a first channel and a second channel, including: through each user device, processing some tasks in the target tasks of each user device in the first part of each user device to obtain intermediate data of each user device, where the intermediate data represents the feature information after the target task is processed by the front-end computing layer of the target model; according to the first channel and the transmission power determined by each user device, transmitting the intermediate data corresponding to each user device to the second server determined by each user device; based on the intermediate data of each user device, through the second server determined by each user device, performing the remaining tasks in the target tasks in the second part of each user device to obtain the inference result corresponding to each user device; transmitting the inference result of each user device to the corresponding user device through the second channel determined by each user device.
[0009] Optionally, the bidirectional graph linked list structure of the target model is obtained through the following steps: converting the initial structure of the target model into a bidirectional graph linked list according to the target algorithm, where the initial structure uses the form of a unidirectional graph to represent the dependency relationship between each computing layer in the target model and the unidirectional flow path of data from the front-end layer to the back-end layer. The bidirectional graph linked list includes multiple node data and edge data, where the node data is used to represent each computing layer of the target model, and the edge data is used to represent the bidirectional flow path of data between each computing layer from the front-end layer to the back-end layer and from the back-end layer to the front-end layer.
[0010] Optionally, determining the target policy of each user device according to the target algorithm, the status information of each user device, and the target model of each user device includes: processing the status information of each user device and the global network environment information through the first network of the target algorithm to obtain global information, where the first network is used to integrate and transform the individual status information of the user device and the information of the entire network environment to generate global information shared by all user devices; determining the target policy of each user device according to the second network of the target algorithm, the global information, the inference task queue of each user device, the status information, and the target model, where the inference task queue is used to represent the set of inference tasks currently to be processed by the user device, and the second network is used to generate the target policy corresponding to each user device based on the information associated with the user device and the global information.
[0011] Optionally, the processing method of the inference task further includes: updating the inference task queue of the user device before each decision cycle, where the inference task queue at least includes: newly added inference tasks, inference tasks to be processed, adjusted inference task priorities, and adjusted status information of the inference tasks.
[0012] Optionally, the method for processing the inference task further includes: after each decision cycle ends, collecting feedback information, where the feedback information is used to characterize the performance parameters of each user device in completing the inference task during the current decision cycle; calculating the reward value of each user device according to the feedback information, where the reward value is used to quantitatively evaluate the performance of the user device during the inference process; updating the first network and the second network of the target algorithm based on the reward value of each user device.
[0013] Optionally, calculating the reward value of each user device according to the feedback information includes: determining the target inference delay of each user device according to the first delay, the second delay, and the third delay of each user device in the feedback information, where the first delay is used to characterize the delay of the user device in locally processing the inference task, the second delay is used to characterize the delay during the data transmission process, and the third delay is used to characterize the delay of the second server in processing the inference task; determining the target energy consumption of each user device according to the first energy consumption and the second energy consumption of each user device in the feedback information, where the first energy consumption is used to characterize the energy consumption during the data transmission process, and the second energy consumption is used to characterize the energy consumption of the user device in locally processing the inference task; determining the expected value of the task window of each user device, where the expected value is used to characterize the average number of inference tasks processed by the user device in each task window; determining the fourth delay, where the fourth delay is used to characterize the time for the user device to wait for other devices to complete the collaborative processing of the inference task; determining the reward value of each user device according to the target inference delay, the target energy consumption, the expected value, and the fourth delay of each user device.
[0014] According to another aspect of the present application, there is also provided a processing device for an inference task, including: a receiving unit, in each decision cycle, receiving the state information corresponding to N user devices through a first server, where N is an integer greater than or equal to 1, and the state information is used to characterize the operating environment state of the user device; a determining unit, determining the target policy of each user device according to the target algorithm, the state information of each user device, and the target model of each user device, where the target policy at least includes the model splitting point of each user device, and the model splitting point is used to divide the calculation part locally processed by the user device and the calculation part remotely processed by the second server in the inference task, the target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the state information of the user device and the global network environment information, the structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm to perform bidirectional traversal to select and adjust the model splitting point, and the target model is used to calculate the inference task; a processing unit, processing the inference tasks on each user device respectively according to the target policy of each user device.
[0015] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the processing method of the above-mentioned inference task.
[0016] In the present application, in each decision cycle, first, the first server receives the status information corresponding to N user devices, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user devices. Then, according to the target algorithm, the status information of each user device, and the target model of each user device, the target policy of each user device is determined. The target policy at least includes the model splitting point of each user device, where the model splitting point is used to divide the computational part processed locally by the user device and the computational part remotely processed by the second server in the inference task. The target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information. The structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm to perform bidirectional traversal to select and adjust the model splitting point. The target model is used to calculate the inference task. Then, the inference tasks on each user device are processed according to the target policy of each user device respectively. That is, through the method of intelligent resource scheduling and model splitting strategy, the purpose of optimizing resource allocation and model splitting point selection is achieved, thereby realizing the technical effect of efficiently and dynamically managing multi-user DNN inference tasks in a wireless edge network environment, and further solving the technical problems in the prior art that in a wireless edge network environment, the multi-user DNN-based inference tasks have deficiencies in aspects such as model splitting and resource competition among users, resulting in increased inference latency and energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0018] Figure 1 is the flowchart of an optional processing method for an inference task according to an embodiment of the present application Figure 1 ;
[0019] Figure 2 is the flowchart of an optional processing method for an inference task according to an embodiment of the present application Figure 2 ;
[0020] Figure 3 is the schematic diagram of an optional processing device for an inference task according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the protection scope of this application.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0023] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) are information and data that have been authorized by the user or fully authorized by all parties. And the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between this system and relevant users or institutions to provide corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0024] According to the embodiments of this application, a method embodiment for processing an inference task is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described here can be executed in a different order than here.
[0025] It should be noted that an intelligent processing system can be used as the execution subject of the method for processing an inference task in the embodiments of this application. It can be understood that the method for processing an inference task provided by the embodiments of this application can also be executed by other systems or devices as the execution subject, and the embodiments of this application do not make specific limitations in this regard.
[0026] Figure 1 is the process of a processing method for an optional inference task according to an embodiment of the present application Figure 1 , such as Figure 1 shown, the method includes the following steps:
[0027] Step S101, in each decision cycle, receive the status information corresponding to N user devices through the first server.
[0028] In step S101, N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user device.
[0029] Optionally, the decision cycle can be regarded as a time window during which a complete resource allocation and task scheduling decision is made to adapt to the dynamic changes of the wireless network.
[0030] Optionally, the first server is a scheduling MEC (Multi-Access Edge Computing) server, which is responsible for coordinating resources, scheduling tasks, and controlling the collaboration in the entire inference process.
[0031] Optionally, the N user devices represent multiple mobile devices participating in collaborative inference in the intelligent processing system, and these devices may have different computing capabilities, battery capacities, and distances from the first server.
[0032] Optionally, the status information includes but is not limited to the remaining inference task quantity of the user device, the computing ability of the device, the battery status of the device, the communication channel quality between the device and the first server, etc. These information are used to characterize the operating environment status of the user device and provide a basis for resource allocation and task scheduling.
[0033] Optionally, the intelligent processing system collects the status information of the user device, enabling the scheduling MEC server to understand the network environment and the capabilities of the user device in real time, providing basic data for subsequent resource allocation and task scheduling.
[0034] Step S102, determine the target strategy of each user device according to the target algorithm, the status information of each user device, and the target model of each user device.
[0035] In step S102, the target strategy at least includes the model splitting point of each user device.
[0036] In step S102, the model splitting point is used to divide the computational part processed locally by the user device and the computational part processed remotely by the second server in the inference task. The target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information. The structure of the target model is a bidirectional graph linked list, which supports the target algorithm to traverse bidirectionally to select and adjust the model splitting point. The target model is used to compute the inference task.
[0037] Optionally, the inference task can be a DNN inference task; the target algorithm is an algorithm based on multi-agent PPO (Proximal Policy Optimization), which is a deep reinforcement learning algorithm used to learn the optimal policy in a dynamic and uncertain environment.
[0038] Optionally, the target model represents the DNN model that needs to be inferred on the user device. This model is modeled as a bidirectional graph linked list, which enables the target algorithm to flexibly perform forward and backward traversals to determine the optimal splitting point of the model.
[0039] Optionally, the target policies include but are not limited to model splitting point selection, channel selection, edge server (referring to the second server, which can also be called the computing MEC server) selection, transmit power selection, and task window size selection. These policies work together to reduce the inference latency and energy consumption of the user device.
[0040] Optionally, the model splitting point is used to split the DNN inference task into two parts, one processed locally by the user device and the other processed remotely by the second server. By precisely selecting the splitting point, the total inference latency and energy consumption can be minimized.
[0041] Optionally, the intelligent processing system combines deep reinforcement learning and a bidirectional graph linked list model. The target algorithm can analyze and predict the optimal resource allocation and task execution strategy under the current network conditions, realizing personalized intelligent scheduling for each user device.
[0042] Step S103, process the inference tasks on each user device according to the target policies of each user device.
[0043] Optionally, each user device performs part of the computation of the inference task locally according to the corresponding target policy, and the remaining tasks will be offloaded to the second server (computing MEC server) for remote processing. This process usually includes local computing, intermediate data transmission, remote computing, and result aggregation.
[0044] Optionally, the intelligent processing system realizes efficient distributed processing of inference tasks in a multi-user environment by executing target policies, reduces the overall inference latency, improves energy efficiency, and at the same time ensures the accuracy of inference tasks and the reasonable utilization of network resources.
[0045] Optionally, Figure 2 is the flow of an optional processing method for inference tasks according to an embodiment of the present application Figure 2 , such as Figure 2 shown, mainly including three modules: user equipment, computing MEC server, and scheduling MEC server. Among them, in the user equipment, first initialize the DNN inference task queue: there is a queue of deep neural network (DNN) inference tasks to be processed on the user equipment. These tasks may involve image recognition, speech analysis, etc., and need to perform inferences to generate results; then collect static information: before starting the inference task, the device collects static information about its own state, including but not limited to the current task queue length, the computing power of the device, the battery state, etc.; after collecting the static information, the scheduling MEC server determines the model splitting point, transmission channel, collaborative computing MEC server, transmission power, and task window size of the user equipment based on the collected static information; then the user equipment executes the start DNN inference task in each task window: the user equipment locally processes part of the DNN inference task based on the model splitting point to obtain intermediate data; when the task is executed to the model splitting point, the user equipment transmits the intermediate data (such as feature maps) generated at this time to the computing MEC server through the determined transmission channel and transmission power; after receiving the remaining inference tasks, the computing MEC server processes them, and after processing, returns the value to the user equipment according to the determined transmission channel and transmission power. The user equipment receives the inference result, and then judges whether all DNN inference tasks are completed. If all are completed, it ends. If there are still remaining DNN tasks to be processed, it collects the static information of the current user equipment again and performs a new round of task processing, that is, according to the currently collected static information, determines the strategy for the new round of inference task processing (the model splitting point, transmission channel, collaborative computing MEC server, transmission power, and task window size of the user equipment) through the scheduling MEC server, and processes the inference tasks according to the strategy until all DNN inference tasks are processed.
[0046] It should be noted that the scheduling MEC server conducts in-depth analysis based on the collected static information to determine the distribution execution strategy of the DNN inference task. Specifically, the analysis result may indicate that the splitting point of the model is at its starting point, which means that the front-end calculation of the inference task should be fully executed on the local device, reducing the overhead of data transmission while making full use of local resources. On the contrary, if the splitting point is determined to be the end point of the model, the scheduling mechanism will offload the entire DNN inference task to the computing MEC server, taking advantage of its powerful computing power and optimized resource allocation to achieve efficient inference. These two extreme cases: local processing starting from the model starting point and server-side centralized processing completed at the end point respectively meet the scenario requirements of sufficient local computing power or resource constraints, ensuring the flexibility and efficiency of DNN inference in different environments.
[0047] As can be seen from the content of steps S101 to S103, in this application, in each decision cycle, first, the first server receives the status information corresponding to N user devices, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user device. Then, according to the target algorithm, the status information of each user device, and the target model of each user device, the target strategy of each user device is determined. The target strategy at least includes the model splitting point of each user device, where the model splitting point is used to divide the computational part of the inference task processed locally by the user device and the computational part remotely processed by the second server. The target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information. The structure of the target model is a bidirectional graph linked list, which supports the target algorithm for bidirectional traversal to select and adjust the model splitting point. The target model is used to calculate the inference task. Then, the inference tasks on each user device are processed according to the target strategy of each user device. That is, through the intelligent resource scheduling and model splitting strategy, the purpose of optimizing resource allocation and model splitting point selection is achieved, thereby realizing the technical effect of efficiently and dynamically managing multi-user DNN inference tasks in a wireless edge network environment, and further solving the technical problems in the prior art that the multi-user DNN-based inference tasks in a wireless edge network environment have deficiencies in aspects such as model splitting and resource competition among users, resulting in increased inference latency and energy consumption.
[0048] In an alternative embodiment, the intelligent processing system first takes the inference task of each user device determined according to the task window in the target policy of each user device as the target task of each user device, and then splits the target model corresponding to each user device into a first part and a second part according to the model split point of each user device, where the first part is used to represent the front-end computing layer of the target model determined based on the model split point, and the second part is used to represent the back-end computing layer of the target model determined based on the model split point, and then processes the target task of each user device according to the first part and the second part of the target model corresponding to each user device.
[0049] Optionally, the intelligent processing system first determines the inference tasks to be processed within the current decision cycle of each user device according to the task window size in the target policy of each user device. This process ensures that each device operates within the reasonable range of its computing and communication capabilities, avoiding over-consumption or insufficiency of resources. Next, according to the model split point determined in the target policy of each user device, the intelligent processing system splits the target model of each user device into a front-end computing layer (the first part) and a back-end computing layer (the second part). Among them, the front-end computing layer includes the input processing and part of the computing layers of the model, which are processed locally by the user device, while the back-end computing layer includes the remaining deep computing layers, which are remotely processed by the second server. Then, the intelligent processing system processes the target tasks of each user device according to the first part and the second part of the split model. Specifically, the user device locally executes the first part (front-end computing layer), processes the input data and generates intermediate data. Subsequently, the intermediate data is transmitted to the corresponding second server through the wireless channel, and the second server then executes the second part (back-end computing layer), performs deep inference based on the received intermediate data, and finally generates the inference result. This process realizes the distributed processing of tasks, in which the local computing and remote computing are optimally coordinated.
[0050] As can be seen from the above, the above method emphasizes the flexibility and efficiency of model splitting and task processing. The intelligent processing system ensures that the inference tasks of each user device can be efficiently processed with the minimum latency and energy consumption in the wireless edge environment by intelligently selecting the model split point. The advantage of this mechanism is that it not only considers the local computing power of the user device, but also optimizes the data transmission and computing resource allocation with the remote server (the second server), thereby maximizing the execution efficiency of the inference task under resource-constrained conditions. Overall, this embodiment improves the performance of DNN collaborative inference in the wireless edge network, reduces the computing burden and energy consumption of the user device, and at the same time reduces the inference latency.
[0051] In an alternative embodiment, the intelligent processing system processes partial tasks in the target tasks of each user device in the first part of each user device to obtain intermediate data of each user device, where the intermediate data represents the feature information after the target model front-end calculation layer processes the target tasks. Then, according to the first channel and transmission power determined by each user device, the intermediate data corresponding to each user device is transmitted to the second server determined by each user device. Then, based on the intermediate data of each user device, the second server determined by each user device executes the remaining tasks in the target tasks in the second part of each user device to obtain the inference result corresponding to each user device. Finally, the inference result of each user device is transmitted to the corresponding user device through the second channel determined by each user device.
[0052] Optionally, the intelligent processing system first guides each user device to process partial tasks in its respective target tasks in the first part (front-end calculation layer) of its target model. This part of the calculation usually involves data preprocessing and the initial several layers of the model, aiming to extract basic feature information from the original input data. The processing result of the front-end calculation layer is encapsulated as intermediate data, which contains the key feature information generated after the target model preliminarily processes the target tasks and provides a basis for subsequent in-depth inference. Next, according to the transmission power and first channel selection determined in the target policy of each user device, the user device transmits the intermediate data obtained in the previous step to the second server it determines. Among them, the determination of the first channel selection and transmission power is based on the current network environment conditions and the communication capabilities of the user device, ensuring the efficiency and stability of data transmission and avoiding the impact of communication bottlenecks on the overall inference performance. Then, the second server receives the intermediate data from the user device and, based on this data, executes the remaining inference tasks in the second part (back-end calculation layer) of the target model of each user device. Among them, the back-end calculation layer usually involves the deep calculation layer of the model, which needs to process the feature information generated by the intermediate data and perform complex inference calculations to obtain the inference result corresponding to each user device. This process makes full use of the computing resources of the second server, alleviates the computing burden of the user device, and shortens the inference latency.
[0053] Optionally, the inference result of each user device is transmitted back to the corresponding user device through the second channel (which may be different from or the same as the first channel, dynamically adjusted and optimized based on the network). The accuracy and timeliness of the inference result ensure a seamless experience for artificial intelligence applications, and through transmission via the second channel, the system can flexibly respond to the dynamic changes of the wireless network, ensuring the reliability and efficiency of data transmission.
[0054] As can be seen from the above, the intelligent processing system effectively realizes the distributed processing of DNN inference tasks in the wireless edge network environment. By decomposing the tasks into a front-end computing layer and a back-end computing layer, and intelligent scheduling between user devices and the second server, it significantly improves the resource utilization efficiency and inference performance. This mechanism not only reduces the computing burden and power consumption of user devices, but also reduces communication latency and network congestion through optimized data transmission strategies, ensuring fast response and high-precision results for real-time inference. In addition, the intelligent processing system can dynamically adjust the task processing strategy according to the network state and user device capabilities, enabling the entire system to remain efficient and stable in the face of multi-user environments and resource constraints, providing strong technical support for artificial intelligence applications in wireless edge networks.
[0055] In an alternative embodiment, the intelligent processing system converts the initial structure of the target model into a bidirectional graph linked list according to the target algorithm. Among them, the initial structure uses the form of a unidirectional graph to represent the dependency relationship between each computing layer in the target model, as well as the unidirectional data flow path from the front-end layer to the back-end layer. The bidirectional graph linked list includes multiple node data and edge data. Among them, the node data is used to represent each computing layer of the target model, and the edge data is used to represent the bidirectional data flow path between each computing layer from the front-end layer to the back-end layer and from the back-end layer to the front-end layer.
[0056] Optionally, the intelligent processing system applies the target algorithm to convert the initial structure of the target model from a unidirectional graph form into a bidirectional graph linked list. This conversion process makes full use of the dependency relationship between computing layers in the target model and the data flow path information from the front-end layer to the back-end layer, providing broader possibilities for subsequent model splitting.
[0057] Optionally, the bidirectional graph linked list structure of the target model constructed by the intelligent processing system contains multiple node data and edge data. Among them, the node data corresponds to each computing layer of the target model and stores the attribute information of the computing layer, such as layer type, input / output data format, computing requirements, etc. The edge data not only describes the unidirectional data flow from the front-end layer to the back-end layer, but also particularly introduces the reverse data flow path from the back-end layer to the front-end layer, supporting flexible traversal based on the model splitting point. The construction process of the bidirectional graph linked list is actually a deep parsing and reconstruction of the target model to better meet the requirements of collaborative inference.
[0058] Optionally, after converting the structure of the target model into a bidirectional graph linked list structure, the intelligent processing system can perform model splitting more efficiently, that is, determine the model splitting point, split the model into a front-end computing layer (the first part) and a back-end computing layer (the second part), and dynamically adjust the model splitting point according to the network environment. The structure of the bidirectional graph linked list allows the algorithm to traverse the model from front to back or from back to front to find the optimal splitting position, so as to achieve a balance between local computing and remote computing, and reduce the overall inference latency and energy consumption.
[0059] As can be seen from the above, the flexibility of DNN model splitting and the optimization level of resource allocation have been significantly improved through the above embodiments. By converting the target model into a bidirectional graph linked list, the intelligent processing system can not only clearly identify the dependency relationships between computing layers in the model, but also more flexibly adjust the model splitting point to adapt to the changing resource conditions and user requirements in the wireless edge network. The core of this mechanism lies in that it enhances the dynamic management ability of computing resources in the collaborative inference scenario, enabling the intelligent processing system to more accurately adjust the task processing strategy according to the real-time network state and computing load, thereby reducing the energy consumption of user equipment and network transmission overhead while ensuring the execution efficiency of inference tasks.
[0060] In an alternative embodiment, the intelligent processing system processes the status information of each user equipment and the global network environment information through the first network of the target algorithm to obtain global information. The first network is used to integrate and transform the individual status information of the user equipment and the information of the entire network environment to generate global information shared by all user equipment. Then, according to the second network of the target algorithm, the global information, the inference task queue of each user equipment, the status information, and the target model, the target strategy of each user equipment is determined. The inference task queue is used to represent the set of inference tasks currently to be processed by the user equipment, and the second network is used to generate the corresponding target strategy for each user equipment based on the information associated with the user equipment and the global information.
[0061] Optionally, at the beginning of each decision cycle, the intelligent processing system processes the status information of each user equipment and the current global network environment information through the first network (embedding network) of the target algorithm. The task of the embedding network is to integrate and transform the scattered individual status information, including but not limited to the computing power, battery status, task queue length of the user equipment, and channel quality data related to wireless communication, while analyzing the resource allocation situation of the entire network, such as the computing power of edge servers and the distribution of network bandwidth. The output of the embedding network is global information, which is a comprehensive data structure containing the current status of all user equipment and the network environment for subsequent decision-making networks to refer to in order to generate optimized resource allocation and task scheduling strategies.
[0062] Optionally, after obtaining the global information, the intelligent processing system determines the target policy for each user device according to the second network (Actor Network) of the target algorithm, the global information obtained in the previous step, the current inference task queue of each user device, the status information, and the target model. Among them, the inference task queue records the set of DNN inference tasks that the user device needs to process, including the priority, status, and required resources of the tasks. The Actor Network then generates personalized policies for each user device based on this information related to the user device and the global information. These policies may include the selection of model splitting points, the allocation of communication resources (such as the setting of channels and transmission power), the assignment of edge servers, and the adjustment of task window sizes. The decision of the Actor Network aims to balance computing resources, communication bandwidth, and user requirements to minimize the inference latency and energy consumption while ensuring the efficient execution of inference tasks and the fair allocation of resources.
[0063] It should be noted that in the context of reinforcement learning, especially in policy gradient methods and Actor-Critic architectures, the Actor Network is responsible for generating actions according to the current environmental state. Briefly speaking, the Actor Network is the component responsible for decision-making in the reinforcement learning system. It learns a policy function to map the state of the environment to actions, thereby guiding the agent to take optimal actions in the environment to achieve specific goals. In multi-agent reinforcement learning, each agent may have its own Actor Network to make decisions independently.
[0064] Optionally, in the multi-agent deep reinforcement learning framework, the actions and states in this embodiment can be represented as a 5-tuple and a 2-tuple respectively. The action is A k =(s k , c k , m k , P k , W k ), which respectively represent the selection of DNN model splitting points (s k ), channel selection (c k ), edge server selection (m k ), transmission power selection (P k ), and task window size selection (W k ); the state is S k =(X k , d k ), which respectively represent the remaining number of inference tasks (X k ) and the distance between the user and the base station (d k). The target algorithm uses the Actor-Critic architecture and combines the multi-dimensional representation of actions and states to achieve deep learning optimization of multi-user DNN collaborative reasoning scheduling and resource allocation in complex wireless edge network environments, thereby improving the resource utilization efficiency and task processing speed of the entire system. Specifically in this solution, the Actor network is responsible for k Output Action A k , the states of all users are connected and first processed by the embedded network (first network) to achieve global information sharing, ensuring that each Agent network (referring to each user device) can take into account the states of all other users, so as to make more coordinated action decisions. After the state information of each user device passes through the embedded network, it enters multiple Agent networks respectively. The corresponding Agent network generates an action A according to the state information it processes. k , these actions include the selection of model split points, channels, edge servers, transmission power, and task window size, guiding user devices on how to efficiently perform reasoning tasks. The Critic network is responsible for evaluating the quality of the actions generated by the Actor network, that is, in a given state S k Next, Action A k The output of the Critic network is the value of the task within the task window, which will serve as feedback to guide the training of the Actor network and help the Actor network adjust its strategy to generate more optimized actions.
[0065] From the above content, it can be seen that through the collaborative work of the embedded network and the actor network, the intelligent processing system can achieve real-time optimization control of multi-user DNN collaborative reasoning in the wireless edge network. The embedded network is responsible for integrating and converting scattered state information to generate a global information reflecting the state of the entire system, providing comprehensive and in-depth data support for decision-making. Based on this information, the actor network generates personalized target strategies to guide the resource usage and task execution of each user device, ensuring that all reasoning tasks can be processed most effectively under limited resource conditions. The core advantage of this mechanism is that it uses deep learning technology to achieve intelligent conversion from individual state information to global optimization strategies, significantly improves resource utilization efficiency, reduces reasoning latency and energy consumption, and thus provides an efficient and flexible solution for DNN reasoning in wireless edge computing environments. At the same time, it also improves the system's adaptability, enabling it to quickly respond to changes in network conditions and fluctuations in user demand.
[0066] In an optional embodiment, before each decision cycle, the intelligent processing system updates the reasoning task queue of the user device, wherein the reasoning task queue includes at least: newly added reasoning tasks, pending reasoning tasks, adjusted reasoning task priorities, and adjusted reasoning task status information.
[0067] Optionally, to ensure that the intelligent processing system can make optimal resource allocation and task scheduling decisions in each decision cycle, the system updates the inference task queue before the start of each cycle to reflect the latest changes in the current network environment and user device status.
[0068] Optionally, first, the intelligent processing system adds the newly added inference tasks of each user device in the current cycle to its inference task queue. These newly added tasks may originate from new data input of the user device or subsequent inference requirements derived from other tasks. The addition of newly added tasks ensures that the queue can timely reflect the actual workload of the user device. Subsequently, the system updates the unfinished inference tasks in the inference task queue, including their status information and priorities. Among them, the status information may include the current completion degree of the task, the types and quantities of required resources, the estimated completion time, etc. Updating the priority is based on the urgency of the task, resource requirements, and the capability status of the user device to determine the order of task execution. This update process enables the system to dynamically adjust resources and prioritize critical or urgent tasks, thereby improving the overall response speed and efficiency.
[0069] Optionally, the intelligent processing system adjusts the priorities of inference tasks according to network conditions, device status, and task characteristics. For example, if the battery power of a certain user device is low or the network connection is unstable, the priority of its inference tasks may be increased to ensure preferential service under resource-constrained conditions. By dynamically adjusting task priorities, the system can more effectively balance the needs of multiple users, optimize resource allocation, and reduce task processing delays. Finally, the system also updates the status information of the tasks to reflect the current status of the tasks, such as whether the processing has started, the processing progress, and whether there are resource bottlenecks. The update of status information is crucial for the intelligent processing system because it provides the key basis for decision-making, helps the system monitor the task execution situation in real time, and timely adjust strategies.
[0070] As can be seen from the above, by updating the inference task queue before each decision cycle, the intelligent processing system can achieve fine-grained control of multi-user DNN collaborative inference in the wireless edge network. This mechanism not only ensures that the system can timely respond to the changes in newly added tasks and user requirements, but also optimizes resource allocation and improves the efficiency and accuracy of task execution by dynamically adjusting task priorities and status information. Compared with static or preset queue management methods, the method of dynamically updating the inference task queue significantly enhances the adaptability and flexibility of the system, can better cope with complex and changeable network environments, ensure that all inference tasks can be processed in a timely and effective manner under limited resources, and thus improve the overall performance of artificial intelligence applications in the wireless edge network.
[0071] In an alternative embodiment, after each decision cycle, the intelligent processing system collects feedback information, where the feedback information is used to characterize the performance parameters of each user device in completing the inference task during the current decision cycle, and then calculates the reward value for each user device according to the feedback information, where the reward value is used to quantitatively evaluate the performance of the user device during the inference process, and finally updates the first network and the second network of the target algorithm based on the reward value of each user device.
[0072] Optionally, after each decision cycle, the intelligent processing system actively collects feedback information, which is derived from the performance of each user device in processing the inference task during the current cycle. The feedback information mainly includes performance parameters such as inference latency, energy consumption, and task completion rate, etc., which directly reflect the efficiency and resource usage of the user device when executing the inference task. Collecting feedback information is an important way for the intelligent processing system to understand the effect of the current strategy and changes in the network environment, providing a data basis for subsequent strategy adjustment.
[0073] Optionally, based on the collected feedback information, the intelligent processing system calculates the reward value for each user device. The design of the reward value aims to quantitatively evaluate the performance of the user device during the inference process. It is usually a comprehensive indicator that reflects the trade-off between latency, energy consumption, and task completion rate. A high reward value indicates that the user device has performed well in the current decision cycle, while a low reward value indicates that it is necessary to adjust the resource allocation and task scheduling strategy to cope with network fluctuations or device capacity limitations.
[0074] Optionally, the intelligent processing system updates the first network and the second network of the target algorithm based on the reward value of each user device. Through the feedback of the reward value, the intelligent processing system can adjust the decision-making strategy of the second network to make it more adaptable to the current network environment and user device status.
[0075] As can be seen from the above, by periodically collecting feedback information, calculating reward values, and updating the algorithm network, the system can continuously learn and optimize the resource scheduling strategy to cope with changing network conditions and diverse user requirements. This approach not only improves the execution efficiency of the inference task but also can dynamically adjust the model splitting point, transmission parameters, and computing resource allocation, minimizing latency and energy consumption to the greatest extent and enhancing the inference performance of user devices. In addition, by introducing the reward value, the intelligent processing system can quantitatively evaluate the quality of the strategy, providing an effective learning signal for deep reinforcement learning, accelerating the convergence of the algorithm, ensuring the fairness and efficiency of resource allocation in a multi-user environment, and overall improving the intelligent management level of DNN collaborative inference in the wireless edge network.
[0076] In an alternative embodiment, the intelligent processing system first determines the target inference latency of each user device based on the first latency, second latency, and third latency of each user device in the feedback information, where the first latency is used to characterize the latency of the user device in locally processing the inference task, the second latency is used to characterize the latency during data transmission, and the third latency is used to characterize the latency of the second server in processing the inference task. Then, it determines the target energy consumption of each user device based on the first energy consumption and second energy consumption of each user device in the feedback information, where the first energy consumption is used to characterize the energy consumption during data transmission, and the second energy consumption is used to characterize the energy consumption of the user device in locally processing the inference task. At the same time, it determines the expected value of the task window of each user device, where the expected value is used to characterize the average number of inference tasks processed by the user device within each task window, determines the fourth latency, where the fourth latency is used to characterize the time for the user device to wait for other devices to complete the collaborative processing of the inference task, and finally determines the reward value of each user device based on the target inference latency, target energy consumption, expected value, and fourth latency of each user device.
[0077] Optionally, the intelligent processing system first extracts three key latency metrics for each user device based on the feedback information, namely the local processing latency (the first latency), the data transmission latency (the second latency), and the second server processing latency (the third latency). These latency metrics reflect the complete process of the inference task being processed locally on the user device, transmitted to the second server, and processed on the server. By analyzing these metrics, the system determines the target inference latency of each user device, which ensures the fast and efficient completion of the inference task. Then, the intelligent processing system sets a target energy consumption for each user device based on the first energy consumption (data transmission energy consumption) and the second energy consumption (local processing energy consumption) in the feedback information. By considering the two main sources of energy consumption, the system can optimize the overall energy consumption while maintaining the processing speed of the inference task and reducing the energy consumption of the user device.
[0078] Optionally, the intelligent processing system also calculates the expected value of the task window of each user device, that is, the average number of inference tasks processed by the user device within each task window. This expected value helps to balance the task allocation, enabling the user device to process a reasonable number of tasks within a limited time window and avoiding overuse or waste of resources.
[0079] Optionally, to further optimize the efficiency of collaborative processing, the intelligent processing system introduces the concept of the fourth latency, that is, the time for the user device to wait for other devices to complete the collaborative processing of the inference task. By calculating this waiting time, the system can evaluate and adjust the resource allocation strategy to reduce the additional latency caused by waiting for collaborative processing.
[0080] Optionally, the intelligent processing system comprehensively considers all of the above objectives, namely the target inference latency, target energy consumption, expected value of the task window, and the fourth latency, to determine the reward value for each user equipment. The calculation of the reward value is based on the multi-agent PPO algorithm, which reflects the efficiency and energy consumption of the user equipment in completing tasks, as well as the overall resource utilization efficiency of the system. By continuously adjusting and optimizing these parameters, the system can guide the user equipment and the second server to make more optimal decisions in subsequent inference tasks.
[0081] Optionally, the rate R at which user equipment n sends intermediate data to the edge server n (π) is calculated as shown in Equation (1):
[0082]
[0083] where P n (π) represents the transmit power of user equipment n, G n represents the channel gain of user equipment n, B c represents the bandwidth of channel c, σ c represents the noise power of the channel, ∑ k∈{1,…,N},k≠n P k (π)G k represents the sum of the products of the transmit powers of all user equipment other than user equipment n and their respective channel gains to the edge server, representing the total intensity of channel interference caused by other user equipment when user equipment n is performing data transmission.
[0084] Optionally, the intermediate data transmission latency (second latency) L corresponding to user equipment n trans,n (π) is calculated as shown in Equation (2):
[0085]
[0086] where I n (π) represents the amount of intermediate data that user equipment n needs to transmit, that is, the data size generated at a certain checkpoint during the execution of the DNN inference task and needs to be sent to the edge server for subsequent processing.
[0087] Optionally, the target inference latency of each user equipment is calculated as shown in Equation (3):
[0088] L n (π) = L local,n (π) + L trans,n (π) + L server,n (π) (3)
[0089] where L n (π) represents the target inference latency of the nth user, Llocal,n (π) represents the local inference latency (the first latency) of the user equipment, L trans,n (π) represents the intermediate data transmission latency (the second latency), L server,n (π) represents the inference latency of the second server (the third latency), where, L local,n (π) is equal to the amount of data that the DNN model needs to process on the user equipment divided by the computing rate of the equipment, reflecting the time required for the user equipment n to process a specific model task, L server,n (π) is directly proportional to the baseline inference latency of the server and inversely proportional to the number of tasks executed on the current server.
[0090] Optionally, the intermediate data transmission energy consumption (the first energy consumption) E corresponding to the user equipment n trans,n The calculation formula of (π) is as shown in (4):
[0091] E trans,n (π) = P n (π) × L trans,n (π) (4)
[0092] Where, P n (π) represents the transmission power of the user equipment n when transmitting intermediate data to the edge server.
[0093] Optionally, the target energy consumption E corresponding to the user equipment n n The calculation formula of (π) is as shown in (5):
[0094] E n (π) = E trans,n (π) + ζM n (π) (5)
[0095] Where, ζ is a balance parameter used to adjust the relative importance between the intermediate data transmission energy consumption and the user local inference calculation energy consumption, M n (π) represents the energy consumption of the user equipment n during local inference calculation.
[0096] Optionally, in this embodiment, the goal of multi-user equipment DNN collaborative inference is to minimize the sum of the latency and energy consumption of all user equipment to complete all their tasks, then the above goal can be expressed as formula (6):
[0097]
[0098] Where, is used to represent the inference task queue of the user equipment n, N is the total number of user equipment, is used to represent the maximized user inference latency, represents the average energy consumption of user equipment. The constraint C1 indicates that the balance parameter β and the computing load balance parameter ζ must be greater than 0, ensuring their positive impact in the optimization problem; the constraint C2 determines that the transmit power P n (π) of each user equipment n must be between 0 and P limit , where P limit is the upper limit of the power, which ensures that the transmit power of the user equipment is within a reasonable and safe operating range;
[0099] Optionally, since the formula (6) does not represent a Markov decision process, the above objective needs to be further transformed. Considering each state as a task window containing a small number of inference tasks to be executed, the objective can be transformed into minimizing the average inference delay and energy consumption within the task window. The transformed objective can be expressed as formula (7):
[0100]
[0101] where K(π) is the number of task windows required to complete all inference tasks, L n,k (π) represents the total inference delay of user equipment n within the k-th task window, and L wait,k (π) represents the average waiting time (the fourth delay) within the k-th task window, which refers to the average time that the user equipment waits for other users to complete inference within this window; E k (π) represents the total energy consumption of all users within this window; W(π) is the expected value (expectation) of the task window size; the constraint C1 indicates that the delay penalty factor α, the energy consumption balance factor β, and the computing load balance parameter ζ must be greater than 0, ensuring the positive role of these parameters in the optimization problem and they belong to the real number domain; the constraint C2 determines that the transmit power (π) of each user equipment n must be between 0 and P limit , ensuring that the power control of the user equipment is within a reasonable range, where P limit is the upper limit of the transmit power; the constraint C3 ensures that the size W(π) of the task window must be greater than 0, which is the basic condition for effective scheduling.
[0102] Optionally, based on the above, the three elements of the target algorithm in this embodiment can be determined: reward, action, and state. In reinforcement learning, the reward is the feedback obtained by the agent (equivalent to each user equipment in this embodiment) after taking an action, used to guide the agent on how to adjust the strategy to optimize its performance. The reward value for each iteration can be expressed as the negative value of the objective, as shown in formula (8):
[0103]
[0104] where r k$R^{(n)}_k(\pi)$ represents the reward value of user equipment $n$ in the $k$-th task window; $L$ n,k $D^{(n)}_k(\pi)$ represents the total inference latency of user equipment $n$ within the $k$-th task window; $L$ wait,k $W^{(n)}_k(\pi)$ represents the average waiting time (the fourth latency) within the $k$-th task window, which refers to the average time that user equipment waits for other users to complete inference within this window; $E$ k $P(\pi)$ represents the total energy consumption of all users within this window; $W(\pi)$ is the expectation (expected value) of the task window size; $\alpha$ is a latency penalty factor; $\beta$ is an energy consumption balance factor; $N$ is the total number of user equipment.
[0105] As can be seen from the above, the intelligent processing system dynamically adjusts and optimizes the target inference latency, target energy consumption, task window expected value, and collaborative processing waiting time of user equipment, achieving efficient and energy-saving multi-user DNN collaborative inference in the wireless edge network. This method ensures the satisfaction of real-time requirements, reduces the additional latency caused by waiting for collaborative processing, improves resource utilization efficiency, and ultimately provides a better DNN inference experience for users. Through the design of the reward mechanism, the overall performance of the system is further promoted, reflecting the adaptability and optimization ability of the algorithm in complex wireless edge environments.
[0106] The embodiment of this application also provides a processing device for inference tasks. It should be noted that the processing device for inference tasks in the embodiment of this application can be used to execute the processing method for inference tasks provided in the embodiment of this application. The following introduces the processing device for inference tasks provided in the embodiment of this application.
[0107] According to the embodiment of this application, there is also provided a device for implementing the above-mentioned processing method for inference tasks. Figure 3 FIG. is a schematic diagram of an optional processing device for inference tasks according to the embodiment of this application, as Figure 3 shown, the device includes: a receiving unit 301, a determining unit 302, and a processing unit 303.
[0108] Optionally, a receiving unit 301 receives, in each decision cycle, status information corresponding to N user devices through a first server, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user devices; a determining unit 302 determines, according to a target algorithm, the status information of each user device, and the target model of each user device, a target policy for each user device, where the target policy includes at least a model splitting point of each user device, where the model splitting point is used to divide the computational part processed locally by the user device and the computational part remotely processed by a second server in an inference task, the target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information, the structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm for bidirectional traversal to select and adjust the model splitting point, and the target model is used to calculate the inference task; a processing unit 303 processes the inference tasks on each user device according to the target policy of each user device.
[0109] Optionally, the processing unit 303 includes: a first determining subunit, a first splitting subunit, and a first processing subunit. Among them, the first determining subunit is configured to use the inference task of each user device determined according to the task window in the target policy of each user device as the target task of each user device; the first splitting subunit is configured to split the target model corresponding to each user device into a first part and a second part according to the model splitting point of each user device, where the first part is used to characterize the front-end computational layer of the target model determined based on the model splitting point, and the second part is used to characterize the back-end computational layer of the target model determined based on the model splitting point; the first processing subunit is configured to process the target task of each user device according to the first part and the second part of the target model corresponding to each user device.
[0110] Optionally, the first processing subunit includes: a first processing module, a first transmission module, a second processing module, and a first determination module. The first processing module is configured to process, through each user device, partial tasks in the target task of each user device in the first part of each user device to obtain intermediate data of each user device, where the intermediate data represents the feature information after the front-end computing layer of the target model processes the target task; the first transmission module is configured to transmit the intermediate data corresponding to each user device to the second server determined by each user device according to the first channel and transmission power determined by each user device; the second processing module is configured to, based on the intermediate data of each user device, execute the remaining tasks in the target task in the second part of each user device through the second server determined by each user device to obtain the inference result corresponding to each user device; the first determination module is configured to transmit the inference result of each user device to the corresponding user device through the second channel determined by each user device.
[0111] Optionally, the determination unit 302 includes: a first conversion subunit, configured to convert the initial structure of the target model into a bidirectional graph linked list according to the target algorithm, where the initial structure represents the dependency relationship between each computing layer in the target model and the unidirectional data flow path from the front-end layer to the back-end layer in the form of a unidirectional graph, and the bidirectional graph linked list includes a plurality of node data and edge data, where the node data is used to represent each computing layer of the target model, and the edge data is used to represent the bidirectional data flow path between each computing layer from the front-end layer to the back-end layer and from the back-end layer to the front-end layer.
[0112] Optionally, the determination unit 302 includes: a second processing subunit and a second determination subunit. The second processing subunit is configured to process the status information of each user device and the global network environment information through the first network of the target algorithm to obtain global information, where the first network is used to integrate and transform the individual status information of the user device and the information of the entire network environment to generate global information shared by all user devices; the second determination subunit is configured to determine the target policy of each user device according to the second network of the target algorithm, the global information, the inference task queue of each user device, the status information, and the target model, where the inference task queue is used to represent the set of inference tasks currently to be processed by the user device, and the second network is used to generate the target policy corresponding to each user device based on the information associated with the user device and the global information.
[0113] Optionally, the processing device for inference tasks further includes: a first update unit, configured to update the inference task queue of the user device before each decision period, where the inference task queue at least includes: newly added inference tasks, inference tasks to be processed, adjusted inference task priorities, and adjusted status information of the inference tasks.
[0114] Optionally, the processing device for the inference task further includes: a first collection unit, a first calculation unit, and a second update unit. Among them, the first collection unit is configured to collect feedback information after the end of each decision cycle, where the feedback information is used to characterize the performance parameters of each user device in completing the inference task during the current decision cycle; the first calculation unit is configured to calculate the reward value of each user device according to the feedback information, where the reward value is used to quantitatively evaluate the performance of the user device during the inference process; the second update unit is configured to update the first network and the second network of the target algorithm based on the reward value of each user device.
[0115] Optionally, the first calculation unit includes: a third determination subunit, a fourth determination subunit, a fifth determination subunit, a sixth determination subunit, and a seventh determination subunit. Among them, the third determination subunit is configured to determine the target inference delay of each user device according to the first delay, the second delay, and the third delay of each user device in the feedback information, where the first delay is used to characterize the delay of the user device in locally processing the inference task, the second delay is used to characterize the delay during data transmission, and the third delay is used to characterize the delay of the second server in processing the inference task; the fourth determination subunit is configured to determine the target energy consumption of each user device according to the first energy consumption and the second energy consumption of each user device in the feedback information, where the first energy consumption is used to characterize the energy consumption during data transmission, and the second energy consumption is used to characterize the energy consumption of the user device in locally processing the inference task; the fifth determination subunit is configured to determine the expected value of the task window of each user device, where the expected value is used to characterize the average number of inference tasks processed by the user device in each task window; the sixth determination subunit is configured to determine the fourth delay, where the fourth delay is used to characterize the time for the user device to wait for other devices to complete collaborative processing of the inference task; the seventh determination subunit is configured to determine the reward value of each user device according to the target inference delay, the target energy consumption, the expected value, and the fourth delay of each user device.
[0116] According to another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory storing an executable program; a processor configured to run the program, where when the program runs, it executes the above-mentioned processing method for the inference task.
[0117] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0118] In the above embodiments of the present application, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0119] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0120] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0121] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0122] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0123] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A processing method for an inference task, characterized in that, Including: In each decision cycle, receive the status information corresponding to N user devices through a first server, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user devices; Determine the target policy for each user device according to the target algorithm, the status information of each user device, and the target model of each user device. The target policy at least includes the model splitting point of each user device, where the model splitting point is used to divide the computational part processed locally by the user device and the computational part remotely processed by a second server in the inference task. The target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information. The structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm for bidirectional traversal to select and adjust the model splitting point. The target model is used to calculate the inference task; Process the inference tasks on each user device according to the target policy of each user device respectively.
2. The processing method of the inference task according to claim 1, wherein The target policy of each user device further includes a task window. Process the inference tasks on each user device according to the target policy of each user device respectively, where the task window is used to characterize the number of inference tasks processed in each decision cycle, including: Take the inference task of each user device determined according to the task window in the target policy of each user device as the target task of each user device; According to the model splitting point of each user device, split the target model corresponding to each user device into a first part and a second part, where the first part is used to characterize the front-end computational layer of the target model determined based on the model splitting point, and the second part is used to characterize the back-end computational layer of the target model determined based on the model splitting point; Process the target task of each user device according to the first part and the second part of the target model corresponding to each user device.
3. The method for processing an inference task according to claim 2, wherein The target policy of each user device further includes a transmission channel and a transmission power. Process the target task of each user device according to the first part and the second part of the target model corresponding to each user device, where the transmission channel includes a first channel and a second channel, including: Through each user device, process part of the tasks in the target task of each user device in the first part of each user device to obtain the intermediate data of each user device, where the intermediate data characterizes the feature information after the front-end computational layer of the target model processes the target task; Transmit the intermediate data corresponding to each user device to the second server determined by each user device according to the first channel determined by each user device and the transmission power; Based on the intermediate data of each of the user devices, the second server determined by each of the user devices executes the remaining tasks in the target task in the second part of each of the user devices, and obtains the inference result corresponding to each of the user devices; Transmit the inference result of each of the user devices to the corresponding user device through the second channel determined by each of the user devices.
4. The method for processing an inference task according to claim 1, wherein The bidirectional graph linked list structure of the target model is obtained through the following steps: Convert the initial structure of the target model into the bidirectional graph linked list according to the target algorithm, where the initial structure represents the dependency relationship between each computing layer in the target model and the unidirectional data flow path from the forward layer to the backward layer in the form of a unidirectional graph. The bidirectional graph linked list includes a plurality of node data and edge data, where the node data is used to represent each computing layer of the target model, and the edge data is used to represent the bidirectional data flow path between each computing layer from the forward layer to the backward layer and from the backward layer to the forward layer.
5. The method for processing the inference task according to claim 1, wherein Determine the target strategy of each user device according to the target algorithm, the status information of each user device, and the target model of each user device, including: Process the status information of each user device and the global network environment information through the first network of the target algorithm to obtain global information, where the first network is used to integrate and transform the individual status information of the user device and the information of the entire network environment to generate the global information shared by all user devices; Determine the target strategy of each user device according to the second network of the target algorithm, the global information, the inference task queue of each user device, the status information, and the target model, where the inference task queue is used to represent the set of inference tasks currently to be processed by the user device, and the second network is used to generate the target strategy corresponding to each user device based on the information associated with the user device and the global information.
6. The method for processing the inference task according to claim 1, wherein The method further includes: Before each decision cycle, update the inference task queue of the user device, where the inference task queue at least includes: newly added inference tasks, inference tasks to be processed, adjusted inference task priorities, and status information of adjusted inference tasks.
7. The method for processing an inference task according to claim 5, wherein The method further includes: After each decision cycle, collect feedback information, where the feedback information is used to represent the performance parameters of each user device in completing the inference task during the current decision cycle; Calculate the reward value of each user device according to the feedback information, where the reward value is used to quantitatively evaluate the performance of the user device during the inference process; Update the first network and the second network of the target algorithm based on the reward value of each user device.
8. The method for processing the inference task according to claim 7, wherein Calculating the reward value of each user device according to the feedback information includes: Determine the target inference latency of each user device according to the first latency, the second latency, and the third latency of each user device in the feedback information, where the first latency is used to characterize the latency of the user device in locally processing the inference task, the second latency is used to characterize the latency in the data transmission process, and the third latency is used to characterize the latency of the second server in processing the inference task; Determine the target energy consumption of each user device according to the first energy consumption and the second energy consumption of each user device in the feedback information, where the first energy consumption is used to characterize the energy consumption in the data transmission process, and the second energy consumption is used to characterize the energy consumption of the user device in locally processing the inference task; Determine the expected value of the task window of each user device, where the expected value is used to characterize the average number of inference tasks processed by the user device in each task window; Determine the fourth latency, where the fourth latency is used to characterize the time for the user device to wait for other devices to complete the collaborative processing of the inference task; Determine the reward value of each user device according to the target inference latency, the target energy consumption, the expected value, and the fourth latency of each user device.
9. A processing device for an inference task, characterized in that, Comprising: A receiving unit, in each decision cycle, receives the status information corresponding to N user devices through a first server, where N is an integer greater than or equal to 1, and the status information is used to characterize the operating environment status of the user device; A determining unit, according to the target algorithm, the status information of each user device, and the target model of each user device, determines the target policy of each user device, where the target policy at least includes the model splitting point of each user device, and the model splitting point is used to divide the computational part of the inference task processed locally by the user device and the computational part remotely processed by the second server. The target algorithm determines the resource allocation for the inference task of each user device by integrating and analyzing the status information of the user device and the global network environment information. The structure of the target model is a bidirectional graph linked list, and this structure supports the target algorithm to perform bidirectional traversal to select and adjust the model splitting point. The target model is used to calculate the inference task; A processing unit, processes the inference tasks on each user device respectively according to the target policy of each user device.
10. An electronic device, characterized in that, Comprises one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the processing method of the inference task according to any one of claims 1 to 8.
Citation Information
Patent Citations
Motion sensing edge reasoning method based on model division and service migration
CN116483564A
Distributed model reasoning acceleration method and system under edge computing network
CN117933388A
DNN reasoning acceleration method under edge cloud collaboration
CN118228782A
Elastic configuration and two-way collaborative adaptation system for application computing task
CN119668849A