Multi-User Edge-Cloud Collaborative Inference Multi-Objective Optimization Method and System for Mobile Platforms
Through the multi-objective optimization model, the resource configuration of the mobile platform is optimized in multi-user scenarios, and the dynamic segmentation and collaborative reasoning of neural networks are realized, which solves the problems of resource constraints and task competition, and improves the efficiency and performance of the system.
Patent Information
- Application Number
- CN202510019416.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In multi-user scenarios, the different computing resources, network bandwidth, power and communication distances from servers of the mobile platform have different tasks and resource optimizations, which have led to challenges in task allocation and resource optimization. The traditional centralized computing model cannot fully utilize the computing power of edge devices, and there is large delay and energy consumption, resulting in system performance degradation.
Through the multi-objective optimization model, based on the computing power, communication conditions and distance from the server of the mobile platform, the optimal decision variables of the neural network are determined, dynamic segmentation and parameter configuration are performed, so as to realize the coordinated reasoning between the mobile platform and the server, and optimize time and energy consumption.
It achieves the optimal balance in time, energy consumption and resource utilization, improves the efficiency and performance of the system, and is suitable for application scenarios such as smart mobile terminals and Internet of Things devices.
Smart Images

Figure CN119420759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of edge intelligent computing, and particularly to a multi-objective optimization method and system for multi-user edge-cloud collaborative inference on a mobile platform. Background Art
[0002] With the rapid development of edge computing, the deployment demand for deep learning applications on resource-constrained mobile platform devices is increasing day by day. Mobile platforms are widely used in scenarios such as intelligent manufacturing, drone control, and smart city management. These applications usually rely on complex neural network inference tasks to achieve real-time data processing and intelligent decision-making. However, the computing resources, storage space, and battery capacity of mobile platforms are relatively limited, making it difficult to support the independent execution of high-computation and low-latency deep learning tasks. This limitation makes edge devices often need to work in cooperation with edge servers when processing complex inference tasks, and offload some computing tasks to the edge servers to make up for the insufficient computing power of a single device.
[0003] In a multi-user scenario, multiple mobile platforms simultaneously access the edge server to jointly complete inference tasks. However, due to the differences in the computing power, network bandwidth, power, and communication distance from the server of mobile platforms, how to perform task allocation and resource optimization in a multi-user scenario has become a major challenge. The traditional centralized computing mode cannot fully utilize the computing power of edge devices, and due to the limitations of network bandwidth and power, large delays and energy consumption may occur during task transmission, resulting in a decline in the overall performance of the system.
[0004] Therefore, how to improve the efficiency and performance of the system is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-objective optimization method and system for multi-user edge-cloud collaborative inference on a mobile platform to solve the problems of resource-constrained inference, multi-user task competition, and the balance between energy consumption and inference efficiency.
[0006] To achieve the above purpose, the following technical solutions are adopted:
[0007] In a first aspect, the present invention provides a multi-objective optimization method for multi-user edge-cloud collaborative inference on a mobile platform, and the method includes:
[0008] When multiple mobile platforms or edge device clients are detected within the communication radius of a multi-user collaborative inference server, the mobile platforms or edge devices are automatically added to the edge collaborative inference system; wherein, the communication radius refers to the effective transmission distance between the server and the mobile platform or edge device.
[0009] When a mobile platform or edge device joins an edge collaborative inference system, obtain the platform performance parameters and neural network information; where the platform performance parameters include the computing power, communication conditions, and the distance between the platform and the server of the platform, and the neural network information includes the computational complexity, the number of parameters, and the feature map size between each layer;
[0010] Based on the platform performance parameters and neural network information, according to the different computing capabilities, communication conditions, and the distance from the server of the mobile platform, use a multi-objective optimization model to determine the Pareto solution set of the optimal decision variables of the neural network; where the decision variables include the model segmentation point, communication bandwidth, transmission power, and the main frequency of the platform device;
[0011] Based on the decision variables, through a collaborative inference framework, dynamically segment the neural network and perform corresponding parameter configuration;
[0012] The mobile platform and the server start to execute the collaborative inference task; during the collaborative inference process, the mobile platform and the server continuously interact. The mobile platform transmits some calculation results to the server, and the server executes the remaining complex calculations. The entire inference process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm.
[0013] Furthermore, when a mobile platform or edge device joins an edge collaborative inference system, obtain the platform performance parameters and neural network information, including:
[0014] Obtain the platform performance parameters through the self-report of the platform or the query of the central controller;
[0015] Obtain the model structure of the neural network, layer the model structure of the neural network, and extract the neural network information; where the model structure of the neural network is layered through the following method:
[0016] Ensure that there are no complex connections between two neural network layers to split the network into multiple sub-networks;
[0017] Multiple sub-networks operate and train independently;
[0018] Use a distributed computing framework to allocate the sub-networks to different mobile platforms for processing.
[0019] Furthermore, the multi-objective optimization model is constructed through the following method:
[0020] Taking time consumption and energy consumption as optimization objectives, a time model and an energy consumption model are established respectively; among them, the time consumption includes the calculation time and transmission time of the platform side, which are calculated according to the computing power of the mobile platform and the bandwidth conditions of wireless transmission respectively, and the energy consumption includes the computing energy consumption and transmission energy consumption of the mobile platform, which are calculated according to the energy consumption efficiency and transmission power of the mobile platform respectively;
[0021] Based on the time model and the energy consumption model, the multi-objective optimization model is constructed.
[0022] Furthermore, the multi-objective optimization model is expressed as:
[0023] ;
[0024] In the formula, represents the preference of the decision maker, represents the decision variable corresponding to the minimum objective function value, represents the number of mobile platforms, represents the total collaborative inference delay of the th mobile platform, represents the total energy consumption of the
[0025] th mobile platform.
[0026] ;
[0027] In the formula, represents the total collaborative inference delay, represents the computing delay of the platform, represents the transmission delay of the platform;
[0028] ;
[0029] In the formula, represents the model segmentation point of the platform, , represents the number of floating-point operations of the data on the th layer of the platform, represents the number of floating-point operations per second of the platform, which is obtained by multiplying the platform device frequency by the number of floating-point operations per cycle, represents the number of floating-point operations per cycle,
[0030] ;
[0031] In the formula, represents the data volume size of the output feature variable at the model segmentation point of the platform, Represents the transmission rate;
[0032] ;
[0033] Wherein, and respectively represent the bandwidth and power of the platform, represents the noise power spectrum, represents the average channel gain;
[0034] ;
[0035] Wherein, represents the weight of the channel gain, and and and respectively represent the radio propagation speed, antenna gain, carrier frequency and path loss exponent, represents the distance between the platform and the server.
[0036] Furthermore, the energy consumption model is expressed as:
[0037] ;
[0038] Wherein, represents the total energy consumption of collaborative inference, represents the computing energy consumption of platform data, represents the transmission energy consumption of the platform's intermediate feature variables transmitted to the server, represents the energy efficiency coefficient of the processor chip, represents the chip main frequency of the platform, represents the power of the platform.
[0039] Furthermore, based on the platform performance parameters and neural network information, calculations are performed using a multi-objective optimization model, including:
[0040] Solving the multi-objective optimization model based on a multi-objective optimization algorithm. The multi-objective optimization algorithm generates multiple candidate solutions by performing Gaussian sampling in the solution space; each candidate solution represents a configuration, including the cut point, bandwidth allocation, power consumption, and operation frequency; the optimal configuration is selected by evaluating the time consumption and energy consumption of each candidate solution.
[0041] Furthermore, the collaborative inference framework includes an edge platform and a cloud server. Based on the decision variables, the neural network is dynamically segmented through the collaborative inference framework, and corresponding parameter configurations are performed, including:
[0042] According to the current system resource situation, select the model cut point, and divide the neural network model into two parts, namely ME and M S , M E is the part of the model running on the edge platform, and M S is the part of the model running on the cloud server;
[0043] After the edge platform receives the image data, it uses M E to perform forward inference on the image to obtain the intermediate feature variable IF. After the edge platform compresses and encodes IF, it transmits it to the cloud server through the wireless network;
[0044] After the cloud server receives IF, it decodes and restores it, and uses M S to perform forward inference to obtain the detection result PR; where PR contains the category and location information of the weed targets in the image; the cloud server transmits the detection result PR back to the edge platform.
[0045] In a second aspect, the present invention provides a multi-user edge-cloud collaborative inference multi-objective optimization system for a mobile platform. The system includes a processor configured to:
[0046] When multiple mobile platforms or edge device clients are detected within the communication radius of the multi-user collaborative inference server, automatically add the mobile platforms or edge devices to the edge collaborative inference system; where the communication radius refers to the effective transmission distance between the server and the mobile platform or edge device;
[0047] When a mobile platform or edge device joins the edge collaborative inference system, obtain the platform performance parameters and neural network information; where the platform performance parameters include the computing power, communication conditions, and the distance between the platform and the server of the platform, and the neural network information includes the computational complexity, the number of parameters, and the feature map size between each layer;
[0048] Based on the platform performance parameters and neural network information, use a multi-objective optimization model to calculate the decision variables for each mobile platform; where the decision variables include the model cut-off point, communication bandwidth, transmission power, and the main frequency of the platform device, and the multi-objective optimization model determines the optimal cut-off point of the neural network and the load of each mobile platform in the inference task according to the different computing capabilities, communication conditions, and the distance from the server of the mobile platform;
[0049] Based on the decision variables, dynamically divide the neural network through a collaborative inference framework and perform corresponding parameter configuration;
[0050] Start the execution of the collaborative inference task through the mobile platform and the server; during the collaborative inference process, the mobile platform and the server continuously interact, and the mobile platform transmits some calculation results to the server, and the server executes the remaining complex calculations. The entire inference process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm.
[0051] Further, the processor is further configured to:
[0052] Obtain the platform performance parameters through the self-report of the platform or the query of the central controller;
[0053] Obtain the model structure of the neural network, and layer the model structure of the neural network to extract neural network information; among them, the model structure of the neural network is layered by the following method:
[0054] Ensure that there are no complex connections between two neural network layers to split the network into multiple sub-networks;
[0055] Multiple sub-networks run and train independently;
[0056] Use the distributed computing framework to allocate the sub-networks to different mobile platforms for processing.
[0057] The beneficial effects of the present invention are:
[0058] The present invention dynamically optimizes parameters such as computing resources, transmission bandwidth, and power of multiple platforms through the multi-objective optimization algorithm to ensure that the system reaches the optimal balance in terms of time, energy consumption, and resource utilization. Through the splitting of the neural network and task offloading, the mobile platform edge device and the server cooperate to complete the inference task, realizing the improvement of the system performance with high efficiency and low power consumption. The present invention is applicable to intelligent mobile terminals, Internet of Things devices, and other application scenarios that require efficient edge inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Shows a schematic diagram of an application scenario according to the prior art.
[0060] Figure 2 Shows a flowchart of a multi-objective optimization method for multi-user edge cloud collaborative inference of a mobile platform according to an embodiment of the present invention.
[0061] Figure 3 Shows the optimization principle diagram of the multi-objective optimization model according to an embodiment of the present invention.
[0062] Figure 4 Shows a flowchart of the multi-objective optimization algorithm of dynamic Gaussian sampling according to an embodiment of the present invention.
[0063] Figure 5Shows a schematic diagram of an out-of-bounds processing strategy according to an embodiment of the present invention.
[0064] Figure 6 Shows a structural diagram of a single-user collaborative inference framework according to an embodiment of the present invention.
[0065] Figure 7 Shows a structural diagram of a multi-user edge-cloud collaborative inference multi-objective optimization system for a mobile platform according to an embodiment of the present invention. Detailed implementation manners
[0066] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0067] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention.
[0068] Multi-user collaborative inference refers to the scenario of edge computing where multiple mobile platforms collaborate with an edge server to execute neural network inference tasks. Neural network tasks are usually difficult to complete on a single mobile platform due to their high computational complexity. To reduce the computational load and energy consumption of a single device, collaborative inference offloads some levels of the deep learning model to the edge server for execution, thereby improving the inference efficiency. Each platform dynamically selects to execute some tasks according to its own computing power and communication conditions, and the remaining tasks are completed by the edge server.
[0069] In actual application scenarios, multiple platforms need to participate in the inference task simultaneously. Since the computing resources and communication bandwidth of the edge server are limited, how to reasonably allocate resources among multiple platforms becomes a key issue. In addition, there are also differences in the computing power, bandwidth conditions, and communication power of different platforms, and a unified resource allocation strategy is difficult to meet the personalized needs in multi-user scenarios. Therefore, the embodiments of the present invention provide a multi-objective optimization method for multi-user edge-cloud collaborative inference of a mobile platform to solve the following problems:
[0070] 1) The inference problem with resource constraints: The computing resources and communication bandwidth of the mobile platform are limited and cannot complete complex deep learning inference tasks alone. It is necessary to improve the inference efficiency through reasonable task offloading and resource allocation.
[0071] 2) Multi-user task competition problem: When participating in inference tasks on multiple platforms simultaneously, the limited server computing resources and network bandwidth need to be reasonably allocated to avoid system performance degradation caused by resource contention.
[0072] 3) Balance between energy consumption and inference efficiency: The battery capacity of mobile platforms is limited. Therefore, the system needs to minimize the energy consumption of the platform while ensuring the timely completion of inference tasks.
[0073] Please refer to Figure 1 , which is a schematic diagram of an application scenario according to the prior art. In this application scenario, it includes four devices and a server. The four devices are Device 1, Device 2, Device 3, and Device 4 respectively. The server can act as a collaborative inference server. Figure 1 The dashed circles in
[0074] An embodiment of the present invention provides a multi-objective optimization method for multi-user edge-cloud collaborative inference on a mobile platform. Through the mobile platform multi-user edge collaborative inference technology based on model offloading, multi-objective optimization is achieved to maximize system efficiency. The system mainly consists of multiple mobile platforms and edge servers. Among them, the mobile platform is responsible for performing the preliminary processing of deep learning tasks and transferring the computationally intensive part to the edge server for further processing through model offloading technology. This multi-objective optimization method for multi-user edge-cloud collaborative inference on a mobile platform can be applied to an application scenario as Figure 1 shown. Figure 2 shows a flowchart of a multi-objective optimization method for multi-user edge-cloud collaborative inference on a mobile platform according to an embodiment of the present invention. As Figure 2 shown, this method includes steps S10 to S50, which are introduced in detail as follows.
[0075] S10: Multiple platforms join the edge collaborative inference system.
[0076] In this embodiment, multiple mobile platforms or edge device clients are detected within the communication radius of the multi-user collaborative inference server and automatically join the edge collaborative inference system. The communication radius refers to the effective transmission distance between the server and the mobile platform or edge device. Each platform establishes a communication link with the edge server and transmits the platform's performance parameters, such as the platform main frequency, available computing resources, bandwidth, etc. The server needs to monitor the communication status of the platform in a timely manner to ensure its smooth joining into the system. The main task of this stage is to provide the basic information of the platform-side resources for the subsequent allocation of inference tasks.
[0077] S20: Obtain platform performance parameters and neural network information.
[0078] In this embodiment, once a mobile platform or an edge device joins the system, the server begins to collect various performance parameters of the platform and specific information on network communication. These parameters include the computing power of the platform, communication conditions, the distance between the platform and the server, etc. The communication distance between the platform and the server affects the latency and rate of data transmission and is an important indicator for dynamic optimization. The server also obtains the model structure of the deep neural network, including information such as the computational complexity, the number of parameters, and the size of the feature maps between layers. This information provides the basic data support for the segmentation and offloading strategies of neural network tasks.
[0079] In some embodiments, the platform performance parameters and neural network information are obtained by the following method:
[0080] First, it is necessary to query the platform list of the server to obtain the detailed performance parameters of each platform in the system. These parameters include the number of platforms, the channel gain of wireless transmission, the floating-point operations per second (FLOPs) of the platform, the maximum bandwidth (Bmax) of the platform, the maximum power (Pmax), and the maximum frequency (Fmax), etc. These parameters are obtained through the self-report of the platform or the query of the central controller and are used for subsequent model layering and multi-objective optimization calculations.
[0081] The main methods for neural network layering are as follows:
[0082] No complex connections between neural network layers or blocks: Ensure that there are no complex connections between two neural network layers, which can facilitate the segmentation of the network into multiple sub-networks. For example, multi-layer fully connected neural networks and convolutional neural networks.
[0083] Modular design: Each segmented sub-network can operate and train independently. This method helps to improve the maintainability and scalability of the neural network model.
[0084] Distributed computing: Utilize the distributed computing framework to allocate different parts of the neural network to different mobile platforms for processing. This method can significantly improve the computing efficiency, especially when dealing with large-scale data.
[0085] Through the above methods, the neural network structure and parameters can be represented as a sequence, with each element being a neural network block. The segmentation points are between the elements.
[0086] S30: The multi-objective optimization algorithm solves the Pareto solution set corresponding to the optimization objective.
[0087] After obtaining the performance parameters of the platform and the neural network structure information, the server uses the multi-objective optimization model to calculate the decision variables for each platform. The decision variables include: model segmentation point (C), communication bandwidth (B), transmission power (P), and the main frequency of the platform device (F).
[0088] The multi-objective optimization model combines the inference time and energy consumption objectives of multiple platforms to form a multi-objective optimization problem. Its main purpose is to reasonably allocate resources among different platforms to ensure that the time and energy consumption during the inference process reach Pareto optimality. Specifically, the multi-objective optimization model will determine the optimal splitting point of the neural network and the optimal resource allocation plan for each platform in the inference task according to the different computing capabilities, communication conditions of the platforms, and the distance from the server.
[0089] In some embodiments, please refer to Figure 3 , which is the optimization principle diagram of the multi-objective optimization model provided by the embodiments of the present invention. The multi-objective optimization model takes time (T) and energy (E) as optimization objectives. The time consumption mainly includes the computing time (T_dc) and transmission time (T_tx) at the platform end, and the calculation formulas are based on the computing capabilities of the platforms and the bandwidth conditions of wireless transmission respectively. The energy consumption includes the computing energy consumption (E_dc) and transmission energy consumption (E_tx) at the platform end, which also depend on the energy consumption efficiency and transmission power of the platforms.
[0090] In the scenario of multi-user collaborative inference, the model splitting point (C), platform device bandwidth (B), platform device power (P), and platform device frequency (F) are determined as decision variables affecting the performance of collaborative inference. Based on these variables, this embodiment further elaborates on how to construct a multi-objective optimization model from two dimensions of time consumption and energy consumption.
[0091] Specifically, by establishing a time model and an energy consumption model, and then combining the time model and the energy consumption model, the multi-objective optimization model can be obtained.
[0092] For the time model, it is expressed as:
[0093] ;
[0094] In the formula, represents the total delay of collaborative inference, represents the computing delay of the platform, represents the transmission delay of the platform.
[0095] ;
[0096] In the formula, represents the model splitting point of this platform, . represents the th layer of this platform's data floating-point operation count. represents the cycle floating-point operation count. Represents the floating-point operations per second of the platform, which can be obtained by multiplying the platform device frequency by the floating-point operations per cycle (FPC).
[0097] ;
[0098] In the formula, Represents the data volume size of the output feature variables at the model segmentation point of the platform, Represents the transmission rate;
[0099] ;
[0100] In the formula, 、 Represent the bandwidth and power of the platform respectively, Represents the noise power spectrum, Represents the average channel gain.
[0101] h represents the average channel gain. ω refers to the weight of the channel gain.
[0102] ;
[0103] In the formula, Represents the weight of the channel gain, 、 、 and Represent the radio propagation speed, antenna gain, carrier frequency and path loss exponent respectively, Represents the distance between the platform and the server.
[0104] For the energy model, it is expressed as:
[0105] ;
[0106] In the formula, Represents the total energy consumption of collaborative inference, Represents the computing energy consumption of the platform data. Represents the transmission energy consumption of the platform intermediate feature variables transmitted to the server.
[0107] In the energy model, the data computing energy consumption of the edge platform takes into account the influence of the chip energy efficiency coefficient, chip main frequency and computing delay on the data computing energy consumption. The calculation method is shown in the following formula:
[0108] ;
[0109] In the formula, Is the processor chip energy efficiency coefficient, which refers to the energy consumed by the chip when executing a unit of computing task and reflects the energy efficiency level of the chip. f DThe chip main frequency of a certain platform refers to the frequency at which the chip executes instructions, which directly affects the computing speed of the chip.
[0110] The transmission energy consumption of the edge platform involves transmission power and the transmission of intermediate characteristic variables. When designing a user collaborative inference system, it is necessary to reasonably select the transmission scheme and optimize the transmission strategy to reduce the transmission energy overhead:
[0111] ;
[0112] In the formula, P represents the power of the platform, with the unit of watt (W), that is, joule per second (j / s).
[0113] By comprehensively considering the two objectives of time and energy consumption, the following multi-objective optimization model can be finally established:
[0114] ;
[0115] In the formula, represents the preference of the decision maker. When is greater than 0.5, it indicates that the decision maker has a higher requirement for the real-time performance of the system. On the contrary, it means that the decision maker hopes for less energy consumption. represents the decision variable corresponding to the minimum objective function value. represents the number of mobile platforms. represents the total collaborative inference delay of the th mobile platform.
[0116] In some embodiments, in order to solve the above multi-objective optimization model, a multi-objective optimization algorithm based on dynamic Gaussian sampling is proposed. In the process of algorithm design, in order to improve the search efficiency and the quality of the solution, normalization processing, boundary rebound strategy and differential solution update strategy are introduced. Normalization processing ensures that decision variables in different dimensions are considered on the same scale. The boundary rebound strategy prevents the solution from exceeding the predetermined search space, while the differential solution update strategy enhances the global search ability of the algorithm.
[0117] The multi-objective optimization algorithm based on dynamic Gaussian sampling generates multiple candidate solutions by performing Gaussian sampling in the solution space. Each candidate solution represents a possible configuration, including cut-off points, bandwidth allocation, power consumption and operation frequency. The algorithm selects the optimal configuration by evaluating the time consumption and energy consumption of each candidate solution. Specifically, as Figure 4As shown. This algorithm divides the overall goal of the system into a two-goal problem and uses a multi-objective optimization problem based on dynamic Gaussian sampling. In the algorithm, 100 individuals are randomly initialized as a population. In each population evolution, multiple neighbors of the individual are referred to and Gaussian mutation is performed, and the Chebyshev distance between the result of Gaussian sampling and the original individual is compared. If the result of Gaussian mutation is better than the original individual, then the value of Gaussian mutation is used as the new solution and the Pareto front is updated. After several rounds of sampling, the population will gradually move to the Pareto front surface under the action of two goals. At this time, the decision maker can select the optimal solution according to his own preference. Finally, the overall time consumption and energy consumption of the system reach the minimum, which means that the efficiency of the system is improved.
[0118] During the optimization process, since the solution update strategy adopts the Gaussian mutation algorithm, the out-of-bounds handling strategy as shown in Figure 5 is adopted for the solution. At the same time, in order to accelerate the convergence speed of the algorithm, when optimizing a certain solution, the neighbors with worse performance will be updated at the same time, and the mean point is used to replace the poor solution.
[0119] After the optimization is completed, the algorithm will generate multiple non-dominated solutions to form the Pareto front. The solutions on the Pareto front represent multiple optimal solutions between time and energy consumption, that is, corresponding to the α in the multi-objective optimization model, for the user to select the optimal configuration according to actual needs.
[0120] In a certain embodiment, mathematical libraries such as Python, PyTorch, and NumPy can be used to construct the above multi-objective optimization model and optimization algorithm.
[0121] S40: Perform dynamic partitioning and parameter configuration according to user preferences.
[0122] After calculating the Pareto solution set of the optimization plan, according to the user's personal preference, a suitable optimization plan is selected from the Pareto solution set. The server dynamically partitions the neural network according to the optimization plan through the collaborative inference framework and performs corresponding parameter configuration. According to the performance and communication conditions of the platform, the system offloads some computing tasks of the neural network to the edge server for execution, and the remaining part is completed locally by the mobile platform. This process ensures that the computing load of each platform matches its capabilities and avoids overloading a single platform.
[0123] Neural network task partitioning: According to the differences in platform performance and communication conditions, the server will select to execute some layers of the neural network on the platform and offload the more complex layers to the server for processing. For example, for a platform with relatively rich computing resources, more neural network layers can be allocated for local processing, while for a resource-constrained platform, only a small number of shallow networks are processed, and the rest are processed by the server.
[0124] Parameter configuration and optimization: Based on the segmentation task, the system configures each platform according to parameters such as bandwidth, power, and main frequency obtained by the optimization algorithm. The server dynamically adjusts the bandwidth allocation and transmission power of each platform to improve the efficiency of data transmission. At the same time, the main frequency of the platform is also appropriately adjusted according to the task load to ensure the best balance between energy consumption and computing efficiency.
[0125] In some embodiments, the collaborative inference framework includes a server side and a client side.
[0126] The server side mainly has three functions:
[0127] Record and manage platform information and model library in the system: The server side is responsible for recording and managing platform information and model library in the system, including performance parameters of the platform, structure and parameters of the model, etc.
[0128] Run a multi-objective optimization algorithm based on dynamic Gaussian sampling: The server side runs a multi-objective optimization algorithm based on dynamic Gaussian sampling to calculate the optimal solution of the decision variables of each platform at each moment.
[0129] Complete system inference: The server side receives intermediate feature variables, completes the inference, and returns the result to the client side.
[0130] The client side has two functions:
[0131] Synchronize platform information and receive decision variables distributed by the server: The client side is responsible for synchronizing platform information and receiving decision variables distributed by the server.
[0132] System inference: The client side performs system inference, sends intermediate feature variables, receives the inference result from the server side, and performs the next operation.
[0133] Through the realization of the above functions, the collaborative inference framework of the present invention can implement the multi-user edge collaborative inference technology based on model offloading for mobile platforms, maximize the system efficiency, and meet the requirements of multi-objective optimization.
[0134] In another embodiment, to solve the problem of insufficient computing resources of the edge platform, this embodiment proposes a single-user collaborative inference framework. Please refer to Figure Figure 6 , which is a structural diagram of a single-user collaborative inference framework provided by an embodiment of the present invention. As a collaborative inference framework, taking the weed recognition task as an example, the single-user collaborative inference framework divides the weed recognition task into two parts, one part is executed on the edge platform, and the other part is executed on the cloud server, so as to realize collaborative inference between the edge platform and the cloud server. The collaborative inference framework includes the following steps:
[0135] S41. Model Splitting. Based on the current system resource situation, select an appropriate model splitting point to split the deep neural network model into two parts, denoted as M E and M S . M E is the part of the model running on the edge platform, and M S is the part of the model running on the cloud server. The selection of the model splitting point needs to balance the computing time and transmission time of the edge platform to achieve the optimal collaborative inference efficiency.
[0136] S42. Edge Platform Inference. After receiving the image data, the edge platform uses M E to perform forward inference on the image to obtain the intermediate feature variable IF. After compressing and encoding IF, the edge platform transmits it to the cloud server via a wireless network.
[0137] S43. Cloud Server Inference. After receiving IF, the cloud server decodes and restores it, and then uses M S to perform forward inference to obtain the final detection result PR. PR contains the category and location information of the weed targets in the image. The cloud server transmits it back to the edge platform to complete the collaborative inference process.
[0138] S50: Execution of Platform-Server Collaborative Inference.
[0139] After the resource allocation and parameter configuration in the aforementioned steps S10 to S40, the platform and the server start to execute the collaborative inference task. During the collaborative inference process, the mobile platform and the edge server continuously interact. The mobile platform transmits some calculation results to the server, and the server performs the remaining complex calculations. The execution process of the collaborative inference needs to use an efficient data transmission protocol to ensure that the data exchange between the platform and the server is not affected by factors such as network bandwidth and latency. The entire inference process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm to minimize the system energy consumption on the premise of meeting the real-time requirements of the task.
[0140] The server continuously monitors the running status and communication situation of the platform. If a certain platform shows performance fluctuations or network instability, the system will automatically adjust the task allocation or recalculate the optimization parameters to ensure the stability and reliability of the inference process.
[0141] In summary, the method provided by the present invention combines the performance parameters of the mobile platform and the network communication conditions, and adopts a multi-objective optimization algorithm to achieve dynamic segmentation and load balancing of neural network inference tasks. Compared with the traditional centralized inference method, the present invention significantly reduces the total energy consumption of the system, improves the inference efficiency and the resource utilization rate of the platform. By optimizing the bandwidth, power, and main frequency configuration of the platform, it is ensured that the collaborative inference task between the platform and the server can achieve the optimal in terms of time and energy consumption. In addition, the multi-objective optimization algorithm of the present technology effectively solves the resource contention problem when multiple platforms participate in the inference task simultaneously, and improves the overall inference performance of the system. Specifically, it is reflected in the following aspects:
[0142] First, the model segmentation point: determines which part of the neural network each platform executes. In edge collaborative inference, the neural network is usually divided into different layers, and some layers are executed by the mobile platform, while the remaining parts are offloaded to the server for execution. By reasonably selecting the segmentation point, it is ensured that while ensuring the inference accuracy, the overall energy consumption and time consumption of the system are minimized.
[0143] Second, the communication bandwidth: In collaborative inference, data transmission between the platform and the server requires network bandwidth. Optimizing the bandwidth allocation can effectively reduce the transmission delay and ensure efficient data interaction. The server dynamically allocates bandwidth resources according to the communication conditions of each platform to maximize the overall transmission efficiency of the system.
[0144] Third, the transmission power: The transmission power determines the signal strength and data transmission rate between the platform and the server. In the case of a long transmission distance, appropriately increasing the transmission power can enhance the signal quality and reduce the data transmission delay. However, the increase in power will also lead to an increase in energy consumption. Therefore, in the optimization process, the power parameter needs to be balanced with the energy consumption target.
[0145] Fourth, the main frequency of the platform device: The main frequency of the platform device determines its computing power. The higher the main frequency, the faster the computing speed of the platform, but the corresponding energy consumption will also increase. By dynamically adjusting the main frequency of the platform, the inference efficiency and energy consumption can be balanced. The multi-objective optimization algorithm dynamically adjusts the main frequency according to the computing task load of different platforms to achieve the best inference performance.
[0146] The embodiment of the present invention also provides a multi-objective optimization system for multi-user edge cloud collaborative inference of a mobile platform. Please refer to Figure 7 . This system includes a processor 701, and the processor 701 is configured to:
[0147] When multiple mobile platforms or edge device clients are detected within the communication radius of the multi-user collaborative inference server, the mobile platforms or edge devices are automatically added to the edge collaborative inference system; wherein, the communication radius refers to the effective transmission distance between the server and the mobile platform or edge device.
[0148] When the mobile platform or edge device is added to the edge collaborative inference system, platform performance parameters and neural network information are obtained; wherein, the platform performance parameters include the computing power, communication conditions, and the distance between the platform and the server of the platform, and the neural network information includes the computational complexity, the number of parameters, and the feature map sizes between layers of each layer.
[0149] Based on the platform performance parameters and neural network information, using a multi-objective optimization model, the decision variables of each mobile platform are calculated; wherein, the decision variables include the model splitting point, communication bandwidth, transmission power, and the main frequency of the platform device, and the multi-objective optimization model determines the optimal splitting point of the neural network and the load of each mobile platform in the inference task according to the different computing powers, communication conditions, and the distance from the server of the mobile platform.
[0150] Based on the decision variables, the neural network is dynamically split through a collaborative inference framework, and corresponding parameter configurations are performed.
[0151] The execution of the collaborative inference task is started through the mobile platform and the server; during the collaborative inference process, the mobile platform and the server continuously interact, and the mobile platform transmits some calculation results to the server, and the server executes the remaining complex calculations, and the entire inference process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm.
[0152] In some embodiments, the processor 701 is further configured to:
[0153] Obtain platform performance parameters through the self-report of the platform or the query of the central controller.
[0154] Obtain the model structure of the neural network, and layer the model structure of the neural network to extract neural network information; wherein, the model structure of the neural network is layered by the following method:
[0155] Ensure that there are no complex connections between two neural network layers to split the network into multiple sub-networks.
[0156] Multiple sub-networks run and train independently.
[0157] Use a distributed computing framework to allocate the sub-networks to different mobile platforms for processing.
[0158] It should be noted that the multi-user edge cloud collaborative inference multi-objective optimization system for the mobile platform belongs to the same technical concept as the previously described method, and has the same technical principle and beneficial effects, so it will not be elaborated here.
[0159] The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Those of ordinary skill in the relevant technical fields can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention, and the patent protection scope of the present invention should be defined by the claims.
Claims
1. A multi-objective optimization method for multi-user edge-end cloud collaborative reasoning on a mobile platform, characterized in that: The method comprises: When clients of multiple mobile platforms or edge devices are detected within the communication radius of the multi-user collaborative reasoning server, the mobile platforms or edge devices are automatically added to the edge collaborative reasoning system; wherein the communication radius refers to the effective transmission distance between the server and the mobile platforms or edge devices; When a mobile platform or edge device is added to the edge collaborative reasoning system, platform performance parameters and neural network information are obtained; wherein the platform performance parameters include the computing power of the platform, communication conditions, and the distance between the platform and the server; and the neural network information includes the computational complexity of each layer, the amount of parameters, and the size of the feature graph between each layer; Based on the platform performance parameters and neural network information, and according to the different computing capabilities, communication conditions and distances between the mobile platforms and the server, a multi-objective optimization model is used to determine the optimal Pareto solution set of the neural network decision variables; wherein the decision variables include the model split point, communication bandwidth, transmission power and platform device main frequency; Based on the decision variables, the neural network is dynamically segmented through a collaborative reasoning framework, and corresponding parameter configuration is performed; The mobile platform and the server start to perform the collaborative reasoning task. During the collaborative reasoning process, the mobile platform and the server continue to interact. The mobile platform transmits part of the calculation results to the server, and the server performs the remaining complex calculations. The entire reasoning process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm. The multi-objective optimization model is constructed by the following method: Taking time consumption and energy consumption as optimization targets, a time model and an energy consumption model are established respectively; wherein the time consumption includes the computing time and transmission time of the platform end, which are calculated according to the computing power of the mobile platform and the bandwidth conditions of the wireless transmission respectively, and the energy consumption includes the computing energy consumption and transmission energy consumption of the mobile platform, which are calculated according to the energy consumption efficiency and transmission power of the mobile platform respectively; Based on the time model and the energy consumption model, construct the multi-objective optimization model; Based on the platform performance parameters and the neural network information, the decision variables of each mobile platform are calculated using a multi-objective optimization model, including: The multi-objective optimization model is solved using a multi-objective optimization algorithm based on dynamic Gaussian sampling. The multi-objective optimization algorithm based on dynamic Gaussian sampling generates multiple candidate solutions by performing Gaussian sampling in the solution space. Each candidate solution represents a possible configuration, including a split point, bandwidth allocation, power consumption, and operating frequency. The optimal configuration is selected by evaluating the time consumption and energy consumption of each candidate solution.
2. The method according to claim 1, characterized in that When a mobile platform or edge device is added to the edge collaborative inference system, platform performance parameters and neural network information are obtained, including: Obtain platform performance parameters through platform self-reporting or central controller query; The model structure of the neural network is obtained, and the model structure of the neural network is layered to extract the neural network information; wherein the model structure of the neural network is layered by the following method: Make sure there are no complex connections between two neural network layers to split the network into multiple sub-networks; Multiple sub-networks run and train independently; Using the distributed computing framework, the sub-networks are assigned to different mobile platforms for processing.
3. The method according to claim 1, characterized in that The collaborative reasoning framework includes an edge platform and a cloud server. Based on the decision variables, the neural network is dynamically segmented through the collaborative reasoning framework, and corresponding parameter configuration is performed, including: According to the current system resource situation, select the model split point and divide the neural network model into two parts, namely M E and M S , M E It is the part of the model that runs on the edge platform. S It is the part of the model that runs on the cloud server; After the edge platform receives the image data, it uses M E Perform forward reasoning on the image to obtain the intermediate feature variable IF. The edge platform compresses and encodes IF and transmits it to the cloud server via the wireless network. After receiving the IF, the cloud server decodes and restores it using M S Perform forward reasoning to obtain the detection result PR, where PR contains the category and location information of the target in the image; the cloud server transmits the detection result PR back to the edge platform.
4. The method according to claim 1, characterized in that The time model is expressed as: ; In the formula, represents the total latency of collaborative reasoning, represents the computing latency of the platform, Indicates the transmission delay of the platform; ; In the formula, Indicates the model splitting point of the platform, , Indicates the platform The number of floating point operations of the layer data, Indicates the number of floating-point operations per second of the platform, which is obtained by multiplying the platform device frequency by the number of periodic floating-point operations. Indicates the number of floating-point operations in a cycle. Indicates the device main frequency of the mobile platform; ; In the formula, Indicates the data size of the output feature variables at the model splitting point of the platform, Indicates the transmission rate; ; In the formula, , Represent the bandwidth and power of the platform respectively, represents the noise power spectrum, represents the average channel gain; ; In the formula, represents the weight of the channel gain, , , and represent radio propagation speed, antenna gain, carrier frequency and path loss index respectively, Indicates the distance between the platform and the server; The energy consumption model is expressed as: ; In the formula, represents the total energy consumption of collaborative reasoning, represents the computing energy consumption of platform data, represents the transmission energy consumption of the platform intermediate characteristic variables to the server, Represents the processor chip energy efficiency coefficient, Indicates the chip main frequency of the platform. Indicates the power of the platform.
5. A mobile platform multi-user edge-end cloud collaborative reasoning multi-objective optimization system, characterized in that: The system includes a processor configured to: When clients of multiple mobile platforms or edge devices are detected within the communication radius of the multi-user collaborative reasoning server, the mobile platforms or edge devices are automatically added to the edge collaborative reasoning system; wherein the communication radius refers to the effective transmission distance between the server and the mobile platforms or edge devices; When a mobile platform or edge device is added to the edge collaborative reasoning system, platform performance parameters and neural network information are obtained; wherein the platform performance parameters include the computing power of the platform, communication conditions, and the distance between the platform and the server; and the neural network information includes the computational complexity of each layer, the amount of parameters, and the size of the feature graph between each layer; Based on the platform performance parameters and the neural network information, the decision variables of each mobile platform are calculated using a multi-objective optimization model; wherein the decision variables include a model split point, a communication bandwidth, a transmission power, and a main frequency of a platform device, and the multi-objective optimization model determines the optimal split point of the neural network and the load of each mobile platform in the reasoning task according to the different computing capabilities, communication conditions, and distances between the mobile platforms and the server; Based on the decision variables, the neural network is dynamically segmented through a collaborative reasoning framework, and corresponding parameter configuration is performed; The collaborative reasoning task is started through the mobile platform and the server. During the collaborative reasoning process, the mobile platform and the server continuously interact. The mobile platform transmits part of the calculation results to the server, and the server performs the remaining complex calculations. The entire reasoning process is carried out according to the Pareto optimal conditions set by the multi-objective optimization algorithm. The multi-objective optimization model is constructed by the following method: Taking time consumption and energy consumption as optimization targets, a time model and an energy consumption model are established respectively; wherein the time consumption includes the computing time and transmission time of the platform end, which are calculated according to the computing power of the mobile platform and the bandwidth conditions of the wireless transmission respectively, and the energy consumption includes the computing energy consumption and transmission energy consumption of the mobile platform, which are calculated according to the energy consumption efficiency and transmission power of the mobile platform respectively; Based on the time model and the energy consumption model, construct the multi-objective optimization model; Based on the platform performance parameters and the neural network information, the decision variables of each mobile platform are calculated using a multi-objective optimization model, including: The multi-objective optimization model is solved using a multi-objective optimization algorithm based on dynamic Gaussian sampling. The multi-objective optimization algorithm based on dynamic Gaussian sampling generates multiple candidate solutions by performing Gaussian sampling in the solution space. Each candidate solution represents a possible configuration, including a split point, bandwidth allocation, power consumption, and operating frequency. The optimal configuration is selected by evaluating the time consumption and energy consumption of each candidate solution.
6. The system according to claim 5, characterized in that The processor is further configured to: Obtain platform performance parameters through platform self-reporting or central controller query; The model structure of the neural network is obtained, and the model structure of the neural network is layered to extract the neural network information; wherein the model structure of the neural network is layered by the following method: Make sure there are no complex connections between two neural network layers to split the network into multiple sub-networks; Design the neural network into multiple modules, each of which runs and is trained independently; Using the distributed computing framework, different parts of the neural network are assigned to different computing nodes for processing.
Citation Information
Patent Citations
Edge computing task unloading and resource allocation method based on deep learning
CN114585006A
Space-air-ground communication resource multi-objective optimization method based on Pareto optimal solution set
CN116915315A
Multi-device multi-neural network application splitting reasoning method for energy consumption optimization in edge-end environment
CN118964006A