Edge computing-based collaborative optimization method and system, device, and medium

By building a deep neural network directed acyclic graph and parameter dictionary, combining server allocation, model division and data batch algorithms, edge server resource allocation is optimized, communication overhead and latency problems in multi-service and multi-user scenarios are solved, and the system throughput and availability is improved.

WO2025138355A1PCT designated stage expired Publication Date: 2025-07-03SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072287
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-01-15
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the edge computing scenario of multi-service and multi-user, the existing technology lacks effective collaborative optimization methods, and cannot efficiently divide server resources, model resources and data batches on edge servers with limited resources, resulting in large communication overhead, system delay and insufficient throughput.

Method used

By building a directional acyclic graph and parameter dictionary in deep neural networks, combining server allocation algorithms, model division algorithms and data batch algorithms, we optimize resource allocation and model division of edge servers, reduce data transmission overhead, and enhance multi-service collaborative reasoning capabilities.

Benefits of technology

It realizes fairness in resource allocation under system resource limitations, minimizes service latency, makes full use of parallel processing potential, and improves system availability and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072287_03072025_PF_FP_ABST
    Figure CN2024072287_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are an edge computing-based collaborative optimization method and system, a device, and a medium. Module parameters of sub-modules can be directly forwarded to the sub-modules by means of a parameter dictionary, achieving separation of the parameters and a deep neural network directed acyclic graph, and reducing data transmission overhead. Then, a server allocation algorithm, a model partitioning algorithm, and a data batch processing algorithm are sequentially executed in a loop manner by means of a deep neural network model, so that the effects of respective algorithm results can be collaboratively enhanced, and multi-service collaborative reasoning is comprehensively supported. Taken independently, the server allocation algorithm ensures the fairness of resource allocation under the constraints of system resources; the model partitioning algorithm minimizes service delay; and the data batch processing algorithm makes full use of parallel processing potential and controls introduced communication overhead. After loop execution, service objectives of the three algorithms are maximized by means of the maximum feasible number of batches to obtain the optimal solutions of the three algorithms, improving the overall availability and throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Collaborative optimization method, system, equipment and medium based on edge computing Technical Field

[0001] The present invention relates to the field of edge computing technology, and in particular to a collaborative optimization method, system, device and medium based on edge computing. Background Art

[0002] Edge AI is the combination of AI and edge computing, aiming to bring intelligence to heterogeneous edge devices, thereby achieving ubiquitous intelligence. Edge AI brings model-based services to users, but the challenge is that edge servers are typically resource-constrained. For multi-service, multi-user collaborative processing in edge AI, model partitioning is part of a collaborative inference approach that leverages multiple edge servers (each hosting a subset of a complete model) to perform distributed model inference at the edge. While this approach increases complexity and communication overhead, it enables scalable deployment and execution of large models on resource-constrained edge servers and maximizes server utilization. Multi-service refers to the fact that edge frameworks typically deploy multiple DNN models, each serving user end devices. For example, a CAV (connected autonomous vehicle) and road infrastructure can be considered an edge framework, with the CAV's cameras acting as end devices and the RSUs (roadside units) acting as edge servers. A hypothetical service (also known as a decision maker) proposes a policy that determines whether the CAV should stop based on the traffic image streams collected by the CAV's cameras. Two object detection services are provided to identify traffic lights and pedestrians in image streams, respectively. Multi-user refers to the fact that different user end devices generate fewer requests. When multiple requests utilize the same model-based service, the data from these requests can be combined as a data batch to achieve a single batch inference. Choosing an appropriate batch size not only prevents the model from being misled by outliers but also accelerates the inference process with the hardware support of GPUs (graphics processing units) and NPUs (neural network processing units).

[0003] Therefore, in the collaborative processing scenario of multiple services and multiple users, there is currently no practical solution for the collaborative optimization of server resource partitioning, model partitioning, and data batch partitioning. This solution combines the above scenarios into a whole and explores the potential for improving system latency and throughput.

[0004] Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a collaborative optimization method, system, device, and medium based on edge computing, which can synergistically enhance the solutions for server resource partitioning, model partitioning, and data batch partitioning, reduce communication overhead, increase throughput, and improve edge computing performance.

[0006] In a first aspect, an embodiment of the present invention provides a collaborative optimization method based on edge computing, comprising:

[0007] Obtain local configuration files, web UI data, local model parameters, and client data;

[0008] Constructing a deep neural network directed acyclic graph according to the local configuration file and the web UI data, and constructing a module parameter dictionary according to the local model parameters;

[0009] A deep neural network model is constructed using the parameter dictionary and the deep neural network directed acyclic graph; the following steps are looped through the deep neural network model: a server allocation algorithm is executed by the deep neural network model according to the client data to obtain an edge server allocation result; in response to the edge server allocation result, a model partitioning algorithm is executed on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, a data batch processing algorithm is executed by the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; a server allocation management table in the server allocation algorithm is obtained, and if the server allocation management table does not contain an allocation task and the duration reaches a preset time threshold, the loop is stopped;

[0010] Obtaining module parameters corresponding to the submodule division result through the parameter dictionary in the deep neural network model, and distributing the module parameters to the working edge server corresponding to the edge server allocation result;

[0011] The service request of the client is responded to through the module parameters, the allocation result according to the edge server, the sub-module division result and the data batch division result.

[0012] The method according to the embodiment of the present invention has at least the following beneficial effects:

[0013] This method first constructs a deep neural network model through a parameter dictionary and a deep neural network directed acyclic graph. The parameter dictionary can be used to directly forward the module parameters of the sub-module to the sub-module, thereby realizing the separation of parameters from the deep neural network directed acyclic graph, reducing data transmission overhead and the overhead of subsequent server allocation algorithm, model partitioning algorithm and data batch processing algorithm. The subsequent algorithm only needs to use the deep neural network directed acyclic graph and does not need the module parameters of the sub-module; then the server allocation algorithm, model partitioning algorithm and data batch processing algorithm are executed in sequence through the deep neural network model loop, which can synergistically enhance the effects of their respective algorithm results and fully support multi-business collaborative reasoning. Taken separately, the server allocation algorithm ensures the fairness of resource allocation under the constraints of system resources, the model partitioning algorithm minimizes service delay, and the data batch processing algorithm fully utilizes the potential of parallel processing and controls the communication overhead introduced. After the loop execution, the service goals of the three algorithms are maximized by passing effective information such as the maximum feasible batch number, obtaining the optimal solution for the three algorithms and improving the overall availability and throughput.

[0014] According to some embodiments of the present invention, a forward result of the deep neural network directed acyclic graph is calculated by automatically inferring a model forward function, and the distribution of the module parameters is supervised by the forward result.

[0015] According to some embodiments of the present invention, executing a server allocation algorithm based on the client data using the deep neural network model to obtain an edge server allocation result includes:

[0016] Record the basic information of the registered edge servers through the server configuration table;

[0017] Obtaining a server deployment request from the client;

[0018] In response to the server deployment request, obtaining an edge server to be deployed corresponding to the server deployment request by calculating through the server allocation algorithm;

[0019] Obtaining a logical edge cluster according to the server configuration table, and selecting a cluster leader of the logical edge cluster;

[0020] The cluster leader guides the client to the leading edge server corresponding to the cluster leader, and performs the deployment transaction of the edge server to be deployed on the leading edge server to obtain the edge server allocation result.

[0021] According to some embodiments of the present invention, in response to the edge server allocation result, executing a model partitioning algorithm according to the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model includes:

[0022] Calculating a first inference delay record on each of the working edge servers by a preset first inference analyzer; the first inference delay record represents a plurality of first inference delays corresponding to different sub-module combinations deployed on the working edge server;

[0023] The first inference delay record is input into the model partitioning algorithm to obtain the sub-module partitioning result.

[0024] According to some embodiments of the present invention, in response to the submodule division result, executing a data batch processing algorithm through the deep neural network model to obtain a data batch division result corresponding to the submodule division result includes:

[0025] Obtaining request data from the client;

[0026] Obtain multiple data batches by performing different combinations of the request data;

[0027] Calculating multiple inference delays of each of the submodules under the multiple data batches by a preset second inference analyzer to obtain a second inference delay record of the second inference analyzer;

[0028] The second inference delay record is input into the data batch processing algorithm to calculate the data batch division result.

[0029] According to some embodiments of the present invention, the first reasoning profiler and the second reasoning profiler are both obtained by training a regression model, and the training step of the regression model includes:

[0030] Construct raw material data parameters and raw material module parameters;

[0031] Obtaining configuration parameters of the working edge server, and constructing an input sample using the raw material data parameters, the raw material module parameters, and the configuration parameters;

[0032] constructing pseudo data according to the raw material data parameters, and constructing a pseudo module model according to the raw material module parameters;

[0033] Obtaining a simulated inference delay corresponding to the input sample by measuring the pseudo data and the pseudo module model;

[0034] Obtaining a training set by dividing the simulated inference delay and the input samples;

[0035] The regression model is trained using the training set to obtain the first inference parser or the second inference parser.

[0036] According to some embodiments of the present invention, inputting the first inference delay record into the model partitioning algorithm to obtain the sub-module partitioning result includes:

[0037] Inputting the first inference delay record into the model partitioning algorithm to obtain a sub-module partitioning combination;

[0038] Obtaining a submodule replication strategy corresponding to the data batch partitioning result;

[0039] Constructing a module stub on the working edge server, and storing model parameters corresponding to the sub-module partitioning combination and the sub-module replication strategy through the module stub;

[0040] A distributed module model corresponding to the working edge server is constructed according to the module stub to obtain the sub-module division result.

[0041] In a second aspect, an embodiment of the present invention provides a collaborative optimization system based on edge computing, wherein the collaborative optimization system based on edge computing includes:

[0042] Data acquisition unit, used to obtain local configuration files, web UI data, local model parameters and client data;

[0043] A parameter dictionary and deep neural network directed acyclic graph construction unit, configured to construct a deep neural network directed acyclic graph according to the local configuration file and the web UI data, and to construct a module parameter dictionary according to the local model parameters;

[0044] A model deployment execution planning unit is configured to construct a deep neural network model using the parameter dictionary and the deep neural network directed acyclic graph; looping through the deep neural network model to execute the following steps: executing a server allocation algorithm based on the client data using the deep neural network model to obtain an edge server allocation result; in response to the edge server allocation result, executing a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, executing a data batch processing algorithm using the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; obtaining a server allocation management table in the server allocation algorithm, and stopping the loop if no allocation task exists in the server allocation management table;

[0045] A model deployment execution implementation unit, configured to obtain module parameters corresponding to the submodule division result through the parameter dictionary in the deep neural network model, and distribute the module parameters to the working edge server corresponding to the edge server allocation result;

[0046] The service request execution unit is configured to respond to the service request of the client through the module parameters, the edge server allocation result, the submodule division result, and the data batch division result.

[0047] In a third aspect, an embodiment of the present invention provides an electronic device comprising at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the edge computing-based collaborative optimization method as described in the first aspect.

[0048] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the edge computing-based collaborative optimization method as described in the first aspect.

[0049] It should be noted that the beneficial effects between the second to fourth aspects of the present invention and the prior art are the same as the beneficial effects of the collaborative optimization method based on edge computing in the first aspect, and will not be described in detail here.

[0050] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments with reference to the following drawings, in which:

[0052] FIG1 is a flow chart of a collaborative optimization method based on edge computing provided by one embodiment of the present invention;

[0053] FIG2 is a flow chart showing an edge server allocation result obtained by executing a server allocation algorithm based on client data using a deep neural network model according to an embodiment of the present invention;

[0054] FIG3 is a flowchart of an embodiment of the present invention providing a response to an edge server allocation result, executing a model partitioning algorithm on a deep neural network model, and obtaining a sub-module partitioning result corresponding to the deep neural network model;

[0055] FIG4 is a flowchart of a data batch partitioning result obtained by executing a data batch processing algorithm through a deep neural network model in response to a submodule partitioning result provided by an embodiment of the present invention;

[0056] FIG5 is a flowchart of training a regression model according to an embodiment of the present invention;

[0057] FIG6 is a flowchart of constructing and deploying a distributed module model according to an embodiment of the present invention;

[0058] FIG7 is a schematic diagram of the construction and use of an edge intelligence-based aggregated deep neural network model provided by one embodiment of the present invention;

[0059] FIG8 is a framework diagram of a collaborative optimization method based on edge intelligence provided by an embodiment of the present invention;

[0060] 9 is a flowchart of the server allocation manager according to an embodiment of the present invention;

[0061] 10 is a timing diagram of the server allocation manager working and interacting with other entities according to an embodiment of the present invention;

[0062] 11 is a flowchart of an inference profiler training a regression model according to an embodiment of the present invention;

[0063] FIG12 is a flowchart of a collaborative reasoning optimizer using a first reasoning profiler to obtain a first reasoning delay record according to an embodiment of the present invention;

[0064] 13 is a flowchart of a sorting center using a second reasoning analyzer to obtain a second reasoning delay record according to an embodiment of the present invention;

[0065] FIG14 is a schematic diagram of the initial deployment of an edge environment including a decision process X described in information interaction examples 1 and 2 according to an embodiment of the present invention;

[0066] FIG15 is a schematic diagram illustrating a deployment of a new decision process Y just added to an edge environment, as described in Information Interaction Example 1 provided by an embodiment of the present invention, and an adjustment plan made by the collaborative reasoning optimizer based on the maximum feasible batch number information provided by the sorting center;

[0067] FIG16 is a schematic diagram of the deployment of the collaborative reasoning optimizer described in the information interaction example 1 provided by an embodiment of the present invention after adjustments are made to the new decision process Y added to the edge environment;

[0068] FIG17 is a schematic diagram of a deployment situation in which a new decision-making process Y has just been added to an edge environment and an adjustment plan made by a server allocation manager based on the maximum feasible batch number provided by a sorting center, as described in information interaction example 2 provided by an embodiment of the present invention;

[0069] 18 is a schematic diagram of a deployment situation after the server allocation manager makes adjustments in response to the new decision process Y joining the edge environment as described in the second information interaction example provided by an embodiment of the present invention;

[0070] FIG19 is a schematic diagram of an input and output table of an algorithm for information exchange using the maximum feasible number of data batches provided by one embodiment of the present invention;

[0071] FIG20 is a structural diagram of a collaborative optimization system based on edge computing provided by an embodiment of the present invention;

[0072] FIG21 is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0073] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0074] In the description of the present invention, if there is a description of first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.

[0075] In the description of the present invention, it should be understood that descriptions involving orientation, such as the orientation or positional relationship indicated by up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0076] In the description of the present invention, it should be noted that, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.

[0077] The technical solutions of the present invention will be described clearly and completely below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of the present invention, not all embodiments.

[0078] 1 and 7 , in some embodiments of the present invention, a collaborative optimization method based on edge computing is provided, including:

[0079] Step S100: Obtain local configuration files, web UI data, local model parameters and client data.

[0080] Step S200: construct a deep neural network directed acyclic graph based on the local configuration file and web UI data, and construct a module parameter dictionary based on the local model parameters.

[0081] Step S300, construct a deep neural network model through a parameter dictionary and a deep neural network directed acyclic graph; loop through the deep neural network model to execute the following steps: execute a server allocation algorithm according to client data through the deep neural network model to obtain an edge server allocation result; in response to the edge server allocation result, execute a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, execute a data batch processing algorithm through the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; obtain a server allocation management table in the server allocation algorithm, and stop the loop if there is no allocation task in the server allocation management table and the duration reaches a preset time threshold.

[0082] It should be noted that in the loop, the information interaction of the three algorithms is affected step by step from bottom to top. For example, the maximum feasible batch number is used as the information interaction data, and the allocated server cluster determines the solution space of the model partition, further limiting the maximum feasible batch number of data batch processing. Therefore, when discussing algorithm interaction, reverse information transmission (i.e., from the bottom to the top) is generally considered. It should also be noted that the server allocation management table in the server allocation algorithm comes from the server allocation manager. All client task requests will be recorded in the server allocation management table of the server allocation manager. When there is no task request in the server allocation management table, it means that all client task requests have been completed, and the monitoring duration is monitored to see if it reaches the preset time threshold. If it reaches the threshold, the loop can be stopped.

[0083] Step S400: Obtain module parameters corresponding to the sub-module division result through the parameter dictionary in the deep neural network model, and distribute the module parameters to the working edge server corresponding to the edge server allocation result for model training to obtain sub-modules.

[0084] Step S500: respond to the client's service request through module parameters, edge server allocation results, sub-module division results, and data batch division results.

[0085] First, a deep neural network model is constructed through the parameter dictionary and the deep neural network directed acyclic graph. The parameter dictionary can be used to directly forward the module parameters of the sub-module to the sub-module, realizing the separation of parameters from the deep neural network directed acyclic graph, reducing data transmission overhead and the overhead of subsequent server allocation algorithm, model partitioning algorithm and data batch processing algorithm. The subsequent algorithms only need to use the deep neural network directed acyclic graph and do not need the module parameters of the sub-module; then, the server allocation algorithm, model partitioning algorithm and data batch processing algorithm are executed in sequence through the deep neural network model loop, which can synergistically enhance the effects of their respective algorithm results and fully support multi-business collaborative reasoning. Taken separately, the server allocation algorithm ensures the fairness of resource allocation under the constraints of system resources, the model partitioning algorithm minimizes service delay, and the data batch processing algorithm fully utilizes the potential of parallel processing and controls the communication overhead introduced. After the loop execution, the service goals of the three algorithms are maximized through the maximum feasible batch number to obtain the optimal solution of the three algorithms, thereby improving the overall availability and throughput.

[0086] In some embodiments of the present invention, the forward result of the deep neural network directed acyclic graph is calculated by automatically inferring the model forward function, and the distribution of module parameters is supervised by the forward result.

[0087] By automatically inferring the model forward function, the need for manual coding of traditional deep learning frameworks is eliminated. Instead, the deep neural network directed acyclic graph is derived in a general way, reducing the workload of workers.

[0088] 2 , in some embodiments of the present invention, a server allocation algorithm is executed based on client data using a deep neural network model to obtain edge server allocation results, including:

[0089] Step S311: Record basic information of the registered edge server through the server configuration table.

[0090] It should be noted that, referring to FIG9 , the server allocation algorithm is executed by the server allocation manager. Meanwhile, the server configuration table updates the registered edge servers through the following steps:

[0091] When a new edge server registers with its basic information, the server allocation manager verifies and responds with the credential token for re-identification. After the relevant events of the registered edge server are completed or migrated, the server allocation manager removes the edge server from the server configuration table according to the deregistration request. When the server allocation manager fails to receive a heartbeat reply from a certain edge server, it marks the edge server as inactive. After recovery, the edge server can use the provided credential token to try to reconnect.

[0092] Step S312: Obtain the server deployment request from the client.

[0093] Step S313: respond to the server deployment request and obtain the edge server to be deployed corresponding to the server deployment request through calculation using a server allocation algorithm.

[0094] It should be noted, referring to Figure 10 , that each time a server deployment request is received from a client, the server allocation manager runs an installed server allocation algorithm. This server allocation algorithm can be based on game theory, where clients are defined as players and edge servers are defined as resources. Each player provides a program to be deployed and an objective function to be maximized. The game theory algorithm aims to allocate limited edge servers to the program so that each player can maximize their objective function. Once the system's equilibrium state is found, a corresponding edge server to be deployed corresponding to the server deployment request is obtained. To evaluate the objective function, the server configurations of registered edge servers in the environment are collected. To this end, each edge server is installed with a configuration manager component that is responsible for collecting the required local server configurations and the network bandwidth to other servers, which can be obtained through a network performance analyzer. The configuration manager processes configuration request information from the server allocation manager on demand. In addition, the configuration manager also participates in subsequent model profiling records (obtained by the first inference profiler or the second inference profiler), which are used by both the model partitioning algorithm and the data batching algorithm. Therefore, during the cycle, the server allocation algorithm can benefit from other information, such as the program inference request rate from the sorting center. The program request rate is a quantity used to measure the access frequency to adjust the weight of the deployed programs and perform reallocation.

[0095] Step S314: Obtain a logical edge cluster according to the server configuration table, and select a cluster leader of the logical edge cluster.

[0096] Step S315: The cluster leader guides the client to the leading edge server corresponding to the cluster leader, and performs a deployment transaction of the edge server to be deployed on the leading edge server to obtain an edge server allocation result.

[0097] By deploying transactions through server configuration tables, logical edge clusters, and leading edge servers, we can unify the deployment order and improve deployment efficiency. In addition, the configuration manager collects impact information for the server allocation algorithm to ensure that the server allocation algorithm is continuously optimized in the cycle.

[0098] 3 , in some embodiments of the present invention, in response to the edge server allocation result, a model partitioning algorithm is executed on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model, including:

[0099] Step S321: Calculate a first inference delay record on each working edge server through a preset first inference analyzer; the first inference delay record represents a plurality of first inference delays corresponding to different sub-module combinations deployed on the working edge server.

[0100] It should be noted that the sub-module combination refers to cutting the deep neural network model through the deep neural network directed acyclic graph to obtain multiple sub-modules, and combining the sub-modules through different divisions to obtain the sub-module combination.

[0101] Step S322: Input the first inference delay record into the model partitioning algorithm to obtain a sub-module partitioning result.

[0102] It should be noted that, referring to FIG2, through steps S311 to S315, the client is allocated an edge cluster consisting of multiple edge servers, and communicates with the leading edge server to complete the sub-module allocation of the edge cluster. Specifically, the client registers the decision program in the model receiver, which is used to store the sub-modules in the local database and notify the collaborative reasoning optimizer to run the model partitioning algorithm. When running the model partitioning algorithm, the deep neural network directed acyclic graph is obtained from the local database, and the first reasoning delay record is obtained for model partitioning. Referring to FIG12, at this time, the first reasoning analyzer is located in the configuration manager of each working edge server, obtains the reasoning data of different sub-module combinations deployed in the configuration manager, and calculates the first reasoning delay record through the regression model in the first reasoning analyzer. At the same time, similar to the server allocation algorithm, the collaborative reasoning optimizer will also optimize and adjust based on the information provided by the sorting center when running the model partitioning algorithm. The information provided by the sorting center is, for example, the sub-module replication strategy of the sorting center.

[0103] The first inference analyzer is used to first analyze the delay of different sub-module combinations. The first inference analyzer provides an inference template that can perform inference based on the information brought by different sub-module combinations, adapting to various task scenarios. At the same time, it also avoids the actual delay measurement operation of the sub-module combination and reduces communication overhead.

[0104] 4 , in some embodiments of the present invention, in response to the submodule division result, a data batch processing algorithm is executed through a deep neural network model to obtain a data batch division result corresponding to the submodule division result, including:

[0105] Step S331: Obtain the client's request data.

[0106] Step S332: obtain multiple data batches by performing different combinations of the requested data.

[0107] Step S333: Calculate multiple inference delays of each submodule under multiple data batches through a preset second inference analyzer to obtain a second inference delay record of the second inference analyzer.

[0108] Step S334: Input the second inference delay record into the data batch processing algorithm to calculate the data batch division result.

[0109] It should be noted that the data batch partitioning algorithm is executed by the sorting center, which collects client request data through a request controller and forwards this request data to the sorting center. The sorting center runs the data batch processing algorithm, combining multiple requests into data batches to improve system throughput. Referring to Figure 13, the data batch is then sent to the second inference analyzer for execution using module stubs arranged as a distributed module model. This generates a second inference delay record for the second inference analyzer, which is then input into the data batch processing algorithm to calculate the data batch partitioning result. Finally, the sorting center performs partitioning based on the data batch partitioning result and responds to the user terminal device with the partitioned request results.

[0110] Similarly, by first performing inference delays on different data batches through the second inference analyzer, it can adapt to various task scenarios. At the same time, it can also reduce communication overhead by performing delay measurement on actual data batches.

[0111] 5 , in some embodiments of the present invention, the first reasoning analyzer and the second reasoning analyzer are both obtained through regression model training. The training steps of the regression model include:

[0112] Step S341: Construct raw material data parameters and raw material module parameters.

[0113] Step S342: Obtain configuration parameters of the working edge server, and construct an input sample through raw material data parameters, raw material module parameters and configuration parameters.

[0114] Step S343: construct pseudo data according to the raw material data parameters, and construct a pseudo module model according to the raw material module parameters.

[0115] Step S344: obtain the simulated inference delay corresponding to the input sample through the pseudo data and pseudo module model measurement.

[0116] Step S345: obtain a training set by simulating inference delay and input sample division.

[0117] Step S346: Train the regression model using the training set to obtain the first inference parser or the second inference parser.

[0118] It should be noted that the results of the first and second inference analyzers are identical. This is also intended to provide a universal model for inference latency prediction. For ease of explanation, both the first and second inference analyzers will be referred to as inference analyzers. Referring to Figure 11 , taking the first inference analyzer as an example, the inference analyzer consists of two parts: one for training the corresponding regression model input and the other for the cooking steps involved in training the corresponding regression model. The input for training the corresponding regression model specifically involves using the data parameter sampler and module parameter sampler in the recipe component to construct and sample the corresponding raw material data parameters and raw material module parameters. Then, under the guidance of the recipe component, the inference analyzer combines the raw material data parameters, raw material module parameters, and server configuration parameters to construct the input for the regression model. The cooking steps for training the corresponding regression model specifically involve constructing simulated data and simulated modules based on the raw material data parameters and raw material module parameters in the recipe. The simulated data and simulated modules are used to simulate and measure the corresponding simulated inference latency. The simulated inference latency of each recipe input and output is combined as a sample. The recipe then specifies how to further split the dataset into training, validation, and test sets. Finally, the inference profiler trains a regression model using the generated dataset and saves it to local storage.

[0119] Through the training of regression models, a unified way to generate inference profilers is provided, thereby generalizing the prediction service of model inference latency. It can promote modular and reconfigurable profiling workflows by allowing effortless updates of regression model types and the required data, modules, and server-related parameters used to train regression models. At the same time, by constructing pseudo data and pseudo models, the cost of actual operation measurement is reduced.

[0120] 6 , in some embodiments of the present invention, the first inference delay record is input into a model partitioning algorithm to obtain a submodule partitioning result, including:

[0121] Step S3341: Input the first inference delay record into the model partitioning algorithm to obtain a sub-module partitioning combination.

[0122] Step S3342: Obtain the sub-module replication strategy corresponding to the data batch division result.

[0123] It should be noted that the sub-module replication strategy is to determine the bottleneck module through the delay results of the sorting center monitoring. Assigning the bottleneck module to a more powerful working edge service can optimize the overhead. The bottleneck module determines the upper limit of the available batch size. Therefore, the sub-module replication strategy can be adjusted based on the introduction of the bottleneck module. The sub-module replication strategy is shared with the previous model partitioning algorithm to determine the model parameters for constructing the sub-module.

[0124] Step S3343: construct a module stub on the working edge server, and store model parameters corresponding to the sub-module division combination and the sub-module replication strategy through the module stub.

[0125] Step S3344: construct a distributed module model corresponding to the working edge server according to the module stub to obtain the sub-module division result.

[0126] It should be noted that the DNN partitioner obtains the parameters of the replication submodule corresponding to the submodule replication strategy and constructs the corresponding replication submodule package. The DNN partitioner then assigns the replication submodule package and the module model partition combination to the designated working edge server for execution. Afterwards, the DNN partitioner uses the module stub received from the working edge server to construct the distributed module model. The distributed module model in this embodiment has the same structure as the deep neural network model, which is equivalent to storing the parameters of the distributed module model in the module stub, while the parameters of the deep neural network model are stored in the parameter dictionary.

[0127] The submodule partitioning result, determined by the submodule replication strategy and the submodule partitioning combination, can maximize throughput. At the same time, building a distributed module model through module stubs also simplifies the construction of submodules, reducing computing power waste and communication overhead.

[0128] 8 , to facilitate understanding by those skilled in the art, a specific embodiment of a collaborative optimization method based on edge computing is provided, including:

[0129] The framework of the deep neural network model of this embodiment consists of two major modules: one is composed of a deep neural network directed acyclic graph, which represents the graph topology of the model through the deep neural network directed acyclic graph; the other is composed of a parameter dictionary, which stores the corresponding model parameters according to different modules. Specifically, the deep neural network directed acyclic graph is constructed by a web user interface or a predefined local YAML / JSON file, and the initial model parameters are loaded or initialized from local storage. The following operations can be performed using the deep neural network model:

[0130] (1) The server allocation manager, the initial component in the edge environment, is connected through a deep neural network model. The server allocation manager is used to record the configuration details of the registered edge servers. Therefore, the server allocation manager can maintain the server configuration table. The specific maintenance method is as follows:

[0131] When a new edge server registers with its basic information, the server allocation manager validates and responds with a credential token for re-identification. After the relevant event is completed or migration occurs, the server allocation manager removes the edge server from the server configuration table based on the deregistration request. When the server allocation manager fails to receive a heartbeat reply from a certain edge server, it marks the edge server as inactive. Upon recovery, the edge server can attempt to reconnect using the provided credential token. In addition to the server configuration table, the server allocation manager also maintains a server allocation map that specifies which set of edge servers is assigned to which decision-making program (i.e., the program that is deployed by the client for decision-making). Each time a program client registers a program deployment request, the server allocation manager runs the installed server allocation algorithm and responds with a server allocation plan. Upon client confirmation, a logical edge cluster is formed, an edge cluster leader is elected, and the server allocation plan is notified to the relevant edge servers. It then directs the client to the cluster leader edge server for subsequent deployment transactions.

[0132] The server allocation algorithm can be based on game theory, where application clients are considered players and edge servers are considered resources. Each player provides a program to be deployed and an objective function to be maximized. Essentially, the algorithm aims to allocate limited edge servers to programs so that each player can maximize their own objective function. A solution is achieved by finding the equilibrium state of the system. To evaluate the objective function, the server configurations of registered edge servers in the environment should be collected. To this end, each edge server is installed with a configuration manager component that collects the required local server configurations and the network bandwidth to other servers, as determined by a network performance profiler. The configuration manager processes configuration requests from the server allocation manager on demand. Furthermore, the configuration manager maintains model profile records, which are utilized by both the model partitioning algorithm and the data batching algorithm. Last but not least, the server allocation algorithm can benefit from additional information, such as the program inference request rate from the sorting center, to adjust the weights of deployed programs and perform reallocation.

[0133] (2) Now, the client's decision program is assigned an edge cluster consisting of multiple edge servers, and the client's decision program will communicate with the leader edge server of the cluster. Specifically, the program client registers its decision program with the model receiver, which stores the model in the local database and notifies the collaborative reasoning optimizer to run the model partitioning algorithm. The algorithm obtains the deep neural network directed acyclic graph from the model database and obtains the first inference latency record from the first analyzer manager, including the bandwidth table storing the data transmission bandwidth between the edge servers, the module recipe and the inference analyzer stub. The module recipe is initially provided by the recipe provider, which can be the client or a third-party registration agency. In essence, each module recipe specifies: i) the input (also called raw materials) for training the corresponding regression model, and ii) the cooking operation for training the corresponding regression model. These two steps are described in detail in the above embodiment and will not be repeated here. When the training is completed, the first inference analyzer returns its own RPC stub to the first analyzer manager, so that the regression model can be remotely called to provide latency prediction services. Figure 2 shows that the collaborative reasoning optimizer first uses the recipe in the first analyzer manager to process and cache the raw material module parameters. It then uses the first inference profiler stub to remotely construct the raw data parameters for a single data batch. Finally, using the raw data parameters, server configuration, and processed raw module parameters, it remotely invokes a regression model to predict inference latency and responds to the collaborative inference optimizer. The model partitioning algorithm outputs a partitioning scheme that groups the module set into submodules and recommends deploying the submodules to the corresponding assigned edge servers. Similarly, the algorithm can also benefit from additional information provided by the sorting center, such as the submodule replication strategy, to identify execution bottleneck submodules and adjust the partitioning scheme accordingly.

[0134] (3) Once the partitioning scheme is calculated, the model receiver forwards it to the program client as a response. After confirmation, the DNN partitioner executes the partitioning scheme using the deep neural network directed acyclic graph. Together with the submodule replication strategy from the sorting center, the DNN partitioner also obtains the replication module parameters of each submodule representing the partition replication and builds the corresponding replication submodule package. The DNN partitioner then distributes the replication submodule package to the execution runtime of the designated worker edge server. After that, the DNN partitioner uses the RPC module stub received from the worker edge server to build a distributed DNN model. The structure of the distributed DNN model is the same as the DNN model, except that the module dictionary stores the module stub instead of the module parameters. Finally, the distributed DNN model registers the built distributed DNN model with the request scheduler. Now, the setup phase of the decision program will be marked as completed, followed by the in-service phase. The client's decision program will track the deployment status through the server allocation manager and the leader edge server.

[0135] (4) In the in-service phase of the decision procedure, an inference request is sent from the user terminal device to the request controller, which is located in the execution manager of the leader edge server. The request controller processes and records the metadata of the request and then forwards this information to the sorting center. The sorting center runs the data batching algorithm, which specifies how to combine multiple requests into data batches, calculates the second inference latency record of the data batch through the second inference analyzer, and performs the data batching algorithm on the second inference latency record to obtain the data batch processing result. Finally, the sorting center divides the returned data batch processing result and outputs the divided request results to the request controller, which then responds to the user terminal device. As a supplementary note, executing the data batching algorithm requires the estimated inference latency of each module under different data batch sizes. Similarly, the construction of the second inference analyzer is the same as the first inference analyzer, except that the simulated data type is different. Last but not least, the sorting center monitors the actual execution latency results through the request scheduler to help determine the bottleneck submodule, which determines the upper limit of the available batch size. The sorting center then attempts to introduce replicas of the submodules and adjusts the submodule replication strategy, which is shared with the model acceptance and model partitioning algorithms. The sorting center also monitors the program request rate to help the server allocation manager improve the server allocation plan by providing more resources to busier programs.

[0136] (5) In essence, the algorithm interaction in this embodiment refers to the output of one algorithm being used as the input of another algorithm, thereby affecting the output of the latter. Referring to Figure 19, taking the maximum feasible batch number as the information interaction data as an example, in general, the first execution order of the three algorithms follows: (bottom layer) server allocation, submodule division, and data batch processing (top layer). It is not difficult to see that the output of the algorithm is affected from bottom to top: the allocated server cluster determines the solution space of the model division, which further limits the maximum feasible batch number of data batch processing. Therefore, when discussing algorithm interaction, the reverse (bottom layer to top layer) information transmission is generally considered.

[0137] For example, in the first step, referring to Figure 14 , within the assigned cluster X, decision X is divided into two submodules, X1 and X2, and distributed to the execution runtimes of worker edge servers A and B. In the execution runtime, the space containing the submodules represents the memory consumed by storing these submodules, while the remaining space represents the memory available for future request data storage. This also determines the maximum feasible data batches that submodule X1 can consume in real time (the larger the remaining space, the more data that can be stored, and the larger the maximum feasible data batches). This, in turn, determines the maximum feasible batch size for decision process X, thus affecting the final system throughput. In the second step, referring to Figures 15 and 16 , building on the deployment in the first step, examples 1 and 2 illustrate the solution adjustments after the addition of a new decision process Y. In Example 1, worker edge server A is assigned to a new decision process Y, and submodule Y1 of process Y is subsequently partitioned onto this server. This reduces server A's available memory, thereby limiting the maximum feasible batch size for process X. The sorting center discovered this issue by calculating the maximum feasible batch size during runtime. One solution is for the sorting center to pass this maximum feasible batch size as information to the collaborative reasoning optimizer, allowing it to repartition process X or move the bottleneck submodule to a server with more free memory. Example 2: Instead of having the collaborative reasoning optimizer adjust the model partitioning scheme within cluster X, the maximum feasible batch size information is passed to the underlying server allocation manager. Referring to Figure 17 , the server allocation algorithm will search for a working edge server (server C) with more free memory in another cluster (cluster Z) and assign it to cluster X. Subsequently, referring to Figure 18 , as described in the first step, the collaborative reasoning optimizer will move submodule X1 to server C. In this example, cluster X's sorting center passes the maximum feasible batch size as information to the collaborative reasoning optimizer, which then passes this information down to the server allocation manager, instructing the manager to assign server C from another cluster with more free memory to cluster X. The optimizer then moves submodule X1 to server C, thereby evenly distributing the free memory between servers A and C. This shows that compared to the first step, the maximum feasible batch size for process X has been further increased. If submodule Z1 in server C is not the bottleneck submodule of the process, then Example 2 will outperform Example 1 in terms of execution results. In the third step, all three decision processes participate in the deployment. If process X has a significantly higher request rate than processes Y and Z, then the sorting centers in the three clusters can transmit the number of requests per unit time (such as hours or minutes) as interaction information to the server allocation manager. This way, when the server allocation algorithm runs, if server A can be assigned to any process, process X, with a higher request rate, will be given priority for server A.

[0138] 20 , an embodiment of the present invention further provides a collaborative optimization system based on edge computing, including a data acquisition unit 1001, a parameter dictionary and deep neural network directed acyclic graph construction unit 1002, a model deployment execution planning unit 1003, a model deployment execution implementation unit 1004, and a service request execution unit 1005, wherein:

[0139] The data acquisition unit 1001 is used to acquire local configuration files, web UI data, local model parameters and client data.

[0140] The parameter dictionary and deep neural network directed acyclic graph construction unit 1002 is used to construct a deep neural network directed acyclic graph according to the local configuration file and web UI data, and to construct a module parameter dictionary according to the local model parameters.

[0141] The model deployment execution planning unit 1003 is used to construct a deep neural network model through a parameter dictionary and a deep neural network directed acyclic graph; the deep neural network model loop executes the following steps: executes a server allocation algorithm according to client data through the deep neural network model to obtain an edge server allocation result; in response to the edge server allocation result, executes a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, executes a data batch processing algorithm through the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; obtains a server allocation management table in the server allocation algorithm, and stops the loop if there is no allocation task in the server allocation management table and the duration reaches a preset time threshold.

[0142] The model deployment execution implementation unit 1004 is used to obtain the module parameters corresponding to the sub-module division result through the parameter dictionary in the deep neural network model, and distribute the module parameters to the working edge server corresponding to the edge server allocation result.

[0143] The service request execution unit 1005 is used to respond to the service request of the client through module parameters, edge server allocation results, sub-module division results and data batch division results.

[0144] It should be noted that since the collaborative optimization system based on edge computing in this embodiment and the collaborative optimization method based on edge computing mentioned above are based on the same inventive concept, the corresponding content in the method embodiment is also applicable to the device embodiment and will not be described in detail here.

[0145] 21 , another embodiment of the present invention further provides an electronic device, wherein the electronic device 6000 may be any type of smart terminal, such as a mobile phone, a tablet computer, a personal computer, etc.

[0146] Specifically, the electronic device 6000 includes: one or more control processors 6001 and a memory 6002. Figure 21 takes one control processor 6001 and one memory 6002 as an example. The control processor 6001 and the memory 6002 can be connected via a bus or other means. Figure 21 takes the connection via a bus as an example.

[0147] The memory 6002 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to an electronic device in an embodiment of the present invention;

[0148] The control processor 6001 executes various functional applications and data processing of a collaborative optimization method based on edge computing by running non-transient software programs, instructions and modules stored in the memory 6002, that is, implements a collaborative optimization method based on edge computing of the above-mentioned method embodiment.

[0149] The memory 6002 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a collaborative optimization method based on edge computing, etc. In addition, the memory 6002 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 6002 may optionally include a memory remotely located relative to the control processor 6001, and these remote memories may be connected to the electronic device 6000 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] One or more modules are stored in the memory 6002. When executed by the one or more control processors 6001, a collaborative optimization method based on edge computing in the above method embodiment is executed, for example, the method steps of Figures 1 to 6 described above are executed.

[0151] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0152] It should be noted that since an electronic device in this embodiment and the above-mentioned collaborative optimization method based on edge computing are based on the same inventive concept, the corresponding content in the method embodiment is also applicable to the device embodiment and will not be described in detail here.

[0153] One embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute: a collaborative optimization method based on edge computing as described in the above embodiment.

[0154] It should be noted that since the computer-readable storage medium in this embodiment and the above-mentioned collaborative optimization method based on edge computing are based on the same inventive concept, the corresponding content in the method embodiment is also applicable to the device embodiment and will not be described in detail here.

[0155] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing data (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired data and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any data delivery media.

[0156] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0157] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.

Claims

1. A collaborative optimization method based on edge computing, characterized in that The collaborative optimization method based on edge computing includes: Obtain the local configuration file, web UI data, local model parameters, and client data; Construct a directed acyclic graph of a deep neural network according to the local configuration file and the web UI data, and construct a module parameter dictionary according to the local model parameters; Construct a deep neural network model through the parameter dictionary and the directed acyclic graph of the deep neural network; the following steps are cyclically executed through the deep neural network model: execute a server allocation algorithm according to the client data through the deep neural network model to obtain an edge server allocation result; in response to the edge server allocation result, execute a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, execute a data batch processing algorithm through the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; obtain the server allocation management table in the server allocation algorithm, and if there is no allocation task in the server allocation management table and the duration reaches a preset time threshold, stop the loop; Obtain the module parameters corresponding to the sub-module partitioning result through the parameter dictionary in the deep neural network model, and distribute the module parameters to the working edge servers corresponding to the edge server allocation result; Respond to the service request of the client through the module parameters, according to the edge server allocation result, the sub-module partitioning result, and the data batch partitioning result.

2. The collaborative optimization method based on edge computing according to claim 1, wherein Calculate the forward result of the directed acyclic graph of the deep neural network through the automatic inference model forward function, and supervise the distribution of the module parameters through the forward result.

3. The collaborative optimization method based on edge computing according to claim 1, wherein The step of executing a server allocation algorithm according to the client data through the deep neural network model to obtain an edge server allocation result includes: Record the basic information of the registered edge servers through a server configuration table; Obtain the server deployment request of the client; In response to the server deployment request, calculate the edge server to be deployed corresponding to the server deployment request through the server allocation algorithm; Obtain a logical edge cluster according to the server configuration table, and select the cluster leader of the logical edge cluster; Guide the client to the leading edge server corresponding to the cluster leader through the cluster leader, and perform the deployment transaction of the edge server to be deployed on the leading edge server to obtain the edge server allocation result.

4. The collaborative optimization method based on edge computing according to claim 1, characterized in that, The step of, in response to the edge server allocation result, executing a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model includes: Calculate the first inference delay records on each of the working edge servers through a preset first inference profiler; the first inference delay records represent multiple first inference delays corresponding to different sub-module combinations deployed on the working edge servers; Input the first inference delay records into the model partitioning algorithm to obtain the sub-module partitioning result.

5. The collaborative optimization method based on edge computing according to claim 4, wherein In response to the sub-module partitioning result, execute a data batch processing algorithm through the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result, including: Obtain the request data of the client; Obtain multiple data batches through different combinations of the request data; Calculate multiple inference latencies of each sub-module under the multiple data batches through a preset second inference profiler to obtain a second inference latency record of the second inference profiler; Input the second inference latency record into the data batch processing algorithm to calculate and obtain the data batch partitioning result.

6. The collaborative optimization method based on edge computing according to claim 5, wherein, Both the first inference profiler and the second inference profiler are obtained through training of a regression model. The training steps of the regression model include: Construct raw data parameters and raw module parameters; Obtain the configuration parameters of the working edge server, and construct an input sample through the raw data parameters, the raw module parameters, and the configuration parameters; Construct pseudo data according to the raw data parameters, and construct a pseudo module model according to the raw module parameters; Measure the simulated inference latency corresponding to the input sample through the pseudo data and the pseudo module model; Divide the training set through the simulated inference latency and the input sample; Train the regression model through the training set to obtain the first inference profiler or the second inference profiler.

7. The collaborative optimization method based on edge computing according to claim 4, wherein The step of inputting the first inference latency record into the model partitioning algorithm to obtain the sub-module partitioning result includes: Input the first inference latency record into the model partitioning algorithm to obtain a sub-module partitioning combination; Obtain the sub-module replication strategy corresponding to the data batch partitioning result; Construct a module stub on the working edge server, and store the model parameters corresponding to the sub-module partitioning combination and the sub-module replication strategy through the module stub; Construct a distributed module model corresponding to the working edge server according to the module stub to obtain the sub-module partitioning result.

8. A collaborative optimization system based on edge computing, characterized in that, The collaborative optimization system based on edge computing includes: A data acquisition unit for acquiring a local configuration file, web UI data, local model parameters, and client data; A parameter dictionary and a deep neural network directed acyclic graph construction unit for constructing a deep neural network directed acyclic graph according to the local configuration file and the web UI data, and constructing a module parameter dictionary according to the local model parameters; A model deployment execution planning unit is configured to build a deep neural network model through the parameter dictionary and the deep neural network directed acyclic graph; and cyclically execute the following steps through the deep neural network model: execute a server allocation algorithm according to the client data through the deep neural network model to obtain an edge server allocation result; in response to the edge server allocation result, execute a model partitioning algorithm on the deep neural network model to obtain a sub-module partitioning result corresponding to the deep neural network model; in response to the sub-module partitioning result, execute a data batch processing algorithm through the deep neural network model to obtain a data batch partitioning result corresponding to the sub-module partitioning result; obtain a server allocation management table in the server allocation algorithm, and if there is no allocation task in the server allocation management table and the duration reaches a preset time threshold, stop the loop; A model deployment execution implementation unit is configured to obtain module parameters corresponding to the sub-module partitioning result through the parameter dictionary in the deep neural network model, and distribute the module parameters to the working edge servers corresponding to the edge server allocation result; A service request execution unit is configured to respond to a service request from the client through the module parameters, based on the edge server allocation result, the sub-module partitioning result, and the data batch partitioning result.

9. An electronic device, characterized in that: It includes at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the edge computing-based collaborative optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the edge computing-based collaborative optimization method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-user deep neural network model segmentation and resource allocation optimization method in edge computing scene

    CN112822701A

  • Accelerated resource allocation techniques

    US20200104184A1