A deep learning optimization method and device for cloud computing and edge computing
By introducing model selection and data compression modules into deep learning services, the model complexity and data transmission volume are dynamically configured, and the impact of input data volume on service cost in the prior art and the lack of dynamic optimization configuration are solved, thus achieving effective reduction of latency and energy consumption cost of deep learning services.
Patent Information
- Application Number
- CN202111333858.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-11
AI Technical Summary
The prior art fails to effectively consider the impact of input data on service cost when deploying deep learning services, and lacks dynamic optimization configuration capabilities, resulting in difficult to effectively reduce the cost of delay and energy consumption.
By introducing model selection and data compression modules, dynamic configuration of deep learning model complexity and data transmission volume is realized, and optimal configuration is achieved through model capability query tables, achieving the purpose of dynamic real-time optimization.
On the premise of ensuring model capabilities, dynamically adjust the data compression ratio and model compression ratio, significantly reduce the latency and energy consumption cost of deep learning services, improve user experience and reduce service operation costs.
Smart Images

Figure CN114036155B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an optimization method and device for deep learning for cloud computing and edge computing. Background Art
[0002] Deep learning has a wide range of applications in the Internet of Things, Internet of Vehicles, drones, industrial Internet of Things, and other applications. It is one of the key supporting technologies for edge computing and cloud computing. In cloud computing or edge computing environments, artificial intelligence services based on deep learning are sensitive to latency and energy costs. How to reduce latency or energy costs as much as possible while ensuring model capability requirements is a problem that must be considered when deploying deep learning services. The capabilities of deep learning models gradually increase with their complexity, and as the capabilities gradually increase, the performance gain brought by the increase in complexity gradually decreases until the maximum value is reached. The cost of deploying the highest performance deep learning service is very high, and its latency is relatively high. In actual deployment, it is generally necessary to compress the deep learning model to reduce its computational complexity and reduce the latency or energy cost of the service while slightly reducing the model capability.
[0003] Research has found that although the number of deep learning model parameters is very large, some parameters in the model have little effect on the model's capabilities, and deleting a certain proportion of low-value weights has little effect on the model's capabilities. Therefore, by pruning the model according to certain rules, the purpose of compressing the computational complexity can be achieved, thereby reducing the model's latency and energy consumption. Related patents include "A deep learning network model compression method based on network layer pruning" (CN2020101779121), "Deep learning model compression method and device" (CN2020106919481), and "A deep learning model compression method and device" (CN202110245695X).
[0004] Existing structural optimization usually adds multiple output ports to the middle layer of the deep learning model, so that the model has the ability to output results in advance. The model complexity corresponding to the relatively shallow output is also lower, but its model capability will also be reduced to a certain extent. The purpose of reducing model complexity can be achieved by selecting different early output ports to obtain results.
[0005] Problems with existing technologies include not considering the impact of the amount of input data on the service cost: the quality of deep learning input data will affect the capabilities of the model. Specifically, the higher the quality, the higher the model capability (or the quality of the results obtained), but the amount of data will also be larger, which in turn affects the data transmission delay and energy consumption cost. In addition, if the input data dimension is reduced after the data quality is reduced, the complexity of the corresponding model will also be reduced to a certain extent, ultimately reducing the model's computational delay and energy consumption cost. In cloud computing and edge computing scenarios, data is often transmitted from the user's sending node to the cloud center or edge computing server, and the data transmission process will also generate non-negligible delay and energy consumption costs.
[0006] No support for dynamic optimization configuration: Existing methods only support static configuration, that is, only one deep learning model is provided during deployment, and its computational complexity and model capabilities are fixed, and dynamic optimization is no longer possible. However, due to the dynamic changes in the capacity of the data transmission link, existing technologies cannot reduce latency or energy consumption costs by dynamically configuring model complexity and data transmission volume.
[0007] In response to the above problems, the present invention realizes dynamic configuration of deep learning model complexity and data transmission volume by simultaneously introducing model selection and data compression modules, and further realizes optimal configuration of model complexity and data transmission volume by introducing model capability query table, thereby achieving the purpose of dynamic real-time optimization. Summary of the invention
[0008] The present invention proposes the present invention to solve the problem of how to dynamically configure the model complexity and data transmission volume of deep learning to achieve dynamic real-time optimization.
[0009] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0010] On the one hand, the present invention provides an optimization method for deep learning for cloud computing and edge computing, the method is implemented by an optimization system for deep learning, the optimization system for deep learning includes a sending node and a server node, and the method includes:
[0011] S1. The sending node compresses the data of the task to be executed and sends it to the server node.
[0012] S2. The server node receives the compressed data and inputs it into the deep learning optimization model.
[0013] S3, the server node outputs the optimal dynamic optimization result of the task to be executed based on compressed data and deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0014] Optionally, the sending node in S1 compresses the data of the task to be executed including:
[0015] The sending node compresses the data of the task to be executed according to the data compression ratio; the data compression ratio is expressed as a, 0≤a<1, the larger the value of a, the greater the data compression ratio, and the smaller the amount of data required to be transmitted.
[0016] Optionally, the service cost in S3 includes a delay cost T(a, b) and an energy consumption cost E(a, b).
[0017] The calculation formula of service cost C(a, b) is expressed by the following formula (1):
[0018] C(a,b)=sT(a,b)+(1-s)E(a,b) (1)
[0019] Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of delay cost and energy consumption cost, 0≤s≤1.
[0020] The calculation formula of the delay cost T(a,b) is expressed by the following formula (2):
[0021] T(a,b)=T tr (a)+T cp (b) = (1-a)T tr (0)+(1-b)T cp (0) (2)
[0022] Among them, T tr (a) represents the transmission delay, T cp (b) represents the computation delay, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the calculation delay when b=0.
[0023] The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3):
[0024] E(a,b)=E tr (a)+E cp (b) = (1-a) E tr (0)+(1-b)E cp (0) (3)
[0025] Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b=0.
[0026] Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4):
[0027] C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4)
[0028] Let the service cost coefficient x = s[T tr (0)+E tr (0)], service cost coefficient y = (1-s)[T cp (0)+E cp (0)], the above formula (4) can be further simplified as the following formula (5):
[0029] C(a,b)=(1-a)x+(1-b)y (5)
[0030] According to the above formula (5), when the service cost coefficients x, y are known, the service cost is obtained by configuring a and b.
[0031] Optionally, the deep learning optimization model includes a training module and a deployment module.
[0032] The server nodes in S3 output the optimal dynamic optimization results of the tasks to be executed based on compressed data and deep learning optimization models, including:
[0033] S31. Input the compressed data into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool. The mapping function of the model capability query table is Q(a, b).
[0034] S32. Based on the model pool, the model capability query table and the deployment module, the optimal dynamic optimization result of the task to be executed is obtained.
[0035] Optionally, in S31, the compressed data is input into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool, including:
[0036] S311. Set M data compression ratios and N model compression ratios according to the tasks to be performed, and determine the data compression method and the initial deep learning model; wherein M and N are both positive integers.
[0037] S312. According to M data compression ratios and N model compression ratios, the initial deep learning model is trained to obtain a model pool, where the model pool consists of MxN different models, each model having a different data compression ratio and model complexity compression ratio.
[0038] S313. Test the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error.
[0039] S314. According to the model capability value, model complexity compression ratio, and data compression ratio, a model capability query table between the model capability value and the two indicators of model complexity compression ratio and data compression ratio is obtained.
[0040] Optionally, the data compression method in S311 includes an autoencoder method, a downsampling method, and a dimensionality reduction method.
[0041] Optionally, the deep learning model in S311 includes a residual neural network and an attention mechanism network.
[0042] Optionally, obtaining the optimal dynamic optimization result based on the model pool, the model capability query table, and the deployment module in S32 includes:
[0043] S321, model initialization, including deployment of model pool, model capability query table, setting model capability threshold Q th .
[0044] S322. Obtain the highest transmission delay T under the current environment tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a).
[0045] S323, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculate the service cost coefficients x, y, and calculate the service cost C(a, b) corresponding to the model capability query table.
[0046] S324: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a,b)≥Q th The (a,b) value combination of the condition.
[0047] S325. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * ,b * ), that is, C(a * ,b * )=minC(a,b).
[0048] S326, according to (a * ,b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S322 to perform the next round of dynamic optimization process of service cost until the task is completed.
[0049] Optionally, the server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
[0050] On the other hand, the present invention provides an optimization device for deep learning for cloud computing and edge computing, which is used to implement an optimization method for deep learning for cloud computing and edge computing, and the device includes:
[0051] The data sending unit is used for the sending node to compress the data of the task to be executed and send it to the server node.
[0052] The input unit is used for the server node to receive compressed data and input it into the deep learning optimization model.
[0053] The output unit is used for the server node to output the optimal dynamic optimization result of the task to be executed based on compressed data and deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0054] Optionally, the data sending unit is further used to:
[0055] The sending node compresses the data of the task to be executed according to the data compression ratio; the data compression ratio is expressed as a, 0≤a<1, the larger the value of a, the greater the data compression ratio, and the smaller the amount of data required to be transmitted.
[0056] Optionally, the service cost includes a delay cost T(a, b) and an energy consumption cost E(a, b).
[0057] The calculation formula of service cost C(a, b) is expressed by the following formula (1):
[0058] C(a,b)=sT(a,b)+(1-s)E(a,b) (1)
[0059] Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of delay cost and energy consumption cost, 0≤s≤1.
[0060] The calculation formula of the delay cost T(a,b) is expressed by the following formula (2):
[0061] T(a,b)=T tr (a)+T cp (b) = (1-a)T tr(0)+(1-b)T cp (0) (2)
[0062] Among them, T tr (a) represents the transmission delay, T cp (b) represents the computation delay, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the calculation delay when b=0.
[0063] The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3):
[0064] E(a,b)=E tr (a)+E cp (b) = (1-a) E tr (0)+(1-b)E cp (0) (3)
[0065] Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b=0.
[0066] Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4):
[0067] C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4)
[0068] Let the service cost coefficient x = s[T tr (0)+E tr (0)], service cost coefficient y = (1-s)[T cp (0)+E cp (0)], the above formula (4) can be further simplified as the following formula (5):
[0069] C(a,b)=(1-a)x+(1-b)y (5)
[0070] According to the above formula (5), when the service cost coefficients x, y are known, the service cost is obtained by configuring a and b.
[0071] Optionally, the deep learning optimization model includes a training module and a deployment module.
[0072] Based on compressed data and deep learning optimization models, the server node outputs the optimal dynamic optimization results of the tasks to be executed, including:
[0073] S31. Input the compressed data into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool. The mapping function of the model capability query table is Q(a, b).
[0074] S32. Based on the model pool, the model capability query table and the deployment module, the optimal dynamic optimization result of the task to be executed is obtained.
[0075] Optionally, the output unit is further configured to:
[0076] S311. Set M data compression ratios and N model compression ratios according to the tasks to be performed, and determine the data compression method and the initial deep learning model; wherein M and N are both positive integers.
[0077] S312. According to M data compression ratios and N model compression ratios, the initial deep learning model is trained to obtain a model pool, where the model pool consists of MxN different models, each model having a different data compression ratio and model complexity compression ratio.
[0078] S313. Test the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error.
[0079] S314. According to the model capability value, model complexity compression ratio, and data compression ratio, a model capability query table between the model capability value and the two indicators of model complexity compression ratio and data compression ratio is obtained.
[0080] Optionally, the data compression method includes an autoencoder method, a downsampling method, and a dimensionality reduction method.
[0081] Optionally, the deep learning model includes a residual neural network and an attention mechanism network.
[0082] Optionally, the output unit is further configured to:
[0083] S321, model initialization, including deployment of model pool, model capability query table, setting model capability threshold Q th .
[0084] S322. Obtain the highest transmission delay T under the current environment tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a).
[0085] S323, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculate the service cost coefficients x, y, and calculate the service cost C(a, b) corresponding to the model capability query table.
[0086] S324: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a,b)≥Q th The (a,b) value combination of the condition.
[0087] S325. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * ,b * ), that is, C(a * ,b * )=minC(a,b).
[0088] S326, according to (a * ,b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S322 to perform the next round of dynamic optimization process of service cost until the task is completed.
[0089] Optionally, the server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
[0090] On the one hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned optimization method for deep learning for cloud computing and edge computing.
[0091] On the one hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement the above-mentioned optimization method for deep learning for cloud computing and edge computing.
[0092] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0093] In the above scheme, an optimization method that supports real-time dynamic configuration is disclosed for the optimization problem of latency, energy consumption and other costs of deep learning services. First, multiple optional models with different data transmission volumes and computational complexities are trained, and a model pool and data compression module are built on this basis. Subsequently, a mapping table between model capability and two indicators, model complexity and data transmission volume, is obtained through performance testing. After the model is deployed, for a deep learning task, the optimal model complexity and data volume configuration in the current environment is obtained by querying the model capability mapping table, and the optimal data transmission volume and model computational complexity are dynamically configured under the premise of meeting the model capability constraints, thereby reducing the latency and energy consumption costs of deep learning services, improving user experience, and reducing service operating costs.
[0094] Under the premise of ensuring that the capabilities of the deep learning model meet the given threshold constraints, the present invention dynamically adjusts the data compression ratio on the sending end and the model compression ratio on the server end to minimize the latency or energy consumption cost of the deep learning service, thereby improving the user experience quality or reducing the operating cost of the service.
[0095] In addition to being applied to cloud computing and edge computing scenarios, the present invention can also be applied to various types of collaborative Internet of Things, Internet of Vehicles, and UAV systems. Specifically, the receiving node does not need to be a cloud computing or edge computing server, but is any node with data decompression and model selection capabilities. At this time, the receiving node provides deep learning services, and the present invention can optimize its service cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0097] Figure 1 It is a schematic flow chart of an optimization method for deep learning for cloud computing and edge computing of the present invention;
[0098] Figure 2 It is a schematic flow chart of an optimization method for deep learning for cloud computing and edge computing of the present invention;
[0099] Figure 3 It is a block diagram of the optimization system of deep learning of the present invention;
[0100] Figure 4 It is a block diagram of an optimization device for deep learning of cloud computing and edge computing of the present invention;
[0101] Figure 5It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0102] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0103] like Figure 1 As shown, an embodiment of the present invention provides an optimization method for deep learning for cloud computing and edge computing. The method is implemented by an optimization system for deep learning. The optimization system for deep learning includes a sending node and a server node. The method includes:
[0104] S11. The sending node compresses the data of the task to be executed and sends it to the server node.
[0105] S12. The server node receives the compressed data and inputs it into the deep learning optimization model.
[0106] S13. The server node outputs the optimal dynamic optimization result of the task to be executed based on compressed data and deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0107] Optionally, the sending node in S11 compresses the data of the task to be executed, including:
[0108] The sending node compresses the data of the task to be executed according to the data compression ratio; the data compression ratio is expressed as a, 0≤a<1, the larger the value of a, the greater the data compression ratio, and the smaller the amount of data required to be transmitted.
[0109] Optionally, the service cost in S13 includes a delay cost T(a, b) and an energy consumption cost E(a, b).
[0110] The calculation formula of service cost C(a, b) is expressed by the following formula (1):
[0111] C(a,b)=sT(a,b)+(1-s)E(a,b) (1)
[0112] Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of delay cost and energy consumption cost, 0≤s≤1.
[0113] The calculation formula of the delay cost T(a,b) is expressed by the following formula (2):
[0114] T(a,b)=T tr (a)+T cp (b) = (1-a)T tr (0)+(1-b)T cp (0) (2)
[0115] Among them, T tr (a) represents the transmission delay, T cp (b) represents the computation delay, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the calculation delay when b=0.
[0116] The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3):
[0117] E(a,b)=E tr (a)+E cp (b) = (1-a) E tr (0)+(1-b)E cp (0) (3)
[0118] Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b=0.
[0119] Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4):
[0120] C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4)
[0121] Let the service cost coefficient x = s[T tr (0)+E tr (0)], service cost coefficient y = (1-s)[T cp (0)+E cp (0)], the above formula (4) can be further simplified as the following formula (5):
[0122] C(a,b)=(1-a)x+(1-b)y (5)
[0123] According to the above formula (5), when the service cost coefficients x, y are known, the service cost is obtained by configuring a and b.
[0124] Optionally, the deep learning optimization model includes a training module and a deployment module.
[0125] The server nodes in S13 output the optimal dynamic optimization results of the tasks to be executed based on compressed data and deep learning optimization models, including:
[0126] S131. Input the compressed data into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool. The mapping function of the model capability query table is Q(a, b).
[0127] S132. Based on the model pool, the model capability query table and the deployment module, the optimal dynamic optimization result of the task to be executed is obtained.
[0128] Optionally, in S131, the compressed data is input into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool, including:
[0129] S1311. Set M data compression ratios and N model compression ratios according to the tasks to be performed, and determine the data compression method and the initial deep learning model; wherein M and N are both positive integers.
[0130] S1312. According to M data compression ratios and N model compression ratios, the initial deep learning model is trained to obtain a model pool, where the model pool consists of MxN different models, each model having a different data compression ratio and model complexity compression ratio.
[0131] S1313. Test the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error.
[0132] S1314. According to the model capability value, model complexity compression ratio, and data compression ratio, a model capability query table between the model capability value and the two indicators of model complexity compression ratio and data compression ratio is obtained.
[0133] Optionally, the data compression method in S1311 includes an autoencoder method, a downsampling method, and a dimensionality reduction method.
[0134] Optionally, the deep learning model in S1311 includes a residual neural network and an attention mechanism network.
[0135] Optionally, obtaining the optimal dynamic optimization result based on the model pool, the model capability query table, and the deployment module in S132 includes:
[0136] S1321, model initialization, including deployment of model pool, model capability query table, setting model capability threshold Q th .
[0137] S1322. Obtain the highest transmission delay T under the current environmenttr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a).
[0138] S1323, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculate the service cost coefficients x, y, and calculate the service cost C(a, b) corresponding to the model capability query table.
[0139] S1324: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a,b)≥Q th The (a,b) value combination of the condition.
[0140] S1325. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * ,b * ), that is, C(a * ,b * )=minC(a,b).
[0141] S1326, according to (a * ,b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S1322 to perform the next round of dynamic optimization process of service cost until the task is completed.
[0142] Optionally, the server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
[0143] In the above scheme, an optimization method that supports real-time dynamic configuration is disclosed for the optimization problem of latency, energy consumption and other costs of deep learning services. First, multiple optional models with different data transmission volumes and computational complexities are trained, and a model pool and data compression module are built on this basis. Subsequently, a mapping table between model capability and two indicators, model complexity and data transmission volume, is obtained through performance testing. After the model is deployed, for a deep learning task, the optimal model complexity and data volume configuration in the current environment is obtained by querying the model capability mapping table, and the optimal data transmission volume and model computational complexity are dynamically configured under the premise of meeting the model capability constraints, thereby reducing the latency and energy consumption costs of deep learning services, improving user experience, and reducing service operating costs.
[0144] Under the premise of ensuring that the capabilities of the deep learning model meet the given threshold constraints, the present invention dynamically adjusts the data compression ratio on the sending end and the model compression ratio on the server end to minimize the latency or energy consumption cost of the deep learning service, thereby improving the user experience quality or reducing the operating cost of the service.
[0145] In addition to being applied to cloud computing and edge computing scenarios, the present invention can also be applied to various types of collaborative Internet of Things, Internet of Vehicles, and UAV systems. Specifically, the receiving node does not need to be a cloud computing or edge computing server, but is any node with data decompression and model selection capabilities. At this time, the receiving node provides deep learning services, and the present invention can optimize its service cost.
[0146] like Figure 2 As shown, an embodiment of the present invention provides an optimization method for deep learning for cloud computing and edge computing, and the method is implemented by an optimization system for deep learning, such as Figure 3 The deep learning optimization system shown includes a sending node and a server node, and the method includes:
[0147] S21. The sending node compresses the data of the task to be executed and sends it to the server node.
[0148] Optionally, the sending node compresses the data of the task to be executed according to a data compression ratio; DCR (Data Compression Ratio) is represented by a, 0≤a<1, and the larger the value of a is, the larger the data compression ratio is, and the smaller the amount of data required to be transmitted is.
[0149] In a feasible implementation, data compression refers to reducing data storage space to reduce data transmission, storage and processing efficiency. In the present invention, in order to pursue the data compression effect, the compression process may cause a certain amount of information loss. The transmission ratio of compressed data is 1-a, so the larger a is, the less data needs to be transmitted, thereby reducing transmission delay and energy consumption. When a=0, it corresponds to a situation without compression. At this time, the sending node sends the original data completely to the receiving point and then processes it. Note that lossy data compression will cause information loss, thereby reducing the quality of the output result.
[0150] Among them, the sending node can be any qualified sensor node, mobile terminal, vehicle, drone, etc. The sending node has a data compression function, that is, it can compress the data to be executed according to the data compression ratio DCR configuration value. After compression, the data transmission delay and energy consumption are reduced. There may be a certain amount of information loss after the data information is restored. For example, after the image data is downsampled, the clarity will decrease after it is restored again. In addition, after the compressed data is delivered to the server, it can be directly input into the deep learning model for calculation without restoration, but this mode requires the deep learning model to support the compressed data dimension.
[0151] Optionally, the server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
[0152] S22, the server node receives the compressed data and inputs it into the training module to obtain the model pool and the MCQT (Model Capability Query Table) corresponding to each model in the model pool, including S221-S224:
[0153] S221. Set M data compression ratios and N MCRs (Model Compression Ratios) according to the tasks to be performed, and determine the data compression method and the initial deep learning model; wherein M and N are both positive integers.
[0154] In a feasible implementation, model compression refers to reducing the number of parameters of deep learning to reduce its storage space and computational complexity. Model compression will cause a decrease in model capabilities. Generally speaking, the greater the compression, the greater the degree of model capability loss. The model compression ratio is expressed as b, 0≤b<1. The larger the value of b, the greater the model compression ratio. The complexity ratio of the compressed model is 1-b, so the larger b is, the lower the corresponding computational complexity. When b=0, the model with the maximum complexity is used for processing. Note that model compression will reduce the model's capabilities to a certain extent, thereby reducing the quality of the output results.
[0155] Optionally, the data compression method in S221 may include an autoencoder method, a downsampling method, and a dimensionality reduction method.
[0156] In a feasible implementation manner, the data compression method may be a general data compression method, such as an autoencoder, downsampling, dimensionality reduction, etc. If necessary, the original data may be restored at the receiving end to keep the data dimension consistent.
[0157] Optionally, the deep learning model in S221 may include a residual neural network and an attention mechanism network.
[0158] In a feasible implementation, the deep learning model can adopt any suitable model method, such as residual neural network, attention mechanism network, etc., and a targeted network structure can also be designed according to actual needs.
[0159] S222. According to M data compression ratios and N model compression ratios, the initial deep learning model is trained to obtain a model pool, where the model pool consists of MxN different models, each model having a different data compression ratio and model complexity compression ratio.
[0160] In a feasible implementation, the data compression ratio M is set to 3, and the specific data compression ratios are 0.1, 0.2, and 0.3; the model complexity compression ratio N is set to 4, and the specific model complexity compression ratios are 0.4, 0.5, 0.6, and 0.7; then 3×4=12 different models can be obtained, which may include model 1 with a data compression ratio of 0.1 and a model complexity compression ratio of 0.4, model 2 with a data compression ratio of 0.1 and a model complexity compression ratio of 0.5, model 3 with a data compression ratio of 0.1 and a model complexity compression ratio of 0.6, model 4 with a data compression ratio of 0.1 and a model complexity compression ratio of 0.7, model 5 with a data compression ratio of 0.2 and a model complexity compression ratio of 0.4... model 12 with a data compression ratio of 0.3 and a model complexity compression ratio of 0.7.
[0161] Model compression can adopt more common pruning and early output methods.
[0162] S223. Test the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error.
[0163] In a feasible implementation manner, the present invention measures the capability value of the model by the quality of the output results of the deep learning model. The relevant quality indicators may include classification accuracy, mean square error, etc.
[0164] The model capability refers to the performance indicators of the deep learning model used, which may include classification accuracy, detection rate, minimum mean square error, etc.
[0165] S224. According to the model capability value, model complexity compression ratio, and data compression ratio, a model capability query table between the model capability value and the two indicators of model complexity compression ratio and data compression ratio is obtained.
[0166] In a feasible implementation, a mapping table is used to record the model capability and the DQR and MCR, and the mapping function is assumed to be Q(a, b). Table 1 is a model capability query table for image classification applications, which gives an example of MCQT for image recognition applications based on deep learning, where Q(a, b) decreases monotonically with a or b. After a capability threshold is given, the feasible value range of the DQR and MCR that meet the requirements can be queried.
[0167] Table 1
[0168] Q(a,b) b=0 b=0.1 b=0.2 b=0.3 b=0.4 a=0 0.98 0.97 0.95 0.92 0.89 a=0.1 0.97 0.96 0.94 0.91 0.88 a=0.2 0.96 0.95 0.93 0.90 0.87 a=0.3 0.95 0.94 0.92 0.89 0.86 a=0.4 0.94 0.93 0.91 0.88 0.85
[0169] S23, based on the model pool, model capability query table and deployment module, the optimal dynamic optimization results are obtained including S231-S236:
[0170] S231, model initialization, including deployment of model pool, model capability query table, setting model capability threshold Q th .
[0171] S232. Obtain the highest transmission delay T under the current environment tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a).
[0172] In a feasible implementation manner, the above-mentioned relevant information may be obtained through monitoring software, or may be estimated based on the network and server status.
[0173] S233, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculate the service cost coefficients x, y, and calculate the service cost C(a, b) corresponding to the model capability query table.
[0174] Optionally, the service cost in S233 includes a delay cost T(a, b) and an energy consumption cost E(a, b).
[0175] The calculation formula of service cost C(a, b) is expressed by the following formula (1):
[0176] C(a,b)=sT(a,b)+(1-s)E(a,b) (1)
[0177] Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of delay cost and energy consumption cost, 0≤s≤1.
[0178] In a feasible implementation, when s=1, only the delay cost is considered, and when s=0, only the energy consumption cost is considered. There are different calculation methods for delay and energy consumption costs, which can be targeted according to the actual scenario. The present invention provides a simple and effective method. First, the calculation formula of the delay cost T(a, b) is expressed by the following formula (2):
[0179] T(a,b)=T tr (a)+T cp (b) = (1-a)T tr (0)+(1-b)T cp (0) (2)
[0180] Among them, T tr (a) represents the transmission delay, T cp (b) represents the computational delay, for example, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the calculation delay when b=0.
[0181] The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3):
[0182] E(a,b)=E tr (a)+E cp (b) = (1-a) E tr (0)+(1-b)E cp (0) (3)
[0183] Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b=0.
[0184] Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4):
[0185] C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4)
[0186] Let the service cost coefficient x = s[T tr (0)+E tr (0)], service cost coefficient y = (1-s)[T cp (0)+E cp(0)], the above formula (4) can be further simplified as the following formula (5):
[0187] C(a,b)=(1-a)x+(1-b)y (5)
[0188] According to the above formula (5), when the service cost coefficients x, y are known, the service cost is obtained by configuring a and b.
[0189] S234: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a,b)≥Q th The (a,b) value combination of the condition.
[0190] S235. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * ,b * ), that is, C(a * ,b * )=minC(a,b).
[0191] S236, according to (a * ,b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S232 to perform the next round of dynamic optimization process of service cost until the task is completed.
[0192] In a feasible implementation, the server node optimizes the configuration of the sending data compression ratio and the deep learning model complexity compression ratio according to the service performance requirements; the server node configures the model structure and parameters according to the model compression ratio to obtain the deep learning optimization model; the sending node compresses the data of the task to be executed and sends it to the server node; the server node receives the compressed data and directly inputs it into the deep learning model or inputs it into the deep learning optimization model after restoration and reconstruction; the server node outputs the optimal dynamic optimization result of the task to be executed based on the input data and the configured deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0193] The present invention is applied to actual industrial production, the deep learning application type of image recognition is selected, and the service cost coefficients x=1.5, y=1 are set, Q th =0.95, then the service cost is C(a,b)=1.5*(1-a)+(1-b).
[0194] Table 2
[0195] Q(a,b) b=0 b=0.1 b=0.2 b=0.3 b=0.4 a=0 <![CDATA[ 0.98 ]]> <![CDATA[ 0.97 ]]> <![CDATA[ 0.95 ]]> 0.92 0.89 a=0.1 <![CDATA[ 0.97 ]]> <![CDATA[ 0.96 ]]> 0.94 0.91 0.88 a=0.2 <![CDATA[ 0.96 ]]> <![CDATA[ 0.95 ]]> 0.93 0.90 0.87 a=0.3 <![CDATA[ 0.95 ]]> 0.94 0.92 0.89 0.86 a=0.4 0.94 0.93 0.91 0.88 0.85
[0196] Table 3
[0197] C(a,b) b=0 b=0.1 b=0.2 b=0.3 b=0.4 a=0 <![CDATA[ 2.50 ]]> <![CDATA[ 2.40 ]]> <![CDATA[ 2.30 ]]> 2.20 2.10 a=0.1 <![CDATA[ 2.35 ]]> <![CDATA[ 2.25 ]]> 2.15 2.05 1.95 a=0.2 <![CDATA[ 2.20 ]]> <![CDATA[ 2.10 ]]> 2.00 1.90 1.80 a=0.3 <![CDATA[ 2.05 ]]> 1.95 1.85 1.75 1.65 a=0.4 1.90 1.80 1.70 1.60 1.50
[0198] At this time, if only one option of model compression is provided, that is, b has different options, and the value of a is fixed to 0. When the threshold constraint is Q(a,b)≥0.95, it can be seen that the minimum value of C(a,b) is C(a=0,b=0.2)=1.5+0.8=2.3.
[0199] By using the method of the present invention, it can be seen that there are 8 feasible solutions, which are the underlined elements in Table 2, and the corresponding service costs are shown in Table 3. At this time, it can be seen that the service cost C (a = 0.3, b = 0) = 2.05 is the minimum.
[0200] Compared with only adjusting the model complexity, the overall service cost of the present invention is reduced by (2.3-2.05) / 2.3=10.87%. While ensuring the completion of the image recognition task, the amount of data transmitted is reduced, the computational complexity is low, and dynamic optimization configuration is supported. Depending on the different x, y coefficient values, the performance improvement achieved will vary, and the service cost achieved by the joint optimization method proposed in the present invention will definitely not exceed that of the existing method. In most cases, the present invention can effectively reduce latency and energy consumption costs, thereby improving user experience quality, and reducing service costs and carbon emissions.
[0201] In the above scheme, an optimization method that supports real-time dynamic configuration is disclosed for the optimization problem of latency, energy consumption and other costs of deep learning services. First, multiple optional models with different data transmission volumes and computational complexities are trained, and a model pool and data compression module are built on this basis. Subsequently, a mapping table between model capability and two indicators, model complexity and data transmission volume, is obtained through performance testing. After the model is deployed, for a deep learning task, the optimal model complexity and data volume configuration in the current environment is obtained by querying the model capability mapping table, and the optimal data transmission volume and model computational complexity are dynamically configured under the premise of meeting the model capability constraints, thereby reducing the latency and energy consumption costs of deep learning services, improving user experience, and reducing service operating costs.
[0202] Under the premise of ensuring that the capabilities of the deep learning model meet the given threshold constraints, the present invention dynamically adjusts the data compression ratio on the sending end and the model compression ratio on the server end to minimize the latency or energy consumption cost of the deep learning service, thereby improving the user experience quality or reducing the operating cost of the service.
[0203] In addition to being applied to cloud computing and edge computing scenarios, the present invention can also be applied to various types of collaborative Internet of Things, Internet of Vehicles, and UAV systems. Specifically, the receiving node does not need to be a cloud computing or edge computing server, but is any node with data decompression and model selection capabilities. At this time, the receiving node provides deep learning services, and the present invention can optimize its service cost.
[0204] like Figure 4 As shown, an embodiment of the present invention provides an optimization device 400 for deep learning for cloud computing and edge computing. The device 400 is applied to implement an optimization method for deep learning for cloud computing and edge computing. The device 400 includes:
[0205] The data sending unit 410 is used for the sending node to compress the data of the task to be executed and send it to the server node.
[0206] The input unit 420 is used for the server node to receive compressed data and input it into the deep learning optimization model.
[0207] The output unit 430 is used for the server node to output the optimal dynamic optimization result of the task to be executed based on the compressed data and the deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0208] Optionally, the data sending unit 410 is further configured to:
[0209] The sending node compresses the data of the task to be executed according to the data compression ratio; the data compression ratio is expressed as a, 0≤a<1, the larger the value of a, the greater the data compression ratio, and the smaller the amount of data required to be transmitted.
[0210] Optionally, the service cost includes a delay cost T(a, b) and an energy consumption cost E(a, b).
[0211] The calculation formula of service cost C(a, b) is expressed by the following formula (1):
[0212] C(a,b)=sT(a,b)+(1-s)E(a,b) (1)
[0213] Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of delay cost and energy consumption cost, 0≤s≤1.
[0214] The calculation formula of the delay cost T(a,b) is expressed by the following formula (2):
[0215] T(a,b)=T tr (a)+T cp (b) = (1-a)T tr (0)+(1-b)T cp(0) (2)
[0216] Among them, T tr (a) represents the transmission delay, T cp (b) represents the computational delay, for example, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the calculation delay when b=0.
[0217] The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3):
[0218] E(a,b)=E tr (a)+E cp (b) = (1-a) E tr (0)+(1-b)E cp (0) (3)
[0219] Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b=0.
[0220] Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4):
[0221] C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4)
[0222] Let the service cost coefficient x = s[T tr (0)+E tr (0)], service cost coefficient y = (1-s)[T cp (0)+E cp (0)], the above formula (4) can be further simplified as the following formula (5):
[0223] C(a,b)=(1-a)x+(1-b)y (5)
[0224] According to the above formula (5), when the service cost coefficients x, y are known, the service cost is obtained by configuring a and b.
[0225] Optionally, the deep learning optimization model includes a training module and a deployment module.
[0226] Based on compressed data and deep learning optimization models, the server node outputs the optimal dynamic optimization results of the tasks to be executed, including:
[0227] S31. Input the compressed data into the training module to obtain the model pool and the model capability query table MCQT corresponding to each model in the model pool. The mapping function of the model capability query table is Q(a, b).
[0228] S32. Based on the model pool, the model capability query table and the deployment module, the optimal dynamic optimization result of the task to be executed is obtained.
[0229] Optionally, the output unit 430 is further configured to:
[0230] S311. Set M data compression ratios and N model compression ratios according to the tasks to be performed, and determine the data compression method and the initial deep learning model; wherein M and N are both positive integers.
[0231] S312. According to M data compression ratios and N model compression ratios, the initial deep learning model is trained to obtain a model pool, where the model pool consists of MxN different models, each model having a different data compression ratio and model complexity compression ratio.
[0232] S313. Test the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error.
[0233] S314. According to the model capability value, model complexity compression ratio, and data compression ratio, a model capability query table between the model capability value and the two indicators of model complexity compression ratio and data compression ratio is obtained.
[0234] Optionally, the data compression method includes an autoencoder method, a downsampling method, and a dimensionality reduction method.
[0235] Optionally, the deep learning model includes a residual neural network and an attention mechanism network.
[0236] Optionally, the output unit 430 is further configured to:
[0237] S321, model initialization, including deployment of model pool, model capability query table, setting model capability threshold Q th .
[0238] S322. Obtain the highest transmission delay T under the current environment tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a).
[0239] S323, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculate the service cost coefficients x, y, and calculate the service cost C(a, b) corresponding to the model capability query table.
[0240] S324: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a,b)≥Q th The (a,b) value combination of the condition.
[0241] S325. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * ,b * ), that is, C(a * ,b * )=minC(a,b).
[0242] S326, according to (a * ,b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S322 to perform the next round of dynamic optimization process of service cost until the task is completed.
[0243] Optionally, the server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
[0244] In the above scheme, an optimization method that supports real-time dynamic configuration is disclosed for the optimization problem of latency, energy consumption and other costs of deep learning services. First, multiple optional models with different data transmission volumes and computational complexities are trained, and a model pool and data compression module are built on this basis. Subsequently, a mapping table between model capability and two indicators, model complexity and data transmission volume, is obtained through performance testing. After the model is deployed, for a deep learning task, the optimal model complexity and data volume configuration in the current environment is obtained by querying the model capability mapping table, and the optimal data transmission volume and model computational complexity are dynamically configured under the premise of meeting the model capability constraints, thereby reducing the latency and energy consumption costs of deep learning services, improving user experience, and reducing service operating costs.
[0245] Under the premise of ensuring that the capabilities of the deep learning model meet the given threshold constraints, the present invention dynamically adjusts the data compression ratio on the sending end and the model compression ratio on the server end to minimize the latency or energy consumption cost of the deep learning service, thereby improving the user experience quality or reducing the operating cost of the service.
[0246] In addition to being applied to cloud computing and edge computing scenarios, the present invention can also be applied to various types of collaborative Internet of Things, Internet of Vehicles, and UAV systems. Specifically, the receiving node does not need to be a cloud computing or edge computing server, but is any node with data decompression and model selection capabilities. At this time, the receiving node provides deep learning services, and the present invention can optimize its service cost.
[0247] Figure 5 is a schematic diagram of the structure of an electronic device 500 provided in an embodiment of the present invention. The electronic device 500 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 501 and one or more memories 502, wherein the memory 502 stores at least one instruction, and the at least one instruction is loaded and executed by the processor 501 to implement the following steps of the optimization method for deep learning for cloud computing and edge computing:
[0248] S1. The sending node compresses the data of the task to be executed and sends it to the server node.
[0249] S2. The server node receives the compressed data and inputs it into the deep learning optimization model.
[0250] S3, the server node outputs the optimal dynamic optimization result of the task to be executed based on compressed data and deep learning optimization model; the optimal dynamic result is the result when the service cost is minimized.
[0251] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions, which can be executed by a processor in a terminal to complete the above-mentioned optimization method for deep learning for cloud computing and edge computing. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a tape, a floppy disk, an optical data storage device, etc.
[0252] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0253] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An optimization method for deep learning for cloud computing and edge computing, characterized in that: The method is implemented by a deep learning optimization system, the deep learning optimization system includes a sending node and a server node, and the method includes: S1, the sending node compresses the data of the task to be executed and sends it to the server node; S2. The server node receives the compressed data and inputs it into the deep learning optimization model; S3, the server node outputs the optimal dynamic optimization result of the task to be executed based on the compressed data and the deep learning optimization model; the optimal dynamic optimization result is the result when the service cost is minimum; The sending node in S1 compresses the data of the task to be executed, including: The sending node compresses the data of the task to be executed according to the data compression ratio; the data compression ratio is represented by a, 0≤a<1, and the larger the value of a is, the larger the data compression ratio is, and the smaller the amount of data required to be transmitted is; The service cost in S3 includes a delay cost r(a, b) and an energy consumption cost E(a, b); The calculation formula of service cost C(a, b) is expressed by the following formula (1): C(a,b)=sr(a,b)+(1-s)E(a,b) (1) Wherein, b is the model compression ratio, 0≤b<1; s is the weight coefficient of the delay cost and the energy consumption cost, 0≤s≤1; The calculation formula of the delay cost r(a, b) is expressed by the following formula (2): r(a,b)=T tr (a)+T cp (b)=(1-a)T tr (0)+(1-b)T cp (0) (2) Among them, T tr (a) represents the transmission delay, T cp (b) represents the computation delay, T tr (0) represents the transmission delay when a=0, T cp (0) indicates the computation delay when b = 0; The calculation formula of energy consumption cost E(a, b) is expressed by the following formula (3): E(a,b)=E tr (a)+E cp (b)=(1-a)E tr (0)+(1-b)E cp (0) (3) Where E tr (a) represents the transmission energy consumption, E cp (b) represents the computing energy consumption, E tr (0) represents the transmission energy consumption when a=0, E cp (0) represents the transmission energy consumption when b = 0; Substituting formula (2) and (3) into formula (1) yields the service cost calculation formula, as shown in formula (4): C(a,b)=(1-a)s[T tr (0)+E tr (0)]+(1-b)(1-s)[T cp (0)+E cp (0)] (4) Let the service cost coefficient x = s[Ttr(0)+Et r (0)], service cost coefficient y = (1-s)[T cp (0)+E cp (0)], the above formula (4) can be further simplified as the following formula (5): C(a, b) = (1-a)x + (1-b)y (5) According to the above formula (5), when the service cost coefficients x and y are known, the service cost is obtained by configuring a and b; The deep learning optimization model includes a training module and a deployment module; The server node in S3 outputs the optimal dynamic optimization result of the task to be executed based on the compressed data and the deep learning optimization model, including: S31, inputting the compressed data into the training module to obtain a model pool and a model capability query table MCQT corresponding to each model in the model pool, wherein the mapping function of the model capability query table is Q(a, b); S32, obtaining the optimal dynamic optimization result of the task to be executed based on the model pool, the model capability query table and the deployment module; The step of inputting the compressed data into the training module in S31 to obtain a model pool and a model capability query table MCQT corresponding to each model in the model pool includes: S311. According to the task to be performed, M data compression ratios and N model compression ratios are set, and a data compression method and an initial deep learning model are determined; wherein M and N are both positive integers; S312, training the initial deep learning model according to the M data compression ratios and the N model compression ratios to obtain the model pool, wherein the model pool consists of MxN different models, each of which has a different data compression ratio and model complexity compression ratio; S313, testing the capability value of each model according to the quality of the output result of each model in the model pool; the quality of the model output result includes: classification accuracy and mean square error; S314. According to the capability value of the model, the model complexity compression ratio, and the data compression ratio, a model capability query table between the capability value of the model and the two indicators of the model complexity compression ratio and the data compression ratio is obtained.
2. The method according to claim 1, characterized in that The data compression method in S311 includes an autoencoder method, a downsampling method, and a dimensionality reduction method.
3. The method according to claim 1, characterized in that The deep learning model in S311 includes a residual neural network and an attention mechanism network.
4. The method according to claim 1, characterized in that The obtaining of the optimal dynamic optimization result based on the model pool, the model capability query table and the deployment module in S32 includes: S321, model initialization, including deploying the model pool, model capability query table, setting model capability threshold Q th ; S322. Obtain the highest transmission delay T under the current environment tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a); S323, according to the highest transmission delay T tr (a) Transmission energy consumption E tr (a) Calculate the delay T cp (a) Calculate energy consumption E cp (a), calculating the service cost coefficients x, y, and calculating the service cost C(a, b) corresponding to the model capability query table; S324: Based on the mapping function Q(a, b) of the model capability query table and the model capability threshold Q th , query all the MCQTs that satisfy Q(a, b)≥Q th The (a, b) value combination of the condition; S325. According to the (a, b) value combination, select the value combination (a) that minimizes the service cost C(a, b). * , b * ), that is, C(a * , b * ) = minC(a, b); S326, according to (a * , b * ) Select the corresponding data compression ratio and model complexity compression ratio, and then adjust the data transmission volume and the computational complexity of the model; repeat S322 to perform the next round of dynamic optimization process of service cost until the task is completed.
5. The method according to claim 1, characterized in that The server node also includes a data recovery module, which is used to decompress the compressed data received by the server node and input it into the deep learning optimization model.
6. An optimization device for deep learning for cloud computing and edge computing, the optimization device for deep learning for cloud computing and edge computing is used to implement the optimization method for deep learning for cloud computing and edge computing as described in any one of claims 1 to 5, characterized in that: The device comprises: The data sending unit is used for the sending node to compress the data of the task to be executed and send it to the server node; An input unit, used for the server node to receive the compressed data and input it into the deep learning optimization model; An output unit is used for the server node to output the optimal dynamic optimization result of the task to be executed based on the compressed data and the deep learning optimization model; the optimal dynamic optimization result is the result when the service cost is minimized.
Citation Information
Patent Citations
Deep learning target detection system based on server-embedded cooperation
CN111709522A
Data compression method and computing equipment
CN113055017A