Dynamic optimization method of deep neural network model blocks for edge computing

By extracting, pruning, retraining and scaling the deep neural network model, we generate descendant blocks suitable for edge devices, and selecting combination methods according to task characteristics, the problem of low deployment efficiency of DNN models on edge devices is solved, and efficient computing and flexible deployment are achieved.

CN119783744BActive Publication Date: 2025-06-06QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510291378.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-06
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Deep neural network models are difficult to deploy efficiently on resource-constrained edge devices due to complex structure and large computing volume.

Method used

By extracting model blocks with different focus from the deep neural network model, pruning, retraining and scaling optimization, generating descendant blocks, and selecting the descendant block combination method according to the task characteristics, the dynamic optimization of the model is achieved.

Benefits of technology

It significantly improves the computing efficiency and deployment flexibility of DNN models on edge devices, improves resource utilization, enhances model flexibility and scalability, and improves the performance of the model in multitasking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783744B_ABST
    Figure CN119783744B_ABST
Patent Text Reader

Abstract

The present invention relates to a dynamic optimization method for deep neural network model blocks for edge computing, which belongs to the field of deep learning technology, including: extracting model blocks with different focuses from a deep neural network model; pruning the model blocks to generate offspring blocks; retraining the offspring blocks to improve accuracy and obtain labeled offspring blocks; scaling and optimizing the labeled offspring blocks based on the current resource availability and delay requirements of the system; selecting the offspring block combination method according to the task characteristics to complete the deployment of the deep neural network model. The present invention aims to solve the problem that deep neural network models are difficult to efficiently deploy on resource-constrained edge devices due to their complex structure and large amount of calculation. By deeply analyzing and processing the deep learning model, the computing efficiency and deployment flexibility of the model in different computing environments are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dynamic optimization method for a deep neural network model block for edge computing, which belongs to the field of deep learning technology, and in particular to an edge computing method based on dynamic optimization of a deep neural network (DNN) model block, aiming to improve the computing efficiency and deployment flexibility of the DNN model on the edge device. Background Art

[0002] Deep neural networks (DNNs) have been widely used in many fields due to their complex structure and high-performance solutions. However, DNN models have large computational complexity, long training time, high hardware requirements, and are difficult to deploy on edge devices with limited computing resources. Deep neural networks (DNNs) have achieved remarkable results in many fields due to their powerful feature learning and task processing capabilities. However, their complex structure and huge computational complexity lead to strict requirements on hardware resources, and deployment in edge computing devices or scenarios with limited resources faces challenges. Traditional monolithic DNN models cannot be flexibly adjusted according to different computing environments, which limits their application scope and efficiency. Therefore, it is crucial to develop a technology that can effectively split, optimize, and dynamically select and combine DNN models.

[0003] With the widespread application of neural network technology in various fields, the scale of models continues to increase, and the demand for computing resources and storage resources is also increasing. In practical applications, different tasks have different focuses on models. Traditional monolithic neural network models are inefficient and inflexible when processing multiple tasks. Therefore, how to effectively decompose the neural network model so that it can be flexibly combined according to different tasks has become an important direction of current research. In the existing technology, although there are some methods for model compression and optimization, most of them only focus on a single technical means, such as simple weight pruning or quantization technology, and fail to consider the structural decomposition, module optimization and dynamic combination of the model as a whole, and cannot fully meet the comprehensive requirements for model efficiency and flexibility in edge computing scenarios. Therefore, there is an urgent need for an innovative technical solution to break through these limitations and realize the efficient deployment and application of DNN models on edge devices. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a dynamic optimization method for deep neural network model blocks for edge computing, aiming to solve the problem that deep neural network models are difficult to efficiently deploy on resource-constrained edge devices due to their complex structure and large amount of computation. By in-depth analysis and processing of deep learning models, the computing efficiency and deployment flexibility of the models in different computing environments are improved.

[0005] The present invention adopts the following technical solution:

[0006] A method for dynamic optimization of deep neural network model blocks for edge computing includes the following steps:

[0007] S1: Extract model blocks with different focuses from the deep neural network model;

[0008] S2: Prune the model block to generate offspring blocks;

[0009] S3: Retrain the offspring block to improve accuracy and obtain the labeled offspring block;

[0010] S4: Optimize the scaling of labeled descendant blocks based on the current resource availability and latency requirements of the system;

[0011] S5: Select the combination method of the offspring blocks according to the characteristics of the task and complete the deployment of the deep neural network model.

[0012] The present invention accurately extracts model blocks with different focuses from the deep learning model, and then uses a variety of advanced optimization techniques to deeply compress and improve the performance of these model blocks, generate descendant blocks, and strictly retrain the descendant blocks to improve accuracy, and finally optimizes them into a model that perfectly adapts to the specific computing environment, thereby significantly improving the computing efficiency and deployment flexibility of the DNN model on the edge device.

[0013] Preferably, the process of extracting the model block in step S1 includes:

[0014] S11: Through computational graph analysis technology, the neural network model is represented as a directed acyclic graph (DAG), where nodes represent computing operations (such as convolution, full connection, etc.), edges represent data flow, and the back propagation algorithm is used to calculate the gradient contribution of each node to the final classification result. , according to the gradient contribution The size will satisfy The node collection of Represents the gradient contribution threshold;

[0015] S12: Define the functional similarity metric function And data association function , which are used to measure the two sub-modules in the core computing module and Functional similarity and data association closeness, set the function similarity threshold Threshold for close correlation with data ,when and When and into different model blocks, and in other cases and Divide into the same model blocks;

[0016] S13: Construct the common knowledge module and task-specific knowledge module of the model block. For the construction of the common knowledge module, the principal component analysis (PCA) technique is used. Suppose the input matrix of the model block is , calculated by principal component analysis algorithm The covariance matrix of :

[0017] ;

[0018] in, Represents the input matrix The number of samples;

[0019] Covariance matrix Perform eigenvalue decomposition to obtain eigenvalues and the corresponding eigenvector , select before The eigenvectors corresponding to the largest eigenvalues ​​form the projection matrix , the output of the public knowledge module ,The task-specific module is built on the output of the common knowledge module by adding task-specific neural network layers;

[0020] S14: setting a selection basis for the model block, and obtaining a model block that has been preliminarily decomposed and processed;

[0021] S15: For the model blocks obtained in step S14, further extract the model blocks based on the graph convolutional network (GCN) technology, combine the model blocks based on the clustering analysis method, and use reinforcement learning to optimize the combined model blocks to obtain the final model blocks.

[0022] Preferably, in step S12, the functional similarity measurement function The data correlation function is calculated based on the operation type of the submodule, input and output data characteristics and other factors. It is determined by calculating indicators such as the data flow size and data dependency between the two sub-modules.

[0023] Preferably, in step S14, the available resource vector of the designed computing environment is ,in Indicates The amount of resources (such as memory, computing power, etc.) available for each model block , establish a resource requirement vector, which represents the amount of various resources required for the operation of the model block, where Representation Model Nugget Run the required nThe amount of resources; at the same time, define the performance score function of the model block on a specific task , used to measure the model block On Task t Performance on ; Computing resource adaptation function and performance weight function :

[0024] ;

[0025] ;

[0026] Comprehensively evaluate the applicability of the model block in the current computing environment and select The model nugget or combination of model nuggets with the largest value is deployed.

[0027] Preferably, in step S15, the extraction process based on graph convolutional network (GCN) technology is:

[0028] Abstracting deep neural network models into graph structures ,in represents a collection of nodes (corresponding to the layers or modules in the model), Represents a set of edges (reflecting the data flow and dependency between modules);

[0029] Introduce graph convolutional neural network (GCN) to process the graph structure to extract key model blocks: For graph structure The adjacency matrix of is the node feature matrix (its elements may include information such as module type, parameter scale, computational complexity, etc.), is the degree matrix, ,in is the adjacency matrix If the elements in i and j If there is an edge between =1, otherwise 0; the formula for propagation through a layer of graph convolutional neural network (GCN) is as follows:

[0030] ;

[0031] in, Indicates that after The first layer of graph convolutional neural network obtained after propagation The hidden layer feature matrix of the layer, ,in is the identity matrix, used to add self-connection, for The degree matrix of Indicates The hidden layer features of the layer, Indicates The result of self-connection of the hidden layer feature matrix of the layer. represents the learnable weight matrix, No. The learnable weight matrix of the layer is the result of self-connection processing, is the activation function;

[0032] After propagation through a multi-layer graph convolutional neural network (GCN), the importance score of the node (which can be based on the final hidden layer features) is calculated. Specific dimension values ​​or comprehensive evaluations of the model), filter out key nodes and their associated subgraph structures, which are the extracted model blocks;

[0033] After extracting multiple model blocks, they need to be reasonably combined. The combination process based on the cluster analysis method is as follows:

[0034] Define the similarity measure function between model blocks , comprehensively consider the similarity of input and output data features of the model blocks (such as data dimension, data distribution, etc.), functional complementarity (through predefined functional labels or performance correlation analysis based on model blocks on specific tasks) and computing resource requirement compatibility; then use the K-Means clustering algorithm to cluster the model blocks.

[0035] Preferably, in step S15, the process of optimizing the combined model blocks by using reinforcement learning is:

[0036] The environment of reinforcement learning is defined as the operating environment of the model, including input data characteristics, computing resource constraints, and task objectives. The action space of the agent is to select different model block combination strategies;

[0037] Set state Indicates that at time step Environmental status, action Represents the selected model block combination strategy, reward function Comprehensively consider the model's performance indicators (such as accuracy, recall, etc.), resource utilization efficiency (such as memory usage, the ratio of computing time to resource constraints), and task completion under the combined strategy; the agent performs the task according to the strategy network. To select actions, the policy network can be built based on a deep neural network and optimized by training on a large number of tasks and environment samples. During the training process, the parameters of the policy network are updated according to the Bellman equation:

[0038] ;

[0039] in is a discount factor used to weigh the importance of future rewards against current rewards; for Q Value function, which means that in state Next action After that, the agent is expected to receive a cumulative reward, which takes into account the immediate reward obtained by the current action. And from the next state At the beginning, the future rewards that can be obtained by following the optimal strategy; by continuously training the intelligent agent, the intelligent agent can dynamically select the optimal model block combination strategy according to different environmental states, thereby maximizing model performance and optimizing resource utilization.

[0040] Preferably, step S2 specifically includes:

[0041] S21: Prune model blocks based on weights and gradients;

[0042] For weight pruning, first, assume that the weight matrix of a layer of the neural network is ,in Indicates Neuron to The connection weights of neurons, calculate the absolute value of each weight ( ), set a threshold ,if , the connection corresponding to the weight is cut off; then, for the neuron output amplitude pruning, during the training process, for the neurons, input After samples, the output value is , calculate the average output amplitude of the neuron , set the threshold ,when When Prune neurons and their connections;

[0043] For gradient pruning, during the back propagation process, the gradient amplitude of the weight is calculated. , whose gradient is , the gradient amplitude is ,in is the loss function, sorts and prunes the weights according to the gradient magnitude, and sets the threshold ,when When , cut off the corresponding connection;

[0044] S22: Reduce the number of layers of model blocks;

[0045] Multiple adjacent layers are merged into a new layer through layer fusion technology. At the same time, the layer depth is dynamically adjusted according to the task requirements and data characteristics of the model block;

[0046] S23: quantify model parameters;

[0047] Quantization technology is used to convert high-precision model parameters into low-precision representations, and different parts are quantized with different precisions according to the importance of the parameters;

[0048] S24: model block parameter sharing;

[0049] Find the part of the model block that can share parameters, and perform parameter sharing and factorization.

[0050] Preferably, step S3 specifically includes:

[0051] S31: offspring block retraining;

[0052] An independent training environment is constructed for each offspring block. During the training process, the stochastic gradient descent (SGD) optimization algorithm is used. Assume that the parameters of the offspring block are , the loss function is , for a containing Mini-batches of training samples , the stochastic gradient descent optimization algorithm updates the parameters according to the following formula at each iteration:

[0053] ;

[0054] in Indicates The parameter value at the iteration, represents the learning rate, Represents the loss function with respect to the parameter In the sample The gradient on the training data is continuously iterated to gradually adjust the parameters of the offspring block ;

[0055] In order to accurately measure the accuracy loss caused by pruning and effectively compensate for it, the accuracy index is introduced. The accuracy of the model block before pruning on the test set is , the accuracy of the offspring block after training on the same test set is , by adjusting the hyperparameters during training (such as the learning rate , Iterations etc.) and optimize the selection and processing of training data so that As small as possible until the preset accuracy approximation threshold is met , that is, when The iteration stops when

[0056] S32: label establishment of descendant blocks;

[0057] ① For the determination of runtime memory usage, assume that the descendant block contains Parameters , the memory space occupied by each parameter is (The memory usage here is related to the data type of the parameter. For example, for a 32-bit floating point parameter, it occupies 4 bytes of memory). Calculated by the following formula:

[0058] ;

[0059] ② In terms of accuracy analysis, assume that the offspring block is in the test data set Reasoning on is the input sample, is the corresponding true label, and the predicted label is obtained after inference , then the accuracy Calculated by the following formula:

[0060] ;

[0061] in, It is an indicator function that returns 1 if the condition in the brackets is true, otherwise it returns 0;

[0062] against Different representative model sparsity , for each sparsity model form, in the validation dataset Perform reasoning on it, and assume that the predicted label under the corresponding sparsity is , then the accuracy under this sparsity is:

[0063] ;

[0064] in Represents the validation set The number of samples;

[0065] Mean accuracy loss By comparing the accuracy of the original unpruned model on the same validation set Comparative calculations;

[0066] ;

[0067] ③ In terms of delay analysis, let the original block size be (The original block refers to the model block that has not been pruned and optimized, and the offspring block is the model block after pruning and optimization), the offspring block size is , then the reduction ratio of processing delay Calculated by the following formula:

[0068] ;

[0069] Multi-dimensional labels for runtime memory usage, mean accuracy loss, and reduction ratio of processing delay of descendant blocks are constructed, providing an indispensable key basis for dynamic and reasonable selection of descendant blocks under different resource and performance requirement scenarios.

[0070] Preferably, the scaling optimization process in step S4 is:

[0071] S41: Real-time monitoring of the current system resource status, including available memory capacity and latency requirements ;

[0072] S42: constructing an optimization objective function according to the labels of the generated offspring blocks to select the optimal offspring block or block combination;

[0073] The optimization objective function comprehensively considers three factors: runtime memory usage, accuracy loss, and processing delay. as follows:

[0074] ;

[0075] in, α , β , γ They are the weight coefficients for accuracy loss, runtime memory usage, and processing delay, and their values ​​directly affect the optimization objective function. The balance between , thus determining the selection strategy of the offspring blocks in the dynamic scaling optimization mechanism;

[0076] S43: dynamically selecting a descendant block or block combination that best suits current system resources and task requirements based on the optimization objective function;

[0077] When input data or resource status changes, the optimal offspring block or block combination is intelligently selected in real time based on the predefined optimization objective function.

[0078] Preferably, in step S5, for scenarios with clear task flows and fixed data processing order, a rule-based combination is selected; for tasks that require sequential processing of data, a pipeline-based combination is selected; for tasks in which different parts of the data have different importances and model block processing weights need to be dynamically allocated, a combination based on the attention mechanism is selected.

[0079] For any details not provided in the present invention, please refer to the prior art.

[0080] The beneficial effects of the present invention are:

[0081] The present invention achieves the goal of minimizing accuracy loss while meeting latency requirements by dynamically selecting the optimal offspring block, thereby improving the computational efficiency and deployment flexibility of the DNN model on edge devices, which is reflected in:

[0082] 1. Improve resource utilization: Through model decomposition and optimization, unnecessary computing and storage overhead is reduced, resource utilization efficiency is improved, and the model can run more efficiently in resource-limited environments, such as edge computing devices.

[0083] 2. Enhanced flexibility and scalability: The composability of model blocks allows customized models to be quickly built according to different task requirements without retraining the entire model, which greatly improves the flexibility and scalability of the system and can better adapt to changing task requirements and business scenarios.

[0084] 3. Improve model performance: The optimized model blocks can process relevant information more attentively in their respective tasks. At the same time, through reasonable combination and coordination mechanisms, the advantages of each model block can be fully utilized, thereby improving the performance of the overall model in multi-task processing, including accuracy, processing speed and other aspects.

[0085] 4. Easy to manage and maintain: The decomposition of model blocks makes knowledge modular and easy to manage and maintain. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0087] Figure 1 This is an overall flow chart of the dynamic optimization method of the deep neural network model block for edge computing of the present invention;

[0088] Figure 2 Detailed diagram of each step of the method for dynamic optimization of deep neural network model blocks for edge computing of the present invention;

[0089] Figure 3 Extract and simplify schematics for model blocks;

[0090] Figure 4 Schematic diagram of the scaling optimization process. DETAILED DESCRIPTION

[0091] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the drawings in the implementation of this specification, but are not limited to this. Anything not fully described in the present invention shall be based on the conventional technology in the art.

[0092] Terminology explanation:

[0093] 1. Weight pruning: Weight pruning refers to identifying and removing weight parameters in a neural network that have little impact on the final output. These pruned weights are usually considered redundant or unimportant. The pruned model has fewer parameters, making it more efficient in storage and computing while maintaining the original performance as much as possible.

[0094] 2. Neuron output amplitude pruning: Pruning is performed by comparing the average output amplitude of neurons with the threshold, and neurons and connections with low output are removed.

[0095] 3. Gradient pruning: Pruning is done by sorting the back-propagation gradient amplitude. Weights with small amplitudes are updated slowly and contribute little, and are often used in complex model training.

[0096] 4. Layer fusion technology: merge adjacent layers, such as convolution, activation and batch normalization layer fusion, optimize calculation and transmission, and adjust layer depth as needed.

[0097] 5. Quantization technology: Convert high-precision model parameters to low-precision. Mixed-precision quantization can also be used to balance compression and performance. For example, different parameters of a speech model use different precisions.

[0098] 6. Parameter sharing and factorization: Sharing of the same parameters such as convolution kernels, factorization of parameter matrices, and data-driven discovery of actionable structures, such as sharing parameters of some convolution kernels in image convolution networks.

[0099] 7. Stochastic Gradient Descent (SGD): When training a neural network, it randomly selects small batches of samples to calculate the gradient and update the parameters. It is faster than traditional gradient descent and is used for large-scale model training.

[0100] 8. Soft attention mechanism: Dynamically assign model block weights based on input data, such as weighting semantic model blocks according to content when generating text.

[0101] 9. Hard attention mechanism: select specific model blocks for processing according to data features, such as medical image segmentation, which selects blocks to process different areas according to regional features.

[0102] Example 1

[0103] A dynamic optimization method for deep neural network model blocks for edge computing, such as Figure 1 , Figure 2 As shown, the following steps are included:

[0104] S1: Extract model blocks with different focuses from the deep neural network model;

[0105] S2: Prune the model block to generate offspring blocks;

[0106] S3: Retrain the offspring block to improve accuracy and obtain the labeled offspring block;

[0107] S4: Optimize the scaling of labeled descendant blocks based on the current resource availability and latency requirements of the system;

[0108] S5: Select the combination method of the offspring blocks according to the characteristics of the task and complete the deployment of the deep neural network model.

[0109] Example 2

[0110] A method for dynamic optimization of a deep neural network model block for edge computing, as shown in Example 1, except that the process of extracting the model block in step S1 includes:

[0111] S11: Through computational graph analysis technology, the neural network model is represented as a directed acyclic graph (DAG), where nodes represent computing operations (such as convolution, full connection, etc.) and edges represent data flow. For a given task type, such as image classification tasks, the back propagation algorithm is used to calculate the gradient contribution of each node to the final classification result. , let the loss function of the model be , for a point in the model , whose gradient contribution It can be calculated by the formula:

[0112] ;

[0113] According to the gradient contribution The size will satisfy The node collection of represents the gradient contribution threshold. In this embodiment, ; Usually choose the gradient contribution exceeding a certain threshold The nodes and their associated nodes and edges are used as the core calculation path, that is, The node set of the core computing module and the data interaction path is the basis. This method can effectively identify the computational part that is critical to the task and provide a basis for subsequent decomposition. When different types of neural networks (such as recurrent neural networks) are applied to computational graph analysis, the loop structure needs to be specially processed. The loop part can be expanded a certain number of steps to construct a DAG, or special nodes and edges can be used to mark the cyclic dependency to ensure the accuracy of calculating the gradient contribution and determining the core module.

[0114] S12: Model decomposition based on functional independence and data association;

[0115] When decomposing a large model into several model blocks, a cluster analysis algorithm is used to define a functional similarity measurement function And data association function , which are used to measure the two sub-models in the core computing module and Functional similarity and data association closeness, set the function similarity threshold Threshold for close correlation with data ,when and When and into different model blocks, and in other cases and Divide into the same model blocks;

[0116] S13: Construct the common knowledge module and task-specific knowledge module of the model block. For the construction of the common knowledge module, the principal component analysis (PCA) technique is used. Suppose the input matrix of the model block is , calculated by principal component analysis algorithm The covariance matrix of :

[0117] ;

[0118] in, Represents the input data matrix The number of samples;

[0119] Using the PCA function of the sklearn library in Python, the covariance matrix Perform eigenvalue decomposition to obtain eigenvalues and the corresponding eigenvector , select before The eigenvectors corresponding to the largest eigenvalues ​​form the projection matrix ,in Determined by the proportion of explained variance, usually the cumulative explained variance exceeds a certain proportion (such as 80%). Value, the output of the public knowledge module ,The task-specific module is constructed by adding task-specific neural network layers based on the output of the common knowledge module. For example, in the image classification task, the task-specific module can add a fully connected layer to perform classification calculations based on the features extracted by the common knowledge module;

[0120] S14: setting a selection basis for the model block, and obtaining a model block that has been preliminarily decomposed and processed;

[0121] S15: For the model blocks obtained in step S14, further extract the model blocks based on the graph convolutional network (GCN) technology, combine the model blocks based on the clustering analysis method, and use reinforcement learning to optimize the combined model blocks to obtain the final model blocks.

[0122] Example 3

[0123] A method for dynamic optimization of deep neural network model blocks for edge computing, as shown in Example 2, except that in step S12, the functional similarity measurement function The data correlation function is calculated based on the operation type of the submodule, input and output data characteristics and other factors. It is determined by calculating indicators such as the data flow size and data dependency between the two sub-modules.

[0124] For example, in the image recognition and processing model, the feature extraction convolution layer and the classification fully connected layer belong to different blocks if the conditions are met. The functional similarity measurement function can be specifically designed as follows: for the operation type, set the similarity scores of different operation types, such as the high similarity between convolution and convolution operations, and the low similarity between convolution and fully connected operations; for the input and output data features, calculate the differences in statistics such as data dimension, mean, variance, etc. as similarity measurement indicators, and combine these factors to obtain the functional similarity. The data association function can set the association score for modules with direct data transmission according to the data flow direction and traffic size. The correlation is high when the traffic is large, and the correlation is low when there is no direct transmission, so as to accurately divide the model blocks. Functional similarity threshold Threshold for close correlation with data , can be flexibly selected according to the actual situation, for the function similarity threshold For highly correlated operations (such as convolutional layers and convolutional layers), a higher value (such as 0.7-0.9) can be used. For less correlated operations (such as convolutional layers and fully connected layers), a lower value (such as 0.3-0.5) can be used.

[0125] For the data correlation density threshold For directly dependent modules (such as feature extraction and motion detection), a higher value (such as 0.7-0.9) can be used. For indirectly dependent modules (such as feature extraction and classification), a lower value (such as 0.3-0.5) can be used.

[0126] Example 4

[0127] A method for dynamic optimization of deep neural network model blocks for edge computing, as shown in Example 3, except that in step S14, the available resource vector of the design computing environment is ,in Indicates The amount of resources (such as memory, computing power, etc.) available for each model block , establish the resource demand vector , which indicates the amount of various resources required to run the model block. Representation Model Nugget Run the required n The specific value can be obtained by evaluating and measuring the computational complexity, memory usage, storage requirements, etc. of the model block. For example, for memory requirements, assuming that the model block have k parameters, each parameter occupies s bytes, the memory requirement is .

[0128] At the same time, define the performance score function of the model block on a specific task , used to measure the model block On Task t Performance on performance; performance score function You can do this in the task t The performance indicators of the test model block, such as accuracy and recall, are calculated. The specific calculation method can be determined according to the requirements of the task and the evaluation criteria. For example, for classification tasks, the performance score function is the calculation formula of accuracy: ; For regression tasks, the performance score function is the calculation formula of the mean square error.

[0129] Computing resource fitness function and performance weight function :

[0130] ;

[0131] ;

[0132] Comprehensively evaluate the applicability of the model block in the current computing environment and select The model nugget or combination of model nuggets with the largest value is deployed.

[0133] Example 5

[0134] A method for dynamic optimization of deep neural network model blocks for edge computing, as shown in Example 4, except that in step S15, the extraction process based on graph convolutional network (GCN) technology is:

[0135] Given the complexity of the deep neural network model structure, the deep neural network model can be abstracted into a graph structure. ,in represents a collection of nodes (corresponding to the layers or modules in the model), Represents a set of edges (reflecting the data flow and dependency between modules);

[0136] Introduce graph convolutional neural network (GCN) to process the graph structure to extract key model blocks: For graph structure The adjacency matrix of is the node feature matrix (its elements may include information such as module type, parameter scale, computational complexity, etc.), is the degree matrix, ,in is the adjacency matrix If the elements in i and j If there is an edge between =1, otherwise 0; the formula for propagation through a layer of graph convolutional neural network (GCN) is as follows:

[0137] ;

[0138] in, Indicates that after The first layer of graph convolutional neural network obtained after propagation The hidden layer feature matrix of the layer, ,in is the identity matrix, used to add self-connection, for The degree matrix of Indicates The hidden layer features of the layer, Indicates The result of self-connection of the hidden layer feature matrix of the layer. represents the learnable weight matrix, No. The learnable weight matrix of the layer is the result of self-connection processing, is the activation function;

[0139] After propagation through a multi-layer graph convolutional neural network (GCN), the importance score of the node (which can be based on the final hidden layer features) is calculated. Specific dimension values ​​or comprehensive evaluations of the model), filter out key nodes and their associated subgraph structures, which are the extracted model blocks; for example, set the importance score threshold , when the node Importance score of When a node is found, it and its adjacent nodes are included in the model block. For example, in image classification tasks with different focuses, assuming that in medical image segmentation tasks, the accuracy requirements for the model block are high, the threshold can be set to 0.8. Assuming that in ordinary daily image classification tasks, the computational efficiency requirements are high, the threshold can be set to 0.4.

[0140] After extracting multiple model blocks, they need to be reasonably combined. The combination process based on the cluster analysis method is as follows:

[0141] Define the similarity measure function between model blocks , comprehensively consider the similarity of input and output data features of the model block (such as data dimension, data distribution, etc.), functional complementarity (through predefined functional labels or based on the performance correlation analysis of the model block on a specific task), and compatibility of computing resource requirements (such as memory usage, computing time, etc.); for data feature similarity, the cosine similarity of the input data can be calculated:

[0142] ;

[0143] Representation Model Nugget and The similarity of input data features, Representation Model Nugget ), Representation Model Nugget The numerator is the dot product and the denominator is the modulo product.

[0144] Functional complementarity can be quantified by testing the performance improvement of different combinations of model blocks on a benchmark dataset. For example, in an image recognition task, the recognition accuracy is defined as the performance indicator of the model. Assuming that there is a model block and Their performance indicators are and . Combine the two model blocks into a new model block , the combined model block is tested on the same benchmark dataset, and its performance index is recorded as , the performance improvement can be defined as:

[0145] ;

[0146] To calculate resource requirement compatibility, calculate the sum of the reciprocals of the resource requirement ratios:

[0147] ;

[0148] in Representation Model Nugget About Resources demand.

[0149] Combining the similarity of input and output data features, functional complementarity, and computing resource requirement compatibility of the above model blocks, a weighted calculation is performed to obtain the similarity matrix , is the number of model blocks:

[0150] ;

[0151] in, is the weight of data feature similarity, is the weight of functional complementarity, is the weight of the compatibility of computing resource requirements. For example, for image recognition of road traffic signs, this task has high requirements for accuracy and processing time. Since the characteristics of traffic sign images are obvious, the resource requirements for processing are relatively small. Therefore, more attention should be paid to the impact of the similarity of image data dimensions, distribution and other features on model performance. The close coordination of model blocks with different functions has a greater impact on the final recognition accuracy. , , .

[0152] Then the K-Means clustering algorithm is used to cluster the model blocks and randomly initialize Cluster Centers , by iteratively updating the cluster center and assigning the model block to the nearest cluster center until the cluster center no longer changes: the cluster center is the "representative point" of each cluster cluster, K represents the number of cluster centers (cluster clusters), k represents the index of the cluster, k The value range is 1 to K :

[0153] ;

[0154] in Indicates k The model blocks in the same cluster will be combined together to form a combination of model blocks with collaborative working capabilities to adapt to different task requirements and computing environments.

[0155] Preferably, in step S15, the process of optimizing the combined model blocks by using reinforcement learning is:

[0156] In order to further improve the adaptability of model block combinations in different tasks and environments, the reinforcement learning method is introduced. The environment of reinforcement learning is defined as the operating environment of the model, including input data characteristics, computing resource constraints, and task goals. The action space of the agent is to select different model block combination strategies. In reinforcement learning, the agent refers to an entity that can take actions according to the state of the environment and learn the optimal strategy.

[0157] Set state Indicates that at time step Environmental status, action Represents the selected model block combination strategy, reward function Comprehensively consider the model's performance indicators (such as accuracy, recall, etc.), resource utilization efficiency (such as memory usage, the ratio of computing time to resource constraints), and task completion under the combined strategy; the agent performs the task according to the strategy network. To select actions, the policy network can be built based on a deep neural network and optimized by training on a large number of tasks and environment samples. During the training process, the parameters of the policy network are updated according to the Bellman equation:

[0158] ;

[0159] in is a discount factor used to weigh the importance of future rewards against current rewards; for Q Value function, which means that in state Next action After that, the agent is expected to receive a cumulative reward, which takes into account the immediate reward obtained by the current action. And from the next state At the beginning, the future rewards that can be obtained by following the optimal strategy; by continuously training the intelligent agent, the intelligent agent can dynamically select the optimal model block combination strategy according to different environmental states, thereby maximizing model performance and optimizing resource utilization.

[0160] Through the above-mentioned GCN-based extraction technology, the combined method of cluster analysis and the optimization strategy of reinforcement learning, a comprehensive and efficient model block extraction and combination system is constructed, which provides strong support for the application of deep neural network models in resource-constrained environments such as edge computing.

[0161] Example 6

[0162] A method for dynamic optimization of a deep neural network model block for edge computing, as shown in Example 5, is different in that the importance of neuron connections in the model block is analyzed, connections with little impact on model output are removed, and key connections are retained to ensure model performance, such as Figure 3 As shown, step S2 specifically includes:

[0163] S21: Prune model blocks based on weights and gradients;

[0164] For weight pruning, first, assume that the weight matrix of a layer of the neural network is ,in Indicates Neuron to The connection weights of neurons, calculate the absolute value of each weight ( ), set a threshold ,if , the connection corresponding to the weight is pruned, that is, the connection with a smaller weight is pruned, such as the connection with an absolute value of weight less than 0.001 and an activation frequency less than 10%; then, for neuron output amplitude pruning, during the training process, for the neurons, input After samples, the output value is , calculate the average output amplitude of the neuron , set the threshold ,when When Prune neurons and their connections;

[0165] For gradient pruning, during the back propagation process, the gradient amplitude of the weight is calculated. , whose gradient is , the gradient amplitude is ,in is the loss function, sorts and prunes the weights according to the gradient magnitude, and sets the threshold ,when When , the corresponding connection is pruned; because the weights with small gradient amplitudes are updated slowly during training, their contribution to the model learning is relatively small, and the impact on model performance after pruning is relatively small. In the training of some complex deep learning models, this method can be used to simplify the model structure and improve the inference speed without significantly reducing the model accuracy.

[0166] S22: Reduce the number of layers of model blocks;

[0167] Through layer fusion technology, multiple adjacent layers are merged into a new layer. For example, the convolution layer and the subsequent activation layer and batch normalization layer are fused to optimize the data transmission and calculation process between layers. At the same time, according to the task requirements and data characteristics of the model block, the layer depth is dynamically adjusted to avoid performance problems caused by excessive compression or retaining too many layers. After repeated experiments, 8-12 layers can be reserved for the feature extraction model block to ensure rich feature extraction; the classification and target detection model blocks are set with 3-5 layers and 4-6 layers respectively to prevent overfitting and improve computing efficiency.

[0168] S23: quantify model parameters;

[0169] Quantization technology is used to convert high-precision model parameters into low-precision representations, such as 8-bit quantization or lower-precision quantization methods. On this basis, a mixed-precision quantization strategy is developed to quantize different parts with different precisions according to the importance of the parameters, such as using higher precision for key weight parameters and lower precision for bias parameters or parameters in less sensitive layers, in order to balance the relationship between model compression and performance.

[0170] In this embodiment, a mixed precision quantization technique is adopted. After sensitivity analysis of different parameters during model training and inference, 4-bit quantization is adopted for the convolution kernel parameters with low correlation with texture features in the feature extraction model block, and 8-bit quantization is adopted for the key contour feature extraction convolution kernel and the key weights of the classification and target detection model blocks.

[0171] S24: model block parameter sharing;

[0172] Find the part of the model block that can share parameters, and perform parameter sharing and factorization. For example, in a convolutional neural network, for convolution kernels with the same receptive field and feature map size, force them to share parameters; factorize the parameter matrix, decompose the large matrix into the product of multiple small matrices, and reduce the parameter storage requirements, such as reducing the storage volume by 40%. At the same time, based on a data-driven approach, automatically discover the shareable and decomposable parameter structure to improve the compression effect.

[0173] The offspring blocks obtained after a series of processing have a more concise structure and fewer parameters than the original model blocks, and can run efficiently while ensuring accuracy.

[0174] Example 7

[0175] A method for dynamic optimization of a deep neural network model block for edge computing, as shown in Example 6, except that step S3 specifically includes:

[0176] S31: offspring block retraining;

[0177] An independent training environment is constructed for each offspring block. During the training process, the stochastic gradient descent (SGD) optimization algorithm is used. Assume that the parameters of the offspring block are , the loss function is , for a containing Mini-batches of training samples , the stochastic gradient descent optimization algorithm updates the parameters according to the following formula at each iteration:

[0178] ;

[0179] in Indicates The parameter value at the iteration, represents the learning rate, Represents the loss function with respect to the parameter In the sample The gradient on the training data is continuously iterated to gradually adjust the parameters of the offspring block ;

[0180] In order to accurately measure the accuracy loss caused by pruning and effectively compensate for it, the accuracy index is introduced. The accuracy of the model block before pruning on the test set is , the accuracy of the offspring block after training on the same test set is , by adjusting the hyperparameters during training (such as the learning rate , Iterations etc.) and optimize the selection and processing of training data so that As small as possible until the preset accuracy approximation threshold is met , that is, when The iteration stops when

[0181] In this embodiment, the initial learning rate is set to 0.001. According to the change of the loss function and the fluctuation of the accuracy of the validation set during the training process, the learning rate is adjusted at a decay rate of 0.9 every 10 epochs, and it is iterated 50-100 times.

[0182] S32: label establishment of descendant blocks;

[0183] After retraining the offspring blocks, building a comprehensive and sophisticated multi-dimensional labeling system is crucial for accurately evaluating and effectively utilizing these offspring blocks.

[0184] ① For the determination of runtime memory usage, assume that the descendant block contains Parameters , the memory space occupied by each parameter is (The memory usage here is related to the data type of the parameter. For example, for a 32-bit floating point parameter, it occupies 4 bytes of memory). Calculated by the following formula:

[0185] ;

[0186] ② In terms of accuracy analysis, assume that the offspring block is in the test data set Reasoning on is the input sample, is the corresponding true label, and the predicted label is obtained after inference , then the accuracy Calculated by the following formula:

[0187] ;

[0188] in, It is an indicator function that returns 1 when the condition in the brackets is true, otherwise it returns 0. This accuracy rate is used as an important dimension of the label to intuitively reflect the reasoning accuracy of the offspring block.

[0189] against Different representative model sparsity , for each sparsity model form, in the validation dataset Perform reasoning on it, and assume that the predicted label under the corresponding sparsity is , then the accuracy under this sparsity is:

[0190] ;

[0191] in Represents the validation set The number of samples;

[0192] Mean accuracy loss By comparing the accuracy of the original unpruned model on the same validation set Comparative calculations;

[0193] ;

[0194] This accuracy loss mean is incorporated into the label to fully reflect the accuracy changes of offspring blocks under different sparsity.

[0195] ③ In terms of delay analysis, let the original block size be (The original block refers to the model block that has not been pruned and optimized, and the offspring block is the model block after pruning and optimization), the offspring block size is , then the reduction ratio of processing delay Calculated by the following formula:

[0196] ;

[0197] Since in general, block size is closely related to processing latency, smaller blocks tend to have lower processing latency. In this way, the reduction in processing latency of subsequent blocks is accurately predicted and used as a key dimension of the label.

[0198] By constructing multi-dimensional labels for the runtime memory usage, mean accuracy loss, and reduction ratio of processing delay for the offspring blocks, we provide an indispensable key basis for dynamic and reasonable selection of offspring blocks in different resource and performance requirement scenarios, ensuring that the optimization and application of the entire deep learning model can be carried out efficiently and accurately.

[0199] Example 8

[0200] A dynamic optimization method for deep neural network model blocks for edge computing, as shown in Example 7, is different in that in an edge computing environment, system resources (such as memory capacity and computing power) and task latency requirements are dynamically changing. In order to achieve optimal model performance under different resource and performance constraints, a dynamic scaling optimization mechanism is designed, such as Figure 4 As shown, the scaling optimization process in step S4 is:

[0201] S41: Real-time monitoring of the current system resource status, including available memory capacity and latency requirements ; Available memory capacity The system's current memory size used for model operation, with a monitoring accuracy of 0.1MB; delay requirement is the upper limit of the processing time required for the task, that is, the maximum allowed delay of model inference.

[0202] S42: constructing an optimization objective function according to the labels of the generated offspring blocks to select the optimal offspring block or block combination;

[0203] The optimization objective function comprehensively considers three factors: runtime memory usage, accuracy loss, and processing delay. as follows:

[0204] ;

[0205] in, α , β , γ They are the weight coefficients for accuracy loss, runtime memory usage, and processing delay, and their values ​​directly affect the optimization objective function. The balance between the weight coefficients determines the selection strategy of the next generation blocks in the dynamic scaling optimization mechanism; α , β , γ The value of needs to be determined according to the specific application scenario, task requirements, and the importance of system resources. For example, in the image recognition task on the edge device, since the image recognition task has high accuracy requirements, accuracy loss is a key indicator. α =0.5. The memory resources of edge devices are limited, and the memory usage needs to be controlled. β =0.3. The processing delay requirement for non-real-time tasks is not strict, so γ =0.2.

[0206] S43: According to the optimization objective function, dynamically select the offspring block or block combination that best suits the current system resources and task requirements, specifically:

[0207] A. Filter descendant blocks: According to the current available memory capacity of the system, filter out descendant blocks that meet the memory constraints.

[0208] B. Calculate the optimization objective function value: For each selected offspring block or block combination, calculate the optimization objective function The value of .

[0209] C. Select the optimal offspring block: Select the offspring block or block combination with the smallest optimization objective function value as the optimal solution for the current task.

[0210] D. Real-time adjustment: When the system resource status or task requirements change, re-execute the above steps and dynamically adjust the selected descendant blocks or block combinations to ensure the optimal balance between model performance and resource utilization.

[0211] Through the above dynamic scaling optimization mechanism, the present invention can automatically select the most appropriate descendant block or block combination under different system resources and task delay requirements to achieve:

[0212] Minimize accuracy loss: Select the offspring block with the smallest accuracy loss while meeting resource constraints, ensuring that the model performance is as close to the original model as possible.

[0213] Optimize resource utilization: Dynamically adjust the model's memory usage and computing requirements based on the system's available memory and computing power to avoid wasting resources.

[0214] Meeting latency requirements: Ensure that the model inference time meets the maximum allowed latency of the task by selecting the descendant block that has the highest reduction in processing latency.

[0215] Example 9

[0216] A method for dynamic optimization of deep neural network model blocks for edge computing, as shown in Example 8, except that in step S5, for scenarios with clear task flows and fixed data processing order, a rule-based combination is selected; for tasks that require sequential processing of data, a pipeline-based combination is selected; for tasks in which different parts of the data have different importances and require dynamic allocation of processing weights for subsequent blocks, a combination based on an attention mechanism is selected.

[0217] Specifically, rule-based combination: According to the task flow, the combination order of the descendant blocks is pre-determined, and they are connected in a fixed order to form a processing pipeline. For example, in a natural language processing task, the text is first input into the word segmentation descendant block, and then processed by the descendant blocks such as part-of-speech tagging and syntactic analysis.

[0218] Set conditional rules to select different descendant block combination paths based on the type, length or characteristics of the input data. For example, for short text input, a simple descendant block combination is used; for long text input, a descendant block for processing long sequence data is added to adapt to task requirements in different data scenarios.

[0219] Pipeline-based combination: Build a linear pipeline, arrange the descendant blocks in sequence, and process the data through each descendant block in sequence. Each descendant block operates on the output of the previous descendant block and passes the result to the next descendant block. For example, in the image recognition process, the image first passes through the feature extraction descendant block, then the classification descendant block, and finally the result post-processing descendant block, forming a linear data processing process.

[0220] Set up branch pipelines to direct data to different branch descendant blocks for processing according to specific conditions. For example, in multi-language processing tasks, text is directed to the corresponding language-specific descendant block branches for processing according to the language type, and the processed results are merged into the main pipeline to achieve targeted processing of different types of data.

[0221] Combination based on Attention mechanism:

[0222] Using the soft attention mechanism, attention weights are dynamically assigned according to different parts of the input data to determine the contribution of each descendant block to the final result. For example, in the text generation task, different attention weights are assigned to descendant blocks responsible for different semantic topics according to the current generated text content, so that the most relevant descendant blocks play a greater role, thereby generating more accurate and coherent text. In the object detection model block, the soft attention mechanism is used to assign attention weights to different submodules responsible for object detection according to the feature significance of different regions in the image. For example, for areas in the image where objects may appear, higher attention weights are assigned to the corresponding detection submodules to improve the accuracy of object detection.

[0223] The hard attention mechanism is used to select specific offspring blocks to process the current data part according to the data features. For example, in the image segmentation task, the offspring block that is most suitable for processing the area is selected according to the characteristics of the image area to improve the processing efficiency and accuracy. The hard attention mechanism is used to select specific convolution kernel combinations in the feature extraction model block according to the local features of the image to process the feature extraction of different areas, so that the model can focus more on the features of the key areas and improve the effectiveness of feature extraction.

[0224] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A dynamic optimization method for deep neural network model blocks for edge computing, characterized in that: The steps include: S1: Extract model blocks with different focuses from the deep neural network model; S2: Prune the model block to generate offspring blocks; S3: Retrain the offspring block to improve accuracy and obtain the labeled offspring block; S4: Optimize the scaling of labeled descendant blocks based on the current resource availability and latency requirements of the system; S5: Select the combination of offspring blocks according to the task characteristics and complete the deployment of the deep neural network model; The process of extracting the model block in step S1 includes: S11: Through computational graph analysis technology, the neural network model is represented as a directed acyclic graph, where nodes represent computational operations and edges represent data flows. The back propagation algorithm is used to calculate the gradient contribution G of each node to the final classification result. n , according to the gradient contribution G n The size of G n >T g The node set of T g Represents the gradient contribution threshold; S12: Define the functional similarity measurement function S f (m i ,m j ) and data correlation function S d (m i ,m j ), which are used to measure the two submodules m in the core computing module. i and m j The functional similarity and data association degree of the function are set to the functional similarity threshold T f The threshold value T of the degree of correlation with the data d , when S f (m i ,m j ) <T f And S d (m i ,m j ) <T d When m i and m j Divide into different model blocks, and in other cases, m i and m j Divide into the same model blocks; S13: Construct the common knowledge module and task-specific knowledge module of the model block. For the construction of the common knowledge module, the principal component analysis technology is used. Let the input matrix of the model block be X1, and the covariance matrix C of X1 is calculated by the principal component analysis algorithm: Where N represents the number of samples in the input matrix X1; Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalue λ i and the corresponding eigenvector v i , the eigenvectors corresponding to the first k largest eigenvalues ​​are selected to form the projection matrix P. The output of the common knowledge module is Z = X1·P. The task-specific module is constructed by adding a neural network layer for a specific task based on the output of the common knowledge module; S14: setting a selection basis for the model block, and obtaining a model block that has been preliminarily decomposed and processed; S15: for the model blocks obtained in step S14, further extract the model blocks based on the graph convolutional network technology, combine the model blocks based on the clustering analysis method, and optimize the combined model blocks using reinforcement learning to obtain the final model blocks; Step S3 specifically includes: S31: offspring block retraining; An independent training environment is constructed for each offspring block. During the training process, the stochastic gradient descent optimization algorithm is used. Assuming that the parameter of the offspring block is θ and the loss function is L(θ), for a small batch of data containing m training samples, The stochastic gradient descent optimization algorithm updates parameters according to the following formula at each iteration: where θ t represents the parameter value at the tth iteration, η represents the learning rate, It represents the loss function about the parameter θ in the sample (x (i) ,y (i) ) on the training data, and gradually adjust the parameters θ of the offspring block by continuously iterating on the training data; In order to measure the accuracy loss caused by pruning and compensate effectively, the accuracy index is introduced. The accuracy of the model block before pruning on the test set is set to Accuracy pre , the accuracy of the offspring block after training on the same test set is Accuracy post , by adjusting the hyperparameters during training, Accuracy pre -Accuracy post As small as possible until the preset accuracy approximation threshold ∈ is met, that is, when |Accuracy pre -Accuracy post The iteration stops when |≤∈; S32: label establishment of descendant blocks; ① For the determination of runtime memory usage, assume that the descendant block contains m parameters p1, p2, ···, p m , the memory space occupied by each parameter is s1,s2,···,s m , then the runtime memory usage Memory is calculated by the following formula: ② In terms of accuracy analysis, assume that the offspring block is in the test data set Reasoning on i is the input sample, y i is the corresponding true label, and the predicted label is obtained after inference Accuracy original Calculated by the following formula: Among them, I() is an indicator function, which returns 1 when the condition in the brackets is met, otherwise it returns 0; For k different representative model sparsity sparsity1, sparsity2, ···, sparsity k , for each sparsity model form, in the validation dataset Perform reasoning on it, and assume that the predicted label under the corresponding sparsity is The accuracy under this sparsity is: Where n v Denotes the validation set D v The number of samples; Accuracy loss mean Loss Accuracy The accuracy of the original unpruned model on the same validation set is original Comparative calculations; ③ In terms of delay analysis, let the original block size be S original , the offspring block size is S offspring , then the reduction ratio of processing delay is Reduction delay Calculated by the following formula: Multi-dimensional labels for the reduction in runtime memory usage, mean accuracy loss, and processing latency for descendant blocks.

2. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 1, characterized in that: In step S12, the functional similarity measurement function S f (m i ,m j ) Based on the comprehensive calculation of the operation type and input and output data characteristics of the submodule, the data correlation function S d (m i ,m j ) is determined by calculating the data flow size and data dependency indicators between the two sub-modules.

3. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 2, characterized in that: In step S14, the available resource vector of the designed computing environment is R=[r1,r2,···,r i ,···,r n ], where r i Represents the available amount of the i-th resource. For each model block m b , establish the resource demand vector Indicates the amount of various resources required to run the model block, where Represents the model block m b The amount of n-th resource required to run; at the same time, define the performance score function of the model block on a specific task Used to measure the model block m b Performance on task t; Computing resource fitness function and performance weight function Comprehensively evaluate the applicability of the model block in the current computing environment and select The model nugget or combination of model nuggets with the largest value is deployed.

4. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 3, characterized in that: In step S15, the extraction process based on graph convolutional network technology is: The deep neural network model is abstracted into a graph structure G(V,E), where V represents the node set and E represents the edge set; A graph convolutional neural network is introduced to process the graph structure to extract key model blocks: Let A be the adjacency matrix of the graph structure G(V,E), X be the node feature matrix, D be the degree matrix, and D ii =∑ j A ij , where A ij is an element in the adjacency matrix A. If there is an edge between i and j, then A ij =1, otherwise 0; the formula for propagation through a layer of graph convolutional neural network is as follows: Among them, H (l+1) represents the hidden layer feature matrix of the l+1th layer obtained after propagation through the lth layer of graph convolutional neural network, Where I is the identity matrix, used to add self-connection, for The degree matrix, H (l) represents the hidden layer features of the lth layer, represents the result of the hidden layer feature matrix of the lth layer after self-connection processing, W (l) represents the learnable weight matrix, The learnable weight matrix of layer l is the result of self-connection processing, and σ is the activation function; After propagating through the multi-layer graph convolutional neural network, the key nodes and their associated subgraph structures are screened out according to the importance scores of the nodes. These subgraph structures are the extracted model blocks. After extracting multiple model blocks, they need to be reasonably combined. The combination process based on the cluster analysis method is as follows: Define the similarity measurement function Sim(M i ,M j ), comprehensively considering the similarity of input and output data features, functional complementarity, and computing resource requirement compatibility of the model blocks; then the K-Means clustering algorithm is used to cluster the model blocks.

5. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 4, characterized in that: In step S15, the process of optimizing the combined model blocks by using reinforcement learning is as follows: The environment of reinforcement learning is defined as the operating environment of the model, including input data characteristics, computing resource constraints, and task objectives. The action space of the agent is to select different model block combination strategies; Assume state s t represents the state of the environment at time step t, action a t represents the selected model block combination strategy, the reward function r(s t ,a t ) comprehensively considers the performance indicators, resource utilization efficiency and task completion degree of the model under the combined strategy; the agent follows the strategy network π(a t s t ) Select an action. During the training process, the parameters of the policy network are updated according to the Bellman equation: Where γ1 is the discount factor used to weigh the importance of future rewards and current rewards; Q(s t ,a t ) is the Q value function, which means that in state s t Next, perform action a t After that, the agent is expected to obtain the cumulative reward; by continuously training the agent, the agent can dynamically select the optimal model block combination strategy according to different environmental states to maximize model performance and optimize resource utilization.

6. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 5, characterized in that: Step S2 specifically includes: S21: Prune model blocks based on weights and gradients; For weight pruning, first, assume that the weight matrix of a certain layer of the neural network is W, where w IJ represents the connection weight from the I-th neuron to the J-th neuron, calculate the absolute value of each weight W j , set a threshold T. If |W ij | < T, then cut off the connection corresponding to the weight; then, for neuron output amplitude pruning, during the training process, for the k-th neuron, after inputting m samples, the output values o k1 , o k2 , ···, o km are obtained. Calculate the average output amplitude of the neuron Set the threshold T′. When A k < T′, prune the k-th neuron and its connections; For gradient pruning, during the back propagation process, the gradient amplitude of the weight is calculated. For the weight w, its gradient is The gradient amplitude is Where L is the loss function, the weights are sorted and pruned according to the gradient amplitude, and the threshold T is set ″ ,when When , cut off the corresponding connection; S22: Reduce the number of layers of model blocks; Multiple adjacent layers are merged into a new layer through layer fusion technology. At the same time, the layer depth is dynamically adjusted according to the task requirements and data characteristics of the model block; S23: quantify model parameters; Quantization technology is used to convert high-precision model parameters into low-precision representations, and different parts are quantized with different precisions according to the importance of the parameters; S24: model block parameter sharing; Find the part of the model block that can share parameters, and perform parameter sharing and factorization.

7. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 6, characterized in that: The scaling optimization process in step S4 is: S41: Real-time monitoring of the current resource status of the system, including available memory capacity system and delay requirement reduction system ; S42: constructing an optimization objective function according to the labels of the generated offspring blocks to select the optimal offspring block or block combination; The optimization objective function comprehensively considers the runtime memory usage, accuracy loss and processing delay. k as follows: Among them, α, β, and γ are the weight coefficients of accuracy loss, runtime memory usage, and processing delay respectively; S43: dynamically selecting a descendant block or block combination that best suits current system resources and task requirements based on the optimization objective function; When input data or resource status changes, the optimal offspring block or block combination is intelligently selected in real time based on the predefined optimization objective function.

8. The method for dynamic optimization of deep neural network model blocks for edge computing according to claim 7, characterized in that: In step S5, for scenarios with clear task flows and fixed data processing order, select a rule-based combination; for tasks that require sequential data processing, select a pipeline-based combination; for tasks where different parts of the data have different importances and model block processing weights need to be dynamically allocated, select an attention-based combination.

Citation Information

Patent Citations

  • A neural network pruning quantization method based on retraining

    CN109635936A

  • Neural network automatic pruning method and device and electronic equipment

    CN111967591A