A method for automatically generating lightweight models in a cloud-edge-device collaborative system
By utilizing NAS optimization strategies and hypernetwork structures in the cloud edge-end collaborative system, a lightweight image processing model suitable for industrial monitoring tasks is automatically generated, solving the problems of edge-end deployment and real-time processing, and achieving efficient and reliable model construction and deployment.
Patent Information
- Application Number
- CN202310059697.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-01-16
AI Technical Summary
In an industrial production environment, image processing models have problems such as resource incompatibility and limited processing speed in edge-end deployment and real-time processing. The existing lightweight model construction technology relies on manual design, which is time-consuming and suboptimal.
Design an automatic generation method of lightweight models in cloud edge-end collaborative systems. Based on the existing hypernet structure, NAS is used as the optimization strategy for subnet search, and adapted lightweight models are generated and trained based on actual task goals and equipment resources.
By introducing temperature factor optimization search strategy, the model search time is shortened, the generated model is more reliable and reliable, the number of model layers is reduced, lightweight construction is realized, model accuracy is improved, and parameter quantity is reduced.
Smart Images

Figure CN116011520B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of automatic machine learning, and specifically relates to a method for automatically generating lightweight models in a cloud-edge-end collaborative system. Background Art
[0002] In industrial production environments, data collection and monitoring are important foundations for high-end intelligent manufacturing. Their theories, technologies, and development methods face huge challenges in the transformation of industry to digitalization, networking, and intelligence. In actual industrial monitoring system environments, the main problem is that image processing models in industrial environments are difficult to deploy on the edge and process in real time. In actual applications, there is an incompatibility between the size of edge device resources and the size of image processing models. Models that are too large cannot be deployed on the edge with limited resources, and the processing speed of the model will also be affected due to the limitation of edge computing power. According to the research on model compression and acceleration by Deng Yunbin et al. in the paper "Deeplearning on mobile devices: a review", the construction technology of traditional lightweight models relies on manually designed heuristic models, which requires designers to explore a large design space and make trade-offs between size, speed, and accuracy through manual design, which is usually suboptimal and time-consuming.
[0003] Commonly used construction schemes for lightweight models include pruning, subnet search, model distillation, etc. For example, the model compression technology using NAS for pruning, in the paper "HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers", proposed to use NAS (Neural Architectures Search) to search for lightweight high-resolution network models to find effective and accurate networks for different tasks. Its main method is to update the search space and search strategy of NAS according to the task objectives, and design a lightweight transformer whose computational complexity can change dynamically according to different objective functions and resource budgets. And the model compression technology built by subnet, in the paper "Exploration and Estimation for Model Compression", it mainly defines the direct search subnet as a non-smooth, non-convex, complex non-deterministic polynomial integer programming problem. The subnet is used as a sample of multi-source Bernoulli distribution, and the optimal subnet is found in a way to solve an approximate continuous problem.
[0004] In "HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers", NAS search is effective, but it often requires a lot of training time. Due to the large amount of monitoring data in industrial environments, it will affect the efficiency of the search. Compared with direct NAS search, the method of using supernet to search subnets in "Exploration and Estimation for Model Compression" is more reliable and faster because there are prior effective models in general industrial environments. However, in the environment of cloud-edge collaboration, the optimization strategy of alternating exploration and estimation using probabilistic models is time-consuming and prone to falling into local optimality. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention designs a method for automatically generating lightweight models in a cloud-edge-end collaborative system. Based on an existing or a priori artificially designed supernet structure, NAS is used as an optimization strategy for subnet search, so that an adapted lightweight model can be obtained and trained according to actual task objectives and equipment resources, thereby completing the deployment of a lightweight environment in a cloud-edge-end collaborative industrial monitoring system and helping to realize intelligent industrial production.
[0006] A method for automatically generating a lightweight model in a cloud-edge-device collaborative system, specifically comprising the following steps:
[0007] Step 1: Build a supernet model and select it;
[0008] According to the target tasks and the effects of the original model in actual industrial scenarios, a supernet model structure for NAS is designed as the actual search space; the computational operations of the effective computational units of the original model in actual industrial scenarios are extracted, including convolution, pooling, and residual computational operations, with the purpose of forming a complete network model through effective computational operations; the specific computational operations in each unit of the designed NAS supernet model structure are obtained by subsequent searches; the only thing that needs to be determined manually is the number of layers in the unit, that is, the number of nodes; the connection between nodes is the specific operation, which is represented as a sub-operation set in the supernet and a single sub-operation in the subnet;
[0009] The supernet model as a whole is a backbone network structure composed of 8 stacked units, each unit is a mixture of CNN and GCN parts: the CNN part is a supernet model structure designed according to GoogleNet, each unit is calculated by four layers, and the calculation operations used include: no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution and residual connection; the GCN part is a lightweight structure of glorify_unit, and the number of channels of the input image feature after dimensionality reduction is the value that determines the region size when the search is executed, so it is used as one of the optimization targets of the search;
[0010] Step 2: Train, evaluate and update the parameters of the supernet model obtained in step 1;
[0011] Step 2.1: Process the image features input into the supernet model;
[0012] Step 2.1.1: If the unit for image feature input this time is the first unit, the first unit input includes a randomly initialized feature sequence and original image features;
[0013] Step 2.1.2: If the unit of this image feature input is not the first unit, the unit input is the output of the previous unit and the original image feature;
[0014] Step 2.1.3: If the unit number of the image feature input is divisible by 3, the number of input channels of the unit is divided by 2 and a downsampling is performed;
[0015] Step 2.2: The image features input in the unit are propagated from top to bottom in the nodes. The weighted sum of all calculation operations is used to obtain the output value and pass it to the next node. The temperature factor in distillation learning is introduced to improve the discrimination of the weights. The calculation formula is as follows:
[0016]
[0017] Where i, j represent the i-th node and j-th point; o is the selected sub-operation, O is the set of all sub-operations, o(x) represents the result of input x after sub-operation o; T is the temperature factor; Represents the weighted sum of the outputs of all sub-operations between two nodes:
[0018] Step 2.2.1: The input features are subjected to no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution, and residual calculation to obtain 9 values;
[0019] Step 2.2.2: If the weight values of all operations are undefined, initialize the weight values of the operations and assign a weight value to each operation;
[0020] Step 2.2.3: If the weight values of all operations are defined, do not change their values;
[0021] Step 2.2.3: Multiply each independent operation by its corresponding weight value and sum them to get the output value;
[0022] Step 2.2.4: Output the value and assign it to the next node;
[0023] Step 2.2.5: Repeat step 2.2.1 for each subsequent node of the next node until there are no subsequent nodes;
[0024] Step 2.3: After the feature calculation is completed, the feature value of each intermediate node is processed using a lightweight unit of GCN to obtain the contextual relationship between the feature regions of the farther image;
[0025] Step 2.3.1: After the feature calculation is completed, the eigenvalues of each intermediate node are calculated through two convolutional layers to obtain the reduced eigenvalues and mapping matrices;
[0026] Step 2.3.2: Multiply the reduced features and the mapping matrix to obtain a feature vector, which contains all the information of the graph;
[0027] Step 2.3.3: Perform a second one-dimensional convolution on the feature vector itself, i.e., graph convolution, to obtain the contextual relationship between different regions;
[0028] Step 2.3.4: Multiply the result of the graph convolution in step 2.3.3 by the transpose of the mapping matrix, and then perform a dimensionality-up convolution to obtain a feature map of the same size as before dimensionality reduction. This feature map is added to the input feature as the output feature of this operation;
[0029] Step 2.4: The weighted sum of the output features of different channel numbers in GCN is sent to the next node;
[0030] Step 2.5: Concatenate the eigenvalues of all nodes without subsequent nodes as the output of this unit;
[0031] Step 2.6: The output of the previous unit is used as one of the inputs of the next unit, and step 2.1 is repeated until there is no next unit;
[0032] Step 2.7: Complete the supernet training and obtain the supernet prediction results. The results are cross-entropy loss with the original results of the ImageNet dataset to update all sub-parameters in the supernet model, including model structure parameters and convolution parameters.
[0033] Step 2.8: Set the unit count to 0, restart training, and repeat the training process starting from step 2.1 until the preset round is reached;
[0034] Step 3: After completing the supernet training in step 2, select and train the subnet according to the weights in the supernet;
[0035] Step 3.1: Integrate all weights initialized in step 2, including the weights of all preceding edges of the node, the weights of all sub-operations between two nodes, and the weights of different numbers of channels of glore_unit;
[0036] Step 3.2: Process the edge weights in the supernet model to obtain the model structure;
[0037] Step 3.2.1: The input is the edge weight tensor, and the edge weight tensor is normalized using the softmax function;
[0038] Step 3.2.2: Based on the processed edge weights, sort and split according to the number of nodes, determine the weight values of all the leading edges of each intermediate node, and the sum of the weights of the leading edges of the same node should be 1;
[0039] Step 3.2.3: Select the two leading edges before each node to form the overall structure of the subnet;
[0040] Step 3.3: Process the sub-operation weights in the supernet model to obtain the model operation process;
[0041] Step 3.3.1: The input is the sub-operation weight tensor, and the sub-operation weight tensor is normalized using the softmax function;
[0042] Step 3.3.2: Based on the processed sub-operation weights, divide them according to the total number of sub-operation types, corresponding to different head and tail nodes; determine the weight values of all sub-operations between every two nodes, and the sum of the weights of all sub-operations between two nodes should be 1;
[0043] Step 3.3.3: On the selected edge, select the largest sub-operation according to the sub-operation weight to form the sub-network operation process;
[0044] Step 3.4: According to the number of channels of the same group of glore_unit, select the glore_unit structure with the largest weight;
[0045] Step 3.4.1: The input is a weight tensor with different numbers of channels, and the channel weight tensor is normalized using the softmax function;
[0046] Step 3.4.2: According to the processed channel weights, each intermediate node output has a set of glore_unit channel weights, and their sum should be 1;
[0047] Step 3.4.3: According to the weights of different numbers of channels on each glore_unit, select the glore_unit structure with the highest weight;
[0048] Step 3.5: Determine the model structure based on the selected edges, operations, and glore_unit structures, and train the sub-model on the ImageNet dataset;
[0049] Step 3.6: Decide whether to adopt the sub-model based on the sub-network requirements;
[0050] Step 4: Optimize the supernet model;
[0051] According to the actual sub-network effect, if the task matching is poor, the super-network needs to be optimized. This optimization usually includes the selection of the number of edges, corresponding to the number of layers of the model in the actual model structure; the type of sub-computation operation, including the size of the convolution, the convolution method, etc.; and then repeat steps 1 to 3 to obtain a better sub-model.
[0052] Beneficial technical effects of the present invention:
[0053] In terms of model search, compared with the prior art, the present invention can complete the same model search task in a shorter search time by introducing the temperature factor as a search strategy optimization. At the same time, in the process of sub-model selection, different edges, sub-operations, etc. will not have similar weights, making the resulting model more credible and reliable.
[0054] At the same time, the present invention also reduces the number of layers of the sub-model, thereby achieving the purpose of lightweight model construction, and by introducing additional lightweight graph neural network units, it makes up for the problem of insufficient feature information acquisition of conventional single CNN image processing network models at long distances, thereby improving the overall accuracy of the model without causing a large change in the model parameters. In addition, the automated lightweight model generation and training can be applied to different industrial monitoring tasks, and the sub-operation set can be added or modified according to the actual usage scenario, or the number of connecting edges of the model can be added or modified to meet the actual task's demand for computing resources or computing speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 The lightweight model of the embodiment of the present invention automatically generates the overall pre-training process;
[0056] Figure 2 The hybrid model search space and unit structure of the embodiment of the present invention; wherein Figure a is the overall space structure, and Figure b is the unit structure;
[0057] Figure 3Flowchart of supernet model training, evaluation and parameter update according to an embodiment of the present invention;
[0058] Figure 4 Flowchart of the subnet selection and training process according to an embodiment of the present invention. DETAILED DESCRIPTION
[0059] A method for automatically generating lightweight models in a cloud-edge-device collaborative system, as shown in the attached Figure 1 As shown, the specific steps include:
[0060] Step 1: Build a supernet model and select it;
[0061] In actual industrial scenarios, there are some original models that already have good effects, but these original models are usually not suitable for edge deployment. Therefore, according to the target tasks and the effects of the models in actual industrial scenarios, a supernet model structure for NAS is designed as the actual search space. In actual design, previous work can provide a basis for the design of supernets. The computational operations of the effective computing units of the original models in actual industrial scenarios are extracted, including convolution, pooling, and residual computing operations. The purpose is to form a complete network model through effective computing operations; the specific computing operations in each unit of the designed NAS supernet model structure are obtained by subsequent searches; the only thing that needs to be determined manually is the number of layers in the unit, that is, the number of nodes; the connection between nodes is the specific operation, which is expressed as a set of sub-operations in the supernet and a single sub-operation in the subnet;
[0062] like Figure 2 As shown in a, the supernet model imitates the structure of InceptionNet, and the overall backbone network structure is composed of 8 stacked units. Figure 2 b shows the specific structure within the unit. Each unit is a mixture of CNN and GCN parts: the CNN part is a supernet model structure designed according to GoogleNet. Each unit consists of four layers of calculations, and the calculation operations used include: no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution and residual connection; the GCN part is a lightweight structure of glorify_unit. The number of channels of the input image feature after dimensionality reduction is the value that determines the size of the region when the search is executed, so it is used as one of the optimization targets of the search;
[0063] Step 2: Train, evaluate and update the parameters of the supernet model obtained in step 1. The overall process is as follows: Figure 3 As shown;
[0064] Step 2.1: Process the image features input into the supernet model;
[0065] Step 2.1.1: If the unit for image feature input this time is the first unit, the first unit input includes a randomly initialized feature sequence and original image features;
[0066] Step 2.1.2: If the unit of this image feature input is not the first unit, the unit input is the output of the previous unit and the original image feature;
[0067] Step 2.1.3: If the unit number of the image feature input is divisible by 3, the number of input channels of the unit is divided by 2 and a downsampling is performed;
[0068] Step 2.2: The image features input in the unit are propagated from top to bottom in the nodes. The weighted sum of all calculation operations is used to obtain the output value and pass it to the next node. The temperature factor in distillation learning is introduced to improve the discrimination of the weights. The calculation formula is as follows:
[0069]
[0070] Where i, j represent the i-th node and j-th point; o is the selected sub-operation, O is the set of all sub-operations, o(x) represents the result of input x after sub-operation o; T is the temperature factor, which is set to 0.6 in this paper; Represents the weighted sum of the outputs of all sub-operations between two nodes:
[0071] Step 2.2.1: The input features are subjected to no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution, and residual calculation to obtain 9 values;
[0072] Step 2.2.2: If the weight values of all operations are undefined, initialize the weight values of the operations and assign a weight value to each operation;
[0073] Step 2.2.3: If the weight values of all operations are defined, do not change their values;
[0074] Step 2.2.3: Multiply each independent operation by its corresponding weight value and sum them to get the output value;
[0075] Step 2.2.4: Output the value and assign it to the next node;
[0076] Step 2.2.5: Repeat step 2.2.1 for each subsequent node of the next node until there are no subsequent nodes;
[0077] Step 2.3: After the feature calculation is completed, the feature value of each intermediate node is processed using a lightweight unit of GCN to obtain the contextual relationship between the feature regions of the farther image;
[0078] Step 2.3.1: After the feature calculation is completed, the eigenvalues of each intermediate node are calculated through two convolutional layers to obtain the reduced eigenvalues and mapping matrices;
[0079] Step 2.3.2: Multiply the reduced features and the mapping matrix to obtain a feature vector, which contains all the information of the graph;
[0080] Step 2.3.3: Perform a second one-dimensional convolution on the feature vector itself, i.e., graph convolution, to obtain the contextual relationship between different regions;
[0081] Step 2.3.4: Multiply the result of the graph convolution in step 2.3.3 by the transpose of the mapping matrix, and then perform a dimensionality-up convolution to obtain a feature map of the same size as before dimensionality reduction. This feature map is added to the input feature as the output feature of this operation;
[0082] Step 2.4: The weighted sum of the output features of different channel numbers in GCN is sent to the next node;
[0083] Step 2.5: Concatenate the eigenvalues of all nodes without subsequent nodes as the output of this unit;
[0084] Step 2.6: The output of the previous unit is used as one of the inputs of the next unit, and step 2.1 is repeated until there is no next unit;
[0085] Step 2.7: Complete the supernet training and obtain the supernet prediction results. The results are cross-entropy loss with the original results of the ImageNet dataset to update all sub-parameters in the supernet model, including model structure parameters and convolution parameters.
[0086] Step 2.8: Set the unit count to 0 and restart training. Repeat the training process in step 2.1 until the preset round is reached;
[0087] Step 3: After completing the supernet training in step 2, select and train the subnets according to the weights in the supernet. The process is as follows: Figure 4 As shown;
[0088] Step 3.1: Integrate all weights initialized in step 2, including the weights of all preceding edges of the node, the weights of all sub-operations between two nodes, and the weights of different numbers of channels of glore_unit;
[0089] Step 3.2: Process the edge weights in the supernet model to obtain the model structure;
[0090] Step 3.2.1: The input is the edge weight tensor, and the edge weight tensor is normalized using the softmax function;
[0091] Step 3.2.2: Based on the processed edge weights, sort and split according to the number of nodes, determine the weight values of all the leading edges of each intermediate node, and the sum of the weights of the leading edges of the same node should be 1;
[0092] Step 3.2.3: Select the two leading edges before each node to form the overall structure of the subnet;
[0093] Step 3.3: Process the sub-operation weights in the supernet model to obtain the model operation process;
[0094] Step 3.3.1: The input is the sub-operation weight tensor, and the sub-operation weight tensor is normalized using the softmax function;
[0095] Step 3.3.2: Based on the processed sub-operation weights, divide them according to the total number of sub-operation types, corresponding to different head and tail nodes; determine the weight values of all sub-operations between every two nodes, and the sum of the weights of all sub-operations between two nodes should be 1;
[0096] Step 3.3.3: On the selected edge, select the largest sub-operation according to the sub-operation weight to form the sub-network operation process;
[0097] Step 3.4: According to the number of channels of the same group of glore_unit, select the glore_unit structure with the largest weight;
[0098] Step 3.4.1: The input is a weight tensor with different numbers of channels, and the channel weight tensor is normalized using the softmax function;
[0099] Step 3.4.2: According to the processed channel weights, each intermediate node output has a set of glore_unit channel weights, and their sum should be 1;
[0100] Step 3.4.3: According to the weights of different numbers of channels on each glore_unit, select the glore_unit structure with the highest weight;
[0101] Step 3.5: Determine the model structure according to the selected edge, operation and glob_unit structure, and train the sub-model on the ImageNet dataset. The present invention uses the public dataset ImageNet, and the actual industrial scene dataset is also applicable;
[0102] Step 3.6: Decide whether to adopt the sub-model based on the sub-network requirements;
[0103] Step 4: Optimize the supernet model;
[0104] According to the actual sub-network effect, if the task matching is poor, the super-network needs to be optimized. This optimization usually includes the selection of the number of edges, corresponding to the number of layers of the model in the actual model structure; the type of sub-computation operation, including the size of the convolution, the convolution method, etc.; and then repeat steps 1 to 3 to obtain a better sub-model.
[0105] The following table compares the effects of image processing models generated by the method of the present invention and other NAS model search methods. The data set is searched using cifar100 and evaluated on imagenet. The results show that the model generated by the method of the present invention meets the requirements of lightweight and high precision, and basically does not increase the search time;
[0106]
[0107] Different from the model compression method of single subnet search or pruning through NAS, the present invention reduces the labor cost in the process of model lightweighting by combining subnet search and NAS, and uses the optimization strategy of NAS to replace the original subnet search strategy, thereby improving the efficiency of model compression. At the same time, the unified supernet ensures that all lightweight models have certain unified task objectives, making model training and deployment faster and easier.
[0108] At the same time, due to the limited resources and computing power of the edge cloud in industrial systems and the need for timely feedback of monitoring results, a lightweight graph neural network unit glore_unit is introduced. The purpose is to obtain regional contextual relationships in a longer range while ensuring that the model has a small number of parameters and a fast calculation speed. Network models with higher numbers of layers usually have better results, so there is an inevitable loss of accuracy in the lightweight process. The addition of the graph neural network unit glor_unit can still effectively improve the accuracy of the model when the number of layers of the model is reduced. This unit mainly performs graph convolution on image data, and the area size of the graph convolution is usually manually trial-and-error, which costs a lot of money and is prone to errors. Therefore, the design of the present invention uses it as an optimization target in NAS, and uniformly evaluates and automatically selects the area size of the graph convolution within a certain range.
[0109] The automatic generation and training of the lightweight model can automatically generate a small model with fewer operations and parameters. Unlike previous work, the present invention introduces a temperature factor T as a weight correction, making T less than 1 so that the softmax function that distinguishes weights can have a higher degree of distinction. Compared with other image processing models and common NAS search models, the results of the present invention are more accurate. On this basis, it can help realize the automatic generation and deployment of lightweight models in cloud-edge-end collaborative systems.
Claims
1. A method for automatically generating lightweight models in a cloud-edge-device collaborative system, characterized in that: The specific steps include: Step 1: Build a supernet model and select it; Step 2: Train, evaluate and update the parameters of the supernet model obtained in step 1; Step 3: After completing the supernet training in step 2, select and train the subnet according to the weights in the supernet; Step 4: Optimize the supernet model; Among them, step 2 is specifically as follows: Step 2.1: Process the image features input into the supernet model; Step 2.2: The image features input in the unit are propagated from top to bottom in the nodes. The weighted sum of all calculation operations is used to obtain the output value and pass it to the next node. The temperature factor in distillation learning is introduced to improve the discrimination of the weights. The calculation formula is as follows: Where i, j represent the i-th node and j-th point; o is the selected sub-operation, O is the set of all sub-operations, o(x) represents the result of input x after sub-operation o; T is the temperature factor; Represents the weighted sum of the outputs of all sub-operations between two nodes: Step 2.3: After the feature calculation is completed, the feature value of each intermediate node is processed using a lightweight unit of GCN to obtain the contextual relationship between the feature regions of the farther image; Step 2.4: The weighted sum of the output features of different channel numbers in GCN is sent to the next node; Step 2.5: Concatenate all feature values without subsequent nodes as the output of this unit; Step 2.6: The output of the previous unit is used as one of the inputs of the next unit, and step 2.1 is repeated until there is no next unit; Step 2.7: Complete the supernet training and obtain the supernet prediction results. The results are cross-entropy lost with the original results of the ImageNet dataset to update all sub-parameters in the supernet model, including model structure parameters and convolution parameters. Step 2.8: Set the unit count to 0, restart training, and repeat the training process starting from step 2.1 until the preset round is reached.
2. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 1: Construct a supernet model and select it as follows: According to the target tasks and the effects of the original model in actual industrial scenarios, a supernet model structure for NAS is designed as the actual search space; the computational operations of the effective computational units of the original model in actual industrial scenarios are extracted, including convolution, pooling, and residual computational operations, with the purpose of forming a complete network model through effective computational operations; the specific computational operations in each unit of the designed NAS supernet model structure are obtained by subsequent searches; the only thing that needs to be determined manually is the number of layers in the unit, that is, the number of nodes; the connection between nodes is the specific operation, which is represented as a sub-operation set in the supernet and a single sub-operation in the subnet; The overall supernet model is a backbone network structure composed of 8 stacked units, each unit of which is a mixture of CNN and GCN parts: the CNN part is a supernet model structure designed according to GoogleNet, and each unit consists of four layers of calculations. The calculation operations used include: no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution and residual connection; the GCN part is a lightweight structure of glor_unit, and the number of channels of the input image feature after dimensionality reduction is the value that determines the region size when the search is executed, so it is used as one of the optimization targets of the search.
3. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 2.1 is as follows: Step 2.1.1: If the unit for image feature input this time is the first unit, the first unit input includes a randomly initialized feature sequence and original image features; Step 2.1.2: If the unit of this image feature input is not the first unit, the unit input is the output of the previous unit and the original image feature; Step 2.1.3: If the unit number of the image feature input is divisible by 3, the number of input channels of the unit is divided by 2 and a downsampling is performed.
4. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 2.2 is as follows: Step 2.2.1: The input features are subjected to no operation, 3x3 average pooling, 3x3 maximum pooling, 3x3 separable convolution, 5x5 separable convolution, 7x7 separable convolution, 3x3 dilated convolution, 5x5 dilated convolution, and residual calculation to obtain 9 values; Step 2.2.2: If the weight values of all operations are undefined, initialize the weight values of the operations and assign a weight value to each operation; Step 2.2.3: If the weight values of all operations are defined, do not change their values; Step 2.2.3: Multiply each independent operation by its corresponding weight value and sum them to get the output value; Step 2.2.4: Output the value and assign it to the next node; Step 2.2.5: Repeat from step 2.2.1 for each successor node of the next node until there are no successor nodes.
5. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 2.3 is as follows: Step 2.3.1: After the feature calculation is completed, the eigenvalues of each intermediate node are calculated through two convolutional layers to obtain the reduced eigenvalues and mapping matrices; Step 2.3.2: Multiply the reduced features and the mapping matrix to obtain a feature vector, which contains all the information of the graph; Step 2.3.3: Perform a second one-dimensional convolution on the feature vector itself, i.e., graph convolution, to obtain the contextual relationship between different regions; Step 2.3.4: Multiply the result of the graph convolution in step 2.3.3 by the transpose of the mapping matrix, and then perform a dimensionality-enhancing convolution to obtain a feature map of the same size as before dimensionality reduction. This feature map is added to the input feature as the output feature of this operation.
6. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 3 is as follows: Step 3.1: Integrate all weights initialized in step 2, including the weights of all preceding edges of the node, the weights of all sub-operations between two nodes, and the weights of different numbers of channels of glore_unit; Step 3.2: Process the edge weights in the supernet model to obtain the model structure; Step 3.2.1: The input is the edge weight tensor, and the edge weight tensor is normalized using the softmax function; Step 3.2.2: Based on the processed edge weights, sort and split according to the number of nodes, determine the weight values of all the leading edges of each intermediate node, and the sum of the weights of the leading edges of the same node should be 1; Step 3.2.3: Select the two leading edges before each node to form the overall structure of the subnet; Step 3.3: Process the sub-operation weights in the supernet model to obtain the model operation process; Step 3.4: According to the number of channels of the same group of glore_unit, select the glore_unit structure with the largest weight; Step 3.5: Determine the model structure based on the selected edges, operations, and glore_unit structures, and train the sub-model on the ImageNet dataset; Step 3.6: Decide whether to adopt this sub-model based on the sub-network requirements.
7. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 6, it is characterized in that: Step 3.3 is as follows: Step 3.3.1: The input is the sub-operation weight tensor, and the sub-operation weight tensor is normalized using the softmax function; Step 3.3.2: Based on the processed sub-operation weights, divide them according to the total number of sub-operation types, corresponding to different head and tail nodes; determine the weight values of all sub-operations between every two nodes, and the sum of the weights of all sub-operations between two nodes should be 1; Step 3.3.3: On the selected edge, select the sub-operation with the largest weight to form the sub-network operation process.
8. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 6, it is characterized in that: Step 3.4 is as follows: Step 3.4.1: The input is a weight tensor with different numbers of channels, and the channel weight tensor is normalized using the softmax function; Step 3.4.2: According to the processed channel weights, each intermediate node output has a set of glore_unit channel weights, and their sum should be 1; Step 3.4.3: According to the weights of different numbers of channels on each glore_unit, select the glore_unit structure with the highest weight.
9. According to the method for automatically generating lightweight models in a cloud-edge-device collaborative system according to claim 1, it is characterized in that: Step 4 optimizes the supernet model as follows: According to the actual sub-network effect, if the task matching is poor, the super-network needs to be optimized. The optimization includes the selection of the number of edges, which corresponds to the number of layers of the model in the actual model structure; the type of sub-computation operation, including the size of the convolution and the convolution method; and then repeating steps 1 to 3 to obtain a better sub-model.
Citation Information
Patent Citations
SAR optical image mapping model lightweight method based on conditional generative adversarial network
CN114202017A
Tire condition evaluation method and system
CN114838961A