Adjustment Method, Device, and Computer-Readable Storage Medium of Neural Network Model
By adjusting the network layer of the neural network model and optimizing its structure to adapt to the memory limitations of the MCU, the problem of insufficient memory of the MCU is solved and the neural network model is efficiently run on the MCU.
Patent Information
- Application Number
- CN202210820863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-13
AI Technical Summary
MCU memory resources are limited and it is difficult to support neural network models with high memory consumption demand. The neural network model needs to be adjusted and optimized to reduce memory consumption.
By obtaining the output results of the test data, determine the adjustable variables and target input data, and adjust the network layer of the neural network model according to the network layer adjustment rules, including deleting and merging variables, reordering the network layer, and optimizing the network structure.
It effectively reduces the memory consumption of the neural network model, allowing it to be better deployed on the MCU, performs data computing and IO tasks in parallel, and reduces IO delay.
Smart Images

Figure CN115081624B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to, but is not limited to, the field of artificial intelligence technology, and in particular, to a method, device, and computer-readable storage medium for adjusting a neural network model. Background Art
[0002] In recent years, with the rapid development of artificial intelligence (AI), neural network models have been increasingly widely used in the field of artificial intelligence, gradually showing their advantages in fields such as computer vision, speech recognition, and natural language processing. At the same time, with the development of intelligent electronic devices, there are more and more types of electronic devices that can run neural network models. Microcontrollers (MCUs) have become increasingly popular due to their small size, portability, low cost, and low power consumption. The combination of MCUs and neural network models has become a popular trend. Neural network models generate a large amount of data during operation, and the network structure of the neural network models themselves also requires memory occupancy. However, MCUs do not have the traditional cache-memory-hard disk storage architecture and have less memory resources, making it difficult to support neural network models with high memory consumption requirements. Therefore, when deploying neural network models to MCUs for operation, it is necessary to adjust and optimize the neural network models. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present invention provide a method, device, and computer-readable storage medium for adjusting a neural network model, which can effectively reduce the memory consumption of the neural network model by adjusting the neural network model.
[0005] In a first aspect, embodiments of the present invention provide a method for adjusting a neural network model, including:
[0006] Obtain test data, input the test data into a preset first neural network model, and obtain a first output result;
[0007] Obtain a first adjustable variable from the first output result;
[0008] Obtain first target input data corresponding to the first adjustable variable;
[0009] Adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model.
[0010] In some embodiments, obtaining the first target input data corresponding to the first adjustable variable includes:
[0011] Obtaining a preset relationship mapping table, where the relationship mapping table represents the mapping relationship between the first output result and the network output layer, and the network output layer is the network layer that generates the output result;
[0012] Determining the network output layer corresponding to the first adjustable variable in the relationship mapping table as the target network output layer;
[0013] Obtaining the first target input data according to the target network output layer, where the first target input data is the data input to the target network output layer.
[0014] In some embodiments, the first adjustable variable includes at least two first variables; adjusting at least two first network layers in the neural network model according to the first adjustable variable, the target input data, and a preset network layer adjustment rule to obtain a second neural network model includes:
[0015] Determining a first target variable from at least two of the first variables, and obtaining first input data corresponding to the first target variable from the first target input data;
[0016] Determining the network output layer corresponding to the first target variable in the relationship mapping table as the first network output layer, and obtaining a first network layer identifier, where the first network layer identifier is a parameter that uniquely identifies the first network output layer;
[0017] Determining the sum of the memory occupancy value of the first adjustable variable and the memory occupancy value of the first input data as the first memory occupancy value;
[0018] Deleting the first target variable from the first adjustable variable, and performing data merging processing on the first adjustable variable after deleting the first target variable and the first input data to obtain a first intermediate data set, where the first intermediate data set includes a second adjustable variable, and the second adjustable variable includes at least two second variables;
[0019] Determining a second target variable from at least two of the second variables, and obtaining second target input data corresponding to the second target variable;
[0020] Determining the sum of the memory occupancy value of the second adjustable variable and the memory occupancy value of the second input data as the second memory occupancy value;
[0021] Adjust at least two of the first network layers in the neural network model according to the first memory occupancy value, the second memory occupancy value, and the first network layer identifier to obtain the second neural network.
[0022] In some embodiments, after adjusting at least two of the first network layers in the neural network model according to the first memory occupancy value, the second memory occupancy value, and the first network layer identifier to obtain the second neural network, the method further includes:
[0023] When the number of the remaining second variables in the second adjustable variable is greater than a preset first threshold value, determine the second adjustable variable as a new first adjustable variable;
[0024] Obtain new first target input data corresponding to the new first adjustable variable;
[0025] Adjust at least two second network layers in the second neural network model according to the new first adjustable variable, the new first target input data, and the network layer adjustment rule to obtain a new second neural network model.
[0026] In some embodiments, the second neural network model includes at least two second network layers connected in sequence. After adjusting at least two first network layers in the first neural network model according to the first adjustable variable, the target input data, and the preset network layer adjustment rule to obtain the second neural network model, it further includes:
[0027] Obtain the third memory occupancy value corresponding to each of the second network layers;
[0028] Determine a boundary network layer from at least two of the second network layers, and the ratio of the sum of the third memory occupancy values corresponding to each of the second network layers located before the boundary network layer to the sum of the third memory occupancy values corresponding to each of the second network layers located after the boundary network layer is greater than a preset second threshold value.
[0029] In some embodiments, after determining the boundary network layer from at least two of the second network layers, the method further includes:
[0030] Obtain data to be processed, input the data to be processed into the second neural network model to obtain a second output result, and the second output result includes at least two optional output data;
[0031] Obtain optional input data corresponding to each of the optional output data;
[0032] Obtain target output data from at least two of the optional output data, where the network output layer corresponding to the target output data is the second network layer ranked before the boundary network layer;
[0033] Obtain third target input data, where the third target input data is the input data corresponding to the target output data;
[0034] Perform data merging processing on the target output data and the third target input data to obtain a second intermediate data set;
[0035] Classify the data in the second intermediate data set according to a preset data classification rule to obtain first intermediate data and second intermediate data, where the task objective of the first intermediate data is an IO task, and the task objective of the second intermediate data is a computing task;
[0036] Divide the storage device into a first storage area and a second storage area, where the storage device is used to run the second neural network model;
[0037] Send the first intermediate data to the first storage area to perform input / output IO operations;
[0038] Send the second intermediate data to the second storage area to perform data operation operations.
[0039] In some embodiments, before inputting the data to be processed into the second neural network model, the method further includes:
[0040] Obtain a preset data cleaning rule;
[0041] Clean the data to be processed according to the data cleaning rule to obtain new data to be processed.
[0042] In a second aspect, an embodiment of the present invention provides an adjustment device for a neural network model, including:
[0043] A first output result acquisition module, which is used to acquire test data, input the test data into a preset first neural network model, and obtain a first output result;
[0044] A first adjustable variable acquisition module, which is used to acquire a first adjustable variable from the first output result;
[0045] A first target input data acquisition module, which is used to acquire first target input data corresponding to the first adjustable variable;
[0046] A neural network model adjustment module, which is used to adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model.
[0047] In a third aspect, an embodiment of the present invention further provides an adjustment device for a neural network model, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the adjustment method for the neural network model as described in the first aspect.
[0048] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, where the computer program is used to execute the adjustment method for the neural network model as described in the first aspect.
[0049] Embodiments of the present invention include: obtaining test data, inputting the test data into a preset first neural network model to obtain a first output result; obtaining a first adjustable variable from the first output result; obtaining first target input data corresponding to the first adjustable variable; and adjusting at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model. According to the solution provided by the embodiments of the present invention, by adjusting the network layers of the neural network model through the first adjustable variable, the first target input data, and the network layer adjustment rule, a new neural network model with effectively reduced memory consumption can be obtained.
[0050] Other features and advantages of the present invention will be described in the subsequent description, and part of them will become obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the description, the claims, and the drawings. Description of the Drawings
[0051] The drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the description. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation to the technical solutions of the present invention.
[0052] Figure 1 is a flowchart of the steps of the adjustment method for the neural network model provided by an embodiment of the present invention;
[0053] Figure 2 is a flowchart of the steps of obtaining the first target input data provided by another embodiment of the present invention;
[0054] Figure 3It is a flowchart of the steps of the method for adjusting a neural network model provided by another embodiment of the present invention;
[0055] Figure 4 It is a flowchart of the steps of the method for adjusting a neural network model provided by another embodiment of the present invention;
[0056] Figure 5 It is a flowchart of the steps for determining a boundary network layer provided by another embodiment of the present invention;
[0057] Figure 6 It is a flowchart of the steps for running a second neural network model provided by another embodiment of the present invention;
[0058] Figure 7 It is a flowchart of the steps for performing data cleaning on data to be processed provided by another embodiment of the present invention;
[0059] Figure 8 It is a schematic diagram of network layer reordering provided by another embodiment of the present invention;
[0060] Figure 9 It is a schematic diagram of the modules of the device for adjusting a neural network model provided by another embodiment of the present invention;
[0061] Figure 10 It is a structural diagram of the device for adjusting a neural network model provided by another embodiment of the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first", "second", etc. in the description, claims or the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0064] The present invention provides a method, an apparatus, and a computer-readable storage medium for adjusting a neural network model. The adjustment method includes: obtaining test data, inputting the test data into a preset first neural network model to obtain a first output result; obtaining a first adjustable variable from the first output result; obtaining first target input data corresponding to the first adjustable variable; and adjusting at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model. According to the technical solution of the present application, by adjusting the network layers of the neural network model through the first adjustable variable, the first target input data, and the network layer adjustment rule, a new neural network model with reduced memory consumption can be effectively obtained.
[0065] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0066] As Figure 1 shown, Figure 1 FIG. is a flowchart of steps of a method for adjusting a neural network model according to an embodiment of the present invention. The service metric reporting method includes but is not limited to the following steps:
[0067] Step S110: Obtain test data, input the test data into a preset first neural network model to obtain a first output result;
[0068] Step S120: Obtain a first adjustable variable from the first output result;
[0069] Step S130: Obtain first target input data corresponding to the first adjustable variable;
[0070] Step S140: Adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model.
[0071] It should be noted that the embodiments of the present application do not limit the specific model structure of the first neural network model, which may be a MobileNet lightweight network model, an AlexNet network model, a VGG16 network model, etc.
[0072] It should be noted that the embodiments of the present application do not limit the data formats of the first adjustable variable and the first target input data. The data formats of the first adjustable variable and the first target input data may be tensors, and at the same time, the specific tensor forms are not limited, which may be 0-dimensional tensors, i.e., scalars, 1-dimensional tensors, i.e., vectors, two-dimensional tensors, i.e., matrices, or high-dimensional tensors, which will not be elaborated here.
[0073] It can be understood that in the embodiments of the present application, by obtaining the first adjustable variable from the first output result of the first neural network model, the first target input data corresponding to the first adjustable variable, and the preset network layer adjustment rule to adjust and optimize at least two first network layers in the first neural network model, a second neural network model with a reduced memory occupancy rate is obtained, so that the neural network model can be better deployed on an electronic device. The embodiments of the present application do not limit the specific structure of the electronic device running the second neural network model, which may be a mobile phone terminal or an MCU.
[0074] It should be noted that the embodiments of the present application do not limit the specific operations for adjusting at least two first network layers of the first neural network model. Adjusting the first network layer based on the preset network layer adjustment rule may be to fuse multiple first network layers, or delete a certain first network layer, or reorder the first network layers.
[0075] In addition, referring to Figure 2 , in an embodiment, Figure 1 the steps in the embodiment shown include but are not limited to the following steps:
[0076] Step S210, obtaining a preset relationship mapping table, where the relationship mapping table represents the mapping relationship between the first output result and the network output layer, and the network output layer is the network layer that generates the output result;
[0077] Step S220, determining the network output layer corresponding to the first adjustable variable in the relationship mapping table as the target network output layer;
[0078] Step S230, obtaining the first target input data according to the target network output layer, where the first target input data is the data input to the target network output layer.
[0079] It can be understood that by obtaining the relationship mapping table representing the mapping relationship between the first output result and the network output layer, determining the target network output layer corresponding to the first adjustable variable from the relationship mapping table, and obtaining the first target input data input to the target network output layer, an effective data basis can be provided for adjusting the network layer in the first neural network model.
[0080] It should be noted that the first target input data in the embodiments of the present application may be the output data of the network layer located in the upper layer of the target network output layer in terms of sorting, or the test data initially input to the first neural network model, and no further limitation is made here.
[0081] In addition, referring to Figure 3 , in an embodiment, the first adjustable variable includes at least two first variables; Figure 1 the steps in the embodiment shown include but are not limited to the following steps:
[0082] Step S310: Determine a first target variable from at least two first variables, and obtain first input data corresponding to the first target variable from the first target input data;
[0083] Step S320: Determine the network output layer corresponding to the first target variable in the relationship mapping table as the first network output layer, and obtain a first network layer identifier, where the first network layer identifier is a parameter that uniquely identifies the first network output layer;
[0084] Step S330: Determine the sum of the memory occupancy value of the first adjustable variable and the memory occupancy value of the first input data as the first memory occupancy value;
[0085] Step S340: Delete the first target variable in the first adjustable variable, and perform data merging processing on the first adjustable variable after deleting the first target variable and the first input data to obtain a first intermediate data set, where the first intermediate data set includes a second adjustable variable, and the second adjustable variable includes at least two second variables;
[0086] Step S350: Determine a second target variable from at least two second variables, and obtain second target input data corresponding to the second target variable;
[0087] Step S360: Determine the sum of the memory occupancy value of the second adjustable variable and the memory occupancy value of the second input data as the second memory occupancy value;
[0088] Step S370: Adjust at least two first network layers in the neural network model according to the first memory occupancy value, the second memory occupancy value, and the first network layer identifier to obtain a second neural network.
[0089] In addition, referring to Figure 4 , in an embodiment, after performing step S370 in the embodiment shown in Figure 3 , the method for adjusting the neural network model further includes but is not limited to the following steps:
[0090] Step S410: When the number of remaining second variables in the second adjustable variable is greater than a preset first threshold value, determine the second adjustable variable as the new first adjustable variable;
[0091] Step S420: Obtain new first target input data corresponding to the new first adjustable variable;
[0092] Step S430: Adjust at least two second network layers in the second neural network model according to the new first adjustable variable, the new first target input data, and the network layer adjustment rule to obtain a new second neural network model.
[0093] It can be understood that, in order to more clearly describe the steps of adjusting the first network layer in the first neural network model and the second network layer in the second neural network model, the following uses a specific example for illustration: First, obtain the first output result. The first output result S1 includes a constant data set CS and a first adjustable variable AS1. Select a first target variable x1 from AS1, obtain the network output layer I1 corresponding to x1, determine the first network layer identifier of l1, and obtain the input data set IS1 of I1. Calculate the sum of the memory occupancies of IS1 and AS1 to obtain the first memory occupancy value P; delete x1 from AS, merge the AS1 after deleting x1 with IS1 to obtain the first intermediate data set S2. AS2 is included in S2. Obtain a second target variable x2 from AS2, obtain the network output layer I2 corresponding to x2, and obtain the input data set IS2 of I2. Calculate the sum of the memory occupancies of IS2 and AS2 to obtain the second memory occupancy value Q. Compare and process P and Q. When P is less than Q, then determine that the network layer l1 is the target network layer. When sorting and scheduling the first neural network model, l1 will be selected through the first network layer identifier of l1, and l1 is confirmed as the first network layer of the second neural network model; then continue to delete x2 from AS2, repeat the above data merging, calculate the new first memory occupancy value and the new second memory occupancy value. When the new first memory occupancy value is less than the new second memory occupancy value, obtain the second network layer identifier of the network output layer l2 corresponding to x2, determine the network output layer l2 corresponding to x2 as the new target network layer, and confirm l2 as the second network layer of the second neural network model, until the adjustable variables are all deleted, that is, the number of the second variable x22 in AS2 is equal to the first threshold value, and the first threshold value is 0. At this time, output the memory occupancy value of CS, and the sorting of the first network layer of the first neural network model is completed, and the second neural network model is obtained.
[0094] The network layer sorting in the embodiments of this application can refer to Figure 8, assume that the memory occupancy value of network layer 810 is 50, the memory occupancy value of network layer 820 is 40, the memory occupancy value of network layer 830 is 30, the memory occupancy value of network layer 840 is 35, the memory occupancy value of network layer 850 is 40, the memory occupancy value of network layer 860 is 50, the memory occupancy value of network layer 870 is 25, and the memory occupancy value of network layer 880 is 20; the output activation value of network layer 810 will be sent to two branches and finally converge at network layer 880. The previous network layer cannot be released until all the subsequent network layers are scheduled. When using the scheduling order of 810, 820, 830, 840, 850, 860, 870, 880, the peak memory occupancy will occur at network layer 860 and be 125 (40 of network layer 850 + 50 of network layer 860 + 35 of network layer 840); if the order of 810, 850, 860, 870, 820, 830, 840, 880 is used, the peak memory occupancy will also occur at network layer 860, and the memory occupancy value will be 140 (40 of network layer 850 + 50 of network layer 860 + 50 of network layer 810). It can be seen that by reasonably arranging the operator reordering, we can effectively reduce the peak memory occupancy.
[0095] In addition, referring to Figure 5 , in an embodiment, the second neural network model includes at least two second network layers connected in sequence. After performing Figure 1 the steps in the embodiment shown in
[0096] Step S510, obtain the third memory occupancy value corresponding to each second network layer;
[0097] Step S520, determine the boundary network layer from at least two second network layers, and the ratio of the sum of the third memory occupancy values corresponding to each second network layer located before the boundary network layer to the sum of the third memory occupancy values corresponding to each second network layer located after the boundary network layer is greater than a preset second threshold value.
[0098] It can be understood that the network layers in a convolutional neural network can generally be divided into computation-intensive and I / O-intensive layers. The computation-intensive layers mainly include convolutional layers, and the I / O-intensive layers mainly include fully-connected layers and depthwise separable convolutional layers. Other network layers such as pooling layers and non-linear layers have relatively low overhead, and their memory occupancy values do not need to be considered. The execution time of a convolutional neural network is jointly determined by the computation time of the computation-intensive layers and the I / O time of the I / O-intensive layers. The computation-intensive layers are mainly concentrated in the first few layers of the entire network, and it is these layers that account for most of the memory occupancy of the entire network. Therefore, determining the boundary network layer between the computation-intensive network layers with large memory occupancy values and the network layers with small memory occupancy values can provide an effective data basis for determining the network layers with relatively large memory occupancy values.
[0099] It should be noted that the embodiments of the present application do not limit the specific value of the second threshold, and those skilled in the art can determine it according to the actual situation.
[0100] In addition, referring to Figure 6 , in one embodiment, after executing Figure 5 the steps in the embodiment shown in
[0101] Step S610: Obtain the data to be processed, input the data to be processed into the second neural network model, and obtain a second output result. The second output result includes at least two alternative output data.
[0102] Step S620: Obtain the alternative input data corresponding to each alternative output data.
[0103] Step S630: Obtain the target output data from at least two alternative output data. The network output layer corresponding to the target output data is the second network layer whose ranking is before the boundary network layer.
[0104] Step S640: Obtain the third target input data, where the third target input data is the input data corresponding to the target output data.
[0105] Step S650: Perform data merging processing on the target output data and the third target input data to obtain a second intermediate data set.
[0106] Step S660: Classify the data in the second intermediate data set according to a preset data classification rule to obtain a first intermediate data and a second intermediate data. The task target of the first intermediate data is an I / O task, and the task target of the second intermediate data is a data operation task.
[0107] Step S670: Divide the storage device into a first storage area and a second storage area, where the storage device is used to run the second neural network model.
[0108] Step S680: Send the first intermediate data to the first storage area to perform input / output (I / O) operations.
[0109] Step S690: Send the second intermediate data to the second storage area to perform data operation operations.
[0110] It can be understood that when the MCU runs the second neural network model, the memory of the MCU is divided into a first storage area and a second storage area. The target output data and the third target input data of the second network layer located in front of the boundary network layer are obtained, and the target output data and the third target input data are subjected to data merging processing to obtain a second intermediate data set. The data in the second intermediate data set is classified according to a preset data classification rule to obtain the first intermediate data and the second intermediate data. The first intermediate data is sent to the first storage area to perform input / output (I / O) operations (i.e., data exchange between the memory and the flash memory, including tasks of reading inputs and weights from the external flash memory into the memory, and writing outputs from the memory to the external flash memory); the second intermediate data is sent to the second storage area to perform data operation operations, so as to perform data operation tasks and I / O tasks in parallel. The solution provided in this embodiment may be that the MCU maintains two queues, one is an I / O queue and the other is a data operation (compute) queue; the calculation process of a network layer located before the boundary network layer position is divided into several read tasks, several write tasks, and several data operation tasks. There are dependencies between each task. A compute task depends on the read task to read in the relevant input data and weights first. The write task depends on the output data of the compute task. The task of reading weights depends on the task of reading inputs. When all the dependencies of a task are satisfied, it enters the ready state. The read task and the write task enter the I / O queue, and the compute task enters the compute queue. Continuously check these two queues in parallel. When it is detected that a task is in the ready state, execute the task, and kick it out of the queue after execution. In the embodiment of the present application, on the basis of obtaining the second neural network model with a reduced memory occupancy value, the second neural network model performs data operation tasks and I / O tasks in parallel on the MCU, which can further reduce the I / O latency.
[0111] It should be noted that the specific model of the MCU is not limited in the embodiments of the present application. It can be a small-scale MCU of STM32. Through the technical solution of the present application, by optimizing the scheduling order of network layers in the neural network model and performing data exchange between the memory and the external flash for the first few computationally intensive network layers sorted in the boundary network layer of the neural network, the peak memory can be reduced without introducing excessive latency overhead.
[0112] In addition, referring to Figure 7 , in one embodiment, before step S610 in the embodiment shown in Figure 6 , the method for adjusting the neural network model further includes but is not limited to the following steps:
[0113] Step S710, obtaining a preset data cleaning rule;
[0114] Step S720, cleaning the data to be processed according to the data cleaning rule to obtain new data to be processed.
[0115] It can be understood that cleaning the data to be processed before inputting it into the second neural network model can make the output result of the second neural network model more usable, and the new data to be processed after data cleaning can provide an effective data basis for adjusting the second neural network model.
[0116] In addition, referring to Figure 9 , Figure 9 is a schematic diagram of the modules of an apparatus for adjusting a neural network model provided by another embodiment of the present invention. An embodiment of the present invention also provides an apparatus 900 for adjusting a neural network model. The apparatus 900 for adjusting a neural network model includes:
[0117] A first output result obtaining module 910, which is used to obtain test data, input the test data into a preset first neural network model, and obtain a first output result;
[0118] A first adjustable variable obtaining module 920, which is used to obtain a first adjustable variable from the first output result;
[0119] A first target input data obtaining module 930, which is used to obtain first target input data corresponding to the first adjustable variable;
[0120] A neural network model adjusting module 940, which is used to adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model.
[0121] In addition, furthermore, reference Figure 10 , Figure 10 FIG. Figure 10 is a structural diagram of an adjustment device for a neural network model provided by another embodiment of the present invention. An embodiment of the present invention also provides an adjustment device 1000 for a neural network model. The adjustment device 1000 for the neural network model includes: a memory 1010, a processor 1020, and a computer program stored on the memory 1010 and executable on the processor 1020.
[0122] The processor 1020 and the memory 1010 can be connected through a bus or other means.
[0123] The non-transitory software program and instructions required to implement the adjustment method of the neural network model in the above embodiments are stored in the memory 1010. When executed by the processor 1020, the adjustment method of the neural network model in the above embodiments is executed. For example, the method steps S110 to S140 described above are executed, Figure 1 the method steps S210 to S230 in Figure 2 the method steps S310 to S370 in Figure 3 the method steps S410 to S430 in Figure 4 the method steps S510 to S520 in Figure 5 the method steps S610 to S690 in Figure 6 the method steps S710 to S720 in Figure 7 are executed.
[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0125] In addition, an embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a processor or a controller, for example, executed by a processor 920 in the embodiment of the above adjustment device 1000 for the neural network model, the above processor can execute the adjustment method of the neural network model in the above embodiments. For example, the method steps S110 to S140 described above are executed, Figure 1 the method steps S210 to S230 in Figure 2 the method steps S310 to S370 in Figure 3 the method steps S410 to S430 in Figure 4Method steps S410 to S430 in Figure 5 Method steps S510 to S520 in Figure 6 Method steps S610 to S690 in Figure 7 Method steps S710 to S720 in. Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0126] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more programs for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based device for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0128] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0129] It should be noted that although several modules or units of devices for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0130] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0131] The terminal in this embodiment may include components such as a Radio Frequency (RF) circuit, a memory, an input unit, a display unit, sensors, an audio circuit, a wireless fidelity (WiFi) module, a processor, and a power supply. The RF circuit can be used for receiving and transmitting signals during information reception or calls. Specifically, after receiving the downlink information from the base station, it is given to the processor for processing; in addition, the uplink data designed is sent to the base station. Generally, the RF circuit includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc. The memory can be used to store software programs and modules. The processor executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the terminal (such as audio data, a phone book, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices. The input unit can be used to receive input digital or character information and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit can include a touch panel and other input devices. The touch panel, also known as a touch screen, can collect touch operations on or near it (such as operations using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller.Among them, the touch detection device detects the touch orientation and the signals brought by the touch operation, and transmits the signals to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor, and can also receive and execute the commands sent by the processor. In addition, multiple categories such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel. In addition to the touch panel, the input unit may further include other input devices. Specifically, the other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), trackballs, mice, joysticks, etc. The display unit can be used to display the input information or the provided information as well as various menus of the terminal. The display unit may include a display panel. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD for short), an organic light-emitting diode (OLED for short), etc. Further, the touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor to determine the category of the touch event. Subsequently, the processor provides corresponding visual output on the display panel according to the category of the touch event. The touch panel and the display panel are implemented as two independent components to realize the input and output functions of the terminal. However, in some embodiments, the touch panel and the display panel can be integrated to realize the input and output functions of the terminal. The terminal may further include at least one sensor, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or the backlight when the terminal is moved close to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of the acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the terminal (such as horizontal / vertical screen switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc.; as for other sensors that can also be configured in the terminal, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here. The audio circuit, speaker, and microphone can provide an audio interface. The audio circuit can convert the received audio data into an electrical signal and transmit it to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit and then converted into audio data. After the audio data is output and processed by the processor, it is sent through the RF circuit to, for example, another terminal, or the audio data is output to the memory for further processing.
[0132] Other embodiments of the present application will be readily contemplated by those skilled in the art after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0133] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for adjusting a neural network model, characterized in that Including: Obtain test data, input the test data into a preset first neural network model, and obtain a first output result; Obtain a first adjustable variable from the first output result; Obtain first target input data corresponding to the first adjustable variable; Adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model; The second neural network model includes at least two sequentially connected second network layers. After adjusting at least two first network layers in the first neural network model according to the first adjustable variable, the target input data, and a preset network layer adjustment rule to obtain a second neural network model, it further includes: Obtain a third memory occupancy value corresponding to each of the second network layers; Determine a boundary network layer from at least two of the second network layers, and the ratio of the sum of the third memory occupancy values corresponding to each of the second network layers sorted before the boundary network layer to the sum of the third memory occupancy values corresponding to each of the second network layers sorted after the boundary network layer is greater than a preset second threshold; After determining the boundary network layer from at least two of the second network layers, the method further includes: Obtain data to be processed, input the data to be processed into the second neural network model, and obtain a second output result, where the second output result includes at least two optional output data; Obtain optional input data corresponding to each of the optional output data; Obtain a target output data from at least two of the optional output data, and the network output layer corresponding to the target output data is the second network layer sorted before the boundary network layer; Obtain third target input data, where the third target input data is the input data corresponding to the target output data; Perform data merging processing on the target output data and the third target input data to obtain a second intermediate data set; Classify the data in the second intermediate data set according to a preset data classification rule to obtain first intermediate data and second intermediate data, where the task objective of the first intermediate data is an IO task, and the task objective of the second intermediate data is a computing task; Divide a storage device into a first storage area and a second storage area, where the storage device is used to run the second neural network model; Send the first intermediate data to the first storage area to perform input / output (IO) operations; Send the second intermediate data to the second storage area to perform data operation operations.
2. The method according to claim 1, characterized in that, The obtaining the first target input data corresponding to the first adjustable variable includes: Obtain a preset relationship mapping table, where the relationship mapping table represents the mapping relationship between the first output result and the network output layer, and the network output layer is the network layer that generates the output result; Determine the network output layer corresponding to the first adjustable variable in the relationship mapping table as the target network output layer; The first target input data is obtained according to the target network output layer, and the first target input data is the data input to the target network output layer.
3. The method according to claim 2, characterized in that, The first adjustable variable includes at least two first variables; adjusting at least two first network layers in the neural network model according to the first adjustable variable, the target input data, and a preset network layer adjustment rule to obtain a second neural network model, including: Determine a first target variable from at least two of the first variables, and obtain first input data corresponding to the first target variable from the first target input data; Determine the network output layer corresponding to the first target variable in the relationship mapping table as the first network output layer, and obtain a first network layer identifier, where the first network layer identifier is a parameter that uniquely identifies the first network output layer; Determine the sum of the memory occupancy value of the first adjustable variable and the memory occupancy value of the first input data as the first memory occupancy value; Delete the first target variable in the first adjustable variable, and perform data merging processing on the first adjustable variable after deleting the first target variable and the first input data to obtain a first intermediate data set, where the first intermediate data set includes a second adjustable variable, and the second adjustable variable includes at least two second variables; Determine a second target variable from at least two of the second variables, and obtain second target input data corresponding to the second target variable; Determine the sum of the memory occupancy value of the second adjustable variable and the memory occupancy value of the second target input data as the second memory occupancy value; Adjust at least two of the first network layers in the neural network model according to the first memory occupancy value, the second memory occupancy value, and the first network layer identifier to obtain the second neural network.
4. The method according to claim 3, characterized in that, After adjusting at least two of the first network layers in the neural network model according to the first memory occupancy value, the second memory occupancy value, and the first network layer identifier to obtain the second neural network, the method further includes: When the number of the remaining second variables in the second adjustable variable is greater than a preset first threshold, determine the second adjustable variable as a new first adjustable variable; Obtain new first target input data corresponding to the new first adjustable variable; Adjust at least two second network layers in the second neural network model according to the new first adjustable variable, the new first target input data, and the network layer adjustment rule to obtain a new second neural network model.
5. The method according to claim 1, wherein Before inputting the data to be processed into the second neural network model, the method further includes: Obtain a preset data cleaning rule; Perform data cleaning on the data to be processed according to the data cleaning rule to obtain new data to be processed.
6. An adjustment device for a neural network model, characterized in that, Including: A first output result acquisition module, which is used to obtain test data, input the test data into a preset first neural network model, and obtain a first output result; The first adjustable variable acquisition module, which is used to acquire the first adjustable variable from the first output result; The first target input data acquisition module, which is used to acquire the first target input data corresponding to the first adjustable variable; The neural network model adjustment module, which is used to adjust at least two first network layers in the first neural network model according to the first adjustable variable, the first target input data, and a preset network layer adjustment rule to obtain a second neural network model; The second neural network model includes at least two sequentially connected second network layers. After adjusting at least two first network layers in the first neural network model according to the first adjustable variable, the target input data, and a preset network layer adjustment rule to obtain a second neural network model, it further includes: Obtain the third memory occupancy value corresponding to each of the second network layers; Determine the boundary network layer from at least two of the second network layers, and the ratio of the sum of the third memory occupancy values corresponding to each of the second network layers sorted before the boundary network layer to the sum of the third memory occupancy values corresponding to each of the second network layers sorted after the boundary network layer is greater than a preset second threshold; After determining the boundary network layer from at least two of the second network layers, it further includes: Obtain the data to be processed, input the data to be processed into the second neural network model to obtain a second output result, and the second output result includes at least two optional output data; Obtain the optional input data corresponding to each of the optional output data; Obtain the target output data from at least two of the optional output data, and the network output layer corresponding to the target output data is the second network layer sorted before the boundary network layer; Obtain the third target input data, which is the input data corresponding to the target output data; Perform data merging processing on the target output data and the third target input data to obtain a second intermediate data set; Classify the data in the second intermediate data set according to a preset data classification rule to obtain a first intermediate data and a second intermediate data, the task target of the first intermediate data is an IO task, and the task target of the second intermediate data is a computing task; Divide the storage device into a first storage area and a second storage area, where the storage device is used to run the second neural network model; Send the first intermediate data to the first storage area to perform input / output IO operations; Send the second intermediate data to the second storage area to perform data operation operations.
7. An adjustment device for a neural network model, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, it implements the neural network model adjustment method according to any one of claims 1 to 5.
8. A computer-readable storage medium stores computer-executable instructions for performing the method for adjusting a neural network model according to any one of claims 1 to 5.
Citation Information
Patent Citations
Network convergence method, network convergence device, electronic equipment and storage medium
CN111340216A
Optimization method and device for deep learning network
CN112116081A