A model generation method and an electronic device

By building an initial model and combining virtual paths and loss function training, a lightweight model is generated, which solves the problems of high memory usage and long inference time caused by complex models in terminal devices, and achieves the simplification and performance improvement of the model.

CN119360046BActive Publication Date: 2025-07-25HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411930038.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-07-25
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Due to the excessive complexity of existing models on terminal devices, high memory usage, high computing volume and extended forward inference time are severely hindered their application in real scenarios, especially in low and medium-performance terminals.

Method used

Build the initial model and use pre-processing and post-processing modules connected by direct feed-through, train the model in combination with virtual paths and loss functions, simplify the model structure and generate a lightweight model.

Benefits of technology

Generate a lightweight model with a simple structure, improve the model training effect, reduce forward inference time and power consumption, and expand model application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360046B_ABST
    Figure CN119360046B_ABST
Patent Text Reader

Abstract

The present application provides a model generation method and an electronic device. First, an initial model is determined. In the preprocessing module and the postprocessing module of the initial module, each sub-module adopts a direct feed connection. For example, the preprocessing module includes N convolutional blocks (N≥2) connected in direct feed in sequence. Then, a virtual path is constructed for the initial model. The virtual path includes connecting the outputs of some or all sub-modules in the preprocessing module after accumulation to a virtual postprocessing module. Some sub-modules in the preprocessing module at least include non-last sub-modules in the preprocessing module. The virtual postprocessing module is the same as the postprocessing module. Finally, the initial model is trained using the weighted sum of the model loss and the virtual path loss, and the weights in the model are gradually updated until the weighted sum of the new model loss and the virtual path loss meets a preset condition, and then the new model is applied in the NPU. In this way, a lightweight model with a simple structure is generated and the accuracy of the forward inference of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminals, and in particular, to a model generation method and an electronic device. Background Art

[0002] The magnitude of a model is crucial for applications on terminals. When the model magnitude is too large, problems such as high memory occupancy, high computational complexity, and extended forward inference time will occur. Especially for low- and medium-performance terminals, complex models will seriously hinder their applications in real scenarios.

[0003] How to design a lightweight model is an urgent problem to be solved. Summary of the Invention

[0004] This application provides a model generation method and an electronic device. The method includes first constructing a simple straight-through initial model, and then constructing a virtual path for the model, and training the initial model by combining the loss of the model and the loss of the virtual path to obtain a finally updated simple straight-through initial model.

[0005] In a first aspect, this application provides, for example, a model generation method, which includes: generating an initial model, the initial model including a preprocessing module and a postprocessing module, and both the preprocessing module and the postprocessing module including a plurality of sub-modules connected in sequence in a straight-through manner; generating a virtual path for the initial model, the virtual path including: an accumulation module and a virtual postprocessing module; the input of the accumulation module is a first input, the first input being the same as the output of one or more sub-modules in the preprocessing module, the output of the accumulation module being the input of the virtual postprocessing module, and the virtual postprocessing module having the same structure as the postprocessing module; updating the initial model and updating the virtual path based on the loss of the initial model and the loss of the virtual path until the total loss of the updated model and the virtual path is less than a first threshold.

[0006] Implementing the method provided in the first aspect can generate a lightweight model and improve the model training effect, ensuring the output performance of the lightweight model.

[0007] Combined with the method described in the first aspect, the accumulation module is connected to one or more sub-modules in the preprocessing module, and the first input is the output of one or more sub-modules in the preprocessing module; or, the virtual path further includes a virtual preprocessing module, the virtual preprocessing module having the same structure as part or all of the preprocessing module, the accumulation module being connected to one or more sub-modules in the virtual preprocessing module, and the first input being the output of one or more sub-modules in the virtual preprocessing module.

[0008] In this way, the virtual path can be constructed in multiple ways, including reusing some or all of the sub-modules in the model, or independently constructing virtual paths identical to some or all of the sub-modules in the model outside the sub-modules in the model, providing multiple implementable methods.

[0009] Combined with the method described in the first aspect, update the initial model based on the loss of the initial model and the loss of the virtual path, specifically including: updating the weights of the sub-modules in the initial model based on the loss of the initial model and the loss of the virtual path, specifically including updating the weights of each sub-module in the pre-processing module and the post-processing module.

[0010] In this way, by updating the weights of the sub-modules in the model during the training process, the output performance of each intermediate sub-module in the model can be optimized, thereby improving the overall output performance of the model.

[0011] Combined with the method described in the first aspect, update the initial model based on the loss of the initial model and the loss of the virtual path, and update the virtual path until the total loss of the updated model and the virtual path is less than the first threshold, specifically including: updating the initial model and the virtual path based on the loss of the initial model and the loss of the virtual path until the weighted sum of the loss of the updated model and the loss of the virtual path is less than the first threshold.

[0012] In this way, by updating the model through the weighted summation of the model loss and the virtual path loss, the weights corresponding to the two losses can be set according to actual needs.

[0013] Combined with the method described in the first aspect, the weight of the loss of the updated model is greater than the weight of the loss of the virtual path.

[0014] In this way, the loss of the model can be mainly considered, and the loss of the virtual path can be considered to a lesser extent to update the model, thereby playing an additional supervisory role for the pre-processing module in the model to a certain extent.

[0015] Combined with the method described in the first aspect, the method is applied to the first device, and the method further includes: storing the updated model in the second device.

[0016] In this way, after the model is generated in the first device at the local end, the model can also be applied in more other second devices, expanding the coverage of the model.

[0017] Combined with the method described in the first aspect, the method further includes: inputting a first image into the updated model, and outputting a second image after being processed by the updated model, and the resolution of the second image is higher than the resolution of the first image.

[0018] In this way, the models in the existing image processing field can be optimized. In addition, the model generation method provided in this application can also be adopted in other fields.

[0019] Combined with the method described in the first aspect, the sub-module in the preprocessing module includes at least two convolutional blocks; the previous convolutional block in the two convolutional blocks is used to extract features from the image; the latter convolutional block in the two convolutional blocks is used to extract features from the output of the previous convolutional block.

[0020] In a second aspect, this application provides an electronic device, including: a model, a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the method described in any one of the first aspect.

[0021] In a third aspect, this application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the first aspect is implemented.

[0022] In a fourth aspect, this application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspect is implemented. Description of the Drawings

[0023] Figure 1 It is a schematic structural diagram of a complex model provided by an embodiment of this application;

[0024] Figure 2 It is a schematic structural diagram of a lightweight model provided by an embodiment of this application;

[0025] Figure 3 It is a schematic diagram of the effect of equivalently optimizing the model structure by using a loss function provided by an embodiment of this application;

[0026] Figure 4 It is a schematic diagram of the principle of training a lightweight model provided by an embodiment of this application;

[0027] Figure 5 It is a schematic diagram of the hardware architecture of the electronic device provided by an embodiment of this application;

[0028] Figure 6 It is a schematic diagram of the software architecture of the electronic device provided by an embodiment of this application. Detailed Embodiments

[0029] The technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings. In the present application, referring to "embodiment" means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described in the present application may be combined with other embodiments.

[0030] The complexity of the model greatly affects its application effect on the terminal. A complex model will increase the time consumption of its forward inference on the terminal and increase the power consumption of its forward inference on the terminal, resulting in fewer types of terminals that the model can cover, seriously hindering the application of the model in more scenarios. Among them, the forward inference of the terminal refers to the process in which the terminal uses the model to process data, including the entire process from receiving input data to generating output data. The time consumption and power consumption of this process directly affect the terminal response speed and user experience. Therefore, the complexity of the model is crucial for real-time applications on the terminal.

[0031] In order to design a lightweight model, in a feasible manner, the model can be simplified by reducing the number of parameters and the amount of computation in the model, improving its forward inference performance on the terminal, so that it can be applied to more types of terminals and cover more application scenarios. However, the optimization of the number of parameters and the amount of computation of the model is limited. How to further lightweight the model is an urgent problem to be solved.

[0032] By comparing and analyzing models with similar numbers of parameters and amounts of computation and different model structures, the present application can find that the model structure will also seriously affect the time and power consumption of the forward inference of the model on the mid-end. Therefore, by designing the model structure, the limitation of optimizing the lightweight model can be further broken through. Based on this, in order to further lightweight the model, the present application provides a model generation method and an electronic device. The method includes: First, determine an initial model. Each sub-module in the pre-processing module and the post-processing module of the initial module uses a direct feed connection. For example, the pre-processing module includes N convolutional blocks (N≥2) connected in direct feed in sequence. Then, construct a virtual path for the initial model. The virtual path includes accessing the output of some or all sub-modules in the pre-processing module after accumulation to a virtual post-processing module. Some sub-modules in the pre-processing module at least include non-last sub-modules in the pre-processing module. The virtual post-processing module is the same as the post-processing module. Finally, train the initial model using the weighted sum of the model loss and the virtual path loss, and gradually update the weights in the model until the weighted sum of the new model loss and the virtual path loss meets the preset conditions, and then apply the new model to the NPU. In this way, a lightweight model with a simple structure is generated and the accuracy of forward inference is improved.

[0033] The task type performed by the model in the embodiments of the present application is specifically to output an image after super-resolution processing.

[0034] The pre-processing module and the post-processing module of the model described in the present application do not refer to specific entity sub-modules, but are general terms after dividing the entity sub-modules in the model from a functional perspective. Among them, the pre-processing module is mainly responsible for performing necessary pre-processing on the original input data, such as operations like feature extraction, to ensure that the data is suitable for subsequent processing of the model. Among them, the post-processing module is mainly used to further process the output result of the model to obtain a high-resolution image. In the embodiments of the present application, the image input to the lightweight super-resolution model can also be referred to as the first image, and the image output after the lightweight super-resolution model performs super-resolution processing on the first image can also be referred to as the second image.

[0035] Adopting the model generation method provided by the present application can bring the following beneficial effects:

[0036] (1) Provide a lightweight model with a simple structure. Specifically, in the determined initial model and the final model obtained after training the initial model, only multiple sub-modules using the direct feed connection method are included. Compared with the existing models with complex structures, additional residual connections, merge (contact) operations, etc. are removed.

[0037] (2) Improve the training effect of the lightweight model. Specifically, during the process of training the model, avoid the drop in model performance caused by removing complex connection structures. By constructing a virtual path and jointly training the model by combining the loss of the virtual path and the loss of the model, considering that the loss of the virtual path can supervise multiple sub-modules input to the virtual path in the model, avoid the deviation of the intermediate state of the model, and thus play an additional supervision role in the model training process and improve the model training effect.

[0038] (3) Expand the application scenarios of the model and improve the application effect of the model. Specifically, by simplifying the structure of the model, the time consumption and power consumption of the forward inference of the model at the terminal can be reduced, covering more terminal types, and expanding the application scenarios of the model.

[0039] Next, the structure of the complex model before optimization, the structure of the optimized lightweight model provided by the present application, achieving the effect of equivalently optimizing the model structure by designing a loss function, and the principle of training the lightweight model, etc. will be introduced in detail in sequence.

[0040] Figure 1 Exemplarily show the structure of a complex model.

[0041] Figure 1 Specifically, taking the model for implementing image super-resolution processing (abbreviated as super-resolution model) as an example, a complex model is shown. AsFigure 1 As shown, the pre - processing module of this complex model includes Convolution Block 1, Convolution Block 2, and Convolution Block 3, and the post - processing module of this complex model includes convolution, a truncation activation function, and an up - sampling module (such as PixelShuffle). When designing a model, it is usually necessary to design the basic component modules of the model, including the number of input channels, the number of output channels, and the convolution kernel size of the convolution block, etc., and it is also necessary to determine the number of modules, which will determine the number of model parameters and the computational complexity. After designing the basic component modules of the model, it is also necessary to design the connection structure of these modules. Introducing connection structures other than direct feed - through connections can enhance the representational ability of the model, but at the same time, it will also bring additional computational operations. For example, introducing residual connections, concat operations, etc. This application does not make specific restrictions on the number of component modules in the model, and only focuses on introducing the connection structure between each module.

[0042] Among them, the pre - processing module is located at the input end of the model. The pre - processing module refers to the module that performs pre - processing operations on the input data. Usually, the following pre - processing operations are included: data augmentation, which improves the generalization ability of the model by increasing the quantity and diversity of the data set; feature extraction, which selects useful features for model training from the original features and removes irrelevant or redundant features.

[0043] Among them, the post - processing module is located at the output end of the model. The post - processing module refers to the module that performs optimization operations on the output data of the model so as to better apply it to the actual scenario. Usually, the following optimization operations are included: feature extraction, which extracts useful information from the output of the model for further analysis; data post - processing, which constrains the data within a certain range to reduce the model quantization error.

[0044] In practical applications, the specific implementation methods of the pre - processing and post - processing modules will vary according to the specific tasks of the model and the data characteristics. For example, in the image super - resolution processing task, the pre - processing may include operations such as image scaling and cropping.

[0045] Continue to refer to Figure 1 , the Convolution Block 1, Convolution Block 2, and Convolution Block 3 in the pre - processing module adopt the residual connection method.

[0046] Among them, the residual connection (also known as the skip connection) means that the output of the previous layer or multiple previous layers jumps over one or multiple intermediate layers and is accumulated to the output of the subsequent layer to form a residual path. In deep models, the residual connection method is relatively complex and increases the computational complexity of the model, but the residual connection can be used to solve the problems of gradient vanishing and gradient explosion. Specifically, in Figure 1In the shown model structure, residual connections can be used to directly output the output of convolutional block 1 across the intermediate convolutional block 2 and convolutional block 3 to the output of convolutional block 3 for accumulation, and also to directly output the output of convolutional block 2 across the intermediate convolutional block 3 to the output of convolutional block 3 for accumulation, which helps to solve the problems of gradient vanishing and gradient explosion in deep convolutions. Optionally, there are various derivative schemes of residual connections in model design. Figure 1 Only taking the example that convolutional block 1 spans multiple convolutional blocks and convolutional block 2 spans one convolutional block is shown. In other residual connection methods, there can be different connection methods, such as accessing additional modules in the residual path. Although these connection methods usually do not generate additional model parameters, they will bring additional computational cost and will inevitably increase the forward inference latency when the model is deployed on mobile devices. In addition, Figure 1 The concat operation not shown needs to make structural changes to specific convolutional layers to make this operation effective. It will directly modify the number of input channels of the connected convolutional layer, so it will also introduce additional computational overhead and will inevitably increase the forward inference latency when the model is deployed on mobile devices.

[0047] Among them, each of convolutional blocks 1 - 3 contains one or more convolutional layers and one or more activation function modules, and an activation function is connected after each convolutional layer. Among them, the convolutional layer is the core component of the convolutional block and is responsible for extracting local features from the input data. Specifically, the convolutional layer processes the input data by applying filters (or called convolutional kernels) to extract local features. Each convolutional layer can contain one or more convolutional kernels, and the size (such as 3x3, etc.), stride, and padding method of these convolutional kernels can be set according to specific requirements. The result of the convolutional operation of the convolutional layer is a feature map, and each pixel value of the feature map is obtained by multiplying the weights of the filter with the pixel values at the corresponding positions of the input image and then summing all the product results. Among them, the role of the activation function is to introduce non - linear transformation. Common activation functions include ReLU, Sigmoid, and Tanh, etc., which enable the model to learn more complex feature representations. In short, the convolutional block is a structural unit composed of multiple convolutional layers and activation functions, and through the combination of multi - layer convolution and non - linear transformation, it further extracts the complex features of the input data. Taking the super - resolution model as an example, convolutional blocks 1, 2, and 3 can be used to extract the features of the image.

[0048] Continue to refer to Figure 1 , in the post - processing module, the convolutional layer, truncated activation function, and upsampling module are connected in a feed - through manner in sequence.

[0049] Among them, the feed-through connection means that the output of the previous layer is directly used as the input of the next layer, and the output of each layer is calculated based on the output of the previous layer without the intervention of the intermediate layer. In the deep model, the feed-through connection method is simple and direct, which is the basis for constructing a multi-layer model. However, the feed-through connection may cause the problem of gradient disappearance or gradient explosion, thus affecting the inference effect of the model.

[0050] Among them, the convolution in the post-processing module is used to perform channel transformation on the feature map output by the pre-processing module, and transform it into the number of channels convenient for subsequent upsampling processing. The truncation activation function in the post-processing module is used to directly truncate the output of the model between 0 and 1 during the process of the model quantifying the original 32-bit floating-point weights and responses into 8-bit integer weights and responses, so as to reduce the error during model quantization. The upsampling module in the post-processing module is used to perform upsampling processing on the image to achieve the final high-resolution effect.

[0051] It can be understood that Figure 1 The structure of the complex model shown is only an example. In addition to the feed-through connection and residual connection methods, the complex model before optimization can also include other connection methods such as merge (concat) connection. Figure 1 For the time being, a complex model is shown only by taking the example of including feed-through connection and residual connection. In addition, in the structures of other optional complex models, the number of convolutional blocks can be more or less, and there can be other functional modules at the backend. The embodiments of the present application do not limit the connection structure, the number of modules, etc. in the model.

[0052] Figure 2 An exemplary structure of a lightweight model is shown.

[0053] Figure 2 Also specifically taking the model for realizing image super-resolution processing (abbreviated as super-resolution model) as an example, a lightweight model is shown. As Figure 2 shown, the pre-processing module of this lightweight model includes convolutional block 1, convolutional block 2, and convolutional block 3, and the post-processing module of this lightweight model includes convolution, truncation activation function, and upsampling module. Figure 2 The basic component modules of the lightweight model shown are the same as Figure 1 those shown, and it is obtained by simplifying the structure based on the complex model shown in Figure 1 For the specific structure simplification method, reference can be made to the detailed introduction of Figures 3 - 4 later.

[0054] Continue to refer to Figure 2 , the convolutional block 1, convolutional block 2, and convolutional block 3 in the pre-processing module adopt the feed-through connection method in sequence, and the convolution, truncation activation function, and upsampling module in the post-processing module also adopt the feed-through connection method in sequence. Specifically, inFigure 2 In the shown model structure, a direct feed connection means that the output of the previous layer is directly used as the input of the next layer. The output of each layer is calculated based on the output of the previous layer without the intervention of an intermediate layer. In a deep model, the direct feed connection method is simple and straightforward and is the basis for constructing a multi-layer model. However, the direct feed connection may cause problems such as gradient disappearance or gradient explosion, thus affecting the inference effect of the model.

[0055] Among them, the convolution in the post-processing module is used to perform channel transformation on the feature map output by the pre-processing module, transforming it into the number of channels convenient for subsequent upsampling processing. The truncation activation function in the post-processing module is used to directly truncate the output of the model between 0 and 1 during the process of the model lightweighting the original 32-bit floating-point weights into 8-bit integer weights, thereby reducing the error during model quantization. The upsampling module in the post-processing module is used to perform upsampling processing on the image to achieve the final high-resolution effect. Among them, the core idea of model quantization is to reduce the precision of model parameters from a higher bit width (such as FP32) to a lower bit width (such as int8). The main purpose of doing this is to reduce the size of the model and improve the inference speed. During the quantization process, data type conversion is usually involved, that is, the original floating-point tensor is converted into an integer tensor for processing. Specifically, 8-bit quantization can compress the weights and activation values that originally required 32-bit storage space to only 8 bits for storage, thus significantly reducing the size of the model and memory usage. This quantization method can not only reduce the model size and improve the inference speed, but also adapt to edge computing devices because some microprocessors on terminals usually have low power consumption requirements and the hardware accelerator only supports int8.

[0056] It can be understood that Figure 2 The structure of the lightweight model shown is only an example. In the structures of other optional lightweight models, the number of convolution blocks can be more or less, and there can be other functional modules at the backend. The embodiments of the present application do not limit the connection structure, the number of modules, etc. in the model.

[0057] The foregoing only combines Figure 1 and Figure 2 to introduce the structure of the complex model before optimization and the structure of the optimized lightweight model. Next, the effect of equivalently optimizing the model structure by designing a loss function and the principle of training the lightweight model will be introduced in detail.

[0058] Figure 3 Exemplarily shows the effect of equivalently optimizing the model structure using a loss function.

[0059] Figure 3Specifically, two effects brought about by model optimization are shown. The lightweight model provided in this application can have any one or both of the following effects: Compared with the lightweight model trained using the original loss function, the lightweight model optimized using the new loss function can improve the output performance of the model. Compared with the complex model, the lightweight model optimized using the new loss function has a simpler structure and its output performance is close to that of the complex model.

[0060] Taking Figure 3 the shown Effect 1 as an example, when the output performance of the lightweight model (i.e., the model shown in Figure 2 ) trained by constructing a virtual path and using the new loss function Ls is better than that of the lightweight model trained without constructing a virtual path and using the original loss function L, it indicates that constructing a virtual path and using the new loss function can improve the training effect of the lightweight model. Therefore, construct a virtual path and use the new loss function to train the lightweight model. Among them, the new loss function includes the loss of the lightweight model and the loss of the virtual path, and can be expressed as: Ls = 1 * L(model output 0, target) + w * L(virtual output 0, target). Among them, the original loss function of the lightweight model only includes the loss of the lightweight model and can be expressed as: L(model output 1, target). Regarding the specific implementation method of using the new loss function to train the lightweight model, reference can be made to the following description of Figure 4 , which will not be elaborated here for the time being.

[0061] Taking Figure 3 the shown Effect 2 as an example, when the gap between the output performance of the lightweight model (i.e., the model shown in Figure 2 ) trained by constructing a virtual path and using the new loss function Ls and the output performance of the complex model (i.e., the model shown in Figure 1 ) trained using the original loss function L is less than the threshold, it indicates that the performance of the lightweight model after simplifying the structure is close to that of the complex model before simplifying the structure. Therefore, simplify the structure of the complex model and apply the simplified lightweight model. Among them, the original loss function of the complex model only includes the loss of the complex model and can be expressed as: L(model output 2, target).

[0062] It can be seen that meeting the above two model optimization effects is equivalent to both simplifying the structure of the complex model and training the simplified lightweight model using the new loss function. In this way, the purpose of simplifying the model structure can be achieved on the basis of ensuring the model performance, so as to meet the requirements of model simplification.

[0063] It can be understood that Figure 3The two effects shown can be known by comparing with other models after training the lightweight model with the new loss function, and they are not necessary stages in the model generation method described in this application.

[0064] Figure 4 An example shows the principle of training a lightweight model.

[0065] Based on the previous introduction of the complex model before optimization and the lightweight model before optimization shown in Figures 1 - 2 it can be known that the lightweight model has a simpler structure than the complex model. Therefore, before training the lightweight model, it is necessary to first design a lightweight model with a simple structure as the initial model to be trained. Since a model with a simple structure may have performance problems, based on this, when training on the basis of this initial model, additional supervision of the initial model is also carried out by constructing a virtual path and combining the loss of the virtual path.

[0066] As Figure 4 shown, the architecture for training the lightweight model includes an initial model (also called the main path) and a virtual path.

[0067] Among them, the initial model has the same basic component modules as the lightweight model shown in the previous introduction of Figure 2 and the connection method is also the same. The difference is the weights of each sub-module in the model. After training and updating the weights of the initial model, the final lightweight model shown in Figure 2 can be obtained. Optionally, the initial model can also be other models with a simple structure, as long as it meets the condition of a feed-forward network composed of multiple sub-modules. For example, the pre-processing module in the initial model includes N convolutional blocks (N≥2) connected in sequence by feed-forward, where the last convolutional block is connected to the post-processing module of the initial model. The post-processing module in the initial model includes M sub-modules (M≥1) connected in sequence by feed-forward. The number and functions of the M sub-modules are related to the specific task of the model. In the embodiments of this application, the number of convolutional blocks in the pre-processing module and the number and types of functional modules in the post-processing module are not limited.

[0068] Among them, the virtual path is constructed to avoid introducing other connection structures and only maintain the main path as the structure of the model, and is used to play a role similar to that of other connection structures. The virtual path includes: an accumulation module and a virtual post-processing module. The accumulation module is used to add the outputs of convolution block 1, convolution block 2, and convolution block 3 in the pre-processing module of the initial model and output them to the virtual post-processing module, and this virtual post-processing module is the same as the post-processing module in the initial module. Optionally, the virtual path can also be a path with other structures. For example, the accumulation module in the virtual path is used to add the outputs of all or part of the convolution blocks in the pre-processing module and then connect them to the virtual post-processing module. The embodiments of the present application do not limit the number of outputs of the convolution blocks that need to be processed by accumulation. Optionally, in addition to multiplexing Figure 4 the outputs of multiple sub-modules in the pre-processing module of the initial model shown, the input of the accumulation module in the virtual path can also construct a virtual pre-processing module separately, and this virtual pre-processing module is the same as part or all of the results of the pre-processing module in the model, where the same part specifically depends on what the input of the accumulation module includes. In the embodiments of the present application, the input of the virtual path can also be referred to as the first input.

[0069] Continue to refer to Figure 4 , after establishing the initial model (i.e., the main path) and constructing a virtual path for the initial model, a new loss function needs to be used to train the initial model to update the weights in the model until the new loss reaches the minimum or meets certain conditions, then the training is completed.

[0070] Specifically, when using Figure 4During the process of training the model using the shown training architecture, not only is a loss function used to calculate the loss of the main path, but also a loss function is needed to calculate the loss of the virtual path. The loss of the main path and the loss of the virtual path are combined to jointly train the model. Taking a specific example, when the loss function of the model is L, the output of the model (the output of the main path) is denoted as "model output", the output of the virtual path is "virtual output", and the output in the ideal situation is "target (or target)". Then the loss of the main path can be expressed as L(model output, target), the loss of the virtual path can be expressed as L(virtual output, target), and the loss of the entire training architecture can be expressed as Ls = 1 * L(model output, target) + w * L(virtual output, target). Among them, the target is the label corresponding to the input data in the training data pair, and w is the weight of the loss of the virtual path, which is used to balance the importance of the virtual path in the entire training architecture. When w is equal to 0, the virtual path does not take effect. At this time, it is equivalent to only the original main path in the training architecture; when w is equal to 1, the virtual path and the main path take effect simultaneously and have equivalent effects. After determining w and obtaining Ls, the training phase can be directly started using Ls to obtain the updated weights in the model. Usually, w ≥ 0. In this way, during the training process, in addition to largely considering the hierarchical output of convolutional block 1 passing through convolutional block 2 and then through convolutional block 3 in the main path, the outputs of convolutional block 1, convolutional block 2, and convolutional block 3 are also considered to a certain extent in the virtual path.

[0071] In this way, since the loss used in the training process takes into account the respective outputs of convolutional block 1 - convolutional block 3 in the virtual path, this is similar to the skip - type direct transmission of the outputs of convolutional block 1 - convolutional block 3 in the residual connection manner. Equivalently, it supervises the intermediate states such as convolutional block 1 - convolutional block 3 to avoid the deviation of the intermediate states. After the training is completed, the virtual path is removed, and only the main path is used as the trained model. It can be seen that the training method of the lightweight model provided by this application is equivalent to achieving the purpose of simplifying the model structure by means of the design of the loss function.

[0072] Next, the product form, software - hardware architecture, etc. of the terminal on which the lightweight model described in this application is deployed will be introduced.

[0073] The terminal on which the lightweight model can be deployed is also called an electronic device, such as being deployed on a mobile phone, etc.

[0074] Figure 5 An exemplary illustration shows the hardware architecture of the electronic device provided by this application.

[0075] As Figure 5As shown, the electronic device may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a sensor module 180, a camera 193, a display screen 194, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, etc.

[0076] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0077] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a graphics processing unit (GPU), an image signal processor (ISP), a memory, and a neural-network processing unit (NPU), etc. Among them, due to the relatively low computing power of the NPU, applying a lightweight model to the NPU can greatly accelerate the output efficiency of the model. Therefore, the lightweight model provided in the present application may be specifically applied to the NPU of the terminal.

[0078] In the embodiments of the present application, the electronic device may run a lightweight model as Figure 2 shown through the processor 110 to perform super-resolution processing on the input image. Among them, the lightweight model of the electronic device may be generated and stored by this electronic device, or pre-installed in this electronic device after being generated by other electronic devices. The embodiments of the present application do not specifically limit the electronic device that generates and applies the lightweight model. Whether it is generated by the application-side electronic device or other electronic devices, its generation method is the same as the method introduced for Figures 3 - 4 above, and will not be elaborated here.

[0079] In the embodiments of the present application, the electronic device may display the originally input image, the high-resolution image output after being processed by the lightweight super-resolution model, etc. through the display screen 194. The embodiments of the present application do not limit this.

[0080] In an embodiment of the present application, the electronic device can capture an image through the camera 193. When the resolution of the captured image is low, the electronic device can use a lightweight super-resolution model to process the image to output a high-resolution image. In addition, in addition to performing super-resolution processing on the images captured by the camera, the lightweight super-resolution model can also perform super-resolution processing on the images downloaded from the browser and the images in the application. The embodiments of the present application do not limit this.

[0081] The internal memory 121 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM). The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to implement the data storage function.

[0082] In an embodiment of the present application, the electronic device can store the lightweight model as described above through the aforementioned memory, Figure 2 as shown, and store the data output using the lightweight model, etc.

[0083] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. The touch sensor 180K, also known as the "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby.

[0084] In an embodiment of the present application, the electronic device can receive an operation for capturing an image, or an operation for inputting a low-resolution image into the lightweight super-resolution model, or an operation for storing the high-resolution image output after being processed by the lightweight super-resolution model, etc. through the pressure sensor 180A or the touch sensor 180K. The embodiments of the present application do not limit this and will not elaborate.

[0085] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In an embodiment of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the electronic device.

[0086] Figure 6 Exemplarily shows the software architecture of the electronic device provided by the embodiments of the present application.

[0087] Such as Figure 6As shown, the layered architecture divides software into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0088] The application layer may include a series of application packages. For example, Figure 6 As shown, the application packages may include image processing applications, cameras, galleries, calendars, calls, maps, navigation, Bluetooth, videos, text messages, and other applications. Among them, the image processing application may integrate the aforementioned lightweight super-resolution model to provide users with image super-resolution processing. The name of this image processing application is not limited in this application.

[0089] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, Figure 6 As shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.

[0090] Android runtime includes core libraries and virtual machines. The system libraries may include multiple functional modules, such as a surface manager, a 3D graphics processing library, a 2D graphics engine, and a media library, etc.

[0091] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.

[0092] It should be understood that each step in the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in software form. Combining the method steps disclosed in the embodiments of the present application can be directly embodied as the NPU performing the super-resolution operation of the lightweight model and collaborating with the CPU to complete the entire super-resolution process.

[0093] The present application also provides a chip system, which includes at least one processor for implementing the method executed by the electronic device in any one of the above embodiments. In a possible design, the chip system further includes a memory for storing program instructions and data, and the memory is located inside or outside the processor.

[0094] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method executed by the electronic device in any one of the above embodiments.

[0095] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the method executed by the electronic device in any of the above embodiments.

[0096] The various embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0097] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in the text is merely a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0098] The terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more than two.

[0099] In summary, the above description is only for the embodiments of the technical solution of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made according to the disclosure of the present application shall be included within the protection scope of the present application.

Claims

1. A model generation method, characterized in that, The method includes: Generating an initial model, where the initial model includes a preprocessing module and a postprocessing module, and both the preprocessing module and the postprocessing module include a plurality of sub-modules connected in a direct-feed manner in sequence; Generating a virtual path of the initial model, where the virtual path includes: an accumulation module and a virtual postprocessing module; the accumulation module is connected to the plurality of sub-modules in the preprocessing module, the input of the accumulation module is a first input, the first input is the same as the outputs of the plurality of sub-modules in the preprocessing module, the accumulation module is used to accumulate the outputs of the plurality of sub-modules in the preprocessing module and use the result as the input of the virtual postprocessing module, and the virtual postprocessing module has the same structure as the postprocessing module; Updating the initial model based on the loss of the initial model and the loss of the virtual path until the total loss of the updated model and the virtual path is less than a first threshold; Inputting a first image into the updated model, and outputting a second image after being processed by the updated model, where the resolution of the second image is higher than that of the first image.

2. The method according to claim 1, wherein Updating the initial model based on the loss of the initial model and the loss of the virtual path specifically includes: updating the weights of the sub-modules in the initial model based on the loss of the initial model and the loss of the virtual path.

3. The method according to claim 1, characterized in that, Updating the initial model based on the loss of the initial model and the loss of the virtual path until the total loss of the updated model and the virtual path is less than a first threshold specifically includes: updating the initial model based on the loss of the initial model and the loss of the virtual path until the weighted sum of the loss of the updated model and the loss of the virtual path is less than a first threshold.

4. The method according to claim 3, characterized in that, The weight of the loss of the updated model is greater than the weight of the loss of the virtual path.

5. The method according to claim 1, wherein The method is applied to a first device, and the method further includes: storing the updated model in a second device.

6. The method according to claim 1, wherein The sub-modules in the preprocessing module include at least two convolutional blocks; The previous convolutional block of the two convolutional blocks is used to extract features from an image; The latter convolutional block of the two convolutional blocks is used to extract features from the output of the previous convolutional block.

7. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the method according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-6.

9. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Target detection model based on deep learning and training method thereof

    CN108182456A

  • Object classification method and device

    CN111104954A