Target Student Network Model Training Method and Low-Resolution Image Recognition Method

Through joint training of target student network models and the use of FSR-W modules, the recognition difficulties of the prior art under low resolution and harsh conditions are solved, and efficient image recognition on edge computing devices is achieved.

CN116310717BActive Publication Date: 2025-06-20CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310179846.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-06-20
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle visual understanding tasks under harsh conditions, such as low resolution, low brightness and complex backgrounds, and is difficult to implement on edge computing devices.

Method used

Through the target student network model training method, the initial student network model and the target assist network model are used to jointly train, and combined with the FSR-W module, the initial student network model recognizes low-resolution images.

Benefits of technology

The recognition effect of low-resolution images is enhanced, the image recognition performance on edge computing devices is improved, and the recognition difficulties of the prior art under harsh conditions are overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310717B_ABST
    Figure CN116310717B_ABST
Patent Text Reader

Abstract

The present application provides a method for training a target student network model, including inputting a high-resolution image into an initial student network model and a target assistance network model to obtain corresponding first intermediate features and second intermediate features; inputting the first intermediate features and the second intermediate features into a first loss function to obtain a first loss value; inputting a low-resolution image into the initial student network model, obtaining third intermediate features through an FSR-W module, and obtaining a first prediction loss value through a prediction loss function; training the initial student network model based on the first loss value and the first prediction loss value to obtain a target student network model. The present application uses the target assistance network model to train the initial student network model, adds an FSR-W module to the initial student network model to obtain a target student network model, and this model can enhance the recognition effect on low-resolution images in the low-resolution image recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a method for training a target student network model, a method for recognizing low-resolution images, an apparatus, a device, and a storage medium. Background Art

[0002] In the past few years, image recognition methods have been greatly improved, from traditional palmprint recognition methods using handcrafted features to deep learning methods that can automatically learn feature representations from input data. However, in practical applications, CNN (Convolutional Neural Networks) always fails to handle visual understanding tasks under harsh conditions such as low resolution, low brightness, and complex backgrounds, and it is difficult to be easily equipped on edge computing devices such as smartphones, embedded devices, small drones, etc. This is because these methods are usually designed by complex network architectures and learn on large-scale high-quality training images. However, edge computing devices are limited in computing capacity and memory usage, and they may capture natural low-resolution images.

[0003] Therefore, how to enhance the recognition effect of low-resolution images has become a problem to be solved.

[0004] The above information disclosed in the background art is only used to enhance the understanding of the background of this application, and therefore it may contain information on prior art that is not known to those of ordinary skill in the art. Summary of the Invention

[0005] This application provides a method for training a target student network model, a method for recognizing low-resolution images, an apparatus, a device, and a storage medium to solve the problems existing in the prior art.

[0006] In a first aspect, this application provides a method for training a target student network model, including the following steps:

[0007] S11. Input a high-resolution image into an initial student network model and a target assistance network model. The initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature;

[0008] S12. Input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value. The first loss function is used to make the initial student network model simulate the second intermediate feature of the target assistance network model;

[0009] S13. Input the low-resolution image into the initial student network model, and obtain the third intermediate feature through the FSR-W module in the convolutional layer. The third intermediate feature is used to obtain the first prediction loss value of the initial student network model through the prediction loss function;

[0010] S14. Based on the first loss value and the first prediction loss value, train the initial student network model to obtain the target student network model.

[0011] In some embodiments, the training process of the target assistance network model includes the following steps:

[0012] S101. Input the high-resolution image into the target teacher network model, and obtain the fourth intermediate feature through the convolutional layer;

[0013] S102. Input the high-resolution image into the initial assistance network model, and obtain the fifth intermediate feature through the convolutional layer;

[0014] S103. Calculate the loss between the fourth intermediate feature and the fifth intermediate feature to obtain the feature loss value;

[0015] S104. The initial assistance network model obtains the assistance prediction score according to the feature loss value;

[0016] S105. Based on the teacher prediction score generated by the target teacher network model and the assistance prediction score, calculate and obtain the second prediction loss value;

[0017] S106. Perform regression training on the initial assistance network model according to the second prediction loss value to obtain the target assistance network model.

[0018] In some embodiments, the obtaining of the third intermediate feature through the FSR-W module in the convolutional layer includes the following steps:

[0019] S131. The intermediate feature of the low-resolution image enters the FSR-W module, and the intermediate feature includes the target feature map;

[0020] S132. Perform upsampling processing on the intermediate feature to obtain the upsampled feature map;

[0021] S133. Calculate the residual between the target feature map and the upsampled feature map to obtain the residual feature map;

[0022] S134. Obtain the weight map according to the second intermediate feature of the target assistance network model;

[0023] S135. Add the upsampled feature map, the residual feature map and the weight map to obtain the third intermediate feature.

[0024] In some embodiments, the training process of the target teacher network model includes:

[0025] Input a high-resolution image into the initial teacher network model, perform regression training on the initial teacher network model, and obtain the target teacher network model.

[0026] In a second aspect, the present application provides a low-resolution image recognition method, including the following steps:

[0027] S21. Obtain an image to be recognized, where the image to be recognized is a low-resolution image;

[0028] S22. Input the image to be recognized into the target student network model for recognition processing to obtain an image recognition result;

[0029] Wherein, the target student network model is trained by the target assistance network model, and the target student network model includes an FSR-W module.

[0030] In some embodiments, the FSR-W module is a high-frequency content feature super-resolution module, which is used to reduce the performance difference caused by the low-resolution image.

[0031] In a third aspect, the present application provides a target student network model training device, including:

[0032] An input module, configured to input a high-resolution image into the initial student network model and the target assistance network model. The initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature;

[0033] A first processing module, configured to input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value. The first loss function is used to make the initial student network model simulate the second intermediate feature of the target assistance network model;

[0034] A second processing module, configured to input a low-resolution image into the initial student network model, obtain a third intermediate feature through the FSR-W module in the convolutional layer, and obtain a first prediction loss value of the initial student network model through the prediction loss function;

[0035] A training module, configured to train the initial student network model based on the first loss value and the first prediction loss value to obtain the target student network model.

[0036] In a fourth aspect, the present application provides a low-resolution image recognition device, including:

[0037] An acquisition module, configured to acquire an image to be recognized, where the image to be recognized is a low-resolution image;

[0038] An identification module, configured to input the image to be recognized into a target student network model for identification processing to obtain an image recognition result; wherein, the target student network model is trained by a target assistance network model, and the target student network model includes an FSR-W module.

[0039] In a fifth aspect, the present application provides a terminal device, including:

[0040] A memory, configured to store a computer program;

[0041] A processor, configured to read the computer program in the memory and execute the operations corresponding to the low-resolution image recognition method.

[0042] In a sixth aspect, the present application further provides a computer-readable storage medium. A computer-executable instruction is stored in the computer-readable storage medium, and when the computer-executable instruction is executed by a processor, it is used to implement the low-resolution image recognition method.

[0043] The target student network model training method provided by the present application includes the following steps: Input a high-resolution image into an initial student network model and a target assistance network model. The initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature; Input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value. The first loss function is used to make the initial student network model simulate the second intermediate feature of the target assistance network model; Input a low-resolution image into the initial student network model, and obtain a third intermediate feature through the FSR-W module in the convolutional layer. The third intermediate feature passes through a prediction loss function to obtain a first prediction loss value of the initial student network model; Based on the first loss value and the first prediction loss value, train the initial student network model to obtain a target student network model. The technical solution involved in the present application uses a target assistance network model to train an initial student network model, improves the learning effect of the initial student network model on high-resolution images, and adds an FSR-W module to the convolutional layer of the initial student network model, so that the initial student network model can learn better feature representations for classification. Through the above method, the target student network model obtained can enhance the recognition effect on low-resolution images in the low-resolution image recognition method. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] Figure 1 Schematic diagram of the target student network model training method provided by this application;

[0046] Figure 2 Schematic diagram of the processing of the FSR-W module involved in the target student network model training method provided by this application;

[0047] Figure 3 Schematic diagram of the prediction-based distillation method;

[0048] Figure 4 Schematic diagram of the feature-based distillation method;

[0049] Figure 5 Schematic diagram of the pixel distillation method;

[0050] Figure 6 Schematic diagram of the teacher-assisted-student distillation method;

[0051] Figure 7 Is the pixel neighborhood in a 3×3 true heart rate image.

[0052] Through the above-mentioned drawings, the clear embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a" and "the" used in the embodiments of this application are also intended to include the plural forms unless the context clearly indicates otherwise.

[0055] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality" and "several" mean two or more unless otherwise specifically defined.

[0056] It should be noted that the structures, proportions, sizes, etc. shown in the attached drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the conditions under which this application can be implemented. Therefore, they do not have any substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that this application can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in this application.

[0057] It should also be noted that the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such commodity or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the commodity or system including the said element.

[0058] For the problems that in practical applications, CNNs cannot handle visual understanding tasks under harsh conditions and edge computing devices are limited in computing capacity and memory usage, there are the following two solutions:

[0059] One is to preprocess the low-quality images captured under adverse conditions. This connects low-level image processing methods and high-level visual understanding methods. Many experiments have been completed in this area in the past few years. Their work shows that current preprocessing methods can improve the performance of visual algorithms on some low-quality images, but also reduce their performance on some other images. For example, image super-resolution can enlarge low-resolution images to high-resolution images, which should improve the accuracy of visual classification, but sometimes it also introduces artifacts and affects the judgment of the classifier. Generally speaking, existing preprocessing methods can only slightly improve the classification performance, or even reduce it.

[0060] The other is the knowledge distillation technique. In existing knowledge distillation methods, the design of the student network model uses a smaller network architecture than the teacher network model, usually with fewer network layers or smaller channel dimensions, which can reduce the requirements for floating-point operations (FLOPs) and memory space. Figure 3 Schematic diagram of the prediction-based distillation method, Figure 4 Schematic diagram of the feature-based distillation method, such as Figure 3 、 Figure 4As shown, according to current research, two typical distillation schemes are prediction-based distillation and feature-based distillation. Specifically, the prediction-based distillation method uses the soft prediction labels generated from the teacher network model as knowledge to guide the student network model. In contrast, the feature-based distillation method uses intermediate features from different layers to transfer semantic knowledge. Figure 5 is a schematic diagram of the pixel distillation method, as Figure 5 shown, in recent research, a new distillation framework, namely pixel distillation, has also been proposed. Compared with previous distillation frameworks, pixel distillation not only extracts knowledge from the features and predictions of the teacher network's work, but also extracts knowledge from the network input resolution.

[0061] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0062] Figure 1 is a schematic diagram of the training method for the target student network model provided by the present application, as Figure 1 shown, the present application provides a training method for a target student network model, including the following steps:

[0063] S11. Input a high-resolution image into the initial student network model and the target assistance network model. The size of the high-resolution image is w×h. The initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature;

[0064] In some embodiments, the training process of the target assistance network model includes the following steps:

[0065] S101. Input a high-resolution image into the target teacher network model, and through the convolutional layer, obtain a fourth intermediate feature F t ;

[0066] Specifically, the training process of the target teacher network model includes:

[0067] Input a high-resolution image into the initial teacher network model, perform regression training on the initial teacher network model to obtain the target teacher network model, that is, the trained teacher network model.

[0068] S102. Input a high-resolution image into the initial assistance network model, and through the convolutional layer, obtain a fifth intermediate feature F a ;

[0069] S103. Based on the feature loss function L FKD, calculate the fourth intermediate feature F t and the fifth intermediate feature F a to obtain a feature loss value;

[0070] S104. Through the initial assistance network model and according to the feature loss value, obtain an assistance prediction score x a ;

[0071] S105. Based on the teacher prediction score x t generated by the target teacher network model and the assistance prediction score, calculate a second prediction loss value;

[0072] Specifically, in the embodiment of the present application, the second prediction loss value is calculated through a prediction-based loss function L PKD .

[0073] It should be noted that after passing through the initial assistance network model and the target teacher network model and before obtaining the prediction score, a prediction classification label will be generated.

[0074] S106. Perform regression training on the initial assistance network model according to the second prediction loss value to obtain a target assistance network model.

[0075] S12. Input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value, where the first loss function is used to enable the initial student network model to simulate the second intermediate feature of the target assistance network model;

[0076] S13. Input the low-resolution image into the initial student network model, and obtain a third intermediate feature through the FSR-W module in the convolutional layer. The third intermediate feature passes through a prediction loss function to obtain a first prediction loss value of the initial student network model;

[0077] Figure 2 is a schematic diagram of the processing of the FSR-W module involved in the target student network model training method provided by the present application. As Figure 2 shown, in some embodiments, obtaining the third intermediate feature through the FSR-W module in the convolutional layer includes the following steps:

[0078] S131. The intermediate feature of the low-resolution image enters the FSR-W module, and the intermediate feature includes a target feature map;

[0079] S132. Perform upsampling processing on the intermediate feature to obtain an upsampled feature map;

[0080] S133. Calculate the residual between the target feature map and the upsampled feature map to obtain a residual feature map;

[0081] S134. Obtain a weight map based on the second intermediate feature of the target assistance network model;

[0082] S135. Add the upsampled feature map, the residual feature map, and the weight map to obtain a third intermediate feature.

[0083] It should be noted that the low-resolution image enters the convolutional layer of the student network to generate an intermediate feature F s , F s enters the FSR-W module and is divided into two branches. In one branch, upsampling is performed, and in the other branch, the residual between the target feature map and the upsampled feature map is learned. Among them, the weight map is generated by the fifth intermediate feature F a of the assistance network model. Finally, the residual feature map, the upsampled feature map, and the weight map are added together to form a new feature representation.

[0084] S14. Based on the first loss value and the first prediction loss value, train the initial student network model to obtain a target student network model.

[0085] Specifically, in the embodiments of the present application, the training iteration of the initial student network model includes two types: implicit training settings and explicit training settings.

[0086] Among them, the implicit training setting is that the initial student network model uses the same high-resolution input image I as the target assistance network model to generate intermediate features F a and F s , and uses the L HRT loss to let the initial student network model simulate the intermediate features from the target assistance network model;

[0087] The explicit training setting is that the initial student network model directly performs training with the low-resolution image as the input of the explicit training setting. The pixel size of the image is Using FSR-W in the convolutional layer of the initial student network model, when the intermediate feature F s passes through the FSR-W module and the linear layer, a student prediction score x s is generated, and the explicit loss functions L PKD and L FSR-W are calculated;

[0088] Finally, the loss of the student network model is the sum of the losses of these two training settings.

[0089] This application uses a target-assisted network model to train an initial student network model, improving the learning effect of the initial student network model on high-resolution images, and adding an FSR-W module to the convolutional layer of the initial student network model, enabling the initial student network model to learn better feature representations for classification. Through the above methods, the obtained target student network model can enhance the recognition effect on low-resolution images in the low-resolution image recognition method. The low-resolution image is enhanced through a super-resolution method, and a teacher-assistant-student distillation framework is used to combine a cross-resolution learning mechanism and an off-the-shelf knowledge distillation method; at the same time, high-resolution representation transfer learning and a feature super-resolution module improve the performance of low-resolution inputs in implicit and explicit ways respectively; by using the FSR-W module, the obtained feature maps contain richer location, appearance information, and prominent high-frequency content information, enhancing the recognition effect on low-resolution images.

[0090] The following modules are specifically described:

[0091] 1. Image super-resolution training framework

[0092] The mainstream deep learning-based image super-resolution method is to learn a mapping model to reconstruct high-resolution images from low-resolution images. During training, the super-resolution CNN takes a low-resolution LR image X as its input and outputs a reconstructed high-resolution HR image Y'. Then, the difference between the two, Y' and the real image Y, is calculated as the training loss, which can be used to guide backpropagation to optimize the weight parameters of the network. The most commonly used loss functions in image super-resolution methods are the L1 or L2 distances.

[0093] In traditional deep learning-based image super-resolution methods, all image pixels are treated equally in the loss function. In fact, for network optimization, pixels in the high-frequency region of the image play a more important role than those in the low-frequency region. Therefore, giving priority to optimizing the high-frequency region can make the method more effective. To achieve this, a weight map W is designed and added to the training framework. The weight map has the same shape as Y, which can help the network filter out the high-frequency region from the real heart rate image, enabling the network to focus on optimizing the loss in these regions and enhancing its ability to obtain more high-frequency content. After integrating with the weight map, the loss function of the new network is:

[0094]

[0095] where Y represents the low-resolution image, Y' represents the high-resolution image, W represents the weight map, w, h, and c represent the width, height, and number of channels of Y' and Y respectively, and m represents the total number of sampled samples.

[0096] 2. Weight map

[0097] To select pixels containing more high-frequency content, we propose a method of assigning weights to each pixel position. Figure 7 For the pixel neighborhood in a 3×3 true heart rate image, as Figure 7 shown, in this pixel neighborhood, the central pixel p c has four nearest neighboring pixels, located above, below, to the left, and to the right of it respectively. For each pixel neighborhood, the following formula is defined to calculate the initial weight value w:

[0098]

[0099]

[0100] w = D[arg max(D a )]

[0101] where p u , p d , p l , p r correspond to the neighboring pixels above, below, to the left, and to the right of the central pixel p c respectively, D is the difference between the central pixel p c and the four nearest neighboring pixels, D a is the absolute value of these differences, and w is the initial weight value of this central pixel.

[0102] First, calculate the differences between the central pixel p c and the four nearest neighboring pixels. Second, calculate the absolute values of these differences. Then, find the maximum absolute value and select the corresponding difference as the initial weight value w. And so on, calculate the initial weight value for each pixel. For those pixels located at the corners and edges, if they exist, we calculate their initial weight values using their two or three nearest neighboring pixels. Since the image usually has three channels, all channels are processed one by one in the same way.

[0103] The high-frequency regions in the original image are significantly represented in this generated weight map. After the weight map is initialized, the next step is to select pixels for training according to this weight map.

[0104] Since the part of the ground truth HR image most needed for network training is the high-frequency part, it is necessary to screen out the pixels containing high-frequency information from the true heart rate image. For this purpose, we modify the initial weight map according to the Gaussian distribution of the weight values. First, normalize the initial weight map values to the range of [0, 1], and then calculate their distribution with the mean μ and standard deviation σ. Then the weight values are modified by the following formula:

[0105]

[0106] Among them, w' is the modified weight value, μ is the average value of the weight value, σ is the standard deviation of the weight value, and α is a control parameter used to control and limit the total number of zero values in the weight map.

[0107] The bright part of the modified weight map represents the high-frequency content in the ground truth image. Finally, this modified weight map is applied to train the network. As a result, in the loss function, only the loss of high-frequency pixels is multiplied by the weight value 1, which means they have been selected. On the other hand, the loss of low-frequency pixels is multiplied by 0 and ignored. Since only the high-frequency part of the true heart rate image contributes to backpropagation, the network can reconstruct more high-frequency content.

[0108] 3. Teacher-Assistant-Student Framework

[0109] Figure 6 is a schematic diagram of the teacher-assistant-student distillation method. As Figure 6 shown, the traditional knowledge distillation method trains the student network by learning the output of the teacher network model. Based on the method of obtaining supervision information, we classify the previous work into two categories: prediction-based methods and feature-based methods. The prediction-based knowledge distillation method uses the class scores predicted by the teacher to train the student; the predicted class scores are:

[0110]

[0111] L pkd (x t ,x s )=(1 - α)L cls (y,x s )+αT 2 L kl (p t ,p s )

[0112] Among them, y is the true map, x t and x s are the predicted classification scores of the teacher and student network models respectively, T is a temperature parameter, and α is a hyperparameter to balance the classification loss L cls and the Kullback–Leibler divergence loss L kl .

[0113] Different from the prediction-based method, the feature-based knowledge distillation method further extracts supervision from the intermediate features of the teacher to guide the learning of the student:

[0114]

[0115] Among them, F s ={Fs (1), F s (2),..., F s (M)}, F t = {F t (1), F t (2),..., F t (M)} represent the features of the teacher and the student respectively, M is the number of network blocks, g t (·) and g s (·) denote that the function extracts information from the intermediate features, δ(·) is the distance metric function, and β is a hyperparameter to balance the classification loss and the feature distillation loss.

[0116] The goal of Teacher-Assisted-Student (TAS) is to train a student network model with the help of a teacher network model, where the network architectures and input resolutions are different. The teacher network model takes a high-resolution HR image as input and uses a heavyweight network, while the student network model takes a low-resolution LR image as input and uses a lightweight network.

[0117] This application uses the pixel distillation Teacher-Assisted-Student (TAS) framework. To reduce the performance degradation caused by the LR input and the compact structure in pixel distillation, an auxiliary network, namely the assistance network model, is introduced into the classical teacher-student framework. The assistance network model adopts the same input resolution as the teacher network model and maintains the same network architecture as the student network model. Compared with the traditional teacher-student structure, the proposed teacher-assisted student framework can bring the following advantages in pixel distillation:

[0118] 1) There is no resolution difference between the teacher network model and the assistance network model. In this case, any off-the-shelf knowledge distillation method can be directly used here. In contrast, if we use the traditional teacher-student framework here, some knowledge distillation methods cannot be used because they require the features of the teacher and the student to have the same spatial resolution.

[0119] 2) When the assistance network model and the student network model use the same input resolution, their feature maps will have the same dimensions, both in terms of spatial resolution and number of channels, which makes it easier and more effective for the student network model to learn from the HR image.

[0120] 4. High-frequency content feature super-resolution

[0121] Although the proposed high-resolution representation transfer learning scheme can provide intermediate supervision for the student network, its learning process is still not simple because the low-resolution input is not used during the training of the student network model. To explicitly mitigate the performance gap caused by the input resolution difference, we propose a Feature Super-Resolution with Weighting (FSR-W) module to enable the student network model to learn better feature representations for classification. The FSR-W module is inspired by the success of super-resolution for improving low-resolution image classification by enhancing high-frequency content and weighted reconstruction of high-resolution (HR) natural images from low-resolution (LR) images. In traditional image super-resolution tasks, the super-resolution (SR) model is expressed as:

[0122] I hr =f sr (W sr ,I lr )

[0123] where f sr (·) and W sr represent the network inference operation and trainable weights of the SR model. In the method we proposed, we have the high-resolution feature map extracted by the auxiliary network, i.e., F hr =F a . The low-resolution feature map F lr ={F lr (1),F lr (2),...,F lr (M)} and the classification scores of the predicted low-resolution input

[0124]

[0125] Since using the FSR-W module after the intermediate layer of the neural network will greatly increase the computational cost of the subsequent layers, in this work, we add the FSR-W module after the final convolutional layer. Then, the low-resolution feature map F lr (M) endeavors to approach the high-resolution feature map F hr (M) through the proposed FSR-W module:

[0126]

[0127] where f fsr-w and W fsr are the network inference operation and trainable weights of the FSR-W module, is the predicted feature map of FSR.

[0128] We use the loss function of the new network after integration with the weight map to train the FSR-W module as follows:

[0129]

[0130] Among them, N×M represents the number of elements in the feature map, and W fsr-w is the high-frequency content weight map.

[0131] The detailed structure of the FSR-W module is as Figure 2 shown. An FSR-W module consists of several residual blocks, at most 2 in this work. For a single block, we first transfer the spatial resolution in one branch. Then, in another branch, we learn the residual between the target feature map and the upsampled feature map, and the middle branch is the weight map of the target feature map. Finally, we add together the residual feature map, the upsampled feature map, and the weight map to form a new feature representation.

[0132] To reduce the computational cost and the number of parameters, we use group convolution, group deconvolution, and channel shuffling in the FSR-W module. Note that if necessary, cropping will be used to match the spatial size of the target feature map, for example, from 8×8 to 7×7.

[0133] Improving the spatial resolution of the last layer will bring the following two benefits:

[0134] First, since the learned high-resolution feature map contains more high-resolution information, it helps to improve the classification performance of the student network.

[0135] Second, the high-resolution feature map makes it possible to implement many previous image recognition algorithms at a small input resolution. By using the FSR-W module, we can obtain a feature map with a spatial resolution of 7×7, which contains richer location, appearance information, and prominent high-frequency content information.

[0136] Specifically, in the embodiments of this application, the assistance network model decomposes the learning process into the following two stages:

[0137] In the first stage, the assistance network model f t (W t , ·) is trained by learning the representation provided by the teacher f a (W a , ·), where f t (·) and f a (·) represent network inference operations, and W t and W a represent trainable weights. This is a standard knowledge distillation problem, and we can directly use a non-helicopter knowledge distillation method here to reduce the performance degradation caused by the compact network architecture. In this stage, both the teacher network model and the assistance network model take the high-resolution image I hrAs the input. Assist the network model L a The loss of can be calculated by the formula in the teacher-assistant-student framework.

[0138] In the second stage, under the guidance of the assist network model, learn the student network model f s (W s , ·). As Figure 1 shown, each training iteration of the student network model has two settings: one is the implicit training setting, where the student network uses the same high-resolution input image as the auxiliary network. In this training setting, we use a high-resolution representation transfer (HRT) learning scheme to let the student network simulate the intermediate features from the auxiliary network. Since the student network model has the same network architecture as the assistant model, when using the high-resolution image as the input, it should produce the same intermediate representation as the assistant model. In this case, we use the L2 loss to guide the intermediate features L of the student network model hrt :

[0139]

[0140] where, F a ={F a (1), F a (2),..., F a (M)} are the intermediate features of the assist network model, are the intermediate features of the student network model, N i is the number of elements of the i-th feature.

[0141] The loss in the implicit training stage is:

[0142]

[0143] where, L pkd is the loss function based on prediction, x t is the predicted classification score in the teacher network model, is the predicted classification score in the implicit learning stage, η is a hyperparameter, L hrt is the loss function used to let the student network model simulate the intermediate features from the assist network model.

[0144] Different from the implicit training setting, the student network will be directly trained with the LR image as the input in the explicit training setting. This training setting will enable the network model to perform well explicitly on the LR input. To achieve this goal, a new high-frequency content feature super-resolution FSR-W module is introduced here to explicitly reduce the performance gap caused by the LR input. The loss of this explicit training setting is defined as:

[0145]

[0146] Finally, the loss of the student network is the sum of these two training settings, L s :

[0147]

[0148] where is the loss in the implicit training phase, the loss of the explicit training setting.

[0149] This application provides a low-resolution image recognition method, including the following steps:

[0150] S21. Obtain an image to be recognized, where the image to be recognized is a low-resolution image;

[0151] S22. Input the image to be recognized into the target student network model for recognition processing to obtain an image recognition result;

[0152] where the target student network model is trained by a target assistance network model, and the target student network model includes an FSR-W module.

[0153] In some embodiments, the FSR-W module is a high-frequency content feature super-resolution module, which is used to reduce the performance difference caused by the low-resolution image.

[0154] This application provides a target student network model training device, including:

[0155] An input module, configured to input a high-resolution image into an initial student network model and a target assistance network model, where the initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature;

[0156] A first processing module, configured to input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value, where the first loss function is used to enable the initial student network model to simulate the second intermediate feature of the target assistance network model;

[0157] A second processing module, configured to input a low-resolution image into the initial student network model, obtain a third intermediate feature through the FSR-W module in the convolutional layer, and obtain a first prediction loss value of the initial student network model through a prediction loss function;

[0158] A training module, configured to train the initial student network model based on the first loss value and the first prediction loss value to obtain a target student network model.

[0159] The present application provides a low-resolution image recognition device, including:

[0160] An acquisition module, configured to acquire an image to be recognized, where the image to be recognized is a low-resolution image;

[0161] A recognition module, configured to input the image to be recognized into a target student network model for recognition processing to obtain an image recognition result; wherein, the target student network model is trained by a target assistance network model, and the target student network model includes an FSR-W module.

[0162] The present application provides a terminal device, including:

[0163] A memory, configured to store a computer program;

[0164] A processor, configured to read the computer program in the memory and execute the operations corresponding to the low-resolution image recognition method.

[0165] The present application further provides a computer-readable storage medium, where computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the low-resolution image recognition method.

[0166] It should be understood that although the steps in the flowcharts in the above embodiments are displayed in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least some of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments, and their execution order does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0167] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0168] After considering the specification and practicing the application disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0169] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for training a target student network model, characterized in that, It includes the following steps: S11. Input the high-resolution image into the initial student network model and the target assistance network model. The initial student network model obtains the first intermediate feature, and the target assistance network model obtains the second intermediate feature; S12. Input the first intermediate feature and the second intermediate feature into the first loss function to obtain the first loss value. The first loss function is used to make the initial student network model simulate the second intermediate feature of the target assistance network model; S13. Input the low-resolution image into the initial student network model, and obtain the third intermediate feature through the FSR-W module in the convolutional layer. The third intermediate feature passes through the prediction loss function to obtain the first prediction loss value of the initial student network model; S14. Based on the first loss value and the first prediction loss value, train the initial student network model to obtain the target student network model; The training process of the target assistance network model includes the following steps: S101. Input the high-resolution image into the target teacher network model, and obtain the fourth intermediate feature through the convolutional layer; S102. Input the high-resolution image into the initial assistance network model, and obtain the fifth intermediate feature through the convolutional layer; S103. Calculate the loss between the fourth intermediate feature and the fifth intermediate feature to obtain the feature loss value; S104. The initial assistance network model obtains the assistance prediction score according to the feature loss value; S105. Based on the teacher prediction score generated by the target teacher network model and the assistance prediction score, calculate and obtain the second prediction loss value; S106. According to the second prediction loss value, perform regression training on the initial assistance network model to obtain the target assistance network model; The process of obtaining the third intermediate feature through the FSR-W module in the convolutional layer includes the following steps: S131. The intermediate feature of the low-resolution image enters the FSR-W module, and the intermediate feature includes the target feature map; S132. Perform upsampling processing on the intermediate feature to obtain the upsampled feature map; S133. Calculate the residual between the target feature map and the upsampled feature map to obtain the residual feature map; S134. Obtain the weight map according to the second intermediate feature of the target assistance network model; S135. Add the upsampled feature map, the residual feature map, and the weight map to obtain the third intermediate feature; After being integrated with the weight map, the loss function of the new network is used to train the FSR-W module as follows: Among them, N×M represents the number of elements in the feature map, and W fsr-w is the high-frequency content weight map.

2. The method for training a target student network model according to claim 1, characterized in that, The training process of the target teacher network model includes: Input the high-resolution image into the initial teacher network model, and perform regression training on the initial teacher network model to obtain the target teacher network model.

3. A method for low-resolution image recognition, characterized in that, Applying the training method of the target student network model according to claim 1 includes the following steps: S21. Obtain the image to be recognized, and the image to be recognized is a low-resolution image; S22. Input the image to be recognized into the target student network model for recognition processing to obtain the image recognition result; Among them, the target student network model is trained by a target assistance network model, and the target student network model includes an FSR-W module.

4. The method for low-resolution image recognition according to claim 3, characterized in that, The FSR-W module is a high-frequency content feature super-resolution module, which is used to reduce the performance difference caused by low-resolution images.

5. A device for training a target student network model, characterized in that, The method for training the target student network model according to claim 1 is applied, including: An input module, configured to input a high-resolution image into an initial student network model and a target assistance network model. The initial student network model obtains a first intermediate feature, and the target assistance network model obtains a second intermediate feature. A first processing module, configured to input the first intermediate feature and the second intermediate feature into a first loss function to obtain a first loss value. The first loss function is used to make the initial student network model simulate the second intermediate feature of the target assistance network model. A second processing module, configured to input a low-resolution image into the initial student network model, obtain a third intermediate feature through the FSR-W module in the convolutional layer, and obtain a first prediction loss value of the initial student network model through a prediction loss function. A training module, configured to train the initial student network model based on the first loss value and the first prediction loss value to obtain a target student network model.

6. A low-resolution image recognition device, characterized in that, The method for recognizing a low-resolution image according to claim 3 is applied, including: An acquisition module, configured to acquire an image to be recognized, and the image to be recognized is a low-resolution image. A recognition module, configured to input the image to be recognized into the target student network model for recognition processing to obtain an image recognition result. Among them, the target student network model is trained by a target assistance network model, and the target student network model includes an FSR-W module.

7. A terminal device, characterized in that, Including: A memory, configured to store a computer program. A processor, configured to read the computer program in the memory and execute the operations corresponding to the method for recognizing a low-resolution image according to any one of claims 3-4.

8. A computer-readable storage medium, characterized in that, A computer-executable instruction is stored in the computer-readable storage medium, and when the computer-executable instruction is executed by the processor, it is used to implement the method for recognizing a low-resolution image according to any one of claims 3-4.

Citation Information

Patent Citations

  • Compositions and formulations for maintaining and increasing muscle mass, strength, and performance and methods of production and use thereof

    CN107223019A

  • Super-resolution restoration network model generation method, image super-resolution restoration method and image super-resolution restoration device

    CN113177888A