Low-resolution pedestrian re-identification algorithm based on scale learning

By using a rate-adaptive super-resolution network and a distillation recognition sub-network, the problem of feature extraction and recovery of low-resolution pedestrian images is solved, achieving higher recognition accuracy and efficiency, and is applicable to fields such as criminal investigation and intelligent security.

CN116030495BActive Publication Date: 2026-01-20DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211615584.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2026-01-20
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

Existing pedestrian re-identification algorithms suffer from decreased retrieval accuracy in complex scenarios due to image resolution mismatch, especially the challenges of feature extraction and restoration for low-resolution images.

Method used

A rate-adaptive super-resolution network and a distillation recognition sub-network are used to process low-resolution pedestrian images in a non-end-to-end manner, recover pixel information and extract discriminative features, and combine the feature structure learning function to eliminate differences in image feature distribution.

Benefits of technology

It improves the accuracy and efficiency of pedestrian re-identification tasks, especially in terms of recognition accuracy on low-resolution datasets, reaching the state-of-the-art level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030495B_ABST
    Figure CN116030495B_ABST
Patent Text Reader

Abstract

The application belongs to the field of computer vision, and provides a low-resolution pedestrian re-identification algorithm based on magnification learning, and proposes a new non-end-to-end magnification adaptive fusion and discriminative feature regression network model, which learns to restore different magnification low-resolution pedestrian images, and can extract pedestrian discriminative information for image retrieval. Specifically, a multi-task super-resolution identification learning network is designed, which includes a magnification adaptive super-resolution sub-network and a distillation identification sub-network. The magnification adaptive super-resolution sub-network can adaptively fuse contour information and detail information, and can restore pixel information for low-resolution images of different magnifications, reducing noise interference in pedestrian images. The distillation identification sub-network has an identification loss and a feature structure loss to guide the network to extract pedestrian discriminative features. A large number of experiments show that the RAMN model proposed by the application has reached the latest level in accuracy and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and aims to improve the recognition accuracy of low-resolution pedestrian images, and proposes a low-resolution pedestrian re-identification algorithm based on magnification learning. BACKGROUND

[0002] Pedestrian re-identification aims to match the images of the same person in multiple cameras, and is widely used in criminal investigation, intelligent security, target searching, and computer forensics. In recent years, the field of pedestrian re-identification has made great progress with the help of the powerful representation ability of deep learning. However, most of these methods only use high-resolution images as input, and match pedestrians with the same identity by extracting robust pedestrian features. In complex scenes, due to factors such as shooting angle, camera parameters, and background clutter, the image resolution is usually low and inconsistent, which will inevitably interfere with the final retrieval accuracy.

[0003] Current deep learning-based person re-identification algorithms are developing rapidly, and researchers have proposed many robust and more discriminative feature extraction methods to solve the difficulty of feature extraction in person re-identification. ABD-Net (Chen T, Ding S, Xie J, et al. Abd-net: Attentive but diverse person re-identification. ICCV, 2019) integrates channel attention and spatial attention mechanisms, and introduces SVDO regularization terms to reduce the coupling between different features of the network and improve the global diversity feature representation ability of the network. PCB (Sun Y, Zheng L, Yang Y, et al. Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline). ECCV, 2018) guides the network to focus on the feature representation of each local region of the pedestrian by horizontally cutting the feature map. Spindle Net (Zhao H, Tian M, Sun S, et al. Spindle net: Person re-identification with human body region guided feature decomposition and fusion. IEEE, 2017) introduces a semantic information extraction module to accurately align the local features of the body parts using pose estimation. PFD (Wang T, Liu H, Song P, et al. Pose-guided feature disentangling for occluded person re-identification based on transformer. AAAI, 2022) proposes a pose-guided feature disentangling method based on transform, which explicitly matches the visible body part information of the pedestrian to eliminate the interference of occlusion noise. However, these studies focus on how to produce more accurate pedestrian feature representation, while ignoring the problem of mismatched resolution of pedestrian images collected in actual scenes.

[0004] For the resolution mismatch problem, there are mainly two research directions: one is to design super-resolution and multi-task learning framework, and to recover the details of the blurred image before recognition. PRI(Han K, Huang Y, Chen Z, et al. Prediction and recovery for adaptive low-resolution person re-identification. ECCV, 2020) aims to guide the network to self-supervised prediction of the best scale factor by generating dynamic soft labels, and to recover the missing image details. INTACT(Cheng Z, Dong Q, Gong S, et al. Inter-task association critic for cross-resolution person re-identification. CVPR, 2020) aims to realize the association mechanism between tasks by using parameterized constraints, and guide the super-resolution image to learn the distribution information of the real high-resolution image. MRJL(Zhang G, Chen Y, Lin W, et al. Low resolution information also matters: Learning multi-resolution representations for person re-identification. IJCAI, 2021) aims to highlight the global information provided by the low-resolution image by multi-scale reconstruction and double feature fusion of the low-resolution image. APSR(Han K, Huang Y, Song C, et al. Adaptive super-resolution for person re-identification with low-resolution images. PR, 2021) proposes a multi-scale pixel information adaptive fusion method to suppress image artifact interference. The other research direction is image feature decoupling. DDGAN(Huang Y, Zha Z J, Fu X, et al. Real-world person re-identification via degradation invariance learning. CVPR, 2020) obtains the invariant features of the picture by multi-information separation method.Tian et al. proposed a method of mutual information learning with self-distillation and bottleneck (Tian X, Zhang Z, Lin S, et al. Farewell to mutual information: Variational distillation for cross-modal person re-identification. CVPR, 2021) to maximize the shared features to ensure feature invariance. Compared with these solutions, the present invention proposes an image magnification adaptive super-resolution method and a feature structure learning function, which can accurately predict the optimal scaling factor for super-resolution, while guiding the distribution consistency of super-resolution images and real images to improve the accuracy of pedestrian image recognition. SUMMARY

[0005] The purpose of the present invention is to provide a pedestrian re-identification network based on magnification adaptive super-resolution, which can improve the accuracy of pedestrian re-identification tasks by recovering high-quality low-resolution pedestrian images and extracting discriminative features. In this work, we propose a magnification adaptive distillation network (RAMN) to process pixel and semantic information in a non-end-to-end manner, as shown in Figure 1 Specifically, the RAMN network includes a magnification adaptive super-resolution network and a distillation recognition subnetwork. The magnification adaptive super-resolution network can recover pixel information for low-resolution images of different magnifications, reducing noise interference in pedestrian images. The distillation recognition subnetwork uses recognition loss and feature correlation loss to guide the extraction of discriminative features for pedestrians. Extensive experiments show that the proposed RAMN model has reached the latest level in terms of accuracy and efficiency.

[0006] Technical solutions of the present invention:

[0007] A low-resolution pedestrian re-identification algorithm based on magnification learning,

[0008] (1) Magnification adaptive super-resolution network

[0009] The magnification adaptive super-resolution network is composed of seven stacked units, each unit structure having three branches: H branch, M branch and a branch, which respectively realize the learning of super-resolution information, fine-grained supplementary information and fusion coefficients for low-resolution image generalization; the pedestrian image is extracted to the shallow feature F c after the convolution layer, and is input into the H branch and the M branch respectively; different magnification input images are assigned different magnification codes a in in the data processing stage, and the fusion coefficient is obtained by inputting a branch;

[0010] Shallow feature extraction:

[0011] F c = f(xL ) (1)

[0012] where x L ∈R h×w×3 , represents a low-resolution image; f is a 3x3 convolution; F c ∈R h×w×c , where h, w represent the height and width of the image features, and c is the number of channels;

[0013] The H branch is composed of a set of downward and upward projection units in the DBPN, and the branch alleviates the projection error by correcting the residual error of the projected image and the original image; this image residual re-projection method can restore the edge contour information of the pedestrian image, i.e., the generalized pedestrian information;

[0014] H t =GB(H t-1 ),H 0 =F c (2)

[0015] where GB is a projection unit, H t-1 and H t are the input and output of the tth unit, and the input of the first unit is the shallow feature F c ;

[0016] In the pedestrian re-identification task, the edge contour information cannot accurately identify a specific pedestrian, and the fine-grained information also needs to be completed; the M branch is a feature aggregation branch composed of a convolution layer and a ReLU activation function; the convolution layer realizes the extraction of fine-grained information of the image, and the M branch finally outputs the feature fusion of the pedestrian edge contour information and the fine-grained information, so as to obtain the pedestrian discriminative information;

[0017] M t =Block(H t-1 , M t-1 ), M 0 =F c (3)

[0018] where Block is a convolution unit, H t-1 and M t-1 are the input and output of the tth unit, and the input of the first unit is the shallow feature F c ;

[0019] In order to ensure that the magnification adaptive supermolecular network can accurately super for input images of various magnifications, a linear guide branch a is introduced to dynamically control the proportion of different feature information according to the different magnifications of the image input;

[0020] a t =T(a t-1 ), a 0 =a in, a t = σ(α t ), b t = 1-a t (4)

[0021] M t = a t H t +b t (DB(M t-1 )) (5)

[0022] where T is a linear unit consisting of three fully connected layers, α t-1 and α t are the input and output of the t-th linear unit; the first unit input is the scale code α in ; σ is a sigmoid function that maps the output of the fully connected layer to the proportion of the H branch and the M branch a t and b t ; formula (5) is the fusion formula of the unit structure discriminative feature;

[0023] The scale adaptive super-molecular network integrates the fusion features of all stages of the M branch as the final output, and the integrated features are reduced to the final high-resolution image through 3x3 convolution;

[0024] x SR = f Rec (M 1 , M 2 ...M 7 ) (6)

[0025] where f Rec is a concat operation, f Rec is a 3x3 convolution dimension reduction operation, and x SR is a pedestrian super-resolution image;

[0026] (2) Distillation recognition sub-network

[0027] HRNet-ReID network is used as the feature extraction backbone network, and HRNet-ReID network learns different resolution feature information from high to low in four parallel branches; to adapt to the re-ID image retrieval requirement, the last feature map of each branch is compressed into a feature sequence as the final output through average pooling and maximum pooling; two high-resolution information branches learn spatial position information, and two low-resolution information branches highlight pedestrian semantic understanding; HRNet-ReID network maintains the high spatial sensitivity of the learned features by interacting and fusing the information of the four branches, and highlights the expression of fine-grained semantic information;

[0028] To eliminate the distribution domain difference between super-resolution images and real images in the feature space, a distillation structure model is designed. The distillation structure model is composed of two identical HRNet-ReID networks: a real high-resolution image discriminative feature extraction branch and a super-resolution high-resolution image discriminative feature extraction branch, and the network parameters are not shared. At the same time, a feature structure learning formula is proposed to guide the super-resolution image branch and the real image branch to learn the consistency of feature distribution. The feature structure learning formula is as follows:

[0029]

[0030] where h is the feature map learned by the network, h(u) and h(v) are the spatial positions of any two points in the feature map, ||h(u)||2 is the L2 norm of the feature map where u is located, represents the normalization of the feature map where u is located, is the transpose operation on the obtained normalized matrix; The solving process is the same as Through this formula, the mutual relationship between any two points in the feature map can be calculated, and a feature structure relationship graph with a size of HWxHW is learned.

[0031] (3) Design of loss function

[0032] The complete multi-task loss is represented as follows:

[0033]

[0034] where, is the super-resolution image pixel loss; and are the cross-entropy loss and the hard sample mining triplet loss commonly used in pedestrian re-identification tasks; and are the strong and weak correlation losses of the distillation network; the definitions and constraint ranges of each loss function are as follows, and λ1, λ2, λ3, λ4, λ5 are the weighting parameters.

[0035]

[0036] where x SR is the pedestrian super-resolution image, x HR is the corresponding high-resolution image, and the target result of the task;

[0037]

[0038] The final output of the HRNet-ReID network includes four feature sequences of different resolutions, the whole sequence formed by concatenating all branches, and the feature classification sequence after the classification layer. The final output of the six feature sequences is denoted as [PART]​1-4 ,[GLB],[CLS]; is the probability that x belongs to the real label;

[0039]

[0040] where d p ,d n ∈{[PART] 1-4 ,[GLB]},d p is the positive sample sequence, d n is the negative sample sequence;Alpha is the threshold value;The high-resolution image feature extraction branch and the super-resolution image feature extraction branch of the distillation recognition subnetwork are both subjected to and loss constraints;

[0041] To narrow the feature distribution difference between the super-resolution image and the real image, a structure correlation loss is proposed to transfer the feature relationship of the real image to the super-resolution image, guiding the recognition network to accurately extract the discriminative features of the super-resolution image;Formula (12) is a strong connection of feature structure, S h is the output of the high-resolution image feature extraction branch, S l is the output of the super-resolution image feature extraction branch, S h ,S l ∈{[PART] 1-4 ,[GLB]},represent the mutual learning of the six sequences of the two branches of the distillation network;

[0042]

[0043]

[0044] According to the proposed feature structure learning formula, a feature structure weak connection loss function formula (13) is designed;High-level semantics can more accurately express feature structure information, so low-resolution features are selected as the structure learning benchmark;Therefore, f h and f s respectively represent the feature maps of [PART]3 and [PART]4 branches in the high-resolution image feature extraction branch and the super-resolution image feature extraction branch of the distillation network.

[0045] The beneficial effects of the application are as follows:

[0046] (1) We propose a recognition learning network based on adaptive magnification, which learns image super-resolution and image retrieval from both pixel and feature aspects, to solve the serious problem of low resolution pedestrian dataset retrieval accuracy in pedestrian re-identification algorithm.

[0047] (2) We propose an adaptive super-resolution guidance mechanism based on the difference in magnification, which guides the overall image recovery of pedestrians by self-adjusting the supplementary image contour and fine-grained information.

[0048] (3) We design a feature structure loss function to guide the consistent expression of different magnification images of the same pedestrian, and eliminate the difference in image feature distribution. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 The network framework for the low-resolution pedestrian re-identification algorithm proposed in the application.

[0050] Figure 2 The super-resolution network unit structure framework proposed in the application.

[0051] Figure 3 The distillation structure framework proposed in the application.

[0052] Figure 4 The feature structure formula diagram proposed in the application.

[0053] Figure 5 The total loss function curve of the application.

[0054] Figure 6 The super-resolution result comparison diagram of different training strategies in the application. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical scheme and advantages of the application more clear, the specific embodiments of the application are described in further detail below.

[0056] Our experiments involve four data sets, including two high-resolution data sets Market1501 and DukeMTMC-reid and two synthetic low-resolution data sets MLR-Market1501 and MLR-DukeMTMC-reid. The experimental evaluation is carried out on two synthetic low-resolution data sets: MLR-Market1501 includes 32,668 images of 1,501 pedestrians taken by 6 cameras. MLR-DukeMTMC-reid includes 36,411 images of 1,404 pedestrians taken by 8 cameras. The low-resolution data set image adopts the unified MLR processing standard: for all the images of the same pedestrian under the same camera, a camera is randomly selected, and the image under the camera is randomly down-sampled, the down-sampling coefficient r∈{2,3,4}, as the low-resolution image query library, and the rest of the camera pictures remain unchanged, as the high-resolution image library.

[0057] Alpha inSetting requirements: in the image preprocessing process, the down-sampled image is encoded and paired. When γ = 2, the encoding setting is [1, 0, 0]; when γ = 3, the encoding setting is [0, 1, 0]; and when γ = 4, the encoding setting is [0, 0, 1]. In the training stage, the image and the corresponding encoding are input into the super-resolution network, and the network is trained to have super-resolution capability for different magnification images. In the test stage, all the encoding settings are [1, 1, 1], and the network can adaptively input the magnification information of the image to super-resolve the accurate image.

[0058] Implementation details: in this embodiment, the training process is divided into three stages. In the first stage, we only train the super-resolution module; in the second stage, we train the identification performance of the distillation upper branch network using the traditional high-resolution dataset; and in the third stage, we perform overall training of the super-resolution and identification network on the low-resolution dataset. We use the momentum SGD optimizer, set the initial learning rate to 0.003, and multiply it by 0.1 after every 30 epochs. The weight decay rate is set to 5e-4, and the network is trained for a total of 50 epochs. The hyperparameters λ1, λ2, λ3, λ4 and λ5 are 1, 1.15, 0.2, 0.5 and 0.02, respectively.

[0059] Comparative experiment: as shown in Table 1, we show the recognition accuracy results of different methods on two commonly used datasets. The last row is the method of the present application, and the first to ninth rows are the experimental results of the current low-resolution pedestrian re-identification mainstream methods, which almost cover all the top methods in the field. According to the comparison results of the experiment, it can be seen that the method proposed in the present application achieves the most advanced performance on the two datasets. Compared with the PS-HRNet method, we perform learning of the magnification adaptive super-resolution image and constraint of the structure loss function, and the Rank1 accuracy is improved by 1.7% and 4.4% on the two datasets, respectively, which confirms the effectiveness of the method of the present application.

[0060] Table 1 Results of different methods on MLR-Market1501 and MLR-DukeMTMC-reid

[0061]

[0062]

[0063] Ablation experiment: to verify the effectiveness of each module, the present application performs an ablation experiment on all loss functions on the MLR-Market1501 dataset to illustrate the effectiveness of the modules proposed by the present application, and the experimental results are shown in Table 2. represents to ensure that the network realizes the basic pedestrian identification function. to introduce the super-resolution module. to the distillation network feature structure strength connection loss Ablation experiments show that the proposed super-resolution module restores the discriminative information of pedestrian images, and the recognition accuracies of Rank1, Rank5 and Rank10 are improved by 1.8%, 0.9% and 0.4% respectively. The feature structure loss proposed in the application reduces the distribution difference of similar images, and the recognition accuracies of Rank1, Rank5 and Rank10 are improved by 2.1%, 0.8% and 0.5% respectively on the basis of the original.

[0064] Table 2 Ablation experiment recognition results of different loss functions on MLR-Market1501

[0065]

[0066] Iteration number analysis: In order to facilitate model training and evaluate network efficiency, the loss function change curve is as shown in Figure 5 The test set results and the graph prove that the experimental results tend to be stable at the 50th epoch, which proves that the scheme can quickly achieve good results in the pedestrian re-identification task.

[0067] Non-end-to-end training strategy: Due to the invisible characteristics of the training set and the test set of the pedestrian re-identification task, the discriminative regions of the pedestrians located by the training and the test set will be different. Previous end-to-end training methods cause the network to only focus on the recovery of the discriminative region of the training set, resulting in poor recovery effect of the discriminative region of the test set. The application adopts a non-end-to-end training strategy, which restores all regions of the low-resolution image by training the super-resolution network to avoid the loss of important information. As shown in Figure 6 The comparison chart of super-resolution output of different training strategies can be seen that the non-end-to-end training strategy can more clearly restore the pedestrian image region and line contour. The experimental results in Table 3 also prove that the non-end-to-end strategy is better than the end-to-end in the pedestrian re-identification task.

[0068] Table 3 Recognition results of different training strategies of the application on MLR-Market1501

[0069] Training strategy Rank 1 Rank 5 Rank 10 End-to-end 91.7 97.1 97.9 Non-end-to-end 93.2 97.6 98.3

[0070] The above is the specific embodiment of the application and the technical principle used, if the change is made according to the concept of the application, the function generated still does not exceed the spirit covered by the specification and drawings, and still belongs to the protection scope of the application.

Claims

1. A low-resolution pedestrian re-identification algorithm based on magnification learning, characterized in that, (1) Ratio-adaptive supramolecular network The rate-adaptive supramolecular network consists of units stacked seven times, each unit having three branches: the H branch, the M branch, and... The branches respectively implement the super-resolution information, fine-grained supplementary information, and fusion coefficients for learning low-resolution image generalization; pedestrian images are processed through convolutional layers to extract shallow features. Then, the H branch and M branch are input separately; different magnification input images are assigned different magnification codes during the data processing stage. ,enter The branch obtains the fusion coefficient; Shallow feature extraction: (1); in, , representing a low-resolution image; It is convolution; , where h and w represent the height and width of the image feature, and c is the number of channels; The H branch consists of a set of downward and upward projection units in the DBPN. This branch alleviates projection errors by correcting the residuals between the projected image and the original image. This image residual reprojection method can recover the edge contour information of the pedestrian image, that is, the generalized pedestrian information. (2); in, It is a projection unit. and These are the input and output of the t-th unit, where the input of the first unit is the shallow features. ; In pedestrian re-identification tasks, edge contour information cannot accurately identify specific pedestrians, and fine-grained information also needs to be supplemented. The M branch is a feature aggregation branch, which consists of convolutional layers and ReLU activation functions. The convolutional layers realize the extraction of fine-grained information from the image, and the M branch finally outputs feature fusion of pedestrian edge contour information and fine-grained information to obtain pedestrian discriminative information. (3) ; in, It is a convolutional unit. and These are the input and output of the t-th unit, where the input of the first unit is the shallow features. ; To ensure accurate super-resolution of input images at various magnifications by the magnification-adaptive supramolecular network, a linear guiding branch is introduced. The proportion of different feature information is dynamically controlled according to different image input magnification. (4) ; (5) ; in, It is a linear unit, consisting of three fully connected layers. and These are the inputs and outputs of the t-th linear unit; the input of the first unit is the rate-encoded value. ; It is the sigmoid function, which maps the output of a fully connected layer to the proportions of the H-branch and M-branch. and Formula (5) is the fusion formula for the discriminative features of unit structure; The rate-adaptive supramolecular network integrates the fusion features of all M-branch stages as the final output. The integrated features are then processed... Convolutional dimensionality reduction yields the final high-resolution image; (6) ; It's a concat operation. yes Convolution dimensionality reduction operation, It is a pedestrian over-resolution image; (2) Distillation recognition subnetwork The HRNet-ReID network is used as the backbone network for feature extraction. The HRNet-ReID network learns feature information from different resolutions from high to low in parallel through four branches. To adapt to the requirements of re-ID image retrieval, the feature map of each branch is compressed into a feature sequence as the final output after average pooling and max pooling. Two high-resolution information branches learn spatial location information, and two low-resolution information branches highlight the semantic understanding of pedestrians. The HRNet-ReID network maintains the high spatial sensitivity of the learned features and highlights the expression of fine-grained semantic information by interacting and fusing the information from the four branches. To eliminate the difference in feature space distribution between super-resolution and real images, a distillation structure model is designed. This model consists of two identical HRNet-ReID networks: a discriminative feature extraction branch for the real high-resolution image and a discriminative feature extraction branch for the super-resolution high-resolution image, with no parameter sharing. A feature structure learning formula is also proposed to guide the super-resolution image branch and the real image branch in learning feature distribution consistency. The feature structure learning formula is shown below: (7) ; in, These are feature maps learned by the network. and It is the spatial position of any two points in the feature map. and , It is the feature map where u is located. Norm, This represents normalization of the feature map containing u. This involves transposing the obtained normalized matrix. Solution process and Similarly, this formula (7) can be used to calculate the relationship between any two points in the feature map, and learn the relationship between them. A diagram showing the structural relationship between size characteristics; (3) Design of the Loss function Complete multi-task loss It is expressed as follows: (8) ; in, It is the pixel loss of the super-resolution image; and These are the cross-entropy loss and hard sample mining triplet loss commonly used in pedestrian re-identification tasks; and This refers to the strong and weak correlation loss of the distillation network; the definitions and constraints of each loss function are as follows. It is a trade-off parameter; (9) ; It is a pedestrian over-resolution image. This is the corresponding high-resolution image, representing the target result of the task. (10); The final output of the HRNet-ReID network includes four feature sequences at different resolutions, an integral sequence formed by concatenating all branches, and a feature classification sequence after passing through a classification layer. The six final output feature sequences are denoted as follows: ; yes The probability of belonging to a true label; (11) ; in, , It is a positive sample sequence. It is a negative sample sequence; It is a threshold; both the high-resolution image feature extraction branch and the super-resolution image feature extraction branch of the distillation recognition subnetwork are affected. and Loss constraints; To reduce the difference in feature distribution between the super-resolution image and the real image, a structural correlation loss is proposed, which transfers the feature relationships of the real image to the super-resolution image, guiding the recognition network to accurately extract the discriminative features of the super-resolution image; Formula (12) represents a strong connection of feature structure. It is the output of the high-resolution image feature extraction branch. It is the output of the super-resolution image feature extraction branch. This represents the mutual learning of six sequences across two branches of the distillation network; (12) ; (13) ; Based on the proposed feature structure learning formula, a feature structure weak connection loss function formula (13) is designed; high-level semantics can more accurately express feature structure information, so low-resolution features are selected as the structure learning benchmark; therefore and These represent the high-resolution image feature extraction branch and the super-resolution image feature extraction branch of the distillation network, respectively. and Characteristic diagram of a branch road.

Citation Information

Patent Citations

  • Lightweight image super-resolution reconstruction method based on multi-dimensional knowledge distillation

    CN113240580A

  • Target detection method and device based on medium and low resolution remote sensing images and equipment

    CN113705532A