Lightweight stereo image super-resolution reconstruction method and device
The lightweight stereo image super-resolution reconstruction network is constructed through adaptive pruning and cross-viewpoint distillation mechanisms, which solves the problems of network redundancy and computational overhead, and realizes efficient stereo image super-resolution reconstruction.
Patent Information
- Application Number
- CN202510410987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In the existing three-dimensional image super-resolution reconstruction method, the network structure has redundancy and the computing overhead of the mutual attention mechanism is large, resulting in high computational complexity and it is difficult to achieve a balance between computational complexity and performance.
Adaptive pruning strategy and cross-viewpoint distillation mechanism are adopted to build a teacher network, pruning residual blocks and mutual attention mechanism through the gated unit of importance estimation, and explore the correlation between viewpoints through the cross-viewpoint hierarchical attention module, providing supervision signals for the student network and building a lightweight student network.
While reducing the network computing complexity, the consistency of matching relationships between viewpoints is maintained, improving the super-resolution reconstruction performance and pixel-level accuracy of stereoscopic images.
Smart Images

Figure CN120355573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing and deep learning, and particularly to a lightweight stereoscopic image super-resolution reconstruction method and device based on adaptive pruning and cross-viewpoint distillation. Background Art
[0002] In recent years, stereoscopic image super-resolution reconstruction, as a key technology for obtaining high-quality and high-resolution stereoscopic images, has gradually become a research hotspot in the academic community. This technology aims to reconstruct high-resolution stereoscopic images by using low-resolution stereoscopic images captured by a binocular camera. Currently, stereoscopic image super-resolution reconstruction plays an important role in many 3D edge devices and applications. For example: smartphones, smart glasses, tablets, etc. However, existing stereoscopic image super-resolution reconstruction methods mainly focus on the quality of the reconstructed image, ignoring the high computational complexity of deep learning models, thus resulting in high computational costs. Therefore, how to study lightweight stereoscopic image super-resolution reconstruction methods for stereoscopic image characteristics to achieve a balance between computational complexity and performance still needs further exploration.
[0003] Early on, researchers proposed super-resolution reconstruction methods for single images, that is, converting the input low-resolution image into a high-resolution image through a pre-set mathematical model. Typical methods include: interpolation-based methods and sparse representation-based methods, such as: linear interpolation, bilinear interpolation, sparse learning dictionaries, etc. Compared with traditional methods, deep learning has shown excellent performance in the field of single-image super-resolution reconstruction. For example: Dai et al. designed a second-order channel attention module to adaptively modulate the features of different channels, and at the same time proposed a non-local enhanced residual group to capture spatial context information, effectively improving the quality of super-resolution reconstruction. Fang et al. proposed a super-resolution reconstruction network assisted by soft edges, and reconstructed high-resolution images by fusing rough super-resolution feature maps and super-resolution soft edges. Compared with using single-image methods to reconstruct the left and right views separately, the super-resolution reconstruction method for stereoscopic images can further improve the performance of super-resolution reconstruction by exploring the inter-viewpoint relationship between the left and right views. For example, Wang et al. integrated the left and right viewpoint information by exploring the inter-viewpoint consistency relationship on the horizontal epipolar line of the left and right viewpoints, effectively improving the performance of stereoscopic image super-resolution reconstruction. Ying et al. used existing single-image super-resolution reconstruction methods to extract features, and combined a stereoscopic attention module to mine inter-viewpoint information, thereby reconstructing high-resolution stereoscopic images. Chen et al. overcame the limitation that the disparity attention only obtains information on the epipolar line by expanding the receptive field, and used a single model to achieve super-resolution reconstruction with different magnification factors.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following disadvantages and deficiencies in the prior art:
[0005] When constructing a network, existing methods usually set the same number of channels for different layers of the network. However, considering that different layers in the network have different degrees of importance for the reconstruction quality, setting the same number of channels for each layer of the network will result in redundancy in the network structure. In addition, existing methods often explore the information between viewpoints based on the mutual attention mechanism. When calculating the inter-viewpoint correlation of the features of each layer of the left and right viewpoints, the mutual attention mechanism will generate a huge computational overhead. Summary of the Invention
[0006] The present invention provides a lightweight method and device for super-resolution reconstruction of stereo images. The present invention first proposes an adaptive pruning strategy, constructs a teacher network by designing a gating unit based on importance estimation, and adaptively prunes the residual blocks and mutual attention mechanism in the corresponding gating unit to construct a lightweight student network; then, a cross-viewpoint distillation mechanism is proposed. By constructing a cross-viewpoint hierarchical attention module, the inter-viewpoint correlation of multi-level features of the left and right viewpoints is explored, so as to provide an additional supervision signal for the student network, as described in detail below:
[0007] In a first aspect, a lightweight method for super-resolution reconstruction of stereo images, the method includes:
[0008] Propose an adaptive pruning strategy, construct a teacher network based on a gating unit for importance estimation, prune the residual blocks and mutual attention mechanism in the corresponding gating unit, and obtain an optimized teacher network to construct a student network;
[0009] Propose a cross-viewpoint distillation mechanism, and explore the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the optimized teacher network, as well as the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the student network by constructing a cross-viewpoint hierarchical attention module, and obtain correlation matrices respectively;
[0010] Use the root mean square error loss function to measure the difference between the correlation matrices, provide an additional supervision signal for the student network, and obtain a lightweight student network;
[0011] In the inference stage, only the lightweight student network is used to reconstruct the high-resolution stereo image.
[0012] Among them, the adaptive pruning strategy is used to obtain a lightweight student network by adaptively pruning the residual blocks and mutual attention mechanism at different levels in the teacher network. Among them, the gating unit for importance estimation is: the importance of the residual block for extracting single-viewpoint features: using the importance β n of the mutual attention mechanism as a weight, and using soft gating to perform weighted fusion on the input features and output features of the mutual attention mechanism, which is expressed as follows:
[0013]
[0014] Among them, CrossAttn(·) represents the cross-attention mechanism, and f l n+1 and respectively represent the left-viewpoint feature and the right-viewpoint feature output by the nth gating unit;
[0015] Among them,
[0016]
[0017] Among them, ReLU(·) represents the ReLU activation function, Cos(·) represents the cosine similarity calculation, and β n ∈[0,1] represents the importance of the cross-attention mechanism in the nth gating unit obtained by calculation, represents the importance of the cross-attention mechanism in the nth gating unit of the left-viewpoint feature; represents the importance of the cross-attention mechanism in the nth gating unit of the right-viewpoint feature.
[0018] Among them, the left-viewpoint feature and the right-viewpoint feature output by the gating unit are:
[0019] Taking the importance α n as the weight, and using soft gating to perform weighted fusion on the input feature and the output feature of the residual block, which is expressed as follows:
[0020]
[0021] Among them, Res(·) represents the residual block, and ⊙ represents pixel-wise multiplication, and respectively represent the left-viewpoint feature and the right-viewpoint feature output by the residual block;
[0022]
[0023] Among them, Conv(·) represents the convolutional layer, GP(·) represents the global pooling layer, Sigmoid(·) represents the Sigmoid activation function, Avg(·) represents calculating the average value, and α n ∈[0,1] represents the importance of the residual block in the nth gating unit finally obtained by calculation, represents the importance of the residual block in the nth gating unit of the left-viewpoint feature; represents the importance of the residual block in the nth gating unit of the right-viewpoint feature.
[0024] Among them, pruning the residual block and the cross-attention mechanism in the corresponding gating unit, the optimized teacher network obtained is:
[0025]
[0026] Among them, represents the number of channels of the residual block in the n-th gating unit in the teacher network, represents the number of channels of the n-th residual block in the pruned student network, and × represents the multiplication operation, represents the ceiling operation, represents the flag bit indicating whether to introduce the mutual attention mechanism after the n-th residual block in the student network, Net t and Net s represent the teacher network and the student network respectively. Prun(·) represents the process of pruning the network according to the calculated number of channels and the flag bit, and N represents the total number of gating units.
[0027] Among them, the proposed cross-viewpoint distillation mechanism explores the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the optimized teacher network by constructing a cross-viewpoint hierarchical attention module as follows:
[0028] For the multi-level features of the left and right viewpoints obtained by the student network and By constructing a cross-viewpoint hierarchical attention module, the inter-viewpoint attention map M s ;
[0029] Construct an inter-viewpoint correlation distillation constraint, and use the inter-viewpoint attention map M t of the teacher network to constrain the inter-viewpoint attention map M s of the student network, and improve the accuracy of the inter-viewpoint relationship of each level of features extracted by the student network. The inter-viewpoint correlation distillation constraint is expressed as follows:
[0030]
[0031] Among them, represents the mean square error loss, and L distillation represents the inter-viewpoint correlation distillation constraint.
[0032] Among them, the inter-viewpoint attention map is:
[0033]
[0034] Among them, M t ∈R HW×N×N represents the inter-viewpoint attention map of the multi-level features of the left and right viewpoints in the teacher network, represents matrix multiplication, · T represents the matrix transpose operation, and R(·) represents the dimension transformation operation, which transforms the dimensions of Q l and Q r from R N×C×H×W to R HW×N×C .
[0035] Among them, the
[0036]
[0037] Among them, Q l and Q r respectively represent the query vectors of the left view point and the right view point, R(·) represents a dimensionality transformation operation, which and are respectively transformed into feature vectors with a dimension of R N×C×H×W , C represents the number of channels, H and W respectively represent the width and height of the feature, and proj(·) represents a linear mapping.
[0038] In a second aspect, a lightweight stereoscopic image super-resolution reconstruction device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method described in any one of the first aspect.
[0039] In a third aspect, a computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is enabled to execute the method described in any one of the first aspect.
[0040] The beneficial effects of the technical solution provided by the present invention are:
[0041] 1. This method can effectively reduce the network calculation complexity while obtaining better stereoscopic image super-resolution reconstruction performance;
[0042] 2. In the process of reconstructing high-quality and high-resolution stereoscopic images, this method can maintain the consistency of the viewpoint matching relationship while paying attention to the pixel-level accuracy within each view. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a schematic diagram of the network structure of the gating unit based on importance estimation;
[0044] Figure 2 is a schematic diagram of the qualitative comparison result of the stereoscopic image super-resolution reconstruction result;
[0045] Figure 3 is a flowchart of a lightweight stereoscopic image super-resolution reconstruction method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0047] Example 1
[0048] An embodiment of the present invention designs a lightweight-based three-dimensional image super-resolution reconstruction method. See Figures 1 - 3 , and this method includes the following steps:
[0049] 101: Propose an adaptive pruning strategy, use the designed gating unit based on importance estimation to construct a teacher network, calculate the importance of each layer in the teacher network, and prune the residual blocks and mutual attention mechanisms in the corresponding gating units, so as to obtain a more efficient network structure to construct a student network;
[0050] 102: Propose a cross-viewpoint distillation mechanism. By constructing a cross-viewpoint hierarchical attention module, explore the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the teacher network and the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the student network , obtain the correlation matrices M and the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the student network ; obtain the correlation matrix M t and M s ;
[0051] 103: By using the root mean square error loss function to measure the difference between the matrices M t and M s , thereby providing an additional supervision signal for the student network and further improving the three-dimensional image super-resolution reconstruction performance of the student network;
[0052] 104: In the inference stage, only use the lightweight student network to reconstruct high-resolution three-dimensional images with high quality.
[0053] In summary, through the above steps 101-104, the embodiment of the present invention can effectively reduce the network calculation complexity while obtaining better three-dimensional image super-resolution reconstruction performance, meeting various needs in practical applications.
[0054] Embodiment 2
[0055] Next, the solution in Embodiment 1 will be further described in combination with specific calculation formulas and calculation examples. See the following description for details:
[0056] I. Adaptive Pruning Strategy
[0057] Since different layers in the network contribute differently to the overall performance, more channels should be allocated to those layers that have a more significant impact on performance to enhance their feature expression ability. Conversely, for layers with relatively lower importance and less contribution to performance, the number of channels should be appropriately reduced to optimize the allocation of computing resources and improve the overall efficiency of the network. At the same time, calculating the viewpoint - to - viewpoint correlation of each level of features in the network using the mutual - attention mechanism requires a huge amount of computational effort. Considering introducing the mutual - attention mechanism only at specific levels of the network to achieve a balance between computational complexity and performance. Therefore, the embodiments of the present invention propose an adaptive pruning strategy, which obtains a lightweight student network by adaptively pruning the residual blocks and mutual - attention mechanisms at different levels of the teacher network.
[0058] To achieve the adaptive pruning of the teacher network, a gating unit based on importance estimation is designed, and the designed gating unit is used to construct the teacher network. Figure 1 The network structure of the gating unit based on importance estimation is shown. Taking the nth gating unit as an example, for the input left - viewpoint feature f l n and the right - viewpoint feature First, estimate the importance of the residual block used to extract single - viewpoint features in this gating unit. This process can be expressed by the following formula:
[0059]
[0060] where Conv(·) represents the convolutional layer, GP(·) represents the global pooling layer, Sigmoid(·) represents the Sigmoid activation function, and Avg(·) represents calculating the average. α n ∈[0,1] represents the importance of the residual block in the nth gating unit finally calculated, represents the importance of the residual block in the nth gating unit for the left - viewpoint feature; represents the importance of the residual block in the nth gating unit for the right - viewpoint feature.
[0061] After calculating the importance α n of the residual block, using α n as the weight, the input features and output features of the residual block are weighted and fused using a soft gating. This process can be expressed by the following formula:
[0062]
[0063] where Res(·) represents the residual block, and ⊙ represents element - by - element multiplication. and respectively represent the left - viewpoint feature and right - viewpoint feature output by the residual block.
[0064] After that, based on the left-viewpoint and right-viewpoint features and estimate the importance of the mutual attention mechanism in the gating unit. This process can be expressed by the following formula:
[0065]
[0066] where ReLU(·) represents the ReLU activation function, and Cos(·) represents the cosine similarity calculation. β n ∈[0,1] represents the importance of the mutual attention mechanism in the nth gating unit obtained by calculation, represents the importance of the mutual attention mechanism in the nth gating unit of the left-viewpoint feature; represents the importance of the mutual attention mechanism in the nth gating unit of the right-viewpoint feature.
[0067] After calculating the importance β n of the mutual attention mechanism, use β n as the weight and perform weighted fusion on the input features and output features of the mutual attention mechanism using soft gating. This process can be expressed by the following formula:
[0068]
[0069] where CrossAttn(·) represents the mutual attention mechanism. f l n+1 and respectively represent the left-viewpoint feature and right-viewpoint feature output by the nth gating unit.
[0070] Finally, after completing the training of the teacher network, prune the residual blocks and mutual attention mechanisms in the gating unit according to the importance calculated by each gating unit in the teacher network to obtain a lightweight student network. This process is expressed by the following formula:
[0071]
[0072] where represents the number of channels of the residual block in the nth gating unit of the teacher network, represents the number of channels of the nth residual block in the pruned student network. × represents the multiplication operation, represents the ceiling operation. represents the flag indicating whether to introduce the mutual attention mechanism after the nth residual block in the student network. If the flag is False, the mutual attention mechanism is not introduced; if the flag is True, the mutual attention mechanism is introduced. Net t and Net srespectively represent the teacher network and the student network, Prun(·) represents the process of pruning the network according to the calculated number of channels and flag bits, and N represents the total number of gating units.
[0073] II. Cross-Viewpoint Distillation Mechanism
[0074] Considering that there is a disparity between different viewpoints of stereo images, reconstructing high-quality and high-resolution stereo images requires maintaining the consistency of the matching relationship between viewpoints while paying attention to the pixel-level accuracy within each view. In view of this, the embodiments of the present invention propose a cross-viewpoint distillation mechanism to further improve the reconstruction performance of the student network by constructing a feature-level inter-viewpoint correlation distillation constraint.
[0075] Since the residual blocks used to extract left and right viewpoint features in the teacher network have a higher number of channels, and the mutual attention mechanism is used to explore the inter-viewpoint information of each layer of features, the left and right viewpoint features obtained by the teacher network have a more accurate inter-viewpoint relationship compared to the student network. Therefore, by calculating the inter-viewpoint correlation matrix of different viewpoint features in the teacher network to guide the student network to extract more accurate left and right viewpoint features, it helps to improve the performance of the student network. Specifically, for the multi-level left-viewpoint features and multi-level right-viewpoint features obtained by the teacher network, a cross-viewpoint hierarchical attention module is constructed to explore the inter-viewpoint correlation. In the cross-viewpoint hierarchical attention module, first, and are subjected to dimensionality transformation, and the features after dimensionality transformation are projected into query vectors using a linear mapping. This process can be expressed by the following formula:
[0076]
[0077] where Q l and Q r represent the query vectors of the left and right viewpoints respectively. R(·) represents the dimensionality transformation operation, which transforms and into feature vectors with a dimension of R N×C×H×W respectively, C represents the number of channels, and H and W represent the width and height of the feature respectively. proj(·) represents the linear mapping.
[0078] Then, the query vectors of the left viewpoint and the query vectors of the right viewpoint are subjected to dimensionality transformation and matrix multiplication to calculate and obtain the inter-viewpoint attention map of the multi-level left and right viewpoint features in the teacher network. This process can be expressed by the following formula:
[0079]
[0080] where M t ∈R HW×N×NIndicates the inter-viewpoint attention map of the left and right view multi-level features in the teacher network. Indicates matrix multiplication, · T Indicates the matrix transpose operation. R(·) indicates the dimension transformation operation, which transforms Q l and Q r from dimension R N×C×H×W to R HW×N×C .
[0081] Similarly, for the left and right view multi-level features obtained by the student network and By constructing a cross-view hierarchical attention module, the inter-viewpoint attention map M s can be calculated. Finally, an inter-viewpoint correlation distillation constraint is constructed, and the inter-viewpoint attention map M t of the teacher network is used to constrain the inter-viewpoint attention map M s of the student network, thereby improving the accuracy of the inter-viewpoint relationship of each level of features extracted by the student network. The inter-viewpoint correlation distillation constraint can be expressed by the following formula:
[0082]
[0083] where, represents the mean square error loss, and L distillation represents the inter-viewpoint correlation distillation constraint.
[0084] III. Loss Function
[0085] For the teacher network, in the embodiments of the present invention, the mean square error loss is used to measure the difference between the reconstructed image and the high-resolution ground truth. The loss function of the teacher network can be expressed as follows:
[0086]
[0087] where, L SR represents the loss function of the teacher network. and respectively represent the high-resolution left and right view ground truths in the super-resolution reconstruction task. and respectively represent the left and right views reconstructed by the teacher network.
[0088] For the student network, in order to ensure the training stability of the student network, in the first 15 epochs, only the high-resolution ground truth is used to constrain the image reconstructed by the network based on the mean square error loss; afterwards, on the basis of the mean square error loss, the inter-viewpoint correlation distillation constraint is introduced to train the student network together, and the weight of the inter-viewpoint correlation distillation constraint is set to 0.1.
[0089] In the teacher network and the student network, the residual block adopts the intra-frame block structure in the non-linear non-activation network and shares weights between the left and right viewpoints. The reconstruction layer uses a sub-pixel convolutional layer to obtain a high-resolution image. For the teacher network, it contains a total of 32 gating units based on importance estimation, and the number of channels of the residual blocks in each gating unit is set to 96, that is, N = 32, For the student network, a convolutional layer with a convolution kernel of 1×1 is inserted between the residual blocks to ensure the concatenation between the residual blocks. For the proposed cross-viewpoint distillation mechanism, the teacher network and the student network share weights in the cross-viewpoint hierarchical attention module.
[0090] Figure 2 The qualitative comparison results of the method proposed in the embodiment of the present invention and the comparison method for 4-fold super-resolution reconstruction are shown. The comparison algorithms include: the StereoSR method and the CVCnet method. As Figure 2 shown, compared with other lightweight methods, the images reconstructed by the method proposed in the present invention are closer to the real high-resolution images. It can be seen from the enlarged areas that both the left view and the right view reconstructed by the method proposed in the embodiment of the present invention have higher quality. Taking Figure 2 the first row and the second row as an example, the stereo images reconstructed by the method proposed in the embodiment of the present invention have clearer edges; taking Figure 2 the third row and the fourth row as an example, the method proposed in the embodiment of the present invention can reconstruct more accurate texture details. The reason is that the method proposed in the embodiment of the present invention adaptively prunes different network layers, and at the same time uses the inter-viewpoint correlation obtained by the teacher network to provide additional constraints for the student network, while significantly reducing the model calculation complexity, maintaining the super-resolution reconstruction performance of stereo images. Figure 3 The technical flow chart of the embodiment of the present invention is given, which mainly includes: an adaptive pruning strategy, a cross-viewpoint distillation mechanism, and a loss function.
[0091] Embodiment 3
[0092] A lightweight-based stereo image super-resolution reconstruction device, the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the following method steps in Embodiment 1:
[0093] Propose an adaptive pruning strategy, construct a teacher network based on the gating unit of importance estimation, prune the residual blocks and the mutual attention mechanism in the corresponding gating units, and obtain an optimized teacher network to construct a student network accordingly;
[0094] A cross-viewpoint distillation mechanism is proposed. By constructing a cross-viewpoint hierarchical attention module, the correlations between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the optimized teacher network, as well as the correlations between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the student network, are explored to obtain correlation matrices respectively.
[0095] The root mean square error loss function is used to measure the difference between the correlation matrices, providing an additional supervision signal for the student network to obtain a lightweight student network.
[0096] In the inference stage, only the lightweight student network is used to reconstruct high-resolution stereo images.
[0097] Among them, the adaptive pruning strategy is used to adaptively prune the residual blocks and mutual attention mechanisms at different levels in the teacher network to obtain a lightweight student network.
[0098] Among them, the gating unit for importance estimation is: the importance of the residual block for extracting single-viewpoint features:
[0099] Taking the importance β of the mutual attention mechanism n as the weight, soft gating is used to perform weighted fusion on the input features and output features of the mutual attention mechanism, which is expressed as follows:
[0100]
[0101] Among them, CrossAttn(·) represents the mutual attention mechanism, f l n+1 and represent the left-viewpoint feature and the right-viewpoint feature output by the nth gating unit respectively;
[0102] Among them,
[0103]
[0104] Among them, ReLU(·) represents the ReLU activation function, Cos(·) represents the cosine similarity calculation, β n ∈[0,1] represents the importance of the mutual attention mechanism in the nth gating unit obtained by calculation, represents the importance of the mutual attention mechanism in the nth gating unit of the left-viewpoint feature; represents the importance of the mutual attention mechanism in the nth gating unit of the right-viewpoint feature.
[0105] Among them, the left-viewpoint feature and the right-viewpoint feature output by the gating unit are:
[0106] Taking the importance α n as the weight, soft gating is used to perform weighted fusion on the input features and output features of the residual block, which is expressed as follows:
[0107]
[0108] Among them, Res(·) represents the residual block, and ⊙ represents element-wise multiplication. and represent the left-viewpoint feature and the right-viewpoint feature output by the residual block, respectively.
[0109]
[0110] Among them, Conv(·) represents the convolutional layer, GP(·) represents the global pooling layer, Sigmoid(·) represents the Sigmoid activation function, Avg(·) represents calculating the average value, and α n ∈[0,1] represents the importance of the residual block in the nth gating unit obtained by the final calculation. represents the importance of the residual block in the nth gating unit of the left-viewpoint feature. represents the importance of the residual block in the nth gating unit of the right-viewpoint feature.
[0111] Among them, pruning the residual block and the mutual attention mechanism in the corresponding gating unit, the optimized teacher network obtained is:
[0112]
[0113] Among them, represents the number of channels of the residual block in the nth gating unit in the teacher network. represents the number of channels of the nth residual block in the pruned student network. × represents the multiplication operation. represents the ceiling operation. represents the flag bit indicating whether to introduce the mutual attention mechanism after the nth residual block in the student network. Net t and Net s represent the teacher network and the student network respectively. Prun(·) represents the process of pruning the network according to the calculated number of channels and the flag bit. N represents the total number of gating units.
[0114] Among them, a cross-viewpoint distillation mechanism is proposed. By constructing a cross-viewpoint hierarchical attention module, the correlation between the left-viewpoint multi-level features and the right-viewpoint multi-level features in the optimized teacher network is explored as:
[0115] For the left and right viewpoint multi-level features obtained by the student network and By constructing a cross-viewpoint hierarchical attention module, the inter-viewpoint attention map M s ;
[0116] Construct the inter-viewpoint correlation distillation constraint, and use the inter-viewpoint attention map M of the teacher network to constrain the inter-viewpoint attention map M of the student network s , improving the accuracy of the inter-viewpoint relationship of each level of features extracted by the student network. The inter-viewpoint correlation distillation constraint is expressed as follows:
[0117]
[0118] Among them, represents the mean squared error loss, and L distillation represents the inter-viewpoint correlation distillation constraint.
[0119] Among them, the inter-viewpoint attention map is:
[0120]
[0121] Among them, M t ∈R HW×N×N represents the inter-viewpoint attention map of the multi-level features of the left and right viewpoints in the teacher network, represents matrix multiplication, · T represents the matrix transpose operation, and R(·) represents the dimension transformation operation, which transforms the dimensions of Q l and Q r from R N×C×H×W to R HW×N×C .
[0122] Among them,
[0123]
[0124]
[0125] Among them, Q l and Q r respectively represent the query vectors of the left and right viewpoints, and R(·) represents the dimension transformation operation, which respectively transforms and into feature vectors with dimensions of R N×C×H×W , C represents the number of channels, H and W respectively represent the width and height of the feature, and proj(·) represents the linear mapping.
[0126] It should be noted here that the device description in the above embodiments corresponds to the method description in the embodiments, and the embodiments of the present invention will not be elaborated here.
[0127] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as computers, single-chip microcomputers, and microcontrollers. Specifically, in implementation, the embodiments of the present invention do not limit the execution subjects, and they are selected according to the needs in actual applications.
[0128] Data signals are transmitted between the memory and the processor through a bus, which is not elaborated in the embodiments of the present invention.
[0129] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored program that controls the device where the storage medium is located to execute the method steps in the above embodiments when the program runs.
[0130] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.
[0131] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the description of the method in the embodiments, and the embodiments of the present invention will not elaborate on this here.
[0132] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.
[0133] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium or a semiconductor medium, etc.
[0134] In the embodiments of the present invention, except for those with special specifications, the models of other devices are not limited, as long as the devices can perform the above functions.
[0135] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0136] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight method for super-resolution reconstruction of stereo images, characterized in that, The method includes: Proposing an adaptive pruning strategy, constructing a teacher network based on a gating unit for importance estimation, pruning the residual blocks and mutual attention mechanisms in the corresponding gating unit, obtaining an optimized teacher network, and thus constructing a student network; Proposing a cross-viewpoint distillation mechanism, by constructing a cross-viewpoint hierarchical attention module, exploring the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the optimized teacher network, and the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the student network, respectively obtaining correlation matrices; Using the root mean square error loss function to measure the difference between the correlation matrices, providing an additional supervision signal for the student network, and obtaining a lightweight student network; Only using the lightweight student network in the inference stage to reconstruct a high-resolution stereo image.
2. A lightweight stereoscopic image super-resolution reconstruction method according to claim 1, characterized in that, The adaptive pruning strategy is used to obtain a lightweight student network by adaptively pruning the residual blocks and mutual attention mechanisms at different levels in the teacher network.
3. A lightweight three-dimensional image super-resolution reconstruction method according to claim 1, characterized in that The gating unit for importance estimation is: the importance of the residual block for extracting single-viewpoint features: Take the importance β of the mutual attention mechanism n as a weight, and use soft gating to perform weighted fusion on the input features and output features of the mutual attention mechanism, which is expressed as follows: Among them, CrossAttn(·) represents the cross-attention mechanism, and f l n+1 and represent the left-viewpoint feature and the right-viewpoint feature output by the m-th gating unit, respectively; Wherein, Among them, ReLU(·) represents the ReLU activation function, Cos(·) represents the cosine similarity calculation, and β n ∈[0,1] represents the importance of the mutual attention mechanism in the nth gating unit obtained by calculation, represents the importance of the mutual attention mechanism in the nth gating unit of the left view feature; represents the importance of the mutual attention mechanism in the nth gating unit of the right view feature.
4. A lightweight three-dimensional image super-resolution reconstruction method according to claim 3, characterized in that The left-viewpoint features and right-viewpoint features output by the gating unit are: Take the importance α n As the weight, use soft gating to perform weighted fusion on the input features and output features of the residual block, as shown below: Among them, Res(·) represents the residual block, and ⊙ represents element-wise multiplication, and respectively represent the left-viewpoint feature and the right-viewpoint feature output by the residual block; Among them, Conv(·) represents the convolutional layer, GP(·) represents the global pooling layer, Sigmoid(·) represents the Sigmoid activation function, Avg(·) represents calculating the average value, and α n ∈[0,1] represents the importance of the residual block in the nth gating unit obtained by the final calculation, represents the importance of the residual block in the nth gating unit of the left-viewpoint feature; represents the importance of the residual block in the nth gating unit of the right-viewpoint feature.
5. A lightweight stereoscopic image super-resolution reconstruction method according to claim 3, characterized in that The pruning of the residual blocks and mutual attention mechanisms in the corresponding gating unit to obtain an optimized teacher network is: Among them, represents the number of channels of the residual block in the nth gating unit in the teacher network, represents the number of channels of the nth residual block in the pruned student network, and × represents the multiplication operation, represents the ceiling operation, represents the flag bit indicating whether to introduce the mutual attention mechanism after the nth residual block in the student network, Net t and Net s represent the teacher network and the student network respectively, Prun(·) represents the process of pruning the network according to the calculated number of channels and flag bits, and N represents the total number of gating units.
6. A lightweight three-dimensional image super-resolution reconstruction method according to claim 1, characterized in that The proposed cross-viewpoint distillation mechanism, by constructing a cross-viewpoint hierarchical attention module, exploring the correlation between the multi-level features of the left viewpoint and the multi-level features of the right viewpoint in the optimized teacher network is: For the left and right view multi-level features obtained by students through the network and By constructing a cross-view hierarchical attention module, an inter-view attention map M is obtained s ; Construct the inter-viewpoint correlation distillation constraint, and utilize the inter-viewpoint attention map \(M\) of the teacher network t to constrain the inter-viewpoint attention map \(M\) of the student network s , and improve the accuracy of the inter-viewpoint relationship of each level of features extracted by the student network. The inter-viewpoint correlation distillation constraint is expressed as follows: Among them, represents the mean squared error loss, L distillation represents the viewpoint - to - viewpoint correlation distillation constraint.
7. A lightweight stereoscopic image super-resolution reconstruction method according to claim 6, characterized in that The inter-viewpoint attention map is: Among them, M t ∈R HW×N×N represents the inter-viewpoint attention map of the multi-level features of the left and right viewpoints in the teacher network, represents matrix multiplication, · T represents the matrix transpose operation, and R(·) represents the dimensionality transformation operation, which transforms the dimensions of Q l and Q r from R N×C×H×W to R HW ×N×C .
8. A lightweight three-dimensional image super-resolution reconstruction method according to claim 7, characterized in that The Among them, Q l and Q r represent the query vectors of the left view point and the right view point respectively. R(·) represents a dimensional transformation operation that transforms and into feature vectors with a dimension of R N×C×H×W respectively. C represents the number of channels, H and W represent the width and height of the feature respectively, and proj(·) represents a linear mapping.
9. A lightweight stereoscopic image super-resolution reconstruction device, characterized in that, The device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in any one of claims 1-8.
Citation Information
Patent Citations
Distillation method and device based on online cooperation and feature fusion
CN116257751A
Face super-resolution reconstruction method and system based on dual generalized distillation
CN116452424A
Automatic inspection method and system for interior of building based on VR and unmanned aerial vehicle
CN119739199A
Automatic channel pruning via graph neural network based hypernetwork
US20230084203A1
Data processing method and apparatus
WO2024245061A1