Image interpolation method based on deep bilateral learning
Through the image interpolation method based on deep bilateral learning, the weight calculation is performed using more anchor pixels and double domain information, the problem of inaccurate estimation when the missing pixels in image interpolation is outlier, and higher interpolation accuracy and quality are achieved.
Patent Information
- Application Number
- CN202510274358.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
Prior Art In image interpolation, when the missing pixel is an 'outlier' relative to the anchor pixel, the estimation result may be inaccurate.
A method of image interpolation based on deep bilateral learning is proposed, using more anchor pixels for estimation, combining the intensity domain and spatial domain information to calculate the weight, and generating the final predicted pixel value through weighting.
It effectively solves the problem of inaccurate estimation when missing pixels are outliers, improves the accuracy and quality of image interpolation, and performs better than existing methods in terms of peak signal-to-noise ratio, structural similarity, feature similarity and edge similarity.
Smart Images

Figure CN120219154A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image interpolation method based on deep bilateral learning. Background Art
[0002] Interpolation is a basic mathematical operation, the purpose of which is to improve the resolution of a signal by generating new samples from a given input signal and then inserting them into the input signal. For example, to interpolate a one-dimensional discrete-time signal by a factor of 2, a new sample needs to be estimated and inserted between each pair of adjacent data points. It should be noted that the original data points of the input signal should remain unchanged after interpolation, because the purpose of interpolation is only to improve the resolution of the signal, rather than to change its original content.
[0003] In the problem of image interpolation, the input image is usually represented as a low-resolution (LR) image, while the output image is usually represented as a high-resolution (HR) image. Existing image interpolation methods can be roughly divided into two categories, namely traditional methods and deep learning methods. Figure 1 Image interpolation is shown for discussion without loss of generality. First, the original pixel intensity values of the input LR image are copied to the corresponding pixel positions (i.e., the colored areas) on the HR image grid, and the values of the missing pixels (i.e., the areas without color) need to be estimated based on the known pixels. To achieve this goal, existing deep learning-based methods, refer to Figure 1 (a), are to predict the pixel intensity difference generated between each missing pixel (marked by a question mark) and its associated anchor pixel (red) from the same interpolation window (delimited by the dashed line). Then, the predicted difference is added to the value of the pixel to obtain the final estimated pixel value. It should be noted that predicting the pixel intensity difference instead of directly estimating the intensity value of the missing pixel is to effectively limit the estimation dynamic range. This will reduce the estimation variance, thereby obtaining higher interpolation accuracy.
[0004] Although the above scheme has achieved remarkable success, when the missing pixel is an "outlier" of the anchor pixel, the estimation result may not be satisfactory. Summary of the Invention
[0005] To solve the problems existing in the prior art, the present invention considers using more anchor pixels surrounding each missing pixel for estimation, refer to Figure 1(b)-(d). In the case of involving multiple anchor pixels, inspired by the basic idea of the classical bilateral filter, the present invention proposes an image interpolation method based on deep bilateral learning, designs a new deep bilateral learning (DBL) method, which not only uses the information from the intensity domain but also uses the information from the spatial domain (i.e., the distance between pixels on the image grid) to calculate weights, and then calculates the weighted sum as the final filtered pixel value; the present invention also integrates DBL into a convolutional neural network to develop a deep bilateral learning network (DBLN) for performing image interpolation. A feature extraction method that can generate better image features is constructed in DBLN, and a module that can execute DBL in parallel is designed to map the features to the pixels to be estimated, thereby generating a high-precision HR image.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides an image interpolation method based on deep bilateral learning for predicting missing pixels, including:
[0008] S1: For each missing pixel, introduce the known pixels around it as anchor pixels;
[0009] S2: In terms of the intensity domain, obtain a set of candidate values by predicting the intensity difference between the missing pixel and each anchor pixel;
[0010] S3: In terms of the spatial domain, generate spatial weights corresponding to the candidate values based on the Euclidean distance and the Gaussian function;
[0011] S4: Estimate the missing pixel by means of weighted sum based on the candidate values and their corresponding spatial weights;
[0012] S5: Use the input LR image and the estimated missing pixels to generate the target HR image through pixel rearrangement operation.
[0013] Optionally, in step S1, p i represents the i-th missing pixel in a certain interpolation window on the HR image grid, where i = 1,..., k 2 -1, and k is the interpolation factor; find a minimum set of anchor pixels that exactly covers the missing pixel to be estimated as the set of anchor pixels around the missing pixel p i represented as:
[0014]
[0015]
[0015] where, e jDenote the j-th anchor pixel, and J is the total number of anchor pixels.
[0016] Optionally, in step S2, the calculation process of the set of candidate values is as follows:
[0017] For each anchor pixel e from the set of anchor pixels estimate the intensity difference Δ(p j , e i ) between it and the missing pixel p i through the following formula: j )
[0018]
[0019] where includes a set of network modules for extracting a set of image features F from the input LR image X, is a non-linear mapping function that maps the obtained feature F to Δ(p i , e j );
[0020] Generate candidate values through the following steps
[0021]
[0022] and then obtain a set of candidate values corresponding to the missing pixels pi J is the total number of anchor pixels.
[0023] Optionally, in step S3, the calculation formula of the spatial weight is as follows:
[0024]
[0025] where f ω (·) represents the calculation function of the spatial weight, d(p i , e j ) represents the Euclidean distance between the missing pixel p i and the anchor pixel e j , represents the set of anchor pixels, and σ is the standard deviation of the learnable Gaussian kernel function.
[0026] Optionally, in step S4, the calculation formula of the missing pixel is as follows:
[0027]
[0028] where ω i,j represents the spatial weight corresponding to the candidate value , and e j represents the missing pixel p iThe anchor pixels, represent the set of anchor pixels.
[0029] Optionally, the process of extracting a set of image features from the input LR image X is as follows:
[0030] Use a convolutional layer to extract a shallow image feature set F0 from the input LR image X;
[0031] Then, a series of cascaded network module groups are used to generate deeper image feature sets, where the input and output feature sets of the nth network module group are respectively denoted as F n-1 and F n , n = 1, 2,..., N, and N represents the total number of groups;
[0032] In the nth group of network modules, first use the asynchronous multi-scale module to extract multi-scale information from F n-1 , then execute the cross attention module to enhance the features using spatial similarity and output the enhanced feature set F n ; after executing all network module groups, use the adaptive feature fusion module to fuse the obtained feature sets F0, F1,..., F N to generate F as the output.
[0033] Optionally, the asynchronous multi-scale module captures multi-scale information corresponding to different window sizes in an asynchronous manner and consists of two rounds of asynchronous multi-scale feature extraction operations. The first round is implemented as follows:
[0034]
[0035] S 12 = σ r (Conv 3×3 (S 11 ))
[0036] where Conv 3×3 (·) is a 3×3 convolutional layer, and σ r (·) is the ReLU function. The input and output of the mth asynchronous multi-scale module in the nth group are respectively denoted as and where m = 1, 2,..., M, and M is the number of asynchronous multi-scale modules included in the network module group;
[0037] The obtained S 11 and S 12 After concatenation, a second round of asynchronous multi-scale feature extraction operation is performed to generate S 21 and S 22 ;
[0038] Use a 1×1 convolutional layer to fuse S21 and S 22 , and add skip connections to generate
[0039] Optionally, the cross - attention module includes two rounds of attention operations, each round only focusing on the similarity in the cross - region of each position. For the position (s, t) on, its cross - region Ω(s, t) is defined as follows:
[0040] Ω(s, t) = {(s, 1),..., (s, W)} U {(1, t),..., (H, t)},
[0041] where s = 1, 2,..., H, t = 1, 2,..., W, and H and W represent the height and width of the image respectively;
[0042] The first - round attention operation is expressed as:
[0043]
[0044] where and are obtained by reducing the number of channels of using three 1×1 convolutional layers respectively. C represents the number of channels of the feature maps in F, and f s is a similarity metric function, and the formula is as follows:
[0045]
[0046] Input the obtained in the first round into the second - round attention operation to establish global connections between any two positions on the grid.
[0047] Optionally, the adaptive feature fusion module automatically identifies the importance of each feature map in the input feature set, and completes feature fusion after weighting the feature maps according to the obtained importance. Specifically:
[0048] is obtained by concatenating F0, F1,..., F N . H and W represent the height and width of the image respectively, and C represents the number of channels of the feature maps in F. The following formula is used to assign corresponding importance scores to the (N + 1)C feature maps in U:
[0049]
[0050] where Pool avg (·) represents the average pooling layer, Full(·) represents the fully - connected layer, and σ rDenote the ReLU function, importance score;
[0051] Feature fusion is achieved through the following formula:
[0052]
[0053] Wherein, represents element-wise multiplication.
[0054] Optionally, classify the missing pixels, each category adopts an independent branch, and perform deep bilateral learning in parallel; wherein, for each category of pixels, obtain the anchor pixels by using the observation window with different offsets compared to the original window, and use the neural network model to predict the intensity difference corresponding to these anchor pixels at one time, and then obtain the candidate values of this category of pixels through matrix addition; finally, perform weighting on the candidate values through matrix multiplication to generate the final predicted value of this category of pixels.
[0055] In a second aspect, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the image interpolation method based on deep bilateral learning as described in the first aspect.
[0056] The beneficial effects of the present invention are: Inspired by the basic idea of the classical bilateral filter, the present invention proposes a new DBL method, which not only uses the information from the intensity domain, but also uses the information from the spatial domain (i.e., the distance between pixels on the image grid) to calculate the weights, and then calculates the weighted sum as the final predicted pixel value. The learnable weights adopted have significant advantages in providing better image interpolation performance. Based on DBL, the present invention also proposes a DBLN, which includes a new feature extraction structure to generate better image features and promote DBL to produce better prediction results. The present invention effectively solves the problem that inaccurate estimation will occur when the missing pixels are "outliers" relative to the anchor pixels, and is superior to the existing image interpolation methods in terms of performance such as peak signal-to-noise ratio, structural similarity, feature similarity, and edge similarity, and has excellent performance in edge reconstruction and other aspects, and can generate high-precision interpolated HR images. Description of the Drawings
[0057] Figure 1 Anchor pixels involved for each missing pixel (marked with a question mark) in the case of image interpolation with a magnification factor of 2×2: (a) is the existing deep learning-based method; (b)-(d) are the deep bilateral learning (DBL) methods proposed by the present invention.
[0058] Figure 2 For estimating the missing pixel p at different positions i the required anchor pixel ej : (a) is for the 2×2 case; (b) is for the 3×3 case.
[0059] Figure 3 shows the visualization results of weights generated by the Gaussian weighting function for different σ values. Here, only the positions of three missing pixels (i.e., p1, p3, and p8, all represented by red dots) in the 3×3 interpolation case shown in Figure 2 are discussed. The influence degree of each anchor pixel is represented by a blue line segment. The darker the line segment color, the greater the influence.
[0060] Figure 4 is a diagram of the deep bilateral learning network (DBLN) developed for image interpolation in the present invention: (a) demonstrates the overall network structure including two sequentially performed stages (including feature extraction and feature-to-image mapping); (b), (c), and (d) further describe three key modules, namely the asynchronous multi-scale block (AMB), the cross cross attention block (CCAB), and the adaptive feature fusion block (AFFB) used in the feature extraction stage; the symbols and represent element addition and multiplication operations respectively, and the symbol is the weighted summation operation.
[0061] Figure 5 demonstrates the parallel execution of the proposed DBL in the feature-to-image mapping stage of the DBLN for 2×2 image interpolation. The symbol is element addition.
[0062] Figure 6 demonstrates the pixel collection operation designed in the present invention: (a) shows the pixels constituting the (m, n)th interpolation window, and (b) describes how to parallelly extract all the of the interpolation windows to construct Φ2.
[0063] Figure 7 is a set of comparison diagrams of 2×2 enhancement results generated by different image interpolation methods.
[0064] Figure 8 is a set of comparison diagrams of 2×2 enhancement results generated by different image interpolation methods.
[0065] Figure 9 is a set of comparison diagrams of 2×2 enhancement results generated by different deep learning image interpolation methods.
[0066] Figure 10 is a set of comparison diagrams of 3×3 enhancement results generated by different interpolation methods. Detailed implementation manner
[0067] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0068] In one embodiment, the present invention proposes an image interpolation method based on deep bilateral learning, which mainly includes the following parts.
[0069] 1. A new feature-to-pixel mapping method for image interpolation, aiming to achieve missing pixel prediction.
[0070] In this embodiment, based on the basic principle of the existing bilateral filter, learning is carried out in two aspects: the intensity domain and the spatial domain. For each missing pixel, the surrounding known pixels are introduced as anchor points. Subsequently, in terms of the intensity domain, the features extracted in the DBLN are used to predict the intensity difference between the missing pixel and each anchor pixel, obtaining a set of candidate values. At the same time, in terms of the spatial domain, based on the Euclidean distance and the Gaussian function, spatial weights are generated to achieve the weighted effect of near-large and far-small. Specifically, the variance in the used Gaussian function is learnable. Finally, based on the predicted candidate values and the corresponding spatial weights, the value of the missing pixel is generated by means of weighted summation. The specific operation is as shown below.
[0071] A. Learning to predict intensity differences
[0072] As Figure 2 shown, let p i represent the i-th missing pixel in an interpolation window on the HR image grid, where i = 1,..., k 2 -1, and k is the interpolation (or magnification) factor. In this embodiment, k = 2 and k = 3 are taken as examples. In these two cases, there will be 3 and 8 missing pixels to be estimated in the interpolation window respectively. The set of anchor pixels (i.e., the local support domain) around the missing pixel p i can be further expressed as:
[0073]
[0074] where, e j represents the j-th anchor pixel, and J is the total number of anchor pixels. In the selection of the support domain, the strategy adopted in this embodiment is to find the smallest set of anchor pixels that exactly covers the missing pixel to be estimated. Based on this strategy, the actual support domain is as Figure 2As shown. Specifically, for the missing pixels located on the upper and left edges of the interpolation window, its support domain consists of 6 anchor pixels (i.e., J = 6); while for other regions, the corresponding support domain only contains 4 pixels (i.e., J = 4).
[0075] For each anchor pixel e from the set j , it is necessary to estimate the intensity difference Δ(p i , e i ) between it and the missing pixel p j . With the help of deep learning technology, this can be achieved through the following two steps: )
[0076]
[0077] Among them, includes a set of network modules for extracting a set of image features F from the input LR image X, is a non-linear mapping function that maps the obtained feature F to Δ(p i , e j ). Then, candidate values can be generated through the following steps:
[0078]
[0079] For this embodiment, when multiple anchor pixels are involved, the image features extracted by (2) can be shared, while operations (3) and (4) need to be repeated for each anchor pixel in . That is to say, by substituting into (3) and (4), a set of candidate values for the missing pixel p i can be obtained, that is
[0080] B. Learning to generate spatial weights
[0081] Based on the candidate values generated in the previous stage, this stage will generate a set of spatial weights, denoted as to weight these candidate values to generate the final estimated value of p i . Following the idea of the bilateral filter, a Gaussian function is also used here for weight calculation. Specifically, based on the distance d(p i , e j ) (for j = 1, 2,..., J) between the missing pixel p i and the anchor pixel e j , the adopted weight calculation formula is expanded as follows:
[0082]
[0083] Where σ is the standard deviation of the Gaussian kernel function. Note that the main innovation here is that the standard deviation σ of the Gaussian kernel function in the original bilateral filter is fixed, while the standard deviation σ of the Gaussian function used in DBL is learnable, thereby generating a set of learnable spatial weighting parameters.
[0084] Figure 3 The visualization results of the weights obtained when σ takes different values when the interpolation coefficient is 3×3 are shown in . It can be seen that when σ takes a small value (such as σ=0.5), the weights assigned to those anchor pixels that are farther away are almost zero, resulting in the estimation of missing pixels relying only on the nearest anchor pixels. On the contrary, when σ is larger (for example, σ=4.0), more uniform weights will be produced, resulting in an approximate average effect. When σ=2.0, each anchor pixel is covered and the obtained weights are in a relationship of large near and small far with their distance to the missing pixel. Obviously, the ideal optimal value of σ is close to 2.0. In this regard, this embodiment will find the optimal value by adjusting it during the training process. This adjustment process needs to be carried out simultaneously with the intensity prediction learning in the previous part so that the two can work better together in DBL.
[0085] Using the learned σ, (1) and (5) can be applied to Figure 2 Multiple sets of weights are generated for different missing pixel locations on the HR grid shown. Subsequently, based on the candidate pixel values generated by (4) and the weights generated by (5), the missing pixel p can be estimated by calculating the following weighted sum i :
[0086]
[0087] Note that (1) to (6) above are only the steps that need to be performed in the local interpolation window. To interpolate the entire image, this series of operations should be repeated for all interpolation windows on the HR grid. Since different interpolation windows are completely independent, these operations can be implemented in parallel and thus accelerated by the GPU. The following sections will provide more details on how to integrate DBL into the deep neural network DBLN for image interpolation in a GPU-friendly way.
[0088] Based on the image feature set F generated in the feature extraction stage, this stage will estimate all missing pixels on the HR image grid by using the proposed DBL, i.e., using (3), (4), (5) and (6). As emphasized above, a significant advantage of DBL is that it can be applied to each interpolation window on the HR grid separately, so the rate can be increased through parallel computing. In this regard, taking the case where the interpolation coefficient is 2×2 as an example, the specific approach in this embodiment is as follows: Figure 5 Note that a similar processing scheme can also be used when the interpolation coefficient is 3×3.
[0089] According to Figure 2 For the 2×2 interpolation case shown in (a), to complete the interpolation over the entire HR grid, the three cases demonstrated need to be repeatedly applied to all interpolation windows. In other words, all the missing pixels on the HR grid can be classified into three categories, where the missing pixels in each group share the same pattern. To this end, this embodiment establishes three parallel branches, each dedicated to separately processing a specific group, as Figure 5 shown. The following will take the second branch as an example for illustration.
[0090] The purpose of the second branch is to generate the matrix Each element P2(m, n) represents the estimated value of the missing pixel p2 in the (m, n)-th interpolation window on the HR grid, where 1 ≤ m ≤ H and 1 ≤ n ≤ W. To achieve this goal, the first step is to collect 6 anchor pixels for each missing pixel p2 in every interpolation window on the HR image grid according to (1), so as to form the anchor pixel matrix According to Figure 6 (a), it can be known that
[0091]
[0092] Although Φ2 can be obtained by enumerating all H×W cases corresponding to (7), the efficiency is extremely low. To solve this problem, this embodiment proposes a more efficient pixel aggregation operation, that is, to quickly obtain the required pixels by moving the observation window on X with different offsets, as Figure 6 (b) shown.
[0093] Figure 6 In (b), multiple observation windows are marked with different color dashed boxes, and the corresponding vertical and horizontal offsets are shown below each case. Specifically, each offset observation window will generate one channel in Φ2. For example, for m = 1, 2,..., H and n = 1, 2,..., W, the first channel of the Φ2 set composed of X(m, n - 1) can be obtained by moving the observation window downward by (0, -1) units, see Figure 6 the yellow dashed box in the middle of (b), and the obtained channel is highlighted in the same color on the right. Note that when the moving observation window exceeds the image boundary, the unobservable parts need to be filled. Similarly, the remaining channels of Φ2 can be observed with offsets of (0, 0), (0, 1), (1, -1), (1, 0), and (1, 1) respectively. Finally, all the obtained channels form the matrix Φ2, see Figure 6 the right side of (b).
[0094] Based on the generated Φ2, two 3×3 convolutional layers are then applied to perform the intensity difference estimation operation given in (3) in parallel, as follows:
[0095] D2 = Conv 3×3 (σ r (Conv 3×3 (F))), (8)
[0096] Among them, the first layer compresses the C channels in, into channels, and the second layer maps the obtained channels to generate Note that in the design here, the output of the first layer will be shared with the other two branches to reduce the overall complexity of the model. The obtained D2(m, n) contains the intensity difference between the missing pixel p2 in the (m, n)th interpolation window and the corresponding anchor pixel in Φ2(m, n). Subsequently, by adding Φ2 and D2 pixel by pixel, the matrix is obtained as follows:
[0097]
[0098] Among them, contains multiple estimation candidates for the missing pixel p2 in the (m, n)th interpolation window on the HR grid. Then, the final predicted value can be generated by calculating the weighted sum of all channels in 2,1 , ω 2,2 ,..., ω 2,6} generated by (5). This step can be achieved through a 1×1 convolutional operation with the parameter set to ω2, that is
[0099]
[0100] By performing similar operations from (7) to (10) in the other two branches, two sets of estimated missing pixels (i.e., P1 and P3) will be generated.
[0101] Finally, using the input LR image X and the obtained sets of missing pixels P1, P2, and P3, the target high-resolution image Y can be generated in the following way:
[0102]
[0103] Among them, represents the pixel rearrangement operation.
[0104] 2. A new neural network feature extraction method that combines multi-scale and spatial and channel attention mechanisms.
[0105] In this embodiment, multi-scale and cross-scale information in an image is captured by using cascaded multi-scale feature extraction, and then spatial attention is used to enhance the features, and the two constitute a mode of feature extraction. By continuously using the above mode, image feature sets at different levels are obtained. Then, channel attention is used to perceive the importance of different feature maps in the obtained set, and the features are weighted based on the importance, and then feature fusion is completed.
[0106] As Figure 4 (a) shows, in this embodiment, a 3×3 convolutional layer is first used to extract a shallow image feature set F0 from the input image X. Then, a series of cascaded network module groups are adopted to generate deeper image feature sets. The input and output feature sets of the nth group are respectively denoted as F n-1 and F n , where n = 1, 2,..., N, and N is the total number of groups. In the nth group, first, a sequence of asynchronous multi-scale blocks (AMB) is used to extract multi-scale information from F n-1 , and then a cross-criss attention block (CCAB) is executed to enhance the features by using spatial similarity and output the enhanced feature set F n . Compared with only using AMB for feature extraction, this combination of multiple AMBs and one CCAB is more effective in generating information-rich image features. After all module groups are executed, an adaptive feature fusion block (AFFB) is used to fuse the obtained image feature sets (i.e., F0, F1,..., F N ) to generate F as the output of this stage. The following part provides the detailed information of three network modules.
[0107] (1) Asynchronous multi-scale block (AMB):
[0108] AMB is an efficient network module designed to capture multi-scale information corresponding to different window sizes in an asynchronous manner (for example, the 3×3 window captures low-scale information, and the 5×5 window captures high-scale information). The "asynchronous" here means that a small-size convolution is first executed to achieve low scale, and then another small-size convolution is executed on this basis to achieve high scale. Compared with the "synchronous" version that executes 3×3 and 5×5 convolutions in parallel, this asynchronous method requires much fewer parameters.
[0109] As Figure 4 (b) shows, the input and output of the mth AMB in the nth group are respectively denoted as and where \(m = 1, 2, \cdots, M\), and \(M\) is the number of AMB modules included in the feature extraction module group. AMB consists of two rounds of asynchronous multi-scale feature extraction operations. The first round is implemented as follows:
[0110]
[0111] S 12 =\(\sigma\) r (Conv 3×3 (S 11 )),(13)
[0112] where Conv 3×3 (·) is a 3×3 convolutional layer, and \(\sigma\) r (·) is the ReLU function. Then, the obtained \(S\) 11 and \(S\) 12 will be input into the second-round multi-scale operation similar to (12) and (13) after concatenation to generate \(S\) 21 and \(S\) 22 . Note that since \(S\) 11 and \(S\) 12 are extracted from different scales, the second round can actually be equivalent to capturing cross-scale information. Finally, a 1×1 convolutional layer is used to fuse \(S\) 21 and \(S\) 22 , and skip connections are added to generate
[0113] 2) Cross-Cross Attention Block (CCAB):
[0114] Images captured from real-world scenarios naturally contain many self-similar structures, which inspires this embodiment to utilize spatial similarity to enhance the features in each feature extraction group. To avoid introducing a large computational burden, this embodiment introduces CCAB into the design of DBLN because this module has high execution efficiency.
[0115] As shown in Figure 4 (c), to make full use of the similarity in the entire spatial domain of the input image features, CCAB contains two rounds of attention operations, and each round only focuses on the similarity in the cross-cross area of each position. Specifically, for the position \((s, t)\) on , its cross-cross area is defined as follows:
[0116] \(\Omega(s, t)=\{(s, 1), \cdots, (s, W)\} \cup \{(1, t), \cdots, (H, t)\}\), (14)
[0117] where \(s = 1, 2, \cdots, H\), \(t = 1, 2, \cdots, W\). Based on this, the first round of attention operation can be expressed as:
[0118]
[0119] Among them, and are respectively obtained by reducing the number of channels of using three 1×1 convolutional layers. f s is a similarity metric function, and the formula is as follows:
[0120]
[0121] Among them, the superscript represents matrix transpose. By further inputting the obtained into another round of attention operations similar to (14) and (15), the global connection between any two positions on the grid can be effectively established, so as to realize feature enhancement based on spatial similarity in the generated F . n
[0122] 3) Attention-Aware Feature Fusion Block (AFFB):
[0123] The AFFB can automatically identify the importance of each feature map in the input feature set, and then complete feature fusion after weighting the feature maps according to the obtained importance degree. Based on this advantage, in this embodiment, it is selected to be used for fusing the feature sets generated by different feature extraction module groups. The specific structure is as shown in Figure 4 (d). Let be obtained by concatenating F0, F1,..., F N . In the AFFB, first, the corresponding importance scores are assigned to the (N + 1)C feature maps in U through the following two steps:
[0124]
[0125] In (17), Pool avg (·) represents the average pooling layer, which is used to calculate the average value of each feature map, and the obtained average value is stored in . In (18), Full(·) represents a fully connected layer. By using two fully connected layers and a ReLU function σ r , an efficient non-linear mapping function can be established. Taking the mean as the input, a set of importance scores is output, that is, More specifically, the two fully connected layers used will change the number of channels of the features. Among them, the first one reduces the number of channels from (N + 1)C in The second restores this quantity to the original (N+1)C in the generated W. Using the importance scores obtained in W, feature fusion is achieved as follows:
[0126]
[0127] where the symbol represents element-wise multiplication.
[0128] 3. A new network training method.
[0129] To promote the proposed DBLN to achieve good training results, this embodiment proposes a new network training method, including how to initialize the parameters and how to update them. Specifically, through theoretical analysis, the initialization is selected as σ = 2.0, and the visualization results affected by this parameter can be seen in Figure 3 At the same time, other parameters in the network are randomly initialized. Subsequently, two-stage training is adopted: in the first stage, σ is fixed and other parameters are optimized, aiming to enable the network to obtain stable feature extraction ability; in the second stage, the restriction on σ is released and all parameters are optimized synchronously to learn the optimal value of σ.
[0130] Next, the beneficial effects of this embodiment are evaluated in combination with examples.
[0131] (1) Objective evaluation: Four image quality assessment metrics are used for objective evaluation, including peak signal-to-noise ratio (PSNR), structural similarity (SSIM), feature similarity (FSIM), and edge similarity (ESIM). Given a color image, these four metrics are calculated on its luminance component.
[0132] Referring to Table 1, by comparing the DBLN of this embodiment with eleven available interpolation methods including Bicubic, NEDI, SAI, RSAI, AGSI, CGI, FCGI, PCI, NARM, FIRF, and AIN, the evaluation is first carried out in the case of 2×2 interpolation. It can be seen from Table 1 that the DBLN of this embodiment obtains the highest scores on all used evaluation metrics and all test datasets. Specifically, compared with ten traditional methods, the DBLN proposed in this embodiment shows PSNR gains of 3.41dB, 3.53dB, 2.68dB, 2.55dB, 2.83dB, 2.78dB, 2.74dB, 2.59dB, 2.31dB, and 1.77dB, and is 0.26dB better than the performance of the deep learning-based method AIN. When using other evaluation metrics (i.e., SSIM, FSIM, and ESIM) to conduct performance evaluation alone, significant improvements can also be observed.
[0133] Table 1 compares the performance of DBLN in this embodiment with eleven available interpolation methods for codes
[0134]
[0135] (2) Subjective evaluation: First, a subjective evaluation is conducted by comparing the existing available code methods with DBLN in this embodiment, and the interpolation factor is set to 2×2. The test images are the gray image "Cameraman" obtained from Set15 and the color image "img061" obtained from Urban100. The results of the two test images are respectively as Figure 7 and Figure 8 shown. To process the RGB image "img06l", in this embodiment, different image interpolation methods are only applied to the "Y" channel of each test image in its YCbCr color space, and the Cb and Cr channels are simply processed by Bicubic.
[0136] From Figure 7 and Figure 8 it can be seen that although the conventional methods (i.e., the first two columns of each group of images) have made a lot of efforts in using edge and contrast information to guide the interpolation process, they still cannot recover the correct edges in these two images, and serious blurring and edges with wrong directions can be easily observed. Two traditional learning-based methods (i.e., NARM and FIRF) also cannot recover details and produce performance similar to other conventional methods. The deep learning-based AIN can recover some details, but it still produces heavy artifacts on the straight lines of the building. In contrast, DBLN in this embodiment provides the best subjective effect with less artifacts.
[0137] To further demonstrate the superiority of DBLN in this embodiment among deep learning image interpolation methods, it is compared with VLN and AIN on the test image "082" from the dataset Urban12 with an interpolation factor of 2×2, and the comparison results are as Figure 9 shown. It can be seen that VLN faces difficulties in reconstructing the edges in the test image and produces obvious wavy artifacts.
[0138] AIN achieves better performance and corrects the wrong reconstruction of the left window edge part; however, there are still many artifacts on the right window edge. In contrast, DBLN in this embodiment can correctly reconstruct all the edges.
[0139] In addition, the 3×3 interpolation results of three common interpolation methods (one is Bicubic, one is AGSI, and one is NARM), and the 3×3 interpolation results of two deep learning methods (i.e., AIN and DBLN proposed in this embodiment) are in Figure 10For comparison, the test image "img 073" is taken from the Urban 100 dataset. After careful inspection, it is obvious that the DBLN of this embodiment is superior to other methods, providing significantly superior results, especially in obtaining sharper edges.
[0140] In summary, all of the above comparisons clearly demonstrate the effectiveness of the DBLN of this embodiment in enhancing the visual quality of interpolated images, showing the superiority of the DBLN of this embodiment in subjective evaluation.
[0141] In another embodiment, the present invention proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the image interpolation method based on deep bilateral learning of the foregoing embodiment is implemented.
[0142] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0143] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A deep bilateral learning based image interpolation method for predicting missing pixels, characterized in that: include: S1: For each missing pixel, introduce the known pixels around it as anchor pixels; S2: In the intensity domain, a set of candidate values is obtained by predicting the intensity difference between the missing pixel and each anchor pixel; S3: In the spatial domain, the spatial weights corresponding to the candidate values are generated based on the Euclidean distance and Gaussian function; S4: Estimate missing pixels by weighted sum based on candidate values and their corresponding spatial weights; S5: Generate the target HR image through a pixel rearrangement operation using the input LR image and the estimated missing pixels.
2. The image interpolation method based on deep bilateral learning according to claim 1, characterized in that: In step S1, p i represents the i-th missing pixel in a certain interpolation window on the HR image grid, where i = 1, ..., k 2 -1, k is the interpolation factor; find a minimum anchor pixel set that just covers the missing pixel to be estimated as the missing pixel p i The surrounding anchor pixel set It is expressed as: Among them, e j represents the jth anchor pixel, and J is the total number of anchor pixels.
3. The image interpolation method based on deep bilateral learning according to claim 1, characterized in that: In step S2, the calculation process of the set of candidate values is as follows: For the set of anchor pixels Each anchor pixel e j , and its relationship with the missing pixel p is estimated by the following formula i The intensity difference Δ(p i , e j ): in, It includes a set of network modules for extracting a set of image features F from the input LR image X. It maps the obtained feature F to Δ(p i , e j )’s nonlinear mapping function; Generate candidate values by following the steps below Then obtain a set of corresponding missing pixels p i Candidate value of J is the total number of anchor pixels.
4. The image interpolation method based on deep bilateral learning according to claim 1, characterized in that: In step S3, the calculation formula of the spatial weight is as follows: Among them, f ω (·) represents the calculation function of spatial weight, d(p i , e j ) indicates missing pixel p i and anchor pixel e j The Euclidean distance between represents the set of anchor pixels, and σ is the standard deviation of the learnable Gaussian kernel function.
5. The image interpolation method based on deep bilateral learning according to claim 1, characterized in that: In step S4, the missing pixel calculation formula is as follows: Among them, ω i,j Representation and candidate values The corresponding spatial weight, e j Indicates missing pixel p i The anchor pixel, Represents a set of anchor pixels.
6. The image interpolation method based on deep bilateral learning according to claim 3, characterized in that: The process of extracting a set of image features F from the input LR image X is as follows: Use the convolutional layer to extract the shallow image feature set F0 from the input LR image X; Then, a series of cascaded network module groups are used to generate a deeper image feature set, where the input and output feature sets of the nth network module group are represented as F n-1 and F n , n=1,2,...,N, N represents the total number of groups; In the nth group of network modules, the asynchronous multi-scale module is first used to extract n-1 Then, the cross attention module is executed to enhance the features by using the spatial similarity and output the enhanced feature set F. n ; After executing all network module groups, the adaptive feature fusion module is used to obtain the feature set F0, F1, ..., F N The fusion is performed to generate F as output.
7. The image interpolation method based on deep bilateral learning according to claim 6, characterized in that: The asynchronous multi-scale module captures multi-scale information corresponding to different window sizes in an asynchronous manner and consists of two rounds of asynchronous multi-scale feature extraction operations. The first round is implemented as follows: S 12 =σ r (Conv 3×3 (S 11 )), Among them, Conv 3×3 (·) is a 3×3 convolutional layer, σ r (·) is the ReLU function, and the input and output of the mth asynchronous multi-scale module in the nth group are expressed as and Where m = 1, 2, ..., M, M is the number of asynchronous multi-scale modules included in the network module group; The S 11 and S 12 After concatenation, a second round of asynchronous multi-scale feature extraction is performed to produce S 21 and S 22 ; Use a 1×1 convolutional layer to fuse S 21 and S 22 , and add skip connections to produce 8. The image interpolation method based on deep bilateral learning according to claim 7, characterized in that: The cross attention module includes two rounds of attention operations, each round only focuses on the similarity in the cross region of each position. The position (s, t) on the , its cross-section area Ω(s, t) is defined as follows: Ω(s,t)={(s,1),...,(s,W)}∪{(1,t),...,(H,t)}, Wherein, s = 1, 2, ..., H, t = 1, 2, ..., W, H and W represent the height and width of the image respectively; The first round of attention operation is expressed as: in, and By using three 1×1 convolutional layers The channel dimensionality reduction is performed, C represents the number of channels of the feature map in F, f s is the similarity measurement function, and the formula is as follows: The first round of Input into the second round of attention operation to establish Global connectivity between any two locations on the grid.
9. The image interpolation method based on deep bilateral learning according to claim 6, characterized in that: The adaptive feature fusion module automatically identifies the importance of each feature map in the input feature set, and completes feature fusion after weighting the feature maps according to the obtained importance, specifically: From F0, F1, ..., F N The concatenation results in that H and W represent the height and width of the image, respectively, and C represents the number of channels of the feature map in F. The following formula is used to assign corresponding importance scores to the (N+1)C feature maps in U: Among them, Pool avg (·) represents the average pooling layer, Full(·) represents the fully connected layer, σ r ReLU function, importance score Feature fusion is achieved through the following formula: in, Stands for element-wise multiplication.
10. The image interpolation method based on deep bilateral learning according to claim 1, characterized in that: The missing pixels are classified, and each category uses an independent branch to perform deep bilateral learning in parallel. For each type of pixel, an observation window with a different offset from the original window is used to obtain anchor pixels, and a neural network model is used to predict the intensity difference corresponding to these anchor pixels at one time. Then, the candidate values of this type of pixels are obtained through matrix addition. Finally, the candidate values are weighted through matrix multiplication to generate the final predicted value of this type of pixel.
Citation Information
Cited By
Copper plate defect detection and inspection method, system, equipment and medium
CN120931644A