Image demosaicing method
By optimizing the neural network structure and combining lookup tables, the problem of excessive parameter amount and calculation amount in edge devices is solved, and efficient image demosaic on devices with high computing and storage requirements is achieved.
Patent Information
- Application Number
- CN202510512305.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
The existing deep learning demosaic algorithm is difficult to apply to edge devices with high computing and storage requirements due to the huge amount of parameters and calculations. The lookup table-based method is difficult to apply to image demosaics effectively due to the limited receptive field size and large amount of parameters.
By optimizing the neural network structure, combining deep learning and lookup tables, an image demosaic method is designed, including training neural networks, transforming lookup tables and micro-survey tables, increasing the receptive field size and reducing the number of lookup table parameters and calculation complexity.
Improves the visual effect of demosaics, while reducing the amount of parameters and interpolation calculations of the lookup table, making it suitable for devices with high computing and storage requirements.
Smart Images

Figure CN120387954A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image demosaicing method. Background Art
[0002] By adding a color filter array on the basis of a charge-coupled device or a CMOS image sensor, the color information of an object is obtained. The most common arrangement of the color filter array is the Bayer pattern. In the RAW image of the Bayer pattern obtained by the Bayer color filter array, only the information of one color component among R, G, and B can be obtained at each pixel point, and the G component is twice that of the R component and the B component. In image signal processing, the color components of each pixel point can be complemented through a demosaicing algorithm, and the full-color RGB image corresponding to the object can be obtained through channel merging.
[0003] Deep learning demosaicing algorithms are difficult to apply to edge devices due to their huge number of parameters and computational complexity, but their superior performance is expected to become the subsequent algorithm direction. How to lightweight deep learning is one of the key tasks. At present, to lightweight the deep learning-based demosaicing method, it can be achieved by directly optimizing the neural network structure or by using a look-up table. In the method of directly optimizing the neural network structure, the number of parameters and computational complexity are significantly reduced, but it still cannot be applied to devices with high requirements for computing and storage. The research on the method based on the look-up table is less, and it is difficult to apply due to the limited receptive field size and the large number of parameters of the look-up table. Summary of the Invention
[0004] The purpose of the present invention is to provide an image demosaicing method, which is a lightweight method combining deep learning and a look-up table for demosaicing. By optimizing the design of the neural network structure, a method is provided to increase the receptive field size and optimize the design of the neural network structure, thereby improving the visual effect of demosaicing. At the same time, the method reduces the number of parameters of the look-up table and the computational complexity, reduces the number of parameters of the look-up table, and reduces the computational amount during the interpolation of the look-up table.
[0005] The present invention provides an image demosaicing method, including:
[0006] S1. Convert the RGB three-channel image into a Bayer image; use the Bayer image as the input and the RGB three-channel image as the label to train the demosaicing neural network model; the structure of the demosaicing neural network model includes a first stage to a fourth stage;
[0007] S2. After the demosaicing neural network model converges in training, convert the mapping relationship between the Bayer image and the RGB three-channel image into a look-up table;
[0008] S3. Fine-tune the look-up table;
[0009] S4. Apply the fine-tuned look-up table to the Bayer format image to obtain the demosaicked RGB full-color image.
[0010] Furthermore, in step S1, it specifically includes: the resolution of the RGB three-channel image is H*W*3, where H is the number of rows of the image, W is the number of columns of the image, and 3 is the number of channels; taking the top-left corner of the RGB three-channel image as the original coordinate, from left to right, top to bottom, with a step size of 2, extract image blocks of size 2*2*3.
[0011] Take the B value at the top-left coordinate, the G value at the top-right coordinate, the G value at the bottom-left coordinate, and the R value at the bottom-right coordinate in the image block to form an image block of size 2*2*1. The entire RGB three-channel image is processed in the same way. A Bayer image of H*W*1 is obtained from an RGB three-channel image of H*W*3; establish the data set of the input and the label.
[0012] Furthermore, in step S1, the first stage S11 is divided into 2 layers, and each layer uses multiple linear layers; perform edge padding on the Bayer image. Starting from the top-left corner of the edge-padded Bayer image, from left to right, top to bottom, extract image blocks of 4*4*1 at a step size of 2, that is, 4 pairs of BGGR; recombine the 4 pairs of BGGR, and connect them row by row in order. Recombine from a shape of 4*4 to a pixel value of 1*16; process the 16 pixel values one by one with the 16 linear layers in the first layer. The number of input and output channels of each linear layer in the first layer is 1 and N respectively, where N is an adjustable parameter. After processing, a tensor of 16*N is obtained; then use the 16 linear layers in the second layer to process the 16*N tensor one by one. The number of input and output channels of each linear layer in the second layer is N and 12 respectively, where N is an adjustable parameter. After calculating the mean value and the tanh activation function, a tensor of 1*12 is obtained. Reshape the 1*12 tensor to obtain a tensor of 2*2*3 as the output; a Bayer image of H*W*1 is processed by the first stage to obtain a tensor of H*W*3.
[0013] Furthermore, in step S1, the second stage S12 is divided into 6 layers, and each layer uses a convolutional layer; the receptive field of the first convolutional layer is 1*3, and the number of input and output channels is 1 and N respectively, where N is an adjustable parameter; the receptive fields of the second to fifth convolutional layers are 1*1, and the number of input and output channels is N and N respectively. The receptive field of the sixth convolutional layer is 1*1, and the number of input and output channels is N and 1 respectively; there is a ReLU activation function after the first to fifth convolutional layers, and there is a tanh activation function after the sixth convolutional layer.
[0014] The calculation process of the second stage includes: first, perform edge padding on the H*W*3 tensor obtained after the first stage processing, and then the 6 convolutional layers process the input tensor in sequence; the above operations process one of the 3 channels of the tensor, and the same processing is performed on the other channels using the same method; a tensor of H*W*3 is still a tensor of H*W*3 after the above processing.
[0015] Further, in step S1, the third stage S13 is divided into 2 parallel groups of structures, each group of structures is divided into 2 layers, and each layer uses 25 linear layers; perform edge padding on the H*W*3 tensor obtained after the second stage processing, and input the edge-padded tensor into the 2 parallel groups of structures respectively. Both groups of structures start from the upper left corner of the tensor, and the order is from left to right, from top to bottom, and 5*5*1 tensors are taken out in sequence with a step size of 1; the taken-out tensors are recombined, and connected row by row in units of rows, and recombined from a shape of 5*5 into pixel values of 1*25; the first layer of 25 linear layers in the 2 groups of structures process the corresponding tensors one by one. The number of input and output channels of each linear layer in the first layer is 1 and N respectively, where N is an adjustable parameter, and a 25*N tensor is obtained after processing; then the second layer of 25 linear layers process the corresponding tensors one by one. The number of input and output channels of each linear layer in the second layer is N and 1 respectively, where N is an adjustable parameter, and a 25*1 tensor is obtained after processing. After calculating the mean value and the tanh activation function, a 1*1 tensor is obtained as the output; the above operations are performed on one of the 3 channels by the 2 parallel layers of structures, and the same processing is performed on the other channels using the same method; a tensor of H*W*3 is still a tensor of H*W*3 after the third stage processing; 2 groups of H*W*3 tensors are obtained after the 2 groups of structures are processed.
[0016] Further, in step S1, the fourth stage S14 is divided into 2 parallel groups of structures, each group of structures is divided into 6 layers, and each layer uses a convolutional layer; the receptive field of the first convolutional layer of one group of structures is 1*3, and the receptive field of the first convolutional layer of the other group of structures is 3*1. The number of input and output channels of the first layer of the 2 groups of structures is 1 and N respectively, where N is an adjustable parameter; the receptive fields of the second to fifth convolutional layers are 1*1, and the number of input and output channels are N and N respectively. The receptive field of the sixth convolutional layer is 1*1, and the number of input and output channels are N and 1 respectively; there is a ReLU activation function after the first to fifth convolutional layers, and a tanh activation function after the sixth convolutional layer;
[0017] The calculation in the fourth stage includes: First, perform edge padding on the H*W*3 tensor obtained after the third-stage processing, and then input the two sets of tensors obtained in the third stage into two sets of structures respectively; starting from the upper left corner of one of the tensors, in the order from left to right and from top to bottom, take out 1*3*1 tensors one by one with a step size of 1, and use a set of structures to process them to obtain 1*1 tensors; starting from the upper left corner of the other tensor, in the order from left to right and from top to bottom, take out 3*3*1 tensors one by one with a step size of 1, and then take out the values in the upper right corner, the middle, and the lower left corner of the 3*3*1 tensor and reorder them into a 3*1 tensor, and use another set of structures to process them to obtain 1*1 tensors; the above operations are performed on one channel in the 3 channels of the tensor, and the same processing is performed on the other channels using the same method; the two sets of H*W*3 tensors are processed as above to obtain two sets of H*W*3 tensors; add the two sets of tensors processed by the two sets of structures to obtain the output of the fourth stage.
[0018] Further, rotate the output tensor of the second stage by 0 degrees, 90 degrees, 180 degrees, and 270 degrees in sequence, and process them through the third stage and the fourth stage respectively to obtain the outputs of 4 tensors. Add the 4 tensors to obtain the final output of the neural network.
[0019] Further, the training process in step S1 includes: Select B data pairs, use the Bayer image as the input of the neural network to obtain a prediction, calculate the loss function of the prediction and the label, calculate the backpropagation and update the parameters to complete one iteration. Optimize the parameters and adjust the learning rate during training. After completing multiple iterations, obtain the trained neural network.
[0020] Further, in step S2, a total of 6 lookup tables are converted in the first stage to the fourth stage; step S2 specifically includes: Step S21, establish input values; Step S22, convert the lookup tables; Step S23, calculate the lookup tables.
[0021] Further, step S21, establishing input values specifically includes: For the linear layer, take a value every 2i within the pixel value range of [0, 257), where i is an adjustable parameter; for the convolutional layer, the receptive field is k*k, and at each receptive field position, a value needs to be taken every 2i within the range of [0, 257), and permutations and combinations are formed.
[0022] Further, step S22, converting the lookup tables specifically includes:
[0023] In the first stage S11, there is a set of lookup tables, which are divided into 2 layers, each with 16 linear layers. The 1 input value in step S21 is sequentially input into the first linear layer of the first layer to obtain 1*N values, and then input into the first linear layer of the second layer to obtain 1*12 values as the lookup table; the 1 input value is sequentially input into the second linear layer of the first layer to obtain 1*N values, and then input into the second linear layer of the second layer to obtain 1*12 values as the lookup table; the 1 input value is sequentially input into the third linear layer of the first layer to obtain 1*N values, and then input into the third linear layer of the second layer to obtain 1*12 values as the lookup table; and so on, a total of 1 set of lookup tables with a total of 16*1*12 values is obtained;
[0024] In the second stage S12, there is a set of lookup tables, and one input value is sequentially input into all the calculation units of the second stage to obtain a set of lookup tables with a total of I*1 values;
[0025] In the third stage S13, there are two sets of lookup tables, which are obtained from the two parallel structures. The conversion process is the same as that of the first stage S11, and a total of two sets of lookup tables are obtained. Each lookup table has a total of 25*1*1 values.
[0026] The fourth stage S14 has two sets of lookup tables, which are obtained from two parallel structures. The conversion process is the same as the second stage S12, and a total of two sets of lookup tables are obtained, each set of lookup tables has a total of I*1 value lookup tables.
[0027] Furthermore, step S23, calculating the lookup table specifically includes:
[0028] In the first stage, the input is a H*W*1 Bayer image, the edge of the Bayer image is filled, and the upper left corner of the edge-filled Bayer image is used as the original coordinate. From left to right and from top to bottom, 4*4*1 pixel values are sequentially taken out with a step size of 2, that is, 4 pairs of BGGR. The pixel values are used as indexes to take out 2 groups of data at the corresponding positions of the lookup table, each group of data has 12 values. The two groups of data are linearly interpolated according to the distance between the pixel values and the sampling points to obtain the output of the lookup table, the mean is calculated, the tanh activation function is used, and the shape is reshaped to obtain 4*3 as the output of the area; the entire Bayer image is calculated in sequence to obtain the final output as a tensor of H*W*3;
[0029] In the second stage, perform edge padding on the H*W*3 tensor output in the first stage. Taking the top-left corner of the three-channel tensor after edge padding as the original coordinates, starting from left to right and top to bottom, sequentially extract 1*3 tensors with a step size of 1. Using the 1*3 tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, with the output size of 1*1*1. Sequentially calculate the tensors for the three channels to obtain the final output as an H*W*3 tensor;
[0030] In the third stage, perform edge padding on the H*W*3 tensor output in the second stage. Process each channel of the tensor after edge padding separately. Taking the top-left corner as the original coordinates, starting from left to right and top to bottom, sequentially extract 5*5*1 tensors with a step size of 1. Input the 5*5*1 tensors into two parallel lookup tables respectively. Using the values of the 5*5*1 tensor as an index, extract 2 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate linear interpolation for the 2 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, with the output size of 25*1. Calculate the mean, apply the tanh activation function, and reshape to obtain 1*1 as the output of the tensor. Sequentially calculate the entire tensor to obtain the final output as two H*W*3 tensors;
[0031] In the fourth stage, perform edge padding on the two H*W*3 tensors output in the third stage; process each channel of the tensor after edge padding separately. Starting from the top-left corner of one of the tensors, in the order from left to right and top to bottom, sequentially extract 1*3*1 tensors with a step size of 1. Using this tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, with the output size of 1*1*1; starting from the top-left corner of the other tensor, in the order from left to right and top to bottom, sequentially extract 3*3*1 tensors with a step size of 1, and then extract the values in the upper-right corner, middle, and lower-left corner of the 3*3*1 tensor and reorder them into a 3*1 tensor. Using this tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, with the output size of 1*1*1; the above operations are performed on 1 channel out of the 3 channels of the tensor, and the same processing is performed on the other channels using the same method; an H*W*3 tensor is obtained after the above processing of an H*W*3 tensor; add the two tensors processed by the 2 sets of lookup tables to obtain the output of the fourth stage.
[0032] Furthermore, step S2 further includes:
[0033] Rotate the output tensor of the second stage by 0 degree, 90 degrees, 180 degrees, and 270 degrees in sequence, and obtain the output of 4 tensors after being processed by the third stage and the fourth stage respectively. Add the 4 tensors to obtain the final output of 6 lookup tables.
[0034] Further, in step S3, during the process of fine-tuning the lookup table, use the dataset established in step S1 and the configuration of training the neural network. Take the calculation process of the lookup table as forward propagation, perform backpropagation, and update the values of all the lookup tables. After the iteration is completed, obtain the fine-tuned lookup table.
[0035] Further, steps S1, S2, and S3 are all offline operations. In actual applications, only the parameter quantity of the lookup table and the interpolation calculation are required.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] The present invention provides an image demosaicing method, which is a lightweight method for demosaicing by combining deep learning and lookup tables. The present invention is divided into 4 steps, namely training a neural network, converting it into a lookup table, fine-tuning the lookup table, and applying the lookup table. The present invention provides a method for increasing the receptive field size and optimizing the design of the neural network structure through the optimized design of the neural network structure, thereby improving the visual effect of demosaicing, while reducing the parameter quantity and computational complexity of the lookup table, reducing the parameter quantity of the lookup table and the computational amount during the interpolation of the lookup table. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic flowchart of an image demosaicing method according to an embodiment of the present invention.
[0039] Figure 2 It is a schematic diagram of a neural network structure in an image demosaicing method according to an embodiment of the present invention.
[0040] Figure 3 It is a first schematic diagram of the comparison of processing images between the present invention and an existing demosaicing method.
[0041] Figure 4 It is a second schematic diagram of the comparison of processing images between the present invention and an existing demosaicing method.
[0042] Figure 5 It is a third schematic diagram of the comparison of processing images between the present invention and an existing demosaicing method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of facilitating and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0044] For ease of description, some embodiments of the present application may use spatial relative terms such as "above", "below", "top", "bottom", etc. to describe the relationship between one element or component and another (or other) element or component as shown in the respective drawings of the embodiments. It should be understood that in addition to the orientations described in the drawings, the spatial relative terms are also intended to include different orientations of the device during use or operation. For example, if the device in the drawing is flipped, an element or component described as "below" or "beneath" other elements or components will subsequently be positioned "above" or "on top of" the other elements or components. The terms "first", "second", etc. in the following text are used to distinguish between similar elements and are not necessarily used to describe a specific order or time sequence. It is to be understood that these terms may be replaced where appropriate.
[0045] Embodiments of the present invention provide an image demosaicing method, as Figure 1 shown, including:
[0046] S1. Convert the RGB three-channel image into a Bayer image; use the Bayer image as the input and the RGB three-channel image as the label to train the demosaicing neural network model; the structure of the demosaicing neural network model includes a first stage to a fourth stage;
[0047] S2. After the demosaicing neural network model converges in training, convert the mapping relationship between the Bayer image and the RGB three-channel image into a lookup table;
[0048] S3. Fine-tune the lookup table;
[0049] S4. Apply the fine-tuned lookup table to the Bayer format image to obtain the demosaiced RGB full-color image.
[0050] The following details each step of the image demosaicing method of the embodiments of the present invention.
[0051] Step S1: Convert the RGB three-channel image into a Bayer image; use the Bayer image as the input and the RGB three-channel image as the label to train the demosaicking neural network model. Specifically, establish a dataset corresponding to the input and the label; for example, convert multiple (such as 800 to 1000) RGB three-channel images into Bayer images. Assume the resolution of the RGB three-channel image is H*W*3, where H is the number of rows of the image, W is the number of columns of the image, and 3 is the number of image channels, and the three channels are the R channel, the G channel, and the B channel respectively. The process of converting to a Bayer image is to take an image block of size 2*2*3 with the top-left corner of the RGB three-channel image as the original coordinate, from left to right, top to bottom, with a step size of 2; then take the B value from the top-left coordinate in the image block, the G value from the top-right coordinate, the G value from the bottom-left coordinate, and the R value from the bottom-right coordinate to form an image block of size 2*2*1. The entire image is processed in the same way, and a Bayer image of H*W*1 can be obtained from an RGB three-channel image of H*W*3. Use the Bayer image of H*W*1 as the input and the H*W*3 image as the label to establish the dataset.
[0052] Training the demosaicking neural network model is divided into four stages, that is, the demosaicking neural network structure is divided into four stages. As Figure 2As shown, the first stage S11 is divided into two layers, with 16 linear layers, i.e., fully connected layers, used in each layer. The first linear layer is 1*N, and the second linear layer is N*12, where N is an adjustable parameter. The input to the first stage S11 is the Bayer image. Edge padding is performed on this Bayer image. The rightmost two columns of the image are copied and filled to the rightmost edge, and the bottommost two rows of the image are copied and filled to the bottommost edge. The first stage processes the edge-padded Bayer image. Starting from the upper left corner, in the order from left to right and top to bottom, image blocks of 4*4*1 are sequentially taken out with a step size of 2, that is, 4 pairs of BGGR. The 4 pairs of BGGR taken out are recombined, and connected sequentially in rows, that is, the second row is connected to the first row, the third row is connected to the second row, and the fourth row is connected to the third row, and they are recombined from a shape of 4*4*1 into pixel values of 1*16. The 16 linear layers in the first layer process the pixel values one by one, that is, the first linear layer processes the first pixel value, the second linear layer processes the second pixel value, the third linear layer processes the third pixel value, and so on; after processing, a tensor of 16*N is obtained. Then, the 16 linear layers in the second layer process the tensor one by one, that is, the first N*12 linear layer processes the first 1*N tensor to obtain a 1*12 tensor, the second N*12 linear layer processes the second 1*N tensor to obtain a 1*12 tensor, the third N*12 linear layer processes the third 1*N tensor to obtain a 1*12 tensor, and so on. After processing, a tensor of 16*12 is obtained. After calculating the mean value and the tanh activation function, a 1*12 tensor is obtained, and the shape of this tensor is reshaped to obtain a 2*2*3 tensor as the output of this image block. A Bayer image of H*W*1 is processed by the first stage to obtain a tensor of H*W*3.
[0053] The second stage S12 is divided into six layers, and each layer uses a convolutional layer. The receptive field of the first convolutional layer is 1*3, and the number of input and output channels is 1 and N respectively, where N is an adjustable parameter. The receptive fields of the second to fifth convolutional layers are 1*1, and the number of input and output channels is N and N respectively. The receptive field of the sixth convolutional layer is 1*1, and the number of input and output channels is N and 1 respectively. There is a ReLU activation function after the first to fifth convolutional layers, and a tanh activation function after the sixth convolutional layer. The calculation process of the second stage S12 is as follows: First, edge padding is performed on the H*W*3 tensor obtained after the first stage processing. The leftmost column of the tensor is copied and filled to the leftmost edge, and the rightmost column of the tensor is copied and filled to the rightmost edge. Then, the six convolutional layers process the input tensor in order. The above operations are performed on one channel among the three channels of the tensor, and the same method is used to perform the same processing on the other channels. A tensor of H*W*3 is still obtained after the above processing of a tensor of H*W*3.
[0054] The third stage S13 is divided into two parallel groups of structures. Each group of structures ACUnit1 is further divided into two layers, and each layer uses 25 linear layers. The first linear layer of each group of structures is 1*N, and the second linear layer is N*1, where N is an adjustable parameter. Input the tensor of H*W*3 obtained after the second stage processing and perform edge padding. Copy the rightmost 4 columns of the image and fill them to the rightmost edge, and copy the bottommost 4 rows of the image and fill them to the bottommost edge. Input the edge-padded tensor into two parallel groups of structures respectively. Both groups of structures start from the upper left corner of the tensor, in the order from left to right and from top to bottom, and take out tensors of 5*5*1 in sequence with a step size of 1. Recombine the taken-out tensors and connect them row by row in sequence, that is, the second row is connected to the first row, the third row is connected to the second row, the fourth row is connected to the third row, and recombine from the shape of 5*5*1 to 1*25 pixel values. The first 25 linear layers of each group of structures process the corresponding tensors one by one, that is, the first linear layer processes the first tensor value, the second linear layer processes the second tensor value, the third linear layer processes the third tensor value, and so on. After processing, a tensor of 25*N is obtained. Then use the second 25 linear layers to process the corresponding tensors one by one, that is, the first N*1 linear layer processes the first 1*N tensor to obtain a 1*1 tensor, the second N*1 linear layer processes the second 1*N tensor to obtain a 1*1 tensor, the third N*1 linear layer processes the third 1*N tensor to obtain a 1*1 tensor, and so on. After processing, a tensor of 25*1 is obtained. Calculate the mean value and the tanh activation function to obtain a 1*1 tensor as the output. The above operations are for two parallel layers of structures to process one channel in the three channels, and the same method is used to perform the same processing on the other channels. A tensor of H*W*3 is obtained after the above processing. Two groups of structures will obtain two groups of tensors of H*W*3 after processing.
[0055] The fourth stage S14 is divided into two parallel groups of structures, namely FUnit0 and Funit1. Each group of structures is further divided into six layers, and each layer uses a convolutional layer. Among them, the receptive field of the first convolutional layer of the first group of structures FUnit0 is 1*3, and the receptive field of the first convolutional layer of the second group of structures Funit1 is 3*1. The number of input and output channels of the first layer of the two groups of structures are 1 and N respectively, where N is an adjustable parameter. The receptive fields of the second to fifth convolutional layers are 1*1, and the number of input and output channels are N and N respectively. The receptive field of the sixth convolutional layer is 1*1, and the number of input and output channels are N and 1 respectively. There is a ReLU activation function after the first to fifth convolutional layers, and a tanh activation function after the sixth convolutional layer. The calculation process of the fourth stage S14 is as follows: First, perform edge padding on the H*W*3 tensor obtained after the third stage processing. Copy the leftmost column of one of the tensors and fill it to the leftmost edge, and copy the rightmost column of the tensor and fill it to the rightmost edge; copy the leftmost column of the other tensor and fill it to the leftmost edge, copy the rightmost column of the tensor and fill it to the rightmost edge, copy the topmost row of the tensor and fill it to the topmost edge, and copy the bottommost row of the tensor and fill it to the bottommost edge. Then, input the two groups of tensors obtained in the third stage into the two groups of structures respectively. Starting from the upper left corner of one of the tensors, in the order from left to right and from top to bottom, take out 1*3*1 tensors one by one with a step size of 1, and use a group of structures to process them to obtain 1*1 tensors; starting from the upper left corner of the other tensor, in the order from left to right and from top to bottom, take out 3*3*1 tensors one by one with a step size of 1, and then take out the values in the upper right corner, middle, and lower left corner of the 3*3*1 tensor and reorder them into a 3*1 tensor, and use the other group of structures to process them to obtain 1*1 tensors. The above operations are performed on one channel of the 3 channels of the tensor, and the same processing is performed on the other channels using the same method. After the above processing, two groups of H*W*3 tensors are obtained. Add the two groups of tensors processed by the two groups of structures to obtain the output of the fourth stage.
[0056] The above processing in the third and fourth stages is the processing of one direction of the output tensor of the second stage. Rotate the output tensor of the second stage by 0 degrees, 90 degrees, 180 degrees, and 270 degrees in turn and process them through the third and fourth stages to obtain the output of four tensors. Add the four tensors to obtain the final output of the neural network.
[0057] In training the demosaicing neural network model, training configurations are set. In the present invention, the Adam optimizer and the learning rate adjustment method of cosine annealing are adopted, and a total of E iterations are performed, where E is an adjustable parameter. Randomly select a data block of size C*C*3 from the RGB three-channel images in the training set and normalize it. C is an adjustable parameter. The normalization is to divide the 8-bit data image by 255. This data block serves as the label of the neural network, and then extract it into a Bayer image in the manner of step S1 as the input of the neural network. During the training process, the mini-batch gradient descent method is adopted, and each B data pairs are iterated for one training, where B is an adjustable parameter. Save the neural network structure and parameters after training for the next use. The initialization of the neural network uses the kaiming method, and the mean square error is used as the loss function. The training process is as follows: Select B data pairs, use the Bayer image as the input of the neural network to obtain a prediction, calculate the loss function of the prediction and the label, calculate the backpropagation and update the parameters to complete one iteration. During training, the Adam algorithm is used to optimize the parameters and the learning rate is adjusted in the way of cosine annealing. After completing E iterations, a trained neural network is obtained.
[0058] Step S2: After the training of the demosaicing neural network model converges, convert the mapping relationship between the Bayer image and the RGB three-channel image into a look-up table. Specifically, the neural network structure is divided into 4 stages in total, and 6 groups of look-up tables are converted in total.
[0059] S21: Establish input values; for the linear layer, take a value every 2i within the pixel value range of [0, 257), where i is an adjustable parameter. For example, when i = 1, the values are 0, 2, 4, 6, 8, …, 252, 254, 256; for the convolutional layer, the receptive field is k*k, and at each receptive field position, a value is taken every 2i within the range of [0, 257), and a permutation and combination is formed; for example, when the receptive field is 1*3 and i = 0, the values are (0, 0, 0), (0, 0, 1), (0, 0, 2), (0, 0, 3),.., (0, 0, 256), (0, 1, 0), (0, 1, 1), (0, 1, 2), …, (0, 256, 256), …, (1, 0, 0), …, (256, 256, 256).
[0060] S22, convert the lookup table; in the first stage S11, there is a total of 1 set of lookup tables, which are divided into 2 layers, with 16 linear layers in each layer. The 1 input value in step S21 is sequentially input into the first linear layer of the first layer to obtain 1*N values, and then input into the first linear layer of the second layer to obtain 1*12 values as the lookup table; the 1 input value is sequentially input into the second linear layer of the first layer to obtain 1*N values, and then input into the second linear layer of the second layer to obtain 1*12 values as the lookup table; the 1 input value is sequentially input into the third linear layer of the first layer to obtain 1*N values, and then input into the third linear layer of the second layer to obtain 1*12 values as the lookup table; and so on, a total of 1 set of lookup tables with a total of 16*1*12 values is obtained.
[0061] In the second stage S12, there is a total of 1 set of lookup tables. I input values are sequentially input into all calculation units in the second stage to obtain a set of lookup tables with a total of I*1 values.
[0062] The third stage S13 has two sets of lookup tables, which are obtained from two parallel structures. The conversion process is the same as the first stage S11, and a total of two sets of lookup tables are obtained. Each lookup table has a total of 25*I*1 values.
[0063] The fourth stage S14 has two sets of lookup tables, which are obtained from two parallel structures. The conversion process is the same as the second stage S12, and a total of two sets of lookup tables are obtained. Each set of lookup tables has a total of I*1 value lookup tables.
[0064] S23. Calculate the lookup table. In the first stage, the input is a Bayer image of H*W*1. The Bayer image is edge-filled. The rightmost two columns of the image are copied and filled to the rightmost edge, and the bottom two rows of the image are copied and filled to the bottom edge. Taking the upper left corner as the original coordinate, from left to right, from top to bottom, with a step size of 2, 4*4*1 pixel values are sequentially taken out, i.e., 4 pairs of BGGR. The pixel value is used as the index to take out the 2 sets of data at the corresponding position of the lookup table. Each set of data has 12 values. The 2 sets of data are linearly interpolated according to the distance between the pixel value and the sampling point to obtain the output of the lookup table. The output size is 1*16*12. The mean, tanh activation function and reshape are calculated to obtain 4*3 as the output of the area. The entire Bayer image is calculated in sequence to obtain a final output of a tensor of H*W*3.
[0065] In the second stage, the input is a three-channel tensor of H*W*3. Edge padding is performed on the input tensor. The leftmost column of the tensor is copied and padded to the leftmost edge, and the rightmost column of the tensor is copied and padded to the rightmost edge. Each channel is processed separately. With the upper left corner as the original coordinate, a 1*3 tensor is taken step by step from left to right and top to bottom with a step size of 1. Using this tensor as an index, 4 sets of data are taken from the corresponding positions in the lookup table. Each set of data has 1 value. The 4 sets of data are calculated by tetrahedral interpolation according to the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the output size is 1*1. The tensors of the three channels are calculated in turn to obtain the final output as a tensor of H*W*3.
[0066] In the third stage, the input is a tensor of H*W*3. Edge padding is performed on this tensor. The rightmost 4 columns of the tensor are copied and padded to the rightmost edge, and the bottommost 4 rows of the image are copied and padded to the bottommost edge. Each channel is processed separately. With the upper left corner as the original coordinate, a 5*5*1 tensor is taken step by step from left to right and top to bottom with a step size of 1. This tensor is input into two parallel lookup tables respectively. Using the values of this tensor as indices, 2 sets of data are taken from the corresponding positions in the lookup table. Each set of data has 1 value. The 2 sets of data are calculated by linear interpolation according to the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the output size is 25*1. The mean value, tanh activation function, and reshaping are calculated to obtain 1*1 as the output of this tensor. The entire tensor is calculated in turn to obtain the final output as two tensors of H*W*3.
[0067] In the fourth stage, the input consists of two tensors of size H*W*3. Perform edge padding on the two tensors. For one tensor, copy the leftmost column and pad it to the left edge, and copy the rightmost column and pad it to the right edge. For the other tensor, copy the leftmost column and pad it to the left edge, copy the rightmost column and pad it to the right edge, copy the topmost row and pad it to the top edge, and copy the bottommost row and pad it to the bottom edge. Process each channel separately. Starting from the upper left corner of one of the tensors, in the order from left to right and top to bottom, extract tensors of size 1*3*1 one by one with a step size of 1. Use this tensor as an index to retrieve 4 sets of data from the lookup table, each set having 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, which has a size of 1*1. Starting from the upper left corner of the other tensor, in the order from left to right and top to bottom, extract tensors of size 3*3*1 one by one with a step size of 1. Then, extract the values in the upper right corner, the middle, and the lower left corner from the 3*3*1 tensor and reorder them into a tensor of size 3*1. Use this tensor as an index to retrieve 4 sets of data from the lookup table, each set having 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, which has a size of 1*1. The above operations are performed on one channel out of the 3 channels of the tensor, and the same method is used to process the other channels. After the above processing, a tensor of size H*W*3 is obtained from a tensor of size H*W*3. Add the two sets of tensors processed by the two lookup tables to obtain the output of the fourth stage.
[0068] One processing of the above third and fourth stages is for one direction of the output tensor of the second stage. Rotate the output tensor of the second stage by 0 degrees, 90 degrees, 180 degrees, and 270 degrees in sequence and process them through the third and fourth stages respectively to obtain the outputs of 4 tensors. Add the 4 tensors to obtain the final output of 6 lookup tables.
[0069] Step S3: Fine-tune the lookup table. Specifically, in the process of fine-tuning the lookup table, use the dataset established in step S1 and the configuration of training the neural network. Take the calculation process of the lookup table as the forward propagation, perform backpropagation, and update the values of all lookup tables. After the iteration is completed, obtain the fine-tuned lookup table.
[0070] Step S4: Apply the fine-tuned lookup table to the Bayer format image to obtain the demosaicked RGB full-color image. Use the lookup table calculation process in step S3 and the fine-tuned lookup table as parameters to calculate the Bayer image to obtain the demosaicked RGB three-channel image.
[0071] Steps S1, S2, and S3 are all offline operations. In practical applications, only the parameter quantity of the lookup table and the interpolation calculation need to be considered.
[0072] The present invention provides a lightweight method for demosaicing by combining deep learning and a lookup table. The present invention is divided into 4 steps, namely training a neural network, converting it into a lookup table, fine-tuning the lookup table, and applying the lookup table. The present invention provides a method for increasing the receptive field size and optimizing the design of the neural network structure through the optimized design of the neural network structure, thereby improving the visual effect of demosaicing, while reducing the parameter quantity and computational complexity of the lookup table, reducing the parameter quantity of the lookup table and the computational amount during the interpolation of the lookup table.
[0073] Figure 3 It is the first schematic diagram for comparing the images processed by the present invention and an existing demosaicing method. Figure 4 It is the second schematic diagram for comparing the images processed by the present invention and an existing demosaicing method. Figure 5 It is the third schematic diagram for comparing the images processed by the present invention and an existing demosaicing method. Figures 3 to 5 The existing demosaicing method in [ ] adopts the same method.
[0074] As Figures 3 to 5 shown, the SSIM parameter index and PSNR (peak signal-to-noise ratio) parameter index of the image processed by the demosaicing method of the present invention and the image processed by an existing lookup table demosaicing method are quite the same (basically at the same level); however, the demosaicing method of the present invention greatly reduces the memory size occupied by the lookup table, only occupying dozens of KB (for example, actually measured as 18 KB), while an existing lookup table occupies several thousand KB (for example, actually measured as 1712 KB). The demosaicing method of the present invention maintains the SSIM basically the same while improving the visual effect under the condition of greatly reducing the size of the lookup table. SSIM (structural similarity index) is an index used to measure the similarity between two images and is widely used in fields such as image quality assessment, compression effect analysis, and image reconstruction. SSIM starts from the characteristics of the human visual system, comprehensively considers the brightness, contrast, and structural information of the image, and is more in line with human subjective perception.
[0075] In summary, the present invention provides an image demosaicing method, and the present invention is a lightweight method for demosaicing by combining deep learning and a lookup table. The present invention is divided into 4 steps, namely training a neural network, converting it into a lookup table, fine-tuning the lookup table, and applying the lookup table. The present invention provides a method for increasing the receptive field size and optimizing the design of the neural network structure through the optimized design of the neural network structure, thereby improving the visual effect of demosaicing, while reducing the parameter quantity and computational complexity of the lookup table, reducing the parameter quantity of the lookup table and the computational amount during the interpolation of the lookup table.
[0076] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the methods disclosed in the embodiments, since they correspond to the devices disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0077] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the rights of the present invention in any way. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention all fall within the protection scope of the technical solution of the present invention.
Claims
1. An image demosaicing method, characterized in that, Including: S1. Convert the RGB three-channel image into a Bayer image; Use the Bayer image as the input and the RGB three-channel image as the label to train the demosaicing neural network model; The structure of the demosaicing neural network model includes the first stage to the fourth stage; S2. After the training of the demosaicing neural network model converges, convert the mapping relationship between the Bayer image and the RGB three-channel image into a lookup table; S3. Fine-tune the lookup table; S4. Apply the fine-tuned lookup table to the Bayer format image to obtain the demosaiced RGB full-color image.
2. The image demosaicing method according to claim 1, wherein In step S1, specifically including: the resolution of the RGB three-channel image is H*W*3, where H is the number of rows of the image, W is the number of columns of the image, and 3 is the number of channels; starting from the upper left corner of the RGB three-channel image as the original coordinate, from left to right, from top to bottom, with a step size of 2, extract image blocks of size 2*2*3; Take the B value at the upper left coordinate, the G value at the upper right coordinate, the G value at the lower left coordinate, and the R value at the lower right coordinate in the image block to form an image block of size 2*2*1. The entire RGB three-channel image is processed in the same way, and an H*W*1 Bayer image is obtained from an H*W*3 RGB three-channel image; establish the dataset of the input and the label.
3. The image demosaicing method according to claim 1, wherein In step S1, the first stage S11 is divided into 2 layers, and each layer uses multiple linear layers; perform edge padding on the Bayer image. Starting from the upper left corner of the edge-padded Bayer image, from left to right, from top to bottom, extract image blocks of 4*4*1 with a step size of 2, that is, 4 pairs of BGGR; recombine the 4 pairs of BGGR, and connect them row by row in order. Recombine from a shape of 4*4 to 1*16 pixel values; process the 16 pixel values one by one with 16 linear layers in the first layer. The number of input and output channels of each linear layer in the first layer is 1 and N respectively, where N is an adjustable parameter. After processing, a tensor of 16*N is obtained; then use 16 linear layers in the second layer to process the 16*N tensor one by one. The number of input and output channels of each linear layer in the second layer is N and 12 respectively, where N is an adjustable parameter. After calculating the mean value and the tanh activation function, a tensor of 1*12 is obtained. Reshape the 1*12 tensor to obtain a tensor of 2*2*3 as the output; an H*W*1 Bayer image is processed by the first stage to obtain a tensor of H*W*3.
4. The image demosaicing method according to claim 1, wherein In step S1, the second stage S12 is divided into 6 layers, and each layer adopts a convolutional layer; the receptive field of the first convolutional layer is 1*3, and the number of input and output channels is 1 and N respectively, where N is an adjustable parameter; the receptive fields of the second to fifth convolutional layers are 1*1, and the number of input and output channels is N and N respectively; the receptive field of the sixth convolutional layer is 1*1, and the number of input and output channels is N and 1 respectively; there is a ReLU activation function after the first to fifth convolutional layers, and there is a tanh activation function after the sixth convolutional layer; The calculation process of the second stage includes: first, perform edge padding on the H*W*3 tensor obtained after the first stage is processed, and then the 6 convolutional layers process the input tensor in sequence; the above operations process 1 channel out of the 3 channels of the tensor, and use the same method to perform the same processing on the other channels; a tensor of H*W*3 is still a tensor of H*W*3 after the above processing.
5. The image demosaicing method according to claim 1, wherein In step S1, the third stage S13 is divided into 2 parallel structures, and each structure is further divided into 2 layers, and each layer uses 25 linear layers; perform edge padding on the H*W*3 tensor obtained after the second stage is processed, and input the edge-padded tensor into the 2 parallel structures respectively. The 2 structures both start from the upper left corner of the tensor, and the order is from left to right and from top to bottom, and 5*5*1 tensors are taken out in sequence with a step size of 1; the taken-out tensors are recombined, and are connected row by row in units of rows, and are recombined from a shape of 5*5 into pixel values of 1*25; the first layer of 25 linear layers in the 2 structures process the corresponding tensors one by one. The number of input and output channels of each linear layer in the first layer is 1 and N respectively, where N is an adjustable parameter, and a tensor of 25*N is obtained after processing; then the second layer of 25 linear layers process the corresponding tensors one by one. The number of input and output channels of each linear layer in the second layer is N and 1 respectively, where N is an adjustable parameter, and a tensor of 25*1 is obtained after processing. After calculating the mean value and the tanh activation function, a 1*1 tensor is obtained as the output; the above operations are for 1 channel out of the 3 channels by the 2 parallel layers, and use the same method to perform the same processing on the other channels; a tensor of H*W*3 is still a tensor of H*W*3 after being processed by the third stage; 2 groups of H*W*3 tensors will be obtained after the 2 structures are processed.
6. The image demosaicing method according to claim 1, wherein In step S1, the fourth stage S14 is divided into two parallel groups of structures. Each group of structures is further divided into six layers, and each layer uses a convolutional layer. The receptive field of the first convolutional layer in one group of structures is 1×3, and the receptive field of the first convolutional layer in the other group of structures is 3×1. The number of input and output channels of the first layer of the two groups of structures are 1 and N respectively, where N is an adjustable parameter. The receptive fields of the second to fifth convolutional layers are 1×1, and the number of input and output channels are N and N respectively. The receptive field of the sixth convolutional layer is 1×1, and the number of input and output channels are N and 1 respectively. There is a ReLU activation function after the first to fifth convolutional layers, and a tanh activation function after the sixth convolutional layer. The calculation in the fourth stage includes: first, perform edge padding on the H×W×3 tensor obtained after processing in the third stage, and then input the two groups of tensors obtained in the third stage into the two groups of structures respectively. Starting from the upper left corner of one of the tensors, in the order from left to right and top to bottom, take out 1×3×1 tensors one by one with a step size of 1, and use one group of structures to process them to obtain 1×1 tensors. Starting from the upper left corner of the other tensor, in the order from left to right and top to bottom, take out 3×3×1 tensors one by one with a step size of 1, and then take out the values in the upper right corner, middle, and lower left corner of the 3×3×1 tensor and reorder them into a 3×1 tensor, and use the other group of structures to process them to obtain 1×1 tensors. The above operations are performed on one channel out of the three channels of the tensor, and the same processing is performed on the other channels using the same method. After the above processing, two groups of H×W×3 tensors are obtained. Add the two groups of tensors processed by the two groups of structures to obtain the output of the fourth stage.
7. The image demosaicking method according to claim 1, wherein Rotate the output tensor of the second stage by 0 degree, 90 degrees, 180 degrees, and 270 degrees respectively, and process them through the third stage and the fourth stage to obtain the outputs of four tensors. Add the four tensors to obtain the final output of the neural network.
8. The image demosaicking method according to claim 1, wherein The training process in step S1 includes: selecting B data pairs, using the Bayer image as the input of the neural network to obtain a prediction, calculating the loss function between the prediction and the label, calculating the backpropagation and updating the parameters to complete one iteration. Optimize the parameters and adjust the learning rate during training. After completing multiple iterations, obtain the trained neural network.
9. The image demosaicking method according to claim 1, wherein In step S2, a total of six groups of lookup tables are converted in the first to fourth stages. Step S2 specifically includes: step S21, establishing input values; step S22, converting the lookup tables; step S23, calculating the lookup tables.
10. The image demosaicking method according to claim 9, wherein Step S21 of establishing the input values specifically includes: for the linear layer, take a value every 2^i within the pixel value range of [0, 257), where i is an adjustable parameter; for the convolutional layer, the receptive field is k*k, and at each receptive field position, a value needs to be taken every 2^i within the range of [0, 257), and permutations and combinations are formed.
11. The image demosaicing method according to claim 10, wherein Step S22 of converting the lookup table specifically includes: In the first stage S11, there is a total of 1 group of lookup tables, which are divided into 2 layers, with 16 linear layers in each layer. Input the I input values in step S21 into the first linear layer of the first layer in sequence to obtain I*N values, and then input them into the first linear layer of the second layer to obtain I*12 values as the lookup table; input the I input values into the second linear layer of the first layer in sequence to obtain I*N values, and then input them into the second linear layer of the second layer to obtain I*12 values as the lookup table; input the I input values into the third linear layer of the first layer in sequence to obtain I*N values, and then input them into the third linear layer of the second layer to obtain I*12 values as the lookup table; and so on, a total of 1 group of lookup tables with 16*I*12 values is obtained; In the second stage S12, there is a total of 1 group of lookup tables. Input the I input values into all the computing units in the second stage in sequence to obtain 1 group of lookup tables with a total of I*1 values; In the third stage S13, there are a total of 2 groups of lookup tables, which are obtained from 2 parallel structures. The conversion process is the same as that of the first stage S11, and a total of 2 groups of lookup tables are obtained, with 25*I*1 values in each group of lookup tables; In the fourth stage S14, there are a total of 2 groups of lookup tables, which are obtained from 2 parallel structures. The conversion process is the same as that of the second stage S12, and a total of 2 groups of lookup tables are obtained, with I*1 values in each group of lookup tables.
12. The image demosaicing method according to claim 9, wherein Step S23 of calculating the lookup table specifically includes: In the first stage, the input is a Bayer image of H*W*1. Perform edge padding on the Bayer image. Starting from the upper left corner of the edge-padded Bayer image as the original coordinate, from left to right, from top to bottom, take out 4*4*1 pixel values at a step size of 2, that is, 4 pairs of BGGR. Using the pixel values as indices, take out 2 groups of data at the corresponding positions in the lookup table. Each group of data has 12 values. Calculate the output of the lookup table through linear interpolation according to the distance between the pixel values and the sampling points for the 2 groups of data, calculate the mean value, apply the tanh activation function, and reshape to obtain 4*3 as the output of this area; calculate the entire Bayer image in sequence to obtain a final output of a tensor of H*W*3; In the second stage, perform edge padding on the H*W*3 tensor output in the first stage. Taking the upper left corner of the three-channel tensor after edge padding as the original coordinates, sequentially extract 1*3 tensors from left to right and top to bottom with a step size of 1. Using the 1*3 tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the output size is 1*1*1. Sequentially calculate the tensors of the three channels to obtain the final output as an H*W*3 tensor; In the third stage, perform edge padding on the H*W*3 tensor output in the second stage. Process each channel after edge padding separately. Taking the upper left corner as the original coordinates, sequentially extract 5*5*1 tensors from left to right and top to bottom with a step size of 1, and input the 5*5*1 tensors into two parallel lookup tables respectively; using the values of the 5*5*1 tensor as an index, extract 2 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate linear interpolation for the 2 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the size of this output is 25*1. Calculate the mean, apply the tanh activation function, and reshape to obtain 1*1 as the output of this tensor; sequentially calculate the entire tensor to obtain the final output as two H*W*3 tensors; In the fourth stage, perform edge padding on the two H*W*3 tensors output in the third stage; process each channel after edge padding separately. Starting from the upper left corner of one of the tensors, in the order from left to right and top to bottom, sequentially extract 1*3*1 tensors with a step size of 1. Using this tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the output size is 1*1*1; starting from the upper left corner of the other tensor, in the order from left to right and top to bottom, sequentially extract 3*3*1 tensors with a step size of 1, and then extract the upper right corner, middle, and lower left corner values from the 3*3*1 tensor and reorder them into a 3*1 tensor. Using this tensor as an index, extract 4 sets of data at the corresponding positions in the lookup table. Each set of data has 1 value. Calculate tetrahedral interpolation for the 4 sets of data based on the distance between the pixel value and the sampling point to obtain the output of the lookup table, and the output size is 1*1*1; the above operations are performed on 1 channel out of the 3 channels of the tensor, and the same processing is performed on the other channels using the same method; an H*W*3 tensor is obtained after the above processing on an H*W*3 tensor; add the two tensors processed by the 2 sets of lookup tables to obtain the output of the fourth stage.
13. The image demosaicking method according to claim 9, wherein Step S2 further includes: Rotate the tensor output in the second stage 0 degrees, 90 degrees, 180 degrees, and 270 degrees in sequence, and process them through the third stage and the fourth stage respectively to obtain the outputs of 4 tensors. Add the 4 tensors to obtain the final output of 6 sets of lookup tables.
14. The image demosaicing method according to claim 1, wherein in step S3, in the process of fine-tuning the look-up table, the data set established in step S1 and the configuration for training the neural network are used, the calculation process of the look-up table is regarded as forward propagation, backpropagation is performed, and the values of all the look-up tables are updated. After the iteration is completed, the fine-tuned look-up table is obtained.
15. The image demosaicing method according to claim 1, wherein steps S1, S2 and S3 are all offline operations, and only the parameter quantity of the look-up table and interpolation calculation are required in practical applications.