Depth image pixel enhancement method, apparatus, and computer-readable storage medium

By using guided convolutional networks and target convolutional networks to extract and fuse features from high-resolution color images and low-resolution depth images, the problem of enhancing low-resolution depth images is solved, and the image resolution and quality are improved.

CN114445277BActive Publication Date: 2025-11-04SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111538003.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-11-04
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

Existing low-cost depth cameras produce depth images with low resolution and high noise, making them difficult to apply in practice.

Method used

By acquiring high-resolution color guiding images and low-resolution images to be processed for preprocessing, feature extraction is performed using guiding convolutional networks and target convolutional networks, and image fusion and enhancement are performed by combining pixel adaptive convolutional modules.

Benefits of technology

It improves the resolution of low-resolution depth images, enhances image quality and accuracy, and solves the problem of low-resolution image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445277B_ABST
    Figure CN114445277B_ABST
Patent Text Reader

Abstract

The application discloses a depth image pixel enhancement method and device and a computer readable storage medium. The method comprises the following steps: obtaining a to-be-processed image and a color guide image; wherein the color guide image is a color image of the to-be-processed image, and the resolution of the color guide image is higher than that of the to-be-processed image; pre-processing the to-be-processed image to obtain a pre-processed image; inputting the pre-processed image into a target convolution network to obtain a first feature image; inputting the color guide image into a guide convolution network to obtain a second feature image; and obtaining a pixel-enhanced image based on the first feature image, the second feature image and the to-be-processed image. Through the above method, the low-resolution image can be pixel-enhanced, and the image quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to a depth image pixel enhancement method and device and a computer readable storage medium. BACKGROUND

[0002] With the rapid development of sensor technology, depth cameras have important applications in automatic driving, three-dimensional reconstruction, etc. However, the depth images obtained by low-cost depth cameras have the problems of low resolution and excessive image noise, which brings great difficulties in practical applications. SUMMARY

[0003] The present application mainly provides a depth image pixel enhancement method, device and computer readable storage medium, which solves the problem that it is difficult to enhance the resolution of low-resolution images in the prior art.

[0004] To solve the above technical problems, the first aspect of the present application provides a depth image pixel enhancement method, comprising: obtaining a to-be-processed image and a color guide image; wherein the color guide image is a color image of the to-be-processed image, and the resolution of the color guide image is higher than that of the to-be-processed image; pre-processing the to-be-processed image to obtain a pre-processed image; inputting the pre-processed image into a target convolution network to obtain a first feature image; inputting the color guide image into a guide convolution network to obtain a second feature image; and obtaining a pixel-enhanced image based on the first feature image, the second feature image and the to-be-processed image.

[0005] To solve the above technical problems, the second aspect of the present application provides a depth image pixel enhancement device, comprising: an acquisition module for acquiring a to-be-processed image and a color guide image; wherein the color guide image is a color image of the to-be-processed image, and the resolution of the color guide image is higher than that of the to-be-processed image; a pre-processing module for pre-processing the to-be-processed image to obtain a pre-processed image; a feature extraction module for inputting the pre-processed image into a target convolution network to obtain a first feature image, and inputting the color guide image into a guide convolution network to obtain a second feature image; and a pixel enhancement module for obtaining a pixel-enhanced image based on the first feature image, the second feature image and the to-be-processed image.

[0006] To solve the above technical problems, the third aspect of the present application provides a depth image pixel enhancement device, which comprises a processor and a memory coupled to each other; the memory stores a computer program, and the processor is configured to execute the computer program to implement the depth image pixel enhancement method provided in the first aspect.

[0007] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program data, and the program data is executed by a processor to implement the deep image pixel enhancement method of the first aspect.

[0008] The beneficial effects of the present application are: different from the prior art, the present application first acquires a to-be-processed image and a color guide image, wherein the color guide image is a color image of the to-be-processed image, and the resolution of the color guide image is higher than that of the to-be-processed image, the to-be-processed image is preprocessed to obtain a preprocessed image, the preprocessed image is input into a target convolution network to obtain a first feature image, the color guide image is input into a guide convolution network to obtain a second feature image, and based on the first feature image, the second feature image and the to-be-processed image, a pixel-enhanced image is obtained. The present application provides a high-resolution color image of the to-be-processed image, and uses the guide convolution network and the target convolution network to extract features from the high-resolution color image and the to-be-processed image respectively, and then performs pixel enhancement on the to-be-processed image according to the feature images output by the guide convolution network and the target convolution network, which can effectively improve the resolution of the to-be-processed image and has good processing effect. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0010] Figure 1 is a flowchart of an embodiment of the deep image pixel enhancement method of the present application;

[0011] Figure 2 is a structural schematic diagram of an embodiment of the deep image pixel enhancement network of the present application;

[0012] Figure 3 is a flowchart of an embodiment of step S15 of the present application;

[0013] Figure 4 is a flowchart of an embodiment of step S151 of the present application;

[0014] Figure 5 is a structural schematic diagram of an embodiment of the deep image pixel enhancement device of the present application;

[0015] Figure 6 is a structural schematic diagram of another embodiment of the deep image pixel enhancement device of the present application;

[0016] Figure 7is a structural schematic block diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0018] The terms "first", "second" in the present application are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0019] In this document, the reference to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a particular embodiment that is independent of or alternative to other embodiments. The skilled person explicitly and implicitly understands that the embodiments described herein can be combined with other embodiments.

[0020] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the depth image pixel enhancement method of the present application. It should be noted that the flow order shown in the present embodiment is not limited if there is substantially the same result. The present embodiment includes the following steps: Figure 1

[0021] Step S11: Obtain a to-be-processed image and a color guide image.

[0022] The color guide image is a color image of the to-be-processed image, and the resolution of the color guide image is higher than that of the to-be-processed image.

[0023] ​The embodiment can acquire an image or a video photographed by a mobile terminal, a camera, a video recorder, a monitoring device, or the like. For a video, the image to be processed and the color guide image can be one or more frames of images selected from the video. Of course, the image to be processed and the color guide image can also be acquired from the cloud or a network, can also be acquired from the local storage of the device, and can also be acquired from a mobile hard disk or a U disk. The format of the original image can be PNG, JPG, tiff, or other formats, which are not limited herein.

[0024] Step S12: Preprocessing the image to be processed to obtain a preprocessed image.

[0025] The resolution of the preprocessed image obtained in this step is the same as that of the color guide image.

[0026] Specifically, the color guide image is a high-resolution image with a size of HxWxC, and the image to be processed is a low-resolution image with a size of where H, W, and C are the horizontal resolution, the vertical resolution, and the channel number of the color guide image respectively, and l is a down-sampling factor, which can be 4, 8, 16, or the like. The resolution of the image to be processed needs to be adjusted to be the same as that of the color guide image through preprocessing.

[0027] In an embodiment, this step first determines the resolution of the color guide image, and performs a bicubic interpolation operation on the image to be processed to obtain a preprocessed image with the same resolution as the color guide image.

[0028] Step S13: Inputting the color guide image into a guide convolutional neural network to obtain a first feature image.

[0029] Please refer to Figure 2 , the guide convolutional neural network CNN G includes three convolutional modules, each of which includes 16 convolutional layers, and the size of the filter is 5x5. Guidance shown in the figure represents the input color guide image.

[0030] This step performs a convolution operation on the color guide image using the guide convolutional neural network to extract features and output a first feature image.

[0031] Step S14: Inputting the preprocessed image into a target convolutional neural network to obtain a second feature image.

[0032] where the target convolutional neural network CNN TThe three convolution modules each include 16 convolution layers, and the filter size is 5*5. In addition, a channel attention module is added between each two convolution modules. The use of the channel attention mechanism can help the network learn the correlation between channels, adaptively recalibrate the channel feature response, increase important features, and weaken unimportant features, so that the extracted features are more directional.

[0033] In an embodiment, the channel attention module is a SENET (Squeeze-and-Excitation networks). After each convolution, the global average pooling is performed on all channels. Assuming that the current convolution layer is H*W*C, the global average pooling will form a new vector Z c (1*1*C). Z c is input into the feature extraction subnetwork of FCN-ReLU-FCN. The first fully connected layer is used to reduce the calculation amount and reduce the channel number C to , i.e., the size of the current convolution layer is For example, l is taken as 16. After a ReLU activation function, the whole operation is nonlinear. The second fully connected layer is used to restore the channel number to 1*1*C, and then a sigmoid activation layer is used to obtain the channel attention weight distribution after adding the channel attention mechanism. The weight is multiplied by the result after convolution to obtain the input result of the next convolution layer.

[0034] After the target convolution network is added with the channel attention module, the feature image obtained is more clear and accurate in maintaining the edge structure, and the number of convergence iterations is reduced.

[0035] Step S15: obtaining a pixel-enhanced image based on the first feature image, the second feature image, and the to-be-processed image.

[0036] This step uses the first feature image and the second feature image to perform pixel enhancement on the to-be-processed image.

[0037] Please refer to Figure 3 , Figure 3 is a flowchart diagram of an embodiment of step S15 of the present application. It should be noted that the present embodiment is not limited to the flow order shown in Figure 3 if there is substantially the same result. The present embodiment includes the following steps:

[0038] Step S151: inputting the first feature image and the second feature image into a prediction network to output a prediction image.

[0039] , wherein the prediction network CNNF The prediction image is a feature image obtained by fusing the first feature image and the second feature image.

[0040] In an embodiment, the prediction network CNN F may comprise a pixel adaptive convolution module, and the first feature image and the second feature image are fused by using the pixel adaptive convolution module. Specifically, please refer to Figure 4 , Figure 4 is a flowchart diagram of an embodiment of step S151 of the present application. It should be noted that the present embodiment is not limited to the flow order shown in Figure 4 if there is substantially the same result. The present embodiment comprises the following steps:

[0041] S1511: calculating an adaptive convolution kernel of the pixel adaptive convolution module according to the first feature image.

[0042] In an embodiment, the adaptive convolution kernel is determined according to a Gaussian kernel function, and the expression is as follows:

[0043]

[0044] wherein f i , f j are pixel values of corresponding positions of a pixel point on the first feature image in the adaptive convolution kernel, wherein the size of the adaptive convolution kernel is 5x5.

[0045] S1512: performing weighted processing on the second feature image by using the adaptive convolution kernel and the weight matrix of the last layer of the target convolution network, and adding a bias term to the weighted result to obtain a prediction image.

[0046] Specifically, the prediction image can be obtained according to the following formula:

[0047]

[0048] wherein W j is a value of a pixel point on the second feature image at a corresponding position of the weight matrix of the pixel adaptive convolution module, V j is a pixel value of the pixel point on the second feature image, b is a bias term, Ω(i) represents a convolution window size of 5x5 around the pixel i, and v′ i represents a pixel value of a corresponding pixel i position in the prediction image.

[0049] wherein the prediction network further comprises three convolution modules for performing convolution operation on the prediction image to obtain a final prediction image, and the convolution layers of the three convolution modules are 32, 32 and 1 respectively.

[0050] Step S152: obtaining the pixel-enhanced image based on the prediction image and the image to be processed.

[0051] Specifically, the pixel value of each pixel point of the image to be processed is added by a preset multiple of the pixel value of the corresponding pixel point in the prediction image to obtain the pixel-enhanced image.

[0052] The preset multiple is between 0.8 and 1.2. For example, the preset multiple can be 0.8, 1, or 1.2.

[0053] Since the high-frequency information in the upsampling process is obtained through network learning, pixel-level addition with the image to be processed is required, which can well preserve the high-frequency information and low-frequency information of the image and improve the image enhancement effect, and the enhanced image has the characteristics of higher quality and higher accuracy.

[0054] Please refer to Figure 5 , Figure 5 is a structural schematic block diagram of an embodiment of the depth image pixel enhancement device. The depth image pixel enhancement device 100 comprises an acquisition module 110, a preprocessing module 120, a feature extraction module 130, and a pixel enhancement module 140.

[0055] The acquisition module 110 is configured to acquire an image to be processed and a color guide image. The color guide image is a color image of the image to be processed, and the resolution of the color guide image is higher than that of the image to be processed. The preprocessing module 120 is configured to pre-process the image to be processed to obtain a pre-processed image. The feature extraction module 130 is configured to input the pre-processed image into a target convolutional network to obtain a first feature image, and input the color guide image into a guide convolutional network to obtain a second feature image. The pixel enhancement module 140 is configured to obtain a pixel-enhanced image based on the first feature image, the second feature image, and the image to be processed.

[0056] Optionally, the preprocessing module 120 is further configured to determine the resolution of the color guide image, and perform bicubic interpolation on the image to be processed to obtain a pre-processed image with the same resolution as the color guide image.

[0057] Optionally, the pixel enhancement module 140 is further configured to input the first feature image and the second feature image into a prediction network to output a prediction image, and obtain the pixel-enhanced image based on the prediction image and the image to be processed.

[0058] Optionally, the pixel enhancement module 140 is further configured to calculate an adaptive convolution kernel of a pixel adaptive convolution module according to the first feature image, perform weighted processing on the second feature image by using the adaptive convolution kernel and a weight matrix of the last layer of the target convolutional network, and add a bias term to the weighted result to obtain the prediction image.

[0059] Optionally, the pixel enhancement module 140 is further configured to add a preset multiple of the pixel value of the corresponding pixel in the predicted image to the pixel value of the pixel in the image to be processed to obtain the pixel-enhanced image.

[0060] The specific manners of the steps performed by the processor will be described in the steps of the depth image pixel enhancement method embodiments of the present application.

[0061] Referring to Figure 6 , Figure 6 is a structural schematic block diagram of another embodiment of the depth image pixel enhancement device of the present application. The depth image pixel enhancement device 200 includes a processor 210 and a memory 220 coupled to each other. The memory 220 stores a computer program, and the processor 210 is configured to execute the computer program to implement the depth image pixel enhancement method described in the above embodiments.

[0062] The descriptions of the steps performed by the processor will be described in the steps of the depth image pixel enhancement method embodiments of the present application.

[0063] The memory 220 can be configured to store program data and modules. The processor 210 performs various functional applications and data processing by running the program data and modules stored in the memory 220. The memory 220 can mainly include a program storage area and a data storage area. The program storage area can store an operating system, application programs required by at least one function (such as an image preprocessing function, an image feature extraction function, etc.), etc. The data storage area can store data (such as image data, a first feature image, a second feature image, a predicted image, etc.) created according to the use of the depth image pixel enhancement device 200, etc. In addition, the memory 220 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 220 can also include a memory controller to provide the processor 210 with access to the memory 220.

[0064] In the embodiments of the present application, the disclosed method and device can be implemented in other manners. For example, the above-described embodiments of the depth image pixel enhancement device 200 are merely schematic, and the division of the modules or units is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0065] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0066] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0067] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium.

[0068] Reference Figure 7 , Figure 7 The structure of an embodiment of the computer readable storage medium of the present application is shown in the block diagram. The computer readable storage medium 300 stores program data 310, which is executed to realize the steps of each embodiment of the depth image pixel enhancement method described above.

[0069] For the description of the steps of the processing execution, please refer to the description of the steps of the embodiments of the depth image pixel enhancement method of the present application described above, which will not be repeated here.

[0070] The computer readable storage medium 300 can be a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0071] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. A method for enhancing pixels in a depth image, characterized in that, The method includes: Acquire the image to be processed and the color guide image; wherein the color guide image is a color image of the image to be processed, and the resolution of the color guide image is higher than that of the image to be processed; The image to be processed is preprocessed to obtain a preprocessed image; The color guide image is input into the guide convolutional network to obtain the first feature image; The preprocessed image is input into the target convolutional network to obtain the second feature image; wherein the target convolutional network is different from the guided convolutional network; The process of inputting the first feature image and the second feature image into a prediction network and outputting a prediction image includes: calculating the adaptive convolution kernel of the pixel adaptive convolution module based on the first feature image; weighting the second feature image using the adaptive convolution kernel and the weight matrix of the last layer of the target convolutional network, and adding a bias term to the weighting result to obtain the prediction image; wherein the prediction network includes the pixel adaptive convolution module; the prediction network is distinct from the guided convolutional network and the target convolutional network; Based on the predicted image and the image to be processed, a pixel-enhanced image is obtained; The adaptive convolution kernel is determined by a Gaussian kernel function, as follows: in, , Let be the pixel values ​​of the pixels on the first feature image at the corresponding positions in the adaptive convolution kernel, and let be the pixel values ​​of the pixels at the corresponding positions in the adaptive convolution kernel, wherein the size of the adaptive convolution kernel is 5×5; The predicted image is obtained according to the following formula: in, It is the value of the pixel on the second feature image at the corresponding position in the weight matrix of the pixel adaptive convolution module. It is the pixel value of the pixel point on the second feature image. It is a bias term. This represents the size of the 5×5 convolution window surrounding pixel i. This represents the pixel value at position i in the predicted image.

2. The method according to claim 1, characterized in that, The preprocessing of the image to be processed to obtain a preprocessed image includes: Determine the resolution of the color guide image; The image to be processed is subjected to bicubic interpolation to obtain a preprocessed image with the same resolution as the color guide image.

3. The method according to claim 1, characterized in that, The process of obtaining a pixel-enhanced image based on the predicted image and the image to be processed includes: The pixel values ​​of each pixel in the image to be processed are added to a preset multiple of the pixel values ​​of the corresponding pixels in the predicted image to obtain the pixel-enhanced image.

4. The method according to claim 3, characterized in that, The preset multiple is between 0.8 and 1.

2.

5. A depth image pixel enhancement device, characterized in that, The device includes: An acquisition module is used to acquire an image to be processed and a color guide image; wherein the color guide image is a color image of the image to be processed, and the resolution of the color guide image is higher than that of the image to be processed; The preprocessing module is used to preprocess the image to be processed to obtain a preprocessed image; The feature extraction module is used to input the preprocessed image into the target convolutional network to obtain a first feature image, and to input the color guiding image into the guiding convolutional network to obtain a second feature image; wherein the target convolutional network is different from the guiding convolutional network; A pixel enhancement module is used to input a first feature image and a second feature image into a prediction network and output a prediction image. The module includes: calculating an adaptive convolution kernel of the pixel adaptive convolution module based on the first feature image; weighting the second feature image using the adaptive convolution kernel and the weight matrix of the last layer of the target convolutional network, and adding a bias term to the weighting result to obtain the prediction image; and obtaining a pixel-enhanced image based on the prediction image and the image to be processed. The prediction network includes the pixel adaptive convolution module and is distinct from the guided convolutional network and the target convolutional network. The adaptive convolution kernel is determined by a Gaussian kernel function, as follows: in, , Let be the pixel values ​​of the pixels on the first feature image at the corresponding positions in the adaptive convolution kernel, and let be the pixel values ​​of the pixels at the corresponding positions in the adaptive convolution kernel, wherein the size of the adaptive convolution kernel is 5×5; The predicted image is obtained according to the following formula: in, It is the value of the pixel on the second feature image at the corresponding position in the weight matrix of the pixel adaptive convolution module. It is the pixel value of the pixel point on the second feature image. It is a bias term. This represents the size of the 5×5 convolution window surrounding pixel i. This represents the pixel value at position i in the predicted image.

6. The apparatus according to claim 5, characterized in that, The preprocessing module is also used to determine the resolution of the color guide image and perform bicubic interpolation on the image to be processed to obtain a preprocessed image with the same resolution as the color guide image.

7. A depth image pixel enhancement device, characterized in that, The depth image pixel enhancement device includes a processor and a memory coupled to each other; the memory stores a computer program, and the processor executes the computer program to implement the steps of the method as described in any one of claims 1-4.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data that, when executed by a processor, implements the steps of the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Depth image enhancement method and device, electronic equipment and storage medium

    CN112767294A