Image processing method and device based on depth camera, equipment, medium and product
Patent Information
- Application Number
- CN202211091848.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-07
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2042-09-07
AI Technical Summary
[0004]本公开实施例提供一种基于深度相机的图像处理方法、装置、设备、介质及产品,以克服头戴式设备输出的环境视频较为虚假,不清晰的问题
[0022]The technical solution provided in this embodiment can determine the target pixels that meet the pixel completion conditions based on the image to be processed, thus obtaining the target pixels that need to be completed. Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixels can be projected onto the grayscale image, obtaining the target projection area of the target pixels from the depth camera that meet the completion conditions in the grayscale image. Using the target projection area, pixel completion can be performed on each pixel in the grayscale image within that area, obtaining the target pixel values of each pixel in the target projection area. These target pixel values can be used for pixel fusion processing of the image to be processed and the target projection image to obtain a target depth image. By fusing the pixels in the pixel-completed target projection area into the original image to be processed, the image to be processed incorporates the content acquired from the grayscale image, resulting in a more accurate target depth image and improving the image precision and accuracy of the target depth image.
Smart Images

Figure CN117710222B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an image processing method, apparatus, device, medium, and product based on a depth camera. Background Technology
[0002] See-through technology is the core underlying technology of the MR (Mixed Reality) scanning function in head-mounted virtual reality (VR) devices. Essentially, after a user wears the head-mounted device, the device plays a video of the user's actual environment. Therefore, a camera located in front of the head-mounted device is needed to capture this environmental video.
[0003] To ensure a clear shooting area, depth cameras in head-mounted displays are typically positioned at the center. However, in practical applications, since human eyes are generally positioned at the left or right edges of the head-mounted display, environmental videos captured by a center-mounted depth camera often fail to accurately simulate the real environment as perceived by the human eye. This results in environmental videos output by the head-mounted display appearing artificial and unclear. Summary of the Invention
[0004] This disclosure provides an image processing method, apparatus, device, medium, and product based on a depth camera to overcome the problem of artificial and unclear environmental videos output by head-mounted devices.
[0005] In a first aspect, embodiments of this disclosure provide an image processing method based on a depth camera, comprising:
[0006] The image to be processed is acquired by a depth camera and a grayscale image is acquired by a grayscale camera; the grayscale image and the image to be processed are acquired simultaneously.
[0007] Based on the image to be processed, determine the target pixels that meet the pixel completion conditions;
[0008] Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the target projection area;
[0009] Pixel padding is performed on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region;
[0010] Based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain a target depth image.
[0011] In a second aspect, embodiments of this disclosure provide an image processing apparatus based on a depth camera, comprising:
[0012] An image acquisition unit is used to acquire an image to be processed captured by a depth camera and a grayscale image captured by a grayscale camera; the grayscale image and the image to be processed are acquired simultaneously.
[0013] The target determination unit is used to determine the target pixels that meet the pixel completion conditions based on the image to be processed.
[0014] A pixel projection unit is used to project the target pixel onto the grayscale image based on the pixel projection relationship between the depth camera and the grayscale camera to obtain a target projection area.
[0015] A pixel padding unit is used to perform pixel padding processing on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region.
[0016] The image fusion unit is used to perform pixel fusion processing on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image.
[0017] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0018] The memory stores computer-executed instructions;
[0019] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the depth camera-based image processing method as described in the first aspect and various possible designs of the first aspect.
[0020] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image processing method based on a depth camera as described in the first aspect and various possible designs of the first aspect.
[0021] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the image processing method based on a depth camera as described in the first aspect and various possible designs of the first aspect.
[0022] The technical solution provided in this embodiment can determine the target pixels that meet the pixel completion conditions based on the image to be processed, thus obtaining the target pixels that need to be completed. Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixels can be projected onto the grayscale image, obtaining the target projection area of the target pixels from the depth camera that meet the completion conditions in the grayscale image. Using the target projection area, pixel completion can be performed on each pixel in the grayscale image within that area, obtaining the target pixel values of each pixel in the target projection area. These target pixel values can be used for pixel fusion processing of the image to be processed and the target projection image to obtain a target depth image. By fusing the pixels in the pixel-completed target projection area into the original image to be processed, the image to be processed incorporates the content acquired from the grayscale image, resulting in a more accurate target depth image and improving the image precision and accuracy of the target depth image. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 An example diagram illustrating the differences in data acquisition between a depth camera and the human eye, provided as an embodiment of this disclosure;
[0025] Figure 2 An application network architecture diagram of an image processing method based on a depth camera provided in this disclosure embodiment;
[0026] Figure 3 A flowchart illustrating one embodiment of an image processing method based on a depth camera provided in this disclosure;
[0027] Figure 4 A flowchart of yet another embodiment of an image processing method based on depth images provided in this disclosure;
[0028] Figure 5 An example diagram illustrating the association between a local image and various grayscale images provided in an embodiment of this disclosure;
[0029] Figure 6 A flowchart illustrating one embodiment of an image processing method based on depth images provided in this disclosure;
[0030] Figure 7 An example diagram of a pixel projection provided in an embodiment of this disclosure;
[0031] Figure 8 A structural example diagram of an image processing apparatus based on depth images provided in this disclosure embodiment;
[0032] Figure 9 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0034] The technical solution disclosed herein can be applied to head-mounted devices, using grayscale images captured by a grayscale camera to perform image restoration on depth images captured by a depth camera, thereby obtaining clearer target depth images and improving the clarity of environmental videos output by the head-mounted device.
[0035] In related technologies, a depth camera can be placed at the center of the outer surface of a head-mounted device to capture video of the environment using depth images. The head-mounted device can then output the captured environmental video. Currently, the output environmental video is generally blurry, especially when there are obstructions in front of the head-mounted device, which may prevent the depth camera from capturing content on the left and right sides of the obstruction. For example... Figure 1 As shown, assuming the depth camera of the head-mounted device is C1, under normal circumstances, without the head-mounted device, the human eyes E1 and E2 can be located on either side of the depth camera C1. An occlusion B is located on what can be a background object D, directly in front of the depth camera C1. The area that the depth camera C1 can capture is A1. The area that the human eye E1 can see is A2, and the area that the human eye E2 can see is A3. Normally, based on the signals captured by the human eye E1 and E2, the brain synthesizes a signal containing the content of the occlusion itself, including the left side B3 and the occlusion itself. However, due to the presence of occlusion B1, the video or image corresponding to area A2 captured by the depth camera C1 lacks the content of the left side B2 and the right side B3 of the occlusion. Therefore, the video or image captured by the depth camera C1 lacks some information, resulting in an incomplete image and exhibiting false or unclear phenomena.
[0036] To address the aforementioned technical issues, the inventors considered image completion techniques for depth videos captured by depth cameras to obtain clearer depth videos or images. To achieve depth video completion, cameras can be placed at the four corners of a head-mounted device to capture grayscale images in four directions: up, down, left, and right. The depth images captured by the depth camera are also divided into four equal parts. The positions of the grayscale cameras and their corresponding local images in the depth images are correlated to obtain corresponding local images. Based on the pixel projection relationship between the depth camera and the grayscale camera, the corresponding local images are then used to complete the image completion using the grayscale images, ultimately resulting in a clearer target depth image. By using the grayscale images corresponding to the positional relationships to repair the depth image, the content captured by the grayscale images is integrated into the depth image, resulting in a clearer and more accurate target depth image.
[0037] In the embodiments of this disclosure, an image to be processed captured by a depth camera and a grayscale image captured by a grayscale camera can be acquired simultaneously, ensuring the synchronization of image processing. Based on the image to be processed, target pixels that meet the pixel completion conditions can be identified, obtaining the target pixels that need to be completed. According to the pixel projection relationship between the depth camera and the grayscale camera, the target pixels can be projected onto the grayscale image, obtaining the target projection area of the target pixels from the depth camera that meet the completion conditions in the grayscale image. Using the target projection area, pixel completion can be performed on each pixel in the grayscale image within that area, obtaining the target pixel values of each pixel in the target projection area. These target pixel values can be used for pixel fusion processing of the image to be processed and the target projection image to obtain a target depth image. By fusing the pixels in the pixel-completed target projection area into the original image to be processed, the image to be processed incorporates the content captured by the grayscale image, resulting in a more accurate target depth image and improving the image precision and accuracy of the target depth image.
[0038] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0039] Figure 2 This is an application network architecture diagram of an image processing method based on a depth camera according to the present disclosure. The application network architecture according to embodiments of the present disclosure may include an electronic device and a head-mounted device connected to the electronic device via a local area network (LAN) or a wide area network (WAN). It is assumed that the electronic device can be a personal computer, a regular server, a supercomputer, a cloud server, or other similar type of server; the specific type of electronic device is not limited in this disclosure. Figure 2 As shown, the electronic device is a cloud server 1, and the head-mounted device is a VR (Virtual Reality) glasses 2. A depth camera 21 can be set at the center of the front of the VR glasses 2, which can capture environmental video of the area directly in front of the VR glasses 2. In addition, unlike other VR glasses, grayscale cameras 22 can be set at the upper left, upper right, lower left, and lower right corners of the VR glasses 2 disclosed in this invention.
[0040] In this system, depth image 21 is used to acquire the image to be processed, and grayscale camera 22 is used to acquire grayscale images. Both the image to be processed and the grayscale images can be sent to electronic device 1. Electronic device 1 can be configured with the image processing method based on the depth camera disclosed herein, which maps a local image of the image to be processed to the grayscale image, and uses the image processing method disclosed herein to perform fill-in processing on the corresponding local image using the grayscale image to obtain the target depth image. The target depth image can be fed back from electronic device 1 to VR glasses 2, which then outputs a clearer target depth image or a target environment video corresponding to the target depth image to the user.
[0041] refer to Figure 3 , Figure 3 A flowchart illustrating one embodiment of an image processing method based on a depth camera provided in this disclosure. This method can be configured as an image processing device based on a depth camera, which can be located in an electronic device. The image processing method may include the following steps:
[0042] 301: Acquire the image to be processed from the depth camera and the grayscale image from the grayscale camera; the grayscale image and the image to be processed are acquired simultaneously.
[0043] Optionally, a depth camera can be used to acquire depth video, and a grayscale camera can be used to acquire grayscale video. The image to be processed can be extracted from the depth video. The grayscale image can be extracted from the grayscale video. If the timestamps of the image to be processed and the grayscale image are the same, the grayscale image and the image to be processed are acquired simultaneously.
[0044] Optionally, the electronic device can be a head-mounted device including a depth camera and a grayscale camera. The electronic device may also include a server that establishes a communication connection with the head-mounted device. The electronic device can be configured with the image processing method based on the depth camera disclosed herein to perform image processing on the image to be processed acquired by the depth camera and the grayscale image acquired by the grayscale camera to obtain a target depth image.
[0045] 302: Based on the image to be processed, determine the target pixels that meet the pixel completion conditions.
[0046] The target pixels can be determined based on the image to be processed.
[0047] 303: Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the target projection area.
[0048] Pixel projection relationship refers to the pixel projection function between pixels of a depth camera and a grayscale camera; it's a pixel mapping relationship. Projecting a target pixel onto a grayscale image yields its projected pixel in the grayscale image. The coordinates of the target pixel are projected onto the grayscale image, and the resulting projection coordinates are the pixel points in the grayscale image. Projected pixels include both projection coordinates and pixel values. These pixel values can be assigned the color values of the target pixel.
[0049] 304: Perform pixel padding on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region.
[0050] Pixel padding can refer to the process of padding the depth and color of individual pixels. The target pixel value can include the target depth value and the target color value.
[0051] 304: Based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain the target depth image.
[0052] Optionally, after obtaining the target depth image, the process also includes outputting the target depth video corresponding to the target depth image. The target depth video can be the repaired environmental video, with a clearer picture.
[0053] The image to be processed can be a depth map, which can be a matrix composed of all depth data within the range acquired by the depth camera, with the unit being millimeters. The depth map can include color values and depth values represented in RGB color space.
[0054] The target depth image can also be converted into a target point cloud, which is a data matrix composed of point cloud information of all points captured by the depth camera. The point cloud can be recorded as three-dimensional coordinates (X, Y, Z), where x, y, and z can be float values. The target point cloud can be the coordinate points of the image to be processed after completion, recorded in point cloud image format. Specifically, the conversion from a target depth image to a target point cloud can be achieved through the conversion principles between depth images and point clouds.
[0055] Optionally, after obtaining the target depth map and / or target point cloud, the corresponding environmental video can also be output through the target depth map and / or target point cloud to achieve a more comprehensive environmental display.
[0056] In this embodiment, the image to be processed acquired by a depth camera and the grayscale image acquired by a grayscale camera can be obtained simultaneously, ensuring the synchronization of image processing. Based on the image to be processed, target pixels that meet the pixel completion conditions can be identified, and the target pixels that need to be completed can be obtained. According to the pixel projection relationship between the depth camera and the grayscale camera, the target pixels can be projected onto the grayscale image, obtaining the target projection area of the target pixels from the depth camera that meet the completion conditions in the grayscale image. Using the target projection area, pixel completion can be performed on each pixel in the grayscale image in that area, obtaining the target pixel value of each pixel in the target projection area. The target pixel values of each pixel in the target projection area can be used for pixel fusion processing of the image to be processed and the target projection image to obtain the target depth image. By fusing the pixels in the target projection area after pixel completion into the original image to be processed, the content acquired by the grayscale image is incorporated into the image to be processed, resulting in a more accurate target depth image and improving the image precision and accuracy of the target depth image.
[0057] As an example, such as Figure 4 The flowchart shown is an embodiment of an image processing method based on depth images provided in this disclosure. The difference from the previous embodiments lies in that, based on the image to be processed, the target pixel points that satisfy the pixel padding conditions are determined, including:
[0058] 401: Based on the positional correspondence between the grayscale camera and the depth camera, determine the target local image associated with the grayscale image in the image to be processed.
[0059] Optionally, the target local image associated with the grayscale image in the image to be processed can determine the position of the first camera of the grayscale camera that acquired the grayscale image, map the position of the first camera from the spatial coordinate system to the image coordinate system of the image to be processed, and obtain the image mapping position; divide the image to be processed acquired by the depth camera into at least one local image, and determine the image region of each local image in the image to be processed; determine the target image region where the image mapping position is located from the image regions of each local image; and determine the local image corresponding to the target image region as the target local image.
[0060] The process of dividing the image to be processed acquired by the depth camera into at least one local image may include: dividing the image to be processed into an upper left local image, an upper right local image, a lower left local image, and a lower right local image according to the horizontal central axis and the vertical central axis.
[0061] For ease of understanding, Figure 5An example diagram showing the association between local images and various grayscale images is illustrated. The image to be processed 501 can be divided into a top-left local image 5011, a top-right local image 5012, a bottom-left local image 5013, and a bottom-right local image 5014. Assume that the grayscale cameras located on the rectangular surface 502 of the head-mounted device are grayscale camera 5021 located in the top-left corner, grayscale camera 5022 located in the top-right corner, grayscale camera 5023 located in the bottom-left corner, and grayscale camera 5024 located in the bottom-right corner. (Reference) Figure 5 The depth camera 5025 can be located at the center of the rectangular surface 502. Based on the positional relationship between each grayscale camera and the depth camera, it can be determined that the grayscale image acquired by the grayscale camera 5021 is associated with the upper left local image 5011, the grayscale image acquired by the grayscale camera 5022 is associated with the upper right local image 5012, the grayscale image acquired by the grayscale camera 5023 is associated with the lower left local image 5013, and the grayscale image acquired by the grayscale camera 5024 is associated with the lower right local image 5014.
[0062] Determining the position of each local image in the first image of the image to be processed includes: determining the upper left image region corresponding to the upper left local image as its image region, determining the upper right image region corresponding to the upper right local image as its image region, determining the lower left image region corresponding to the lower left local image as its image region, and determining the lower right image region corresponding to the lower right local image as its image region.
[0063] 402: Extract the first depth image from the local image of the target.
[0064] Optionally, the image to be processed can be an RGBD (Red, Green, Blue Depth) image acquired by a depth camera, and the target local image is a portion of the image to be processed. Image extraction methods can be used to extract a first depth image and a first color image from the target local image. The first depth image is an image containing depth data. The first color image is an image containing RGB (Red, Green, Blue) data.
[0065] 403: Based on the first depth image, determine the target pixel that meets the pixel change condition.
[0066] Optionally, the pixel change condition can refer to pixels whose pixel modulus change is greater than the modulus threshold.
[0067] In this embodiment, based on the positional correspondence between the grayscale camera and the depth camera, a target local image associated with the grayscale image in the image to be processed can be determined, and an association relationship between the local image in the image to be processed and the grayscale image can be established. For the local target image with the association relationship, a first depth image is extracted to accurately obtain the target pixels in the first depth image that meet the pixel change conditions, thereby achieving accurate acquisition of the pixels in the depth image that need to be padded, and thus improving the accuracy of image padding.
[0068] To obtain accurate target pixels, as an example, based on a first depth image, target pixels that satisfy pixel change conditions are determined, including:
[0069] Calculate the target gradient image corresponding to the first depth image; the target gradient image includes the magnitude value of each pixel.
[0070] Based on a preset modulus threshold, determine the center pixels whose modulus values are greater than the modulus threshold from the target gradient image;
[0071] A preset gradient selection region is determined with the center pixel as the center, and the associated pixels located in the gradient selection region and associated with the center pixel are obtained;
[0072] The target pixel is determined by identifying the center pixel and the associated pixels of each center pixel in the local image.
[0073] Optionally, determining the center pixel with a magnitude value greater than the magnitude threshold from the target gradient image according to the preset magnitude threshold may include comparing the magnitude value of each pixel in the target gradient image with the magnitude threshold. If the magnitude value of any pixel is greater than the magnitude threshold, then the pixel is determined as the center pixel, until all pixels in the target gradient image have been traversed to obtain the obtained center pixel.
[0074] A gradient selection region can refer to a pixel area defined around a central pixel. For example, it can be a 3x3 selection region, where 3 represents the number of pixels in the width and height, and the central pixel is the center of the 3x3 pixel region. Associated pixels can be any pixels within the gradient selection region other than the central pixel.
[0075] In this embodiment, a target gradient image can be calculated on a first depth image to obtain the modulus value of each pixel in the first depth image. The modulus value can characterize the amount of change in a pixel. Using a preset modulus threshold, center pixels with modulus values greater than the threshold can be acquired, thus obtaining the center pixels in the first depth image where pixel values change. Using the center pixels as the center, a gradient selection region can be determined, and associated pixels related to the center pixels can be acquired, thereby expanding the pixel range. The acquired center pixels and associated pixels are used as target pixels. By selecting the modulus value, accurate extraction of pixels with large modulus values and their surrounding pixels can be achieved, improving the effectiveness of target pixel selection.
[0076] In one possible design, calculating the target gradient image corresponding to the first depth image includes:
[0077] Based on the image denoising algorithm, the first depth image is denoised to obtain the second depth image;
[0078] Based on the image magnitude algorithm, the gradient magnitude of the second depth image is calculated to obtain the target gradient image.
[0079] Optionally, the image denoising algorithm may include wavelet denoising algorithm, bilateral filtering algorithm, etc. In this embodiment, the specific type of image denoising algorithm is not limited.
[0080] Optionally, the image modulus algorithm may include the modulus calculation algorithm corresponding to the Sobel operator, machine learning algorithm, etc. In this embodiment, the specific type of image modulus algorithm is not limited.
[0081] In this embodiment, an image denoising algorithm is first used to denoise the first depth image to obtain a second depth image, achieving accurate image denoising. Then, an image modulus algorithm is used to calculate the gradient modulus of the second depth image to obtain a target gradient image, achieving accurate gradient calculation. Through the image denoising algorithm and the image modulus algorithm, accurate gradient calculation can be performed on the first depth image, resulting in a more accurate target gradient image.
[0082] like Figure 6 The flowchart shown is a further embodiment of an image processing method based on depth images provided by this disclosure. The difference from the previous embodiments lies in that pixel padding processing is performed on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region. This may include:
[0083] 601: Extract the first color image from the target local image.
[0084] 602: Based on the color value of the target pixel in the first color image, perform color completion processing on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area.
[0085] 603: Calculate the depth value of each pixel in the grayscale image within the target projection area to obtain the target depth value of each pixel in the target projection area.
[0086] 604: Determine the target color value and target depth value of each pixel in the target projection area as the corresponding target pixel value.
[0087] Optionally, the image format of the first color image can be RGB (Red, Green, Blue, the three primary colors). The color value of the target pixel in the first color image can refer to the color value obtained at the target pixel coordinates in the first color image. The color value can refer to the values of the three primary colors of RGB. For example, the color value can be recorded as (R=121, G=134, B=89).
[0088] The target pixel values of each pixel in the target projection area can include the target color value and the target depth value.
[0089] In this embodiment of the disclosure, a first color image can be extracted from a local image of the target. Based on the color values of the target pixels in the first color image, color completion processing can be performed on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area. This achieves color completion for each pixel. Furthermore, depth values can be calculated for each pixel in the grayscale image within the target projection area to obtain the target depth value of each pixel. By obtaining the target color value and target depth value of each pixel in the target projection area as the target pixel value, comprehensive completion of each pixel in terms of both color and depth is achieved, resulting in more comprehensive target pixel values. This improves the accuracy and comprehensiveness of the target pixel values, thereby enhancing the accuracy and precision of the target depth image completion.
[0090] In one possible design, depth values are calculated for each pixel in the grayscale image within the target projection region to obtain the target depth value for each pixel in the target projection region, including:
[0091] Based on the image depth algorithm, the depth value of each pixel in the target projection area of the grayscale image is calculated to obtain the target depth value of each pixel in the target projection area.
[0092] Optionally, image depth algorithms may include Laplacian algorithms, machine learning algorithms, etc. This embodiment does not impose excessive limitations on the specific type of image depth algorithm.
[0093] In this embodiment, an image depth algorithm is used to extract the target depth value of each pixel in the target projection area, thereby improving the extraction efficiency and accuracy of the depth value.
[0094] As one embodiment, based on the color value of the target pixel in the first color image, color completion processing is performed on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area, including:
[0095] Determine the projected pixel corresponding to the target pixel in the target projection area;
[0096] The color value of the target pixel in the first color image is determined to be the first color value of the projected pixel;
[0097] Based on the image coloring algorithm, the first color value of the projected pixel is used to perform color filling processing on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area.
[0098] The target pixels are projected onto the grayscale image to obtain the projected pixels. These projected pixels are then used as contour points to form the target projection region. The projected pixels located within the target projection region are then identified.
[0099] Projecting the target pixel onto a grayscale image yields the projected pixel, and the region formed by connecting these projected pixels is the target projection region. The color value of the target pixel in the first color image can be used as the color value of the corresponding projected pixel.
[0100] The pixel in the target projection region, excluding the pixel projection point, is the center point of the region.
[0101] Based on an image coloring algorithm, the color of the center point of the region can be filled in by combining the first color value of the projected pixels to obtain the second color value of the center point. The target color value of each pixel in the target projection region includes the first color value corresponding to the projected pixel and the second color value corresponding to the center point of the region.
[0102] Optionally, the image coloring algorithm may include the coloring problem solving algorithm involved in The Fast Bilateral Solver, the Poisson editing algorithm, etc. This embodiment does not impose excessive limitations on the specific type of image coloring algorithm.
[0103] In this embodiment, an image coloring algorithm is used to perform depth calculation on the target color of each pixel in the target projection area, so as to accurately obtain the target color value.
[0104] As one embodiment, the difference from the previous embodiments is that the grayscale camera includes at least two; based on the positional correspondence between the grayscale camera and the depth camera, determining the target local image associated with the grayscale image in the image to be processed includes:
[0105] Determine at least two grayscale cameras to be installed on the head-mounted device and determine the camera position of each grayscale camera in the image coordinate system of the image to be processed; the grayscale cameras are evenly distributed on the left and right edges of the head-mounted device.
[0106] The image to be processed is divided into at least two local images along at least one horizontal axis and / or at least one vertical axis, and the local image region of each local image is determined; the number of local images is equal to the number of grayscale cameras.
[0107] Based on the camera position of each grayscale camera and the local image region of each local image, determine the target local image region that matches the camera position of each grayscale camera.
[0108] The local image corresponding to the target local image region matched with the grayscale camera is determined as the target local image associated with the grayscale camera, so as to obtain the target local image associated with the grayscale image of each grayscale camera.
[0109] Optionally, the number of local images divided into the image to be processed is equal to the number of grayscale cameras, for example, both can be 4. The local images and the grayscale images acquired by the grayscale cameras can have a one-to-one mapping relationship.
[0110] In this embodiment, multiple grayscale cameras can be evenly arranged on the left and right sides of the head-mounted device. After the image to be processed is evenly segmented horizontally and / or vertically, the segmented local image and the grayscale image acquired by the grayscale camera can be associated using the camera position of the grayscale camera in the image coordinate system. This achieves a closer combination of the target local image and the grayscale image in terms of position, and allows for more accurate image completion of the target local image with positional association using the grayscale image. This integrates the relevant content acquired by the grayscale image into the image to be processed at the corresponding position, improving the accuracy and precision of image completion.
[0111] In some embodiments, based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain a target depth image, including:
[0112] Based on the target pixel values of each pixel in the target projection region corresponding to each grayscale camera, pixel fusion processing is performed sequentially on the image to be processed and the target projection region corresponding to each grayscale camera to obtain the target depth image.
[0113] In this embodiment, for multiple grayscale cameras, each grayscale camera can sequentially perform pixel fusion processing on the image to be processed to obtain a target depth image that is filled in by each grayscale camera in each direction or angle, thereby obtaining a target depth image with higher and more comprehensive filling effect.
[0114] As one embodiment, based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the target projection area, including:
[0115] Based on the pixel projection functions corresponding to the depth camera and the grayscale camera, the target pixel coordinates of the target pixel are projected onto the grayscale image to obtain the projected pixel.
[0116] The target projection area is determined based on the outline of the region formed by the projected pixels.
[0117] Alternatively, the pixel projection function can be expressed using the following formula:
[0118] p i =π(T) id D(p d )π -1 (p d ))
[0119] Where i represents the camera identifier, T id p is the extrinsic parameter of the depth camera. i =π(P) represents projecting the coordinates of a point in three-dimensional space onto the grayscale image of the i-th camera. This represents projecting the pixels in the image captured by the depth camera onto the Z=1 plane in the depth coordinate system, D(p d ) indicates that the pixel coordinate is p d The depth value.
[0120] For ease of understanding, such as Figure 7 The diagram shows an example of pixel projection. Assuming the target pixels captured in the depth image at points P1 and P2 are 701 and 702, projecting target pixel 701 onto the grayscale image yields projected pixel 703, and projecting target pixel 702 onto the grayscale image yields projected pixel 704. Target pixels 701 and 702 are two adjacent pixels, while projected pixels 703 and 704 are pixels with a certain pixel interval, forming a line segment between them. Therefore, the contour of the region corresponding to the projected pixels can form the target projection region.
[0121] In this embodiment, the target pixel coordinates of the target pixel can be projected onto a grayscale image using a pixel projection function to obtain the projected pixel. Based on the contour of the region corresponding to the projected pixel, the target projection region can be determined, thereby achieving accurate acquisition of the target projection region.
[0122] In one possible design, based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain a target depth image, including:
[0123] Based on the target pixel values of each pixel in the target projection region, the target color value and target depth value of each pixel in the target projection region are fused to obtain the target region image.
[0124] The target region image is stitched to the image to be processed to obtain the target depth image.
[0125] Optionally, fusing the target color value and target depth value of each pixel in the target projection area to obtain the target area image can specifically refer to fusing the target color value and target depth value in the target pixel value of each pixel in the target projection area to obtain depth pixels containing the target color value and depth value. The depth pixels corresponding to the target projection area can constitute the target area image.
[0126] Optionally, stitching the target region image to the image to be processed to obtain a target depth image may include: replacing the local depth image corresponding to the target projection region of the image to be processed with the target region image, and stitching the target region image with other region images to obtain the target depth image. Optionally, the other region images may be target region images corresponding to the target projection regions of other grayscale cameras. For example, when there are four grayscale cameras, the image to be processed is divided into four parts, and the target region images corresponding to the four grayscale images can be stitched together according to their target projection regions to obtain the target depth image.
[0127] Optionally, the target region image can be directly converted into a triangular mesh format to obtain a target depth image in the form of a triangular mesh, and then the target depth image in the triangular mesh format can be rendered.
[0128] In this embodiment of the disclosure, the target color value and target depth value of each pixel in the target projection area can be fused based on the target pixel value of each pixel in the target projection area to obtain a target area image. The target area image is a depth pixel containing the target depth value and target color value. By stitching the target area image to the target projection area corresponding to the image to be processed, the depth image of the image to be processed can be directly replaced and stitched to obtain the target depth image. The efficiency and accuracy of obtaining the target depth image can be improved by image replacement and stitching.
[0129] like Figure 8The diagram shown is a structural schematic of an embodiment of an image processing device based on a depth camera provided in this disclosure. This device can be configured with the above-described method and can be located in an electronic device. The image processing device 800 may include:
[0130] Image acquisition unit 801: used to acquire the image to be processed captured by the depth camera and the grayscale image captured by the grayscale camera; the grayscale image and the image to be processed are acquired simultaneously.
[0131] Target determination unit 802: used to determine target pixels that meet the pixel completion conditions based on the image to be processed;
[0132] Pixel projection unit 803: used to project target pixels onto a grayscale image based on the pixel projection relationship between the depth camera and the grayscale camera to obtain the target projection area;
[0133] Pixel padding unit 804: Used to perform pixel padding processing on each pixel in the target projection area of the grayscale image to obtain the target pixel value of each pixel in the target projection area.
[0134] Image fusion unit 805: Used to perform pixel fusion processing on the image to be processed and the target projection area based on the target pixel values of each pixel in the target projection area to obtain a target depth image.
[0135] As one embodiment, the target determination unit includes:
[0136] The image association module is used to determine the target local image of the grayscale image in the image to be processed based on the positional correspondence between the grayscale camera and the depth camera.
[0137] The first extraction module is used to extract a first depth image from the target local image;
[0138] The pixel selection module is used to determine target pixels that meet the pixel change conditions based on the first depth image.
[0139] In one possible design, the pixel selection module includes:
[0140] The gradient calculation submodule is used to calculate the target gradient image corresponding to the first depth image; the target gradient image includes the magnitude value of each pixel.
[0141] The threshold comparison submodule is used to determine the center pixel point in the target gradient image whose modulus value is greater than the modulus threshold based on the preset modulus threshold.
[0142] The associated acquisition submodule is used to determine a preset gradient selection area with the center pixel as the center, and obtain the associated pixels located in the gradient selection area that are associated with the center pixel;
[0143] The target determination submodule is used to determine the center pixels and associated pixels of each center pixel in the local image of the target as the target pixels.
[0144] In some embodiments, the gradient calculation submodule may specifically be used for:
[0145] Based on the image denoising algorithm, the first depth image is denoised to obtain the second depth image;
[0146] Based on the image magnitude algorithm, the gradient magnitude of the second depth image is calculated to obtain the target gradient image.
[0147] In some embodiments, the pixel padding unit may include:
[0148] The first extraction module is used to extract a first color image from the target local image;
[0149] The color completion module is used to perform color completion processing on each pixel in the target projection area based on the color value of the target pixel in the first color image, so as to obtain the target color value of each pixel in the target projection area.
[0150] The depth completion module is used to calculate the depth value of each pixel in the target projection area of the grayscale image, and obtain the target depth value of each pixel in the target projection area.
[0151] The target determination module is used to determine the target color value and target depth value of each pixel in the target projection area, which are the corresponding target pixel values.
[0152] As an optional implementation, the depth completion module includes:
[0153] The depth calculation submodule is used to calculate the depth value of each pixel in the target projection area of the grayscale image based on the image depth algorithm, so as to obtain the target depth value of each pixel in the target projection area.
[0154] In some embodiments, the color completion module includes:
[0155] The projection determination submodule is used to determine the projected pixel point corresponding to the target pixel point in the target projection area;
[0156] The color determination submodule is used to determine the color value of the target pixel in the first color image as the first color value of the projected pixel;
[0157] The region coloring submodule is used to perform color completion processing on each pixel in the target projection region based on the image coloring algorithm and using the first color value of the projected pixel to obtain the target color value of each pixel in the target projection region.
[0158] In one possible design, the grayscale camera includes at least two; the image association module may include:
[0159] The first determining submodule is used to determine at least two grayscale cameras set on the head-mounted device and determine the camera position of each grayscale camera in the image coordinate system of the image to be processed; each grayscale camera is evenly set on the left and right edges of the head-mounted device.
[0160] The second determining submodule is used to divide the image to be processed into at least two local images according to at least one horizontal axis and / or at least one vertical axis, and to determine the local image region of each local image; the number of local images is equal to the number of grayscale cameras;
[0161] The third determination submodule is used to determine the target local image region that matches the camera position of each grayscale camera based on the camera position of each grayscale camera and the local image region of each local image.
[0162] The fourth determination submodule is used to determine the local image corresponding to the target local image region matched with the grayscale camera as the target local image associated with the grayscale camera, so as to obtain the target local image associated with the grayscale image of each grayscale camera.
[0163] In one possible design, the image fusion unit includes:
[0164] The first fusion module is used to perform pixel fusion processing on the image to be processed and the target projection areas corresponding to each grayscale camera in sequence, based on the target pixel values of each pixel in the target projection area corresponding to each grayscale camera, to obtain a target depth image.
[0165] As another embodiment, the pixel projection unit includes:
[0166] The function projection module is used to project target pixels onto a grayscale image based on the pixel projection functions corresponding to the depth camera and grayscale camera, thereby obtaining the projected pixels.
[0167] The region selection module is used to determine the target projection region based on the region outline formed by the projected pixels.
[0168] As another embodiment, the image fusion unit includes:
[0169] The second fusion module is used to perform fusion processing on the target color value and target depth value of each pixel in the target projection area based on the target pixel value of each pixel in the target projection area to obtain the target area image.
[0170] The image stitching module is used to stitch the target region image to the image to be processed to obtain the target depth image.
[0171] The apparatus provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.
[0172] To implement the above embodiments, this disclosure also provides an electronic device.
[0173] refer to Figure 9 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0174] like Figure 9 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0175] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0176] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0177] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0178] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0179] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0180] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0181] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0182] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0183] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0184] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0185] In a first aspect, according to one or more embodiments of this disclosure, an image processing method based on a depth camera is provided, comprising:
[0186] Acquire the image to be processed captured by the depth camera and the grayscale image captured by the grayscale camera; acquire the grayscale image and the image to be processed simultaneously.
[0187] Based on the image to be processed, determine the target pixels that meet the pixel completion conditions;
[0188] Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the target projection area.
[0189] Pixel padding is performed on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region.
[0190] Based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain the target depth image.
[0191] According to one or more embodiments of this disclosure, determining target pixels that satisfy pixel completion conditions based on the image to be processed includes:
[0192] Based on the positional correspondence between the grayscale camera and the depth camera, the target local image associated with the grayscale image in the image to be processed is determined;
[0193] Extract the first depth image from the local image of the target;
[0194] Based on the first depth image, target pixels that meet the pixel change conditions are identified.
[0195] According to one or more embodiments of this disclosure, determining target pixels that satisfy pixel change conditions based on a first depth image includes:
[0196] Calculate the target gradient image corresponding to the first depth image; the target gradient image includes the magnitude value of each pixel.
[0197] Based on a preset modulus threshold, determine the center pixels whose modulus values are greater than the modulus threshold from the target gradient image;
[0198] A preset gradient selection region is determined with the center pixel as the center, and the associated pixels located in the gradient selection region and associated with the center pixel are obtained;
[0199] The target pixel is determined by identifying the center pixel and the associated pixels of each center pixel in the local image.
[0200] According to one or more embodiments of this disclosure, calculating a target gradient image corresponding to a first depth image includes:
[0201] Based on the image denoising algorithm, the first depth image is denoised to obtain the second depth image;
[0202] Based on the image magnitude algorithm, the gradient magnitude of the second depth image is calculated to obtain the target gradient image.
[0203] According to one or more embodiments of this disclosure, pixel padding processing is performed on each pixel in the target projection region of a grayscale image to obtain the target pixel value of each pixel in the target projection region, including:
[0204] Extract the first color image from the target local image;
[0205] Based on the color value of the target pixel in the first color image, color completion processing is performed on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area.
[0206] Calculate the depth value of each pixel in the grayscale image within the target projection area to obtain the target depth value of each pixel in the target projection area.
[0207] Determine the target color value and target depth value of each pixel in the target projection area to be the corresponding target pixel value.
[0208] According to one or more embodiments of this disclosure, depth value calculation is performed on each pixel of a grayscale image in a target projection region to obtain the target depth value of each pixel in the target projection region, including:
[0209] Based on the image depth algorithm, the depth value of each pixel in the target projection area of the grayscale image is calculated to obtain the target depth value of each pixel in the target projection area.
[0210] According to one or more embodiments of this disclosure, based on the color value of the target pixel in a first color image, color completion processing is performed on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area, including:
[0211] Determine the projected pixel corresponding to the target pixel in the target projection area;
[0212] The color value of the target pixel in the first color image is determined to be the first color value of the projected pixel;
[0213] Based on the image coloring algorithm, the first color value of the projected pixel is used to perform color filling processing on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area.
[0214] According to one or more embodiments of this disclosure, the grayscale camera includes at least two;
[0215] Based on the positional correspondence between the grayscale camera and the depth camera, the target local image associated with the grayscale image in the image to be processed is determined, including:
[0216] Determine at least two grayscale cameras to be installed on the head-mounted device and determine the camera position of each grayscale camera in the image coordinate system of the image to be processed; the grayscale cameras are evenly distributed on the left and right edges of the head-mounted device.
[0217] The image to be processed is divided into at least two local images along at least one horizontal axis and / or at least one vertical axis, and the local image region of each local image is determined; the number of local images is equal to the number of grayscale cameras.
[0218] Based on the camera position of each grayscale camera and the local image region of each local image, determine the target local image region that matches the camera position of each grayscale camera.
[0219] The local image corresponding to the target local image region matched with the grayscale camera is determined as the target local image associated with the grayscale camera, so as to obtain the target local image associated with the grayscale image of each grayscale camera.
[0220] According to one or more embodiments of this disclosure, a pixel fusion process is performed on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image, including:
[0221] Based on the target pixel values of each pixel in the target projection region corresponding to each grayscale camera, pixel fusion processing is performed sequentially on the image to be processed and the target projection region corresponding to each grayscale camera to obtain the target depth image.
[0222] According to one or more embodiments of this disclosure, based on the pixel projection relationship between a depth camera and a grayscale camera, a target pixel is projected onto a grayscale image to obtain a target projection area, including:
[0223] Based on the pixel projection functions corresponding to the depth camera and grayscale camera, the target pixel is projected onto the grayscale image to obtain the projected pixel.
[0224] The target projection area is determined based on the outline of the region formed by the projected pixels.
[0225] According to one or more embodiments of this disclosure, a pixel fusion process is performed on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image, including:
[0226] Based on the target pixel values of each pixel in the target projection region, the target color value and target depth value of each pixel in the target projection region are fused to obtain the target region image.
[0227] The target region image is stitched to the image to be processed to obtain the target depth image.
[0228] Secondly, according to one or more embodiments of this disclosure, an image processing apparatus based on a depth camera is provided, comprising:
[0229] The image acquisition unit is used to acquire the image to be processed captured by the depth camera and the grayscale image captured by the grayscale camera; the grayscale image and the image to be processed are acquired simultaneously.
[0230] The target determination unit is used to determine the target pixels that meet the pixel completion conditions based on the image to be processed.
[0231] The pixel projection unit is used to project the target pixel onto the grayscale image based on the pixel projection relationship between the depth camera and the grayscale camera to obtain the target projection area.
[0232] The pixel padding unit is used to perform pixel padding processing on each pixel in the target projection area of the grayscale image to obtain the target pixel value of each pixel in the target projection area.
[0233] The image fusion unit is used to perform pixel fusion processing on the image to be processed and the target projection area based on the target pixel values of each pixel in the target projection area to obtain a target depth image.
[0234] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0235] The memory stores the instructions that the computer executes;
[0236] At least one processor executes computer execution instructions stored in memory, causing at least one processor to perform the depth camera-based image processing method as described in the first aspect above and various possible designs of the first aspect.
[0237] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions that, when executed by a processor, implement the image processing method based on a depth camera as described in the first aspect and various possible designs of the first aspect.
[0238] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the image processing method based on a depth camera as described in the first aspect above and various possible designs of the first aspect.
[0239] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0240] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0241] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method of image processing based on a depth camera, characterized in that, include: The image to be processed is acquired by a depth camera and a grayscale image is acquired by a grayscale camera; the grayscale image and the image to be processed are acquired simultaneously. Based on the positional correspondence between the grayscale camera and the depth camera, the target local image associated with the grayscale image in the image to be processed is determined; a first depth image is extracted from the target local image; Calculate the target gradient image corresponding to the first depth image; the target gradient image includes the magnitude value of each pixel. Based on a preset modulus threshold, determine the center pixels in the target gradient image whose modulus value is greater than the modulus threshold. A preset gradient selection region is determined with the center pixel as the center, and associated pixels located in the gradient selection region that are associated with the center pixel are obtained; The central pixel and the associated pixels associated with each central pixel in the target local image are determined as target pixels; Based on the pixel projection relationship between the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the target projection area; Pixel padding is performed on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region; Based on the target pixel values of each pixel in the target projection region, pixel fusion processing is performed on the image to be processed and the target projection region to obtain a target depth image.
2. The method of claim 1, wherein, The calculation of the target gradient image corresponding to the first depth image includes: Based on the image denoising algorithm, the first depth image is denoised to obtain the second depth image; Based on the image magnitude algorithm, the gradient magnitude of the second depth image is calculated to obtain the target gradient image.
3. The method according to claim 1, characterized in that, The step of performing pixel padding processing on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region includes: Extract a first color image from the target local image; Based on the color value of the target pixel in the first color image, color completion processing is performed on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area. The depth value of each pixel in the grayscale image in the target projection area is calculated to obtain the target depth value of each pixel in the target projection area. The target color value and target depth value of each pixel in the target projection area are determined to be the corresponding target pixel value.
4. The method according to claim 3, characterized in that, The step of calculating the depth value of each pixel in the grayscale image in the target projection area to obtain the target depth value of each pixel in the target projection area includes: Based on the image depth algorithm, the grayscale values of each pixel in the grayscale image in the target projection area are calculated to obtain the target depth value of each pixel in the target projection area.
5. The method according to claim 3, characterized in that, The step of performing color completion processing on each pixel in the target projection area based on the color value of the target pixel in the first color image to obtain the target color value of each pixel in the target projection area includes: Determine the projection pixel corresponding to the target pixel in the target projection area; The color value of the target pixel in the first color image is determined to be the first color value of the projected pixel; Based on the image coloring algorithm, the first color value of the projected pixel is used to perform color filling processing on each pixel in the target projection area to obtain the target color value of each pixel in the target projection area.
6. The method according to claim 1, characterized in that, The grayscale camera includes at least two; The step of determining the target local image associated with the grayscale image in the image to be processed based on the positional correspondence between the grayscale camera and the depth camera includes: Determine at least two grayscale cameras to be installed on the head-mounted device and determine the camera position of each grayscale camera in the image coordinate system of the image to be processed; the grayscale cameras are evenly distributed on the left and right edges of the head-mounted device. The image to be processed is divided into at least two local images according to at least one horizontal axis and / or at least one vertical axis, and the local image region of each local image is determined; the number of local images is equal to the number of grayscale cameras. Based on the camera position of each grayscale camera and the local image region of each local image, determine the target local image region that matches the camera position of each grayscale camera. The local image corresponding to the target local image region matched with the grayscale camera is determined as the target local image associated with the grayscale camera, so as to obtain the target local image associated with the grayscale image of each grayscale camera.
7. The method according to claim 6, characterized in that, The step of performing pixel fusion processing on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image includes: Based on the target pixel values of each pixel in the target projection region corresponding to each grayscale camera, pixel fusion processing is performed sequentially on the image to be processed and the target projection region corresponding to each grayscale camera to obtain the target depth image.
8. The method according to any one of claims 1-7, characterized in that, The step of projecting the target pixel onto the grayscale image based on the pixel projection relationship between the depth camera and the grayscale camera to obtain the target projection region includes: Based on the pixel projection functions corresponding to the depth camera and the grayscale camera, the target pixel is projected onto the grayscale image to obtain the projected pixel. The target projection area is determined based on the region outline formed by the projected pixels.
9. The method according to any one of claims 1-7, characterized in that, The step of performing pixel fusion processing on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image includes: Based on the target pixel values of each pixel in the target projection region, the target color value and target depth value of each pixel in the target projection region are fused to obtain the target region image. The target region image is stitched to the image to be processed to obtain a target depth image.
10. An image processing device based on a depth camera, characterized in that, include: An image acquisition unit is used to acquire an image to be processed captured by a depth camera and a grayscale image captured by a grayscale camera; the grayscale image and the image to be processed are acquired simultaneously. The target determination unit is used to determine the target pixels that meet the pixel completion conditions based on the image to be processed. A pixel projection unit is used to project the target pixel onto the grayscale image based on the pixel projection relationship between the depth camera and the grayscale camera to obtain a target projection area. A pixel padding unit is used to perform pixel padding processing on each pixel in the target projection region of the grayscale image to obtain the target pixel value of each pixel in the target projection region. An image fusion unit is used to perform pixel fusion processing on the image to be processed and the target projection region based on the target pixel values of each pixel in the target projection region to obtain a target depth image; The target determination unit includes: The image association module is used to determine the target local image of the grayscale image in the image to be processed based on the positional correspondence between the grayscale camera and the depth camera. The first extraction module is used to extract a first depth image from the target local image; The pixel selection module is used to determine target pixels that meet the pixel change conditions based on the first depth image. The pixel selection module includes: The gradient calculation submodule is used to calculate the target gradient image corresponding to the first depth image; the target gradient image includes the magnitude value of each pixel. The threshold comparison submodule is used to determine the center pixel point in the target gradient image whose modulus value is greater than the modulus threshold based on the preset modulus threshold. The associated acquisition submodule is used to determine a preset gradient selection area with the center pixel as the center, and obtain the associated pixels located in the gradient selection area that are associated with the center pixel; The target determination submodule is used to determine the center pixels and associated pixels of each center pixel in the local image of the target as the target pixels.
11. An electronic device, characterized in that, include: Processor, memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, such that the processor is configured with the image processing method based on a depth camera as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the image processing method based on a depth camera as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, The computer program is executed by a processor to configure the image processing method based on a depth camera as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Depth of field barrier avoiding method, equipment and unmanned aircraft
CN106960454A
Indoor synchronous positioning and mapping method
CN110887487A