Depth map generation method, device, and storage medium

By using a fisheye lens and an active depth sensor to generate a parallax map, the problem of insufficient field of view in existing technologies is solved, and dense depth images with a large field of view are generated, meeting the display requirements of AR devices.

CN113840130BActive Publication Date: 2025-12-09ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010591582.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-24
Publication Date
2025-12-09
Estimated Expiration
2040-06-24

AI Technical Summary

Technical Problem

Existing active and passive sensors generate dense depth images in AR devices with small field of view, which cannot meet the needs of display technology.

Method used

A parallax map is generated by acquiring spherical images using a first fisheye lens and a second fisheye lens, and combined with depth information from an active depth sensor, a dense depth map with a large field of view is generated through image fusion.

Benefits of technology

The field of view of the depth image was increased, generating a dense depth image with a large field of view, which meets the display requirements of AR devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113840130B_ABST
    Figure CN113840130B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a kind of depth map generation method, equipment and storage medium, belong to image processing technical field.The method comprises: according to the first spherical image of the first fish-eye lens collected and the second spherical image of the second fish-eye lens collected, the first parallax map of the space region where the terminal device is located is generated;And according to the depth information of the space region collected by the active depth sensor, the second parallax map of the space region is generated;Then according to the first parallax map and second parallax map, the target depth map of the space region is generated.The technical scheme of the present application can improve the field view angle of depth image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to a depth map generation method, device and storage medium. BACKGROUND

[0002] With the rapid development of science and technology, visual three-dimensional (3-Dimension, 3D) perception (i.e. depth information perception within a visual range) technology is increasingly widely used in life. For example, an augmented reality (AR) mobile head-mounted device is an important device for popularization of 3D technology. A dense depth image is an important basis for perceiving fine structures and understanding complete object surfaces, and is also an important core technology for AR based on 3D perception, which is of great significance to AR devices. Most AR mobile head-mounted devices have widely adopted active sensors and passive sensors to detect and obtain depth information in a field of view, and then fuse data of the active sensors and the passive sensors to obtain a dense depth image. However, since the field of view (FOV) of the existing active sensors is generally 65°x40° and the FOV of the passive sensors is generally 69°x42°, the field of view of the dense depth image obtained by fusing data of the active sensors and the passive sensors is relatively small, which does not match the development of display technology. SUMMARY

[0003] The main purpose of the present application is to provide a depth map generation method, device and storage medium, which aims to improve the field of view of a depth image.

[0004] In a first aspect, an embodiment of the present application provides a depth map generation method applied to a terminal device, wherein the terminal device comprises a first fisheye lens, a second fisheye lens and an active depth sensor, and the method comprises the following steps:

[0005] generating a first parallax map of a space region where the terminal device is located according to a first spherical image collected by the first fisheye lens and a second spherical image collected by the second fisheye lens;

[0006] generating a second parallax map of the space region according to depth information of the space region collected by the active depth sensor;

[0007] generating a target depth map of the space region according to the first parallax map and the second parallax map.

[0008] In a second aspect, the embodiments of the present application further provide a terminal device, the depth map generation device comprising a first fisheye lens, a second fisheye lens, an active depth sensor, a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory, wherein the computer program is executed by the processor to realize the steps of any one of the depth map generation methods provided in the specification of the present application.

[0009] In a third aspect, the embodiments of the present application further provide a storage medium for computer readable storage, characterized in that the storage medium stores one or more programs, the one or more programs being executable by one or more processors to realize the steps of any one of the depth map generation methods provided in the specification of the present application.

[0010] The embodiments of the present application provide a depth map generation method, device and storage medium, the embodiments of the present application generate a first disparity map of a space region where a terminal device is located by using a first spherical image collected by a first fisheye lens and a second spherical image collected by a second fisheye lens, generate a second disparity map of the space region according to depth information of the space region collected by an active depth sensor, and finally generate a target depth map of the space region according to the first disparity map and the second disparity map. The above scheme, since the fisheye lens has a large field of view angle, the first spherical image collected by the first fisheye lens and the second spherical image collected by the second fisheye lens can generate a first disparity map with a large field of view angle, and the depth information of the space region collected by the active depth sensor can generate a second disparity map, and finally based on the first disparity map with a large field of view angle and the second disparity map, a dense depth image with a large field of view angle can be generated, thereby improving the field of view angle of the depth image. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a structural schematic diagram of a terminal device for implementing the depth map generation method provided by the embodiments of the present application;

[0012] Figure 2 is a flowchart of a depth map generation method provided by the embodiments of the present application;

[0013] Figure 3 is a flowchart of a depth map generation method provided by the embodiments of the present application; Figure 2 is a flowchart of a depth map generation method provided by the embodiments of the present application;

[0014] Figure 4 is a flowchart of a depth map generation method provided by the embodiments of the present application; Figure 3 is a flowchart of a depth map generation method provided by the embodiments of the present application;

[0015] Figure 5 is a flowchart of a depth map generation method provided by the embodiments of the present application; Figure 2A flowchart illustrating the sub-steps of the depth map generation method in [the document / technology].

[0016] Figure 6 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0019] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0020] This invention provides a depth map generation method, device, and storage medium. The depth map generation method can be applied to terminal devices. Please refer to [link / reference]. Figure 1 , Figure 1 This is a schematic diagram of a terminal device implementing the depth map generation method provided in the embodiments of the present invention, as shown below. Figure 1 As shown, the terminal device 100 includes a first fisheye lens 110, a second fisheye lens 120, and an active depth sensor 130. The mounting positions of the first fisheye lens 110, the second fisheye lens 120, and the active depth sensor 130 on the terminal device, the spacing between the first fisheye lens 110 and the second fisheye lens 120, and the field of view angles of the first fisheye lens 110 and the second fisheye lens 120 can be set according to actual conditions. This embodiment of the invention does not impose specific limitations on these settings. For example, the first fisheye lens 110 and the second fisheye lens 120 may be 5 cm or 10 cm apart, and their field of view angles may both be 150°×180° or 210°×180°. In one embodiment, the terminal device may be an AR (Augmented Reality) head-mounted device.

[0021] Understandable Figure 1The terminal device 100 in the figures and the names of the various components of the terminal device 100 are merely for the purpose of identification and do not limit the embodiments of the present application.

[0022] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0023] Please refer to Figure 2 , Figure 2 A flowchart of a depth map generation method provided by an embodiment of the present application.

[0024] As shown in Figure 2 , the depth map generation method comprises steps S101 to S103.

[0025] Step S101, generating a first disparity map of a space region where the terminal device is located according to a first spherical image collected by a first fisheye lens and a second spherical image collected by a second fisheye lens.

[0026] The depth image generation method is applied to a terminal device, and the terminal device comprises a first fisheye lens, a second fisheye lens and an active depth camera. The active depth sensor comprises a TOF (Time of flight) sensor, a structured light sensor and a Lidar (Laser Radar), etc. The interval distance between the first fisheye lens and the second fisheye lens and the field of view angle of the first fisheye lens and the second fisheye lens can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereon. For example, the first fisheye lens and the second fisheye lens are 8 centimeters apart, and the field of view angle of the first fisheye lens and the second fisheye lens is 145°x180°.

[0027] In an embodiment, an image of the space region where the terminal device is located is collected by the first fisheye lens to obtain a first spherical image, and an image of the space region where the terminal device is located is collected by the second fisheye lens to obtain a second spherical image.

[0028] In an embodiment, as shown in Figure 3 , step S101 comprises sub-steps S1011 to S1014.

[0029] S1011, fusing the first spherical image and the second spherical image to obtain a target planar image of a preset field of view angle.

[0030] The first spherical image and the second spherical image are curved surface images, the first spherical image and the second spherical image are converted from the curved surface images into planar images of a preset field of view angle, and a target planar image of the preset field of view angle is obtained. The preset field of view angle can be set according to actual conditions, and the embodiment of the application does not make specific limitation on this, for example, the preset field of view angle can be set to 150°*180°.

[0031] In an embodiment, the first spherical image is converted into a first stereoscopic image, wherein the first stereoscopic image comprises a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the first spherical image; the second spherical image is converted into a second stereoscopic image, wherein the second stereoscopic image comprises a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the second spherical image; and the first stereoscopic image and the second stereoscopic image are fused to obtain the target planar image of the preset field of view angle. By converting the first spherical image and the second spherical image into stereoscopic images and fusing the two stereoscopic images, the target planar image of the preset field of view angle can be obtained, which is convenient for subsequent generation of a depth image of a large field of view angle based on a planar image of the large field of view angle.

[0032] In an embodiment, the first spherical image is converted into a first stereoscopic image in the following manner: the first spherical image is subjected to normalization processing to obtain a normalized sphere of the first spherical image, the normalized sphere of the first spherical image is split into a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image, and the front mapping image, the left mapping image, the right mapping image, the top mapping image and the bottom mapping image of the first spherical image are spliced to obtain the first stereoscopic image; similarly, the second spherical image is subjected to normalization processing to obtain a normalized sphere of the second spherical image, the normalized sphere of the second spherical image is split into a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image, and the front mapping image, the left mapping image, the right mapping image, the top mapping image and the bottom mapping image of the second spherical image are spliced to obtain the second stereoscopic image.

[0033] In one embodiment, the method for fusing the first stereoscopic image and the second stereoscopic image to obtain a target planar image with a preset field of view can be as follows: stitching together the forward-mapped image, left-mapped image, and right-mapped image in the first stereoscopic image to obtain a first image with a preset field of view; stitching together the forward-mapped image, left-mapped image, and right-mapped image in the second stereoscopic image to obtain a second image with a preset field of view; fusing the first image and the second image to obtain a first fused image, and fusing the top-mapped image in the first stereoscopic image with the top-mapped image in the second stereoscopic image to obtain a second fused image; fusing the bottom-mapped image in the first stereoscopic image with the bottom-mapped image in the second stereoscopic image to obtain a third fused image; and stitching together the first fused image, the second fused image, and the third fused image to obtain the target planar image with the preset field of view.

[0034] In one embodiment, fusing the first image and the second image to obtain a first fused image can be achieved by fusing the first image and the second image using an image fusion algorithm. Similarly, a second fused image can be obtained by fusing the top-viewed image in the first stereo image with the top-viewed image in the second stereo image using an image fusion algorithm, or a third fused image can be obtained by fusing the bottom-viewed image in the first stereo image with the bottom-viewed image in the second stereo image using an image fusion algorithm. The image fusion algorithms include wavelet transform-based image fusion algorithms and pyramid decomposition-based image fusion algorithms.

[0035] In one embodiment, such as Figure 4 As shown, step S1011 includes sub-steps S1011a to S1011b.

[0036] S1011a, Calibrate the first spherical image and the second spherical image.

[0037] When acquiring the first spherical image and the second spherical image through the first fisheye lens and the second fisheye lens, the acquired first spherical image and the second spherical image may be distorted due to the shaking of the terminal device or the movement of objects in the image, resulting in distortion of the first spherical image and the second spherical image. Therefore, it is necessary to calibrate the first spherical image and the second spherical image.

[0038] In an embodiment, the manner of calibrating the first spherical image and the second spherical image can be: converting the first spherical image into a third stereoscopic image, and converting the second spherical image into a fourth stereoscopic image; determining, according to the third stereoscopic image and the fourth stereoscopic image, a plurality of feature point matching pairs corresponding to a plurality of spatial points of a space region in which the terminal device is located, to obtain a plurality of feature point matching pairs; eliminating abnormal feature point matching pairs from the plurality of feature point matching pairs to obtain a plurality of target feature point matching pairs; and calibrating the first spherical image and the second spherical image according to the plurality of target feature point matching pairs. The manner of eliminating abnormal feature point matching pairs from the plurality of feature point matching pairs to obtain a plurality of target feature point matching pairs can be: obtaining a preset mathematical model, and eliminating abnormal feature point matching pairs from the plurality of feature point matching pairs based on the preset mathematical model to obtain a plurality of target feature point matching pairs, wherein the preset mathematical model is determined based on a random sample consensus (RANdom SAmple Consensus, RANSAC) algorithm. By calibrating the first spherical image and the second spherical image, it is convenient to subsequently generate an accurate disparity map based on the calibrated first spherical image and the second spherical image.

[0039] In an embodiment, the manner of converting the first spherical image into a third stereoscopic image and converting the second spherical image into a fourth stereoscopic image can be: performing normalization processing on the first spherical image to obtain a normalized sphere of the first spherical image, and splitting the normalized sphere of the first spherical image into a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image; splicing the front mapping image, the left mapping image, the right mapping image, the top mapping image, and the bottom mapping image of the first spherical image to obtain the third stereoscopic image; and similarly, performing normalization processing on the second spherical image to obtain a normalized sphere of the second spherical image, and splitting the normalized sphere of the second spherical image into a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image; splicing the front mapping image, the left mapping image, the right mapping image, the top mapping image, and the bottom mapping image of the second spherical image to obtain the fourth stereoscopic image.

[0040] In an embodiment, the feature point matching pairs corresponding to the plurality of spatial points of the space region where the terminal device is located are determined according to the third stereoscopic image and the fourth stereoscopic image, and the manner of obtaining the plurality of feature point matching pairs can be: based on a feature point extraction algorithm, extracting the feature points corresponding to the plurality of spatial points of the space region where the terminal device is located from the third stereoscopic image to obtain a plurality of first feature points, and based on the feature point extraction algorithm, extracting the feature points corresponding to the plurality of spatial points of the space region where the terminal device is located from the fourth stereoscopic image to obtain a plurality of second feature points; based on a feature point matching algorithm, matching each first feature point in the plurality of first feature points with a second feature point in the plurality of second feature points respectively to obtain a plurality of feature point matching pairs, and one feature point matching pair includes one first feature point and one second feature point. Wherein, the feature point extraction algorithm and the feature point matching algorithm can be selected according to actual conditions, and the embodiments of the present application do not make specific limitations thereon, for example, the feature point extraction algorithm includes at least one of the following: corner detection algorithm (Harris Corner Detection), scale-invariant feature transform (SIFT) algorithm, speeded-up robust features (SURF) algorithm, and FAST (Features From Accelerated Segment Test) feature point detection algorithm, and the feature point matching algorithm includes at least one of the following: KLT (Kanade-Lucas-Tomasi feature tracker) algorithm and brute-force matching algorithm.

[0041] In an embodiment, the feature point matching pairs corresponding to the plurality of spatial points of the space region where the terminal device is located are determined according to the third stereoscopic image and the fourth stereoscopic image, and the manner of obtaining the plurality of feature point matching pairs can be: converting the third stereoscopic image into a third planar image, i.e., performing extension splicing on the front mapping image, the left mapping image, the right mapping image, the top mapping image and the bottom mapping image in the third stereoscopic image to obtain the third planar image; converting the fourth stereoscopic image into a fourth planar image, i.e., performing extension splicing on the front mapping image, the left mapping image, the right mapping image, the top mapping image and the bottom mapping image in the fourth stereoscopic image to obtain the fourth planar image; extracting the feature points corresponding to the plurality of spatial points of the space region where the terminal device is located from the third planar image based on a feature point extraction algorithm to obtain a plurality of first feature points, and extracting the feature points corresponding to the plurality of spatial points of the space region where the terminal device is located from the fourth planar image based on the feature point extraction algorithm to obtain a plurality of second feature points; and matching each first feature point in the plurality of first feature points with a second feature point in the plurality of second feature points based on a feature point matching algorithm to obtain a plurality of feature point matching pairs, wherein one feature point matching pair includes one first feature point and one second feature point.

[0042] S1011b, fuse the calibrated first spherical image and the second spherical image to obtain a target planar image of a preset field of view angle.

[0043] In an embodiment, the calibrated first spherical image is converted into a first stereoscopic image, wherein the first stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated first spherical image; the calibrated second spherical image is converted into a second stereoscopic image, wherein the second stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated second spherical image; and the first stereoscopic image and the second stereoscopic image are fused to obtain a target planar image.

[0044] In an embodiment, the manner of converting the calibrated first spherical image into a first stereoscopic image can be: performing normalization processing on the calibrated first spherical image to obtain a normalized sphere of the calibrated first spherical image, and splitting the normalized sphere of the calibrated first spherical image into a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image, and splicing the front mapping image, the left mapping image, the right mapping image, the top mapping image, and the bottom mapping image of the calibrated first spherical image to obtain the first stereoscopic image; similarly, performing normalization processing on the calibrated second spherical image to obtain a normalized sphere of the second spherical image, and splitting the normalized sphere of the calibrated second spherical image into a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image, and splicing the front mapping image, the left mapping image, the right mapping image, the top mapping image, and the bottom mapping image of the calibrated second spherical image to obtain the second stereoscopic image.

[0045] S1012, convert the first spherical image into a first planar image, and convert the second spherical image into a second planar image.

[0046] convert the first spherical image into a stereoscopic image, wherein the stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image of the first spherical image; splice the front mapping image, the left mapping image, and the right mapping image in the stereoscopic image to obtain a first image; splice the top mapping image, the bottom mapping image in the stereoscopic image, and the first image to obtain the first planar image. Similarly, convert the second spherical image into a corresponding stereoscopic image, wherein the corresponding stereoscopic image of the second spherical image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image, and a bottom mapping image of the second spherical image; splice the front mapping image, the left mapping image, and the right mapping image in the corresponding stereoscopic image of the second spherical image to obtain a second image; splice the top mapping image, the bottom mapping image in the corresponding stereoscopic image of the second spherical image, and the second image to obtain the second planar image.

[0047] S1013, according to the first planar image and the second planar image, determine a plurality of feature point matching pairs corresponding to a plurality of spatial points in a spatial region where the terminal device is located, to obtain a plurality of feature point matching pairs.

[0048] The feature point extraction algorithm and the feature point matching algorithm can be selected according to actual conditions, and embodiments of the present application do not make specific limitations thereon. For example, the feature point extraction algorithm includes at least one of the following: a corner detection algorithm (Harris Corner Detection), a scale-invariant feature transform (SIFT) algorithm, a speeded-up robust features (SURF) algorithm, and a FAST (Features From Accelerated Segment Test) feature point detection algorithm; and the feature point matching algorithm includes at least one of the following: a KLT (Kanade-Lucas-Tomasi feature tracker) algorithm and a brute-force matching algorithm.

[0049] S1014, generating a first disparity map of the space region where the terminal device is located according to the plurality of feature point matching pairs and the target plane image.

[0050] Based on each feature point matching pair in the plurality of feature point matching pairs, a disparity value of a corresponding target space point in the space region where the terminal device is located is generated, and a pixel coordinate of each target space point on the target plane image is obtained; according to the disparity value of the corresponding target space point in the space region where the terminal device is located, a color of a pixel corresponding to the target space point on the target plane image is determined, and a first disparity map of the space region where the terminal device is located is generated according to the color of the pixel corresponding to the target space point on the target plane image.

[0051] In an embodiment, as shown in FIG. 10, after the sub-step S1011, the method further includes sub-steps S1015 to S1016. Figure 5

[0052] S1015, obtaining a historical plane image, wherein the historical plane image is determined according to the first spherical image and the second spherical image collected at the last time.

[0053] ​The first spherical image and the second spherical image collected at a previous time are acquired from a memory of the terminal device, and the first spherical image and the second spherical image collected at the previous time are fused to obtain a historical planar image; or the historical planar image is acquired from the memory of the terminal device. The time interval between the previous time and the current time can be set according to actual conditions, and the embodiments of the present application do not make specific limitations thereon, for example, the time interval between the previous time and the current time is set to 0.1 seconds.

[0054] S1016, a first parallax map of a space region where the terminal device is located is generated according to the target planar image and the historical planar image.

[0055] Based on a feature point extraction algorithm, feature points corresponding to a plurality of spatial points of the space region where the terminal device is located are extracted from the target planar image to obtain a plurality of fifth feature points; based on the feature point extraction algorithm, feature points corresponding to a plurality of spatial points of the space region where the terminal device is located are extracted from the historical planar image to obtain a plurality of sixth feature points; based on a feature point matching algorithm, each fifth feature point in the plurality of fifth feature points is matched with a sixth feature point in the plurality of sixth feature points to obtain a plurality of feature point matching pairs, and one feature point matching pair includes one fifth feature point and one sixth feature point; and a first parallax map of the space region where the terminal device is located is generated according to the plurality of feature point matching pairs. Through the target planar image and the historical planar image, a parallax map with a large field of view can be generated.

[0056] The feature point extraction algorithm and the feature point matching algorithm can be selected according to actual conditions, and the embodiments of the present application do not make specific limitations thereon, for example, the feature point extraction algorithm includes at least one of the following: corner detection algorithm (Harris Corner Detection), scale-invariant feature transform (SIFT) algorithm, scale and rotation invariant feature transform (SURF) algorithm, and FAST (Features From Accelerated Segment Test) feature point detection algorithm, and the feature point matching algorithm includes at least one of the following: KLT (Kanade-Lucas-Tomasi feature tracker) algorithm and brute-force matching algorithm.

[0057] In an embodiment, the first spherical image is converted into a first planar image of a preset field of view angle, and the second spherical image is converted into a second planar image of the preset field of view angle; and a first disparity map of a space region where the terminal device is located is generated according to the first planar image and the second planar image. The first spherical image can be converted into a first stereoscopic image, and a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image on the first stereoscopic image are spliced to obtain the first planar image of the preset field of view angle. Similarly, the second spherical image can be converted into a second stereoscopic image, and a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image on the second stereoscopic image are spliced to obtain the second planar image of the preset field of view angle. The first disparity map of the space region where the terminal device is located is obtained through the first planar image and the second planar image, thereby improving the efficiency and accuracy of the terminal device in generating the first disparity map.

[0058] In an embodiment, the first disparity map of the space region where the terminal device is located can be generated according to the first planar image and the second planar image in the following manner: a plurality of seventh feature points corresponding to a plurality of space points of the space region where the terminal device is located are extracted from the first planar image based on a feature point extraction algorithm; a plurality of eighth feature points corresponding to the plurality of space points of the space region where the terminal device is located are extracted from the second planar image based on the feature point extraction algorithm; each seventh feature point in the plurality of seventh feature points is matched with an eighth feature point in the plurality of eighth feature points based on a feature point matching algorithm, to obtain a plurality of feature point matching pairs, and each feature point matching pair includes a fifth feature point and a sixth feature point; and the first disparity map of the space region where the terminal device is located is generated according to the plurality of feature point matching pairs.

[0059] It should be noted that the above-mentioned several ways of generating the first disparity map of the space region where the terminal device is located can be used to generate the first disparity map of the space region where the terminal device is located in one way, or can be combined to generate the first disparity map of the space region where the terminal device is located. The embodiments of the present application do not make specific limitations thereon, and the first disparity map can be more accurate by reasonable combination according to actual conditions.

[0060] In step S102, a second disparity map of the space region is generated according to the depth information of the space region collected by the active depth sensor.

[0061] The active sensor operation includes actively sending light pulses or other light rays to the target object, then receiving the reflected light pulses or other light rays, and obtaining the depth information of the target object according to the reflected light pulses or other light rays. The active sensor can be selected according to actual conditions, which is not limited in the present application. For example, the active sensor can be a time of flight (TOF) sensor, a laser radar (Lidar), and a structured light sensor.

[0062] In an embodiment, the active sensor controls the emission of light pulses to the space region where the terminal device is located, and receives the reflected light pulses. The depth information of the space region is determined according to the frequency and return time of the reflected light pulses, and the second parallax map of the space region is obtained according to the depth information of the space region.

[0063] In step S103, the target depth map of the space region is generated according to the first parallax map and the second parallax map.

[0064] In an embodiment, the first parallax map and the second parallax map are fused to obtain a target parallax map, and the target depth map of the space region where the terminal device is located is generated based on the target parallax map. The way to fuse the first parallax map and the second parallax map to obtain the target parallax map can be: obtaining the parallax value of each first pixel point in the first parallax map, and obtaining the parallax value of each second pixel point in the second parallax map; determining the target parallax value of each pixel point according to the parallax value of each first pixel point and the parallax value of each second pixel point, and generating the target parallax map based on the target parallax value of each pixel point.

[0065] In an embodiment, the way to determine the target parallax value of each pixel point can be: obtaining a calculation formula of the target parallax value; determining the target parallax value of each pixel point based on the calculation formula d = w T d T +w S d S , according to the parallax value of each first pixel point and the parallax value of each second pixel point. The calculation formula is d = w T d T +w S d S , d is the target parallax, w T is the weight of the parallax value of the first pixel point, w S is the weight of the parallax value of the second pixel point, d T is the parallax value of the first pixel point, d S is the parallax value of the second pixel point, w T and wS The specific value can be set based on the actual situation, and the embodiments of the present invention do not impose specific limitations on it.

[0066] In one embodiment, the fusion of the first disparity map and the second disparity map can be performed as follows: The confidence level of the disparity value of each pixel in the first disparity map is obtained, and the confidence level of the disparity value of each pixel in the second disparity map is also obtained; pixels in the first disparity map whose confidence level is less than a preset confidence level are filtered out to obtain a first calibrated disparity map, and pixels in the second disparity map whose confidence level is less than the preset confidence level are filtered out to obtain a second calibrated disparity map; the first calibrated disparity map and the second calibrated disparity map are fused to obtain a target disparity map. The preset confidence level can be set based on actual conditions, and this embodiment of the invention does not specifically limit it.

[0067] The depth map generation method provided in the above embodiments generates a first disparity map of the spatial region where the terminal device is located using a first spherical image acquired by a first fisheye lens and a second spherical image acquired by a second fisheye lens. It then generates a second disparity map of the spatial region based on depth information acquired by an active depth sensor. Finally, a target depth map of the spatial region is generated based on the first and second disparity maps. In the above embodiments, because the fisheye lens has a large field of view, the first spherical image acquired by the first fisheye lens and the second spherical image acquired by the second fisheye lens can generate a first disparity map with a large field of view. Furthermore, the depth information of the spatial region acquired by the active depth sensor can generate a second disparity map. Finally, based on the first and second disparity maps with a large field of view, a dense depth image with a large field of view can be generated, thereby improving the field of view of the depth image.

[0068] Please see Figure 6 , Figure 6 This is a schematic block diagram of the structure of a terminal device provided in an embodiment of the present invention.

[0069] like Figure 6 As shown, the depth map generation device 200 includes a first fisheye lens 201, a second fisheye lens 202, an active depth sensor 203, a processor 204, and a memory 205. The first fisheye lens 201, the second fisheye lens 202, the active depth sensor 203, the processor 204, and the memory 205 are connected via a bus 206, such as an I2C (Inter-integrated Circuit) bus.

[0070] Specifically, the processor 204 is configured to provide computing and control capabilities to support the operation of the entire depth map generation device. The processor 204 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0071] Specifically, the memory 205 can be a flash chip, a read-only memory (ROM) disk, an optical disk, a U disk or a mobile hard disk, etc.

[0072] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal device to which the scheme of the present application is applied. Specifically, the server can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0073] The processor is configured to run a computer program stored in the memory, and implement any one of the depth map generation methods provided by the embodiments of the present application when the computer program is executed.

[0074] In an embodiment, the processor is configured to run a computer program stored in the memory, and implement the following steps when the computer program is executed:

[0075] According to the first spherical image collected by the first fisheye lens and the second spherical image collected by the second fisheye lens, a first parallax map of a space region where the terminal device is located is generated;

[0076] According to the depth information of the space region collected by the active depth sensor, a second parallax map of the space region is generated;

[0077] According to the first parallax map and the second parallax map, a target depth map of the space region is generated.

[0078] In an embodiment, the processor, when implementing the fusing of the first spherical image and the second spherical image to obtain a target planar image of a preset field of view, is configured to:

[0079] fuse the first spherical image and the second spherical image to obtain a target planar image of a preset field of view;

[0080] convert the first spherical image into a first planar image and convert the second spherical image into a second planar image;

[0081] determine, according to the first planar image and the second planar image, a plurality of feature point matching pairs corresponding to a plurality of spatial points in a spatial region where the terminal device is located, to obtain the plurality of feature point matching pairs;

[0082] generate a first parallax map of the spatial region where the terminal device is located according to the plurality of feature point matching pairs and the target planar image.

[0083] In an embodiment, the processor, when implementing the fusing of the first spherical image and the second spherical image to obtain a target planar image of a preset field of view, is configured to:

[0084] calibrate the first spherical image and the second spherical image;

[0085] fuse the calibrated first spherical image and the calibrated second spherical image to obtain a target planar image of a preset field of view.

[0086] In an embodiment, the processor, when implementing the fusing of the first spherical image and the second spherical image to obtain a target planar image of a preset field of view, is configured to:

[0087] convert the calibrated first spherical image into a first stereoscopic image, wherein the first stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated first spherical image;

[0088] convert the calibrated second spherical image into a second stereoscopic image, wherein the second stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated second spherical image;

[0089] fuse the first stereoscopic image and the second stereoscopic image to obtain a target planar image of a preset field of view.

[0090] In an embodiment, the processor, when implementing the fusing the first stereoscopic image and the second stereoscopic image to obtain a target planar image of a preset field of view, is configured to:

[0091] stitch a front mapping image, a left mapping image and a right mapping image in the first stereoscopic image to obtain a first image of the preset field of view;

[0092] stitch a front mapping image, a left mapping image and a right mapping image in the second stereoscopic image to obtain a second image of the preset field of view;

[0093] fuse the first image and the second image to obtain a first fused image, and fuse a top mapping image in the first stereoscopic image and a top mapping image in the second stereoscopic image to obtain a second fused image;

[0094] fuse a bottom mapping image in the first stereoscopic image and a bottom mapping image in the second stereoscopic image to obtain a third fused image;

[0095] stitch the first fused image, the second fused image and the third fused image to obtain the target planar image of the preset field of view.

[0096] In an embodiment, the processor, when implementing the calibrating the first spherical image and the second spherical image, is configured to:

[0097] convert the first spherical image into a third stereoscopic image, and convert the second spherical image into a fourth stereoscopic image;

[0098] determine, according to the third stereoscopic image and the fourth stereoscopic image, a plurality of feature point matching pairs corresponding to a plurality of spatial points of a space region in which the terminal device is located, to obtain the plurality of feature point matching pairs.

[0099] In an embodiment, the processor, when implementing the converting the first spherical image into a first planar image, is configured to:

[0100] convert the first spherical image into a stereoscopic image, wherein the stereoscopic image includes a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the first spherical image;

[0101] stitch the front mapping image, the left mapping image and the right mapping image in the stereoscopic image to obtain a first image;

[0102] stitch the top mapping image, the bottom mapping image in the stereoscopic image and the first image to obtain the first planar image.

[0103] In an embodiment, after the processor implements the fusing of the first spherical image and the second spherical image to obtain a target planar image with a preset field of view angle, the processor is further configured to implement:

[0104] acquiring a historical planar image, wherein the historical planar image is determined according to a first spherical image and a second spherical image collected at a previous time;

[0105] generating a first parallax map of a space region where the terminal device is located according to the target planar image and the historical planar image.

[0106] It should be noted that, for the convenience and brevity of description, the specific working process of the terminal device described above can refer to the corresponding process in the foregoing depth map generation method embodiments, which will not be described here.

[0107] The embodiment of the present application further provides a storage medium for computer readable storage, the storage medium storing one or more programs, the one or more programs being executable by one or more processors to implement the steps of any depth map generation method provided in the specification of the present application.

[0108] The storage medium can be an internal storage unit of the terminal device, such as a hard disk or a memory of the terminal device. The storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0109] Those skilled in the art can understand that all or some of the steps in the methods disclosed above, the functional modules / units in the systems and devices can be implemented by software, firmware, hardware, or a combination thereof. In hardware implementation, the split between the functional modules / units referred to in the above description does not necessarily correspond to the split between physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components working together. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer readable media, which can include computer storage media (or non-transitory media), and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is common knowledge to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0110] It should be understood that the term "and / or" as used herein, refers to any combination of associated listed items, including all possible combinations, and includes any of the associated listed items. It is to be noted that the terms "comprising", "including", or any other variant thereof, are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or systems that comprise a list of elements do not include only those elements, but can also include other elements not expressly listed or inherent to such processes, methods, articles, or systems. Without more limitations, an element defined by the phrase "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or system including the element.

[0111] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments. The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of depth map generation, characterized by, The method is applied to a terminal device, the terminal device comprising a first fisheye lens, a second fisheye lens and an active depth sensor, and the method comprises: generating a first parallax map of a space region where the terminal device is located according to a first spherical image collected by the first fisheye lens and a second spherical image collected by the second fisheye lens; generating a second parallax map of the space region according to depth information of the space region collected by the active depth sensor; generating a target depth map of the space region according to the first parallax map and the second parallax map; wherein the generating of the first parallax map of the space region where the terminal device is located according to the first spherical image collected by the first fisheye lens and the second spherical image collected by the second fisheye lens comprises: fusing the first spherical image and the second spherical image to obtain a target planar image of a preset field of view; converting the first spherical image into a first planar image and converting the second spherical image into a second planar image; determining a plurality of feature point matching pairs corresponding to a plurality of space points of the space region where the terminal device is located according to the first planar image and the second planar image to obtain the plurality of feature point matching pairs; generating the first parallax map of the space region where the terminal device is located according to the plurality of feature point matching pairs and the target planar image.

2. The depth map generation method of claim 1, wherein, The fusing of the first spherical image and the second spherical image to obtain the target planar image of the preset field of view comprises: calibrating the first spherical image and the second spherical image; fusing the calibrated first spherical image and the calibrated second spherical image to obtain the target planar image of the preset field of view.

3. The depth map generation method of claim 2, wherein, The fusing of the calibrated first spherical image and the calibrated second spherical image to obtain the target planar image of the preset field of view comprises: converting the calibrated first spherical image into a first stereoscopic image, wherein the first stereoscopic image comprises a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated first spherical image; converting the calibrated second spherical image into a second stereoscopic image, wherein the second stereoscopic image comprises a front mapping image, a left mapping image, a right mapping image, a top mapping image and a bottom mapping image of the calibrated second spherical image; fusing the first stereoscopic image and the second stereoscopic image to obtain the target planar image of the preset field of view.

4. The depth map generation method of claim 3, wherein, The fusing of the first stereoscopic image and the second stereoscopic image to obtain the target planar image of the preset field of view comprises: splicing the front mapping image, the left mapping image and the right mapping image in the first stereoscopic image to obtain a first image of the preset field of view; splicing the front mapping image, the left mapping image and the right mapping image in the second stereoscopic image to obtain a second image of the preset field of view; fusing the first image and the second image to obtain a first fused image, and fusing the top mapping image in the first stereoscopic image and the top mapping image in the second stereoscopic image to obtain a second fused image; fusing a bottom-mapped image in the first stereoscopic image with a bottom-mapped image in the second stereoscopic image to obtain a third fused image; stitching the first fused image, the second fused image and the third fused image to obtain a target planar image with a preset field of view angle.

5. The depth map generation method of claim 2, wherein, The calibration of the first spherical image and the second spherical image comprises: converting the first spherical image into a third stereoscopic image and converting the second spherical image into a fourth stereoscopic image; determining, according to the third stereoscopic image and the fourth stereoscopic image, a plurality of feature point matching pairs corresponding to a plurality of spatial points in a spatial region where the terminal device is located to obtain the plurality of feature point matching pairs; eliminating abnormal feature point matching pairs from the plurality of feature point matching pairs to obtain a plurality of target feature point matching pairs; calibrating the first spherical image and the second spherical image according to the plurality of target feature point matching pairs.

6. The depth map generation method of claim 1, wherein, The conversion of the first spherical image into a first planar image comprises: converting the first spherical image into a stereoscopic image, wherein the stereoscopic image comprises a front-mapped image, a left-mapped image, a right-mapped image, a top-mapped image and a bottom-mapped image of the first spherical image; stitching the front-mapped image, the left-mapped image and the right-mapped image in the stereoscopic image to obtain a first image; stitching the top-mapped image, the bottom-mapped image in the stereoscopic image and the first image to obtain the first planar image.

7. The depth map generation method of claim 1, wherein, After the fusion of the first spherical image and the second spherical image to obtain a target planar image with a preset field of view angle, the method further comprises: acquiring a historical planar image, wherein the historical planar image is determined according to a first spherical image and a second spherical image collected at a previous time; generating a first disparity map of a spatial region where the terminal device is located according to the target planar image and the historical planar image.

8. A terminal device, comprising: The terminal device comprises a first fisheye lens, a second fisheye lens, an active depth sensor, a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein the computer program, when executed by the processor, realizes the steps of the depth map generation method according to any one of claims 1 to 7.

9. A storage medium for computer-readable storage, characterized in that, The storage medium stores one or more programs executable by one or more processors to realize the steps of the depth map generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Electronic device for generating 360-degree three-dimensional image and method therefor

    US10595004B2

  • Depth Information Acquisition Method and Device

    US20200128225A1