3D depth map construction method, device and ar glasses
By combining optical multiplexing, multi-frame super-resolution and a specific type of deformable 3D model, the problem of insufficient high precision and high resolution of dense depth images in lightweight AR glasses is solved, improving the user experience and the quality of 3D depth maps.
Patent Information
- Application Number
- CN202011303757.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-19
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2040-11-19
AI Technical Summary
Existing lightweight AR glasses suffer from insufficient high precision and high resolution when constructing dense depth images, resulting in edge noise defects and insufficient original accuracy, which affects the user experience.
By adopting compressed sensing technology based on optical multiplexing, multi-frame super-resolution technology with sliding time windows, and a specific type of deformable 3D model, combined with multi-domain mutually orthogonal enhancement methods, a high-precision and high-resolution 3D depth map is constructed.
This improves the user experience of thin and light AR glasses without increasing the size of the glasses lens or package, eliminates conventionally visible block defects and perceived information noise, and improves the accuracy and resolution of 3D depth maps.
Smart Images

Figure CN114519763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D depth perception technology, and in particular to a 3D depth map construction method, device and AR glasses. Background Art
[0002] Thin, lightweight augmented reality (AR) glasses are increasingly popular in the AR industry for personal (to-customer, 2C) applications. Three-dimensional (3D) perception (i.e., the perception of depth information within the visible range), an increasingly important technical requirement for AR perception, has become an indispensable technical key. High-precision dense depth images are a crucial basis for perceiving fine structures and understanding the surfaces of complete objects (scenes). They are also a key core technology for 3D-based AR and are of great significance to AR devices. Most lightweight AR glasses already utilize small active (ToF) and passive (Red, Green, Blue, RGB) sensors to construct RGBD (RGB Depth) imaging to detect and acquire depth information within the field of view.
[0003] Active depth sensors, such as Time of Flight (ToF), are an active 3D imaging technology with Figure 1 The active light source array shown in Figure 1. The depth perception image resolution is significantly dependent on the light source array size and is susceptible to multipath propagation, resulting in high noise at the edges. RGBD perception images of varying accuracy and resolution significantly impact user interaction and experience. Lightweight AR glasses currently acquire dense depth images by fusing single-frame RGB and ToF images.
[0004] The core problem with the dense depth image perception solutions currently used in lightweight AR glasses is their lack of high precision and resolution. This inability to effectively suppress noise artifacts caused by edges and insufficient raw precision results in depth image graininess, missing artifacts, and blurred transitions, which are inconsistent with the development trend of AR application experiences. Summary of the Invention
[0005] Embodiments of the present invention provide a 3D depth map construction method, device, and AR glasses to at least solve the problem of insufficient accuracy and resolution of depth perception images in related technologies.
[0006] According to one embodiment of the present application, a 3D depth map construction method is provided, comprising: performing optical multiplexing-based compressed sensing enhancement on an original ToF acquisition signal to obtain a first ToF acquisition signal; performing multi-frame super-resolution enhancement based on a sliding time window on a ToF image corresponding to the first ToF acquisition signal to reconstruct a high-resolution second ToF acquisition signal; constructing a 3D depth map of the ToF image according to the second ToF acquisition signal using a specific class of deformable three-dimensional model; and performing confidence-based multi-orthogonal domain enhancement 3D depth map fusion on the 3D depth map.
[0007] In one exemplary embodiment, performing optical multiplexing-based compressed sensing enhancement on an original ToF acquisition signal to obtain a first ToF acquisition signal comprises: changing the encoding on a time-multiplexed light encoding module during multiple exposures of a ToF acquisition image to achieve time-multiplexing of a scene to reconstruct the first ToF acquisition signal with high-resolution amplitude and depth images.
[0008] In one exemplary embodiment, changing the encoding on a time-multiplexed light encoding module during multiple exposures of a ToF acquisition image to achieve time-multiplexing of a scene to reconstruct the first ToF acquisition signal with high-resolution amplitude and depth images comprises: forming a linear image of a scene on the time-multiplexed light encoding module of an original ToF acquisition path; during multiple exposures of a ToF acquisition image, performing time-multiplexing of a scene by changing the encoding on the time-multiplexed light encoding module; relaying the linear image of the scene to a ToF acquisition plane; and collecting the linear image of the scene by ToF pixels on the ToF acquisition plane to reconstruct the first ToF acquisition signal with high-resolution amplitude and depth images.
[0009] In one exemplary embodiment, performing multi-frame super-resolution enhancement based on a sliding time window on a ToF image corresponding to the first ToF acquisition signal to obtain a second ToF acquisition signal comprises: within a sliding time window, obtaining a low-resolution ToF image from the first ToF acquisition signal, and reconstructing a scene real signal based on a probability distribution of the scene real signal through a phasor diagram of multiple frames of the ToF image to generate a high-resolution second ToF acquisition signal.
[0010] In an example embodiment, the low-resolution ToF image is obtained from the first ToF acquisition signal, and a high-resolution second ToF acquisition signal is generated by reconstructing a scene ground truth based on a probability distribution of the scene ground truth from a phasor diagram of the ToF image of multiple frames, including: obtaining the low-resolution ToF image from the first ToF acquisition signal, and calculating a phasor diagram of the low-resolution ToF image; obtaining a reconstructed high-resolution ToF image sequence from a low-resolution ToF image sequence in the sliding time window; reconstructing a scene ground truth based on a probability distribution of the scene ground truth from the high-resolution ToF image sequence to generate a high-resolution second ToF acquisition signal.
[0011] In an example embodiment, a 3D depth map of a ToF image is constructed from the second ToF acquisition signal using a specific class deformable three-dimensional model, including: learning a class-specific deformable model from 2D annotations of a target detection dataset; combining object detection, instance segmentation, and viewpoint estimation to identify, locate, and estimate the pose of objects in the ToF image; using the learned deformable model to perform 3D reconstruction of the objects in the ToF image in combination with the estimated viewpoint and instantiation information obtained from the viewpoint estimation; and fusing three-dimensional shapes and high-frequency local shape indicators to reconstruct the 3D depth map containing multi-view observation angles.
[0012] In an example embodiment, the multi-orthogonal domain enhancement includes time domain enhancement, phase domain enhancement, and semantic enhancement.
[0013] According to another embodiment of the present application, a 3D depth map construction device is provided, the device comprising: a first enhancement module for performing optical multiplexing-based compressed sensing enhancement on an original ToF acquisition signal to obtain a first ToF acquisition signal; a second enhancement module for performing sliding time window-based multi-frame super-resolution enhancement on a ToF image corresponding to the first ToF acquisition signal to reconstruct a high-resolution second ToF acquisition signal; a construction module for constructing a 3D depth map of a ToF image using a specific class deformable three-dimensional model from the second ToF acquisition signal; and a fusion module for performing confidence-based multi-orthogonal domain enhancement 3D depth map fusion on the 3D depth map.
[0014] In an example embodiment, the first enhancement module includes a time-multiplexed light encoding module arranged in an original ToF acquisition path, configured to achieve time-multiplexing of a scene by changing the encoding on the time-multiplexed light encoding module during multiple exposures of a ToF acquisition image.
[0015] According to another embodiment of the present application, an AR glasses is also provided, which comprises the 3D depth map construction device in the above-mentioned embodiments.
[0016] In an exemplary embodiment, the AR glasses further comprises a calculation module, wherein the calculation module can be an NPU acceleration chip in a smart terminal or a smart watch.
[0017] According to another embodiment of the present application, a computer readable storage medium is also provided, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above-mentioned method embodiments when running.
[0018] According to another embodiment of the present application, an electronic device is also provided, which comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the above-mentioned method embodiments.
[0019] In the above-mentioned embodiments of the present application, the compressed sensing technology based on optical multiplexing, the multi-frame super-resolution technology based on sliding time window, and the single image 3D structure inference based on specific class deformable three-dimensional model are utilized, and the multi-domain orthogonal enhancement is performed, so as to construct a 3D depth map with high precision and high resolution. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is a schematic diagram of active depth sensor ToF active 3D imaging according to the related art;
[0021] Figure 2 is a hardware structure diagram of AR glasses device according to an embodiment of the present application
[0022] Figure 3 is a flowchart of a 3D depth map construction method according to an embodiment of the present application;
[0023] Figure 4 is a structural block diagram of a 3D depth map construction device according to an embodiment of the present application;
[0024] Figure 5 is a flowchart of a depth construction method of orthogonal multi-domain perception fusion enhancement according to an embodiment of the present application;
[0025] Figure 6 is a schematic diagram of original ToF acquisition signal enhancement based on compressed sensing of optical multiplexing according to an embodiment of the present application;
[0026] Figure 7 is a schematic diagram of original ToF acquisition signal enhancement process based on multi-frame super-resolution of sliding time window according to an embodiment of the present application;
[0027] Figure 8 is a schematic diagram of a 3D depth map construction process based on specific class deformable three-dimensional model structure reasoning according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with embodiments.
[0029] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0030] In view of the deficiencies of high precision and high resolution in the dense depth image perception scheme used by the current lightweight AR glasses, the noise defects caused by edges and the like and the deficiencies of original precision cannot be effectively inhibited, and the depth map particle defects and missing defects and transition blur caused thereby do not match the experience development trend of AR applications (2C). Embodiments of the present application provide an enhanced means solution based on multi-domain fusion perception (time domain, phase domain, and semantic prior domain, etc.).
[0031] Figure 2 is a hardware structure diagram of a lightweight AR glass device according to an embodiment of the present application. As shown in Figure 2 The lightweight AR glasses involved in the embodiments of the present application can construct a high-precision high-resolution depth map perception solution for lightweight AR glasses by integrating sensing (one ToF active depth lens, with a linear phase measurement light multiplexing calculation module such as a Digital Micromirror Device (DMD) or Liquid Crystal on Silicon (LCoS), and one RGB passive lens) and a calculation unit (which can be based on an NPU acceleration chip of a smart phone / tablet).
[0032] In the embodiments of the present application, optical multiplexing and compressive sensing based on the principle of capturing the phase representation of an image and resulting in a linear image formation model (phase domain enhancement), multi-frame super-resolution technology based on a sliding time window (time domain enhancement), and a single image 3D structure reasoning method based on a specific class deformable three-dimensional model (semantic prior domain enhancement) are combined through the enhancement methods orthogonal to each other, and an innovative high-precision high-resolution depth map perception solution based on a limited packaging size is proposed to complete the high-precision high-resolution depth map later.
[0033] The light and thin AR glasses based on the embodiment of the application can have high-precision and high-resolution forward depth information without adding extra lenses or increasing the size of the corresponding ToF package, thereby greatly improving the user experience of the light and thin AR glasses.
[0034] A 3D depth map construction method applicable to the AR glasses is provided in the embodiment, Figure 3 is a flowchart of the 3D depth map construction method according to the embodiment of the application, as shown in the figure, the flow includes the following steps: Figure 3
[0035] In step S302, the original ToF acquisition signal is compressed sensing enhanced based on optical multiplexing to obtain a first ToF acquisition signal.
[0036] In step S304, the ToF image corresponding to the first ToF acquisition signal is super-resolution enhanced based on a sliding time window to reconstruct a high-resolution second ToF acquisition signal.
[0037] In step S306, a 3D depth map of the ToF image is constructed according to the second ToF acquisition signal using a specific class of deformable three-dimensional model.
[0038] In step S308, the 3D depth map is fused based on confidence-based multi-orthogonal domain enhancement of the 3D depth map.
[0039] In step S302 of the embodiment, a linear image of a scene is formed on the time division multiplexing light encoding module of the original ToF acquisition path; during the multiple exposure processes of the ToF acquisition image, the time division multiplexing of the scene is performed by changing the encoding on the time division multiplexing light encoding module; the linear image of the scene is relayed to the ToF acquisition plane; the ToF pixels on the ToF acquisition plane collect the linear image of the scene, and the first ToF acquisition signal with high-resolution amplitude and depth image is reconstructed.
[0040] In step S304 of the embodiment, the low-resolution ToF image is obtained according to the first ToF acquisition signal, and a phasor diagram of the low-resolution ToF image is calculated; the reconstructed high-resolution ToF image sequence is calculated according to the low-resolution ToF image sequence in the sliding time window; the high-resolution second ToF acquisition signal is generated by reconstructing the scene real signal based on the probability distribution of the scene real signal according to the high-resolution ToF image sequence.
[0041] In step S306 of the embodiment, a class-specific deformable model is learned from 2D annotations of a target detection dataset; object detection, instance segmentation and viewpoint estimation are combined to identify, localize and estimate the pose of objects in a ToF image; using the learned deformable model, 3D reconstruction of objects in the ToF image is performed, fusing three-dimensional shape and high-frequency local shape indicators, reconstructing the 3D depth map containing multi-view observation angles, based on the estimated viewpoint and instantiation information obtained from the viewpoint estimation.
[0042] In step S308 of the embodiment, 3D depth map fusion is enhanced by multi-confidence-based multi-orthogonal domains (based on multi-frame, multi-phase and semantic perception domains), to complete high-precision high-resolution 3D depth construction for a thin AR glasses ToF sensor with limited packaging size.
[0043] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, an optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.
[0044] In the embodiment, a 3D depth map construction device is also provided, which is used to implement the above embodiments and preferred embodiments, and will not be described again. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware or a combination of software and hardware is also possible and contemplated.
[0045] Figure 4 is a structural block diagram of a 3D depth map construction device according to an embodiment of the present application, as shown in Figure 4 The device includes a first enhancement module 10, a second enhancement module 20, a construction module 30 and a fusion module 40.
[0046] The first enhancement module 10 can perform optical multiplexing-based compressed sensing enhancement on the original ToF acquisition signal to obtain a first ToF acquisition signal;
[0047] The second enhancement module 20 can perform multi-frame super-resolution enhancement based on a sliding time window on the ToF image corresponding to the first ToF acquisition signal to reconstruct a high-resolution second ToF acquisition signal;
[0048] The construction module 30 can construct a 3D depth map of the ToF image according to the second ToF acquisition signal using a specific class of deformable three-dimensional model;
[0049] The fusion module can perform confidence-based multi-orthogonal domain enhanced 3D depth map fusion on the 3D depth map.
[0050] In the embodiment, the first enhancement module can include a time division multiplexing light encoding module arranged in the original ToF acquisition path, for realizing time division multiplexing of the scene by changing the encoding on the time division multiplexing light encoding module in the multiple exposure processes of the ToF acquisition image.
[0051] It should be noted that the above modules can be realized by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: the above modules are located in the same processor; or the above modules are located in different processors in any combination.
[0052] It is known that the ToF sensor of light and thin AR glasses based on limited packaging size cannot obtain high-precision and high-resolution 3D depth, which is because the conventional ToF imager is a focal plane array that simultaneously encodes the intensity and depth information of the scene at each pixel. A large number of transistors are required for each pixel to perform correlation measurement by additional electronic devices. Therefore, the pixel size of the conventional CMOS image sensor is close to 1 micrometer, and the fill factor is greater than 90%. The current ToF sensor can only realize a pixel size close to 10 micrometers, and the fill factor is close to 10%, which also leads to that the 3D depth resolution of the ToF sensor of light and thin AR glasses with strict limitations on device size is very limited.
[0053] Therefore, the embodiment of the present application also provides an AR glass. In the AR glass of the embodiment, the depth construction technology of orthogonal multi-domain (multi-frame, multi-phase, and semantic enhancement) perception fusion enhancement is applied.
[0054] Firstly, by adding a time division multiplexing light encoding module (such as a digital micro-mirror DMD) to the original ToF acquisition path, and using an optical multiplexing and compressive sensing technology, a linear image model is formed by using phase domain enhancement, and an enhanced high-precision ToF acquisition signal is obtained.
[0055] Secondly, the original continuous ToF acquisition image (in a fixed size time window) is adopted, the sensing result of which contains low resolution (LR) intensity image and depth map, and the ToF imaging model output can be expressed as a complex model by the phasor diagram representation. Then, the high resolution (SR) ToF acquisition signal is reconstructed by combining the LR intensity and depth image. The process is as shown in Figure 5
[0056] Then, the specific class deformable three-dimensional model structure inference method is adopted, the geometric consistency information is used to bridge the inferred 3D representation and the available 2D supervision, the 3D shape (i.e., the main observation object depth information) obtained from any view observation angle is constructed by using the convolutional neural network based prediction model combined with gradient differentiable calculation.
[0057] Finally, the 3D depth map fusion is enhanced based on the confidence based multi-orthogonal domain (based on multi-frame, multi-phase and semantic perception domain), so as to complete the high precision high resolution 3D depth construction for the thin AR glasses ToF sensor with limited packaging size.
[0058] The depth construction method of the orthogonal multi-domain (multi-frame, multi-phase, semantic enhancement) perception fusion enhancement in the embodiment will be described in detail below. As shown in Figure 5 , the method mainly includes the following steps:
[0059] Step S502, original ToF acquisition signal enhancement based on optical multiplexing compression sensing.
[0060] 1) Add a time division multiplexing light encoding module (DMD or LCoS) and a relay mirror to the original ToF acquisition path.
[0061] The original near-infrared laser diode of the ToF acquisition path is used to illuminate the scene, the scene is formed on the DMD using the objective lens, and then the high resolution DMD modulated image is relayed to the low resolution ToF acquisition plane, as shown in Figure 6
[0062] 2) The reflection from the object is collected by the ToF pixel, and then correlated with the reference signal to generate the output of the camera. The output follows the formula:
[0063]
[0064] Where m(t) is the active output, is the attenuation, and r(t-ψ) is the camera rolling shutter input. The input amplitude and phase at the pixel point are as shown in the following formula:
[0065]
[0066]
[0067] 3) After adding DMD, the linear image model of the scene formed by the objective on the DMD is as follows:
[0068]
[0069] Where a s is the intensity, Ф s is the phase, x is the input scene, y is the linear image obtained by ToF, C is the mapping between the DMD imaging pixels and the ToF imaging pixels, and M is the modulation mode displayed on the DMD.
[0070] 4) Time division multiplexing of the scene is achieved by changing the code on the DMD in multiple exposures, so that high-resolution amplitude and depth images can be reconstructed from multiple low-resolution ToF measurements.
[0071]
[0072] Where, A t = CM t , t ≈ [1, 2,..., T], M t is the code format displayed on the DMD.
[0073] The reconstructed high-resolution amplitude and depth images are as follows:
[0074]
[0075] Where λ is the regularization parameter, and Ф(x) is the regularizer.
[0076]
[0077] Step S504, original ToF acquisition signal enhancement based on multi-frame super-resolution of sliding time window, as shown in Figure 7 , can include the following steps:
[0078] 1) The phasor diagram representation of each frame of ToF image is as formula (3),
[0079] Where a s is the signal amplitude (i.e. intensity), Ф s is the phase angle (i.e. depth), and Y ∈ C M×N .
[0080] 2) The LR ToF image sequence in the time window and the final reconstructed SR ToF are as follows:
[0081] y l = DW l x + n ll = 1,..., L
[0082] where D is a known down-sampling matrix, W is a moving warp matrix between the scene ground truth signal SR image and the l-th frame ToF measurement signal SR version, and n is an additive noise term.
[0083] 3) The probability distribution of the sequence of measurement signal images given the scene ground truth signal can be expressed as
[0084]
[0085] 4) The reconstruction of the scene ground truth signal, becomes an inference problem based on the posterior probability distribution given the known measurement signals, as follows:
[0086]
[0087] where q(x) is a Gaussian distribution.
[0088] Step S506, 3D depth map construction based on class-specific deformable 3D model structure inference. As shown in FIG. 6, the process can include the following steps Figure 8
[0089] 1) Learning class-specific deformable models from the 2D annotations available in the target detection dataset; using the set of images with additional annotations to estimate the camera projection parameters, which are then used together with the object contours to learn the 3D shape models. The learned shape models are able to deform to capture intra-class shape variations.
[0090] 2) Combining object detection, instance segmentation and view point estimation to identify, localize and estimate the pose of the objects in the image.
[0091] 3) Using the learned deformable 3D shape models, combined with the view point and instantiation information (e.g. observed object contours), to produce a "top-down" 3D reconstruction of the objects, mainly guided by class-level cues. The inferred 3D reconstruction best explains the shape of the observed object contours, respects the original shape of the generic shape (smoothness, continuity), and lies on the linear manifold of the class-level shape.
[0092] 4) Fusing the 3D shape and high-frequency local shape cues (e.g. recovering high-frequency shape details from the shading cues present in the image), to obtain a high-precision 3D reconstruction of the foreground salient objects (both corresponding to accurate dense depth information acquisition).
[0093] Step S508, confidence-based multi-orthogonal domain enhanced 3D depth map fusion. As shown in the following formula,
[0094]
[0095] where Ω'f (d) is the confidence of the f-point depth hypothesis d, P f,g,T (d) is the reasonable confidence of the g-propagation through the multi-frame (temporal domain enhancement) obtained depth data, P f,g,P (d) corresponds to the multi-phase (phase domain enhancement) related data, P f,g,S (d) corresponds to the semantic domain enhancement related data, P T (g) the confidence corresponding to the multi-frame (temporal domain enhancement) obtained depth data.
[0096] Compared with the existing light and thin AR glasses technology, the method and device of the embodiment of the present application first realize a high-precision high-resolution 3D depth construction method based on the very simplified lens composition of the glasses itself (single ToF lens, RGB lens and a DMD optical path), and realize high-precision and high-confidence 3D depth based on multi-orthogonal domain (based on multi-frame, multi-phase and semantic perception domain) enhancement by using the intelligent analysis and calculation unit of the mobile phone and watch. The light and thin AR glasses with this technology can well eliminate the problems of conventional visible block-shaped defects, large noise of perceived information and the like, greatly improve the use experience, and also facilitate the construction of high-precision 3D digital model of objects in the interaction process.
[0097] It should be noted that the embodiment of the present application does not have specific implementation of the active depth sensor. For example, the hardware composition of the embodiment of the present application can be ToF+DMD+RGB or ToF+LCoS+RGB or ToF+DMD+Gray arranged on the AR glasses.
[0098] The embodiment of the present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed, the steps in any one of the method embodiments described above are performed.
[0099] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk and various computer program storage media.
[0100] The embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any one of the method embodiments described above.
[0101] In one example embodiment, the electronic device described above can further include a transmission device connected to the processor and an input / output device connected to the processor.
[0102] The specific examples in the embodiments can refer to the examples described in the above embodiments and exemplary implementation, which will not be repeated here.
[0103] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, which can be realized by program codes executable by the computing devices, so that they can be stored in storage devices and executed by the computing devices, and in some cases, the steps shown or described can be executed in different order, or they can be made into individual integrated circuit modules, or multiple modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0104] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application shall be included in the protection scope of the present application.
Claims
1. A 3D depth map construction method, characterized by, The method comprises: performing optical multiplexing-based compressed sensing enhancement on the original ToF acquisition signal to obtain a first ToF acquisition signal; performing multi-frame super-resolution enhancement based on a sliding time window on a ToF image corresponding to the first ToF acquisition signal to reconstruct a high-resolution second ToF acquisition signal; constructing a 3D depth map of the ToF image according to the second ToF acquisition signal using a specific class of deformable 3D models; performing confidence-based multi-orthogonal domain enhancement 3D depth map fusion on the 3D depth map, wherein the multi-orthogonal domain enhancement comprises time domain enhancement, phase domain enhancement, and semantic domain enhancement; wherein the confidence-based multi-orthogonal domain enhancement 3D depth map fusion on the 3D depth map comprises: determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the phase domain enhancement; determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the time domain enhancement; determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the semantic domain enhancement; for each neighboring point in the plurality of neighboring points of each image point in the 3D depth map, fusing the propagation reasonable confidence under the phase domain enhancement, the propagation reasonable confidence under the time domain enhancement, the propagation reasonable confidence under the semantic domain enhancement, and the confidence under the phase domain enhancement, the confidence under the time domain enhancement, and the confidence under the semantic domain enhancement; for each image point in the 3D depth map, summing the fusion results of the plurality of neighboring points to obtain the fusion result of the image point.
2. The method of claim 1, wherein, The method comprises: changing the encoding on the time-multiplexed light encoding module during the multiple exposure processes of the ToF acquisition image to realize time-multiplexing of the scene, so as to reconstruct the first ToF acquisition signal with high-resolution amplitude and depth images.
3. The method of claim 2, wherein, The method comprises: forming a linear image of the scene on the time-multiplexed light encoding module of the original ToF acquisition path; during the multiple exposure processes of the ToF acquisition image, changing the encoding on the time-multiplexed light encoding module to realize time-multiplexing of the scene; relaying the linear image of the scene to a ToF acquisition plane; the ToF pixels on the ToF acquisition plane collect the linear image of the scene, and reconstruct the first ToF acquisition signal with high-resolution amplitude and depth images.
4. The method of claim 1, wherein, The method comprises: Within the sliding time window, a low-resolution ToF image is obtained according to the first ToF acquisition signal, and a high-resolution second ToF acquisition signal is generated by reconstructing the scene real signal based on the probability distribution of the scene real signal through the phasor diagram of multiple frames of the ToF image.
5. The method of claim 4, wherein, According to the first ToF acquisition signal, a low-resolution ToF image is obtained, and a high-resolution second ToF acquisition signal is generated by reconstructing the scene real signal based on the probability distribution of the scene real signal through the phasor diagram of multiple frames of the ToF image, comprising: According to the first ToF acquisition signal, the low-resolution ToF image is obtained, and the phasor diagram of the low-resolution ToF image is calculated; According to the sequence of low-resolution ToF images within the sliding time window, a sequence of reconstructed high-resolution ToF images is calculated; According to the sequence of high-resolution ToF images, the scene real signal is reconstructed based on the probability distribution of the scene real signal to generate a high-resolution second ToF acquisition signal.
6. The method of claim 1, wherein, According to the second ToF acquisition signal, a 3D depth map of the ToF image is constructed using a specific class deformable three-dimensional model, comprising: Learning a class-specific deformable model from 2D annotations of a target detection dataset; Combining object detection, instance segmentation and viewpoint estimation to identify, locate and estimate the pose of objects in the ToF image; Using the learned deformable model to perform 3D reconstruction of objects in the ToF image in combination with the estimated viewpoint and instantiation information obtained from the viewpoint estimation; Fusing three-dimensional shape and high-frequency local shape indication to reconstruct the 3D depth map containing multi-view observation angles.
7. A 3D depth map construction apparatus characterized by comprising: Comprising: A first enhancement module for performing optical multiplexing-based compressed sensing enhancement on the original ToF acquisition signal to obtain a first ToF acquisition signal; A second enhancement module for performing sliding time window-based multi-frame super-resolution enhancement on the ToF image corresponding to the first ToF acquisition signal to reconstruct a high-resolution second ToF acquisition signal; A construction module for constructing a 3D depth map of the ToF image using a specific class deformable three-dimensional model according to the second ToF acquisition signal; The fusion module is configured to perform confidence-based multi-orthogonal domain enhanced 3D depth map fusion on the 3D depth map, wherein the multi-orthogonal domain enhancement includes time domain enhancement, phase domain enhancement, and semantic domain enhancement; wherein the confidence-based multi-orthogonal domain enhanced 3D depth map fusion on the 3D depth map comprises: determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the phase domain enhancement; determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the time domain enhancement; determining the propagation reasonable confidence and the corresponding confidence of each image point in the 3D depth map under the semantic domain enhancement; for each neighboring point in a plurality of neighboring points of each image point in the 3D depth map, fusing the propagation reasonable confidence under the phase domain enhancement, the propagation reasonable confidence under the time domain enhancement, the propagation reasonable confidence under the semantic domain enhancement, and the confidence under the phase domain enhancement, the confidence under the time domain enhancement, and the confidence under the semantic domain enhancement; and for each image point in the 3D depth map, summing the fusion results of the plurality of neighboring points to obtain the fusion result of the image point.
8. The apparatus of claim 7, wherein, The first enhancement module comprises: The time division multiplexing light encoding module is arranged in an original ToF acquisition path, and is configured to change the encoding on the time division multiplexing light encoding module in a plurality of exposure processes of ToF acquisition images, so as to realize time division multiplexing of a scene.
9. An AR eyeglass, characterized by, The 3D depth map construction device according to any one of claims 7 to 8.
10. The AR glasses of claim 9, wherein, The AR glasses further comprise a calculation module, wherein the calculation module is an NPU acceleration chip in a smart terminal or a smart watch.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6. 12.An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Resolution Enhancement of Video Stream Based on Spatial and Temporal Correlation
US20110057933A1
Machine learning based model localization system
US20180189974A1
Method and system for time-of-flight imaging with high lateral resolution
US20200120299A1