Depth estimation methods, systems, and devices based on ToF and RGB fusion

By estimating the pose change information of the RGB camera and the dynamic extrinsic projection, the geometric consistency problem between the ToF sensor and the RGB camera is solved, enabling the generation of high-resolution depth maps, improving the accuracy and robustness of depth estimation, and meeting the real-time depth perception requirements of mobile devices.

CN122089804APending Publication Date: 2026-05-26LEQING POWER IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LEQING POWER IND CO LTD
Filing Date
2025-12-26
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, ToF sensors suffer from performance bottlenecks in mobile devices, including low signal-to-noise ratio, insufficient accuracy in long-distance depth measurement, and OIS floating lenses that disrupt geometric consistency, making it difficult to directly register and fuse depth information. Existing fusion calibration schemes lack robustness and accuracy, failing to meet the needs of practical applications.

Method used

By acquiring data from the ToF sensor and the RGB camera, the pose change information of the RGB camera caused by OIS technology is estimated, dynamic extrinsic parameters are calculated, and the initial depth map is projected onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters for spatial alignment and fusion to generate a high-resolution depth map.

Benefits of technology

It solves the problem of OIS floating lens destroying geometric consistency, enhances the ability to express depth details and registration accuracy, improves the accuracy and robustness of depth estimation, and meets the real-time depth perception requirements of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089804A_ABST
    Figure CN122089804A_ABST
Patent Text Reader

Abstract

This invention discloses a depth estimation method, system, and device based on Time-of-Flight (ToF) and RGB fusion. The method includes: acquiring an initial depth map from a ToF sensor and an RGB image from an RGB camera; estimating the pose change information of the RGB camera due to OIS technology based on the motion information of multiple feature points across multiple consecutive RGB images; calculating dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information; projecting the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain a projected depth map; spatially aligning the projected depth map with the RGB image; and fusing the aligned projected depth map with the RGB image to generate a high-resolution depth map. This invention solves the problem of geometric consistency degradation caused by floating OIS lenses, enhances depth detail representation and registration accuracy, improves the accuracy and robustness of depth estimation, and meets the real-time depth perception requirements of mobile devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and 3D perception technology, and in particular to a depth estimation method, system and device based on the fusion of Time-of-Flight (ToF) and RGB. Background Technology

[0002] With the rapid development of computational photography and 3D vision technologies, high-precision pixel-by-pixel depth information has become a core foundation for mobile devices to achieve applications such as 3D reconstruction, AR interaction, and depth-sensing image editing. Currently, mobile robots and drones commonly integrate multimodal camera systems, acquiring high-resolution color information through RGB cameras while directly measuring scene depth using ToF (Time-of-Flight) sensors. However, existing technologies still face the following key challenges in terms of hardware performance, geometric consistency preservation, and fusion methods: First, ToF sensors face significant performance bottlenecks. Constrained by the power supply of mobile devices' batteries, the active illumination power of ToF sensors is severely limited, resulting in a low signal-to-noise ratio and a significant decrease in the accuracy of long-distance depth measurement. Their spatial resolution is typically only 50,000 to 300,000 pixels, far lower than the 12 million to 64 million pixels of RGB cameras, resulting in sparse depth maps and insufficient detail representation.

[0003] Second, optical image stabilization technology disrupts the fixed geometric relationship between cameras. To acquire high-quality 2D images, main RGB cameras generally employ OIS (Optical Image Stabilization) technology, using a floating lens to compensate for camera body motion. The six degrees of freedom pose (including translation and rotation) of this floating lens dynamically changes with the focusing state and device orientation, and cannot be measured or read in real time electronically. This causes the rigid body transformation assumption between the ToF sensor and the RGB camera to fail, making it difficult to directly register and fuse depth information.

[0004] Third, existing fusion calibration schemes lack robustness and accuracy. Current mainstream methods either rely on synthetic data for training, resulting in inter-domain transfer errors; or they use offline calibration, which cannot adapt to dynamic extrinsic parameter changes caused by the OIS system; some techniques can only achieve scale-independent relative calibration, without considering the dynamic correction of lens distortion parameters, resulting in low absolute accuracy and poor generalization ability of the fused depth map, making it difficult to meet the needs of practical applications.

[0005] Therefore, there is an urgent need to provide a technical solution to address the above problems. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a depth estimation method, system, and device based on the fusion of ToF and RGB.

[0007] Firstly, this invention provides a depth estimation method based on the fusion of ToF and RGB, the technical solution of which is as follows: Acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image; Based on the motion information of multiple feature points in consecutive RGB images, the pose change information of the RGB camera caused by OIS technology is estimated. Based on the pose change information, calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera; Using the dynamic extrinsic parameters, the initial depth map is projected onto the image coordinate system of the RGB camera to obtain a projected depth map; The projection depth map is spatially aligned with the RGB image to obtain the aligned projection depth map; The aligned projection depth map and the RGB image are fused to generate a high-resolution depth map.

[0008] The beneficial effects of the depth estimation method based on ToF and RGB fusion of the present invention are as follows: The method of this invention solves the problem of geometric consistency disruption caused by OIS floating lenses, enhances the ability to express depth details and registration accuracy, improves the accuracy and robustness of depth estimation, and meets the real-time depth perception requirements of mobile devices.

[0009] Based on the above scheme, the depth estimation method based on ToF and RGB fusion of the present invention can be further improved as follows.

[0010] In one alternative approach, the step of acquiring the initial depth map captured by the ToF sensor and the RGB image captured by the RGB camera includes: The ToF sensor acquires raw depth data of the target scene, and simultaneously acquires the RGB image captured by the RGB camera in the target scene. The original depth data is demodulated based on the dual-frequency modulation principle to obtain the initial depth map and the corresponding confidence map; wherein, the confidence map is used as a weight to guide the fusion of the aligned projected depth map and the RGB image.

[0011] The advantages of adopting the above-mentioned optional approach are as follows: further obtaining the confidence map through dual-frequency modulation and demodulation, and using it as a weight to guide the fusion process, effectively improving the reliability of depth data, enhancing the robustness of the fusion algorithm to noise and uncertain regions, and optimizing the edge details and quality of the depth map.

[0012] In one alternative approach, the step of estimating the pose change information of the RGB camera due to OIS technology based on the motion information of multiple feature points across multiple consecutive RGB images includes: Extract the same multiple feature points from the consecutive multi-frame RGB images, and continuously track the motion trajectory of the multiple feature points in the consecutive multi-frame RGB images; Calculate the corresponding motion information based on the motion trajectory of each feature point to form a set of motion information for multiple feature points; Visual inertial odometry technology is used to separate the device motion of the RGB camera from the lens motion caused by OIS technology from the motion information set; Based on the lens motion of the separated RGB camera, the pose change information of the RGB camera is estimated.

[0013] The advantages of adopting the above-mentioned optional methods are as follows: further using visual inertial odometry technology to separate the overall motion of the device from the independent motion of the OIS lens from the motion of feature points, accurately estimating the pose change of the floating lens, providing accurate input data for dynamic extrinsic parameter calculation, and improving the response accuracy of OIS compensation.

[0014] In one alternative approach, the step of calculating the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information includes: The initial extrinsic parameter matrix between the ToF sensor and the RGB camera is obtained through offline calibration. Based on the rotation and translation parameters in the pose change information, construct the pose transformation matrix of the RGB camera; The dynamic extrinsic parameters between the ToF sensor and the RGB camera are calculated by multiplying the initial extrinsic parameter matrix with the pose transformation matrix.

[0015] The advantages of using the above optional method are: by further combining the initial extrinsic parameter matrix of offline calibration with the real-time pose transformation matrix, the dynamic extrinsic parameters between the ToF and RGB cameras can be calculated and updated quickly, ensuring that the projection transformation can adapt to the OIS lens floating in real time and maintain the accuracy of geometric registration.

[0016] In one alternative approach, the step of projecting the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain a projected depth map includes: Based on the dynamic extrinsic parameters, a projection transformation relationship is established from the ToF sensor coordinate system to the RGB camera image coordinate system; Each depth pixel in the initial depth map is transformed to the image coordinate system of the RGB camera using the projection transformation relationship to obtain the coordinates of each three-dimensional point in the RGB camera image coordinate system; By applying perspective projection, the coordinates of each three-dimensional point are converted into two-dimensional image coordinates corresponding to the RGB camera; The depth value of each depth pixel in the initial depth map is assigned to the corresponding two-dimensional image coordinates to generate the projected depth map.

[0017] The advantages of adopting the above optional method are: further establishing an accurate projection transformation relationship based on dynamic extrinsic parameters, mapping the low-resolution depth map pixel by pixel to the RGB image coordinate system, generating a projection depth map consistent with the RGB image perspective through perspective projection, and providing a spatial reference for subsequent fusion.

[0018] In one alternative approach, the step of spatially aligning the projected depth map with the RGB image to obtain an aligned projected depth map includes: The dense optical flow field between the projected depth map and the RGB image is calculated using the optical flow method, and the dense optical flow field contains the displacement vector of each pixel; Based on the displacement vector of each pixel in the dense optical flow field, pixel-by-pixel displacement compensation is performed on the projection depth map to generate a displacement-compensated projection depth map. The displacement-compensated projection depth map is subjected to hole filling and edge smoothing processing to obtain the processed projection depth map. The processed projection depth map is aligned with the RGB image at the pixel level to obtain the aligned projection depth map.

[0019] The advantages of using the above-mentioned optional methods are as follows: by further utilizing the optical flow method to calculate the dense displacement field and perform pixel-by-pixel compensation, combined with hole filling and edge smoothing, projection misalignment and hole artifacts can be effectively eliminated, achieving fine spatial alignment between the projection depth map and the RGB image, and improving registration accuracy.

[0020] In one alternative approach, the step of fusing the aligned projected depth map and the RGB image to generate a high-resolution depth map includes: Construct a fusion model that includes a multi-scale feature encoder, a feature fusion unit, and a deep regressor; The aligned projection depth map and the RGB image are input into the multi-scale feature encoder to extract depth features and RGB features, respectively. Using the confidence map as weights, the feature fusion unit performs weighted fusion of the depth features and the RGB features to obtain a feature-related volume that includes the correlation between the depth features and the RGB features; The feature-related volume is processed by the depth regressor to generate the high-resolution depth map.

[0021] The advantages of adopting the above optional approach are as follows: a multi-scale feature encoder is further constructed to extract depth and RGB features, and a confidence map weighted fusion strategy is adopted to generate a high-resolution depth map through a depth regressor, making full use of RGB texture details to enhance the spatial resolution and edge sharpness of the depth map.

[0022] Secondly, this invention provides a depth estimation system based on the fusion of ToF and RGB, the technical solution of which is as follows: The acquisition module is used to acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image; The processing module is used to estimate the pose change information of the RGB camera caused by OIS technology based on the motion information of multiple feature points in consecutive RGB images. The calculation module is used to calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information; The projection module is used to project the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain a projected depth map; An alignment module is used to spatially align the projection depth map with the RGB image to obtain an aligned projection depth map; A generation module is used to fuse the aligned projection depth map and the RGB image to generate a high-resolution depth map.

[0023] The beneficial effects of the depth estimation system based on ToF and RGB fusion of the present invention are as follows: The system of this invention solves the problem of geometric consistency disruption caused by OIS floating lenses, enhances the ability to express depth details and registration accuracy, improves the accuracy and robustness of depth estimation, and meets the real-time depth perception requirements of mobile devices.

[0024] Thirdly, the technical solution of an electronic device according to the present invention is as follows: It includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the depth estimation method based on ToF and RGB fusion of the present invention.

[0025] Fourthly, the technical solution of a computer-readable storage medium provided by the present invention is as follows: The computer-readable storage medium stores instructions that, when read, cause the computer-readable storage medium to perform the steps of the depth estimation method based on ToF and RGB fusion of the present invention.

[0026] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0027] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a schematic flowchart of an embodiment of a depth estimation method based on ToF and RGB fusion according to the present invention; Figure 2 This is a schematic diagram illustrating the overall principle. Figure 3 This is a schematic diagram of an embodiment of a depth estimation system based on ToF and RGB fusion according to the present invention; Figure 4 This is a schematic diagram of an embodiment of an electronic device according to the present invention. Detailed Implementation

[0028] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0029] Figure 1 This diagram illustrates a flowchart of an embodiment of a depth estimation method based on Time-of-Flight (ToF) and RGB fusion provided by the present invention. This ToF / RGB fusion-based depth estimation method can be executed by electronic devices such as terminal devices or servers. The terminal device can be any fixed or mobile terminal, such as user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the ToF / RGB fusion-based depth estimation method by having its processor call computer-readable instructions stored in memory. Figure 1 As shown, it includes the following steps: S1. Acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image.

[0030] In this context, a ToF sensor refers to a depth sensing device that measures distance by calculating the time-of-flight of light. For example, a ToF sensor on a drone emits a laser beam at a building below and receives the reflected signal, calculating the round-trip time of the light wave to obtain the distance data between the building's roof and the sensor. An initial depth map refers to a two-dimensional image generated directly by a ToF sensor, containing distance information for various points in the scene. For example, a ToF sensor generates a 320x240 pixel grayscale image of a city street, where each pixel value represents the distance of the corresponding building from the sensor. An RGB camera refers to an imaging device capable of capturing red, green, and blue color information. For example, a drone's main camera captures color images of city streets, including the color and texture information of buildings, roads, and vehicles. An RGB image refers to a two-dimensional image captured by an RGB camera, containing red, green, and blue color information. For example, an RGB camera captures city street images with a resolution of 3840x2160 pixels, clearly showing the texture of building facades and the details of road markings. The resolution of an initial depth map refers to the number of pixels contained within a unit area; for example, a depth map generated by a ToF sensor has a resolution of 320x240 pixels, or 76,800 depth measurement points. The resolution of an RGB image refers to the number of pixels contained within a unit area; for example, an image captured by an RGB camera has a resolution of 3840x2160 pixels, or 8,294,400 color sampling points.

[0031] S2. Based on the motion information of multiple feature points in consecutive RGB images, estimate the pose change information of the RGB camera caused by OIS technology.

[0032] Feature points refer to pixel locations in an image that possess significant texture or corner characteristics; for example, the corners of building windows, road intersections, and vehicle outlines in an urban street image. Continuous multi-frame RGB images refer to multiple images continuously acquired by an RGB camera over time; for example, a drone continuously captures 30 frames per second of urban street images during flight, recording scene changes. Motion information refers to the positional changes of feature points within consecutive image frames; for example, a building corner feature might have coordinates (500, 300) in the first frame and change to (505, 298) in the second frame—this coordinate change constitutes motion information. OIS technology refers to optical image stabilization, which compensates for device shake through a floating lens; for example, when a drone encounters turbulent airflow, the internal lens of the RGB camera moves in the opposite direction to maintain image stability, but this changes the camera's optical center position. Pose change information refers to the camera's position and orientation changes in three-dimensional space; for example, an RGB camera might experience a 2mm translation and a 0.5-degree rotation due to the floating OIS lens and the drone's attitude adjustment—these six degrees of freedom parameters constitute pose change information.

[0033] S3. Calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information.

[0034] Among them, dynamic extrinsic parameters refer to the relative pose relationship between the ToF sensor and the RGB camera that changes in real time with the OIS state; for example, the rotation and translation matrix between the ToF sensor and the RGB camera is updated in real time based on the lens displacement and drone attitude changes caused by OIS.

[0035] S4. Using the dynamic extrinsic parameters, the initial depth map is projected onto the image coordinate system of the RGB camera to obtain a projected depth map.

[0036] The image coordinate system of the RGB camera refers to a two-dimensional coordinate system based on the imaging plane of the RGB camera; for example, a pixel coordinate system with the top left corner of the RGB image as the origin (0,0), the positive X-axis pointing to the right, and the positive Y-axis pointing downwards. The projected depth map refers to the depth data after mapping the initial depth map to the RGB image coordinate system; for example, projecting a 320x240 ToF depth map onto a 3840x2160 RGB image coordinate system generates a depth map with the same viewpoint.

[0037] S5. Spatially align the projection depth map with the RGB image to obtain the aligned projection depth map.

[0038] The aligned projection depth map refers to the depth map that is aligned with the pixels of the RGB image after spatial registration. For example, after optical flow alignment, the depth of the building edge completely coincides with the building outline in the RGB image.

[0039] S6. The aligned projection depth map and the RGB image are fused to generate a high-resolution depth map.

[0040] Among them, high-resolution depth map refers to: depth map with RGB image resolution generated after fusion; for example, the final output 3840x2160 pixel depth map contains both accurate distance information and fine spatial details.

[0041] The technical solution of this embodiment solves the problem of OIS floating lens destroying geometric consistency, enhances the ability to express depth details and registration accuracy, improves the accuracy and robustness of depth estimation, and meets the real-time depth perception requirements of mobile devices.

[0042] In one alternative approach, S1 specifically includes: The ToF sensor acquires raw depth data of the target scene, and simultaneously acquires the RGB image captured by the RGB camera in the target scene.

[0043] The target scene refers to the physical space environment being collected; for example, an urban street environment containing high-rise buildings, roads, and green belts. Raw depth data refers to the unprocessed phase or time measurements directly output by the ToF sensor; for example, undemodulated phase difference data collected by the ToF sensor, which includes noise and multiple reflection interference.

[0044] Specifically, a ToF sensor emits modulated light waves into the target scene and receives reflected signals to collect raw depth data including phase measurements; an RGB camera simultaneously collects red, green, and blue color information of the target scene to generate an RGB image; a hardware trigger signal is used to synchronize the working timing of the ToF sensor and the RGB camera, ensuring that the raw depth data and the RGB image are aligned in terms of acquisition time; the ToF sensor uses global shutter exposure control, and the RGB camera uses global shutter mode to achieve precise timing matching between the sensors.

[0045] The original depth data is demodulated based on the dual-frequency modulation principle to obtain the initial depth map and the corresponding confidence map.

[0046] The confidence map serves as a weight to guide the fusion of the aligned projection depth map and the RGB image.

[0047] The dual-frequency modulation principle refers to a signal processing method that uses two different frequencies to modulate light waves to resolve distance ambiguity. For example, a ToF sensor simultaneously illuminates a scene using 10MHz and 100MHz modulation frequencies, obtaining unambiguous depth through dual-frequency calculation. A confidence map is an auxiliary image characterizing the reliability of depth measurements; for example, a grayscale image of the same size as the depth map, where higher brightness indicates a more reliable depth value at the corresponding location, with higher confidence levels for building wall areas and lower confidence levels for glass curtain wall reflection areas.

[0048] Specifically, the original depth data is demodulated based on the dual-frequency modulation principle. The first modulation frequency and the second modulation frequency are used to perform phase demodulation on the original depth data respectively. The absolute phase value of each pixel is calculated by the phase unwrapping algorithm. The depth value of each pixel is calculated according to the relationship between the absolute phase value and the speed of light to generate an initial depth map. At the same time, the confidence level of each pixel is calculated based on the phase consistency of the two modulation frequencies. The confidence level is quantized into grayscale values ​​to generate a confidence map with the same size as the initial depth map. The first modulation frequency is 10MHz, the second modulation frequency is 100MHz, and the phase consistency is calculated by the variance of the phase difference between the two frequencies.

[0049] In the above-mentioned optional methods, a confidence map is further obtained through dual-frequency modulation and demodulation, and used as a weight to guide the fusion process, which effectively improves the reliability of depth data, enhances the robustness of the fusion algorithm to noise and uncertain regions, and optimizes the edge details and quality of the depth map.

[0050] In one alternative approach, S2 specifically includes: Extract the same multiple feature points from the consecutive multi-frame RGB images, and continuously track the motion trajectory of the multiple feature points in the consecutive multi-frame RGB images.

[0051] Each feature point corresponds to a motion trajectory, which refers to the position sequence of the feature point in consecutive frames. For example, the corner feature of a building forms a movement path from (500,300) to (505,298) and then to (510,295) in 10 consecutive frames of images.

[0052] Specifically, multiple feature points are extracted from the first frame of the RGB image, and the FAST corner detection algorithm is used to identify corner features in the image. In subsequent consecutive frames of RGB images, the Lucas-Kanade optical flow method is used to track the feature points frame by frame, and the pixel position change of each feature point in the sequence of images is calculated. The coordinate sequence of each feature point in all frames is recorded to form the motion trajectory of each feature point. The feature point extraction and tracking process is based on the same feature descriptor to ensure that the same set of feature points is tracked in consecutive frames.

[0053] The motion information corresponding to each feature point is calculated based on its motion trajectory, forming a set of motion information for multiple feature points.

[0054] Motion information refers to the data on the position, velocity, and orientation changes of feature points in consecutive image frames. For example, in consecutive frames captured by a drone, the coordinates of a building corner feature change from (500, 300) in the first frame to (505, 298) in the second frame. The calculated horizontal displacement of this feature point is 5 pixels, the vertical displacement is -2 pixels, the speed is approximately 3.6 pixels per second, and the direction of movement is towards the lower right. The motion information set refers to a dataset summarizing the motion information of multiple feature points. For example, in 10 consecutive frames of images of a city street captured by a drone, the position change sequences, velocity vectors, and trajectories of 200 stable feature points (including building corners, road signs, and vehicle feature points) are extracted. These data together constitute a multidimensional dataset containing spatiotemporal motion features.

[0055] Specifically, motion information is calculated based on the motion trajectory of each feature point. The coordinate sequence of each feature point in multiple consecutive RGB images is differentially calculated to obtain the displacement vector of each feature point between adjacent frames. The velocity magnitude and direction angle of each feature point are calculated based on the displacement vector and the time interval between frames. The displacement vector sequence, velocity sequence and direction angle sequence of each feature point are combined into the motion information of a single feature point. The motion information of all feature points is collected to form a motion information set containing displacement, velocity and direction data. Among them, the displacement vector is calculated by the coordinate difference between adjacent frames, the velocity magnitude is calculated by the ratio of the displacement vector magnitude to the time interval, and the direction angle is calculated by the arctangent of the displacement vector.

[0056] Visual inertial odometry (VIO) technology is used to separate the device motion of the RGB camera from the motion information set, as well as the lens motion caused by OIS technology.

[0057] Visual inertial odometry (VIOMA) technology refers to a method of estimating motion by combining visual features and inertial measurement unit (IMU) data; for example, using changes in RGB image feature points and gyroscope data from the UAV's IMU to calculate the camera's trajectory. Device motion refers to the displacement and rotation of the entire UAV platform; for example, the overall movement of the camera when the UAV flies from east to west of a street. Lens motion refers to the independent movement of the floating lens relative to the camera body in an OIS system; for example, to counteract UAV vibrations, the OIS lens makes minute translations and tilts within the camera.

[0058] Specifically, the displacement vectors of feature points in the motion information set are synchronized with the angular velocity and linear acceleration data collected by the inertial measurement unit (IMU). A visual-inertial joint optimization model is established, including device motion parameters and lens motion parameters. The device motion parameters include a six-degree-of-freedom pose transformation, and the lens motion parameters include a three-degree-of-freedom translation of the OIS lens. The weighted sum of the feature point reprojection error and the IMU pre-integration error is minimized using a nonlinear optimization method, while simultaneously solving for the device motion trajectory and the lens motion trajectory. The visual-inertial joint optimization model models the OIS lens motion as an independent translational transformation relative to the camera body, and the device motion as a rigid body transformation of the camera body.

[0059] Based on the lens motion of the separated RGB camera, the pose change information of the RGB camera is estimated.

[0060] Specifically, the translation vector and rotation component in the lens motion parameters are converted into pose transformation representations in the RGB camera coordinate system; the displacement vector of the RGB camera optical center in three-dimensional space is calculated based on the three-degree-of-freedom translation in the lens motion parameters; the orientation change angle of the RGB camera optical axis is calculated based on the rotation component in the lens motion parameters; and six-degree-of-freedom pose change information is constructed by combining the displacement vector and the orientation change angle. The pose change information includes three translation parameters and three rotation parameters. The pose change information is represented in the form of a transformation matrix or Euler angles and is used for subsequent dynamic extrinsic parameter calculations.

[0061] In the above-mentioned optional methods, visual inertial odometry technology is further used to separate the overall motion of the device from the independent motion of the OIS lens from the motion of the feature points, accurately estimate the pose change of the floating lens, provide accurate input data for dynamic extrinsic parameter calculation, and improve the response accuracy of OIS compensation.

[0062] In one alternative approach, S3 specifically includes: The initial extrinsic parameter matrix between the ToF sensor and the RGB camera is obtained through offline calibration.

[0063] Offline calibration refers to the sensor parameter calibration process performed before the equipment leaves the factory; for example, using a checkerboard calibration board in a calibration laboratory to pre-measure the relative positional relationship between the ToF sensor and the RGB camera. The initial extrinsic parameter matrix refers to the fixed pose transformation matrix between the ToF sensor and the RGB camera obtained through offline calibration; for example, a 4x4 matrix describing the initial rotation and translation relationship of the ToF sensor relative to the RGB camera.

[0064] Specifically, a checkerboard calibration board is placed in the common field of view of the ToF sensor and the RGB camera. Color images of the calibration board are acquired from multiple angles using the RGB camera, and the intrinsic parameter matrix and distortion coefficients of the RGB camera are calculated using the Zhang Zhengyou calibration method. Depth point cloud data of the calibration board is acquired using the ToF sensor, and the intrinsic parameters of the ToF sensor are calculated based on point cloud plane fitting. The two-dimensional pixel coordinates of the checkerboard corner points in the RGB image are extracted, and the three-dimensional world coordinates of the corresponding corner points in the calibration board coordinate system are calculated. The depth point cloud data acquired by the ToF sensor is registered with the three-dimensional model of the calibration board to establish the transformation relationship between the ToF sensor coordinate system and the calibration board coordinate system. Based on the intrinsic parameters of the RGB camera and the ToF sensor, the transformation matrix between the RGB camera coordinate system and the calibration board coordinate system is solved using the perspective n-point algorithm. The initial extrinsic parameter matrix from the ToF sensor coordinate system to the RGB camera coordinate system is calculated. The calibration board uses a 7×9 checkerboard pattern with a corner point spacing of 30 mm, and the acquisition angles include pitch, yaw, and roll attitudes.

[0065] Based on the rotation and translation parameters in the pose change information, the pose transformation matrix of the RGB camera is constructed.

[0066] The rotation parameters refer to the three degrees of freedom describing the change in orientation between coordinate systems; for example, Euler angles (1.2°, 0.8°, -0.3°) represent the yaw, pitch, and roll angles of the RGB camera relative to the ToF sensor. The translation parameters refer to the three degrees of freedom describing the change in position between coordinate systems; for example, a three-dimensional vector (15.2, -8.3, 2.1 mm) represents the position offset of the RGB camera's optical center relative to the ToF sensor reference point. The pose transformation matrix refers to the 4x4 transformation matrix describing the camera pose change caused by OIS and UAV attitude changes; for example, a homogeneous coordinate transformation matrix including rotation and translation caused by OIS lens floating and UAV attitude adjustments.

[0067] Specifically, the rotation parameters are converted into a 3×3 rotation matrix. When the rotation parameters are represented by Euler angles, the rotation matrix is ​​calculated by multiplying the three basic rotation matrices. When the rotation parameters are represented by rotation vectors, the rotation matrix is ​​calculated using the Rodrigues formula. The translation parameters are represented as 3×1 translation vectors. The rotation matrix and translation vectors are combined into a 4×4 homogeneous coordinate transformation matrix. The rotation matrix is ​​placed in the upper left 3×3 submatrix, and the translation vector is placed in the upper right 3×1 submatrix. The last row is padded with [0,0,0,1] to complete the construction of the pose transformation matrix. The pose transformation matrix represents the coordinate transformation relationship of the RGB camera from the initial pose to the current pose due to OIS technology.

[0068] The dynamic extrinsic parameters between the ToF sensor and the RGB camera are calculated by multiplying the initial extrinsic parameter matrix with the pose transformation matrix.

[0069] Specifically, the 4×4 initial extrinsic parameter matrix, representing the initial relative pose relationship between the ToF sensor and the RGB camera, is multiplied with the 4×4 pose transformation matrix, representing the pose change of the RGB camera due to OIS technology, according to matrix multiplication rules. In the matrix multiplication operation, each row element of the initial extrinsic parameter matrix is ​​multiplied with each column element of the pose transformation matrix and then summed to obtain a new 4×4 homogeneous coordinate transformation matrix as the dynamic extrinsic parameter. The upper left 3×3 submatrix of the dynamic extrinsic parameter matrix represents the rotation relationship between the ToF sensor and the current RGB camera coordinate system, the upper right 3×1 submatrix represents the position offset of the ToF sensor from the current RGB camera coordinate system, and the last row maintains the homogeneous coordinate form of [0,0,0,1]. The dynamic extrinsic parameter reflects the relative pose relationship between the ToF sensor and the RGB camera in real time, including the influence of OIS lens motion.

[0070] In the above-mentioned optional methods, the initial extrinsic parameter matrix calibrated offline is further multiplied with the real-time pose transformation matrix to achieve rapid calculation and updating of dynamic extrinsic parameters between the ToF and RGB cameras, ensuring that the projection transformation can adapt to the OIS lens floating in real time and maintain the accuracy of geometric registration.

[0071] In one alternative approach, S4 specifically includes: Based on the dynamic extrinsic parameters, a projection transformation relationship is established from the ToF sensor coordinate system to the RGB camera image coordinate system.

[0072] The projection transformation relationship refers to the mathematical relationship that maps a point from the ToF coordinate system to the RGB image coordinate system; for example, the perspective transformation equation that projects a ToF 3D point onto an RGB 2D image plane based on dynamic extrinsic parameters.

[0073] Each depth pixel in the initial depth map is transformed into the image coordinate system of the RGB camera through the projection transformation relationship to obtain the coordinates of each three-dimensional point in the RGB camera image coordinate system.

[0074] In this context, a depth pixel refers to a single pixel in the depth map and the distance value it contains; for example, a pixel with coordinates (100, 80) in the initial depth map has a grayscale value of 150, corresponding to an actual distance of 25 meters. Three-dimensional point coordinates refer to the position of a point in three-dimensional space; for example, a point (10.5, 8.2, 25.0) in the ToF coordinate system represents a position 10.5 meters in the X direction, 8.2 meters in the Y direction, and 25.0 meters in the Z direction from the origin.

[0075] By applying perspective projection, the coordinates of each three-dimensional point are converted into the two-dimensional image coordinates corresponding to the RGB camera.

[0076] Perspective projection refers to the geometric transformation that maps a point in three-dimensional space to a two-dimensional image plane; for example, projecting a three-dimensional point (10.5, 8.2, 25.0) in the ToF coordinate system onto a two-dimensional point (1920, 1080) in the RGB image coordinate system. Two-dimensional image coordinates refer to the pixel position represented in the two-dimensional image plane; for example, the point (1920, 1080) in the RGB image coordinate system represents a position 1920 pixels to the right and 1080 pixels downwards from the top-left corner of the image.

[0077] Specifically: 1) Obtain the intrinsic parameter matrix of the RGB camera. ,in , Represents the focal length along the x-axis. Indicates the focal length along the y-axis. Indicates the x-coordinate of the principal point. 1) Represent the ordinate of the principal point; all parameters are expressed in pixels; 2) The dynamic extrinsic parameter matrix With intrinsic parameter matrix Multiplication yields the projection transformation matrix. ,Right now ,in It is a 4×4 homogeneous transformation matrix, representing the coordinate transformation relationship from the ToF sensor to the RGB camera; 3) Based on the projection transformation matrix Establish projection transformation relationship ,in It is the homogeneous coordinate (4×1 vector) of a three-dimensional point in the ToF sensor coordinate system. It is the two-dimensional homogeneous pixel coordinate (3×1 vector) in the RGB camera image coordinate system; 4) Perform coordinate transformation calculation through projection transformation relationship: Left multiplication Obtain the coordinates of a 3D point in the RGB camera coordinate system ;Will Left multiplication get ;right Homogeneous coordinate normalization yields two-dimensional pixel coordinates. The projection transformation relationship fully describes the coordinate mapping process from the ToF sensor coordinate system to the RGB camera image coordinate system.

[0078] The depth value of each depth pixel in the initial depth map is assigned to the corresponding two-dimensional image coordinates to generate the projected depth map.

[0079] The depth value of a depth pixel refers to the actual distance measurement value represented by the depth pixel; for example, a grayscale value of 180 for a pixel in the depth map corresponds to an actual distance of 30 meters.

[0080] Specifically, for each depth pixel in the initial depth map, the two-dimensional image coordinates of the depth pixel in the RGB camera image coordinate system are calculated according to the projection transformation relationship; the calculated two-dimensional image coordinates are rounded to the nearest integer pixel position; the depth value of the depth pixel is assigned to the pixel position in the projection depth map corresponding to the rounded two-dimensional image coordinates; wherein, the projection depth map is initialized as a two-dimensional array with the same resolution as the RGB image, and all positions are initialized to zero or invalid values. The assignment process overwrites the initial values ​​to form a projection depth map containing depth information.

[0081] In the above-mentioned optional methods, a precise projection transformation relationship is further established based on dynamic extrinsic parameters, and the low-resolution depth map is mapped pixel by pixel to the RGB image coordinate system. A projection depth map consistent with the RGB image viewpoint is generated through perspective projection, providing a spatial reference for subsequent fusion.

[0082] In one alternative approach, S5 specifically includes: The dense optical flow field between the projected depth map and the RGB image is calculated using the optical flow method. The dense optical flow field contains the displacement vector of each pixel.

[0083] Optical flow refers to a computer vision method that calculates the motion vectors of pixels between adjacent image frames. For example, by comparing two consecutive frames of street images captured by a drone, the direction and distance of movement of each pixel from the first frame to the second frame can be calculated. Dense optical flow field refers to the set of motion vectors of all pixels covering the entire image. For example, an optical flow field with a resolution of 3840x2160 contains 8.29 million motion vectors, each describing the movement of the corresponding pixel.

[0084] Specifically, the Farneback polynomial expansion algorithm is used to construct multi-scale pyramids for the projected depth map and the RGB image respectively. At each level of the pyramid, the displacement vector of each pixel between the two images is calculated, where the displacement vector includes horizontal displacement components and vertical displacement components. Through iterative optimization of the pyramid levels from coarse to fine, the displacement vector estimation of each pixel is refined level by level. Finally, a dense optical flow field with the same resolution as the RGB image is generated. This optical flow field is stored in the form of a two-dimensional vector field, where each pixel position corresponds to a displacement vector, and the two components of the displacement vector represent the displacement of the pixel in the horizontal and vertical directions, respectively.

[0085] Based on the displacement vector of each pixel in the dense optical flow field, pixel-by-pixel displacement compensation is performed on the projection depth map to generate a displacement-compensated projection depth map.

[0086] In this context, the pixel displacement vector refers to a two-dimensional vector describing the magnitude and direction of a single pixel's movement between two frames. For example, a pixel displacement vector of (3.2, -1.8) indicates that the pixel moves 3.2 pixels to the right and 1.8 pixels upward. Pixel-by-pixel displacement compensation refers to the process of correcting the position of each depth pixel based on the optical flow vector. For example, adjusting the position of corresponding points in the depth map according to the displacement vector of each pixel to align it with the RGB image. The displacement-compensated projected depth map refers to the depth map after pixel-by-pixel displacement correction. For example, after adjusting the projected depth map using the optical flow vector, the building outline is basically aligned with the edges in the RGB image, but a small number of holes remain.

[0087] Specifically, for each pixel in the projection depth map, the displacement vector corresponding to the pixel position is obtained from the dense optical flow field. The displacement vector includes a horizontal displacement component and a vertical displacement component. The new coordinates of the pixel after displacement compensation are calculated based on the displacement vector. The new coordinates are obtained by adding the horizontal and vertical components of the displacement vector to the original coordinates. The depth value of the original pixel in the projection depth map is assigned to the pixel position corresponding to the new coordinates in the displacement-compensated projection depth map. The displacement-compensated projection depth map is initialized as a two-dimensional array of the same size as the projection depth map. All positions are initialized with invalid values. The assignment process overwrites the initial values ​​to form the displacement-compensated projection depth map. For multiple original pixels mapped to the same new coordinates, the nearest neighbor priority principle is used to retain the depth value processed first.

[0088] The displacement-compensated projection depth map is then subjected to hole filling and edge smoothing processing to obtain the processed projection depth map.

[0089] Hole filling refers to the data processing procedure of completing missing areas in the depth map; for example, using neighborhood depth values ​​to interpolate and fill areas without depth values ​​caused by projection and displacement compensation. Edge smoothing refers to the signal processing of eliminating jagged edges and noise in the depth map; for example, using bilateral filtering to smooth noise in the depth map while maintaining the sharpness of building edges. The processed projected depth map refers to the depth map after hole filling and edge smoothing; for example, after complete post-processing, the displacement-compensated depth map forms a depth map without holes and with clear edges.

[0090] Specifically, a hole-filling algorithm based on neighborhood propagation is adopted to fill the hole region with the mean of the effective depth values ​​within the pixel's neighborhood. During the hole-filling process, firstly, the positions of pixels with invalid depth values ​​in the displacement-compensated projected depth map are identified as hole pixels. Then, each hole pixel is traversed, and the effective depth values ​​within its 3×3 neighborhood are checked. When the number of effective depth values ​​in the neighborhood exceeds a threshold, the arithmetic mean of these effective depth values ​​is calculated as the filling value for the hole pixel. After hole filling is completed, bilateral filtering is applied to smooth the edges of the depth map. The weight calculation of bilateral filtering considers both pixel spatial distance and depth value similarity. The standard deviation of spatial distance is set to 1.5 pixels, the standard deviation of depth value is set to 0.05 meters, and the filtering window size is 5×5 pixels. After hole filling and edge smoothing, a processed projected depth map with continuous depth distribution and smooth edges is obtained.

[0091] The processed projection depth map is aligned with the RGB image at the pixel level to obtain the aligned projection depth map.

[0092] Pixel-level coordinate alignment refers to the process of fully registering the depth map with the RGB image at the pixel level; for example, adjusting the depth map so that the depth outline of a building window coincides with the window outline in the RGB image at the pixel level.

[0093] Specifically, a coordinate mapping relationship is established between the processed projection depth map and the RGB image, based on the same image coordinate system shared by the two images. Each pixel coordinate in the processed projection depth map is directly mapped to the corresponding pixel coordinate in the RGB image, establishing a one-to-one coordinate relationship. The processed projection depth map is resampled using bilinear interpolation to ensure that the pixel coordinates of the depth map correspond completely to the pixel coordinates of the RGB image. The bilinear interpolation uses the depth values ​​of the four nearest neighbor pixels around the target pixel for weighted calculation, with the weights determined by the distance between pixels. The final aligned projection depth map has the same resolution and pixel coordinate correspondence as the RGB image, and each pixel position in the RGB image contains the corresponding depth value.

[0094] In the above-mentioned optional methods, the optical flow method is further used to calculate the dense displacement field and perform pixel-by-pixel compensation. Combined with hole filling and edge smoothing, projection misalignment and hole artifacts are effectively eliminated, and fine spatial alignment between the projection depth map and the RGB image is achieved, thereby improving the registration accuracy.

[0095] In one alternative approach, S6 specifically includes: Construct a fusion model that includes a multi-scale feature encoder, a feature fusion unit, and a deep regressor.

[0096] In this context, the fusion model refers to a complete neural network architecture that achieves multi-source data fusion; for example, an end-to-end deep learning network consisting of a multi-scale feature encoder, a feature fusion unit, and a depth regressor. A multi-scale feature encoder is a neural network module capable of extracting features at different scales; for example, an encoder network containing convolutional and pooling layers that simultaneously extracts features at different scales, such as the overall shape of a building and window details. A feature fusion unit is a neural network component responsible for fusing multimodal features; for example, a network module using an attention mechanism that weights and fuses depth features and RGB features based on a confidence map. A depth regressor is the output layer of a neural network that regresses depth maps from the fused features; for example, a network tail consisting of fully connected layers and deconvolutional layers that converts the fused features into high-resolution depth maps.

[0097] The aligned projection depth map and the RGB image are input into the multi-scale feature encoder to extract depth features and RGB features, respectively.

[0098] Depth features refer to abstract feature representations extracted from depth maps; for example, feature vectors representing the geometry of buildings extracted from an initial depth map using a convolutional neural network. RGB features refer to abstract feature representations extracted from RGB images; for example, feature vectors representing building materials and road textures extracted from a street RGB image using a convolutional neural network.

[0099] Specifically, the RGB image is input into the RGB feature extraction branch, which consists of three convolutional layers. Each convolutional layer uses a 3×3 convolutional kernel and a ReLU activation function. Feature map downsampling is achieved through a max pooling layer with a stride of 2. The aligned projected depth map is input into the depth feature extraction branch, which has the same structure. The convolutional layer weights of the two branches are initialized independently and do not share parameters. In each downsampling stage of the multi-scale feature encoder, RGB feature maps with 64, 128, and 256 channels are output from the RGB feature extraction branch, respectively, and depth feature maps with the same number of channels are output from the depth feature extraction branch. Finally, three sets of RGB feature maps and depth feature maps at different scales are obtained, with feature map sizes of 1 / 2, 1 / 4, and 1 / 8 of the input image, respectively.

[0100] Using the confidence map as weights, the feature fusion unit performs weighted fusion of the depth features and the RGB features to obtain a feature-related volume that includes the correlation between the depth features and the RGB features.

[0101] The fused multimodal features refer to the joint feature representation after fusing depth features and RGB features; for example, a feature tensor that is attention-weighted after concatenating depth geometric features and RGB texture features in the channel dimension. The feature-related volume refers to the three-dimensional feature tensor formed by concatenating depth features and RGB features in the channel dimension; for example, in the depth estimation of urban blocks by a drone, the 256-dimensional depth features and 128-dimensional RGB features extracted by the multi-scale feature encoder are concatenated in the feature fusion unit to form a 384-channel feature-related volume. This feature-related volume maintains a 128×72 grid structure in the spatial dimension, and the 384-dimensional feature vector at each grid location simultaneously encodes the geometric shape features and exterior wall texture features of the building at that location.

[0102] Specifically, the confidence map is adjusted to the same size as the depth feature map using bilinear interpolation, and the confidence values ​​are normalized to the range of zero to one as weight coefficients. The normalized confidence map and the depth feature map are multiplied element-wise to obtain weighted depth features. At the same time, one is subtracted from the confidence weight to obtain complementary weights, which are multiplied element-wise with the RGB feature map to obtain weighted RGB features. The weighted depth features and weighted RGB features are concatenated in the channel dimension to form a feature correlation volume. The feature correlation volume maintains a two-dimensional grid structure in the spatial dimension and contains both weighted depth features and weighted RGB features in the channel dimension, forming a three-dimensional feature tensor that encodes the correlation between depth information and RGB information.

[0103] The feature-related volume is processed by the depth regressor to generate the high-resolution depth map.

[0104] Specifically, the feature-related volume is input into an upsampling network consisting of three deconvolutional layers. Each deconvolutional layer uses a 4×4 convolutional kernel and employs an upsampling operation with a stride of 2, combined with a ReLU activation function to gradually restore the feature map resolution. After each upsampling, the encoder features at the corresponding scale are concatenated with the upsampling features along the channel dimension through skip connections to preserve detail information. Then, a 1×1 convolutional layer is used to reduce the number of channels to 1, obtaining an initial depth estimation map. The estimated value is then mapped to the 0-1 range using a Sigmoid activation function. Finally, the normalized depth value is multiplied by a preset depth range to convert it into an actual depth value, generating a high-resolution depth map with the same resolution as the input RGB image. The output size of the depth regressor is consistent with the input RGB image, and each pixel stores the converted actual depth value.

[0105] In the above optional approach, a multi-scale feature encoder is further constructed to extract depth and RGB features, and a confidence map weighted fusion strategy is adopted to generate a high-resolution depth map through a depth regressor, making full use of RGB texture details to enhance the spatial resolution and edge sharpness of the depth map.

[0106] In this embodiment, Figure 2 The framework describes the overall process of depth estimation: acquiring RGB image pairs and calculating optical flow; performing online calibration based on 2D and 3D keypoints to generate corrected image pairs; constructing feature matching relationships through stereo matching; optimizing feature matching relationships by combining depth information acquired by a ToF sensor with a confidence map; and finally generating a high-resolution depth map through depth estimation algorithms and planar unfolding processing. Figure 2 It fully demonstrates the processing chain from raw ToF image input to depth map output.

[0107] Figure 3 This diagram illustrates the structure of an embodiment of a depth estimation system 200 based on ToF and RGB fusion provided by the present invention. Figure 3 As shown, the depth estimation system 200 based on ToF and RGB fusion includes: The acquisition module 201 is used to acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image; Processing module 202 is used to estimate the pose change information of the RGB camera caused by OIS technology based on the motion information of multiple feature points in consecutive RGB images. The calculation module 203 is used to calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information; Projection module 204 is used to project the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain a projected depth map; Alignment module 205 is used to spatially align the projection depth map with the RGB image to obtain an aligned projection depth map; The generation module 206 is used to fuse the aligned projection depth map and the RGB image to generate a high-resolution depth map.

[0108] In one alternative embodiment, the acquisition module 201 is specifically used for: The ToF sensor acquires raw depth data of the target scene, and simultaneously acquires the RGB image captured by the RGB camera in the target scene. The original depth data is demodulated based on the dual-frequency modulation principle to obtain the initial depth map and the corresponding confidence map; wherein, the confidence map is used as a weight to guide the fusion of the aligned projected depth map and the RGB image.

[0109] In an alternative embodiment, the processing module 202 is specifically used for: Extract the same multiple feature points from the consecutive multi-frame RGB images, and continuously track the motion trajectory of the multiple feature points in the consecutive multi-frame RGB images; Calculate the corresponding motion information based on the motion trajectory of each feature point to form a set of motion information for multiple feature points; Visual inertial odometry technology is used to separate the device motion of the RGB camera from the lens motion caused by OIS technology from the motion information set; Based on the lens motion of the separated RGB camera, the pose change information of the RGB camera is estimated.

[0110] In an alternative embodiment, the computing module 203 is specifically used for: The initial extrinsic parameter matrix between the ToF sensor and the RGB camera is obtained through offline calibration. Based on the rotation and translation parameters in the pose change information, construct the pose transformation matrix of the RGB camera; The dynamic extrinsic parameters between the ToF sensor and the RGB camera are calculated by multiplying the initial extrinsic parameter matrix with the pose transformation matrix.

[0111] In an alternative embodiment, the projection module 204 is specifically used for: Based on the dynamic extrinsic parameters, a projection transformation relationship is established from the ToF sensor coordinate system to the RGB camera image coordinate system; Each depth pixel in the initial depth map is transformed to the image coordinate system of the RGB camera using the projection transformation relationship to obtain the coordinates of each three-dimensional point in the RGB camera image coordinate system; By applying perspective projection, the coordinates of each three-dimensional point are converted into two-dimensional image coordinates corresponding to the RGB camera; The depth value of each depth pixel in the initial depth map is assigned to the corresponding two-dimensional image coordinates to generate the projected depth map.

[0112] In an alternative embodiment, the alignment module 205 is specifically used for: The dense optical flow field between the projected depth map and the RGB image is calculated using the optical flow method, and the dense optical flow field contains the displacement vector of each pixel; Based on the displacement vector of each pixel in the dense optical flow field, pixel-by-pixel displacement compensation is performed on the projection depth map to generate a displacement-compensated projection depth map. The displacement-compensated projection depth map is subjected to hole filling and edge smoothing processing to obtain the processed projection depth map. The processed projection depth map is aligned with the RGB image at the pixel level to obtain the aligned projection depth map.

[0113] In an alternative embodiment, the generation module 206 is specifically used for: Construct a fusion model that includes a multi-scale feature encoder, a feature fusion unit, and a deep regressor; The aligned projection depth map and the RGB image are input into the multi-scale feature encoder to extract depth features and RGB features, respectively. Using the confidence map as weights, the feature fusion unit performs weighted fusion of the depth features and the RGB features to obtain a feature-related volume that includes the correlation between the depth features and the RGB features; The feature-related volume is processed by the depth regressor to generate the high-resolution depth map.

[0114] It should be noted that the beneficial effects of the depth estimation system 200 based on ToF and RGB fusion provided in the above embodiments are the same as those of the depth estimation method based on ToF and RGB fusion, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0115] The depth estimation system 200 based on ToF and RGB fusion of the present invention can be a computer program (including program code) running on a computer device. For example, the depth estimation system 200 based on ToF and RGB fusion of the present invention is an application software that can be used to execute the corresponding steps in the depth estimation method based on ToF and RGB fusion of the present invention.

[0116] In some embodiments, the depth estimation system 200 based on ToF and RGB fusion of the present invention can be implemented in a combination of hardware and software. As an example, the depth estimation system 200 based on ToF and RGB fusion of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the depth estimation method based on ToF and RGB fusion of the present invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0117] The modules described in the embodiments of this invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.

[0118] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the aforementioned depth estimation methods based on ToF and RGB fusion. That is, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store the computer program; the processor is used to execute the depth estimation method based on ToF and RGB fusion shown in any embodiment of the present invention by calling the computer program.

[0119] In one alternative embodiment, an electronic device is provided, such as Figure 4 As shown, Figure 4 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.

[0120] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0121] Bus 4002 may include a path for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus 4002 is represented by only one thick line, but this does not mean that there is only one bus or one type of bus.

[0122] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0123] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0124] Among them, electronic devices can also be terminal devices. A terminal device can be any terminal device that can install applications and access web pages through applications, including at least one of smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, smart TVs, and smart in-vehicle devices.

[0125] It should be noted that, Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0126] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the aforementioned depth estimation methods based on the fusion of ToF and RGB.

[0127] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0128] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the aforementioned depth estimation method based on ToF and RGB fusion.

[0129] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0131] The computer-readable storage medium provided in this invention can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0132] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.

[0133] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0134] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0135] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this invention can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product contained in one or more computer-readable media, which includes computer-readable program code.

[0136] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A depth estimation method based on ToF and RGB fusion, characterized in that, include: Acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image; Based on the motion information of multiple feature points in consecutive RGB images, the pose change information of the RGB camera caused by OIS technology is estimated. Based on the pose change information, calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera; Using the dynamic extrinsic parameters, the initial depth map is projected onto the image coordinate system of the RGB camera to obtain a projected depth map; The projection depth map is spatially aligned with the RGB image to obtain the aligned projection depth map; The aligned projection depth map and the RGB image are fused to generate a high-resolution depth map.

2. The depth estimation method based on ToF and RGB fusion according to claim 1, characterized in that, The steps of acquiring the initial depth map captured by the ToF sensor and the RGB image captured by the RGB camera include: The ToF sensor acquires raw depth data of the target scene, and simultaneously acquires the RGB image captured by the RGB camera in the target scene. The original depth data is demodulated based on the dual-frequency modulation principle to obtain the initial depth map and the corresponding confidence map; wherein, the confidence map is used as a weight to guide the fusion of the aligned projected depth map and the RGB image.

3. The depth estimation method based on ToF and RGB fusion according to claim 2, characterized in that, The step of estimating the pose change information of the RGB camera due to OIS technology based on the motion information of multiple feature points across multiple consecutive RGB images includes: Extract the same multiple feature points from the consecutive multi-frame RGB images, and continuously track the motion trajectory of the multiple feature points in the consecutive multi-frame RGB images; Calculate the corresponding motion information based on the motion trajectory of each feature point to form a set of motion information for multiple feature points; Visual inertial odometry technology is used to separate the device motion of the RGB camera from the lens motion caused by OIS technology from the motion information set; Based on the lens motion of the separated RGB camera, the pose change information of the RGB camera is estimated.

4. The depth estimation method based on ToF and RGB fusion according to claim 2, characterized in that, The step of calculating the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information includes: The initial extrinsic parameter matrix between the ToF sensor and the RGB camera is obtained through offline calibration. Based on the rotation and translation parameters in the pose change information, construct the pose transformation matrix of the RGB camera; The dynamic extrinsic parameters between the ToF sensor and the RGB camera are calculated by multiplying the initial extrinsic parameter matrix with the pose transformation matrix.

5. The depth estimation method based on ToF and RGB fusion according to claim 4, characterized in that, The step of projecting the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain the projected depth map includes: Based on the dynamic extrinsic parameters, a projection transformation relationship is established from the ToF sensor coordinate system to the RGB camera image coordinate system; Each depth pixel in the initial depth map is transformed to the image coordinate system of the RGB camera using the projection transformation relationship to obtain the coordinates of each three-dimensional point in the RGB camera image coordinate system; By applying perspective projection, the coordinates of each three-dimensional point are converted into two-dimensional image coordinates corresponding to the RGB camera; The depth value of each depth pixel in the initial depth map is assigned to the corresponding two-dimensional image coordinates to generate the projected depth map.

6. The depth estimation method based on ToF and RGB fusion according to any one of claims 2 to 5, characterized in that, The step of spatially aligning the projected depth map with the RGB image to obtain an aligned projected depth map includes: The dense optical flow field between the projected depth map and the RGB image is calculated using the optical flow method, and the dense optical flow field contains the displacement vector of each pixel; Based on the displacement vector of each pixel in the dense optical flow field, pixel-by-pixel displacement compensation is performed on the projection depth map to generate a displacement-compensated projection depth map. The displacement-compensated projection depth map is subjected to hole filling and edge smoothing processing to obtain the processed projection depth map. The processed projection depth map is aligned with the RGB image at the pixel level to obtain the aligned projection depth map.

7. The depth estimation method based on ToF and RGB fusion according to claim 6, characterized in that, The step of fusing the aligned projection depth map and the RGB image to generate a high-resolution depth map includes: Construct a fusion model that includes a multi-scale feature encoder, a feature fusion unit, and a deep regressor; The aligned projection depth map and the RGB image are input into the multi-scale feature encoder to extract depth features and RGB features, respectively. Using the confidence map as weights, the feature fusion unit performs weighted fusion of the depth features and the RGB features to obtain a feature-related volume that includes the correlation between the depth features and the RGB features; The feature-related volume is processed by the depth regressor to generate the high-resolution depth map.

8. A depth estimation system based on ToF and RGB fusion, characterized in that, include: The acquisition module is used to acquire an initial depth map acquired by a ToF sensor and an RGB image acquired by an RGB camera; wherein the resolution of the initial depth map is lower than the resolution of the RGB image; The processing module is used to estimate the pose change information of the RGB camera caused by OIS technology based on the motion information of multiple feature points in consecutive RGB images. The calculation module is used to calculate the dynamic extrinsic parameters between the ToF sensor and the RGB camera based on the pose change information; The projection module is used to project the initial depth map onto the image coordinate system of the RGB camera using the dynamic extrinsic parameters to obtain a projected depth map; An alignment module is used to spatially align the projection depth map with the RGB image to obtain an aligned projection depth map; A generation module is used to fuse the aligned projection depth map and the RGB image to generate a high-resolution depth map.

9. An electronic device, characterized in that, The electronic device includes a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the electronic device to implement the depth estimation method based on ToF and RGB fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which, when executed by a processor, implements the depth estimation method based on ToF and RGB fusion as described in any one of claims 1 to 7.