Scene depth reconstruction method and device, medium and product
By combining low-resolution TOF sensors and high-resolution RGB images, using dual-branch cross attention fusion neural networks for feature fusion, the power consumption and cost problems of deep reconstruction of high-resolution scenes in the prior art are solved, and the deep reconstruction effect with low power consumption and high precision is achieved.
Patent Information
- Application Number
- CN202510275571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve high-resolution scene depth reconstruction, especially in terms of power consumption and cost. At the same time, the depth reconstruction algorithm based solely on RGB images has high complexity and poor real-time performance.
By combining low-resolution TOF sensors with high-resolution RGB images, a two-branch cross-attention fusion neural network is built using deep learning methods to perform feature fusion to generate high-resolution depth maps.
It realizes low-power and high-precision high-resolution scene depth reconstruction, overcomes the problems of high power consumption and high cost of TOF sensors, and improves the accuracy and real-time performance of deep reconstruction.
Smart Images

Figure CN120219458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to a method, device, medium, and product for scene depth reconstruction. Background Art
[0002] In existing scene depth reconstruction, the mainstream methods include two methods based on cameras and TOF sensors. The camera-based method usually uses a high-resolution camera to perform stereo matching and dense matching through stereo pairs or sequential images to complete depth reconstruction. The TOF sensor usually reconstructs by emitting and receiving encoded light.
[0003] For the camera-based reconstruction solution, the power consumption is low and the resolution is high, but the algorithm is complex, and generally uses classical stereo vision-based computing or end-to-end computing of deep learning. Due to the high algorithm complexity, to ensure real-time reconstruction, a processor with a large power consumption must be equipped. Typically, a high-performance processor of the Arm Cortex-A series equipped with an NPU.
[0004] For the TOF-based solution, what you see is what you get, and no complex post-processing calculations are required. However, the design of a large field of view, high-resolution TOF sensors and optical algorithms is very difficult. The power consumption, cost, and algorithm difficulty of TOF sensors will increase significantly as the field of view and resolution increase.
[0005] Currently, multi-point TOF sensors with very low resolution but a certain field of view have begun to appear. Typically, the 8x8 TOF sensor of STMicroelectronics. It can provide depth ranging values for 8x8 points with a 60° FOV. This TOF is neither the high-resolution TOF with tens of thousands of points nor the single-point TOF with extremely low power consumption but only a single-digit number of measurement values. It maintains extremely low power consumption but can provide more information about the scene depth. How to use this low-resolution TOF to fuse with RGB images for high-resolution scene reconstruction has extremely high application value. Summary of the Invention
[0006] An object of this application is to provide a method, device, medium, and product for scene depth reconstruction, at least to solve the technical problem of difficult high-resolution scene depth reconstruction.
[0007] To achieve the above object, some embodiments of this application provide the following aspects:
[0008] In a first aspect, some embodiments of the present application further provide a method for reconstructing scene depth, including collecting high-resolution depth maps of different scenes; in the high-resolution depth maps, intercepting regions corresponding to the field of view of low-resolution TOF and dividing them into multiple grids; assigning the maximum value of the depth within the grid to the grid to form simulated low-resolution TOF data; training a dual-branch cross-attention fusion neural network according to the high-resolution depth maps and the low-resolution TOF data; the dual-branch cross-attention fusion neural network includes an image feature extraction branch, a TOF branch network, a cross-attention fusion module, and an encoder-decoder; when performing scene depth reconstruction, obtaining a high-resolution RGB image and low-resolution TOF data as the input of the trained dual-branch cross-attention fusion neural network; and performing feature fusion on the high-resolution RGB image and the low-resolution TOF data through the dual-branch cross-attention fusion neural network to obtain a high-resolution depth map.
[0009] In a second aspect, some embodiments of the present application further provide an electronic device, which includes: one or more processors; and a memory storing computer program instructions, and when the computer program instructions are executed, the processors execute the steps of the method as described above.
[0010] In a third aspect, some embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method as described above.
[0011] In a fourth aspect, some embodiments of the present application further provide a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described above are implemented.
[0012] Compared with the related art, in the solution provided by the embodiments of the present application, by combining a low-resolution TOF sensor with a high-resolution RGB image and using a deep learning method to construct a dual-branch cross-attention fusion neural network, the problems of high power consumption and high cost of TOF sensors caused by high resolution in traditional methods, as well as the defects of high algorithm complexity and poor real-time performance when reconstructing depth solely based on RGB images, are overcome. It can not only reconstruct high-resolution depth maps with low power consumption, but also significantly improve the accuracy and effect of depth reconstruction through cross-attention mechanism and optimization of specific loss functions, and has extremely high application value. Description of the Drawings
[0013] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.
[0014] Figure 1 It is a flowchart of a scene depth reconstruction method provided according to an embodiment of the present application;
[0015] Figure 2 It is a schematic diagram of a high-resolution Depth and downsampled simulated low-resolution TOF provided according to an embodiment of the present application;
[0016] Figure 3 It is an architecture diagram of a dual-branch cross-attention fusion neural network provided according to an embodiment of the present application;
[0017] Figure 4 It is a flowchart of neural network training and application provided according to an embodiment of the present application;
[0018] Figure 5 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. Specific embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0020] The first embodiment
[0021] The first embodiment of the present application relates to a scene depth reconstruction method. As Figure 1 shown, the method may include the following steps:
[0022] S101, collect high-resolution depth maps of different scenes; in the high-resolution depth maps, intercept regions corresponding to the field of view of the low-resolution TOF and divide them into multiple grids; assign the maximum value of the depth within the grid to the grid to form simulated low-resolution TOF data.
[0023] As Figure 2As shown, on the left is the simulated low-resolution TOF data, and on the right is the Depth in the originally acquired high-resolution RGBD data. By acquiring high-resolution depth maps of different scenes and the corresponding high-resolution RGB images, the diversity of training data is ensured, providing rich feature information for the model. In the high-resolution depth map, an area corresponding to the size of the low-resolution TOF field of view is intercepted and divided into multiple grids. The process of assigning the maximum depth value within the grid to the grid simulates the working mode of an actual low-resolution TOF sensor, generating low-resolution TOF data that matches the actual scene. This method of generating low-resolution TOF data based on real high-resolution depth maps not only saves the high-cost hardware data acquisition requirements but also ensures a high degree of consistency between the simulated data and the real TOF data, improving the effectiveness and adaptability of model training.
[0024] S102. Train the dual-branch cross-attention fusion neural network according to the high-resolution depth map and the low-resolution TOF data; the dual-branch cross-attention fusion neural network includes an image feature extraction branch, a TOF branch network, a cross-attention fusion module, and an encoder-decoder.
[0025] By training the dual-branch cross-attention fusion neural network, the complementary characteristics of the high-resolution RGB image and the low-resolution TOF data are fully utilized. The high-resolution RGB image provides rich spatial texture and boundary information, while the low-resolution TOF data directly provides scene depth information.
[0026] In the network, the cross-attention mechanism uses the TOF feature as the Query, and the image feature as the Key and Value, which can effectively capture the complementary relationship between the TOF feature and the RGB image feature, thereby realizing the in-depth fusion of information and improving the depth accuracy of the reconstruction. The defined multiple loss functions not only improve the prediction accuracy of the model for depth values but also enhance the detail representation ability of the reconstruction result, making the gradient distribution of the predicted depth map smoother and more natural.
[0027] S103. When performing scene depth reconstruction, use the acquired high-resolution RGB image and low-resolution TOF data as the input to the trained dual-branch cross-attention fusion neural network.
[0028] In practical applications, inputting the high-resolution RGB image and the low-resolution TOF data into the trained neural network simplifies the reconstruction process. The neural network has mastered the correlation between the RGB and TOF features through training and can complete the real-time prediction of the high-resolution depth map without additional complex post-processing algorithms. By combining deep learning, the real-time computing requirements for hardware resources (such as processor performance and power consumption) are significantly reduced, making it suitable for power-sensitive embedded devices or mobile terminals.
[0029] S104. Feature fusion of the high-resolution RGB image and the low-resolution TOF data is performed through the dual-branch cross-attention fusion neural network to obtain a high-resolution depth map.
[0030] Through the trained dual-branch cross-attention fusion neural network, feature fusion of the high-resolution RGB image and the low-resolution TOF data is performed to successfully generate a high-resolution depth map. In this process, the RGB image compensates for the shortcomings of the low resolution and insufficient details of the TOF data, while the TOF data provides depth information for the RGB image, making up for the limitations of pure image depth estimation. The finally generated high-resolution depth map has high accuracy and rich details and can be used for depth perception applications in scenarios such as robot vision and augmented reality.
[0031] It is not difficult to find that compared with the related technologies, the solution provided by the embodiment of the present application realizes high-resolution scene depth reconstruction with low power consumption, low cost, and high accuracy. By combining the high-resolution RGB image and the low-resolution TOF data, the efficiency of the TOF sensor is retained, and the limitation of its low resolution is overcome. The finally generated high-resolution depth map can meet the dual requirements of real-time performance and accuracy, providing an efficient solution for scene perception technology.
[0032] Second Embodiment
[0033] The second embodiment of the present application relates to a method for scene depth reconstruction. The second implementation is an improvement based on the first embodiment. The specific improvement lies in:
[0034] A low-power real-time scene depth reconstruction method based on the fusion of 8x8 low-resolution TOF (Low-Definition TOF, D-TOF) and high-resolution RGB image (HD-Image) is proposed. As Figure 3 shown, a deep neural network is adopted, with LD-TOF and HD-Image as inputs, and a dual-branch + cross-attention method is used to achieve feature-level fusion. The output is the high-resolution depth (HD-Depth) reconstruction result of the scene. Since the whole method is end-to-end, the training process of the neural network is mainly described, and the inference process of the neural network is briefly described:
[0035] Furthermore, the dual-branch cross-attention fusion neural network includes: defining the training loss including: the error between the depth prediction value and the high-resolution ground truth, the error between the gradient value of the predicted depth and the gradient of the high-resolution ground truth, and the error between the value of each grid of the simulated low-resolution TOF and the maximum value within the corresponding grid of the predicted high-resolution depth value.
[0036] The error between the predicted depth value and the high-resolution ground truth. This loss calculates the per-pixel error between the predicted depth value and the high-resolution ground truth, usually using L1 loss or L2 loss. The goal is to ensure that the predicted depth value is overall close to the true depth value, guaranteeing the accuracy of the generated high-resolution depth map globally and enabling the depth prediction result to accurately reflect the depth distribution of the actual scene.
[0037] The error between the gradient value of the predicted depth and the gradient of the high-resolution ground truth. By calculating the difference between the gradient value of the predicted depth map and the gradient value of the high-resolution depth map ground truth, the prediction ability of the model for depth gradients is constrained. This gradient error usually calculates the image gradient through the Sobel operator or a similar operator to measure the difference in edges and local depth changes. It enhances the model's ability to capture depth edge details, making the generated depth map more accurate in the boundary region, avoiding overly blurred depth value transitions, and improving the local continuity of the predicted depth map and the authenticity of the scene geometry structure.
[0038] The error between the value of each grid of the simulated low-resolution TOF and the maximum value within the corresponding grid of the predicted high-resolution depth value. The depth value of each grid of the simulated low-resolution TOF data is compared with the maximum depth value within the corresponding grid of the predicted high-resolution depth map to calculate the error. This loss term ensures that the predicted high-resolution depth map is compatible with the actual physical characteristics of the low-resolution TOF data. By constraining the consistency between the model's prediction result and the low-resolution TOF data, the generalization ability of the model is enhanced, enabling it to adapt to the application scenario of the actual low-resolution TOF sensor, improving the effective utilization rate of the depth information in the TOF data by the network, and avoiding the depth reconstruction result deviating from the information provided by the TOF sensor.
[0039] Further, intercepting the region corresponding to the field of view of the LD-TOF and dividing it into multiple grids includes: in the high-resolution depth map, intercepting the region corresponding to the field of view of the LD-TOF and dividing it into 8×8 grids.
[0040] The field of view of the LD-TOF sensor is usually small and fixed. Therefore, according to its FOV range, a rectangular region of the corresponding size needs to be cropped from the high-resolution depth map. Taking the typical 8×8 TOF sensor of STMicroelectronics as an example, it can provide depth ranging values of 8x8 points with a 60° FOV. When training the dual-branch cross-attention fusion neural network, it is necessary to increase the amount of training data to ensure the consistency between the simulated low-resolution TOF data and the depth data collected by the actual TOF sensor.
[0041] The cropped area is evenly divided into 8×8 grids, and each grid represents a sampling point in the LD-TOF sensor. The grid division method is equally spaced and covers the entire cropped area, enabling the depth information within each grid to be represented by a single value. For each grid, the depth data within it is extracted, and the maximum value of the depth within the grid is calculated and used as the simulated LD-TOF depth value. This maximum value assignment rule reflects the farthest ranging information that the LD-TOF sensor may capture at each sampling point and simulates its working characteristics.
[0042] Furthermore, the image feature extraction branch includes: passing the high-resolution depth map through multiple convolutional and pooling layers to extract image features with a dimension of 512×H / 32×W / 32; the number of feature channels is 576, H is the height of the image, and W is the width of the image.
[0043] The image feature extraction branch can be constructed using the standard MobileNetV3-Small as the feature encoder. MobileNetV3-Small is a lightweight deep neural network structure that can be optimized to run on low-power devices while being able to extract image features with high expressiveness. Compared with other large networks (such as ResNet or VGG), MobileNetV3-Small has lower computational complexity and the number of parameters, making it very suitable for power-sensitive real-time application scenarios.
[0044] Through multiple convolutional and pooling layers, the image is extracted into tensor features with a dimension of 512x H / 32x W / 32, where 576 is the number of feature channels, H is the height of the image, and W is the width of the image.
[0045] Furthermore, the TOF branch network includes: performing a 1D unfolding on the low-resolution TOF data, then using 8 1D convolutions to expand the original low-resolution TOF data into a vector of 512×1, and then replicating and expanding the 512×1 vector into TOF features of 512×H / 32×W / 32.
[0046] At this time, the feature dimension of TOF is the same as that of the image. Due to the lower resolution of TOF, the fusion with the image is only performed at the feature layer after 32-fold downsampling.
[0047] Furthermore, the cross-attention fusion module includes: using the TOF features as Query, calculating Key and Value from the image features, calculating the cross-attention between the TOF features and the image features, and forming fusion features of 512×H / 32×W / 32 based on the cross-attention.
[0048] Further, the codec includes: a codec adopting a UNet architecture to encode and decode the fusion features; upsampling the features into features of H / 4×W / 4 in three Blocks to predict the depth at the original resolution; and the decoder in it upsamples the features.
[0049] In the application of the neural network, only by inputting a high-resolution image and corresponding low-resolution TOF data, a high-resolution depth map of the scene can be generated.
[0050] It is not difficult to find that in the embodiments of the present application, the high-resolution spatial texture features of the RGB image and the depth information of the LD-TOF are respectively extracted through the dual-branch structure, and the cross-attention mechanism is adopted for feature-level fusion, so that the depth information in the low-resolution TOF data is fully utilized, and at the same time, the shortcoming of its insufficient resolution is made up. Through the optimized network structure and training strategy, the advantages of the low-resolution TOF data and the high-resolution RGB image are fused, and at the same time, the problems of high power consumption, expensive equipment and poor real-time performance in the prior art are overcome, providing a low-cost and high-performance solution for the popularization of the depth perception technology.
[0051] The third embodiment
[0052] The third embodiment of the present application relates to a method for reconstructing the depth of a scene. The third implementation is an improvement based on the first embodiment. The specific improvement lies in:
[0053] An end-to-end depth reconstruction method based on a high-resolution RGB image and low-resolution TOF data is proposed for scene depth reconstruction. The whole process is as Figure 4 shown and mainly includes the following steps:
[0054] Step 401. Data acquisition and simulation: High-resolution depth maps are collected from different scenes, and the corresponding depth regions are intercepted according to the field of view size of the low-resolution TOF (LD-TOF) sensor, and the region is divided into multiple equally spaced grids (such as an 8×8 grid). The maximum value of the depth is extracted in each grid to generate simulated low-resolution TOF data, so that it can reflect the characteristics of the actual sensor.
[0055] Step 402. Construction of the dual-branch feature extraction module: Construct a dual-branch neural network, including an image feature extraction branch and a TOF feature extraction branch.
[0056] The image feature extraction branch takes a high-resolution RGB image as the input, adopts a MobileNetV3-Small lightweight network as the feature encoder, extracts features through multiple convolutional and pooling layers, and converts the RGB image into image features with a dimension of 512×H / 32×W / 32.
[0057] The TOF feature extraction branch takes the simulated low-resolution TOF data as input. Through one-dimensional unfolding and 8 one-dimensional convolutional operations, the TOF features are expanded into a 512×1 vector, and then replicated and expanded into a 512×H / 32×W / 32 TOF.
[0058] Step 403. Construction of the feature fusion module: Based on the extraction of image features and TOF features, a cross-attention fusion module is constructed. Using the TOF features as Query, calculating Key and Value from the image features, and performing feature fusion using the cross-attention mechanism. The fused features with a dimension of 512×H / 32×W / 32 are generated through cross-attention, which retain both the spatial detail information of the RGB image and the depth information of the TOF.
[0059] Step 404. Construction of the encoder-decoder module: Design an encoder-decoder module based on the UNet architecture to further process the fused features. The encoder performs multi-layer downsampling on the fused features to extract deep-level information. The decoder upsamples layer by layer through three Blocks to the features of H / 4×W / 4, and finally generates the depth map of the original resolution (H×W) to achieve high-resolution depth reconstruction of the scene.
[0060] Step 405. Training of the depth reconstruction neural network: Adopt an end-to-end training method to optimize the depth reconstruction network. Define the loss function, including the error between the depth prediction value and the high-resolution ground truth (such as L1 / L2 loss), the error between the predicted depth gradient and the ground truth gradient (calculated through the Sobel operator), and the error between the low-resolution TOF grid value and the maximum value of the predicted high-resolution depth value grid. The design of the loss function ensures the global accuracy, local details of the predicted depth, and the consistency with the LD-TOF sensor data.
[0061] Step 406. Input high-resolution RGB image: Collect the high-resolution RGB image of the scene as the input of the image feature extraction branch.
[0062] Step 407. Input low-resolution TOF data: The real-time low-resolution TOF data collected by the LD-TOF sensor is used as the input of the TOF feature extraction branch.
[0063] Scene depth reconstruction: Apply the trained neural network to the depth reconstruction of the actual scene to output a high-resolution depth map. After feature extraction, fusion, and depth prediction, the final high-resolution depth map is output. The generated depth map not only retains the detail information of the RGB image but also accurately reconstructs the depth distribution of the scene, and can be widely applied to fields such as real-time scene perception and 3D modeling.
[0064] It should be noted that the third embodiment of this application can also be an improvement based on the second embodiment.
[0065] It is not difficult to find that in the embodiments of the present application, through the design of a lightweight network structure and the introduction of a cross-attention fusion mechanism, the information of low-resolution TOF data and high-resolution RGB images is effectively fused, which not only ensures a high-precision depth reconstruction effect but also meets the requirements of real-time performance and low power consumption. At the same time, the optimization of multiple loss functions during the training process further enhances the generalization ability of the model, making it applicable to a variety of actual application scenarios.
[0066] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of the present application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of the algorithm and process, is within the protection scope of this application.
[0067] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.
[0068] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 5 An exemplary structural diagram of the electronic device is disclosed. As Figure 5 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component is interconnected using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (such as, as a server array, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are only examples and are not intended to limit the implementation of the present application described herein and / or claimed.
[0069] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 may be connected via a bus or other means. Figure 5 Here, taking the connection via the bus as an example.
[0070] The input device 1103 can receive input digital or character information, and generate key signal inputs related to user settings and function controls of the electronic device, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 1104 may include a display device, an auxiliary lighting device (e.g., LED), and a haptic feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touchscreen.
[0071] To provide interaction with the user, the electronic device may be a computer. The computer has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or an LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and the input from the user may be received in any form (including voice input, speech input, or haptic input).
[0072] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by the processor, the steps of the method provided in any one or more of the above embodiments are implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist separately and not be assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.
[0073] The memory 1102 may be a non-transitory computer-readable storage medium, which can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 1101 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.
[0074] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device and the like. In addition, the memory 1102 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely disposed relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0075] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device.
[0076] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0077] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0078] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, application-specific integrated circuits (ASICs), general-purpose computers, or any other similar hardware devices can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, for example, RAM memory, magnetic or optical drives, or floppy disks and similar devices. Additionally, some steps or functions of this application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0079] The computer program product provided by the embodiments of this application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0080] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0081] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the term "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple elements or devices recited in the apparatus claims may also be implemented by one element or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.
[0082] As described above, these are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily conceive of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.
Claims
1. A scene depth reconstruction method, characterized in that: The method comprises: Collect high-resolution depth maps of different scenes; in the high-resolution depth map, intercept an area of a size corresponding to the field of view of the low-resolution TOF and divide it into multiple grids; assign the maximum value of the depth in the grid to the grid to form simulated low-resolution TOF data; According to the high-resolution depth map and the low-resolution TOF data, a dual-branch cross-attention fusion neural network is trained; the dual-branch cross-attention fusion neural network includes an image feature extraction branch, a TOF branch network, a cross-attention fusion module and a codec; When reconstructing the scene depth, a high-resolution RGB image and a low-resolution TOF data are obtained as the input of the trained dual-branch cross-attention fusion neural network; the high-resolution RGB image and the low-resolution TOF data are feature fused by the dual-branch cross-attention fusion neural network to obtain a high-resolution depth map.
2. The method according to claim 1, characterized in that: The dual-branch cross-attention fusion neural network includes: defining training losses including: the error between the depth prediction value and the high-resolution true value, the error between the gradient value of the predicted depth and the gradient of the high-resolution true value, and the error between the value of each grid of the simulated low-resolution TOF and the maximum value within the corresponding grid using the predicted high-resolution depth value.
3. The method according to claim 2, characterized in that The method of intercepting an area of a size corresponding to the field of view of the LD-TOF and dividing it into a plurality of grids includes: In the high-resolution depth map, an area corresponding to the size of the field of view of the LD-TOF is intercepted and divided into 8×8 grids.
4. The method according to claim 3, characterized in that: The image feature extraction branch includes: The high-resolution depth map is subjected to multiple convolution and pooling layers to extract image features with a dimension of 512×H / 32×W / 32; the number of feature channels is 576, H is the height of the image, and W is the width of the image.
5. The method according to claim 4, characterized in that The TOF branch network includes: The low-resolution TOF data is expanded in one dimension, and then eight one-dimensional convolutions are used to expand the original low-resolution TOF data into a 512×1 vector, and then the 512×1 vector is copied and expanded into a 512×H / 32×W / 32 TOF feature.
6. The method according to claim 5, characterized in that The cross attention fusion module includes: The TOF feature is used as a query, the Key and the Value are calculated using the image feature, the cross attention of the TOF feature and the image feature is calculated, and a fusion feature of 512×H / 32×W / 32 is formed according to the cross attention.
7. The method according to claim 6, characterized in that The codec includes: A codec based on the UNet architecture is used to encode and decode the fused features; the features are upsampled to H / 4×W / 4 in three blocks to predict the depth of the original resolution; and the decoder upsamples the features.
8. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 7.
9. A computer readable medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.