Stereoscopic matching for depth estimation using image pairs with arbitrary relative pose configurations

By applying relaxed image correction and stereo matching technology in the monocular camera system, the problem of depth information estimation under the monocular camera is solved, and fast and accurate depth estimation is achieved, suitable for images in any relative poses.

CN120153394APending Publication Date: 2025-06-13创峰科技
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101625.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult to estimate depth information directly from images captured by monocular cameras, especially when the relative camera poses of the image are arbitrary, traditional stereo matching-based depth estimation algorithms are not applicable, and pixel triangulation algorithms will introduce additional computational burden and noise.

Method used

By applying a relaxed image correction technique, the pole lines of the current image and the reference image are parallel and stereo matching is performed along the pole lines to estimate the disparity map and the depth map. Optionally, deep neural networks are used to remove noise.

Benefits of technology

It realizes rapid and accurate estimation of depth information under a monocular camera, and is suitable for images of any relative poses, reducing the demand for computing resources and improving the accuracy of depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153394A_ABST
    Figure CN120153394A_ABST
Patent Text Reader

Abstract

An electronic device obtains a current image and a reference image. A slack image correction operation is applied to the reference image with reference to the current image to generate a corrected reference image such that a current epipolar line of the current image is parallel to a reference epipolar line of the corrected reference image. The electronic device generates a first disparity map corresponding to the current image and a corrected reference image, the disparity map representing a first disparity between a current distance of each current pixel from a current pole in a subset of the current image and a reference distance of a corresponding reference pixel from a reference pole of the corrected reference image. A depth map of the current image is determined based on the first disparity map, and the depth map includes a depth value for each current pixel of a subset of the current image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to data processing technologies, including but not limited to methods, systems, and non-transitory computer-readable media for determining depth information from image data based on stereo matching. Background Art

[0002] For augmented reality (AR) applications that need to create occlusion and collision effects for virtual objects and generate dense mesh maps in a physical environment, depth image estimation from a monocular moving camera is necessary. Most existing depth estimation algorithms are developed based on stereo matching of stereo cameras with two camera lenses having fixed relative poses and cannot be directly applied to different images with arbitrary relative poses captured by a monocular camera. Pixel-based triangulation algorithms are usually applied together with polar rectification to generate depth information. However, this algorithm is not applicable to images with arbitrary relative camera poses and may introduce additional computational burdens and noise into depth estimation. Some other solutions use sparse feature points for dense depth estimation for real-time rendering. Although the computational speed is fast, especially when the number of sparse points is not large enough, this dense depth estimation generates depth images with limited resolution. Having a depth estimation mechanism for determining depth information from image data in a fast, accurate, and efficient manner would be beneficial. Summary of the Invention

[0003] Embodiments of this application aim to determine a depth image or depth map of a current image (e.g., in an extended reality application) based on a reference image having a camera pose different from that of the current image. The current image and the reference image have arbitrary and different relative poses. Relaxed image rectification is applied to one or both of the current image and the reference image to make their epipolar lines parallel to each other and / or move their epipoles to the same position in the same image plane of the two images. Stereo matching is used along the epipolar lines to estimate a disparity map between the current image and the reference image and a depth map of the reference image. Additionally, in some embodiments, a deep neural network is applied to the depth map to remove noise with reference to the current image. In some cases, an augmented reality (AR) software development kit (SDK) uses an image sequence to estimate three-dimensional (3D) geometric or structural information of the environment, and the image sequence is collected by a monocular camera of a mobile device based on the camera poses determined by its simultaneous localization and mapping (SLAM) module.

[0004] In one aspect, a method is implemented on an electronic device having one or more processors and a memory. The method includes obtaining a current image associated with a current camera pose and obtaining a reference image associated with a reference camera pose different from the current camera pose. The method further includes applying a relaxation image rectification operation to the reference image with reference to the current image to generate a rectified reference image. A current epipolar line of the current image is parallel to a reference epipolar line of the rectified reference image. The method also includes generating a first disparity map corresponding to the current image and the rectified reference image. The first disparity map represents a first disparity between (1) a current distance of each current pixel from a current pole in a subset of the current image and (2) a reference distance of a corresponding reference pixel from a reference pole in the rectified reference image. The method also includes determining a depth map of the current image based on the first disparity map, the depth map including a depth value for each current pixel in the subset of the current image. In some embodiments, for each current pixel in the subset of the current image, the depth value is represented as: where D p represents the first disparity corresponding to the respective current pixel in the first disparity map, and A and B are two pixel transformation parameters.

[0005] In some embodiments, the method further includes generating a second disparity map corresponding to the current image and the rectified reference image, the second disparity map representing a second disparity between a reference distance of each reference pixel from a reference pole in a subset of the rectified reference image and a current distance of the corresponding current pixel from a current pole in the current image. The depth map of the current image is determined based on the first disparity map and the second disparity map.

[0006] In some embodiments, determining the first disparity map further includes: for each current pixel in the subset of the current image, determining an epipolar line of a reference pixel on the rectified reference image, selecting an epipolar segment on the epipolar line based on a disparity range, creating a cost volume of a plurality of candidate pixels located on the epipolar segment, and determining a first disparity between the current distance and the reference distance based on the cost volume. The cost volume has a plurality of cost elements and each cost element indicates a Hamming distance between a first feature value of the current pixel in the current image and a second feature value of the corresponding candidate pixel on the epipolar segment in the rectified reference image.

[0007] In another aspect, some embodiments include an electronic device. The electronic device includes one or more processors and a memory storing instructions thereon, which when executed by the one or more processors, cause the processors to perform any of the above methods.

[0008] In another aspect, some embodiments include a non-transitory computer-readable medium storing instructions thereon, which when executed by one or more processors, cause the processors to perform any of the above methods.

[0009] The mention of these illustrative embodiments and examples is not intended to limit or define the present disclosure, but rather to provide examples to aid in understanding the present disclosure. Other embodiments are discussed in the detailed description and further description is provided in the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] To better understand the various embodiments described, reference should be made to the following detailed description in conjunction with the following drawings, in which like reference numerals refer to corresponding parts in all the drawings.

[0011] Figure 1 is an example data processing environment having one or more servers communicatively coupled to one or more client devices according to some embodiments.

[0012] Figure 2 is a block diagram showing an electronic system according to some embodiments.

[0013] Figure 3 is an example data processing environment for training and applying a neural network-based data processing model for processing visual and / or speech data according to some embodiments.

[0014] Figure 4A is an example neural network applied to process content data in a neural network (NN)-based data processing model according to some embodiments, Figure 4B is an example node in a neural network according to some embodiments.

[0015] FIG. 5 is an example stereoscopic vision environment for forming a three-dimensional (3D) image including a current image and a corresponding depth map according to some embodiments.

[0016] Figure 5B is an image pair including an example current image and an example rectified reference image according to some embodiments.

[0017] Figure 6 is a flowchart of an example process for generating a depth map corresponding to a current image according to some embodiments.

[0018] Figure 7 is an example list of formulas applied to convert image data into depth data according to some embodiments.

[0019] Figure 8 is a flowchart of an example process for converting image data into depth information according to some embodiments.

[0020] Figure 9A flowchart of an example process for controlling noise in a disparity map for determining a depth map of a current image according to some embodiments.

[0021] Figure 10 A block diagram of an example data processing model for controlling noise in depth information according to some embodiments.

[0022] Figure 11 Shows a feature extraction scheme for extracting image features from image pixels according to some embodiments.

[0023] Figure 12 A flowchart of a noise filtering process for reducing noise in a depth map according to some embodiments.

[0024] Figure 13 A flowchart of an example depth mapping method according to some embodiments.

[0025] In some views of the drawings, the same reference numerals refer to corresponding parts. Detailed Description

[0026] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to those of ordinary skill in the art that various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, it will be apparent to those of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic systems having digital video capabilities.

[0027] Various embodiments of the present application aim to determine depth images or depth maps corresponding to two images captured at two different camera poses (e.g., in an extended reality application). For example, some embodiments are applied in an augmented reality (AR) software development kit (SDK) to estimate geometric or structural information of the environment using an image sequence collected by a monocular camera of a mobile device based on the camera poses determined by its simultaneous localization and mapping (SLAM) module. Specifically, a stereo matching algorithm is used to estimate a depth map from an image pair, where the images in each image pair have arbitrary and different relative poses. A homograph transformation is applied to one or both images in each image pair to make the epipolar lines of the two images parallel to each other and / or move the epipoles of the two images to the same position in the same image plane of the two images. The tilt angle of each epipolar line is pre-determined and stored in a look-up table. Optionally, the epipolar lines of the two images are arranged to be parallel to or intersect with the horizontal axis.

[0028] Furthermore, a stereo matching algorithm is performed along the epipolar lines to estimate a disparity map associated with the two images of each image pair. In some embodiments, two disparity maps are determined for each image pair for disparity consistency checking to remove noise in the disparity map. Additionally, in some embodiments, a deep neural network is applied to the depth map to remove noise with reference to the original images in each image pair. By these methods, depth estimation does not rely entirely on epipolar rectification and requires limited computational resources that can be conveniently provided by a mobile device.

[0029] Figure 1An example data processing environment 100 according to some embodiments has one or more servers 102 communicatively coupled to one or more client devices 104. The one or more client devices 104 can be, for example, a desktop computer 104A, a tablet computer 104B, a mobile phone 104C, or a smart multi-sensing networked home device (e.g., a depth camera, a visible light camera). In some embodiments, the one or more client devices 104 include a head-mounted display 104D configured to present extended reality content. Each client device 104 can collect data or user behavior, execute user applications, and present outputs on its user interface. The data or user behavior collected can be processed locally at the client device 104 (e.g., for training and / or for prediction) and / or remotely by the server 102. The one or more servers 102 provide system data (e.g., boot files, operating system images, and user applications) to the client devices 104 and, in some embodiments, process data and user behavior received from the client devices 104 when user applications are executed on the client devices 104. In some embodiments, the data processing environment 100 further includes a storage device 106 for storing data related to the servers 102, the client devices 104, and the applications executed on the client devices 104. For example, the storage device 106 can store video content (including visual and audio content), static visual content, and / or inertial sensor data.

[0030] The one or more servers 102 can enable real-time data communication between client devices 104 that are physically separated from each other or client devices 104 that are physically separated from the one or more servers 102. Additionally, in some embodiments, the one or more servers 102 can implement data processing tasks that cannot be or would not be preferentially performed locally by the client devices 104. For example, the client device 104 includes a game console (e.g., formed by the head-mounted display 104D) that executes an interactive online game application. The game console receives user instructions and sends them, together with user data, to the game server 102. The game server 102 generates a video data stream based on the user instructions and user data and provides the video data stream for display on the game console and other client devices participating in the same game session with the game console.

[0031] One or more servers 102, one or more client devices 104, and a storage device 106 are communicatively coupled to each other via one or more communication networks 108, which is a medium for providing communication links between these devices within the data processing environment 100 and computers connected together. The one or more communication networks 108 may include connections such as wired connections, wireless communication links, or fiber optic cables. Examples of the one or more communication networks 108 include a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination thereof. The one or more communication networks 108 are optionally implemented using any known network protocol, including various wired or wireless protocols such as Ethernet, universal serial bus (USB), FireWire, long term evolution (LTE), global system for mobile communications (GSM), enhanced data GSM environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wi-Fi, voice over Internet protocol (VoIP), Wi-MAX, or any other suitable communication protocol. Connections to the one or more communication networks 108 can be established directly (e.g., using a 3G / 4G connection to a wireless carrier), or through a network interface 110 (e.g., a router, switch, gateway, hub, or smart home control node), or any combination thereof. Thus, the one or more communication networks 108 can represent the Internet, which is a global collection of networks and gateways that communicate with each other using the Transmission Control Protocol / Internet Protocol (TCP / IP) protocol suite. The core of the Internet lies in the backbone of high-speed data communication lines between the main nodes or host computers, including thousands of commercial, government, educational, and other electronic systems that route data and messages.

[0032] The head-mounted display 104D (also referred to as AR glasses 104D) includes one or more cameras (e.g., visible light cameras), microphones, speakers, one or more inertial sensors (e.g., gyroscopes, accelerometers), and a display. The cameras and microphones are used to capture video and audio data from the scene where the AR glasses 104D are located, while the one or more inertial sensors are used to capture inertial sensor data. In some cases, the cameras capture the gestures of the user wearing the AR glasses 104D. In some cases, the microphones record ambient sounds, including the user's voice commands. In some cases, both the video or static visual data captured by the visible light cameras and the inertial sensor data measured by the one or more inertial sensors are used to determine and predict the device pose. The video, static images, audio, or inertial sensor data captured by the AR glasses 104D are processed by the AR glasses 104D, the server 102, or both to identify the device pose. The device pose is used to control the AR glasses 104D itself or to interact with the applications (e.g., game applications) executed by the AR glasses 104D. In some embodiments, the display of the AR glasses 104D displays a user interface, and the identified or predicted device pose is used to render virtual objects or interact with user-selectable display items on the user interface with high fidelity.

[0033] In some embodiments, SLAM technology is applied to the data processing environment 100 to process the video data or static image data captured by the AR glasses 104D using inertial sensor data. The device pose is identified and predicted, and the scene where the AR glasses 104D are located is mapped and updated. The SLAM technology can optionally be implemented independently by the AR glasses 104D or jointly by the server 102 and the AR glasses 104D.

[0034] In some embodiments, video data or still image data captured by client device 104 (e.g., AR glasses 104D, mobile phone 104C) is also applied to determine a depth image or depth map. The client device 104 obtains a current image captured by a camera at a current camera pose and a reference image captured at a reference camera pose. A relaxed image rectification operation is applied to the reference image with respect to the current image to generate a rectified reference image. A current epipolar line of the current image is parallel to a reference epipolar line of the rectified reference image. Each current pixel of at least one subset of the current image has a current pixel position in the current image and corresponds to a corresponding reference pixel of the reference image. For each current pixel of at least one subset of the current image, a plurality of pixel transformation parameters (e.g., two pixel transformation parameters A and B) are predetermined based on the current pixel position of the current pixel and stored in an electronic device (e.g., in lookup table 250). During depth mapping, for each current pixel of the subset of the current image, a plurality of predetermined pixel transformation parameters are directly extracted. The electronic device determines a disparity between a current distance of the current pixel from a current pole of the current image and a reference distance of the corresponding reference pixel from a reference pole of the reference image, and converts the disparity of the current distance and the reference distance into a depth value of the current pixel based on the predetermined pixel transformation parameters. A depth map of the current image is created to include the depth value of each current pixel of the subset of the current image.

[0035] Figure 2is a block diagram showing an electronic system 200 according to some embodiments. The electronic system 200 includes a server 102, a client device 104, a storage device 106, or a combination thereof. Examples of the electronic system 200 include a mobile phone 104C or AR glasses 104D. The electronic system 200 generally includes one or more processing units (CPUs) 202, one or more network interfaces 204, a memory 206, and one or more communication buses 208 for interconnecting these components (sometimes referred to as a chipset). The electronic system 200 includes one or more input devices 210 that facilitate user actions, such as a keyboard, a mouse, a voice command action unit or a microphone, a touch screen display, a touch-sensitive action pad, a gesture capture camera, or other action buttons or controls. Additionally, in some embodiments, the client device 104 of the electronic system 200 uses a microphone for voice recognition or a camera 260 for gesture recognition to supplement or replace the keyboard. In some embodiments, the client device 104 includes one or more cameras 260 (e.g., RGB cameras), scanners, or light sensor units for capturing images of, for example, a graphic sequence code printed on the electronic device. The electronic system 200 also includes one or more output devices 212 for presenting a user interface and displaying content, including one or more speakers and / or one or more visual displays.

[0036] Optionally, the client device 104 includes a position detection device, such as a GPS (Global Positioning System) or other geographical location receiver, for determining the position of the client device 104. Optionally, the client device 104 includes an inertial measurement unit (IMU) 280 that integrates sensor data captured by multi-axis inertial sensors to estimate the position and orientation of the client device 104 in space. Examples of one or more inertial sensors of the IMU 280 include, but are not limited to, gyroscopes, accelerometers, magnetometers, and inclinometers.

[0037] The memory 206 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices, and optionally includes non-volatile memory, such as one or more disk memories, one or more optical disk memories, one or more flash devices, or one or more other non-volatile solid-state memories. Optionally, the memory 206 includes one or more memories remote from one or more processing units 202. The memory 206 or optionally the non-volatile memory within the memory 206 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 206 or the non-transitory computer-readable storage medium of the memory 206 stores the following programs, modules, and data structures or subsets or supersets thereof: · An operating system 214, including processes for handling various basic system services and for performing hardware-related tasks; · A network communication module 216, for connecting each server 102 or client device 104 to other devices (e.g., servers 102, client devices 104, or storage devices 106) via one or more (wired or wireless) network interfaces 204 and one or more communication networks 108 such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc.; · A user interface module 218, for enabling the presentation of information (e.g., the graphical user interface of an application 224, widgets, websites and their web pages, and / or games, audio and / or video content, text, etc.) at each client device 104 via one or more output devices 212 (e.g., a display, speakers, etc.); · An input processing module 220, for detecting one or more user behaviors or interactions from one of the one or more input devices 210 and interpreting the detected behaviors or interactions; · A web browser module 222, for navigating, requesting (e.g., via HTTP) and displaying websites and their web pages, including a web interface for logging into a user account associated with the client device 104 or another electronic device, for controlling the client or electronic device when associated with the user account, and for editing and viewing settings and data associated with the user account; · One or more user applications 224 (e.g., games, social network applications, smart home applications, extended reality applications, and / or other web-based or non-web-based applications for controlling another electronic device and viewing data captured by such device) executed by the electronic system 200; · A model training module 226, for receiving training data and establishing a data processing model 240 for processing content data (e.g., video, image, audio, or text data) to be collected or obtained by the client device 104; · A data processing module 228, for processing content data (e.g., image data) using the data processing model 240 to identify information contained in the content data, match the content data with other data, classify the content data, or synthesize related content data, wherein in some embodiments, the data processing module 228 is associated with an extended reality application and includes a depth determination module 236 for determining the depth map of an image; · A pose determination and prediction module 230 for determining and predicting the pose of the client device 104 (e.g., the AR glasses 104D), where in some embodiments, the pose determination and prediction module 230 includes a SLAM module 232 for mapping the scene where the client device 104 is located and identifying the pose of the client device 104 within the scene using image and IMU sensor data; · A pose-based rendering module 234 for rendering virtual objects above the field of view of the camera 260 of the client device 104, or creating hybrid, virtual, or augmented reality content using the images captured by the camera 260, where virtual objects are rendered and hybrid, virtual, or augmented reality content is created from the perspective of the camera 260 based on the camera pose of the camera 260; · One or more databases 242 for storing at least data including one or more of the following: ○ Device settings 244, including common device settings of one or more servers 102 or client devices 104 (e.g., service layer, device model, storage capacity, processing power, communication capabilities, etc.); ○ User account information 246 of one or more user applications 224, such as username, security questions, account history data, user preferences, and predefined account settings; ○ Network parameters 248 of one or more communication networks 108, such as IP address, subnet mask, default gateway, DNS server, and hostname; ○ Training data 238 for training one or more data processing models 240; ○ Data processing models 240 for processing content data (e.g., video, image, audio, or text data) using deep learning techniques, where the data processing model 240 includes an encoder-decoder network for enhancing the quality of depth images with reference to the current image; and ○ Lookup table 250 for the current image, where each current image is associated with a lookup table 250 that stores one or more of the following: the distance of each pixel from the pole of the current image, the tilt angle of the epipolar line in the corrected reference image, the parallax range of the parallax between this distance and the reference distance, and multiple pixel transformation parameters; ○ One or more reference selection criteria 252 for selecting a reference image from multiple candidate images for each current image; ○ Content data 254 includes at least image data (e.g., current image, reference image) and depth data determined from the image data.

[0038] Optionally, one or more databases 242 are stored in one of the server 102, the client device 104, and the storage device 106 of the electronic system 200. Optionally, one or more databases 242 are distributed across more than one of the server 102, the client device 104, and the storage device 106 of the electronic system 200. In some embodiments, more than one copy of the above data is stored at different devices. For example, two copies of the data processing model 240 are respectively stored at the server 102 and the storage device 106.

[0039] Each of the above-identified elements may be stored in one or more of the previously mentioned memory devices and corresponds to an instruction set for performing the above functions. The above-identified modules or programs (i.e., instruction sets) need not be implemented as separate software programs, processes, modules, or data structures, and thus subsets of these modules may be combined or otherwise rearranged in various embodiments. In some embodiments, the memory 206 optionally stores a subset of the above-identified modules and data structures. Additionally, the memory 206 optionally stores additional modules and data structures not described above.

[0040] Figure 3 is an example data processing system 300 for training and applying a neural network based (NN-based) data processing model 240 for processing content data 254 (e.g., video, image, audio, or text data) according to some embodiments. The data processing system 300 includes a model training module 226 for establishing the data processing model 240 and a data processing module 228 for processing the content data 254 using the data processing model 240. In some embodiments, both the model training module 226 and the data processing module 228 are located on the client device 104 of the data processing system 300, and a training data source 304 different from the client device 104 provides training data 238 to the client device 104. Optionally, the training data source 304 may be the server 102 or the storage device 106. Alternatively, in some embodiments, both the model training module 226 and the data processing module 228 are located on the server 102 of the data processing system 300. The training data source 304 that provides the training data 238 may be the server 102 itself, another server 102, or the storage device 106. Additionally, in some embodiments, the model training module 226 and the data processing module 228 are respectively located on the server 102 and the client device 104, and the server 102 provides the trained data processing model 240 to the client device 104.

[0041] The model training module 226 includes one or more data preprocessing modules 308, a model training engine 310, and a loss control module 312. The data processing model 240 is trained according to the type of the content data 254 to be processed. The training data 238 is of the same type as the content data 254, so the data preprocessing module 308 is applied to process the training data 238 of the same type as the content data 254. For example, the image preprocessing module 308A is configured to process the image training data 238 into a predefined image format. For example, a region of interest (ROI) is extracted from each training image, and each training image is cropped to a predefined image size. Alternatively, the audio preprocessing module 308B is configured to process the audio training data 238 into a predefined audio format. For example, each training sequence is converted to the frequency domain using Fourier transform. The model training engine 310 receives the preprocessed training data 238 provided by the data preprocessing module 308, further processes the preprocessed training data 238 using the existing data processing model 240, and generates an output from each training data item. During this process, the loss control module 312 can monitor the loss function comparing the output related to the training data item with the ground truth of the training data item. The model training engine 310 modifies the data processing model 240 to reduce the loss function until the loss function meets the loss criterion (e.g., the comparison result of the loss function is minimized or reduced below the loss threshold). The modified data processing model 240 is provided to the data processing module 228 for processing the content data 254.

[0042] In some embodiments, the model training module 226 provides supervised learning, where the training data 238 is fully labeled, and the supervised learning includes the expected output (also referred to as the ground truth in some cases) for each training data item. Conversely, in some embodiments, the model training module 226 provides unsupervised learning, where the training data 238 is unlabeled. The model training module 226 is configured to identify patterns in the training data 238 that have not been previously detected, without pre-existing labels and with little or no human supervision. Additionally, in some embodiments, the model training module 226 provides semi-supervised learning, where the training data 238 is partially labeled.

[0043] The data processing module 228 includes a data preprocessing module 314, a model-based processing module 316, and a data postprocessing module 318. The data preprocessing module 314 preprocesses the content data 254 based on the type of the content data 254. The function of the data preprocessing module 314 is consistent with that of the preprocessing module 308, and converts the content data 254 into a predefined content format acceptable to the stream of the model-based processing module 316. Examples of the content data 254 include one or more of the following: video, image, audio, text, and other types of data. For example, each image is preprocessed to extract the ROI therefrom or crop it to a predefined image size, and each audio clip is preprocessed to convert it into the frequency domain using the Fourier transform. In some cases, the content data 254 includes more than two types, for example, video data and text data. The model-based processing module 316 applies the trained data processing model 240 provided by the model training module 226 to process the preprocessed content data 254. The model-based processing module 316 can also monitor the error indicator to determine whether the content data 254 has been correctly processed in the data processing module 228. In some embodiments, the processed content data 254 is further processed by the data postprocessing module 318 to display the processed content data 254 in a preferred format or provide other relevant information derived from the processed content data 254.

[0044] Figure 4A is an example neural network (NN) 400 applied to process the content data 254 in the data processing model 240 based on a neural network according to some embodiments; and Figure 4B is an example of a node 420 in the neural network (NN) 400 according to some embodiments. The data processing model 240 is established based on the neural network 400. The corresponding model-based processing module 316 applies the data processing model 240 including the neural network 400 to process the content data 254 that has been converted into the predefined content format. The neural network 400 includes a set of nodes 420 connected by links 412. Each node 420 receives one or more node streams and applies a propagation function to generate a node output from the node streams. When the node output is provided to one or more other nodes 420 through one or more links 412, the weight w associated with each link 412 is applied to the node output. Similarly, according to the propagation function, based on the corresponding weight w 1 、w 2 、w 3 、and w 4 ,the node streams are combined. For example, the propagation function is the product of a non-linear activation function and a linear weighted combination of the node streams.

[0045] The set of nodes 420 is organized into one or more layers in the neural network 400. Optionally, one or more layers include a single layer that serves as both the current layer and the output layer. Optionally, one or more layers may include a current layer 402 for receiving a stream, an output layer 406 for providing an output, and zero or more hidden layers 404 (e.g., 404A and 404B) between the current layer 402 and the output layer 406. A deep neural network has more than one hidden layer 404 between the current layer 402 and the output layer 406. In the neural network 400, each layer is connected only to its immediately preceding and / or immediately succeeding layer. In some embodiments, since each node 420 in layer 402 or layer 404B is connected to each node 420 in the immediately succeeding layer, layer 402 or layer 404B is a fully connected layer. In some embodiments, one of the hidden layers 404 includes two or more nodes that are connected to the same nodes in the immediately succeeding layer for downsampling or pooling the nodes 420 between these two layers. In particular, max pooling uses the maximum value of two or more nodes in layer 404B to generate a node in the immediately succeeding layer 406 that is connected to these two or more nodes.

[0046] In some embodiments, a convolutional neural network (CNN) is applied in the data processing model 240 to process content data 254 (particularly video data and image data). The CNN employs convolutional operations and belongs to a type of deep neural network 400, namely a feedforward neural network that only moves data forward from the current layer 402 through the hidden layers to the output layer 406. The hidden layers of the CNN can be convolutional layers that perform convolution using multiplication or dot product. Each node in the convolutional layer receives a stream from a receptive field associated with the previous layer (e.g., five nodes), the receptive field being smaller than the entire previous layer and can vary based on the position of the convolutional layer in the convolutional neural network. The video data or image data is preprocessed into a predefined video / image format corresponding to the stream of the CNN. The preprocessed video data or image data is abstracted by each layer of the CNN into corresponding feature maps. By these methods, the CNN can process video data and image data for video and image recognition, classification, analysis, printing, or synthesis.

[0047] Alternatively and additionally, in some embodiments, a recurrent neural network (RNN) is also applied in the data processing model 240 to process content data 254 (especially text data and audio data). Nodes in consecutive layers of the RNN follow a time series, such that the RNN exhibits time-dynamic behavior. For example, each node 420 of the RNN has a real-valued activation that varies over time. Examples of RNNs include, but are not limited to, long short-term memory (LSTM) networks, fully recurrent networks, Elman networks, Jordan networks, Hopfield networks, bidirectional associative memory (BAM) networks, echo state networks, independently RNN (IndRNN), recurrent neural networks, and neural history compressors. In some embodiments, the RNN can be used for handwriting recognition or speech recognition. It should be noted that, in some embodiments, the data processing module 228 processes two or more types of content data 254 and applies two or more types of neural networks (e.g., CNN and RNN) to jointly process the content data 254.

[0048] The training process is to calibrate all weights w of each layer of the data processing model 240 using the training data set provided in the current layer 402. i The training process generally includes two steps: forward propagation and backward propagation, which are repeated multiple times until a predefined convergence condition is met. In forward propagation, the weight sets of different layers are applied to the current data and the intermediate results of the previous layers. In backward propagation, the error range of the output (such as the loss function) is measured, and the weights are adjusted accordingly to reduce the error. Optionally, the activation function can be a linear function, rectified linear unit (ReLU), sigmoid function, hyperbolic tangent function, or other types. In some embodiments, before applying the activation function, the network bias parameter b is added to the weighted output sum of the previous layer. The network bias parameter b provides a perturbation that helps the NN 400 avoid overfitting the training data. The results of the training include the network bias parameter b for each layer.

[0049] Figure 5A is an example stereoscopic vision environment 500 for forming a three-dimensional (3D) image including the current image 502 and the corresponding depth map (e.g., Figure 6 the depth map 612 in). The depth map 612 is generated based on comparing the current image 502 with the reference image 504. Both the current image 502 and the reference image 504 are captured by the camera 260 (Figure 2 ) Capture. The current image 502 corresponds to the current camera pose of the camera 260. For example, when the camera 260 is at the current camera position O 1 (506C). The reference image 504 corresponds to the reference camera pose of the camera 260. For example, when the camera 260 is at the reference camera position O 2 (506R). At least one subset of the current image 502 records the same content as the subset of the reference image 504. Each current pixel 508C of at least one subset of the current image 202 has a current pixel position in the current image 502 and corresponds to the corresponding reference pixel 508R of the reference image 504. For example, the object point 514 is located at the object position P(508O) in the stereoscopic vision environment 500 and is captured in both the current image 502 and the reference image 504. The object point 514 located at the object position 508O corresponds to the current pixel p(508C) in the current image 502 and the reference pixel p'(508R) in the reference image 504.

[0050] The reference image 504 and the current image 502 are captured by the moving camera 260 at two different times. Optionally, the reference image 504 is captured before or after the current image 502 is captured. The camera 260 is at the current camera position O 1 (506C) and the reference camera position O 2 (506R) has a position offset. The current camera position O 1 (506C) is located in the field of view of the camera 260 at the reference camera position O 2 (506R) and is projected onto the reference pole e'(510R) in the reference image 504. The reference camera position O 2 (506R) is located in the field of view of the camera 260 at the current camera position O 1 (506C) and is projected onto the current pole e(510C) in the current image 502.

[0051] In the current image 502, the current epipolar line 512C connects the current pixel p(508C) and the current pole e(510C) and has a current distance ρ 0 . In the reference image 504, the reference epipolar line 512R connects the reference pixel p'(508R) and the reference pole e'(510R) and has a reference distance ρ 1 . The current distance ρ 0 and the reference distance ρ 1 have a parallax D p0 . The object point located at the object position P(508O) corresponds to the object located at the current camera position O 1The depth D in the field of view of the camera 260 of (506C). The depth D is related to the current distance ρ 0 and the reference distance ρ 1 by the following relationship with the parallax D p0 between them: where C 1 , C 2 , C 3 , and C 4 is a set of pixel conversion parameters associated with the current pixel 508C corresponding to the object point 514. For each individual pixel in the current image 502, the pixel conversion parameter 608 is different. In some embodiments, the set of pixel conversion parameters 608 is pre-determined and stored in a look-up table 250 in association with each pixel of the current image 502 (e.g., the current pixel 508C). Thus, given the current distance ρ 0 corresponding to the current pixel 508C and the parallax of the reference distance ρ 1 , the corresponding depth D is related to the parallax D p0 by the following relationship.

[0052] Figure 5B is an image pair according to some embodiments including an example current image 502 and an example rectified reference image 504'. A subset of the current image 502 includes the current epipole 510C, and a subset of the reference image 504 includes the reference epipole 508R. The subsets of the current image 502 and the reference image 504 both correspond to a portion of the field of view around the object point 514 located at the object position P(508O). In some embodiments, a relaxation image rectification operation is applied to the reference image 504 with respect to the current image 502 to generate the rectified reference image 504' such that the current epipolar line 512C of the current image 502 is parallel to the reference epipolar line 512R' of the rectified reference image 504'. In some cases, the current epipole e(510C) is identified by the current coordinate values (eu 0 , ev 0 ) in the current image 502, and the current epipole e'(510R') is identified by the rectified coordinate values (eu R , ev R ) in the rectified reference image 504'. The current coordinate values (eu 0 , ev 0 ) are equal to the rectified coordinate values (eu R , ev R ). Additionally, in some embodiments, in the relaxation image rectification operation, the reference image 504 is rotated, scaled, and / or tilted.

[0053] Alternatively, in some embodiments not shown, a relaxation image rectification operation is applied to the current image 502 with reference to the reference image 504 such that the epipolar line 512C' of the rectified current image 502' is parallel to the reference epipolar line 512R' of the reference image 504. Alternatively, in some embodiments not shown, two separate relaxation image rectification operations are applied to the current image 502 and the reference image 504 respectively such that the epipolar line 512C' of the rectified current image 502' is parallel to the reference epipolar line 512R' of the rectified reference image 504'. A preliminary depth map is generated from the rectified current image 502', and the preliminary depth map is back-rectified to generate a depth map 612 corresponding to the current image 502.

[0054] It should be noted that, in some embodiments, the current image 502 and / or the reference image 504 are rectified in a relaxed manner. The current epipolar line 512C of the current image 502 and the reference epipolar line 512R' of the rectified reference image 504' are parallel and do not need to be aligned with the horizontal axis. In some embodiments, the current and reference poles 510 are located at the same position in the respective image planes of the current image 502 and the rectified reference image 504' and do not need to share the same vertical coordinate value v 0 and v R .

[0055] In the rectified reference image 504', the reference epipolar line 512R' connects the reference pixel p'(508R') and the reference pole e'(510R') and has a reference distance ρ R . The current distance ρ 0 and the reference distance ρ R have a disparity D p . For each current pixel of the current image 502, the depth D is related to the current distance ρ 0 and the reference distance ρ R by the disparity D p as follows: where f is the focal length of the lens of the camera 260 that captures the current image 502 and the reference image 504, and Base is the baseline. Thus, the depth D of each current pixel of the current image 502 is determined based on the corresponding disparity D p as follows: Wherein, A and B are two pixel conversion parameters associated with the current pixel 508C corresponding to the object point 514. For each individual pixel in the current image 502, these two pixel conversion parameters are different. In some embodiments, these two pixel conversion parameters are pre-determined and stored in a look-up table 250 in association with each pixel of the current image 502 (e.g., the current pixel 508C). The look-up table 250 stores the tilt angles of a plurality of epipolar lines 512C.

[0056] In some embodiments, a planar rectification technique is applied to rectify one or both of two images with arbitrary relative poses such that the two images share a common image plane, the corresponding epipolar lines become parallel to the horizontal axis, and / or the corresponding points in the two images have the same vertical coordinates. Polar coordinate rectification focuses on rectifying a pair of the current image 502 and the reference image 504 with arbitrary relative poses. The current image 502 and the reference image 504 are resampled according to a polar coordinate system centered on the poles 510 of the current image 502 and the reference image 504. In the rectified image 502' or 504', each epipolar line 512 around the pole 510 is distorted into a horizon. After polar coordinate rectification, the vertical direction of the rectified reference image 504' corresponds to the tilt angle of the corresponding epipolar line 512R' of the rectified reference image 504', and the horizontal direction of the rectified reference image 504' corresponds to the distance to the corresponding pole 510R' of the rectified reference image 504'. In this application, polar coordinate rectification and planar rectification are unified in a single framework, and the current image 502 is accessed through the look-up table 250, avoiding distortion in polar coordinate rectification.

[0057] Figure 6 is a flowchart of an example process 800 for generating a depth map 612 corresponding to the current image 502 according to some embodiments. The current image 502 and the reference image 504 are respectively captured by the camera 260 at two different positions in the stereo vision environment 500, and the above two different positions include the current camera position O 1 (506C) and the reference camera position O 2 (506R). A relaxation image rectification operation is applied to the reference image 504 with reference to the current image 502 to generate a rectified reference image 504'. The current epipolar line 512C of the current image 502 is parallel to the reference epipolar line 512R' of the rectified reference image 504'. Each current pixel 508C of the current image 502 has its own disparity 604, and the disparity 604 is in (1) the current distance ρ between the corresponding current pixel 508C and the current pole 510C in the subset of the current image 0 and (2) the reference distance ρ between the corresponding reference pixel 508R' and the reference pole 510R' of the rectified reference image 504' RMeasured between. The respective disparities 604 of the current pixels 508C of the current image 502 form a first disparity map 606.

[0058] For each current pixel 508C in the current image 502, a plurality of pixel conversion parameters 608 (e.g., A and B) are determined and stored. For example, the pixel conversion parameters 608 are determined based on the reference camera pose, the current camera pose, the camera intrinsic matrix K, and the pole position in the current image 502, for example, as Figure 7 described by formulas (5)-(12) in. In some embodiments, the plurality of pixel conversion parameters 608 are related to the current distance ρ0 of the current pixel 508C and the pole e(510C) of the current image 502, the information of the inclination angle of the epipolar line 512C' in the rectified reference image 504', and the current distance ρ 0 and the reference distance ρ R of the disparity 604 are stored together in the look-up table 250. For example, the current distance ρ 0 and the reference distance ρ R of the disparity Dp(604) are determined according to one or more of the following: the current distance ρ in the current image 502 0 , the information of the inclination angle of the epipolar line 512C' in the rectified reference image 504', the current distance ρ 0 and the reference distance ρ R of the disparity range. Further, according to the pixel conversion parameters 608 (e.g., A and B), based on formula (4), the depth value D(610) of each current pixel 508C in a subset of the current image 502 is determined from the disparity 604 of the current distance ρ 0 and the reference distance ρ R . The depth values 610 of the subset of the current image 502 are applied to create a depth map 612 corresponding to the current image 502.

[0059] In some embodiments, after capturing the current image 502, a reference image 504 is selected from a plurality of candidate images 602 (e.g., 602A, 602B, 602C) based on a reference selection criterion 252. In some embodiments, the reference selection criterion 252 requires at least one of the following conditions: the angle 516 between two camera bearings (also known as pointing directions) associated with the current image 502 and the reference image 504 satisfies the bearing angle requirement, the current image 502 and the reference image 504 share at least a threshold number of common feature points, and the error between the projected camera pose and the current camera pose is less than the pose error threshold. Based on the reference camera pose and the common feature points of the reference image 504 and the current image 502, a projected camera pose is determined for the current image 502. Specifically, in some embodiments, for each of a subset of the candidate images 602, an angle 422 between the current camera bearing 524C associated with the current image 502 and a candidate camera bearing (e.g., the reference camera bearing 524R associated with the reference image 504) associated with the candidate image 602 is determined. If the angle corresponding to one of the candidate images 602 satisfies the bearing angle requirement (e.g., the angle 516 falls within a predefined bearing angle range), then that one candidate image 602 is selected as the reference image 504. In one example, the angle 516 corresponding to that one candidate image 602 is closer to the predefined bearing angle than the angles corresponding to any other candidate image 602, and thus, that one candidate image 602 is selected as the reference image 504.

[0060] Alternatively, in some embodiments, for each of a subset of the candidate images 602, a certain number of common feature points exist in both the candidate image 602 and the current image 502. If the number of common feature points determined based on one of the candidate images 602 is within the range of the number of feature points, then that one candidate image 602 is selected as the reference image. An example of the range of the number of feature points is greater than or equal to the feature point number threshold, e.g., ≥500. In some cases, since the number of common feature points determined based on one of the candidate images 602 is greater than the number of common feature points determined based on any other candidate image 602, that one candidate image 602 is selected as the reference image.

[0061] Additionally and alternatively, in some embodiments, a plurality of current feature points are identified in the current image 502. For each of a subset of the candidate images 602, a plurality of candidate feature points are identified in the candidate image 602, and a candidate camera pose corresponding to the current image 502 is estimated by comparing the candidate feature points and the current feature points. A camera pose error is derived between the candidate camera pose of the current image 502 and the current camera pose. If the camera pose error determined based on the candidate feature points of one of the candidate images 602 is less than a pose error threshold, then one of the candidate images 602 is selected as the reference image. Alternatively, in one example, the camera pose error determined based on the candidate feature points of one of the candidate images 602 is the smallest among the camera pose errors determined based on all of the candidate images 602, and one of the candidate images 602 is selected as the reference image 504.

[0062] Figure 7 is a list of examples of formula 700 applied to convert image data to depth data according to some embodiments. A depth map 612 corresponding to the current image 502 is generated based on a disparity map 606 of the current distance of the current image 502 and the reference distance of the rectified reference image 504'. Both the current image 502 and the reference image 504 are captured by the camera 260. The current image 502 corresponds to the current camera pose of the camera 260, and the reference image 504 corresponds to the reference camera pose of the camera 260. For the current image 502, the current camera pose includes a current translation position (i.e., the current camera position T 0 ) and a current rotation position (i.e., the current camera orientation R 0 ). For the reference image 504, the reference camera pose includes a reference translation position (i.e., the reference camera position T 1 ) and a reference rotation position (i.e., the reference camera orientation R 1 ). The camera 260 has a camera intrinsic matrix K.

[0063] The relaxed image rectification operation includes a homography transformation H represented by formula (5). The coordinates of the current epipole e (510C) in the current image 502 are R 0 T (T 1 -T 0 ). The coordinates of the reference epipole e' (510R) in the reference image 504 are R 1 T (T 1 -T 0 ), and are converted to R 0 T (T 1 -T 0)。According to formula (6), based on the camera intrinsic matrix K, the current camera orientation R 0 , the current camera position T 0 , and the reference camera position T 1 determine the vector [t 1 t 2 t 3 T . The coordinate values of the epipoles 510 in both the current image 502 and the rectified reference image 504' are as shown in formula (7). In other words, after the relaxed image rectification, the two epipoles 510C and 510R' have the same coordinate values and are located at the same positions in the corresponding current image 502 and rectified reference image 504' respectively.

[0064] The current epipole e(510C) of the current image 502 is represented as (eu 0 , ev 0 ), and the current pixel 508C of the current image 502 is located at (u, v) of the current image 502. The current pixel 508C of the current image 502 corresponds to the reference pixel 508R in the rectified reference image 504', which is located at the position (u R , v R ) in the rectified reference image 504' and is shifted to the position shown in formula (8) in the rectified reference image 504'. Given that the coordinate values of the epipoles 510 in both the current image 502 and the rectified reference image 504' are as shown in formula (7), the disparity 604 (D p ) is associated with the disparity values Disparity x and Disparity y represented by formulas (9.1) and (9.2). Therefore, for the current pixel 508C located at (u, v) and the current epipole 510C located at (eu 0 , ev 0 ), the disparity D p is represented in formula (10), where the parameter Norm is represented in formula (11). Based on formulas (10) and (11), the current distance ρ 0 of the current frame 502 and the reference distance ρ R of the rectified reference frame 504' with respect to the disparity 604 (D p ) are further determined as follows: where Norm and t 3 are represented in formulas (11) and (6) respectively.

[0065] ​In addition, the tilt angle of the reference epipolar line 512R is α, as described by cos(α) and sin(α) in Equation (12), and it is also established based on the parameter Norm. In some embodiments, the parameters Norm, cos(α), and sin(α) are constants determined by the position (u, v) of each current pixel 508C, pre-calculated and stored in the lookup table 250 to facilitate the determination of the parallax Dp (604) and the depth value D (610). According to Equation (4'-2), the parallax Dp (604) is proportional to the reciprocal of the adjusted depth value D-t 3 where the adjusted depth value is equal to the depth value D (610) minus the parameter t 3 . When t 3 is equal to zero, the pole 510C is located at infinity, and the corresponding tilt angle α of the corresponding epipolar line 512C is 0, indicating that the relaxation correction becomes a normal planar correction. When t 3 is not equal to 0, the relaxation correction is an implicit polar coordinate correction without image distortion, and the depth value of the current pixel of the current image 502 is determined according to the pre-determined lookup table 550.

[0066] Figure 8 is a flowchart of an exemplary process 800 for converting image data (e.g., the current image 502) into depth information (e.g., a depth image or depth map 612) according to some embodiments. Optionally, the process 800 is jointly implemented by the depth determination module 236 and the SLAM module 232 of the electronic system 200, and includes one or more of reference selection 802, homography transformation 803, image feature extraction 804, cost volume construction 806, semi-global matching 808, parallax-to-depth conversion 810, and depth denoising 812. In this process 800, the reference image 504 is rectified and a cost volume 814 is constructed using the current image 502 and the rectified reference image 504', and the cost volume 814 is further processed by a semi-global matching algorithm to estimate the depth map 612. Optionally, the depth map 612 is processed to remove its noise and improve its image quality.

[0067] In some embodiments, for each current pixel 508C in the current image 502, the pixel information of the current image 502 and the rectified reference image 504' is directly stored in the lookup table 250 during the lookup table generation. Examples of such directly stored pixel information include, but are not limited to, for example, the current distance ρ 0 , the reference distance ρ R , the tilt angle of the epipolar line 512C or 512R', the current distance ρ 0 and the reference distance ρ RThe parallax range. Alternatively, in some embodiments, for each pixel in the current image 502, the pixel information of the current image 502 and the reference image 504 is associated according to one or more relevant formulas (e.g., formulas (1)-(12)), and the coefficients of these formulas (e.g., pixel conversion parameters 608 (e.g., A, B)) are stored in the look-up table 250 during look-up table generation. The pixel information of the current image 502 and the corrected reference image 504' is extracted from the look-up table 250 and applied to the parallax-to-depth conversion 810. Reference frame selection

[0068] The current image 502 and the reference image 504 are jointly applied to achieve stereoscopic imaging (i.e., measure parallax and determine the depth map 612), and in the reference selection 802, the reference image 504 is selected from a plurality of candidate images 602 (e.g., Figure 6 images 602A, 602B, and 602C in) based on the reference selection criterion 252. Optionally, the reference image 504 is captured before or after the current image 502 and corresponds to a reference camera pose different from the current camera pose associated with the current image 502. In some embodiments, the reference selection criterion 252 requires at least one condition: the angle 516 between the two camera orientations associated with the current image 502 and the reference image 504 satisfies the azimuth requirement, the current image 502 and the reference image 504 share at least a threshold number of common feature points, and the error between the projected camera pose and the current camera pose is less than the pose error threshold. Based on the reference camera pose and the common feature points of the reference image 404 and the current image 402, the projected camera pose of the current image is determined.

[0069] In some embodiments, a sufficient overlapping region is formed between the current image 502 and the reference image 504. This overlapping region is measured by the angle 516 of the camera orientations 524C and 524R of the current image 502 and the reference image 504. For example, the angle of the camera orientations 524C and 524R is less than the azimuth threshold. Alternatively or additionally, in some embodiments, the translation from the current image 502 to the reference image 504 is within a reasonable range that is neither too small nor too large, i.e., the distance between the camera positions 506C and 506R is within the range between two distance thresholds. In other words, the number of common sparse feature points shared by the current image 502 and the reference image 504 is greater than the threshold number of common feature points. Further, in some embodiments, the relative pose error between the reference image 504 and the current image 502 is measured by the reprojection error of the common sparse feature points and is less than the pose error threshold according to the reference selection criterion 252. Relaxed image rectification (e.g., homography transformation)

[0070] In some embodiments, a relaxation image rectification operation is applied to a reference image 504 with reference to a current image 502 to generate a rectified reference image 504'. The relaxation image rectification operation includes a homography transformation H (803). Additionally, in some embodiments, the reference image 504 is rotated, scaled, and / or tilted in the relaxation image rectification operation. Refer Figure 5B , a current epipolar line 512C of the current image 502 is parallel to a reference epipolar line 512R' of the rectified reference image 504'. In some cases, a current pole e (510C) is identified by current coordinate values (eu 0 , ev 0 ) in the current image 502, and a rectified pole e' (510R') is identified by rectified coordinate values (eu R , ev R ) of the rectified reference image 504'. The current coordinate values (eu 0 , ev 0 ) are equal to the rectified coordinate values (eu R , ev R ). The respective current distance ρ 0 of each current pixel 508C is measured from the current pole 510C in a subset of the current image, and the respective reference distance ρ R of each reference pixel 508R' is measured from the reference pole 510R' of the rectified reference image 504'. For each current pixel 508C in the subset of the current image 502, the respective disparity 604 corresponds to the difference between the current distance ρ 0 and the reference distance ρ R . The respective disparities 604 of a plurality of pixels of the current image 502 form a first disparity map 606 corresponding to the current image 502. Image feature extraction

[0071] During image feature extraction 804 (e.g., 804C and 804R), image features are extracted (e.g., by an image feature extraction network) from the current image 502 and the rectified reference image 504'. In some embodiments, the image features extracted from the current image 502 and the rectified reference image 504' include, but are not limited to, RGB color information, gradients of the image, or Census Transform of the grayscale image. For example, pixel-level image features are determined, and the pixel-level image features include the grayscale level or gradient of each image pixel of the current image 502 or the rectified reference image 504'. The Census Transform is applied to extract the image features of images 502 and 504'. For each image pixel, the corresponding image feature is stored in a 32-bit descriptor generated from a 7×7 region that includes 48 pixels surrounding the corresponding pixel. Optionally, the region is centered on the image pixel or not centered on the image pixel. The grayscale level or gradient of the image pixel is compared with the grayscale levels or gradients of 32 selected adjacent pixels in the region to determine the 32 bits in the 32-bit descriptor. See Figure 11 , the 7×7 region includes 7 rows and 7 columns. The image pixel is marked as "o", and 32 adjacent pixels selected from the remaining 48 pixels are marked as "×". For each of the 32 adjacent pixels, if the grayscale level of the image pixel is greater than the grayscale level of the adjacent pixel, the corresponding bit in the 32-bit descriptor is equal to a first value (e.g., "1"), and conversely, if the grayscale level of the image pixel is equal to or less than the grayscale level of the adjacent pixel, the corresponding bit in the 32-bit descriptor is equal to a second value (e.g., "0"). Thus, each image pixel of the current image or the rectified reference image corresponds to an image feature that is a 32-bit integer descriptor generated from a corresponding region that includes the image pixel and has more than 32 pixels. Lookup table generation

[0072] In some embodiments, the current camera position O 1 (506C) corresponds to the optical center of the current image 502, and the reference camera position O 2 (506R) corresponds to the optical center of the reference image 504. The current epipole e (510C) is located in the current image 502, and the reference epipole e' (510R') is located in the rectified reference image 504'. In the current image 502, the current epipolar line pe (512C) connects the current pixel p (508C) and the current epipole e (510C) and has a current distance ρ 0 . In the rectified reference image 504', the reference epipolar line p'e' (512R') connects the reference pixel p' (508R') and the reference epipole e' (510R') and has a reference distance ρ R。Under the relaxation correction operation, the epipolar lines pe and p'e' are both located in the image plane and parallel to each other, and the inclination angle of the epipolar line p'e' depends on the position of the current pixel p in the image plane. The disparity D between the epipolar lines pe and p'e' p is defined based on formulas (9.1) and (9.2) and converted into the depth value D of the object point P according to formula (4) or (4'-1).

[0073] In some embodiments, the intermediate results of the epipolar lines pe and p'e' are stored in the lookup table 250, which speeds up the cost volume construction 806 and / or the disparity-to-depth conversion 810. For example, the lookup table 250 is an 8-channel matrix and has the same size as the current image 502. Each element of the matrix stores information of the corresponding pixel of the current image 502 or the corrected reference image 504', including but not limited to one or more of the following: the current distance ρ from the current pixel 508C to the current epipole e (510C) in the current image 502 0 、the cosine value and sine value of the inclination angle of the epipolar line p'e' (512R') in the corrected reference frame 504', the current distance ρ 0 and the reference distance ρ R the disparity range (e.g., maximum disparity, minimum disparity) of ρ, and the pixel conversion parameters 608 (e.g., A and B). Based on the lookup table 250, the 3D object point 514 in space is conveniently mapped to the reference pixel p' (508R'), and the reference pixel is located on the epipolar line p'e' (512R') in the corrected reference image 504'. Cost volume construction

[0074] In the cost volume construction 806, a three-dimensional (3D) cost volume with three dimensions, namely image width, image height, and disparity, is constructed. Each cost element is the Hamming distance between the census transform of the current pixel in the current image 502 and the census transform of the candidate pixel in the corrected reference image 504'. For each current pixel 508C in the current image 502, the candidate pixels (including the corresponding reference pixel 508R') of the corrected reference image 504' are located on the epipolar line p'e' that has an inclination angle associated with the current pixel 508C in the lookup table 250. In addition, the minimum disparity value and the maximum disparity value are also extracted from the lookup table 250. The candidate pixels are located on a segment of the epipolar line p'e' (512R') determined by the disparity range defined by the minimum disparity value and the maximum disparity value. In this way, the lookup table 250 can implement an efficient cost volume construction process to locate these candidate pixels, obtain each cost element, and fill them into the 3D cost volume. Semi-global matching

[0075] Identify the reference pixel 508R' of the rectified reference image 504' corresponding to each current pixel 508C of the current image 502 from among the candidate pixels on a segment of the epipolar line p'e' (512R'). Similarly, identify the current distance ρ corresponding to the reference pixel 508R' 0 and the reference distance ρ R of the disparity D p , for depth estimation. In some embodiments, apply semi-global matching 808 to estimate the optimal disparity (D p ) from the 3D cost volume. The semi-global matching method 808 employs 8-way aggregation and selects the optimal disparity from the total cost of the 8-way aggregation. The minimum and maximum disparities of each candidate pixel on the epipolar line p'e' are applied during the aggregation process to set the range of the optimal disparity. Cost aggregation is applied as follows: where C r (x, l) is the aggregated cost of pixel x with disparity l in the neighboring direction r, r ∈ Nr, and Nr is the set of neighboring directions. Eight neighborhoods are used. P 1 and P 2 are penalty values. C(x, l) is the cost value calculated in the previous cost volume construction step. Each cost value applied in Equation (13) corresponds to the image feature of a pixel in the current image or the rectified reference image. According to the semi-global matching method, identify the current distance ρ 0 and the reference distance ρ R of the optimal disparity 604 (D p ) for each current pixel 508C, and this optimal disparity corresponds to the reference pixel 508R' corresponding to the current pixel 508C. The reference pixel 508R' is a candidate pixel on the epipolar line p'e' that provides the optimal disparity 604 (Dp) between the minimum and maximum disparities. Depth image denoising

[0076] After determining the disparity 604 (Dp) of the current distance ρ 0 and the reference distance ρ R for the current pixel 508C, convert (810) the disparity 604 (Dp) to the depth value 610 of the current pixel 508C, and this depth value is used to create the depth map 612 of the current image 502. In some embodiments, determine the inverse adjusted depth value of the current pixel 508C based on Equation (4'-3) from the disparity 604 (Dp) of the current distance and the reference distance and convert it to the depth value 610 (D) of the current pixel 508C. In some embodiments, determine the depth value 610 (D) of the current pixel 508C based on Equation (4'-1) from the disparity 604 (Dp) of the current distance and the reference distance.

[0077] In addition, in some embodiments, one or more guided image filters are applied to reduce the noise level of the inverse adjusted depth map of the current image 502. In some embodiments, the uncertainty level is measured in the semi-global matching 808 and converted into a confidence map. The confidence map is multiplied by the depth map 612 to create a weighted depth map, and fast global image smoothing is applied, e.g., using a guided image filter, to generate a smoothed weighted depth map. The guided filter is applied to the confidence map to generate a smoothed confidence map. The smoothed weighted depth map is divided by the smoothed confidence map to generate a smoothed depth map. The smoothed inverse depth map is further converted into the depth map 612 corresponding to the current image 502.

[0078] In some embodiments, the process 800 is implemented by a mobile device having a single moving monocular camera 260 to provide the depth map 612. The mobile device also includes an augmented reality (AR) software development kit (SDK) that provides more information for developers to implement AR applications. The process can be implemented in the smartphone 104C to estimate the depth image within 100 milliseconds and meet the real-time requirements of AR applications. This depth estimation ability is necessary to provide near-realistic collision and occlusion effects in AR applications. In some cases, AR applications need to create a dense 3D mesh for the environment. The sparse point cloud generated by SLAM is applied to create the dense 3D mesh. In some embodiments, a time-of-flight (TOF) sensor is applied to measure the depth from the smartphone to the environment, which consumes a large amount of battery power. Instead, in some embodiments, to save battery power, the TOF sensor is not used, and a single moving monocular camera 260 is used to estimate the depth image based on the process 800. In these ways, the process 800 achieves near-realistic AR effects in the mobile device and significantly reduces computational and power resources by estimating the depth map from a monocular image sequence.

[0079] In some embodiments, the process 800 is applied in conjunction with SLAM to add more feature points from the descriptor matching process. These SLAM feature points make the previous depth estimation more accurate. In some embodiments, the process 800 is applied based on a neural network model. If the neural network model is trained in an unsupervised manner, a reprojection loss is applied by warping the current image 502 to the rectified reference image 504'. More importantly, the process 800 takes into account the epipolar geometry and determines the depth map 612 based on at least equations (9.1) and (9.2), which avoids duplicate and missing matches and achieves an accurate and effective solution for depth estimation.

[0080] Figure 9is a flowchart of an example process 900 for controlling noise in a disparity map 606 used to determine a depth map 612 of a current image 502. As a result of the relaxation correction operation, the current epipolar line 512C of the current image 502 is parallel to the reference epipolar line 512R' of the corrected reference image 504' ( Figure 5B ). The current image 502 and the corrected reference image 504' correspond to the same look-up table 250. A first disparity map 606A is determined based on stereo matching 902A of the current image 502, and a second disparity map 606B is determined based on stereo matching 902B of the corrected reference image 504'. The difference between the stereo matching operations 902A and 902B is related to the search direction on the corresponding epipolar line 512. For example, the first search direction of stereo matching 902A on the epipolar line is towards the pole, while the first search direction of stereo matching 902B on the epipolar line is away from the pole.

[0081] Alternatively, in some embodiments, a left-to-right disparity check is employed to remove noise in the disparity map 606. The current image 502 corresponds to the first disparity map 606A (D p1 ), and the corrected reference image 504' corresponds to the second disparity map 606B (D p2 ). For a pixel (u, v) in the current image 502, the corresponding tilt angle of the epipolar line pe in the look-up table 250 is b, and the corresponding reference pixel in the corrected reference image 504' is located at (u R , v R ), which is represented as follows: u R = u - D p1 (u, v)·cos(a) v R = v - D p1 (u, v)·sin(a) In some embodiments, D p1 (u, v) is compared with D p2 (u, v) to determine whether the corresponding third disparity (also referred to as the disparity difference map) exceeds a predefined disparity threshold, as follows: |D p1 (u, v) - D p2 (u R , v R )| < 1 or 2 pixels where the third disparity is equal to the absolute value of D p1 (u, v) - D p2 (u R , v R ), and an example of the predefined disparity threshold is 1 or 2. If the two disparity maps 606A and 606B are inconsistent, the corresponding pixel in the disparity map 606 is set to zero.

[0082] In some embodiments, the electronic device generates a second disparity map D corresponding to the corrected reference image 504'. p2 The second disparity map D p2 represents a second disparity between: (1) the reference distance of each reference pixel 508R' to the reference pole 510R' of a subset of the corrected reference image 504', and (2) the current distance of the corresponding current pixel 508C to the current pole 510C in the current image 502. The depth map D of the current image 502 is determined based on the first disparity map D p1 and the second disparity map D p2 . Additionally, in some embodiments, the electronic device compares the first disparity map D p1 and the second disparity map D p2 , and modifies the first disparity map D p1 according to the comparison result to generate a modified disparity map D p . The depth map of the current image 502 is determined according to the modified disparity map D p .

[0083] Moreover, in some embodiments, for each element of the first disparity map D p1 corresponding to the corresponding current pixel 508C of a subset of the current image 502, the electronic device determines the first disparity in the first disparity map D p1 corresponding to the corresponding current pixel 508C and the disparity difference map |D p2 - D p1 | between the second disparity in the second disparity map D p2 corresponding to the corresponding reference pixel 508R. Then, the electronic device determines whether to modify the corresponding element of the first disparity map based on the disparity difference map |D p1 - D p2 |. In some embodiments, according to determining that the corresponding element value of the disparity difference map |D p1 - D p2 | exceeds a predefined disparity threshold (e.g., 1 or 2 pixels), the electronic device sets the corresponding element of the first disparity map D p1 to 0. Conversely, according to determining that the corresponding element value of the disparity difference map |D p1 - Dp2| is lower than the predefined disparity threshold, the corresponding element of the first disparity map D p1 is maintained.

[0084] Figure 10It is a block diagram of an example data processing model 1000 for controlling noise of depth information according to some embodiments. In some embodiments, a depth map 612 is generated from a current image 502 and a reference image 504 in an augmented reality application. The disparity map 606 generated from semi-global matching 808 has an unsatisfactory noise level and needs to be further reduced. In some cases, a deep neural network (DNN) is applied to introduce information of the current image 502 to guide the corresponding denoising process. In one example, referring Figure 10 , the DNN includes "GuideNet", which also includes a first encoder-decoder network 1002. The encoder part and the decoder part of the first encoder-decoder network 1002 each include five and four convolutional blocks respectively. The first encoder-decoder network 1002 is used to process the current image 502 and generate a first sequence of feature maps 1004A - 1004D that gradually shrink (e.g., based on a scaling factor) and a second sequence of feature maps 1006A - 1006C that gradually expand (e.g., based on a scaling factor). The first sequence of feature maps 1004A - 1004D is used as a skip connection and combined with the second sequence of feature maps 1006A - 1006C (e.g., through a direct addition operation) to update the second sequence of feature maps 1006A - 1006C.

[0085] Image features are extracted from the depth map 612 using a second encoder-decoder network 1008. The second encoder-decoder network 1008 is applied to process the corresponding depth map 612 (e.g., the first depth map 612) converted from the disparity map 606 and generate a third sequence of feature maps 1010A - 1010D that gradually shrink (e.g., based on a scaling factor) and a fourth sequence of feature maps 1012A - 1012C that gradually expand (e.g., based on a scaling factor). In some cases, the third sequence of feature maps 1010A - 1010D is used as a skip connection and combined with the fourth sequence of feature maps 1012A - 1012C to update the fourth sequence of feature maps 1012A - 1012C. At least one subset of the first and second sequences of feature maps 1004 and 1006 is applied to guide the operation of the second encoder-decoder network 1008 and generate the third sequence of feature maps 1010 and the fourth sequence of feature maps 1012.

[0086] In one example, the first feature map 1004B, the second feature map 1006B, or a combination thereof is concatenated with the third feature map 1010B to generate a concatenated skip connection 1014 for the decoder portion of the second encoder-decoder network 1008. In other words, at least one in the fourth feature map sequence 1012 is generated based on the skip connection 1014 that combines a subset of the first, second, and third feature map sequences (e.g., the feature maps 1004B, 1006B, and 1010B at the same level). Based on the fourth feature map sequence 1012, the electronic device generates a second depth map 612' with a noise level lower than the first depth map 612. Compared with the traditional guided image filtering (GIF) process for image denoising, the DNN method (e.g., Figure 10 GuideNet in) reduces the need for manual functions and provides an end-to-end solution. In particular, the operation of guiding depth map decoding using multi-level image features effectively reduces the noise in the depth map.

[0087] Figure 11 FIG. 1100 shows a feature extraction scheme for extracting image features from image pixels 1102 according to some embodiments. According to a first census transform, the gray level or gradient of the image pixel 1102 is compared with a first number of adjacent pixels 1104 (e.g., 16 or 32 adjacent pixels) on the current image 502 to determine a first feature value having a first number of bits (e.g., 16 or 32 bits). According to a second census transform, the gray level or gradient of the corresponding candidate pixel is compared with a first number of adjacent pixels on the reference image 504 to determine a second feature value having the first number of bits. In addition, in some embodiments, a first number of adjacent pixels 1104 are selected in a region 1106 including the image pixel 1102, and the adjacent pixels 1104 are evenly distributed in the region 1106. In the example, the first number is 32. For each pixel 1102 on the image 502 or 504', the corresponding image feature is stored in a 32-bit descriptor generated from the region 1106, which has more than the first number of pixels including the corresponding pixel 1102. Optionally, the region 1106 is square, rectangular, or diamond-shaped. Alternatively, the region 1106 may have an irregular shape. Optionally, the region 1106 is centered on the corresponding pixel 1102. Optionally, the region is not centered on the corresponding pixel 1102. For example, when the region includes 7×7 pixels, the pixel 1102 is less than 3 pixels away from the nearest edge of the current image 502. The pixel 1102 may be 2 pixels and 4 pixels away from two opposite edges of the region 1106, respectively.

[0088] See Figure 11, in one example, region 1106 has 7 rows and 7 columns, including 48 pixels surrounding the corresponding pixel 1102. The gray level of the image pixel 1102 is compared with 32 selected adjacent pixels 1104 in region 1106 to determine 32 bits in the 32-bit descriptor. The image pixel 1102 is marked as "o", 32 adjacent pixels 1104 are selected from the remaining 48 pixels and marked as "×". For each of the 32 adjacent pixels 1104, if the gray level of pixel 1102 is greater than the gray level of the corresponding adjacent pixel 1104, the corresponding bit in the 32-bit descriptor is equal to a first value (e.g., "1"), and conversely, if the gray level of pixel 1102 is equal to or less than the gray level of the corresponding adjacent pixel 1104, the corresponding bit in the 32-bit descriptor is equal to a second value (e.g., "0"). Thus, each image pixel 1102 in the current image 502 corresponds to a 32-bit image feature (i.e., 32-bit descriptor) generated from the corresponding region 1106, and each corresponding reference pixel in the reference image 504 also corresponds to a 32-bit reference feature.

[0089] Figure 12 is a flowchart of a noise filtering process 812 for reducing noise in the depth map 612 according to some embodiments. After identifying the current distance ρ of the current epipolar line 512C for the current pixel 508C of the current image 502 0 and the reference distance ρ of the reference epipolar line 512R' R of the disparity 620, the disparity 604 (Dp) is converted to the depth value 610 (D) of the current pixel 508C, and the depth value 610 (D) of each current pixel 508C is applied to create the depth map 612 of the current image 502. An oriented image filter 1202 is applied to reduce the noise level of the depth map 612 of the current image 502. In some embodiments, in semi-global matching, the uncertainty level is measured and converted to a confidence map 1204. The confidence map 1204 is multiplied by the depth map 612 (D) to create a weighted depth map 1206, and a first oriented image filter 1202A (e.g., for fast global image smoothing) is applied to generate a smoothed weighted depth map 1208. A second oriented filter 1202B is applied to the confidence map 1204 to generate a smoothed confidence map 1210. The smoothed weighted depth map 1208 is divided by the smoothed confidence map 1210 to generate a smoothed depth map 612'.

[0090] Figure 13is a flowchart of an example depth mapping method 1300 according to some embodiments. In some embodiments, the method is applied to an AR glasses 104D, a robot system, a vehicle, or a mobile phone. For convenience, the method 1300 is described as being implemented by an electronic device (e.g., the depth determination module 236 of the client device 104). An example of the client device 104 is a head-mounted display 104D or a mobile phone 104C. Optionally, the method 1300 is controlled by instructions stored in a non-transitory computer-readable storage medium and executed by one or more processors of an electronic system. Figure 13 Each operation shown in Figure 2 may correspond to instructions stored in a computer memory or a non-transitory computer-readable storage medium (e.g.,

[0091] The electronic device obtains (1302) a current image 502 associated with a current camera pose and a reference image 504 associated with a reference camera pose different from the current camera pose. A relaxation image correction operation is applied (1304) to the reference image 504 with reference to the current image 502 to generate a corrected reference image 504'. The current epipolar line 512C of the current image 502 is parallel (1306) to the reference epipolar line 512R of the corrected reference image 504'. The electronic device generates (1308) a first disparity map 606 corresponding to the corrected reference image 504' and the current image 502. The first disparity map 606 represents (1310) a first disparity 604 between: (1) a current distance ρ between each current pixel 508C and a current pole 510C in a subset of the current image 502 0 , (2) a reference distance ρ between a corresponding reference pixel 508R' and a reference pole 510R' of the corrected reference image 504'. R The electronic device determines (1312) a depth map 612 of the current image 502 based on the first disparity map 606, and the depth map 612 includes depth values 610 of each current pixel 508C in a subset of the current image 502.

[0092] In some embodiments, reference Figure 9, the electronic device generates (1314) a second disparity map 606B corresponding to the current image 502 and the corrected reference image 504'. The second disparity map 606B represents a second disparity between: (1) a reference distance between each reference pixel 508R' and a reference pole 510R' of a subset of the corrected reference image 504', and (2) a current distance between a corresponding current pixel 508C and a current pole 510C in the current image 502. A depth map 612 of the current image 502 is determined based on both the first disparity map 606A and the second disparity map 606B. Additionally, in some embodiments, the electronic device compares (1316) the first disparity map 606A and the second disparity map 606B, and modifies (1318) the first disparity map 606A according to the comparison result to generate a modified disparity map 606'( Figure 9 ). The depth map 612 of the current image 502 is determined based on the modified disparity map 606'.

[0093] Additionally, in some embodiments, for each element of the first disparity map 606 corresponding to a corresponding current pixel 508C of a subset of the current image 502, the electronic device determines a disparity difference map between a first disparity corresponding to the corresponding current pixel 508C in the first disparity map 606 and a second disparity corresponding to the corresponding current pixel 508C in the second disparity map. Then, the electronic device determines (1320) whether to modify the corresponding element of the first disparity map 606 based on the disparity difference map. In some embodiments, according to determining that the value of the corresponding element of the disparity difference map exceeds a predefined disparity threshold, the electronic device sets (1322) the corresponding element of the first disparity map 606 to 0. Conversely, according to determining (1324) that the value of the corresponding element of the disparity difference map is lower than the predefined disparity threshold, the corresponding element of the first disparity map 606 is maintained.

[0094] In some embodiments, for each current pixel 508C of a subset of the current image 502, the depth value 610 is represented by formula (4), where D prepresents a first disparity 604 corresponding to the corresponding current pixel 508C in the first disparity map 606, and A and B are two pixel transformation parameters 608. Additionally, in some embodiments, for each current pixel 508C in a subset of the current image 502, the electronic device determines two pixel transformation parameters 608 (A and B) based on the current camera pose, the reference camera pose, the camera intrinsic matrix K, and the position of the corresponding current pixel 508C. Further, in some embodiments, the lookup table 250 associates multiple pixel positions with two pixel transformation parameters 608 (A and B) based on the reference camera pose, the current camera pose, and the camera intrinsic matrix. In some embodiments, for each current pixel 508C in a subset of the current image 502, the lookup table 250 includes one or more of the following: the current distance between the current pixel 508C and the current epipole of the current image 502, the tilt angle of the epipolar line 510R' in the rectified reference image 504', the disparity range of the first disparity 604 of the current distance and the reference distance, and two pixel transformation parameters 608 (e.g., A and B).

[0095] In some embodiments, the depth map 612 includes a first depth map 612. The electronic device applies a first encoder-decoder network 1002 to process the current image 502 and generates a first sequence of feature maps 1004 (e.g., 1004A - 1004D) that gradually shrinks based on a scaling factor and a second sequence of feature maps 1006 (e.g., 1006A - 1006C) that gradually enlarges based on the scaling factor. The electronic device applies a second encoder-decoder network 1008 to process the first depth map 612 and generates a third sequence of feature maps 1010 (e.g., 1010A - 1010D) that gradually shrinks based on the scaling factor and a fourth sequence of feature maps 1012 (e.g., 1012A - 1012C) that gradually enlarges based on the scaling factor. The electronic device generates at least one in the fourth sequence of feature maps based on skip connections 1014 that combine subsets of the first, second, and third sequences of feature maps. Based on the fourth sequence of feature maps 1012 (e.g., 1012A - 1012C), the electronic device generates a second depth map 612' with a noise level lower than the first depth map 612.

[0096] In some embodiments, the position of the current epipole 510C in the current image 502 is substantially the same as the position of the reference epipole 510R in the rectified reference image 504'.

[0097] In some embodiments, each current pixel 508C of the current image 502 and the corresponding reference pixel 508R’ of the rectified reference image 504' correspond to the same object in the field of view.

[0098] In some embodiments, by determining from the disparity of the current distance and the reference distance (e.g., using Figure 6The two pixel conversion parameters A and B in 608) the inverse adjustment depth value of the current pixel 508C, and convert the inverse adjustment depth value into the depth value of the current pixel 508C to determine the depth value 610 (D) of the current pixel 508C.

[0099] In some embodiments, the electronic device selects (1326) the reference image 504 from the multiple candidate images 602 according to the reference selection criterion 252. For example, the reference selection criterion 252 requires at least one condition: the angle 516 between the two camera orientations associated with the current image 502 and the reference image 504 is less than the azimuth threshold; the current image 502 and the reference image 504 share at least a threshold number of common feature points; the error between the projected camera pose and the current camera pose is less than the pose error threshold. Based on the reference camera pose and the common feature points of the current image 502 and the reference image 504, determine the projected camera pose of the current image 502.

[0100] In addition, in some embodiments, the reference selection criterion 252 defines the azimuth threshold. For each of the subsets of the candidate images 602, the electronic device determines the angle between the current camera orientation associated with the current image 502 and the candidate camera orientation associated with the candidate image 602. According to the determination that the angle corresponding to one of the candidate images 602 is less than the azimuth threshold, the electronic device selects one of the candidate images 602 as the reference image 504. Alternatively, in some embodiments, the reference selection criterion 252 defines a range of the number of feature points. For each of the subsets of the candidate images, the electronic device determines the number of common feature points that exist in both the candidate image 602 and the current image 502. According to the determination that the number of common feature points determined based on one of the candidate images 602 is within the range of the number of feature points, the electronic device selects one of the candidate images 602 as the reference image 504.

[0101] Alternatively, in some embodiments, the reference selection criterion 252 defines the pose error threshold. The electronic device identifies a plurality of current feature points in the current image 502. For each of the subsets of the candidate images 602, the electronic device identifies a plurality of candidate feature points in the candidate image 602, estimates the candidate camera pose corresponding to the current image 502 by comparing the candidate feature points and the current feature points, determines the camera pose error between the candidate camera pose of the current image 502 and the current camera pose, and according to the determination that the camera pose error determined based on one of the candidate images 602 is less than the pose error threshold, selects one of the candidate images 602 as the reference image 504.

[0102] In some embodiments, for each current pixel 508C, the electronic device determines an epipolar line p'e' (512R') of a reference pixel 508R' on the rectified reference image 504' based on, for example, a reference camera pose, a current camera pose, an intrinsic camera matrix K, a current pixel position, and a reference pixel position. An epipolar segment is selected on the epipolar line p'e' (512R') based on a disparity range. The electronic device creates a cost volume for candidate pixels located on the epipolar segment, and each cost element indicates a Hamming distance between a first eigenvalue of the current pixel 508C in the current image 502 and a second eigenvalue of a corresponding candidate pixel on the epipolar segment in the rectified reference image 504'. The current distance ρ is determined based on the cost volume (e.g., using semi-global matching) 0 and a reference distance ρ R of the disparity, and the disparity corresponds to one of the candidate pixels, thereby identifying one of the candidate pixels as the reference pixel 508R' associated with the current pixel 508C.

[0103] In addition, in some embodiments, according to a first census transform, the electronic device compares the gray level or gradient of the current pixel 508C with a first number of adjacent pixels on the current image 502 to determine a first eigenvalue having a first number of bits (e.g., 8 bits, 16 bits, 32 bits). For each cost element, according to a second census transform, the electronic device compares the gray level or gradient of the corresponding candidate pixel with a first number of adjacent pixels on the rectified reference image 504' to determine a second eigenvalue having the first number of bits. Additionally, refer Figure 11 , in some embodiments, the electronic device selects a first number of adjacent pixels 1104 in the region 1106, where the region 1106 is centered on the current pixel 508C of the current image 502 or on each candidate pixel on the epipolar segment of the rectified reference image 504'. The adjacent pixels 1104 are evenly distributed in the region 1106.

[0104] In some embodiments, by using simultaneous localization and mapping (SLAM) to identify a plurality of sparse feature points in the current image 502, a first depth value (e.g., a minimum depth value) and a second depth value (e.g., a maximum depth value) that define a depth range of the plurality of feature points are determined, and a disparity range of the current image 502 is estimated based on the first depth value and the second depth value.

[0105] The depth mapping method 1300 unifies planar rectification and polar coordinate rectification in a unified framework. Additionally, the depth mapping method 1300 avoids using triangulation to determine a depth map 612 from a disparity 604 between a current distance of each current pixel 508C of the current image 502 and a reference distance of a corresponding reference pixel 508R' of the rectified reference image 504', thereby saving computational resources required to implement the depth mapping method 1300. Triangulation can be avoided by applying a relaxed image rectification operation to the reference image 504. The relaxed image rectification operation can be applied together with various stereo matching methods (e.g., semi-global stereo matching). Additionally, in some embodiments, the depth mapping method 1300 operates based on planar rectification, thereby providing effective results when a monocular camera moves at different positions in a plane perpendicular to the optical axis.

[0106] It should be understood that the specific order of operations described Figure 13 is merely exemplary and is not intended to indicate that the described order is the only order in which the operations can be performed. Those of ordinary skill in the art will recognize various ways of generating a depth map or depth image as described herein. Additionally, it should be noted that the details regarding Figures 1 to 12 the other processes described above also apply in a similar manner to the method 1300 described above regarding Figure 13 For the sake of brevity, these details are not repeated here.

[0107] The terms used in the description of the various embodiments described herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, it should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0108] As used herein, the term "if" is optionally interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting" or "in accordance with a determination", depending on the context. Similarly, the phrase "if determined" or "if [stated condition or event] is detected" is optionally interpreted to mean "when determining" or "in response to determining" or "when [stated condition or event] is detected" or "in response to detecting [stated condition or event]" or "in accordance with determining that [stated condition or event] is detected", depending on the context.

[0109] For illustrative purposes, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the exact forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to practice.

[0110] Although the various figures show multiple logical stages in a particular order, stages that are not order-dependent can be reordered and other stages can be combined or removed. While specific reorderings or other groupings are specifically mentioned, other orderings and groupings will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not an exhaustive list of alternatives. In addition, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.

Claims

1. A depth mapping method implemented on an electronic device having one or more processors and a memory, comprising: obtaining a current image associated with a current camera pose; obtaining a reference image associated with a reference camera pose different from the current camera pose; applying a relaxation image correction operation to the reference image with reference to the current image to generate a corrected reference image, wherein a current epipolar line of the current image is parallel to a reference epipolar line of the corrected reference image; generating a first disparity map corresponding to the current image and the corrected reference image, the first disparity map representing a first disparity between: (1) a current distance of each current pixel from a current pole in a subset of the current image, and (2) a reference distance of a corresponding reference pixel from a reference pole in the corrected reference image; and determining a depth map of the current image based on the first disparity map, the depth map including depth values of each current pixel in the subset of the current image.

2. The method according to claim 1, further comprising: generating a second disparity map corresponding to the corrected reference image, the second disparity map representing a second disparity between: (1) a reference distance of each reference pixel from a reference pole in a subset of the corrected reference image, and (2) a current distance of a corresponding current pixel from a current pole in the current image; wherein the depth map of the current image is jointly determined based on the first disparity map and the second disparity map.

3. The method according to claim 2, further comprising: comparing the first disparity map and the second disparity map; and modifying the first disparity map according to a comparison result to generate a modified disparity map, wherein the depth map of the current image is determined based on the modified disparity map.

4. The method according to claim 2, further comprising: for each corresponding element of the first disparity map corresponding to a corresponding current pixel in the subset of the current image, determining a disparity difference map between a first disparity corresponding to the corresponding current pixel in the first disparity map and a second disparity corresponding to the corresponding current pixel in the second disparity map; and determining whether to modify the corresponding element of the first disparity map based on the disparity difference map.

5. The method according to claim 4, wherein determining whether to modify the corresponding element of the first disparity map based on the disparity difference map further comprises: setting the corresponding element of the first disparity map to 0 according to determining that a value of the corresponding element in the disparity difference map exceeds a predefined disparity threshold; and maintaining the corresponding element of the first disparity map according to determining that the value of the corresponding element in the disparity difference map is lower than the predefined disparity threshold.

6. The method according to any one of the preceding claims, wherein a depth value of each current pixel in the subset of the current image is represented as: Among them, D p represents the first disparity corresponding to the corresponding current pixel in the first disparity map, and A and B are two pixel conversion parameters.

7. The method according to claim 6, further comprising: for each current pixel in the subset of the current image, Determine the two pixel conversion parameters A and B based on the current camera pose, the reference camera pose, the camera intrinsic matrix K, and the position of the corresponding current pixel.

8. The method according to claim 7, further comprises: Based on the reference camera pose, the current camera pose, and the camera intrinsic matrix, create a lookup table that associates multiple pixel positions with the two pixel conversion parameters.

9. The method according to claim 8, wherein, For each current pixel in the subset of the current image, the lookup table includes one or more of the following: the current distance between the current pixel and the current epipole of the current image, the tilt angle of the epipolar line in the rectified reference image, the parallax range of the first parallax between the current distance and the reference distance, and the two pixel conversion parameters.

10. The method according to any one of the preceding claims, wherein, Determining the first disparity map further comprises: for each current pixel in the subset of the current image, Determine the epipolar line of the reference pixel on the rectified reference image; Select an epipolar segment on the epipolar line based on the parallax range; Create a cost volume of multiple candidate pixels located on the epipolar segment, where each cost element of the cost volume indicates the Hamming distance between the first eigenvalue of the current pixel in the current image and the second eigenvalue of the corresponding candidate pixel on the epipolar segment in the rectified reference image; and Determine the first parallax between the current distance and the reference distance based on the cost volume.

11. The method according to claim 10, further comprises: For each current pixel in the subset of the current image, According to the first census transform, compare the gray level or gradient of the corresponding current pixel with the first number of adjacent pixels on the current image to determine the first eigenvalue with the first number of bits; and For each cost element, according to the second census transform, compare the gray level or gradient of the corresponding candidate pixel with the first number of adjacent pixels on the rectified reference image to determine the second eigenvalue with the first number of bits.

12. The method according to claim 11, further comprises: For each cost element, Select the first number of adjacent pixels in the area centered on the corresponding current pixel, and the number of adjacent pixels is evenly distributed in the area.

13. The method according to claim 10, further comprises: Estimate the parallax range of the current image by the following, Use simultaneous localization and mapping (SLAM) to identify multiple sparse feature points in the current image; Determine a first depth value and a second depth value, where the first depth value and the second depth value define the depth range of the multiple feature points; and Estimate the parallax range based on the first depth value and the second depth value.

14. The method according to any one of the preceding claims, the depth map includes a first depth map, the method further comprises: Process the current image using a first encoder-decoder network and generate a first sequence of feature maps that gradually shrink based on a scaling factor and a second sequence of feature maps that gradually expand based on the scaling factor; Process the first depth map using a second encoder-decoder network and generate a third sequence of feature maps that gradually shrink based on the scaling factor and a fourth sequence of feature maps that gradually expand based on the scaling factor, including: Generate at least one feature map in the fourth sequence of feature maps based on a skip connection that combines subsets of the first sequence of feature maps, the second sequence of feature maps, and the third sequence of feature maps; And Generate a second depth map based on the fourth sequence of feature maps, where the noise level of the second depth map is lower than the noise level of the first depth map.

15. The method according to any one of the preceding claims, Wherein, The position of the current pole in the current image is substantially the same as the position of the reference pole in the rectified reference image.

16. The method according to any one of the preceding claims, Wherein, Each current pixel of the current image and the corresponding reference pixel of the rectified reference image correspond to the same object in the field of view.

17. The method according to any one of the preceding claims, further comprising selecting the reference image from a plurality of candidate images according to a reference selection criterion.

18. The method according to claim 17, Wherein, The reference selection criterion requires at least one of the following conditions: The angle between two camera orientations associated with the current image and the reference image satisfies an azimuth requirement; The current image and the reference image share at least a threshold number of common feature points; And The error between the projected camera pose and the current camera pose is less than a pose error threshold, wherein the projected camera pose is determined for the current image based on the reference camera pose and the common feature points of the reference image and the current image.

19. The method according to claim 17 or 18, Wherein, The reference selection criterion defines an azimuth threshold, and selecting the reference image further comprises: For each in a subset of the plurality of candidate images, determining the angle between the current camera orientation associated with the current image and the candidate camera orientation associated with the candidate image; and Selecting one of the plurality of candidate images as the reference image according to determining that the angle corresponding to one of the plurality of candidate images satisfies the azimuth requirement.

20. The method according to any one of claims 17-19, Wherein, The reference selection criterion defines a range of the number of feature points, and selecting the reference image further comprises: For each in a subset of the plurality of candidate images, determining a plurality of common feature points that exist in both the candidate image and the current image; and Selecting one of the plurality of candidate images as the reference image according to determining that the number of common feature points determined based on one of the plurality of candidate images is within the range of the number of feature points.

21. The method according to any one of claims 17-20, Wherein, The reference selection criterion defines a pose error threshold, and selecting the reference image further includes: Identifying a plurality of current feature points in the current image; For each of the subsets of the plurality of candidate images: Identifying a plurality of candidate feature points in the candidate image; Estimating a candidate camera pose corresponding to the current image by comparing the plurality of candidate feature points and the plurality of current feature points; and Determining a camera pose error between the candidate camera pose of the current image and the current camera pose; and Selecting one of the plurality of candidate images as the reference image according to the determined camera pose error determined based on one of the plurality of candidate images being less than the pose error threshold.

22. The method according to any one of the preceding claims, further comprising determining the depth value of the current pixel by: Determining an inverse adjusted depth value of the current pixel from a disparity between the current distance and the reference distance; and Converting the inverse adjusted depth value to the depth value of the current pixel.

23. An electronic device, comprising: One or more processors; and A memory having instructions stored thereon that, when executed by the one or more processors, cause the processors to perform the method of any one of claims 1-22.

24. A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the method of any one of claims 1-22.