Image depth estimation method and device, electronic equipment and storage medium
By performing feature extraction and parallax information processing on the video data collected by the binocular camera, the accuracy problem of the binocular depth estimation method in complex scenarios is solved, and high-quality depth map generation is achieved to meet the needs of various application scenarios.
Patent Information
- Application Number
- CN202311579558.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing binocular depth estimation method is prone to pixel point matching errors in complex image acquisition scenarios, resulting in the output depth map being inaccurate enough and cannot meet the needs of depth maps in various application scenarios.
By preprocessing one frame of binocular images in the video data collected by the binocular camera, feature extraction is performed on the left and right eye images, feature similarity information is calculated, and target parallax information is determined based on the reference parallax information of the previous binocular image, and target depth information is finally obtained.
It improves the accuracy and density of depth maps, and meets the needs of depth maps in various application scenarios such as three-dimensional reconstruction, 3D content generation, and virtual reality.
Smart Images

Figure CN120047510A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image depth estimation method, apparatus, electronic device, and storage medium. Background Art
[0002] With the continuous development of computer vision technology, the application of image depth estimation is becoming more and more extensive. By performing depth estimation on an image, a corresponding depth map can be obtained, and this depth map can be used in various scenarios such as three-dimensional reconstruction, three-dimensional (3D) content generation, and mixed reality in the field of virtual reality (VR).
[0003] In related technologies, a depth estimation camera based on the time-of-flight (TOF) method or a depth estimation camera based on structured light is usually used to obtain the depth map of an image. However, these two types of depth estimation cameras rely on hardware devices with direct depth perception such as near-infrared light or structured light. These hardware devices all require a large amount of power consumption and need to be used in a specific environment, and may not be able to accurately obtain the depth map or reduce the battery life of the hardware device in an unrestricted environment.
[0004] Based on this, the binocular depth estimation method is increasingly applied to image depth estimation. This method obtains a disparity map by matching the binocular images collected by the left and right cameras, that is, for each pixel point in the left-eye image, a matching pixel point is found in the right-eye image, and then the depth map can be obtained based on the disparity map. However, the actual image acquisition scene is usually relatively complex, resulting in pixel point matching errors when matching the binocular images, making the output depth map inaccurate.
[0005] Therefore, the depth map obtained by the existing binocular depth estimation method has a poor effect and cannot meet the requirements of various application scenarios of the depth map. Summary of the Invention
[0006] Embodiments of this application provide an image depth estimation method, apparatus, electronic device, and storage medium to improve the effect of the generated depth map, so as to meet the requirements of various application scenarios of the depth map.
[0007] On the one hand, an image depth estimation method provided by an embodiment of this application includes:
[0008] Extract a frame of binocular image to be processed from the video data collected by the binocular camera, and preprocess the left-eye image and the right-eye image included in the frame of binocular image respectively;
[0009] Feature extraction is respectively performed on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left-image features and right-image features, and similarity calculations are performed on the sub-features of each first pixel point included in the left-image features and the sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information;
[0010] Based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, combined with the image features, the target disparity information of the frame of binocular image is determined; wherein, the image features are the left-image features or the right-image features, and the target disparity information represents the correspondence between each first pixel point and each second pixel point;
[0011] Based on the target disparity information of the frame of binocular image, combined with the camera parameters of the binocular camera, the target depth information of the frame of binocular image is obtained.
[0012] On the one hand, an image depth estimation device provided by an embodiment of the present application, the device includes:
[0013] A preprocessing unit, configured to extract a frame of binocular image to be processed from the video data collected by the binocular camera, and respectively preprocess the left-eye image and the right-eye image included in the frame of binocular image;
[0014] A feature extraction unit, configured to respectively perform feature extraction on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left-image features and right-image features, and perform similarity calculations on the sub-features of each first pixel point included in the left-image features and the sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information;
[0015] A disparity determination unit, configured to determine the target disparity information of the frame of binocular image based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, combined with the image features; wherein, the image features are the left-image features or the right-image features, and the target disparity information represents the correspondence between each first pixel point and each second pixel point;
[0016] A depth determination unit, configured to obtain the target depth information of the frame of binocular image based on the target disparity information of the frame of binocular image, combined with the camera parameters of the binocular camera.
[0017] Optionally, the binocular depth estimation model further includes a second feature extraction layer, and the device further includes:
[0018] A reference feature acquisition unit, configured to input the reference depth information of the obtained frame of binocular image into the second feature extraction layer to obtain reference depth features;
[0019] Then the disparity determination unit is specifically configured to:
[0020] Input the feature similarity information, the reference disparity information, and the reference depth features into the motion encoder to obtain the motion features.
[0021] Optionally, the preprocessing unit is specifically configured to:
[0022] Determine a first correspondence between the image coordinates and the world coordinates of the left-eye image based on the distortion parameters of the left-eye camera in the binocular camera, or determine a second correspondence between the image coordinates and the world coordinates of the right-eye image based on the distortion parameters of the right-eye camera in the binocular camera;
[0023] Perform stereo rectification processing on the left-eye image and the right-eye image based on the first correspondence or the second correspondence, in combination with the relative position relationship between the left-eye camera and the right-eye camera.
[0024] Optionally, when the image feature is the left image feature, the device further includes a conversion unit, configured to:
[0025] Obtain the initial disparity information of the previous frame of binocular image; the initial disparity information includes: the disparity values between each third pixel point in the left-eye image of the previous frame of binocular image and the corresponding fourth pixel point matched in the right-eye image;
[0026] Based on the motion parameters of the binocular camera switching from the frame of binocular image to the previous frame of binocular image, perform coordinate conversion on the pixel coordinates of each third pixel point in the initial disparity information, and use the corresponding disparity values of the converted third pixel points as the reference disparity information; wherein, each of the converted third pixel points corresponds to the corresponding first pixel point in the left image feature.
[0027] Optionally, the first pixel points form multiple rows of first pixel points, the second pixel points form multiple rows of second pixel points, and each row of first pixel points corresponds to a corresponding row of second pixel points;
[0028] Then the feature extraction unit is specifically configured to:
[0029] For each of the multiple first pixel points in each row of first pixel points, perform the following operations respectively: calculate the similarity between the sub-feature of a first pixel point and the sub-features of the multiple second pixel points in the corresponding row of second pixel points to obtain a similarity vector;
[0030] Based on the obtained similarity vectors corresponding to each of the multiple lines of first pixel points, obtain the feature similarity information.
[0031] Optionally, the image feature is a left-eye image feature, and the target disparity information of a frame of binocular images includes: the disparity value between each first pixel point in the left-eye image and the corresponding second pixel point;
[0032] Then the depth determination unit is specifically configured to:
[0033] For each first pixel point in the left-eye image, perform the following operations respectively:
[0034] Based on the disparity value between a first pixel point and the corresponding second pixel point, the principal point offset parameter and focal length of the left-eye camera in the binocular camera, the principal point offset parameter of the right-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of the first pixel point;
[0035] Based on the respective depth values of the first pixel points in the left-eye image, obtain the target depth information of the frame of binocular images.
[0036] Optionally, the image feature is a right-eye image feature, and the target disparity information of a frame of binocular images includes: the disparity value between each second pixel point in the right-eye image and the corresponding first pixel point;
[0037] Then the depth determination unit is specifically configured to:
[0038] For each second pixel point in the right-eye image, perform the following operations respectively:
[0039] Based on the disparity value between a second pixel point and the corresponding first pixel point, the principal point offset parameter and focal length of the right-eye camera in the binocular camera, the principal point offset parameter of the left-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of the second pixel point;
[0040] Based on the respective depth values of the second pixel points in the right-eye image, obtain the target depth information of the frame of binocular images.
[0041] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of any of the above image depth estimation methods.
[0042] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of the above image depth estimation methods.
[0043] An embodiment of the present application provides a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the above image depth estimation methods.
[0044] The embodiments of the present application have at least the following beneficial effects:
[0045] The embodiments of the present application provide an image depth estimation method, device, electronic device, and storage medium. After extracting a frame of binocular image to be processed from the video data collected by a binocular camera, preprocessing is respectively performed on the left-eye image and the right-eye image included in the frame of binocular image, and then feature extraction is respectively performed on the preprocessed left-eye image and right-eye image to obtain left image features and right image features; in order to determine the disparity between the left-eye image and the right-eye image, first, a pixel-level similarity calculation is performed on the left image features and the right image features to obtain feature similarity information, and then, based on the feature similarity information and the reference disparity information of the previous frame of binocular image, combined with the left image features or the right image features, the target disparity information of a frame of binocular image can be accurately determined, that is, the first pixel point of the left-eye image can be accurately matched with the second pixel point of the right-eye image.
[0046] Moreover, since the reference disparity information of the previous frame of binocular image is introduced, the disparity continuity between this frame of binocular image and the previous frame of binocular image can be effectively guaranteed; finally, based on the above target disparity information, combined with the camera parameters of the binocular camera, the target depth information of this frame of binocular image can be accurately obtained, so that the depth map generated based on the target depth information is relatively dense, thereby improving the effect of the depth map to meet the requirements of various application scenarios of the depth map.
[0047] Other features and advantages of the present application will be described in the subsequent specification, and part of them will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0048] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0049] Figure 1 It is a schematic diagram of an application scenario of an image depth estimation method in an embodiment of the present application;
[0050] Figure 2 It is a flowchart of an image depth estimation method in an embodiment of the present application;
[0051] Figure 3A It is a schematic diagram of a binocular image in an embodiment of the present application;
[0052] Figure 3B It is a schematic diagram of a binocular image after image stereo rectification in an embodiment of the present application;
[0053] Figure 4 It is a schematic diagram of the matching process of left image features and right image features in an embodiment of the present application;
[0054] Figure 5 It is a schematic diagram of the structure of a binocular depth estimation model in an embodiment of the present application;
[0055] Figure 6 It is a schematic diagram of the structure of another binocular depth estimation model in an embodiment of the present application;
[0056] Figure 7 It is a schematic diagram of the structure of a related pyramid layer in a binocular depth estimation model in an embodiment of the present application;
[0057] Figure 8 It is a schematic diagram of the structure of another binocular depth estimation model in an embodiment of the present application;
[0058] Figure 9 It is a schematic diagram of a depth map generated based on an image depth estimation method in an embodiment of the present application;
[0059] Figure 10 It is a logical schematic diagram of an image depth estimation method in an embodiment of the present application;
[0060] Figure 11 It is a schematic diagram of an application scenario of a depth map in an embodiment of the present application;
[0061] Figure 12 It is a schematic diagram of another application scenario of a depth map in an embodiment of the present application;
[0062] Figure 13 It is a schematic diagram of the composition structure of an image depth estimation device in an embodiment of the present application;
[0063] Figure 14 It is a schematic diagram of the composition structure of an electronic device in an embodiment of the present application;
[0064] Figure 15 It is a schematic diagram of the composition structure of another electronic device applying the embodiment of the present application. Specific embodiments
[0065] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of the present application.
[0066] The term "exemplary" as used hereinafter means "serving as an example, an embodiment or illustrative". Any embodiment described as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0067] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0068] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0069] Artificial intelligence is an interdisciplinary subject that covers a wide range of fields, including both hardware and software technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, pre-trained model technologies, operation interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0070] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0071] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition and measurement in machine vision, and further performing graphics processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping.
[0072] With the development and progress of artificial intelligence, artificial intelligence is being studied and applied in multiple fields, such as common smart homes, smart wearable devices, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, Artificial Intelligence Generated Content (AIGC), conversational interaction, intelligent healthcare, game AI, etc. It is believed that with the further development of future technologies, artificial intelligence will be applied in more fields and play an increasingly important role.
[0073] The solution provided by the embodiments of this application involves technologies such as machine learning and computer vision in artificial intelligence. A depth map of an image can be obtained through a binocular depth estimation model trained by deep learning technology, and the image can be preprocessed through image processing technology, and then image depth estimation is performed. The specific implementation is described through the following embodiments.
[0074] The design concept of the embodiments of this application is briefly introduced below.
[0075] Currently, binocular depth estimation methods are increasingly applied to image depth estimation. This method obtains a disparity map by matching binocular images collected by left and right cameras. That is, for each pixel point in the left-eye image, a matching pixel point is found in the right-eye image, and then a depth map can be obtained based on the disparity map. However, the actual image acquisition scenario is usually complex, resulting in incorrect pixel point matching when matching the collected binocular images, making the output depth map inaccurate. Therefore, the depth maps obtained by existing binocular depth estimation methods have poor effects and cannot meet the requirements of various application scenarios of depth maps.
[0076] In view of this, the embodiments of this application provide an image depth estimation method, device, electronic device, and storage medium. For a frame of binocular image, by extracting features from the preprocessed left-eye image and right-eye image, left image features and right image features are obtained. Then, based on the pixel-level feature similarity information between the left image features and the right image features, and the reference disparity information of the previous frame of binocular image, and combined with the left image features or the right image features, the target disparity information of this frame of binocular image can be accurately obtained, that is, the first pixel point of the left-eye image is accurately matched with the second pixel point of the right-eye image. And, since the reference disparity information of the previous frame of binocular image is introduced, the disparity continuity between this frame of binocular image and the previous frame of binocular image can be effectively guaranteed. Finally, based on the above target disparity information, combined with the camera parameters of the binocular camera, the target depth information of a frame of binocular image can be accurately obtained, making the depth map generated based on this target depth information relatively dense, thereby improving the effect of the depth map to meet the requirements of various application scenarios of the depth map.
[0077] The preferred embodiments of this application are described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described here are only used to illustrate and explain this application and are not used to limit this application. And without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0078] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of this application. The application scenario diagram includes a terminal device 110 and a server 120.
[0079] In the embodiments of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, etc.; a binocular camera may be installed on the terminal device. The server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0080] In an alternative embodiment, the terminal device 110 and the server 120 may communicate through a communication network. Among them, the communication network may be a wired network or a wireless network.
[0081] It should be noted that the image depth estimation method in each embodiment of the present application may be executed by an electronic device, and the electronic device may be the terminal device 110 or the server 120, that is, the method may be executed independently by the terminal device 110 or the server 120, or jointly executed by the terminal device 110 and the server 120. The following takes the server 120 executing alone as an example for illustration.
[0082] In some embodiments, after the binocular camera in the terminal device 110 captures video data, the video data is uploaded to the server 120, and the server 120 uses the image estimation method of the embodiments of the present application to process each frame of binocular image in the video data, obtains the target depth information of each frame of binocular image, and obtains a depth map based on the target depth information.
[0083] The image depth estimation method of the embodiments of the present application can be used in various application scenarios, including but not limited to 3D reconstruction, 3D content generation, mixed reality in virtual reality (VR), autonomous driving, robots, etc.
[0084] It should be noted that Figure 1 The above is only an example, and actually the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of the present application.
[0085] Next, in combination with the above-described application scenarios, the image depth estimation method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0086] Refer to Figure 2As shown in the figure, it is a flowchart of the implementation of an image depth estimation method provided by an embodiment of the present application. Taking the server as the execution entity as an example, the specific implementation process of this method includes the following steps S21 - S24:
[0087] S21: Extract a frame of binocular image to be processed from the video data collected by the binocular camera, and preprocess the left-eye image and the right-eye image included in the frame of binocular image respectively.
[0088] Among them, in the video data collected by the binocular camera, each frame of binocular image usually has distortion. By preprocessing each frame of binocular image, the distortion is eliminated and the row alignment is performed on each frame of binocular image, so that the imaging origin coordinates of the left-eye image and the right-eye image are the same, the optical axes of the binocular cameras are parallel, the left and right imaging planes are coplanar, and the epipolar lines are row-aligned, etc.
[0089] In specific implementation, according to the monocular (left-eye and right-eye) internal parameter data (focal length, imaging origin, distortion parameters, etc.) obtained after the binocular camera calibration and the relative position relationship of the binocular cameras (including rotation matrix and translation vector), each frame of binocular image can be de-distorted and row-aligned.
[0090] In some embodiments, preprocessing the left-eye image and the right-eye image included in a frame of binocular image respectively may include the following steps A1 - A2:
[0091] A1. Based on the distortion parameters of the left-eye camera in the binocular camera, determine the correspondence between the image coordinates and the world coordinates of the left-eye image, or, based on the distortion parameters of the right-eye camera in the binocular camera, determine the correspondence between the image coordinates and the world coordinates of the right-eye image.
[0092] Among them, when determining the depth map based on the left-eye image, the first correspondence between the image coordinates and the world coordinates of the left-eye image can be determined; when determining the depth map based on the right-eye image, the second correspondence between the image coordinates and the world coordinates of the right-eye image can be determined.
[0093] The distortion parameters of the left-eye camera and the right-eye camera are respectively related to their respective imaging lenses. Taking the left-eye camera as an example, the first correspondence between the image coordinates and the world coordinates is determined according to the distortion parameters of the left-eye camera, specifically as the following formulas (1) - (7):
[0094] x′ = X / Z (1)
[0095] y′ = Y / Z (2)
[0096] r = (x′) 2 +(y′) 2 (3)
[0097] x″ = x′ * (1 + K1 *r 2 +K 2 *r 4 )+2*P 1 *x′*y′+P 2 *(r 2 +2*x′2) (4)
[0098] y″ = y′*(1 + K 1 *r 2 +K 2 *r 4 )+2*P 2 *x′*y′+P 1 *(r 2 +2*y ′2 ) (5)
[0099] u = f x *x″+c x (6)
[0100] v = f y *y″+c y (7)
[0101] Wherein, (X, Y, Z) are world coordinates, (u, v) are image coordinates, K 1 , K 2 are radial distortion parameters, P 1 , P 2 are tangential distortion parameters, f x , f y are the focal length parameters of the left-eye camera, c x , c y are the offset parameters of the optical axis center.
[0102] A2. Based on the first correspondence or the second correspondence, combined with the relative position relationship between the left-eye camera and the right-eye camera, perform stereo rectification processing on the left-eye image and the right-eye image.
[0103] Wherein, the relative position relationship between the left-eye camera and the right-eye camera can be represented by the external parameters of these two cameras relative to the origin, specifically, it can be a rotation matrix and a translation vector. Specifically, using the existing rectification algorithm, based on the first correspondence or the second correspondence, combined with the external parameters of these two cameras relative to the origin, perform stereo rectification processing on the left-eye image and the right-eye image. For example, the Bouguet rectification algorithm can be used. This algorithm is a method for camera rectification. It can project the points in the image to the corresponding points on the camera plane, thereby eliminating the distortion in the left-eye image and the right-eye image, and aligning the rows of the left-eye image and the right-eye image.
[0104] Exemplarily, such as Figure 3AAs shown, the left-eye image and the right-eye image before preprocessing are shown. After preprocessing, the left-eye image and the right-eye image are subjected to distortion elimination and row alignment to obtain Figure 3B the left-eye image and the right-eye image shown.
[0105] In the embodiment of the present application, by performing stereo rectification processing on the left-eye image and the right-eye image to eliminate the distortion in these two images and align the rows of these two images, that is, to align the rows of the image planes of these two images, the purpose is that subsequent pixel point matching only needs to perform one-dimensional search on the same row of pixels, reducing the computational complexity.
[0106] S22: Respectively perform feature extraction on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left-image features and right-image features, and calculate the similarity of the respective sub-features of each first pixel point included in the left-image features and the respective sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information.
[0107] Among them, for the preprocessed left-eye image, the pixel values of each first pixel point in the left-eye image can be feature-encoded to obtain the sub-features (i.e., encoded values) of each first pixel point, constituting the left-image features; similarly, for the preprocessed right-eye image, the pixel values of each second pixel point in the right-eye image can be feature-encoded to obtain the sub-features (i.e., encoded values) of each second pixel point, constituting the right-image features.
[0108] Optionally, a trained feature extraction network can be used to respectively perform feature extraction on the preprocessed left-eye image and the preprocessed right-eye image. For example, the feature extraction network can be a convolutional neural network.
[0109] In order to determine the corresponding relationship between each first pixel point in the left-eye image and each second pixel point in the right-eye image, feature point matching can be performed on the left-image features and the right-image features. Since the preprocessed left-eye image and the preprocessed right-eye image are row-aligned, therefore, the first pixel points in one row of the left-eye image can be feature-matched with the corresponding second pixel points in one row of the right-eye image.
[0110] In some embodiments, each first pixel point in the left-eye image constitutes multiple rows of first pixel points, each second pixel point in the right-eye image constitutes multiple rows of second pixel points, and each row of first pixel points corresponds to the corresponding row of second pixel points;
[0111] In the above S22, calculating the similarity of the respective sub-features of each first pixel point included in the left-image features and the respective sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information may include the following steps B1-B2:
[0112] B1. For multiple first pixel points in each row of the first pixel points, the following operations are respectively performed: The sub-features of a first pixel point are respectively used to calculate the similarity with the sub-features of multiple second pixel points in the corresponding row of the second pixel points, obtaining a similarity vector.
[0113] Among them, the sub-features of each first pixel point can include the feature values of multiple channels, constituting a feature vector. Similarly, the sub-features of each second pixel point are also feature vectors. The similarity measurement method of two vectors can be used to calculate the similarity between the sub-features of each first pixel point and the sub-features of each second pixel point in the corresponding row. For example, the similarity measurement method includes but is not limited to cosine distance, Euclidean distance, etc.
[0114] Through similarity calculation, the similarity values between the sub-features of a first pixel point and the sub-features of multiple second pixel points are obtained, and these similarity values constitute a similarity vector.
[0115] Exemplarily, as Figure 4 shown, assuming that the left image feature is represented as matrix A, each row in matrix A is the sub-feature of multiple first pixel points, and the right image feature is represented as matrix B, each row in matrix B is the sub-feature of multiple second pixel points. For the first row (a 11 , a 12 ...a 1n ) in matrix A, each element in it is respectively used to calculate the similarity with the first row (b 11 , b 12 ...b 1n ) in matrix B, obtaining a similarity vector. By analogy, the similarity vectors corresponding to each row in matrix A are obtained.
[0116] B2. Based on the obtained similarity vectors corresponding to multiple rows of first pixel points, feature similarity information is obtained.
[0117] Specifically, the similarity vectors corresponding to multiple rows of first pixel points can form a similarity matrix. Each element in the similarity matrix represents the similarity between a first pixel point and a second pixel point, and this similarity matrix is the feature similarity information.
[0118] In the embodiments of the present application, by performing feature matching between a row of first pixel points in the left image feature and the corresponding row of second pixel points in the right image feature, pixel-level feature similarity information can be obtained, so as to facilitate subsequent determination of target disparity information based on this feature similarity information.
[0119] S23: Based on the feature similarity information and the reference disparity information of the previous binocular image of a frame of binocular image, and in combination with the image features, determine the target disparity information of a frame of binocular image; wherein, the image features are left image features or right image features, and the target disparity information represents the corresponding relationship between each first pixel point and each second pixel point.
[0120] Among them, the reference disparity information of the previous binocular image includes the corresponding relationship between each pixel point of the previous left-eye image and each pixel point of the previous right-eye image. Since a frame of binocular image is obtained after the previous binocular image is moved, that is to say, the pixel points in this frame of binocular image can be determined after the pixel points in the previous binocular image are moved. Therefore, the reference disparity information of the previous binocular image has a guiding effect on the prediction of the target disparity information of this frame of binocular image.
[0121] In the embodiments of the present application, the feature similarity information of a frame of binocular image, the reference disparity information of the previous binocular image, and the image features are fused to determine the target disparity information of this frame of binocular image. When the image features are left image features, the target disparity information may specifically include: the disparity value between each first pixel point in the left-eye image and the matched second pixel point. The disparity value of a certain first pixel point is specifically: the ordinate of the matched second pixel point on the right-eye image minus the ordinate of this first pixel point on the left-eye image; when the image features are right image features, the target disparity information may specifically include: the disparity value between each second pixel point in the right image and the matched first pixel point. The disparity value of a certain second pixel point is specifically: the ordinate of the matched first pixel point on the left-eye image minus the ordinate of this second pixel point on the right-eye image.
[0122] Specifically, taking the image features in S23 as left image features as an example, the reference disparity information of the previous binocular image includes: the disparity value between each third pixel point in the previous left-eye image and the corresponding fourth pixel point. Based on the similarity between each first pixel point in the feature similarity information and the second pixel points in the corresponding row, and in combination with the disparity values between each third pixel point and the corresponding fourth pixel point, the motion features between the left-eye image and the right-eye image of a frame of binocular image can be predicted. The motion features may include the offset value between each first pixel point and the matched second pixel point. Finally, based on the motion features and the left image features of this frame of binocular image, the target disparity information of this frame of binocular image is predicted.
[0123] When the image features in S23 are right image features, the prediction process of the target disparity information is similar to the above prediction process and will not be elaborated here.
[0124] The following embodiments illustrate the acquisition method of the reference disparity information of the previous binocular image.
[0125] In some embodiments, when the image feature in S23 is a left image feature, the reference disparity information of the previous frame of binocular images can be obtained through the following steps C1 - C2:
[0126] C1. Obtain the initial disparity information of the previous frame of binocular images; the initial disparity information includes: the disparity values between each third pixel point in the left-eye image of the previous frame of binocular images and the corresponding fourth pixel point that matches it in the right-eye image.
[0127] Specifically, the disparity value between each third pixel point and the corresponding fourth pixel point that matches it is: the ordinate of the corresponding fourth pixel point on the right-eye image minus the ordinate of the third pixel point on the left-eye image.
[0128] C2. Based on the motion parameters of the binocular camera when switching from one frame of binocular images to the previous frame of binocular images, perform coordinate transformation on the pixel coordinates of each third pixel point in the initial disparity information, and use the disparity values corresponding to each third pixel point after transformation as the reference disparity information; among them, each third pixel point after transformation corresponds to the corresponding first pixel point in the left image feature respectively.
[0129] Among them, the motion parameters of the binocular camera can be the moving distances of the binocular camera in each direction of the world coordinate system. Based on this motion parameter, the external parameter transformation matrix M of the binocular camera can be determined. First, convert the pixel coordinates of each third pixel point in the initial disparity information to the world coordinate system. Specifically, the conversion can be performed according to equations (1) - (7) in the above embodiments to obtain the world coordinates of each third pixel point; then, based on the above external parameter transformation matrix M, convert the world coordinates of each third pixel point to new world coordinates, as shown in the following equation (8):
[0130]
[0131] Among them, (X, Y, Z) are the world coordinates of the third pixel point, and (X1, Y1, Z1) are the new world coordinates of the third pixel point.
[0132] Based on the above conversion process, each third pixel point after transformation can correspond to the corresponding first pixel point in the left image feature respectively.
[0133] In the embodiments of the present application, after obtaining the initial disparity information of the previous frame of binocular images of a frame of binocular images, in order to use this initial disparity information for the disparity prediction of this frame of binocular images, perform coordinate transformation on the pixel coordinates of each third pixel point in the initial disparity information, so that each third pixel point after transformation corresponds to the corresponding first pixel point in the left image feature of this frame of binocular images respectively, so as to accurately predict the target disparity information of this frame of binocular images.
[0134] S24: Based on the target disparity information of a frame of binocular image and combining with the camera parameters of the binocular camera, obtain the target depth information of a frame of binocular image.
[0135] Among them, the target disparity information can be determined based on the left-eye image in a frame of binocular image or based on the corresponding right-eye image, that is, the image features in the above S23 are left-image features or right-image features. The following will introduce these two cases separately.
[0136] In some embodiments, when the image features in the above S23 are left-eye image features, the target disparity information of a frame of binocular image includes: the disparity values between each first pixel point in the left-eye image and the corresponding second pixel point.
[0137] In the above S24, when obtaining the depth information of a frame of binocular image based on the target disparity information of a frame of binocular image and combining with the camera parameters of the binocular camera, the following operations can be performed respectively for each first pixel point in the left-eye image:
[0138] Based on the disparity value between a first pixel point and the corresponding second pixel point, the principal point offset parameter and focal length of the left-eye camera in the binocular camera, the principal point offset parameter of the right-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of a first pixel point; based on the respective depth values of each first pixel point in the left-eye image, obtain the target depth information of a frame of binocular image.
[0139] Specifically, the respective depth values of each first pixel point in the left-eye image can be calculated based on the following formula (9):
[0140]
[0141] Among them, depth i represents the depth value of the i-th first pixel point, c x1 and c x2 are the principal point offset parameters of the left-eye camera and the right-eye camera respectively, focallength is the focal length of the left-eye camera, and baseline is the origin distance between the two cameras.
[0142] Among them, the depth value of each first pixel point can be understood as the distance between the point corresponding to the first pixel point in the actual scene and the binocular camera.
[0143] In the embodiments of the present application, when generating the depth map of the left-eye image, the target depth information corresponding to the left-eye image can be calculated based on the target disparity information corresponding to the left-eye image and combining with the camera parameters of the binocular camera, that is, obtain the target depth information of a frame of binocular image.
[0144] In some other embodiments, when the image feature is a right-eye image feature, the target disparity information of a frame of binocular images includes: the disparity values between each second pixel point in the right-eye image and the corresponding first pixel point.
[0145] In the above S24, when obtaining the target depth information of a frame of binocular images based on the target disparity information of a frame of binocular images and in combination with the camera parameters of the binocular camera, the following operations may be respectively performed for each second pixel point in the right-eye image:
[0146] Based on the disparity value between a second pixel point and the corresponding first pixel point, the principal point offset parameter and focal length of the right-eye camera in the binocular camera, the principal point offset parameter of the left-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of a second pixel point; based on the respective depth values of each second pixel point in the right-eye image, obtain the target depth information of a frame of binocular images.
[0147] Specifically, the respective depth values of each second pixel point in the right-eye image may be calculated based on the following formula (10):
[0148]
[0149] where depth i ’ represents the depth value of the i-th second pixel point, c x1 and c x2 are respectively the principal point offset parameters of the left-eye camera and the right-eye camera, focallength’ is the focal length of the right-eye camera, and baseline is the origin distance between the two cameras.
[0150] wherein, the depth value of each second pixel point may be understood as the distance between the point corresponding to the second pixel point in the actual scene and the binocular camera.
[0151] In the embodiments of the present application, when generating the depth map of the right-eye image, the target depth information corresponding to the right-eye image may be calculated based on the target disparity information corresponding to the right-eye image and in combination with the camera parameters of the binocular camera, that is, the target depth information of a frame of binocular images is obtained.
[0152] In the above embodiments of the present application, by extracting features from the preprocessed left-eye image and right-eye image, left-image features and right-image features are obtained. Then, based on the pixel-level feature similarity information between the left-image features and the right-image features, and the reference disparity information of the previous frame of binocular image, and combining the left-image features or the right-image features, the target disparity information of this frame of binocular image can be accurately obtained. Moreover, due to the introduction of the reference disparity information of the previous frame of binocular image, the disparity continuity between this frame of binocular image and the previous frame of binocular image can be effectively ensured. Finally, based on the above target disparity information and combining the camera parameters of the binocular camera, the target depth information of this frame of binocular image can be accurately obtained, making the depth map generated based on the target depth information relatively dense, thereby improving the effect of the depth map to meet the requirements of various application scenarios of the depth map.
[0153] The image depth estimation method of the embodiments of the present application can be executed by using a trained binocular depth estimation model. The structure of the binocular depth estimation model will be introduced below.
[0154] As Figure 5 and Figure 6 shown, the binocular depth estimation model includes a first feature extraction layer 501, a correlation pyramid layer 502, a disparity regression layer 503, and an upsampling layer 504 connected in sequence. Among them, the first feature extraction layer 501 can adopt a convolutional neural network, including but not limited to the MobileNet2 network, etc. The disparity regression layer 503 includes a motion encoder 5031 and a recurrent neural network 5032. The motion encoder 5031 can adopt a convolutional neural network, and specifically can include multiple convolutional operators, activation operators, and a concatenation operation Concat operator. Optionally, the specific structure of the disparity regression layer 503 can be a gated recurrent unit (GRU), and the gated recurrent unit can be used to learn and represent sequential data.
[0155] In some embodiments, in the above S22, features are respectively extracted from the preprocessed left-eye image and the preprocessed right-eye image to obtain the corresponding left-image features and right-image features, which specifically may include the following steps D1-D2:
[0156] D1. Respectively normalize the preprocessed left-eye image and the preprocessed right-eye image to obtain the normalized left-eye image and the normalized right-eye image;
[0157] Specifically, the preprocessed left-eye image and the preprocessed right-eye image can be merged and normalized. Among them, merging means splicing the pixel values of each first pixel point of the left-eye image with the pixel values of each first pixel point of the right-eye image, specifically as the following formula (10):
[0158]
[0159] Among them, input 1 and input 2 are respectively the preprocessed left-eye image and the preprocessed right-eye image. The output resolutions of the preprocessed left-eye image and the preprocessed right-eye image can be, but are not limited to, 320*256. Concat is an operation operator for merging. Input includes the normalized left-eye image and the normalized right-eye image.
[0160] D2. Input the normalized left-eye image and the normalized right-eye image into the first feature extraction layer to obtain left image features and right image features.
[0161] Among them, when the first feature extraction layer uses a convolutional neural network, by performing a convolution operation on the normalized left-eye image and the right-eye image, the pixel values of each first pixel point of the left-eye image are encoded to obtain left image features, and the pixel values of each second pixel point of the right-eye image are encoded to obtain right image features.
[0162] Specifically, taking the normalized left-eye image as an example, the convolutional neural network can also perform a pooling operation on the left-eye image and then perform a convolution operation to obtain feature maps of multiple scales, and use the feature maps of multiple scales as left image features. Each scale of image features can correspond to multiple channels, that is, in the image features at each scale, each first pixel point corresponds to feature values of multiple channels.
[0163] In the above embodiments of the present application, in order to facilitate feature extraction of the preprocessed left-eye image and right-eye image, the preprocessed left-eye image and right-eye image can be normalized, and then feature extraction is performed through the first feature extraction layer to obtain left image features and right image features.
[0164] Furthermore, the left image features and the right image features can be input into the correlation pyramid layer to calculate pixel-level feature similarity information. The calculation process of this feature similarity information can refer to the above embodiments of the present application.
[0165] Among them, when both the left image features and the right image features include feature maps of multiple scales, the correlation pyramid layer can be understood as matrices of multiple scales, specifically as Figure 7 shown. Each scale of matrix represents the similarity information between the left feature map and the right feature map at that scale. The similarity information between the left feature map and the right feature map at multiple scales constitutes the above-mentioned feature similarity information.
[0166] Specifically, each element in the matrix of each scale can be calculated by the following formula (11):
[0167] Cijk = ∑ h f ijh * g ikh (11)
[0168] where f, g ∈ R H*W*W , f and g respectively represent the left feature map and the right feature map, H and W are the width and height of the feature map, and f ijh represents the feature value of the h-th channel of the i-th row and j-th column of the left feature map, and g ikh represents the feature value of the h-th channel of the i-th row and k-th column of the right feature map. Figure 7 Each C in 2n is obtained through the pooling operation of C n , and n is an integer greater than or equal to 1.
[0169] In some embodiments, based on the feature similarity information and the reference disparity information of the previous frame of binocular image in the above S23, and combining the image features to determine the target disparity information of a frame of binocular image may include the following steps E1 - E2:
[0170] E1. Input the feature similarity information and the reference disparity information into the motion encoder to obtain the corresponding motion features; wherein, the motion features represent the offset relationship between each first pixel point and each second pixel point.
[0171] Among them, the feature similarity information includes the similarity between each first pixel point in a frame of binocular image and the second pixel points in the corresponding row, and the reference disparity information includes the disparity values of each third pixel point in the previous frame of binocular image and the corresponding fourth pixel points. The motion encoder can predict the motion features between the left-eye image and the right-eye image of a frame of binocular image based on the feature similarity information and in combination with the reference disparity information. The motion features may include the offset values of each first pixel point and the matching second pixel points.
[0172] In some possible implementation manners, as Figure 8 shown, the binocular depth estimation model may further include a second feature extraction layer 505, and the network structure of the second feature extraction layer 505 may be the same as or different from that of the first feature extraction layer. At this time, the reference depth information of a frame of binocular image obtained in advance can be input into the second feature extraction layer to obtain the reference depth features; wherein, the reference depth information can be obtained in advance through other image depth estimation methods.
[0173] In the above step E1, when inputting the feature similarity information and the reference disparity information into the motion encoder to obtain the corresponding motion features, specifically, the feature similarity information, the reference disparity information, and the reference depth features can be input into the motion encoder to obtain the motion features.
[0174] Among them, the concat operator can be used to merge the reference disparity information and the motion features first, and then the merged features and the feature similarity information are input into the motion encoder to output motion features.
[0175] In the above embodiments of the present application, when sparse reference depth information of a binocular image is obtained in advance, after feature extraction of the reference depth information, it can be input into the motion encoder of the disparity regression layer together with the feature similarity information and the reference disparity information, which can improve the accuracy of the output motion features, thereby improving the accuracy of the finally output target disparity information.
[0176] E2. Input the image features and the motion features into a recurrent neural network to obtain candidate disparity information.
[0177] Taking the image features as the left image features as an example, through a recurrent neural network, based on the sub-features of each first pixel point in the left image features, combined with the offset values of each first pixel point and the corresponding second pixel point in the motion features, the candidate disparity information corresponding to the left-eye image can be predicted.
[0178] E3. Through an upsampling layer, perform upsampling processing on the candidate disparity information to obtain the target disparity information.
[0179] Among them, taking the image features as the left image features as an example, the candidate disparity information includes the disparity values of multiple first pixel points and the corresponding second pixel points. The resolution of the disparity map corresponding to the candidate disparity information is smaller than the resolution of the initial left-eye image. By performing upsampling processing on the candidate disparity information, the target disparity information is obtained. The resolution of the disparity map corresponding to the target disparity information is the same as the resolution of the initial left-eye image, so as to generate the depth map corresponding to the left-eye image based on the target disparity information subsequently.
[0180] Finally, based on the target disparity information of a binocular image, combined with the camera parameters of the binocular camera, the target depth information of the binocular image is obtained, and the depth map is obtained according to the target depth information. As Figure 9 shown, assuming the left side is the left-eye image, then the right side is the depth map of the left-eye image.
[0181] In the embodiments of the present application, first, the motion encoder of the disparity regression layer is used to fuse the feature similarity information and the reference disparity information to obtain motion features. Then, the cyclic convolutional network is used to process the image features and the motion features to obtain candidate disparity information. Since the reference disparity information of the previous frame of binocular image is introduced, the discontinuity between the disparity of the current frame of binocular image and the disparity of the previous frame of binocular image can be effectively reduced, so as to more accurately predict the candidate disparity information of this frame of binocular image, and the number of iterations of the disparity regression layer can be reduced to a very small value, which can effectively reduce the overall computational amount of the binocular depth estimation model. Finally, through the upsampling layer, the candidate disparity information is upsampled so that the obtained target disparity information has the same resolution as the original left-eye image or right-eye image.
[0182] The image depth estimation method in the above embodiments of the present application adopts a method based on a binocular depth estimation model. Using a pure software method, it has low requirements for hardware devices and extremely low power consumption. It can support binocular images of various resolutions, and the network structure size of the binocular depth estimation model can be selected according to real-time requirements. The depth map output by the binocular depth estimation model is relatively dense, and the depth value corresponding to each pixel point can be determined. Moreover, the binocular depth estimation model can learn the image data of different scenes, so that the model has strong robustness for complex scenes such as weak texture and light changes, and thus has stronger feasibility in the real environment.
[0183] Next, in combination with Figure 10 an exemplary introduction is given to the overall implementation process of the image depth estimation method in the embodiments of the present application.
[0184] In the embodiments of the present application, after preprocessing the left-eye image and the right-eye image in a frame of binocular image, normalization processing is performed to obtain the normalized left-eye image and right-eye image as shown in Figure 10 . The normalized left-eye image and right-eye image are input into the above-mentioned binocular depth estimation model, and the target disparity information is input. Then, based on the target disparity information and combined with the camera parameters of the binocular camera, the target depth information is calculated, and based on this target depth information, the depth map corresponding to the left-eye image or the right-eye image is obtained.
[0185] The binocular depth estimation model in the above embodiments of the present application can be obtained by performing multiple rounds of iterative training on the binocular depth estimation model to be trained based on a training sample set. Each sample in the training sample set can include a sample binocular image and the corresponding sample disparity information. In addition, it may also include sparse reference depth information.
[0186] The image depth estimation method of the embodiments of the present application can be applied to various application scenarios, including but not limited to 3D reconstruction, 3D content generation, mixed reality in VR, autonomous driving, robotics, etc. Several possible application scenarios are introduced below by way of example.
[0187] In some application scenarios, the image depth estimation method can be used to implement the VR automatic safety boundary function. Specifically, using a binocular camera for image depth estimation can provide the depth of the surrounding environment for the VR automatic safety boundary function, enabling the automatic determination of the user's safe activity boundary after the user wears a VR headset. For example Figure 11 as shown.
[0188] In some other application scenarios, neural path through in MR (mixed reality) is based on the color perspective function of the VST camera. Its main function is to obtain a depth map by using the image depth estimation method for the images output by two cameras. The images converted by each of the two cameras can use the images of the other camera and reference the depth map for pixel point projection completion.
[0189] In still some other application scenarios, the depth map can be used for 3D content generation. Specifically, the depth map can be used as auxiliary information for characters, props, or scenes. By inputting the depth map into a generative artificial intelligence (AIGC) model, 3D assets of different styles can be generated. For example, to generate a scene, the depth map and descriptive text information can be input into the AIGC model to generate the required scene. In this way, users can freely customize the required 3D assets, and also greatly reduce the cost of 3D production.
[0190] For example, in a game scene, it is usually necessary to create characters, props, scenes, etc. in the game through 3D modeling. At this time, the image depth estimation method of the embodiments of the present application can be applied to obtain the depth map of the modeling image, and then based on the depth map and the corresponding descriptive text information, through the above AIGC model, generate the characters, props, scenes, etc. required in the game, improving the efficiency of game production and reducing the cost of game production.
[0191] Exemplarily, as Figure 12 shown, it is a panoramic depth map of a room input into the AIGC model and the related text description "white wall bedroom, beige curtains, patterned bedsheet, yellow headboard wall". On the right is the generated corresponding panoramic RGB (Red Green Blue, primary colors) map, which can then be converted into a 3D scene map.
[0192] In addition, the image depth estimation method can also be applied to the field of autonomous driving or robots. By determining the depth information of objects in binocular images, it can further guide object detection.
[0193] The accuracy error of the binocular depth estimation model in the embodiments of this application is relatively small. The depth map is denser than that obtained by traditional hardware sensors and can be configured with different resolutions according to different computing power requirements. The method of generating depth using algorithm software reduces energy consumption by more than 90% compared with using depth sensor hardware. In existing VR safety boundary functions, it is basically necessary to manually intervene to draw the boundary contour. However, in this method, when used in the safety boundary function, it can automatically delimit an effective safety boundary without manual intervention. In 3D content generation, the 3D configurations corresponding to the environment and objects are reconstructed based on the generated depth map and then used in 3D creation of AIGC, which is simpler and more efficient than the traditional 3D modeling and texturing process.
[0194] Based on the same inventive concept as the method embodiments of this application, an image depth estimation device is also provided in the embodiments of this application. The principle of the device for solving problems is similar to that of the method in the above embodiments. Therefore, the implementation of the device can refer to the implementation of the above method, and the repeated parts will not be elaborated.
[0195] Refer to Figure 13 As shown, the image depth estimation device of the embodiments of this application may include:
[0196] A preprocessing unit 1301, configured to extract a frame of binocular image to be processed from the video data collected by the binocular camera, and perform preprocessing on the left-eye image and the right-eye image included in the frame of binocular image respectively;
[0197] A feature extraction unit 1302, configured to perform feature extraction on the preprocessed left-eye image and the preprocessed right-eye image respectively to obtain corresponding left-image features and right-image features, and calculate the similarity of the respective sub-features of each first pixel point included in the left-image features and the respective sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information;
[0198] A disparity determination unit 1303, configured to determine the target disparity information of a frame of binocular image based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, in combination with the image features; wherein, the image features are left-image features or right-image features, and the target disparity information represents the corresponding relationship between each first pixel point and each second pixel point;
[0199] A depth determination unit 1304, configured to obtain the target depth information of a frame of binocular image based on the target disparity information of the frame of binocular image, in combination with the camera parameters of the binocular camera.
[0200] In the embodiments of the present application, by performing feature extraction on the preprocessed left-eye image and right-eye image, left image features and right image features are obtained. Then, based on the pixel-level feature similarity information between the left image features and the right image features, and the reference disparity information of the previous frame of binocular image, and in combination with the left image features or the right image features, the target disparity information of this frame of binocular image can be accurately obtained. Moreover, due to the introduction of the reference disparity information of the previous frame of binocular image, the disparity continuity between this frame of binocular image and the previous frame of binocular image can be effectively ensured. Finally, based on the above target disparity information, in combination with the camera parameters of the binocular camera, the target depth information of this frame of binocular image can be accurately obtained, making the depth map generated based on this target depth information relatively dense, thereby improving the effect of the depth map to meet the requirements of various application scenarios of the depth map.
[0201] Optionally, the device is executed using a trained binocular depth estimation model, and the binocular depth estimation model at least includes a disparity regression layer and an upsampling layer. The disparity regression layer includes a motion encoder and a recurrent neural network.
[0202] Then, the disparity determination unit 1303 is specifically configured to:
[0203] Input the feature similarity information and the reference disparity information into the motion encoder to obtain corresponding motion features, where the motion features represent the offset relationship between each first pixel point and each second pixel point.
[0204] Input the image features and the motion features into the recurrent neural network to obtain initial disparity information.
[0205] Perform upsampling processing on the candidate disparity information through the upsampling layer to obtain the target disparity information.
[0206] Optionally, the binocular depth estimation model further includes a first feature extraction layer.
[0207] Then, the feature extraction unit 1302 is specifically configured to:
[0208] Normalize the preprocessed left-eye image and the preprocessed right-eye image respectively to obtain the normalized left-eye image and the normalized right-eye image.
[0209] Input the normalized left-eye image and the normalized right-eye image into the first feature extraction layer to obtain left image features and right image features.
[0210] Optionally, the binocular depth estimation model further includes a second feature extraction layer, and the device further includes:
[0211] A reference feature acquisition unit, configured to input the reference depth information of a frame of binocular image obtained in advance into the second feature extraction layer to obtain reference depth features.
[0212] Specifically, the parallax determination unit 1303 is configured to:
[0213] Input the feature similarity information, the reference parallax information, and the reference depth feature into a motion encoder to obtain a motion feature.
[0214] Optionally, the preprocessing unit 1301 is specifically configured to:
[0215] Determine a first correspondence between the image coordinates and the world coordinates of the left-eye image based on the distortion parameters of the left-eye camera in the binocular camera, or determine a second correspondence between the image coordinates and the world coordinates of the right-eye image based on the distortion parameters of the right-eye camera in the binocular camera;
[0216] Perform stereo rectification processing on the left-eye image and the right-eye image based on the first correspondence or the second correspondence, in combination with the relative position relationship between the left-eye camera and the right-eye camera.
[0217] Optionally, when the image feature is the left image feature, the apparatus further includes a conversion unit configured to:
[0218] Obtain initial parallax information of the previous frame of binocular image; the initial parallax information includes: the parallax values of each third pixel point in the left-eye image of the previous frame of binocular image and the corresponding fourth pixel point matched in the right-eye image;
[0219] Based on the motion parameters of the binocular camera switching from the current frame of binocular image to the previous frame of binocular image, perform coordinate conversion on the pixel coordinates of each third pixel point in the initial parallax information, and use the parallax values corresponding to the converted third pixel points as the reference parallax information; wherein, each of the converted third pixel points corresponds to a corresponding first pixel point in the left image feature.
[0220] Optionally, the first pixel points form multiple rows of first pixel points, the second pixel points form multiple rows of second pixel points, and each row of first pixel points corresponds to a corresponding row of second pixel points;
[0221] Specifically, the feature extraction unit 1302 is configured to:
[0222] For each of the multiple first pixel points in each row of first pixel points, perform the following operations respectively: calculate the similarity between the sub-feature of a first pixel point and the sub-features of multiple second pixel points in the corresponding row of second pixel points to obtain a similarity vector;
[0223] Obtain feature similarity information based on the similarity vectors corresponding to each row of first pixel points obtained.
[0224] Optionally, the image feature is the left-eye image feature, and the target disparity information of a frame of binocular image includes: the disparity values between each first pixel point in the left-eye image and the corresponding second pixel point;
[0225] Then the depth determination unit 1304 is specifically configured to:
[0226] For each first pixel point in the left-eye image, perform the following operations respectively:
[0227] Based on the disparity value between a first pixel point and the corresponding second pixel point, the principal point offset parameter and focal length of the left-eye camera in the binocular camera, the principal point offset parameter of the right-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of a first pixel point;
[0228] Based on the respective depth values of the first pixel points in the left-eye image, obtain the target depth information of a frame of binocular image.
[0229] Optionally, the image feature is the right-eye image feature, and the target disparity information of a frame of binocular image includes: the disparity values between each second pixel point in the right-eye image and the corresponding first pixel point;
[0230] Then the depth determination unit 1304 is specifically configured to:
[0231] For each second pixel point in the right-eye image, perform the following operations respectively:
[0232] Based on the disparity value between a second pixel point and the corresponding first pixel point, the principal point offset parameter and focal length of the right-eye camera in the binocular camera, the principal point offset parameter of the left-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera, determine the depth value of a second pixel point;
[0233] Based on the respective depth values of the second pixel points in the right-eye image, obtain the target depth information of a frame of binocular image.
[0234] For the sake of convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present application, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.
[0235] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0236] After introducing the image depth estimation method and apparatus according to the exemplary embodiments of the present application, next, an electronic device according to another exemplary embodiment of the present application will be introduced.
[0237] Based on the same inventive concept as the above method embodiments, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device may be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device may be as Figure 14 shown, including a memory 1401, a communication module 1403, and one or more processors 1402.
[0238] The memory 1401 is used to store the computer program executed by the processor 1402. The memory 1401 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0239] The memory 1401 may be a volatile memory, such as a random-access memory (RAM); the memory 1401 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1401 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1401 may be a combination of the above memories.
[0240] The processor 1402 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1402 is used to implement the above image depth estimation method when calling the computer program stored in the memory 1401.
[0241] The communication module 1403 is used to communicate with the terminal device and other servers.
[0242] In the embodiments of the present application, the specific connection medium between the above memory 1401, communication module 1403, and processor 1402 is not limited. In the embodiments of the present application Figure 14 it is shown that the memory 1401 and the processor 1402 are connected through a bus 1404, and the bus 1404 is inFigure 14 is described by a thick line. The connection manners between other components are only for illustrative purposes and are not restrictive. The bus 1404 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 14 is only described by a thick line, but it does not describe that there is only one bus or one type of bus.
[0243] The computer storage medium is stored in the memory 1401, and computer-executable instructions are stored in the computer storage medium. The computer-executable instructions are used to implement the image depth estimation method of the embodiments of the present application. The processor 1402 is used to execute the above-mentioned image depth estimation method, as Figure 2 shown.
[0244] In another embodiment, the electronic device can also be other electronic devices, such as Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device can be as Figure 15 shown, including components such as a communication component 1510, a memory 1520, a display unit 1530, a camera 1540, a sensor 1550, an audio circuit 1560, a Bluetooth module 1570, and a processor 1580.
[0245] The communication component 1510 is used to communicate with the server. In some embodiments, it can include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.
[0246] The memory 1520 can be used to store software programs and data. The processor 1580 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1520. The memory 1520 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1520 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 1520 can store the operating system and various application programs, and can also store a computer program for executing the image depth estimation method of the embodiments of the present application.
[0247] The display unit 1530 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1530 may include a display screen 1532 disposed on the front of the terminal device 110. Among them, the display screen 1532 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1530 can be used to display images and the like in the embodiments of the present application.
[0248] The display unit 1530 can also be used to receive input numerical or character information and generate signal inputs related to the user settings and function control of the terminal device 110. Specifically, the display unit 1530 may include a touch screen 1531 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.
[0249] Among them, the touch screen 1531 can cover the display screen 1532, or the touch screen 1531 and the display screen 1532 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply called a touch display screen. In the present application, the display unit 1530 can display binocular images and corresponding depth maps.
[0250] The camera 1540 can be used to capture static images, and the user can publish the images captured by the camera 1540 through an application. The camera 1540 can be one or multiple. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1580 to convert it into a digital image signal.
[0251] The terminal device may further include at least one sensor 1550, such as an acceleration sensor 1551, a distance sensor 1552, a fingerprint sensor 1553, a temperature sensor 1554. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, a motion sensor, etc.
[0252] The audio circuit 1560, speaker 1561, and microphone 1562 can provide an audio interface between the user and the terminal device 110. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output. The terminal device 110 can also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560, converted into audio data, and then the audio data is output to the communication component 1510 for sending to, for example, another terminal device 110, or the audio data is output to the memory 1520 for further processing.
[0253] The Bluetooth module 1570 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1570, so as to perform data interaction.
[0254] The processor 1580 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 1520 and calling data stored in the memory 1520, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1580 may include one or more processing units; the processor 1580 can also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1580. In this application, the processor 1580 can run the operating system, application programs, user interface display, and touch response, as well as the image depth estimation method of the embodiments of this application. In addition, the processor 1580 is coupled to the display unit 1530.
[0255] In some possible implementation manners, various aspects of the image depth estimation method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the image depth estimation method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute the steps as shown in Figure 2 shown.
[0256] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0257] The program product of an embodiment of the present application may employ a portable compact disk read-only memory (CD-ROM) and include a computer program, and may be run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0258] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0259] The computer program contained on the readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0260] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The computer program may be executed entirely on the user's electronic device, partially on the user's electronic device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on the remote electronic device or server. In the case of a remote electronic device, the remote electronic device may be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external electronic device (e.g., through the Internet using an Internet service provider).
[0261] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0262] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0263] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.
[0264] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0265] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0266] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for realizing the functions specified in one block or a plurality of blocks.
[0267] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present application.
[0268] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. An image depth estimation method, characterized in that, the method comprises: extracting a frame of binocular image to be processed from the video data collected by a binocular camera, and respectively preprocessing the left-eye image and the right-eye image included in the frame of binocular image; respectively performing feature extraction on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left-image features and right-image features, and calculating the similarity of the respective sub-features of each first pixel point included in the left-image features and the respective sub-features of each second pixel point included in the right-image features to obtain corresponding feature similarity information; based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, and combining image features, determining the target disparity information of the frame of binocular image; wherein, the image features are the left-image features or the right-image features, and the target disparity information represents the corresponding relationship between each first pixel point and each second pixel point; based on the target disparity information of the frame of binocular image, and combining the camera parameters of the binocular camera, obtaining the target depth information of the frame of binocular image.
2. The method according to claim 1, characterized in that, the method is executed by using a trained binocular depth estimation model, the binocular depth estimation model at least includes a disparity regression layer and an upsampling layer, and the disparity regression layer includes a motion encoder and a recurrent neural network; then the determining the target disparity information of the frame of binocular image based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, and combining image features, includes: inputting the feature similarity information and the reference disparity information into the motion encoder to obtain corresponding motion features; wherein, the motion features represent the offset relationship between each first pixel point and each second pixel point; inputting the image features and the motion features into the recurrent neural network to obtain candidate disparity information; performing upsampling processing on the candidate disparity information through the upsampling layer to obtain the target disparity information.
3. The method according to claim 2, characterized in that, the binocular depth estimation model further includes a first feature extraction layer; then the respectively performing feature extraction on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left-image features and right-image features, includes: respectively normalizing the preprocessed left-eye image and the preprocessed right-eye image to obtain a normalized left-eye image and a normalized right-eye image; inputting the normalized left-eye image and the normalized right-eye image into the first feature extraction layer to obtain the left-image features and the right-image features.
4. The method according to claim 2, characterized in that, the binocular depth estimation model further includes a second feature extraction layer, and the method further includes: inputting the previously obtained reference depth information of the frame of binocular image into the second feature extraction layer to obtain reference depth features; Then inputting the feature similarity information and the reference disparity information into the motion encoder to obtain corresponding motion features includes: Inputting the feature similarity information, the reference disparity information, and the reference depth feature into the motion encoder to obtain the motion features.
5. The method according to any one of claims 1-4, wherein, the preprocessing of the left-eye image and the right-eye image included in the frame of binocular images respectively includes: determining a first correspondence between the image coordinates and the world coordinates of the left-eye image based on the distortion parameters of the left-eye camera in the binocular camera, or determining a second correspondence between the image coordinates and the world coordinates of the right-eye image based on the distortion parameters of the right-eye camera in the binocular camera; performing stereo rectification processing on the left-eye image and the right-eye image based on the first correspondence or the second correspondence in combination with the relative position relationship between the left-eye camera and the right-eye camera.
6. The method according to claim 4, wherein, when the image feature is the left-image feature, the reference disparity information of the previous frame of binocular images of the frame of binocular images is obtained by the following method: obtaining the initial disparity information of the previous frame of binocular images; the initial disparity information includes: the disparity values of each third pixel point in the left-eye image of the previous frame of binocular images and the corresponding fourth pixel points matched in the right-eye image; based on the motion parameters of the binocular camera switching from the frame of binocular images to the previous frame of binocular images, performing coordinate conversion on the pixel coordinates of each third pixel point in the initial disparity information, and using the disparity values corresponding to the converted third pixel points as the reference disparity information; wherein, each of the converted third pixel points corresponds to the corresponding first pixel point in the left-image feature.
7. The method according to any one of claims 1-4, wherein, the first pixel points form multiple rows of first pixel points, the second pixel points form multiple rows of second pixel points, and each row of first pixel points corresponds to the corresponding row of second pixel points; then calculating the similarity between the sub-features of each first pixel point included in the left-image feature and the sub-features of each second pixel point included in the right-image feature to obtain corresponding feature similarity information includes: performing the following operations respectively for multiple first pixel points in each row of first pixel points: calculating the similarity between the sub-feature of a first pixel point and the sub-features of multiple second pixel points in the corresponding row of second pixel points to obtain a similarity vector; obtaining the feature similarity information based on the similarity vectors corresponding to the multiple rows of first pixel points obtained.
8. The method according to any one of claims 1-4, wherein, the image feature is the left-eye image feature, and the target disparity information of the frame of binocular images includes: the disparity values of each first pixel point in the left-eye image and the corresponding second pixel points. Then, obtaining the depth information of the frame of binocular image based on the target disparity information of the frame of binocular image and combining the camera parameters of the binocular camera includes: Performing the following operations respectively for each first pixel point in the left-eye image: Determining the depth value of a first pixel point based on the disparity value between the first pixel point and the corresponding second pixel point, the principal point offset parameter and focal length of the left-eye camera in the binocular camera, the principal point offset parameter of the right-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera; Obtaining the target depth information of the frame of binocular image based on the depth values of the respective first pixel points in the left-eye image.
9. The method according to any one of claims 1-4, wherein, the image feature is a right-eye image feature, and the target disparity information of the frame of binocular image includes: the disparity value between each second pixel point in the right-eye image and the corresponding first pixel point; then, obtaining the target depth information of the frame of binocular image based on the target disparity information of the frame of binocular image and combining the camera parameters of the binocular camera includes: Performing the following operations respectively for each second pixel point in the right-eye image: Determining the depth value of a second pixel point based on the disparity value between the second pixel point and the corresponding first pixel point, the principal point offset parameter and focal length of the right-eye camera in the binocular camera, the principal point offset parameter of the left-eye camera in the binocular camera, and the origin distance between the left-eye camera and the right-eye camera; Obtaining the target depth information of the frame of binocular image based on the depth values of the respective second pixel points in the right-eye image.
10. An image depth estimation device, wherein, the device includes: A preprocessing unit, configured to extract a frame of binocular image to be processed from the video data collected by the binocular camera, and respectively perform preprocessing on the left-eye image and the right-eye image included in the frame of binocular image; A feature extraction unit, configured to respectively perform feature extraction on the preprocessed left-eye image and the preprocessed right-eye image to obtain corresponding left image features and right image features, and calculate the similarity of the respective sub-features of each first pixel point included in the left image features and the respective sub-features of each second pixel point included in the right image features to obtain corresponding feature similarity information; A disparity determination unit, configured to determine the target disparity information of the frame of binocular image based on the feature similarity information and the reference disparity information of the previous frame of binocular image of the frame of binocular image, and in combination with the image feature; wherein, the image feature is the left image feature or the right image feature, and the target disparity information represents the corresponding relationship between the respective first pixel points and the respective second pixel points; A depth determination unit, configured to obtain the target depth information of the frame of binocular image based on the target disparity information of the frame of binocular image and in combination with the camera parameters of the binocular camera.
11. The device according to claim 10, wherein, The described device is executed using a trained binocular depth estimation model, which at least includes a disparity regression layer and an upsampling layer. The disparity regression layer includes a motion encoder and a recurrent neural network; Then the disparity determination unit is specifically configured to: Input the feature similarity information and the reference disparity information into the motion encoder to obtain corresponding motion features; wherein, the motion features represent the offset relationship between each first pixel point and each second pixel point; Input the image features and the motion features into the recurrent neural network to obtain candidate disparity information; Perform upsampling processing on the candidate disparity information through the upsampling layer to obtain the target disparity information.
12. The device according to claim 11, wherein, the binocular depth estimation model further includes a first feature extraction layer; Then the feature extraction unit is specifically configured to: Perform normalization processing on the preprocessed left-eye image and the preprocessed right-eye image respectively to obtain a normalized left-eye image and a normalized right-eye image; Input the normalized left-eye image and the normalized right-eye image into the first feature extraction layer to obtain the left image features and the right image features.
13. An electronic device, wherein, it includes a processor and a memory. Among them, the memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of any one of the methods described in claims 1 to 9.
14. A computer-readable storage medium, wherein, it includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of the methods described in claims 1 to 9.
15. A computer program product, wherein, it includes a computer program, and the computer program is stored in a computer-readable storage medium; when the processor of the electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the methods described in claims 1 to 9.
Citation Information
Cited By
Robot control method and device, computer equipment and storage medium
CN120627899A
Depth estimation method and device, computer equipment, storage medium and program product
CN122312729A
Unmanned aerial vehicle sensing method and device and vehicle
CN122391680A