Depth estimation method and apparatus, electronic device, and storage medium
By using depth estimation methods and depth scaling factors, the problem of unknown spacing between binocular stereo cameras or the inability of monocular cameras to calculate depth was solved, enabling monocular cameras to accurately calculate depth values without relying on spacing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2026-03-24
AI Technical Summary
In existing technologies, there are problems with the unknown spacing of binocular stereo cameras or the inability to accurately calculate depth values when using monocular cameras.
A depth estimation method is adopted, which acquires images from a monocular camera and inputs them into a depth estimation model. The depth information of the depth image is calculated by combining the depth scaling factor and distance information of the device.
Even without knowing the distance between the stereo cameras in advance, a monocular camera can accurately calculate the depth value.
Smart Images

Figure CN117218175B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically to a depth estimation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, electronic devices can acquire image information of objects using a stereo camera (i.e., two cameras, such as camera 1 and camera 2). Then, they can identify the same features in the image information acquired by camera 1 and camera 2, calculate the parallax of the same features by camera 1 and camera 2, and then calculate the depth of the point (i.e., the vertical distance from the point to the line connecting the two cameras) based on the parallax and the distance between camera 1 and camera 2.
[0003] However, without the spacing between camera 1 and camera 2, or when using a monocular camera (i.e., one camera), it is impossible to calculate the depth of the target object. Therefore, a solution is urgently needed to address this problem. Summary of the Invention
[0004] In view of the above, it is necessary to propose a depth estimation method, device, electronic device and storage medium that can obtain the actual depth value without knowing the distance between the left and right cameras on a stereo camera in advance, and can also accurately know the actual depth value when using a monocular camera.
[0005] The depth estimation method includes: acquiring a first image, inputting the first image into a depth estimation model to obtain a first depth image, acquiring a depth scaling factor, wherein the depth scaling factor is used to indicate the relationship between the relative depth of a pixel in the first depth image and the depth value of the pixel, and calculating the depth information of the first depth image based on the depth scaling factor and the first depth image.
[0006] Compared with the prior art, the depth estimation method, device, electronic device and storage medium provided by the present invention can obtain the actual depth value without knowing the distance between the left and right cameras on a stereo camera in advance, and can also accurately know the actual depth value when using a monocular camera. Attached Figure Description
[0007] Figure 1 This is a schematic diagram illustrating an application scenario of the depth estimation method provided in the embodiments of this application.
[0008] Figure 2 This is a schematic diagram of a depth estimation method provided in an embodiment of this application.
[0009] Figure 3This is a schematic diagram of the training method for the depth estimation model provided in the embodiments of this application.
[0010] Figure 4 This is a flowchart illustrating the method for obtaining the depth scaling factor provided in an embodiment of this application.
[0011] Figure 5 This is a schematic diagram of a depth estimation device provided in an embodiment of this application.
[0012] Figure 6 A schematic diagram of the structure of the electronic device provided in the application embodiment.
[0013] Explanation of main component symbols
[0014] vehicle 100、120、130 windshield 10 Depth estimation system 20 camera equipment 201 Distance acquisition device 202 processor 203 Horizontal coverage area 110、140 Depth estimation device 51 First Image Acquisition Module 511 Input module 512 Get Module 513 Depth information acquisition module 514 electronic devices 60 memory 61 processor 62
[0015] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] Before providing a detailed explanation of the embodiments of this application, the application scenarios involved in the embodiments of this application will be introduced first.
[0018] Depth estimation of images is an indispensable technique in the field of computer vision, with applications in autonomous driving, scene understanding, robotics, 3D reconstruction, photography, intelligent medicine, intelligent human-computer interaction, spatial mapping, and augmented reality. For example, in autonomous driving, depth information from images can be used to assist in sensor fusion, drivable space exploration, and navigation.
[0019] The following description uses the depth estimation method provided in the embodiments of this application to illustrate its application in an autonomous driving scenario. It should be understood that the depth estimation method provided in the embodiments of this application is not limited to application in autonomous driving scenarios.
[0020] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the depth estimation method provided in the embodiments of this application.
[0021] like Figure 1As shown, the vehicle 100 includes a depth estimation system 20 disposed in an interior compartment behind the windshield 10 of the vehicle 100. The depth estimation system 20 includes a camera device 201, a distance acquisition device 202, and a processor 203. The processor 203 is connected to the camera device 201 and the distance acquisition device 202.
[0022] It is understood that the camera device 201, the distance acquisition device 202, and the processor 203 can be installed in other locations on the vehicle 100, so that the camera device 201 can acquire images of the area in front of the vehicle 100, and the distance acquisition device 202 can detect the distance to objects in front of the vehicle 100. For example, the camera device 201 and the distance acquisition device 202 can be located in the metal grille or front bumper of the vehicle 100. Furthermore, although Figure 1 Only one distance acquisition device 202 is shown, but the vehicle 100 may have multiple distance acquisition devices 202 pointing in different directions (side, front, rear, etc.). Each distance acquisition device 202 may be installed on the windshield, door panel, bumper, or metal grille, etc.
[0023] In this embodiment of the application, the camera device 201 on the vehicle 100 can acquire images of the scenes in front of and to the sides of the vehicle 100. For example... Figure 1 As shown, within the horizontal coverage area 110 (shown by dashed lines) detectable by camera device 201, there are two objects, vehicle 120 and vehicle 130. Camera device 201 can be used to capture images based on the light waves seen through the windshield in the horizontal coverage area 110 and captured, that is, camera device 201 can capture images of vehicle 120 and vehicle 130 in front of vehicle 100.
[0024] In some embodiments, the camera device 201 may be a stereo camera or a monocular camera.
[0025] In some embodiments, the camera device 201 can be implemented as a dashcam. A dashcam is an instrument that records images and sounds, as well as other relevant information, during the driving of a vehicle 100. Once a dashcam is installed on the vehicle 100, it can record images and sounds throughout the entire driving process, one of its core uses being to provide valid evidence in case of traffic accidents. As an example, in addition to the functions described above, the dashcam may also provide functions such as Global Positioning System (GPS) positioning, driving trajectory capture, remote monitoring, electronic dog (speed camera detector), and navigation; this application embodiment does not specifically limit these functions.
[0026] The distance acquisition device 202 can be used to detect objects in front of and to the sides of the vehicle 100 to obtain the distance between the object and the distance acquisition device 202. For example... Figure 1 As shown, within the horizontal coverage area 140 (shown by dashed lines) detectable by the distance acquisition device 202, there are two objects, vehicle 120 and vehicle 130. That is, the distance acquisition device 202 on vehicle 100 can acquire the distance between vehicle 120 and the distance acquisition device 202, as well as the distance between vehicle 120 and the distance acquisition device 202. The distance acquisition device 202 can be an infrared sensor, radar, etc.
[0027] Taking distance acquisition device 202 as an example of radar, radar uses radio frequency (RF) waves to determine the distance, direction, speed, and / or height of objects in front of a vehicle. More specifically, radar includes a transmitter and a receiver. The transmitter emits pulses of RF waves (radar signals), which bounce off any object in their path. The pulses reflected back by the object return a small portion of the RF wave's energy to the receiver, which is typically located at the same location as the transmitter. Figure 1 As shown, the radar is configured to transmit radar signals through the windshield in the horizontal coverage area 140 and receive reflected radar signals reflected by any object within the horizontal coverage area 140, thus obtaining a three-dimensional point cloud image of any object within the horizontal coverage area 140.
[0028] In this embodiment, the horizontal coverage area 110 and the horizontal coverage area 140 may completely overlap, or there may be an overlapping area between the horizontal coverage area 110 and the horizontal coverage area 140 (i.e., Figure 1 The horizontal coverage area 140 can be set to reach a certain threshold so that the camera device 201 and the radar can capture their respective versions of the same scene after the direction is determined.
[0029] In some embodiments, camera device 201 may capture images of a scene within its horizontal coverage area 110 at a certain periodic rate. Similarly, radar may capture three-dimensional point cloud images of the scene within its horizontal coverage area 140 at a certain periodic rate. The periodic rates at which camera device 201 and radar capture their respective frames may be the same or different. The images and three-dimensional point cloud images captured by each camera device 201 may be timestamped. Therefore, where the periodic rates differ, the timestamps can be used to simultaneously or nearly simultaneously select the captured images and three-dimensional point cloud images for further processing (e.g., fusion).
[0030] Among them, the three-dimensional point cloud, also known as laser point cloud (PCD) or point cloud, can be a collection of massive points that represent the spatial distribution and surface characteristics of the target by using laser to acquire the three-dimensional spatial coordinates of each sampling point on the surface of an object in the same spatial reference frame. Compared with images, although the three-dimensional point cloud lacks detailed texture information, it contains rich three-dimensional spatial information, that is, the distance between the object and the distance acquisition device 202.
[0031] For example, such as Figure 1 As shown, at time T0, camera device 201 can acquire images of vehicles 120 and 130. At the same time (time T0), distance acquisition device 202 can also acquire three-dimensional point cloud images within the horizontal coverage area 140, that is, at time T0, it acquires the distance between vehicle 100 and distance acquisition device 202, as well as the distance between vehicle 120 and distance acquisition device 202.
[0032] In this embodiment, the processor 203 may be a general-purpose processor 203, including a central processing unit (CPU), a network processor (NP), etc. The processor 203 may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0033] In this embodiment, the processor 203 (e.g., a digital signal processor 203 (DSP)) analyzes the images of the same scene captured by the camera device 201 (i.e., including the same objects such as vehicle 120 and vehicle 130) at the same time and the distance information collected by the distance acquisition device 202 that captures the same scene (i.e., including the same objects such as vehicle 120 and vehicle 130), in order to identify the depth information of objects within the captured scene. These objects can be other vehicles, pedestrians, road signs, objects on the road, faces, etc.
[0034] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the depth estimation system. In other embodiments of this application, the depth estimation system may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0035] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a depth estimation method provided in an embodiment of this application. This depth estimation method can be applied to… Figure 1 The depth estimation system shown can be derived from... Figure 1 The processor 203 shown below executes the commands, as explained in detail below.
[0036] Step S10: Obtain the first image.
[0037] In this embodiment, the processor can acquire a first image captured by a camera device (such as a monocular camera). For example, the monocular camera can capture a video, and the processor can extract video frames from the video as the first image. Alternatively, the monocular camera can capture an image, and the captured image can be used as the first image.
[0038] Step S11: Input the first image into the depth estimation model and output the first depth image.
[0039] In this embodiment, the processor inputs a first image into a depth estimation model and then obtains a disparity map corresponding to the first image output by the depth estimation model. The first depth image can be obtained by converting the disparity map. The conversion of the disparity map into the corresponding depth image is prior art and will not be described further here.
[0040] In this embodiment, the depth estimation model is a pre-trained depth estimation model used to process the first image to obtain a depth image corresponding to the first image. The depth estimation model can be an autoencoder (AE) network.
[0041] Autoencoders are a class of artificial neural networks (ANNs) used in semi-supervised and unsupervised learning. Their function is to learn representations of the input information by using the input information as the learning target. An autoencoder consists of an encoder and a decoder. Based on learning paradigm, autoencoders can be divided into contractive autoencoders, regularized autoencoders, and variational autoencoders (VAEs), where the first two are discriminative models and the latter is a generative model. Based on architecture, autoencoders can be feedforward or recursive neural networks.
[0042] The training method for depth estimation models is briefly described below.
[0043] Please refer to the following: Figure 3 , Figure 3 This is a schematic diagram of the training method for the depth estimation model provided in the embodiments of this application.
[0044] Step S31: Establish a training dataset based on the images captured by the stereo camera.
[0045] In this embodiment of the application, the processor acquires images captured by the binocular stereo camera during vehicle operation, and establishes a training dataset based on these images, so as to use the training dataset to train the depth estimation model network to be trained.
[0046] The binocular camera includes a first camera and a second camera. The images captured by the binocular camera include images of the same scene (object) captured at the same time, that is, the left image captured by the first camera and the right image captured by the second camera, and the left image and the right image are both images of the same scene captured at the same time.
[0047] Step S32: Input the left image from the training dataset into the depth estimation model network to be trained to obtain the disparity map.
[0048] It is understandable that humans can perceive the three-dimensional world through their eyes because the image position of the same object in three-dimensional space differs in the horizontal direction between the left and right eyes; this is called parallax.
[0049] The disparity map is an image with one image in a stereo image pair as a reference, the size of which is the size of the reference image, and the element value is the disparity value. Disparity estimation is the process of finding the corresponding points between the left and right views, which is the stereo matching process.
[0050] Step S33: Obtain the predicted right image based on the left image and the disparity map.
[0051] In this embodiment, the left image is added to the disparity map to obtain the predicted right image learned by the depth estimation model network to be trained.
[0052] Step S34: Calculate the mean squared error between the right image in the training dataset and the predicted right image.
[0053] Step S35: Use the mean squared error as the loss value.
[0054] Step S36: Iteratively train the depth estimation model network to be trained based on the loss value until a trained depth estimation model is obtained.
[0055] For example, the training dataset includes images captured by a stereo camera. Left image A and right image B are images captured simultaneously by the first and second cameras on the stereo camera, respectively. Left image A is input into an autoencoder network, which outputs a disparity map. The left image A is added to the disparity map to obtain the predicted right image C from the autoencoder network. The mean squared error (MSE) is used to calculate the error between the predicted right image C and the actual right image B, which is then used as the loss value. The autoencoder network is iteratively trained based on this loss value until a fully trained autoencoder network is obtained.
[0056] The mean square deviation is used to calculate the average of the squares of the differences between each data point and the true value. The formula for mean square deviation is: Where MSE is the mean squared error, n is the number of samples, and y i It is real data y i The fitted data, i.e., y i ` is the data of the actual right image B, y i This is the data for predicting the right image C.
[0057] In this embodiment, the pixel value (or grayscale value) of each pixel in the first depth image represents only a relative depth. In some embodiments, relative depth can be understood as the logical relationship between the pixels. The value of each pixel in the first depth image (i.e., the pixel itself) is not a depth value with actual physical meaning; that is, the value of each pixel is not an absolute value provided in a specified unit of measurement (such as meters or centimeters). The distance between the real object corresponding to the pixel and the camera device or reference plane is called the depth value of that pixel, that is, the depth value of the pixel is the vertical distance from the real object corresponding to that point to the aforementioned camera device.
[0058] In other words, the grayscale value of each pixel in the first depth image is not the distance between the real object corresponding to that pixel and the camera device or reference plane. Therefore, the first depth image needs to be combined with the distance between the first camera and the second camera, i.e., other parameters, to calculate the depth value of each pixel in the first depth image. The depth values of each pixel in the first depth image constitute the depth information of the first depth image.
[0059] For a monocular camera (i.e., a single camera), the distance between each feature on an object and the monocular camera (i.e., the depth value of a point) can be the vertical distance between the point where each feature on the object is located and the monocular camera.
[0060] Step S12: Obtain the depth scaling factor.
[0061] In this embodiment, the processor acquires distance information obtained by the distance acquisition device and uses this distance information to calculate a depth scaling factor. The depth scaling factor indicates the relationship between the relative depth of a pixel in the depth image obtained in step S11 and the depth value of that pixel. Multiplying the pixel value of a pixel in the first depth image by the depth scaling factor yields the corresponding depth value for that pixel.
[0062] In some embodiments, the depth scaling factor can be calculated using radar information (depth information with units) obtained by radar. For example, points detected by radar are projected onto corresponding pixels in a depth image, where the depth image represents the radar detecting the same scene at the same time. Then, for a given point in the depth image, the depth scaling factor is calculated based on the distance information provided by the radar and the relative depth (e.g., pixel value) of that point in the depth image. The depth scaling relationships for all pixels in the depth image are obtained, and finally, the depth scaling factor is calculated based on these depth scaling relationships.
[0063] In this embodiment of the application, the depth scaling factor can be calculated based on pre-acquired training data, or it can be calculated based on the first image and the corresponding 3D point cloud image when the first image is acquired.
[0064] The following section details the method for calculating the depth scaling factor based on pre-acquired training data.
[0065] Taking radar as an example for distance acquisition device, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating the method for obtaining the depth scaling factor provided in an embodiment of this application.
[0066] Step S41: Obtain the external parameters between the camera device and the radar.
[0067] In this embodiment, the positional relationship between the camera and the radar can be predetermined. For example, the camera can be placed below the radar, and both positions are fixed. The calibration plate is placed within the overlapping area of the camera's field of view and the radar's field of view (e.g., Figure 1 The horizontal coverage area shown is 110. The surface of the calibration plate can be a checkerboard pattern. Before determining the extrinsic parameters between the camera and radar, multiple sets of calibration images need to be captured using a fixed-position radar and camera as input data for this extrinsic parameter determination method. After the radar or camera captures multiple sets of calibration images, the images are input into the processor, which processes them and determines their extrinsic parameters using the extrinsic parameter determination method between the camera and radar.
[0068] For example, a camera device captures two-dimensional images of a calibration board in multiple poses and sends these images to a processor. These multiple poses refer to several different poses. Similarly, a radar device captures three-dimensional point cloud images of the calibration board in multiple poses and sends these images to the processor. The processor combines the two-dimensional and three-dimensional point cloud images of the calibration board in the same pose, captured by both the camera and radar, into a single set of calibration images, thus determining multiple sets of calibration images. Based on these multiple sets of calibration images, the processor determines the extrinsic parameters between the camera device and the radar.
[0069] It is understandable that acquiring extrinsic parameters between camera equipment and radar is a relatively mature technology, so I will not go into details here.
[0070] In some embodiments, the extrinsic parameter may be stored in memory, and the processor may access the memory to read the extrinsic parameter. In other embodiments, the extrinsic parameter may be stored in the processor.
[0071] Step S42: Obtain the second image and the 3D point cloud image.
[0072] In this embodiment, a target scene can be captured in advance by a camera device to obtain a second image. At the same time, a three-dimensional point cloud image obtained by a distance acquisition device scanning the target scene is acquired. That is, the second image and the three-dimensional point cloud image acquired in step S42 are images captured by the camera device and the distance acquisition device at the same time for the same target scene, respectively. For example, the second image is an image of the scene in front (including vehicles 120 and 130) captured by the camera device on vehicle 100 at time T0, and the three-dimensional point cloud image is an image obtained by the radar on vehicle 100 scanning the scene in front (including vehicles 120 and 130) at time T0.
[0073] In other embodiments, the camera captures a second image of the scene within its horizontal coverage area 110 at a certain periodic rate. Similarly, the radar can capture a three-dimensional point cloud image of the scene within its horizontal coverage area 140 at a certain periodic rate. The timestamps on the second image and the three-dimensional point cloud image captured by the camera are determined, and the second image and the three-dimensional point cloud image with the same timestamp are selected.
[0074] Step S43: Convert the 3D point cloud image into a 2D image based on the extrinsic parameters.
[0075] In this embodiment, the point cloud data on the 3D point cloud image is projected according to extrinsic parameters to obtain a corresponding 2D image. The pixel value of each pixel in the 2D image is the depth value. The conversion of the 3D point cloud image into a 2D image based on extrinsic parameters is prior art and will not be described further here.
[0076] It should be noted that when acquiring the second image and the 3D point cloud image in step S42, the positional relationship between the camera device and the radar is set in the same way as when acquiring the extrinsic parameters. That is, if the camera device is placed below the radar when acquiring the extrinsic parameters, then when acquiring the second image and the 3D point cloud image in step S42, the positional relationship between the camera device and the radar is also set below the radar, referring to the positional relationship between the camera device and the radar when acquiring the extrinsic parameters.
[0077] Step S44: Input the second image into the depth estimation model and output the second depth image.
[0078] The depth estimation model in S44 is the same as the depth estimation model in step S11 above.
[0079] Step S45: Calculate the depth ratio relationship based on the two-dimensional image and the second depth image.
[0080] It's important to note that pixels in a two-dimensional image can be mapped to pixels in a second-depth image. That is, for a point 'a' on an object in a real-world scene, it is represented as pixel a1 in the two-dimensional image and as pixel a2 in the second-depth image; pixel a1 and pixel a2 correspond to each other.
[0081] In this embodiment, each pixel in the two-dimensional image has a depth value, and each pixel in the second depth image has a relative depth. For example, a set of images includes a two-dimensional image and a second depth image, wherein the three-dimensional point cloud image corresponding to the two-dimensional image and the second image corresponding to the second depth image are images of the same target scene obtained at the same time by the camera device and the distance acquisition device, respectively. Taking a set of images including two-dimensional image a and second depth image b as an example, the second image corresponding to the second depth image b and the three-dimensional point cloud image corresponding to the two-dimensional image are both images obtained at the same time for the same target scene. For point A in the target scene, its depth value on two-dimensional image a is 10m, and its relative depth on second depth image b is 2. Therefore, the depth ratio of point A can be calculated as 10cm / 2 = 5cm. Similarly, by calculating the ratio of all pixels on two-dimensional image a or second depth image b, a set of ratios (e.g., [5cm, 6cm, 5.5cm…5cm]) can be obtained. The average value obtained by summing the values in this set of ratios is the depth ratio between two-dimensional image a and second depth image b.
[0082] Step S46: Calculate the depth scaling factor based on the depth scaling relationship.
[0083] In this embodiment of the application, the depth scaling factor can be calculated based on the depth scaling relationship of multiple sets of images.
[0084] For example, as described above, if 100 sets of images are acquired for the same target scene, these 100 sets of images include 100 two-dimensional images and 100 corresponding second depth images. Accordingly, the depth ratio of these 100 sets of images can be calculated. The average of these depth ratios is then summed, and the resulting average value is the depth ratio factor.
[0085] In this embodiment, for each pixel in the depth image, the distance between the pixel and the camera device is obtained, and this distance is used as the pixel depth value. The ratio between the relative depth of the pixel and the pixel depth value is calculated to obtain the depth ratio of the pixels. Based on the depth ratio of the pixels in the depth image, the depth ratio factor of the depth image is obtained.
[0086] The following section explains in detail how to calculate the depth scaling factor based on the first image and the corresponding 3D point cloud image.
[0087] It should be noted that the three-dimensional point cloud image corresponding to the first image is the image scanned by the distance acquisition device, and the timestamp of the three-dimensional point cloud image is the same as that of the first image, and both the three-dimensional point cloud image and the first image are images of the target scene.
[0088] Specifically, a three-dimensional point cloud image corresponding to the first image scanned by the radar is acquired; extrinsic parameters between the camera device and the radar are acquired; the three-dimensional point cloud image is converted into a two-dimensional image based on the extrinsic parameters, wherein the two-dimensional image includes depth values. A depth ratio relationship is calculated based on the first image and the corresponding three-dimensional point cloud image, and a depth ratio factor is further calculated based on the depth ratio relationship; this depth ratio factor is the depth ratio factor of the first image.
[0089] Step S13: Obtain the depth information of the first depth image based on the first depth image and the depth scaling factor.
[0090] In this embodiment, multiplying the pixel value of each pixel in the first depth image by the depth scaling factor yields the depth value corresponding to that pixel, thereby calculating the depth values of all pixels in the first depth image. The depth values of all pixels in the first depth image constitute the depth information of the first depth image.
[0091] In some embodiments, the first depth image is depth-converted according to a depth scaling factor to obtain a third depth image, in which each pixel has a corresponding depth value, that is, the pixels of the first depth image have scale (or size), that is, they have units.
[0092] Please see Figure 5 , Figure 5 This is a schematic diagram of a depth estimation device provided in an embodiment of this application.
[0093] In this embodiment of the application, the depth estimation device 51 includes a first image acquisition module 511, an input module 512, an acquisition module 513, and a depth information acquisition module 514.
[0094] The first image acquisition module 511 is used to acquire the first image.
[0095] The input module 512 is used to input the first image into the depth estimation model and output the first depth image.
[0096] The acquisition module 513 is used to acquire the depth scaling factor.
[0097] The depth information acquisition module 514 is used to obtain the depth information of the first depth image based on the first depth image and the depth scaling factor.
[0098] See Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the application. In a preferred embodiment of the present invention, the electronic device 60 includes a memory 61 and at least one processor 62. Those skilled in the art should understand that... Figure 6 The structure of the computer device shown does not constitute a limitation of the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 60 may also include more or fewer other hardware or software than shown, or different component arrangements.
[0099] In some embodiments, electronic device 60 may further include a camera device and a distance acquisition device.
[0100] In some embodiments, the electronic device 60 includes a terminal capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, the hardware of which includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors, and embedded devices.
[0101] It should be noted that the electronic device 60 is merely an example. Other existing or future electronic products that are suitable for this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0102] In some embodiments, the memory 61 is used to store program code and various data, such as the depth estimation device 51 installed in the electronic device 60, and to enable high-speed, automatic access to programs or data during the operation of the electronic device 60. The memory 61 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable storage medium capable of carrying or storing data.
[0103] In some embodiments, the at least one processor 62 may be composed of integrated circuits, such as a single-packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The at least one processor 62 is the control unit of the electronic device 60, connecting various components of the electronic device 60 via various interfaces and lines. It performs various functions of the electronic device 60 and processes data, such as performing depth estimation, by running or executing programs or modules stored in the memory 61 and calling data stored in the memory 61.
[0104] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0105] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a server, personal computer, etc.) or processor to execute portions of the methods described in the various embodiments of the present invention.
[0106] In a further embodiment, combined with Figure 2 The at least one processor 62 can execute the operating device of the electronic device 60 and various installed applications (such as the depth estimation system 20), program code, etc., for example, the various modules mentioned above.
[0107] The memory 61 stores program code, and the at least one processor 62 can call the program code stored in the memory 61 to execute related functions. For example, Figure 5 The modules described herein are program codes stored in the memory 61 and executed by the at least one processor 62, thereby realizing the functions of the modules to achieve the purpose of depth estimation.
[0108] In one embodiment of the invention, the memory 61 stores one or more instructions (i.e., at least one instruction), which are executed by the at least one processor 62 to implement... Figure 2 The purpose of the depth estimation shown is...
[0109] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0110] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0112] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other elements or, and the singular does not exclude the plural. Multiple elements or devices recited in the apparatus claims may also be implemented by a single element or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A depth estimation method applied to an electronic device, the electronic device comprising an image capturing device for capturing, the method comprising: The method comprises: acquiring a first image; inputting the first image into a depth estimation model to obtain a first depth image; acquiring a depth scale factor, wherein the depth scale factor is used to indicate the relationship between the relative depth of a pixel point on the first depth image and the depth value of the pixel point; the acquisition of the depth scale factor comprises: acquiring the distance between the pixel point and the camera device, taking the distance as the depth value of the pixel point; calculating the ratio between the relative depth of the pixel point and the depth value of the pixel point to obtain the depth scale relationship of the pixel point; and obtaining the depth scale factor according to the depth scale relationship of the pixel point; calculating the depth information of the first depth image according to the depth scale factor and the first depth image.
2. The depth estimation method of claim 1, wherein, The electronic device further comprises a radar, and the acquisition of the distance between the pixel point and the camera device, taking the distance as the depth value of the pixel point comprises: acquiring the distance between the pixel point and the camera device by using the radar, and taking the distance as the depth value of the pixel point.
3. The depth estimation method of claim 2, wherein, The acquisition of the distance between the pixel point and the camera device by using the radar, and taking the distance as the depth value of the pixel point comprises: acquiring a three-dimensional point cloud image corresponding to the first image scanned by the radar; acquiring the external parameter between the camera device and the radar; converting the three-dimensional point cloud image into a two-dimensional image according to the external parameter, wherein the two-dimensional image comprises the depth value.
4. The depth estimation method of claim 1, wherein, The electronic device further comprises a radar, and the acquisition of the depth scale factor comprises: acquiring the external parameter between the camera device and the radar; acquiring a second image captured by the camera device; acquiring a three-dimensional point cloud image corresponding to the second image scanned by the radar; converting the three-dimensional point cloud image into a two-dimensional image according to the external parameter, wherein the two-dimensional image comprises the depth value; inputting the second image into the depth estimation model to output a second depth image; calculating a depth scale relationship according to the two-dimensional image and the second depth image; calculating the depth scale factor according to the depth scale relationship.
5. The depth estimation method of claim 4, wherein, The calculation of the depth scale relationship according to the two-dimensional image and the second depth image comprises: obtaining the depth value corresponding to a pixel point on the second depth image according to the two-dimensional image; calculating the depth scale relationship of the pixel point according to the depth value corresponding to the pixel point and the relative depth of the pixel point; acquiring the depth scale relationship of all pixel points on the second depth image; calculating the depth scale factor according to the depth scale relationship of all the pixel points.
6. The depth estimation method of claim 5, wherein, When the number of the second images is greater than 1, the calculation of the depth scale factor according to the depth scale relationship of all the pixel points comprises: acquiring the depth scale relationship of all the second images; calculating the depth scale factor according to the obtained depth scale relationship.
7. A depth estimation apparatus for implementing the method of claim 1, characterized by comprise: a first image acquisition module, configured to acquire a first image; an input module, configured to input the first image into a depth estimation model to output a first depth image; An acquisition module is configured to acquire a depth scale factor, wherein the depth scale factor is used to indicate a relationship between a relative depth of a pixel point on the first depth image and a depth value of the pixel point. A depth information acquisition module is configured to obtain depth information of the first depth image according to the first depth image and the depth scale factor.
8. An electronic device, comprising: The electronic device comprises a memory and a processor, wherein the memory is configured to store at least one instruction, and the processor is configured to execute the at least one instruction to implement the depth estimation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores at least one instruction, and the at least one instruction is executed by the processor to implement the depth estimation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Point cloud density improving method for carrying out depth prediction based on two-dimensional image gray scale
CN111161338A
Correction method, electronic equipment and computer readable storage medium
CN113298785A