Depth estimation method and apparatus for target object, storage medium, and electronic device
By combining multiple depth estimation algorithms and utilizing feature extraction and prediction branch models, the depth coordinates and uncertainties of the target object are determined, solving the problem of insufficient robustness of a single depth estimation algorithm and achieving higher accuracy and robustness in depth estimation.
Patent Information
- Application Number
- CN202210467512.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-04-29
AI Technical Summary
Existing single depth estimation algorithms are susceptible to noise and misjudgment, lack robustness, and have poor anti-interference ability.
A method combining multiple depth estimation algorithms is adopted. Through feature extraction, multiple prediction branch models and feature fusion, the depth coordinates and uncertainties of the target object are determined, and the final depth estimate is determined by weighted averaging.
This improves the robustness and anti-interference ability of depth estimation, and enhances the accuracy and precision of depth estimation.
Smart Images

Figure CN114782510B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to computer vision technology, and in particular to a target object depth estimation method and device, a storage medium, and an electronic device. BACKGROUND
[0002] Depth estimation is to obtain distance information of each point in the space in the scene in the image to the camera on the camera, and a graph composed of such distance information is called a depth map. Monocular depth estimation is to estimate the distance of a target in an image relative to a camera from an RGB image collected by a single camera. Estimating depth information based on a monocular camera is defined as a monocular camera estimation (MDE) problem. In the field of automotive assisted driving / autonomous driving, depth estimation of a vehicle is a very important and basic task requirement and is a key step for scene reconstruction and understanding tasks. SUMMARY
[0003] Embodiments of the present disclosure provide a target object depth estimation method and device, a storage medium, and an electronic device.
[0004] According to an aspect of an embodiment of the present disclosure, a target object depth estimation method is provided, comprising:
[0005] performing feature extraction on an image collected by a camera and including a target object to obtain a feature map;
[0006] determining a first plane coordinate value of the target object in an image coordinate system in the image based on the feature map;
[0007] determining a plurality of first depth coordinate values of the target object in a local three-dimensional coordinate where the camera is located and a depth uncertainty corresponding to each of the first depth coordinate values based on the feature map and the first plane coordinate value;
[0008] determining a second depth coordinate value of the target object in the local three-dimensional coordinate based on the plurality of first depth coordinate values and the depth uncertainty corresponding to each of the first depth coordinate values.
[0009] According to another aspect of an embodiment of the present disclosure, a target object depth estimation device is provided, comprising:
[0010] a feature extraction module configured to perform feature extraction on an image collected by a camera and including a target object to obtain a feature map;
[0011] a plane coordinate prediction module configured to determine a first plane coordinate value of the target object in an image coordinate system in the image based on the feature map obtained by the feature extraction module;
[0012] an object coordinate estimation module configured to determine a plurality of first depth coordinate values of the target object in a local three-dimensional coordinate in which the camera is located and a depth uncertainty corresponding to each of the first depth coordinate values based on the feature map obtained by the feature extraction module and the first plane coordinate values determined by the plane coordinate prediction module;
[0013] a target object estimation module configured to determine a second depth coordinate value of the target object in the local three-dimensional coordinate based on the plurality of first depth coordinate values and the depth uncertainty corresponding to each of the first depth coordinate values determined by the object coordinate estimation module.
[0014] According to still another aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores a computer program for performing the depth estimation method of the target object according to any of the above embodiments.
[0015] According to still another aspect of the embodiments of the present disclosure, an electronic device is provided, which comprises:
[0016] a processor;
[0017] a memory for storing executable instructions of the processor;
[0018] the processor is configured to read the executable instructions from the memory and execute the instructions to implement the depth estimation method of the target object according to any of the above embodiments.
[0019] The depth estimation method and device of the target object, the storage medium and the electronic device provided by the above embodiments of the present disclosure, by predicting a plurality of first depth coordinate values, each first depth coordinate value corresponding to a different depth estimation algorithm, determining a second depth coordinate value based on the plurality of first depth coordinate values and the uncertainty thereof, realize the depth estimation combining multiple depth estimation algorithms, overcome the problem of dependence on a single depth estimation algorithm caused by the single depth estimation algorithm, and have good robustness and strong anti-interference ability.
[0020] The technical solutions of the present disclosure will be described in further detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0021] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. The drawings provided herein are for illustrative purposes only and constitute part of the detailed description. They serve the purpose of explaining the present disclosure and do not limit the present disclosure. In the drawings, the same reference numerals generally refer to the same components or steps.
[0022] Figure 1ais a network structure schematic diagram of a depth estimation network model involved in an example embodiment of the present disclosure.
[0023] Figure 1b is Figure 1a An optional example structure diagram of a prediction branch model in the depth estimation network model provided by the present disclosure is shown.
[0024] Figure 2 is a flowchart of a depth estimation method of a target object provided by an example embodiment of the present disclosure.
[0025] Figure 3 is a depth estimation network model involved in an example embodiment of the present disclosure. Figure 2 is a flowchart of step 203 in the embodiment shown.
[0026] Figure 4 is a depth estimation network model involved in an example embodiment of the present disclosure. Figure 3 is a flowchart of step 2031 in the embodiment shown.
[0027] Figure 5 is a depth estimation network model involved in an example embodiment of the present disclosure. Figure 2 is a flowchart of step 204 in the embodiment shown.
[0028] Figure 6 is a flowchart of a depth estimation method of a target object provided by another example embodiment of the present disclosure.
[0029] Figure 7 is a structure schematic diagram of a depth estimation device of a target object provided by an example embodiment of the present disclosure.
[0030] Figure 8 is a structure schematic diagram of a depth estimation device of a target object provided by another example embodiment of the present disclosure.
[0031] Figure 9 is a structure diagram of an electronic device provided by an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In the following, example embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the example embodiments described herein.
[0033] It should be noted that: unless otherwise specified, the relative arrangement, numerical expression and numerical value of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0034] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor represent a logical sequence between them.
[0035] It should also be understood that in the embodiments of the present disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more.
[0036] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, it can be understood as one or more in general, without explicit limitation or in the context of the opposite indication.
[0037] In addition, the term "and / or" in the present disclosure is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in the present disclosure generally represents that the front and rear associated objects are in an "or" relationship.
[0038] It should also be understood that the description of the present disclosure for each embodiment emphasizes the differences between each embodiment, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.
[0039] At the same time, it should be understood that, for the convenience of description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship.
[0040] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application or uses.
[0041] The techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods, and devices should be considered as part of the specification.
[0042] It should be noted that: similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in the subsequent drawings.
[0043] The embodiments of the present disclosure can be applied to terminal devices, computer systems, servers and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and the like.
[0044] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0045] Summary of the application
[0046] In the process of implementing the present disclosure, the inventors found that the existing depth estimation method is usually estimated by a single depth estimation algorithm, but the prior art at least has the following problems: since the estimation is based on a single depth estimation algorithm, it is easy to be affected by the accuracy of the single depth estimation algorithm, and it is not robust enough to noise and misjudgment.
[0047] Exemplary network structure
[0048] Figure 1a FIG. 1 is a schematic diagram of the network structure of a depth estimation network model according to an example embodiment of the present disclosure. As shown in FIG. 1, the depth estimation network model includes a feature extraction branch model 101, a prediction branch model 102, and a feature fusion branch model 103.
[0049] Optionally, Figure 1b is Figure 1a An optional example structure diagram of the prediction branch model in the depth estimation network model provided. As Figure 1b shown, the prediction branch model 102 can include but is not limited to: a plane prediction branch model 1021, a depth prediction branch model 1022, a key point branch model 1023, a projection height branch model 1024, and an actual height prediction branch model 1025.
[0050] Utilizing Figure 1a And Figure 1b The process of depth estimation performed by the provided depth estimation network model can include: inputting a single image (e.g., an RGB image) captured by a camera into a feature extraction branch model 101 of the depth estimation network model, performing feature extraction on the image by the feature extraction branch model 101 to obtain a feature map; inputting the feature map into multiple prediction branch models 102 respectively; processing the feature map by a plane prediction branch model 1021 to obtain a first plane coordinate (u, v) and a center offset value of a target object (e.g., a vehicle, etc.) in an image coordinate system corresponding to the image, wherein the first plane coordinate represents a center point coordinate value of the target object, and the first plane coordinate is used to determine a result corresponding to the target object from results of at least one object output by a key point branch model 1023, a projection height branch model 1024, and an actual height prediction branch model 1025; processing the feature map by a depth prediction branch model 1022 to obtain a first depth estimation value and a first uncertainty degree corresponding to the target object, wherein the depth prediction branch model 1022 can output at least one candidate estimation value corresponding to at least one object in the image, and the first depth estimation value is determined by selecting the candidate estimation value closest to the first plane coordinate; processing the feature map by the key point branch model 1023 to obtain a plurality of key point prediction values and a second uncertainty degree of the target object, wherein the key points of the target object in the image can be coordinate values of 8 key points corresponding to a smallest containing cube of the target object in the image, and the specific determination process can include: the key point branch model 1023 can predict at least one group of key points corresponding to at least one object in the image, and the 8 key point coordinates close to the first plane coordinate are determined as the output result of the key point branch model 1023 by selecting (closest to) the plurality of key point prediction values according to the first plane coordinate; for example, when the target object is a vehicle, the key points are 8 vertex coordinate values of a cube containing the vehicle; processing the feature map by the projection height branch model 1024 to obtain a projection height value and a third uncertainty degree of the target object, wherein the projection height represents the height of the target object in the image, and the projection height branch model 1024 can predict a plurality of predicted projection height values (e.g., a plurality of objects included in the image, etc.), and the projection height value of the target object is determined by selecting the plurality of predicted projection height values according to the first plane coordinate; processing the feature map by the actual height prediction branch model 1025 to obtain an actual height prediction value of the target object, for example, when the target object is a vehicle, the actual height prediction value is the predicted actual height of the vehicle.
[0051] Based on the outputs of the plurality of prediction branch models 102, three depth estimation values can be determined, optionally including a first depth estimation value Z1, a second depth estimation value Z2 determined based on the first plane coordinate, the plurality of key point prediction values and the intrinsic matrix of the camera, and a third depth estimation value Z3 determined based on the first plane coordinate, the projected height value and the intrinsic matrix of the camera.
[0052] Optionally, the second depth estimation value can be determined based on the following formula (1):
[0053]
[0054] wherein k1 represents a lower key point coordinate value of the target object in the image (e.g., determined by averaging the coordinate values of the 4 key points belonging to the bottom surface among the 8 key points), k2 represents an upper key point coordinate value of the target object in the image (e.g., determined by averaging the coordinate values of the 4 key points belonging to the top surface among the 8 key points), and specifically, the 8 key points in the embodiment are selected based on the first plane coordinate; f represents a focal length in the camera intrinsic; and H represents an actual height prediction value predicted by the actual height prediction branch model.
[0055] The third depth estimation value can be determined based on the following formula (2):
[0056]
[0057] wherein h represents a projected height value predicted by the projected height branch model, and specifically, the projected height value in the embodiment is determined based on the selection of the first plane coordinate; f represents a focal length in the camera intrinsic; and H represents an actual height prediction value predicted by the actual height prediction branch model.
[0058] The feature fusion branch model 103 determines a target depth estimation value of the target object in the local three-dimensional coordinate of the camera based on the first depth estimation value, the second depth estimation value and the third depth estimation value, and the first uncertainty, the second uncertainty and the third uncertainty. Optionally, the target depth estimation value can be determined based on the following formula (3):
[0059]
[0060] wherein Z represents the target depth estimation value; i is the number of depth estimation values output by the prediction branch model 102, and in the embodiment, i is 3; σ i represents the uncertainty, σ1, σ2 and σ3 respectively represent the first uncertainty, the second uncertainty and the third uncertainty, and the reciprocal of the uncertainty can represent the weight value of the corresponding depth estimation value, i.e., formula (3) realizes weighted average.
[0061] The above can achieve the estimation of the target depth under the local three-dimensional coordinates of the target object. Furthermore, in order to better utilize the target depth estimation, the first plane coordinates (u, v) are inversely projected and transformed in combination with the intrinsic parameter matrix K to obtain the X-axis and Y-axis coordinates of the center point of the target object under the local three-dimensional coordinates. That is, this embodiment realizes the determination of the three-dimensional coordinates of the target object under the local three-dimensional coordinates.
[0062] In this embodiment, the depth value of the target object is estimated using three depth estimation algorithms. Compared with the depth estimation methods based on a single depth estimation algorithm in the prior art, this embodiment relies less on a single depth estimation algorithm and is less sensitive to noise misjudgment, thus achieving greater robustness against interference. Furthermore, compared with direct depth estimation algorithms, predicting key points and projection height from the image is simpler, making depth estimation easier. Combining the three depth estimation algorithms and their uncertainties improves the accuracy of depth estimation.
[0063] Typically, network models need to be trained before application, as described in the embodiments of this disclosure. Figure 1a and Figure 1b Before application, the provided depth estimation network model also requires network training. The training process includes: inputting sample images with known ground truth depth values into the depth estimation network model and outputting depth estimates; determining the network loss based on the depth estimates and uncertainties; and training the depth estimation network model using the network loss. Optionally, the network loss can be based on the branch loss L. i With depth loss L depth The sum is determined; where the branch loss can be determined based on the following formula (4), where i takes the values 1, 2, and 3:
[0064]
[0065] Where, σ i The uncertainty (for different values of i, corresponding to the output uncertainties of depth prediction branch model 1022, key point branch model 1023, and projection height branch model 1024, respectively), z i The z represents the estimated depth value (where different values of i correspond to the depth output of depth prediction branch model 1022, keypoint branch model 1023, and projection height branch model 1024, respectively). * This represents the ground truth depth value corresponding to the sample image.
[0066] Depth loss L depth It can be determined based on the following formula (5):
[0067] L depth =|zz * | Formula (5)
[0068] Where z represents the depth estimate of the target object predicted by the depth estimation network model; z * This represents the ground truth depth value corresponding to the sample image.
[0069] Exemplary method
[0070] Figure 2 This is a schematic flowchart illustrating a depth estimation method for a target object provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 2 As shown, it includes the following steps:
[0071] Step 201: Extract features from the image of the target object captured by the camera to obtain a feature map.
[0072] In this embodiment, it can be achieved through the above... Figure 1a The feature extraction branch model 101 in the depth estimation network model shown performs feature extraction to obtain a feature map. This embodiment can be applied to any scenario. For example, using an in-vehicle camera to collect images around a vehicle, the target object can be other vehicles, pedestrians, buildings, or other obstacles. In this case, the depth estimation method of this embodiment can provide more information for assisted driving or autonomous driving of the vehicle, thereby improving driving safety.
[0073] Step 202: Determine the first plane coordinate value of the target object in the image coordinate system based on the feature map.
[0074] Optionally, this embodiment can be based on the above. Figure 1b The plane prediction branch model 1021 provided in the paper realizes the prediction of the coordinate values of the first plane.
[0075] Step 203: Based on the feature map and the first plane coordinate values, determine multiple first depth coordinate values of the target object in the local three-dimensional coordinates of the camera and the depth uncertainty corresponding to each first depth coordinate value.
[0076] Optionally, this embodiment can be achieved through... Figure 1b Multiple branch models, including depth prediction branch model 1022, key point branch model 1023, projection height branch model 1024, and actual height prediction branch model 1025, are provided in multiple prediction branch models 102 to process the feature map. Combined with the first plane coordinate value, multiple first depth coordinate values of the target object in the local three-dimensional coordinates of the camera and the depth uncertainty corresponding to each first depth coordinate value are output.
[0077] Step 204: Based on multiple first depth coordinate values and the depth uncertainty corresponding to each first depth coordinate value, determine the second depth coordinate value of the target object in local three-dimensional coordinates.
[0078] Optionally, the weight value of each first depth coordinate value can be determined by the depth uncertainty corresponding to each first depth coordinate value, and the second depth coordinate value can be determined by weighted averaging of the weight values.
[0079] The present invention discloses a method for estimating the depth of a target object in the above embodiments. This method predicts multiple first depth coordinate values through multiple prediction branch models. Each first depth coordinate value corresponds to a different depth estimation algorithm in the multiple prediction branch models. A second depth coordinate value is determined based on the multiple first depth coordinate values and their uncertainties. This method achieves depth estimation by combining multiple depth estimation algorithms, overcomes the dependence on a single depth estimation algorithm, and has good robustness and strong anti-interference ability.
[0080] like Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step 203 may include the following steps:
[0081] Step 2031: Based on the feature map and the coordinate values of the first plane, determine multiple predicted values corresponding to the target object, the prediction uncertainty of each predicted value, and the predicted actual height of the target object.
[0082] The predicted values in this embodiment may include, but are not limited to: depth predicted values, key point predicted values, and projected height predicted values; optionally, the actual height predicted value can be obtained through... Figure 1b The projection height was predicted using the provided branch model 1024.
[0083] Step 2032: Based on multiple predicted values and actual height prediction values, determine multiple first depth coordinate values in local three-dimensional coordinates.
[0084] In this embodiment, when the predicted value is the key point predicted value or the projected height predicted value, it can be achieved through... Figure 1b The corresponding first depth coordinate value is determined by calculating using formula (1) or formula (2) in the provided embodiments.
[0085] Step 2033: Based on multiple prediction uncertainties, determine the depth uncertainty corresponding to each first depth coordinate value.
[0086] Since each predicted value corresponds to a prediction uncertainty, the depth uncertainty corresponding to the first depth coordinate value determined by the predicted value is taken as the prediction uncertainty corresponding to the predicted value.
[0087] This embodiment proposes multiple predicted values corresponding to information such as depth, key points, and projection height. Based on these multiple predicted values, multiple first depth coordinate values are determined by different depth estimation algorithms. This adds a depth estimation algorithm, overcomes the problem of over-reliance on a single depth estimation algorithm, avoids inaccurate depth estimation results due to the inaccuracy of a single depth estimation algorithm, and improves the robustness of the depth estimation method provided in this embodiment.
[0088] like Figure 4 As shown above, in the above Figure 3 Based on the illustrated embodiment, step 2031 may include the following steps:
[0089] Step 401: Based on multiple prediction branch models, predict the feature map and determine the candidate prediction result corresponding to each object in at least one object in the image.
[0090] Optionally, each candidate result includes a candidate predicted value and the uncertainty corresponding to the candidate predicted value; this can be achieved through the above... Figure 1b In the embodiment shown, multiple predictions are made using the depth prediction branch model 1022, the key point branch model 1023, and the projection height branch model 1024 provided, to obtain multiple candidate prediction results for each object.
[0091] Step 402: Based on the first plane coordinate values, determine the candidate prediction result corresponding to the target object from at least one candidate prediction result.
[0092] In this embodiment, the first planar coordinate value includes the horizontal and vertical coordinate values of the center point of the target object, which represents the center point coordinates of the target object in the image. Based on the center point coordinates, the candidate prediction results of the corresponding target object can be selected from the candidate prediction results of multiple objects. Optionally, the selection can be performed by the distance from the center point coordinates.
[0093] Step 403: Based on the candidate prediction results corresponding to the target object, obtain multiple predicted values and the prediction uncertainty and the actual height prediction value of the target object corresponding to each predicted value.
[0094] In this embodiment, the candidate prediction result corresponding to the target object is determined from multiple candidate prediction results using planar coordinate values. Since the image may include multiple objects (e.g., a road image may include multiple vehicles, pedestrians, buildings, etc.), a set of prediction values corresponding to the target object is determined by the planar coordinate values of the target object. By filtering with the first planar coordinate values, subsequent operations are based only on the candidate prediction results corresponding to the target object, without having to perform subsequent operations on the candidate prediction results corresponding to each object. This solves the problem that the results are inaccurate because the depth estimation is performed on all candidate prediction results, and the results are not for the target object. This improves the accuracy of at least one prediction value corresponding to the target object.
[0095] Optionally, based on the above embodiments, step 2032 may include at least two of the following:
[0096] a1 is a first depth coordinate value based on the direct depth prediction value.
[0097] In this embodiment, it can be based on Figure 1b The provided depth prediction branch model 1022 determines the first depth coordinate value. Optionally, the depth prediction branch model in this embodiment can adopt any existing network model for depth prediction. This embodiment does not impose specific restrictions on the network structure for depth prediction.
[0098] a2 determines the projected height of the target object in the image based on the predicted values of multiple key points, and determines a first depth coordinate value based on the projected height value, the actual height predicted value, and the camera's intrinsic parameter matrix.
[0099] Optionally, based on the above Figure 1b Formula (1) in the provided embodiment realizes the prediction of a first depth coordinate value, wherein the key points include 8 key points of the cube surrounding the target object, and the difference between the key points of the upper and lower planes determines the projected height value.
[0100] a3 determines a first depth coordinate value based on the projected height prediction value, the actual height prediction value, and the camera's intrinsic parameter matrix.
[0101] Optionally, based on the above Figure 1b Formula (2) in the provided embodiment enables the prediction of a first depth coordinate value.
[0102] In the embodiment, the plurality of predicted values include at least two of the direct depth predicted value in the local three-dimensional coordinate, the plurality of key point predicted values in the local three-dimensional coordinate, and the projection height predicted value in the local three-dimensional coordinate, different predicted values are determined through different branch network models, wherein each predicted value determines a first depth coordinate value through a different depth estimation algorithm, a second depth coordinate value is determined by using the plurality of first depth coordinate values and the uncertainty thereof, the depth estimation combining a plurality of depth estimation algorithms is realized, the problem of dependence on a single depth estimation algorithm caused by the single depth estimation algorithm is overcome, and the robustness of the depth estimation method is improved; and compared with directly estimating the depth value, the key point prediction and the projection height prediction on the target image are simpler and easier to implement, so that the depth estimation of the embodiment is easier to implement, and the depth estimation value obtained by combining a plurality of predicted values is more accurate.
[0103] As shown in Figure 5 the above Figure 2 embodiment, step 204 can include the following steps:
[0104] Step 2041, determining the weight value of each first depth coordinate value based on the plurality of uncertainties.
[0105] Optionally, the reciprocal of each uncertainty is used as the weight value of each first depth coordinate value.
[0106] Step 2042, weighting the plurality of first depth coordinate values based on the plurality of weight values to obtain the second depth coordinate value of the target object in the camera coordinate corresponding to the camera.
[0107] In the embodiment, after determining the weight value of each first depth coordinate value, the second depth coordinate value can be determined based on the formula (3) in the above Figure 1b embodiment, in the embodiment, the second depth coordinate value is determined by weighted average, a plurality of depth estimation algorithms are combined, the problem of dependence on a single depth estimation algorithm in the existing depth estimation method is overcome, and a more robust depth estimation method is provided.
[0108] Figure 6 is a flowchart of a depth estimation method of a target object provided by another exemplary embodiment of the present disclosure. As shown in 6, the embodiment is based on the embodiment shown in Figure 2 , and can further include the following steps:
[0109] Step 601, performing inverse projection transformation from the image coordinate system to the local three-dimensional coordinate system on the first plane coordinate value based on the second depth coordinate value and the intrinsic matrix of the camera.
[0110] Optionally, the intrinsic matrix of the camera can be determined as known data when the camera is determined; in this embodiment, the planar coordinate in the image coordinate system is mapped to the local three-dimensional coordinate where the camera is located through inverse projection transformation combined with the second depth coordinate value and the intrinsic matrix of the camera, and then the coordinate values of the X axis and the Y axis in the local three-dimensional coordinate system are obtained.
[0111] In step 602, the second planar coordinate value of the target object in the local three-dimensional coordinate system is determined.
[0112] In this embodiment, in order to better use the target depth estimation value, the first planar coordinate is subjected to inverse projection transformation combined with the intrinsic matrix K, so as to obtain the coordinate values of the X axis and the Y axis of the center point of the target object in the local three-dimensional coordinate, that is, this embodiment realizes determination of the three-dimensional coordinate value of the target object in the local three-dimensional coordinate, realizes better positioning of the target object, and provides a basis for subsequent other operations.
[0113] Optionally, on the basis of the above-mentioned embodiments, the following can also be included:
[0114] b1, determining the center offset value of the target object in the image coordinate system based on the feature map.
[0115] Optionally, as Figure 1b As shown in the embodiments provided, the planar prediction branch model 1021 can not only output the first planar coordinate, but also output the center offset value.
[0116] b2, correcting the second planar coordinate value based on the center offset value to obtain the third planar coordinate value after correction.
[0117] In this embodiment, the second planar coordinate value is corrected through the output center offset value to obtain the coordinate values of the X axis and the Y axis of the center point of the target object in the local three-dimensional coordinate system after correction; through correction, the prediction accuracy of the center point coordinate of the target object is improved.
[0118] Any one of the target object depth estimation methods provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capability, including but not limited to: terminal devices and servers, etc. Alternatively, any one of the target object depth estimation methods provided in the embodiments of the present disclosure can be executed by a processor, such as a processor executing any one of the target object depth estimation methods mentioned in the embodiments of the present disclosure by calling corresponding instructions stored in a memory. The following will not be described in detail.
[0119] Exemplary apparatus
[0120] Figure 7 is a structural schematic diagram of a target object depth estimation device provided by an exemplary embodiment of the present disclosure. As Figure 7As shown, the device provided by the embodiment includes:
[0121] The feature extraction module 71 is configured to perform feature extraction on the image collected by the camera and including the target object, to obtain a feature map.
[0122] The plane coordinate prediction module 72 is configured to determine a first plane coordinate value of the target object in the image coordinate system based on the feature map obtained by the feature extraction module 71.
[0123] The object coordinate estimation module 73 is configured to determine a plurality of first depth coordinate values of the target object in the local three-dimensional coordinate and a depth uncertainty corresponding to each first depth coordinate value based on the feature map obtained by the feature extraction module 71 and the first plane coordinate value determined by the plane coordinate prediction module 72.
[0124] The target object estimation module 74 is configured to determine a second depth coordinate value of the target object in the local three-dimensional coordinate based on the plurality of first depth coordinate values and the depth uncertainty corresponding to each first depth coordinate value determined by the object coordinate estimation module 73.
[0125] The depth estimation device for the target object provided by the above-mentioned embodiment of the present disclosure predicts a plurality of first depth coordinate values, each first depth coordinate value corresponds to a different depth estimation algorithm, determines a second depth coordinate value based on the plurality of first depth coordinate values and the uncertainty thereof, realizes depth estimation combining a plurality of depth estimation algorithms, overcomes the problem of dependence on a single depth estimation algorithm caused by the single depth estimation algorithm, has good robustness and strong anti-interference ability.
[0126] Figure 8 is a structural schematic diagram of the depth estimation device for the target object provided by another exemplary embodiment of the present disclosure. As Figure 8 shown, in the device provided by the embodiment, the object coordinate estimation module 73 includes:
[0127] The object prediction unit 731 is configured to determine a plurality of prediction values corresponding to the target object and a prediction uncertainty corresponding to each prediction value, and an actual height prediction value of the target object based on the feature map and the first plane coordinate value.
[0128] The coordinate value determination unit 732 is configured to determine a plurality of first depth coordinate values in the local three-dimensional coordinate based on the plurality of prediction values and the actual height prediction value.
[0129] The depth determination unit 733 is configured to determine a depth uncertainty corresponding to each first depth coordinate value based on the plurality of prediction uncertainties.
[0130] Optionally, the object prediction unit 731 is specifically configured to predict the feature map based on the plurality of prediction branch models to determine a plurality of candidate prediction results corresponding to each of the at least one object in the image; determine a candidate prediction result corresponding to the target object from the at least one candidate prediction result based on the first plane coordinate value; and obtain a plurality of prediction values, a prediction uncertainty corresponding to each of the plurality of prediction values, and an actual height prediction value of the target object based on the candidate prediction result corresponding to the target object.
[0131] Optionally, the plurality of prediction values include at least two of: a direct depth prediction value in a local three-dimensional coordinate, a plurality of key point prediction values in the local three-dimensional coordinate, and a projection height prediction value in the local three-dimensional coordinate.
[0132] The coordinate value determination unit 732 is specifically configured to determine a first depth coordinate value based on the direct depth prediction value; and / or determine a projection height value of the target object in the image based on the plurality of key point prediction values, and determine a first depth coordinate value based on the projection height value, the actual height prediction value, and an intrinsic matrix of the camera; and / or determine a first depth coordinate value based on the projection height prediction value, the actual height prediction value, and the intrinsic matrix of the camera.
[0133] In some optional embodiments, the target object estimation module 74 includes:
[0134] The weight value unit 741 is configured to determine a weight value of each of the first depth coordinate values based on the plurality of uncertainties.
[0135] The coordinate value estimation unit 742 is configured to weight the plurality of first depth coordinate values based on the plurality of weight values to obtain a second depth coordinate value of the target object in a camera coordinate corresponding to the camera.
[0136] In some optional embodiments, the target object estimation module 74 further includes:
[0137] The inverse projection module 85 is configured to perform inverse projection transformation of the first plane coordinate value from an image coordinate system to a local three-dimensional coordinate system based on the second depth coordinate value and the intrinsic matrix of the camera, and determine a second plane coordinate value of the target object in the local three-dimensional coordinate system.
[0138] In some optional embodiments, the target object estimation module 74 further includes:
[0139] The coordinate correction module 86 is configured to determine a center offset value of the target object in the image coordinate system based on the feature map, and correct the second plane coordinate value based on the center offset value to obtain a third plane coordinate value after correction.
[0140] Exemplary electronic device
[0141] In the following, reference is made to Figure 9An electronic device according to embodiments of the present disclosure will be described. The electronic device can be either or both of the first device 100 and the second device 200, or a stand-alone device independent from them, which can communicate with the first and second devices to receive the acquired input signals therefrom.
[0142] Figure 9 A block diagram of an electronic device according to embodiments of the present disclosure is illustrated.
[0143] As Figure 9 illustrated, the electronic device 90 includes one or more processors 91 and a memory 92.
[0144] The processor 91 can be a central processing unit (CPU) or other form of processing unit having data processing and / or instruction execution capabilities, and can control other components in the electronic device 90 to perform desired functions.
[0145] The memory 92 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, which the processor 91 can execute to implement the depth estimation method of the target object according to various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, and the like can also be stored in the computer-readable storage media.
[0146] In one example, the electronic device 90 can further include an input device 93 and an output device 94, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0147] For example, when the electronic device is the first device 100 or the second device 200, the input device 93 can be the microphone or the microphone array described above, for capturing the input signal of the sound source. When the electronic device is a stand-alone device, the input device 93 can be a communication network connector, for receiving the acquired input signal from the first device 100 and the second device 200.
[0148] In addition, the input device 93 can further include, for example, a keyboard, a mouse, and the like.
[0149] The output device 94 can output various information including the determined distance information, direction information, etc. to the outside. The output device 94 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, etc.
[0150] Of course, in order to simplify, Figure 9 Only some of the components in the electronic device 90 related to the present disclosure are shown in the figure, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 90 can further include any other appropriate components according to specific application cases.
[0151] Exemplary computer program product and computer readable storage medium
[0152] In addition to the above-mentioned methods and devices, embodiments of the present disclosure can also be a computer program product, which includes computer program instructions that, when executed by a processor, cause the processor to perform the steps in the depth estimation method of the target object according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of the present specification.
[0153] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and a conventional procedural programming language such as "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on a user device, as an independent software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0154] In addition, embodiments of the present disclosure can also be a computer readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the depth estimation method of the target object according to various embodiments of the present disclosure described in the above "Exemplary Methods" section of the present specification.
[0155] The computer readable storage medium can be any combination of one or more computer readable medium(s). The computer readable medium can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0156] The above generally describes the basic principles of the disclosure in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the disclosure are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as the various embodiments of the disclosure must have. In addition, the above specific details of the disclosure are only for the purpose of example and for the purpose of understanding, and the above details do not limit the disclosure to the above specific details.
[0157] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the embodiments can be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0158] The block diagrams of the devices, apparatuses, equipment, systems involved in the disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "include but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0159] The methods and apparatus of the present disclosure can be implemented in a number of ways. For example, the methods and apparatus of the present disclosure can be implemented using software, hardware, firmware or any combination of software, hardware, and firmware. The order of any steps described above is merely exemplary and the steps of the methods of the present disclosure need not be performed in the order described unless otherwise specified. Furthermore, in some embodiments, the present disclosure can also be implemented as a program for running on a computer or a processor to implement the methods according to the present disclosure. Thus, the present disclosure also covers a record medium storing the program in a non-transitory manner. The program can be realized in any of the following forms: an object code, a code composed of a program language that can be interpreted by a computer, or a code composed of a language that can be converted into a machine language by an interpreter or the like.
[0160] It is also noted that the methods of the present disclosure can be implemented by a computer or a processor of a computer. Furthermore, the present disclosure also covers a computer program for running on a computer or a processor to implement the methods according to the present disclosure.
[0161] The above description of disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0162] The above description has been presented for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to forms disclosed herein. Although various example aspects and embodiments have been discussed above, those of skill in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for estimating the depth of a target object, comprising: Feature extraction is performed on images of the target object captured by the camera to obtain a feature map; Based on the feature map, determine the first plane coordinate value of the target object in the image in the image coordinate system; Based on the feature map and the first planar coordinate values, multiple first depth coordinate values of the target object in the local three-dimensional coordinates of the camera and the depth uncertainty corresponding to each first depth coordinate value are determined. Based on the plurality of first depth coordinate values and the depth uncertainty corresponding to each first depth coordinate value, the second depth coordinate value of the target object in the local three-dimensional coordinates is determined; Based on the second depth coordinate value and the camera's intrinsic parameter matrix, the first planar coordinate value is subjected to an inverse projection transformation from the image coordinate system to the local three-dimensional coordinate system; Determine the second plane coordinates of the target object in the local three-dimensional coordinate system.
2. The method according to claim 1, wherein, The step of determining multiple first depth coordinates of the target object in the local three-dimensional coordinates of the camera, and the depth uncertainty corresponding to each first depth coordinate, based on the feature map and the first planar coordinate values, includes: Based on the feature map and the first plane coordinate value, multiple predicted values corresponding to the target object are determined, as well as the prediction uncertainty corresponding to each predicted value and the actual height prediction value of the target object. Based on the multiple predicted values and the actual height predicted value, multiple first depth coordinate values under the local three-dimensional coordinates are determined; Based on the multiple prediction uncertainties, the depth uncertainty corresponding to each of the first depth coordinate values is determined.
3. The method according to claim 2, wherein, The step of determining multiple predicted values corresponding to the target object and the prediction uncertainty corresponding to each predicted value, as well as the predicted actual height of the target object based on the feature map and the first planar coordinate values, includes: Based on multiple prediction branch models, the feature map is predicted to determine the candidate prediction result for each object in at least one object in the image; Based on the first planar coordinate values, the candidate prediction result corresponding to the target object is determined from the at least one candidate prediction result; Based on the candidate prediction results corresponding to the target object, the plurality of predicted values and the prediction uncertainty corresponding to each predicted value, as well as the predicted actual height of the target object, are obtained.
4. The method according to claim 3, wherein, The multiple predicted values include at least two of the following: the direct depth prediction value in the local three-dimensional coordinates, the multiple key point prediction values in the local three-dimensional coordinates, and the projected height prediction value in the local three-dimensional coordinates; The determination of multiple first depth coordinate values in the local three-dimensional coordinate system based on the multiple predicted values and the actual height predicted value includes at least two of the following: The direct depth prediction value is used as a first depth coordinate value; The projected height of the target object in the image is determined based on the predicted values of the multiple key points, and a first depth coordinate value is determined based on the projected height value, the actual height prediction value, and the camera's intrinsic parameter matrix. A first depth coordinate value is determined based on the projected height prediction value, the actual height prediction value, and the camera's intrinsic parameter matrix.
5. The method according to any one of claims 1-4, wherein, The step of determining the second depth coordinate value of the target object in the local three-dimensional coordinates based on the plurality of first depth coordinate values and the uncertainty corresponding to each first depth coordinate value includes: Based on the multiple uncertainties, a weight value is determined for each of the first depth coordinate values; The first depth coordinates are weighted by multiple weight values to obtain the second depth coordinates of the target object in the camera coordinates corresponding to the camera.
6. The method according to any one of claims 1-5, further comprising: Based on the feature map, the center offset value of the target object in the image coordinate system is determined; The second plane coordinate value is corrected based on the center offset value to obtain the corrected third plane coordinate value.
7. A depth estimation device for a target object, comprising: The feature extraction module is used to extract features from images captured by the camera, including the target object, to obtain a feature map; A planar coordinate prediction module is used to determine the first planar coordinate value of the target object in the image in the image coordinate system based on the feature map obtained by the feature extraction module; The object coordinate estimation module is used to determine, based on the feature map obtained by the feature extraction module and the first planar coordinate value determined by the planar coordinate prediction module, multiple first depth coordinate values of the target object in the local three-dimensional coordinates of the camera and the depth uncertainty corresponding to each first depth coordinate value; The target object estimation module is used to determine the second depth coordinate value of the target object in the local three-dimensional coordinates based on multiple first depth coordinate values determined by the object coordinate estimation module and the depth uncertainty corresponding to each first depth coordinate value; The inverse projection module is used to perform an inverse projection transformation on the first planar coordinate values from the image coordinate system to the local three-dimensional coordinate system based on the second depth coordinate values and the intrinsic parameter matrix of the camera; and to determine the second planar coordinate values of the target object in the local three-dimensional coordinate system.
8. A computer-readable storage medium storing a computer program for performing the depth estimation method for a target object according to any one of claims 1-6.
9. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the depth estimation method for the target object according to any one of claims 1-6.
Citation Information
Patent Citations
Underwater monocular vision target depth positioning fusion estimation method based on deep learning
CN111915678A
Image depth estimation method, terminal equipment and computer readable storage medium
CN112070817A