Accurate depth estimation method for dynamic target
By combining object detection, multi-frame image fusion and deep learning technology, an accurate depth estimation method for dynamic targets is designed, which solves the problem of inaccurate depth estimation in dynamic scenarios, and achieves a high-precision and low-complexity real-time depth estimation effect.
Patent Information
- Application Number
- CN202510296950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to accurately capture the depth information of dynamic targets in dynamic scenarios, resulting in large estimation errors. Especially in fast moving targets or complex backgrounds, there are problems such as motion blur, multi-objective interference, high computational complexity and error accumulation.
Using a method combining object detection, multi-frame image fusion and deep learning technology, the position and motion information of dynamic targets are obtained through object detection algorithms and motion tracking technology, multi-frame image fusion is performed using optical flow method or feature point matching method, and a deep learning model based on convolutional neural network is designed for depth estimation, and error correction is performed based on the target's motion characteristics.
It significantly improves the accuracy and robustness of dynamic target depth estimation, reduces the computational complexity, and meets the needs of real-time applications.
Smart Images

Figure CN120219459A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an accurate depth estimation method for dynamic targets, which is applicable to fields such as autonomous driving, robot navigation, and augmented reality. Background Art
[0002] Depth estimation is a key technology in the field of computer vision and is widely used in fields such as autonomous driving, robot navigation, and augmented reality. Traditional methods mainly target static scenes and obtain depth information through monocular or binocular cameras. However, in dynamic scenes, due to the movement of target objects, it is difficult for traditional methods to accurately capture the depth information of dynamic targets, resulting in large estimation errors.
[0003] Existing dynamic target depth estimation methods rely on motion estimation and multi-frame image fusion techniques. Motion estimation infers depth by tracking the target motion trajectory, and multi-frame image fusion uses time series information to improve accuracy. However, these methods face the following challenges when dealing with fast-moving targets or complex backgrounds: First, motion blur: Fast movement causes image blur, affecting the accuracy of target detection and tracking; Second, multi-target interference: The mutual interference of multiple dynamic targets, especially when targets overlap or occlude, increases the difficulty of depth estimation; Third, high computational complexity: Existing methods require a large amount of computing resources and are difficult to meet real-time requirements; Fourth, error accumulation: In a long time series, errors will accumulate, resulting in a larger deviation of the results. In recent years, deep learning has made progress in tasks such as target detection and depth estimation and can automatically learn features from data to overcome some limitations of traditional methods. However, existing deep learning models still face problems such as difficult data annotation and insufficient generalization ability when dealing with dynamic target depth estimation.
[0004] In summary, there are many challenges in existing technologies for dynamic target depth estimation. There is an urgent need for a method that combines target detection, multi-frame image fusion, and deep learning technologies to improve accuracy and robustness, while reducing computational complexity to meet the requirements of real-time applications. Summary of the Invention
[0005] The present invention aims to provide an accurate depth estimation method for dynamic targets to solve the problem of inaccurate depth estimation in existing technologies in dynamic scenarios. To this end, the present invention adopts the following technical solutions.
[0006] An accurate depth estimation method for dynamic targets, the method comprising the following steps:
[0007] S1: Dynamic Target Detection and Tracking
[0008] Detect dynamic objects in the image through object detection algorithms (such as YOLO or Faster R-CNN), and use motion tracking techniques (such as Kalman filtering or SORT algorithm) to obtain the position and motion information of the objects.
[0009] S2: Multi-frame image fusion
[0010] Utilize time series information to fuse the target motion trajectories in multi-frame images through optical flow method or feature point matching method, enhancing the robustness of depth estimation.
[0011] S3: Construction of depth estimation model
[0012] Design a deep learning model based on convolutional neural network (CNN), and combine the object detection results and multi-frame image fusion information to perform depth estimation of dynamic objects.
[0013] S4: Error correction and optimization
[0014] According to the motion characteristics of the object (such as speed and direction), dynamically correct the depth estimation result to reduce the estimation error caused by the object motion.
[0015] S5: Depth estimation of dynamic objects
[0016] Use the trained model to perform real-time depth estimation of dynamic objects, which is applicable to fields such as autonomous driving, robot navigation, and augmented reality. Description of the drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present invention and do not limit the present invention.
[0018] Figure 1 : Schematic diagram of the process of an accurate depth estimation method for dynamic objects;
[0019] Figure 2 : Structure diagram of a depth estimation model based on convolutional neural network (CNN). Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the following further details the present invention in combination with specific embodiments and accompanied by drawings. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0021] The flowchart and structure diagram of the present invention are respectively as Figure 1, Figure 2 As shown, a precise depth estimation method for dynamic targets includes the following steps:
[0022] S1: Dynamic target detection and tracking are the basis of depth estimation. First, use a deep learning-based object detection algorithm (such as YOLO or Faster R-CNN) to detect dynamic targets in the input image and obtain the position information of the targets. Then, use motion tracking techniques (such as Kalman filtering or SORT algorithm) to track the targets and obtain their motion trajectories. The specific implementation formula is as follows:
[0023]
[0024] where, I i is the i-th frame of the input image, B i is the detected target bounding box, and T i is the motion trajectory of the target.
[0025] S2: Multi-frame image fusion enhances the robustness of depth estimation through time-series information. Use optical flow method or feature point matching method to fuse the target motion information in multiple frames. The specific implementation formula is as follows:
[0026]
[0027] where, F i is the optical flow information, and M i is the feature point matching information.
[0028] S3: Design a deep learning model based on convolutional neural network (CNN), combine the object detection results and multi-frame image fusion information, and perform depth estimation of dynamic targets. The input of the model is the object detection results and multi-frame image fusion information, and the output is the depth value of the target. The specific implementation formula is as follows:
[0029] D i = CNN(B i , F i , M i ) (7)
[0030] where, D i is the depth estimation value of the i-th frame target.
[0031] S4: According to the motion characteristics of the target (such as speed and direction), dynamically correct the depth estimation results to reduce the estimation error caused by target motion. The specific implementation formula is as follows:
[0032] D' i = D i + α·Δv i (8)
[0033] Among them, D' i is the corrected depth value, Δv i is the change in the motion speed of the target, and α is the correction coefficient.
[0034] S5: Use the trained model to perform real-time depth estimation on dynamic targets, which is applicable to fields such as autonomous driving, robot navigation, and augmented reality. The specific implementation formula is as follows:
[0035]
[0036] Among them, is the final depth estimation result.
[0037] The present invention significantly improves the accuracy and robustness of dynamic target depth estimation by combining object detection, multi-frame image fusion, and deep learning technologies, while reducing the computational complexity to meet the requirements of real-time applications.
[0038] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for accurate depth estimation of dynamic targets, characterized in that: The method comprises the following steps: S1: Dynamic target detection and tracking. The position and motion information of dynamic targets are obtained through target detection algorithms and motion tracking technology. S2: Multi-frame image fusion. Use time series information to fuse the target motion trajectory in multiple frames to enhance the robustness of depth estimation. S3: Depth estimation model construction. Design a neural network model based on deep learning, and combine target detection and multi-frame image fusion information for depth estimation. S4: Error correction and optimization. Dynamically correct the depth estimation result according to the motion characteristics of the target to reduce the estimation error. S5: Dynamic target depth estimation. Use the trained model to perform real-time depth estimation of dynamic targets.
2. The method for accurate depth estimation of dynamic targets according to claim 1, characterized in that: In step S1, dynamic target detection and tracking uses a deep learning-based target detection algorithm (such as YOLO or Faster R-CNN) to perform target detection, and uses a Kalman filter or SORT algorithm to perform target tracking to obtain the precise position and motion trajectory of the dynamic target.
3. The method for accurate depth estimation of dynamic targets according to claim 1, characterized in that: In step S2, the multi-frame image fusion is performed by using an optical flow method or a feature point matching method to fuse the target motion information in the multi-frame images, thereby enhancing the robustness and accuracy of the depth estimation.
4. The method for accurate depth estimation of dynamic targets according to claim 1, characterized in that: In step S3, the depth estimation model is constructed using a deep learning model based on a convolutional neural network (CNN), the input is the target detection result and multi-frame image fusion information, and the output is the depth value of the dynamic target.
5. The method for accurate depth estimation of dynamic targets according to claim 1, characterized in that: In step S4, error correction and optimization are dynamically corrected by the following loss function: L 校正 =λ1·L 运动 +λ2·L 深度 (1) Among them, L 运动 It is a loss function based on the target motion characteristics and is calculated by the following formula: L 深度 is a loss function based on the depth estimation error and is calculated using the following formula: Among them, v l is the actual speed of the target, is the velocity predicted by the model, d l is the actual depth value of the target, is the depth value predicted by the model, N is the number of samples, and λ1 and λ2 are hyperparameters used to balance the weights of the two parts of the loss.
6. The method for accurate depth estimation of dynamic targets according to claim 1, characterized in that: In step S5, the dynamic target depth estimation uses the trained model to perform real-time depth estimation on the dynamic target, which is applicable to fields such as autonomous driving, robot navigation and augmented reality. The dynamic target depth estimation performs real-time depth calculation using the following formula: Among them, I is the input image, M is the multi-frame image fusion information, and f 模型 is a trained depth estimation model.