Method and apparatus for depth estimation of moving objects, electronic device, and storage medium
By determining the video processing type and selecting an appropriate depth estimation method in the SLAM system, the problem of depth estimation for dynamic objects is solved, achieving accurate depth estimation of moving objects in video frames and improving the user experience.
Patent Information
- Application Number
- CN202211160924.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-09-22
AI Technical Summary
In existing technologies, SLAM systems struggle to effectively estimate the depth of dynamic objects in videos, and can only estimate the depth information of static objects.
By determining the video processing type and selecting real-time or post-processing methods based on the type, depth mean estimation or inverse depth estimation methods are used to accurately estimate the depth information of real-time captured videos or existing videos.
It enables accurate estimation of the depth information of moving objects in video frames, improves the applicability of depth estimation, meets users' personalized needs, and enhances the user experience.
Smart Images

Figure CN117788542B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more particularly to a method, apparatus, electronic device, and storage medium for estimating the depth of a moving object. Background Technology
[0002] With the development of computer vision technology, Simultaneous Localization and Mapping (SLAM) algorithms have been widely used in augmented reality, virtual reality, autonomous driving, and the positioning and navigation of robots or drones.
[0003] In existing technologies, images are input into a SLAM system, and the SLAM system is used to extract scene depth information from the image in order to estimate the depth of objects in the image based on the scene depth information. However, this depth estimation method is only applicable to static objects, and it is difficult to achieve effective depth estimation for dynamic objects in videos. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for estimating the depth of a moving object, so as to achieve the effect of accurately estimating the depth information of a moving object in a video.
[0005] In a first aspect, embodiments of this disclosure provide a depth estimation method for a moving object, the method comprising:
[0006] Determine the video processing type;
[0007] Based on the video processing type, determine the target processing method for depth estimation of the moving object;
[0008] Based on the target processing method, the depth estimate of the moving object in the video frame to be processed is determined.
[0009] Secondly, embodiments of this disclosure also provide a depth estimation device for a moving object, the device comprising:
[0010] The video processing type determination module is used to determine the video processing type.
[0011] The target processing method determination module is used to determine the target processing method for depth estimation of the moving object based on the video processing type.
[0012] The depth estimation determination module is used to determine the depth estimate of a moving object in the video frame to be processed based on the target processing method.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the depth estimation method for a moving object as described in any embodiment of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a depth estimation method for a moving object as described in any of the embodiments of this disclosure.
[0018] The technical solution of this disclosure, by determining the video processing type, further determining the target processing method for depth estimation of moving objects based on the video processing type, and finally determining the depth estimate value of the moving object in the video frame to be processed based on the target processing method, solves the problem that the prior art can only estimate the depth information of static objects, achieves the effect of accurately estimating the depth information of moving objects in video frames, and improves the applicability of depth estimation, meets the personalized needs of users, and enhances the user experience. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1 This is a schematic flowchart of a depth estimation method for a moving object provided in an embodiment of this disclosure;
[0021] Figure 2 This is a schematic flowchart of a depth estimation method for a moving object provided in an embodiment of this disclosure;
[0022] Figure 3 This is a schematic diagram of the structure of a depth estimation device for a moving object provided in an embodiment of this disclosure;
[0023] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0030] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0031] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0032] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0034] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0035] Before introducing this technical solution, an exemplary application scenario can be provided. For example, when a user uses a mobile camera to capture video and uploads it to a SLAM-based system, or selects a target video from a database and actively uploads it to the SLAM-based system, the system can analyze the scene depth information in the video to estimate the depth information of objects contained in the video frame. However, current depth estimation methods can only estimate the depth information of static objects in the video frame and cannot accurately estimate the depth of dynamic objects. In this case, the solution based on this embodiment can utilize the scene depth information and 3D spatial information provided by the SLAM system to estimate the depth information of moving objects in the video frame, thereby achieving the effect of accurately estimating the depth information of dynamic objects in the video frame.
[0036] Figure 1 This is a schematic flowchart of a depth estimation method for a moving object provided in an embodiment of this disclosure. This embodiment of the disclosure is applicable to the situation of estimating the depth information of a moving object in a video frame. The method can be executed by a depth estimation device for a moving object. The device can be implemented in the form of software and / or hardware, or optionally, by an electronic device, such as a mobile terminal, a PC, or a server.
[0037] like Figure 1 As shown, the method includes:
[0038] S110. Determine the video processing type.
[0039] In this embodiment, the apparatus for executing the image rendering method provided in this disclosure can be integrated into application software that supports special effects video processing functions. This software can be installed on an electronic device, optionally a mobile terminal or a PC. The application software can be a type of software for image / video processing; specific application software will not be detailed here, as long as it can perform image / video processing. Alternatively, it can be a specially developed application program to add and display special effects, or it can be integrated into a corresponding page, allowing users to process special effects videos through the integrated page on a PC.
[0040] It should be noted that the technical solution of this embodiment can be executed during real-time video recording on a mobile device, or after the system receives video data actively uploaded by the user. In addition, the solution of this embodiment can be applied to various application scenarios such as augmented reality (AR), virtual reality (VR), and autonomous driving.
[0041] In this embodiment, the video processing type can be determined based on the user's upload method of the video to be processed. Optionally, the video processing type includes real-time processing type and post-processing type. In practical applications, if the video to be processed is captured in real-time by the user using a mobile camera device, and depth estimation is performed on the moving objects contained in the video to be processed based on the mobile device, the current processing type of the video to be processed can be taken as the real-time processing type; if the video to be processed is a video that has already been captured, and is actively uploaded to the system by the user, then depth estimation on the moving objects contained in the received video to be processed can be a post-processing type.
[0042] In practical implementation, if the video data received by the system is captured in real-time by a mobile camera device, the video processing type can be determined as real-time processing; if the system receives complete video data that has already been captured, the video processing type can be determined as post-processing. The advantage of this setting is that it enhances the diversity of moving object depth estimation processing methods, allowing depth estimation of moving objects in the video frame to be processed to be performed both in real-time on the mobile device and in the complete video, thus improving the diversity of video processing and meeting the personalized needs of users.
[0043] S120. Determine the target processing method for depth estimation of moving objects based on the video processing type.
[0044] In this embodiment, when a user triggers a special effect operation, the mobile camera device can face the user in real time to capture the video to be processed. The video is then parsed according to a pre-written program to obtain multiple video frames to be processed. At this point, the video processing type can be determined as real-time processing. Correspondingly, the video frames to be processed may include moving objects. Moving objects can be any object in the frame whose posture or position changes, such as a user or an animal.
[0045] Those skilled in the art will understand that depth estimation can be a subtask within the field of computer vision. Its purpose is to obtain the distance between an object and the shooting point, providing depth information for a range of tasks such as 3D reconstruction, distance perception, SLAM, visual odometry, video interpolation, and image reconstruction. The depth information of a moving object can be the distance between the pixel corresponding to the moving object and the shooting point in the final displayed image, or it can be represented by the position coordinates of each pixel in the camera coordinate system.
[0046] In this embodiment, when the video processing type is determined to be real-time processing, the target processing method for depth estimation of moving objects in the video frame can be the depth mean estimation method corresponding to the real-time processing type. Specifically, the depth mean estimation method involves determining the depth values of some pixels associated with the moving object and averaging these depth values, thereby using the final average depth value as the depth information of the moving object.
[0047] S130. Based on the target processing method, determine the depth estimate of the moving object in the video frame to be processed.
[0048] In this embodiment, the user can capture video of a moving object in real time using the camera device on a mobile terminal and upload it to the mobile terminal in real time. Therefore, it can be understood that the real-time captured video obtained by the system is the video to be processed. Furthermore, based on a pre-written program, the video to be processed is parsed to obtain multiple video frames to be processed. The depth estimation value can be the distance between at least one pixel corresponding to the moving object and the shooting point, or it can be the coordinate value of at least one pixel corresponding to the moving object in the camera coordinate system.
[0049] In this embodiment, the target processing method can be the depth mean estimation method. To determine the depth estimate of the moving object based on the target processing method, the target pixels in the moving object that meet the depth mean estimation conditions can be identified first. Then, the depth mean can be determined based on the depth values of these target pixels, so that the final depth mean can be used as the depth estimate of the moving object.
[0050] Optionally, based on the target processing method, the depth estimate of the moving object in the video frame to be processed is determined, including: determining the shooting parameters corresponding to the video frame to be processed and the pixel parameters of the moving object; determining the target pixel based on the shooting parameters, pixel parameters and constraints; and determining the depth estimate of the moving object based on the point cloud data of the target pixel.
[0051] In this embodiment, the shooting parameters can be the camera pose parameters of the video frame to be processed after pose optimization. It should be noted that camera position and rotation information can be obtained based on the gyroscope and inertial measurement unit in the camera device corresponding to the video frame to be processed. Based on the camera position and rotation information, the initial pose of the camera can be determined. Furthermore, the initial pose is optimized using bundle adjustment, and the optimized pose is used as the shooting parameters corresponding to the video frame to be processed. The advantage of this setup is that it allows the synchronous localization and mapping system to provide a high BA speed, thereby ensuring the real-time processing of each video frame. The pixel parameters can be the pixel coordinates of at least one pixel in the video frame to be processed that constitutes the moving object. It should be noted that when shooting a moving object to obtain multiple video frames to be processed, the video frames may contain not only the moving object but also the scene in which the moving object is located. Therefore, when determining the pixel parameters of the moving object, a mask image of the moving object can be determined first, and the pixel coordinates of at least one pixel constituting the moving object can be determined based on this mask image.
[0052] In this embodiment, the constraint can be a spatial geometric information constraint. That is, when a pixel is observed at a specific location, it is determined whether the state of the pixel corresponds to the specific location. If the state of the pixel corresponds to its observation location, it can be determined that the pixel satisfies the constraint; if the state of the pixel does not correspond to its observation location, it can be determined that the pixel does not satisfy the constraint.
[0053] In practical implementation, after obtaining the video frame to be processed, the initial pose of the video frame can be determined based on the parameters of each sensor of the camera device corresponding to the video frame. Then, the initial pose is optimized using a pose optimization method to obtain the shooting parameters corresponding to the video frame. Simultaneously, the pixel coordinates of the moving object in the video frame are determined as pixel parameters. Further, based on the shooting parameters, pixel parameters, and constraints, the target pixel is determined. Therefore, the depth estimate of the moving object can be determined based on the point cloud data of the target pixel. The advantage of this setup is that it allows the pixels of the moving object to be divided into dynamic and static pixels based on constraints, and the dynamic pixels are selected as tracking pixels for the moving object, improving the accuracy of the depth estimate and enhancing the localization effect of the moving object in the video frame.
[0054] In practical applications, the initial pose of the video frame to be processed can be determined first, and the initial pose can be optimized based on the pose optimization method to obtain the shooting parameters corresponding to the video frame to be processed. At the same time, the pixel coordinates of at least one pixel corresponding to the moving object can be determined to obtain the pixel parameters. Further, based on the shooting parameters, pixel parameters and constraints, the pixels that satisfy the constraints among the pixels corresponding to the moving object can be determined, and these pixels can be used as target pixels.
[0055] Optionally, the target pixel is determined based on the shooting parameters, pixel parameters, and constraints, including: performing triangulation processing based on the shooting parameters and pixel parameters to obtain point cloud data corresponding to the pixel parameters; determining the back-projected pixel parameters based on the point cloud data and constraints; and determining the target pixel based on the pixel parameters and back-projected pixel parameters.
[0056] In this embodiment, triangulation can be used to determine the corresponding point cloud data based on a corner detection algorithm. The corner detection algorithm can be the KLT corner detection method, also known as the KLT optical flow tracing method. Specifically, the KLT corner detection method determines a suitable reference keyframe for tracking in each keyframe and identifies the feature points of that reference keyframe, thereby determining the corresponding point cloud data (PCD) based on these feature points. Point cloud data is commonly used in reverse engineering and is data recorded in the form of points. These points can be coordinates in three-dimensional space, or information such as color or illumination intensity. In practical applications, point cloud data generally also includes point coordinate accuracy, spatial resolution, and surface normal vectors, and is generally saved in PCD format. This format offers greater operability and can improve the speed of point cloud registration and fusion in subsequent processes, which will not be elaborated further in this embodiment.
[0057] In practical applications, after determining the shooting parameters and pixel parameters, triangulation can be performed on these parameters using a corner detection algorithm to obtain 3D point cloud data corresponding to the pixel parameters. Further, based on the point cloud data and constraints, the parameters of this point cloud data in the camera coordinate system are determined; that is, the 3D point cloud data is converted into 2D coordinates. These converted 2D coordinate parameters can be used as back-projected pixel parameters. Since both the point cloud data and the back-projected pixel parameters are determined based on the point cloud data, and both are 2D coordinate parameters, the target pixel can be determined by checking whether the pixel parameters match the corresponding back-projected pixel parameters. In other words, pixels whose pixel parameters do not match the corresponding back-projected pixel parameters are taken as target pixels. It should be noted that the pixels of the moving object are determined based on the mask image. In practical applications, a model deployed on the mobile device is typically used to process the video frame to obtain a mask image corresponding to the moving object. Generally, to improve the processing efficiency of the mobile device and reduce the model's memory footprint, the model deployed on the mobile device is usually a simple model with a fast processing speed. When applying this model to process the moving object mask image of the video frame, the resulting mask image may be larger than the actual size of the moving object, thus including static background points that do not belong to the moving object. Those skilled in the art should understand that static pixels generally satisfy the constraints, while dynamic pixels do not. Therefore, by determining whether the pixels corresponding to the moving object satisfy the constraints, dynamic and static pixels can be distinguished so that different processing methods can be applied to different pixels, ultimately obtaining the depth estimate of the moving object. The advantage of this setting is that it allows for more precise determination of the pixels of a moving object, enabling different processing methods to be applied to different pixels, thereby improving the accuracy of the depth estimation of the moving object.
[0058] For example, based on point cloud data and constraints, the back-projected pixel parameters can be determined using the following formula:
[0059]
[0060] Among them, s i It can represent the depth value of any pixel, (u i v i ) can represent the pixel coordinates of any pixel, K can represent the camera intrinsic parameters, and exp(ξ^) can represent the camera pose, i.e., the R and T matrices. i Y i Zi () can represent the three-dimensional point cloud coordinates of any pixel.
[0061] Furthermore, after determining the target pixel, the depth estimate of the moving object can be determined based on the point cloud data of the target pixel.
[0062] Optionally, the depth estimate of the moving object is determined based on the point cloud data of the target pixel, including: determining at least two video frames to which the target pixel belongs based on the point cloud data of the target pixel; and determining the depth estimate of the moving object based on the depth values of the target pixel in the at least two video frames to be used.
[0063] In this embodiment, after obtaining the target pixels, these target pixels can be triangulated to obtain point cloud data corresponding to the target pixels. Furthermore, the point cloud data is observed in multiple video frames to be processed that contain moving objects, and at least two video frames to be processed in which point cloud data can be observed are selected as video frames to be used.
[0064] In practical applications, after determining at least two video frames to which the target pixel belongs, the depth value of the target pixel in the camera coordinate system can be determined. These depth values are then averaged, and the resulting average depth value can be used as the depth estimate of the moving object. The advantage of this setup is that it allows for a rough estimation of the depth information of a moving object on a mobile device, improving the efficiency of depth estimation for moving objects.
[0065] It should be noted that if the moving object is stationary, then each pixel of the moving object determined based on the mask image satisfies the constraint condition, that is, the pixel parameters of each pixel are consistent with the back-projected pixel parameters. In this case, these pixels can be triangulated to obtain the point cloud data corresponding to these pixels, and these point cloud data can be stored in the SLAM system so that the depth estimate of the moving object can be determined through the SLAM system.
[0066] It should also be noted that this embodiment determines the depth estimate of a moving object in a video frame to be processed when the video processing type is real-time processing. Based on this embodiment, when the video processing type is post-processing, the corresponding target processing method will also change accordingly. The post-processing type will be described in detail below.
[0067] The technical solution of this disclosure, by determining the video processing type, further determining the target processing method for depth estimation of moving objects based on the video processing type, and finally determining the depth estimate value of the moving object in the video frame to be processed based on the target processing method, solves the problem that the prior art can only estimate the depth information of static objects, achieves the effect of accurately estimating the depth information of moving objects in video frames, and improves the applicability of depth estimation, meets the personalized needs of users, and enhances the user experience.
[0068] Figure 2 This is a schematic flowchart of a depth estimation method for a moving object provided in this embodiment. Based on the foregoing embodiments, when the video processing type is post-processing, the corresponding target processing method can be an inverse depth estimation method. Therefore, the depth estimate of the moving object can be determined based on the inverse depth estimation method. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0069] like Figure 2 As shown, the method specifically includes the following steps:
[0070] S210. Determine the video processing type as post-processing type.
[0071] It should be noted that the above embodiments determine the depth estimation value of moving objects in the video frame to be processed when the video processing type is real-time processing. Based on the above embodiments, when the video processing type is post-processing, the corresponding target processing method will also change accordingly. The post-processing type will be described in detail below.
[0072] In this embodiment, a video upload control can be pre-developed. When a user's trigger operation on the video upload control within the application is detected, the user-uploaded video can be received and used as a video to be processed. Further, based on a pre-written program, the video to be processed is parsed to obtain multiple video frames to be processed. Correspondingly, each video frame to be processed contains a moving object, which can be a user, an animal, or any object in the frame whose posture or position changes. When a complete video to be processed is received, the video frames containing moving objects can be used as video frames to be processed, and these video frames can be processed with special effects to obtain corresponding special effects video frames. This video processing method can be used as a post-processing type.
[0073] S220. Based on the post-processing type, the target processing method for depth estimation of moving objects is determined to be the inverse depth estimation method.
[0074] In this embodiment, after receiving the video to be processed and determining that the video processing type is post-processing, the target processing method for depth estimation of moving objects in the video frame to be processed can be determined as inverse depth estimation. Specifically, inverse depth estimation can determine the depth estimate of the moving object based on the inverse depth value of at least one pixel corresponding to the moving object.
[0075] It's important to note that when the video processing method is post-processing, meaning the depth of a moving object is estimated from the complete video data, unlike real-time processing, where the depth information of each pixel in each frame can be determined after receiving the complete video data, and the depth of the moving object is estimated based on this information, the distribution range of the depth information of each pixel in each frame is large and the distribution is unstable. Therefore, inverse depth information corresponding to this depth information can be determined to estimate the depth of the moving object. The advantage of this approach is that the inverse depth distribution is more consistent with a Gaussian distribution, making it more stable and thus resulting in a more accurate depth estimate of the moving object.
[0076] It should also be noted that each video frame to be processed includes both distant and near pixels. For distant pixels, due to the greater distance between them and the shooting point, the parallax of these pixels is smaller. Consequently, the accuracy of the point cloud data will be lower when determining the point cloud data corresponding to these distant pixels. Therefore, an inverse depth approach can be used to reduce the impact of distant pixels on the calculation process. The depth values of distant and near pixels are converted into inverse depth values, and subsequent calculations can be performed based on these inverse depth values, thereby improving the calculation accuracy.
[0077] S230. Based on the inverse depth estimation method, determine the depth estimate of the moving object in the video frame to be processed.
[0078] In this embodiment, after determining that the target processing method is the inverse depth estimation method, the inverse depth value of each pixel in the video frame to be processed can be determined, so that the depth estimate of the moving object can be determined based on these inverse depth values.
[0079] Optionally, based on the inverse depth estimation method, the depth estimate of the moving object in the video frame to be processed is determined, including: performing triangulation processing on each video frame to be processed in the target video to obtain the inverse depth value of each pixel in the video frame to be processed; and determining the depth estimate of the moving object by clustering the inverse depth values in the same video frame to be processed.
[0080] In this embodiment, the target video can be a user-uploaded video that requires the determination of the depth information of moving objects in the video. In practical applications, upon receiving each video frame to be processed from the target video, triangulation can be performed on each frame based on a corner detection algorithm to obtain point cloud data corresponding to each frame. Further, the point cloud data corresponding to each frame can be transformed to the camera coordinate system using a translation and rotation matrix to obtain the depth value of each pixel in the camera coordinate system. Then, these depth values are inverted; that is, the negative first power of each depth value is determined to obtain the inverse depth value of each pixel. Thus, by clustering the inverse depth values in the same video frame, the depth estimate of the moving object can be determined. The advantage of this setup is that estimating the depth of a moving object based on the inverse depth value of each pixel can reduce the influence of distant pixels on the depth estimation, thereby improving the accuracy of the depth estimation and enhancing the display effect of the moving object's fixed points at different timestamps in the target video.
[0081] Clustering can be used to classify each inverse depth value, or it can be binary classification, that is, dividing each inverse depth value into two major categories.
[0082] Optionally, the depth estimate of the moving object is determined by clustering the inverse depth values in the same video frame to be processed, including: sorting the inverse depth values according to their magnitude and determining the depth difference between two adjacent inverse depth values; obtaining the two target inverse depth values with the largest depth difference and determining the depth estimate of the moving object based on the inverse depth values that are greater than the target inverse depth values.
[0083] In practical applications, for each inverse depth value in the same video frame to be processed, the magnitude of each inverse depth value can be determined first, and then sorted according to their magnitude. Next, the difference between two adjacent inverse depth values is determined as the depth difference, and the two adjacent inverse depth values corresponding to the largest depth difference are identified as target inverse depth values. Further, based on these two target inverse depth values, each inverse depth value can be divided into two categories: one category containing inverse depth values greater than the target inverse depth value, and the other category containing inverse depth values less than the target inverse depth value. Finally, the depth estimate of the moving object can be determined based on the inverse depth values greater than the target inverse depth value. The advantage of this setup is that it allows for the classification of near-field and far-field pixels based on the inverse depth values, thus enabling the determination of the depth information of the moving object based on the depth information of the near-field pixels.
[0084] It should be noted that when classifying each inverse depth value based on the target inverse depth value, if the number of inverse depth values in any class is less than a preset threshold, it can be considered that the inverse depth values in this class may have some error. In order to improve the accuracy of the depth estimation of the moving object, these inverse depth values can be deleted, and the remaining inverse depth values can be sorted and classified again. After the reclassification is completed, the depth estimation of the moving object can be determined based on the inverse depth values that are greater than the target inverse depth value in this classification result.
[0085] Based on this, before determining the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value, the method further includes: if the ratio between the number of inverse depth values greater than or less than the target inverse depth value and the total number of inverse depth values is less than a preset ratio, then the corresponding inverse depth value is deleted, and the step of determining the target inverse depth value is re-executed.
[0086] In this embodiment, the preset ratio can be any value, and optionally, it can be 5%.
[0087] In practical applications, after dividing the inverse depth values into those greater than the target inverse depth value and those less than the target inverse depth value, the ratio between the number of inverse depth values in these two categories and the total number of inverse depth values in the current video frame to be processed can be determined. If the ratio of either category is less than a preset proportion, the inverse depth values in that category can be deleted, and the remaining inverse depth values can be reordered based on their magnitude. Then, the difference between two adjacent inverse depth values is determined, and the two inverse depth values with the largest difference are taken as the target inverse depth values. Further, the remaining inverse depth values are classified based on the target inverse depth values, so that the depth estimate of the moving object can finally be determined based on the inverse depth values greater than the target inverse depth value. The advantage of this setting is that it can filter out and delete inverse depth values with large errors, thereby improving the accuracy of the depth estimate of the moving object.
[0088] Optionally, determining the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value includes: averaging each inverse depth value greater than the target inverse depth value to determine the depth estimate of the moving object.
[0089] It should be noted that after obtaining each inverse depth value that is greater than the target inverse depth value, since the pixels corresponding to these inverse depth values are the near-field pixels of the video frame to be processed, those skilled in the art should understand that when calculating based on the near-field pixels, a more accurate calculation result can be obtained. Furthermore, for moving objects, they are generally located in the foreground part of the video frame to be processed. Therefore, when determining the depth estimate of a moving object, calculating based on each inverse depth value that is greater than the target inverse depth value can yield a more accurate depth estimate result.
[0090] In practical applications, the inverse depth values greater than the target inverse depth value can be averaged, and the resulting average inverse depth value can be inverted again to obtain the corresponding average depth value. This average depth value can then be used as the depth estimate of the moving object. The advantage of this setup is that determining the depth information of the moving object based on the depth information of nearby pixels can improve the accuracy of depth estimation.
[0091] It should be noted that for any video frame to be processed in the target video, the above-mentioned technical method can be used to determine the depth estimate of the moving object in each video frame. Then, after obtaining the depth estimate of the moving object in each video frame to be processed, the video frames to be processed can be stitched together to obtain the depth estimate of the moving object in the complete target video.
[0092] The technical solution of this disclosure, by determining the video processing type as post-processing, further determines the target processing method for depth estimation of moving objects as inverse depth estimation based on the post-processing type, and finally, based on the inverse depth estimation method, determines the depth estimate value of the moving object in the video frame to be processed. This solves the problem in the prior art that depth information can only be estimated for static objects, achieving accurate estimation of the depth information of moving objects in video frames. Furthermore, it improves the applicability of depth estimation, meets users' personalized needs, and enhances the user experience.
[0093] Figure 3 This is a schematic diagram of a depth estimation device for a moving object provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, the device includes: a video processing type determination module 310, a target processing method determination module 320, and a depth estimation value determination module 330.
[0094] Among them, the video processing type determination module 310 is used to determine the video processing type;
[0095] The target processing method determination module 320 is used to determine the target processing method for depth estimation of the moving object based on the video processing type.
[0096] The depth estimation determination module 330 is used to determine the depth estimation value of a moving object in the video frame to be processed based on the target processing method.
[0097] Based on the above technical solutions, the video processing types include real-time processing and post-processing.
[0098] Based on the above technical solutions, the target processing method includes a depth mean estimation method corresponding to the real-time processing type, or an inverse depth estimation method corresponding to the post-processing type.
[0099] Based on the above technical solutions, the target processing method includes a depth mean estimation method, and the depth estimation value determination module 330 includes: a shooting parameter determination submodule, a target pixel point determination submodule, and a depth estimation value determination submodule.
[0100] The shooting parameter determination submodule is used to determine the shooting parameters corresponding to the video frame to be processed and the pixel parameters of the moving object;
[0101] The target pixel determination submodule is used to determine the target pixel based on the shooting parameters, pixel parameters, and constraints.
[0102] The depth estimation determination submodule is used to determine the depth estimate of the moving object based on the point cloud data of the target pixels.
[0103] Based on the above technical solutions, the target pixel determination submodule includes: a point cloud data determination unit, a back-projection pixel parameter determination unit, and a target pixel determination unit.
[0104] A point cloud data determination unit is used to perform triangulation processing based on the shooting parameters and the pixel parameters to obtain the point cloud data corresponding to the pixel parameters.
[0105] The back-projection pixel parameter determination unit is used to determine the back-projection pixel parameters based on the point cloud data and the constraint conditions.
[0106] The target pixel determination unit is used to determine the target pixel based on the pixel parameters and the back-projected pixel parameters.
[0107] Based on the above technical solutions, the depth estimation value determination submodule includes: a video frame determination unit and a depth estimation value determination unit.
[0108] The video frame to be used determination unit is used to determine at least two video frames to be used to which the target pixel belongs based on the point cloud data of the target pixel.
[0109] A depth estimation unit is used to determine the depth estimate of the moving object based on the depth values of the target pixel in at least two video frames to be used.
[0110] Based on the above technical solutions, the target processing method includes an inverse depth estimation method, and the depth estimation value determination module 330 further includes an inverse depth value determination submodule and a depth estimation value determination submodule.
[0111] The inverse depth value determination submodule is used to triangulate each video frame to be processed in the target video to obtain the inverse depth value of each pixel in the video frame to be processed.
[0112] The depth estimation determination submodule is used to determine the depth estimate of the moving object by clustering the inverse depth values in the same video frame to be processed.
[0113] Based on the above technical solutions, the depth estimation value determination submodule includes: a depth difference determination unit and a depth estimation value determination unit.
[0114] The depth difference determination unit is used to determine the depth difference between two adjacent inverse depth values after sorting them according to their magnitude.
[0115] A depth estimation unit is used to obtain the two target inverse depth values with the largest depth difference, and to determine the depth estimate of the moving object based on each inverse depth value that is greater than the target inverse depth value.
[0116] Based on the above technical solutions, the device further includes: a reverse depth value deletion module.
[0117] The inverse depth value deletion module is used to delete the corresponding inverse depth value and re-execute the step of determining the target inverse depth value before determining the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value.
[0118] Based on the above technical solutions, the depth estimation value determination unit is specifically used to process the average of all inverse depth values that are greater than the target inverse depth value to determine the depth estimation value of the moving object.
[0119] The technical solution of this disclosure, by determining the video processing type, further determining the target processing method for depth estimation of moving objects based on the video processing type, and finally determining the depth estimate value of the moving object in the video frame to be processed based on the target processing method, solves the problem that the prior art can only estimate the depth information of static objects, achieves the effect of accurately estimating the depth information of moving objects in video frames, and improves the applicability of depth estimation, meets the personalized needs of users, and enhances the user experience.
[0120] The depth estimation device for moving objects provided in this disclosure can execute the depth estimation method for moving objects provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0121] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0122] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 4 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 4 The diagram below shows the structure of the terminal device or server 500. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0123] like Figure 4 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0124] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0125] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0126] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0127] The electronic device provided in this embodiment and the depth estimation method for moving objects provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0128] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the depth estimation method for moving objects provided in the above embodiments.
[0129] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0130] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0131] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0132] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0133] Determine the video processing type;
[0134] Based on the video processing type, determine the target processing method for depth estimation of the moving object;
[0135] Based on the target processing method, the depth estimate of the moving object in the video frame to be processed is determined.
[0136] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0137] Determine the video processing type;
[0138] Based on the video processing type, determine the target processing method for depth estimation of the moving object;
[0139] Based on the target processing method, the depth estimate of the moving object in the video frame to be processed is determined.
[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0142] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0145] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0146] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0147] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for estimating the depth of a moving object, characterized in that, include: Determine the video processing type; Based on the video processing type, determine the target processing method for depth estimation of the moving object; Based on the target processing method, the depth estimate of the moving object in the video frame to be processed is determined; Wherein, if the video processing type is post-processing, the target processing method includes an inverse depth estimation method, and determining the depth estimate of the moving object in the video frame to be processed based on the target processing method includes: Triangulation is performed on each video frame to be processed in the target video to obtain the inverse depth value of each pixel in the video frame to be processed. The depth estimate of the moving object is determined by clustering the inverse depth values in the same video frame.
2. The method according to claim 1, characterized in that, The video processing types include real-time processing and post-processing.
3. The method according to claim 2, characterized in that, The target processing method includes a depth mean estimation method corresponding to the real-time processing type, or an inverse depth estimation method corresponding to the post-processing type.
4. The method according to claim 1, characterized in that, The target processing method includes a depth mean estimation method. The step of determining the depth estimate of a moving object in the video frame to be processed based on the target processing method includes: Determine the shooting parameters corresponding to the video frame to be processed and the pixel parameters of the moving object; Based on the shooting parameters, pixel parameters, and constraints, the target pixel is determined; Based on the point cloud data of the target pixels, the depth estimate of the moving object is determined.
5. The method according to claim 4, characterized in that, The step of determining the target pixel based on the shooting parameters, pixel parameters, and constraints includes: Triangulation is performed based on the shooting parameters and the pixel parameters to obtain the point cloud data corresponding to the pixel parameters; Based on the point cloud data and the constraints, the back-projection pixel parameters are determined; The target pixel is determined based on the pixel parameters and the back-projected pixel parameters.
6. The method according to claim 4, characterized in that, Determining the depth estimate of the moving object based on the point cloud data of the target pixels includes: Based on the point cloud data of the target pixel, determine at least two video frames to be used to which the target pixel belongs; The depth estimate of the moving object is determined based on the depth values of the target pixel in at least two video frames to be used.
7. The method according to claim 1, characterized in that, The step of determining the depth estimate of the moving object by clustering inverse depth values in the same video frame to be processed includes: After sorting the inverse depth values according to their magnitudes, the depth difference between two adjacent inverse depth values is determined. Obtain the two target inverse depth values with the largest depth difference, and determine the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value.
8. The method according to claim 7, characterized in that, Before determining the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value, the method further includes: If the ratio between the number of inverse depth values greater than or less than the target inverse depth value and the total number of inverse depth values is less than a preset ratio, then the corresponding inverse depth value is deleted, and the step of determining the target inverse depth value is re-executed.
9. The method according to claim 7, characterized in that, Determining the depth estimate of the moving object based on each inverse depth value greater than the target inverse depth value includes: The average of all inverse depth values greater than the target inverse depth value is processed to determine the depth estimate of the moving object.
10. A depth estimation device for a moving object, the device comprising: The video processing type determination module is used to determine the video processing type. The target processing method determination module is used to determine the target processing method for depth estimation of the moving object based on the video processing type. The depth estimation determination module is used to determine the depth estimation value of the moving object in the video frame to be processed based on the target processing method. Wherein, when the video processing type is post-processing, the target processing method includes an inverse depth estimation method; the depth estimation value determination module further includes: an inverse depth value determination submodule and a depth estimation value determination submodule; The inverse depth value determination submodule is used to perform triangulation processing on each video frame to be processed in the target video to obtain the inverse depth value of each pixel in the video frame to be processed. The depth estimation determination submodule is used to determine the depth estimate of the moving object by clustering the inverse depth values in the same video frame to be processed.
11. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the depth estimation method for a moving object as described in any one of claims 1-9.
12. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the depth estimation method for a moving object as described in any one of claims 1-9.
Citation Information
Patent Citations
Single-viewpoint video depth obtaining method based on scene classification and geometric dimension
CN105100771A
Image processing method and device, storage medium and electronic equipment
CN111612898A
Image processing method and device, electronic equipment and storage medium
CN113643342A