A 3D scene synthesis method integrating depth estimation and dual-view video
Through depth estimation and inter-frame feature analysis of dual-view video image sequences, the problems of discontinuous reconstruction in occluded areas and insufficient image quality under lighting conditions are solved, and stable and accurate reconstruction of three-dimensional scenes is achieved, which is suitable for lightweight devices and AR applications.
Patent Information
- Application Number
- CN202510935645.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing binocular depth estimation methods lack the continuity of reconstruction in occluded areas, the stability of occlusion recovery in dynamic scenes, and the robustness of image quality under complex lighting conditions, resulting in incomplete three-dimensional scenes or incorrect depth estimation. In particular, under dual-view input conditions, there is a lack of compensation for historical information in occluded scenes and adaptive enhancement of image brightness.
By acquiring a dual-view video image sequence, using an adaptive stereo matching network for depth estimation, combining inter-frame change features and brightness detection, generating an occlusion area mask, executing a residual fusion strategy for depth completion, and performing image enhancement or previous image superposition, constructing a depth confidence map for weighted fusion, and finally generating a dense 3D point cloud.
It effectively enhances the 3D reconstruction integrity and boundary continuity of occluded areas, improves the stability and accuracy of depth estimation in dynamic scenes and complex lighting conditions, and is suitable for lightweight systems such as dual-camera devices and AR glasses.
Smart Images

Figure CN120431274B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image fusion, and in particular to a three-dimensional scene synthesis method that fuses depth estimation and dual-view video. Background Art
[0002] With the widespread adoption of 3D reconstruction and interaction technologies such as augmented reality, virtual reality, and digital twins, the demand for rapidly constructing 3D scenes based on image sequences is increasing. Dual-camera, dual-view video input solutions, due to their ease of hardware deployment and high real-time performance, hold broad application prospects in consumer devices and mobile platforms.
[0003] Existing 3D scene reconstruction methods typically use structured light, LiDAR, or multi-view image arrays to achieve high-precision scene acquisition. However, these methods suffer from high equipment costs, complex deployment, and difficulty adapting to mobile devices. In contrast, 3D reconstruction methods based on binocular video image sequences combined with depth estimation networks are more portable and have become a mainstream research direction.
[0004] Currently, a variety of binocular depth estimation methods based on deep learning have been proposed, such as using a stereo matching network to extract features and estimate disparity between left and right images, and then obtaining a depth map for each frame through depth regression. However, these methods still have significant shortcomings in terms of the continuity of reconstruction of occluded areas, the stability of occlusion recovery in dynamic scenes, and the robustness of image quality under complex lighting conditions. In particular, under dual-view input conditions, the lack of mechanisms such as compensation for historical information in occluded scenes and adaptive enhancement of image brightness often leads to incomplete three-dimensional scenes or incorrect depth estimation. Most current methods process each frame independently, without fully considering the continuity between image frames and the use of redundant information, resulting in weak dependence on occluded area completion; and under conditions such as low light and overexposure, the degradation of image quality further affects the stability of the depth estimation model, and there is a lack of effective brightness enhancement and previous image fusion strategies to compensate for it.
[0005] Therefore, there is an urgent need for a method that integrates depth estimation and dual-view video input to effectively improve the reconstruction quality of occluded areas and the adaptability to complex lighting environments while ensuring lightweight, and to construct a three-dimensional scene with continuity and integrity, which is particularly suitable for indoor AR free-viewpoint interaction application scenarios. Summary of the Invention
[0006] A 3D scene synthesis method integrating depth estimation and dual-view video, comprising:
[0007] Acquire a dual-view video image sequence captured by two synchronous camera devices with fixed view angles, including a synchronized left image sequence and a synchronized right image sequence;
[0008] The left and right images of each frame are paired and input into the depth estimation model built based on the adaptive stereo matching network to extract image features, construct the disparity cost volume and perform depth regression to generate the initial depth map;
[0009] Trigger regional occlusion detection according to the set number of frames or viewpoint motion monitoring, obtain the disparity matching results and inter-frame change characteristics of the dual-view video image of the current frame, and generate an occlusion area mask;
[0010] The fused depth map in the previous frame is jointly spatially aligned with the dual-view video image of the current frame, and the occlusion area in the image of the current frame is located in combination with the occlusion area mask of the current frame;
[0011] For the occluded area, the residual fusion strategy is used to perform depth completion operation based on the aligned fused depth map of the previous frame, the initial depth map of the current frame, and the texture features of the dual-view video image, and the updated depth map is output;
[0012] Perform brightness detection on the dual-view video image of the current frame, select image enhancement or previous image superposition based on the detection result, output the processed dual-view video image, construct a depth confidence map, perform weighted fusion on the updated depth map, and output the fused depth map;
[0013] The fused depth map and the dual-view video image are jointly projected to construct a dense 3D point cloud of the current frame, which is spatially aligned with the 3D scene of the previous frame to output the 3D scene of the current frame.
[0014] As a preferred technical solution of the present invention, generating an initial depth map includes:
[0015] Based on the geometric constraints of the internal and external parameters of the camera device, a feature matching algorithm is used to match the feature points of the left and right images of each frame. Based on the matching results of the feature points, the disparity cost between each pixel pair is calculated, and the initial depth information is generated according to the distribution of the disparity cost. The paired images and initial depth information are used as input, and feature extraction is performed on the input image through the depth estimation model to construct a disparity cost volume and obtain the optimized disparity distribution. Depth regression is performed on the disparity cost volume to generate an initial depth map.
[0016] As a preferred technical solution of the present invention, generating an occlusion area mask includes:
[0017] Perform disparity matching on the dual-view images, identify areas where the disparity matching degree in the left and right images in the current frame exceeds the set threshold and where the matching fails, and use them as the disparity matching results; obtain the inter-frame change features between the current frame and the previous frame, including image brightness difference values, pixel displacement values, and edge structure change values; jointly analyze the disparity matching results and the inter-frame change features through a rule library to determine the occlusion area and generate an occlusion area mask for marking.
[0018] As a preferred technical solution of the present invention, the trigger area occlusion detection includes:
[0019] A frame detection threshold and a view angle movement detection threshold are set to control the triggering timing of occlusion detection; when the number of consecutive image frames reaches a preset threshold, or when the view angle of the camera device is detected to have moved beyond the preset angle range, the occlusion area detection process is triggered; the judgment of the view angle movement is based on analysis of image content changes, camera device position estimation or image registration results.
[0020] As a preferred technical solution of the present invention, locating the blocked area in the image of the current frame includes:
[0021] The fused depth map generated by the previous frame is mapped to the current frame image coordinate system through a spatial transformation method based on the internal and external parameters of the camera; combined with the occlusion area mask, the exact position and boundary of the occlusion area in the dual-view video image of the current frame are determined; and image registration and pixel-level interpolation optimization methods are used to compensate for the occlusion deformation caused by perspective differences or camera movement; the spatial transformation includes three-dimensional reprojection of the depth map of the previous frame and coordinate alignment based on the position and posture parameters of the camera device in the current frame.
[0022] As a preferred technical solution of the present invention, the performing of the depth completion operation includes:
[0023] In the determined occluded area, the aligned fused depth map of the previous frame, the initial depth map of the current frame, and the texture features of the dual-view image of the current frame are fused to construct multi-source residual image features, and pixel-by-pixel depth value optimization is performed to complete the missing depth information of the occluded area in the current frame.
[0024] As a preferred technical solution of the present invention, the brightness detection includes:
[0025] Brightness detection is performed on the left and right images of the current frame respectively. Based on the absolute brightness distribution of the entire image and the brightness deviation of the local area relative to the image mean, the area with abnormal brightness in the image is jointly determined. When the brightness detection result meets the preset trigger condition, the image brightness is compensated by image enhancement or superposition of the previous image. The image enhancement is performed through a neural network model.
[0026] As a preferred technical solution of the present invention, the preceding image superposition includes:
[0027] Select the qualified brightness detection area from the cached previous image, and fuse it with the corresponding area in the current frame image according to the set ratio based on the image expansion domain;
[0028] The image expansion domain is obtained based on the dynamically updated image center point, and the expanded area is marked according to the preset grid rules. The dynamic update is triggered according to the set frame number or the perspective movement detection result; if the camera device supports infrared image acquisition or exposure control function, then after the brightness detection, the device is triggered to perform infrared image acquisition or exposure adjustment operation, obtain auxiliary images and participate in image fusion to replace the previous image superposition result.
[0029] As a preferred technical solution of the present invention, the weighted fusion of the updated depth map includes:
[0030] Obtain the brightness distribution and texture clarity of the dual-view video image after brightness detection processing, and construct the depth confidence map of the image; jointly process the updated depth map and the depth confidence map, assign weighting coefficients to different pixel positions in the updated depth map through the confidence map, process the depth values in a pixel-weighted manner, and output a fused depth map.
[0031] The present invention has the following advantages:
[0032] The present invention constructs an occlusion detection mechanism and an occlusion area mask, spatially aligns the fused depth map of the previous frame with the current frame image, and combines it with a residual fusion strategy to achieve depth completion of the occluded area, effectively enhancing the integrity and boundary continuity of the three-dimensional reconstruction of the occluded area. It is particularly suitable for situations where targets or observation points frequently change in dynamic scenes.
[0033] The present invention determines image illumination abnormality areas through a joint detection method based on absolute brightness and relative brightness, and dynamically improves image quality by combining strategies such as image enhancement, previous image superposition, and device-level image acquisition, effectively reducing the interference of image quality on depth estimation accuracy under low-light or high-contrast conditions.
[0034] The present invention uses the disparity information of dual-view images and the inter-frame motion change characteristics to jointly analyze the occluded area, constructs a depth confidence map and introduces a pixel-by-pixel weighted fusion mechanism, so that the final fused depth map has higher stability and accuracy while retaining structural details, which is significantly better than the existing reconstruction method that only relies on static frame input.
[0035] The present invention uses dual-view image input and is suitable for common dual-camera devices or lightweight systems such as AR glasses. It also provides two compensation strategies, image-level and hardware-level, and can automatically switch between image fusion and infrared acquisition paths according to device capabilities, thereby improving the applicability of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only schematic diagrams of the present invention. Those skilled in the art can also derive other drawings based on the provided drawings without inventive effort.
[0037] Figure 1 This is a flow chart of a 3D scene synthesis method that integrates depth estimation and dual-view video, adopted in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] To make the objectives, technical solutions, and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0039] Embodiment, a 3D scene synthesis method integrating depth estimation and dual-view video, see Figure 1 As shown, the following steps are included:
[0040] Step S1: Obtain a dual-view video image sequence captured by two fixed-view synchronous camera devices. Specifically, a pair of cameras fixed at relative positions are used to respectively capture continuous image sequences of the left view and the right view, where these image sequences are generated by the synchronous capture devices.
[0041] The camera device is typically an industrial camera, a consumer camera, or a virtual reality device (such as AR glasses) having an image sensor (such as a CCD or CMOS sensor) and a corresponding image processing unit.
[0042] During image sequence acquisition, the left and right view frames must be captured in strict synchronization to ensure consistent timing. This synchronization is typically achieved through hardware trigger signals or through software algorithms for time alignment during post-processing. The acquisition frequency of each image frame is adjusted based on actual application requirements, with a common frequency of 30 to 60 fps to ensure smooth and temporal continuity of the image sequence.
[0043] During the synchronous acquisition process, the image data is stored in the memory of the computer or processing unit and encoded and compressed according to the output format of the device. Each frame of the left and right image sequences is transmitted to the subsequent processing unit for depth estimation and 3D reconstruction.
[0044] When capturing the dual-view video image sequence, the relative position and attitude of the cameras must be calibrated to ensure accurate calculation of the parallax between the two cameras. During calibration, the relative position (translation vector) and rotation angle (rotation matrix) between the two cameras are typically measured, and internal and external parameters are calculated using a calibration target or known geometric model. These calibration parameters are used in subsequent depth estimation and spatial alignment operations.
[0045] Step S2: Pair the left and right images of each frame. Specifically, each frame undergoes image preprocessing to remove noise and normalize the image. Image preprocessing includes operations such as color correction, brightness adjustment, and distortion correction to eliminate interference caused by internal and external parameters of the camera equipment or environmental factors such as lighting and reflections. The preprocessed image is then fed into the depth estimation model for further processing.
[0046] The depth estimation model is based on an Adaptive Stereo Matching Network (ASMN). This network uses deep learning to automatically learn feature matching rules between two-view images using training data. The model takes the left and right images as input. The network extracts depth features from the images and calculates the disparity value for each pixel, generating a preliminary depth map for each pixel.
[0047] During model processing, the network first uses a convolutional neural network (CNN) to extract image features, particularly key information such as edges, textures, and depth variations. These image features are then fed into the disparity calculation module. The disparity cost volume is constructed to convert the disparity between each pair of matching pixels into a cost function, which is then optimized using a fully convolutional network. The disparity cost volume optimizes image matching quality based on the distribution of disparity values.
[0048] After constructing the disparity cost volume, the network performs depth regression, which regresses the actual depth value of each pixel based on the disparity cost map. The depth regression stage post-processes the depth map to reduce errors in the disparity calculation and obtain an optimized initial depth map.
[0049] This initially generated depth map will be fused and optimized with the depth information of the previous frame in subsequent steps to further improve the accuracy of the overall depth estimation. It is important to note that depth regression not only relies on the disparity cost map, but also utilizes the local texture information and global feature information of the image to ensure that the depth map can still be accurate and stable in complex scenes (such as occlusion, dynamic objects, etc.).
[0050] Step S3: Triggering occlusion detection based on a set number of frames or viewing angle movement. Specifically, after acquiring each dual-view video frame, the algorithm detects occlusion areas by comparing the disparity matching results of the left and right views. During the disparity matching process for the dual-view images, the algorithm calculates the matching degree of each pixel and determines whether there are any anomalies in the disparity between pixel pairs. If the disparity matching degree between the left and right views in the current frame exceeds a set threshold or indicates a matching failure, the image is marked as a possible occlusion area.
[0051] Obtain inter-frame change features between the current frame and the previous frame. These change features primarily include image brightness difference, pixel displacement, and edge structure change. The image brightness difference reflects the change in brightness between the two frames, the pixel displacement indicates the motion of pixels between the two frames, and the edge structure change reflects the change in the image edge region between the two frames. These change features help the algorithm more accurately determine whether occlusion has occurred in the current frame, especially in dynamic scenes, and effectively distinguish foreground from background.
[0052] The disparity matching results are combined with inter-frame variation characteristics through a rule library. The analysis rules in the rule library combine the image's structural information and motion characteristics to determine which areas are occluded due to disparity inconsistencies or obstructions. Ultimately, combining this information, the algorithm generates an occlusion mask, which marks potentially occluded areas in the current frame.
[0053] The triggering mechanism for regional occlusion detection is determined by setting a frame detection threshold and a view motion detection threshold. The frame detection threshold is used to control occlusion detection after a certain number of consecutive frames, avoiding frequent calculations. The view motion detection threshold is used to determine whether the camera's viewing angle has changed significantly. If the camera's viewing angle exceeds the preset angle range or the image content changes significantly, the occlusion region detection process will be triggered.
[0054] Viewpoint movement can be determined based on changes in image content (such as differences in edges and textures), estimated camera position, or analysis of image registration results. During this process, camera position and angle information must also be estimated using external sensors (such as gyroscopes and accelerometers) or visual odometry from the image sequence. This triggering mechanism effectively controls the computational overhead of occlusion detection, ensuring that computationally intensive occlusion detection is performed only when necessary, improving overall efficiency.
[0055] Step S4: Jointly spatially align the fused depth map generated by the previous frame with the dual-view video image of the current frame. The core of this process is to map the depth map of the previous frame to the image coordinate system of the current frame through spatial transformation, ensuring the consistency of the depth information between the previous and next frames in three-dimensional space.
[0056] Specifically, the spatial transformation operation is based on the intrinsic and extrinsic parameters of the camera device. The internal parameters include focal length, optical center and distortion coefficient, while the extrinsic parameters include the relative position and rotation angle between the two cameras. The three-dimensional coordinates of the depth map of the previous frame are reprojected and aligned with the image coordinate system of the current frame. The specific operation is to combine the depth value of each pixel in the depth map of the previous frame with the corresponding spatial coordinates, and then convert the three-dimensional coordinates according to the device position and rotation angle, and finally obtain depth information consistent with the image coordinate system of the current frame.
[0057] After completing spatial alignment with the previous frame's depth map, the occluded regions in the current frame are further determined by combining the occluded region mask from the current frame. The occluded region mask, generated in step S3, identifies occluded regions in the current frame due to factors such as disparity mismatching and inter-frame variations. This mask is used to pinpoint the specific location and boundaries of occluded regions in the current frame, providing a basis for subsequent depth completion operations.
[0058] To improve alignment accuracy, image registration and pixel-level interpolation methods are used during spatial alignment to optimize deformation caused by differences in viewpoint or camera position. Image registration further adjusts for spatial alignment errors by matching feature points between the current and previous frames. Pixel-level interpolation is used to address holes or discontinuities in the images during the alignment process, ensuring that the aligned depth map remains spatially smooth and avoiding severe depth misalignment or inconsistencies caused by alignment errors.
[0059] Step S5: For the area marked as occluded in the current frame image, a depth completion operation is performed using a residual fusion strategy based on the aligned previous frame fused depth map, the current frame initial depth map, and the texture features of the current frame dual-view image.
[0060] The aligned fused depth map from the previous frame is used as a reference and compared with the initial depth map of the current frame. During this comparison, residual image features are calculated by analyzing the depth information of the same areas in the two frames. These residual image features represent the difference between the depth values in the current frame and the reference depth values in the previous frame. This calculation of residual image features allows for more accurate depth compensation, especially in the presence of occluded areas or missing textures.
[0061] Depth completion is further optimized by extracting texture features from the current frame. Specifically, a convolutional neural network (CNN) is used to extract texture features from the current frame, obtaining information about the texture distribution, edge structure, and local details in each region of the image. These texture features are used to adjust the depth values in the completed region to ensure greater visual realism and coherence in the completed depth map. The introduction of texture features helps enhance the continuity of the depth map in occluded areas, avoiding noticeable gaps or unnatural jumps in the depth map.
[0062] By combining the depth information of the previous frame, the initial depth map of the current frame, and image texture features, a residual fusion strategy is used to optimize the depth values of occluded areas pixel by pixel. In this process, the depth completion operation not only relies on spatial image matching but also incorporates texture features in the image to ensure the visual and geometric accuracy of the depth estimation.
[0063] The result of depth completion is an updated depth map, which not only repairs the missing depth information of the occluded area, but also enhances the continuity and consistency of the image, providing high-quality depth data for subsequent weighted fusion and 3D reconstruction.
[0064] Step S6: Perform brightness detection on the dual-view video image of the current frame. The goal of this step is to identify areas with insufficient brightness or overexposure in the current frame so that appropriate image enhancement or previous image overlay strategies can be selected for subsequent processing.
[0065] Specifically, brightness detection is performed by calculating the global brightness distribution and local brightness deviation of the current frame. Global brightness distribution refers to the average brightness value of the entire image, which helps to assess whether the image is too dark or too bright. Local brightness deviation is calculated by calculating the difference between each region in the image and the average brightness of the entire image to determine the lighting conditions of the local area. This is particularly useful for detecting shadows or highlights in the image.
[0066] During brightness detection, the absolute brightness of the image is first calculated—that is, the grayscale value or brightness value (the luminance component in RGB space) of each pixel in the image. The relative brightness deviation of each pixel from the global mean is then calculated. By setting a preset brightness threshold, abnormal brightness areas in the current frame are automatically detected. These abnormal areas include underlit shadows and overexposed highlights, which can affect the accuracy of subsequent depth estimation.
[0067] Once an area of abnormal brightness is detected, a corresponding processing strategy will be selected based on the detection results. If the overall brightness of the image is normal but there are abnormalities in a local area, an image enhancement method will be selected; if the overall brightness is insufficient or there is a serious lighting problem, the previous image overlay strategy will be used.
[0068] Image enhancement is performed using a neural network model. Specifically, the model uses training data to perform tonal adjustments, contrast enhancement, and exposure compensation on the image to improve the image's brightness distribution and detail. Based on the image's brightness detection results, the neural network model automatically selects appropriate enhancement strategies, such as automatic exposure correction and gamma adjustment.
[0069] Previous image overlay selects areas from a cached previous image that have passed brightness testing and fuses them with the current image at a set ratio. This strategy effectively compensates for areas of insufficient brightness in the current frame, especially in low-light or overexposed scenes. The previous image provides more detailed information, improving overall image quality.
[0070] The dual-view video images after brightness detection and processing are output as processed image pairs, which will provide more reliable input data in the subsequent depth confidence map construction and weighted fusion.
[0071] Step S7: After performing brightness detection and depth estimation on the current frame's dual-view video image, a depth confidence map is constructed. The depth confidence map quantifies the reliability of the depth value of each pixel. It estimates the reliability of the depth value based on the image's brightness distribution and texture clarity.
[0072] Specifically, the construction process of the depth confidence map includes the following steps:
[0073] Brightness distribution calculation: First, the overall brightness distribution of the current frame's dual-view video image is calculated. The brightness distribution of the image reflects the global illumination of the image. Regions with higher brightness typically have stronger texture features, while regions with lower brightness or overexposure can lead to depth estimation errors.
[0074] Texture Clarity Assessment: In addition to brightness distribution, the image's texture clarity also needs to be assessed. Texture clarity is measured by the image's edge information and level of detail. Image edge regions typically contain more depth information, while low-texture areas (such as uniform backgrounds) lead to increased depth estimation errors. Therefore, regions with rich texture in an image are typically assigned higher confidence values.
[0075] Depth confidence map generation: This combines the image's brightness distribution and texture clarity to generate a depth confidence map. A depth confidence map is a two-dimensional matrix where the value of each pixel represents the confidence level of the depth information at that location. Generally, regions with higher confidence levels are given a larger weight, while regions with lower confidence levels are given a smaller weight, avoiding the use of unreliable depth information in these areas.
[0076] After generating the depth confidence map, the next step is to enter the weighted fusion process of the depth map. Weighted fusion is to fuse multiple depth maps (such as the depth map of the previous frame and the depth map of the current frame) according to their respective confidence levels to obtain a more accurate depth map. The specific weighted fusion method is as follows:
[0077] Weight calculation: Each depth map is weighted according to the value of its depth confidence map when it is fused. Regions with higher confidence are given larger weights, ensuring that their depth information has a greater impact on the final depth map.
[0078] Fusion process: During weighted fusion, the depth values of each pixel in multiple depth maps are weighted and averaged according to their corresponding confidence weights. The resulting fused depth map combines depth information from multiple sources, and a weighted strategy ensures that depth information in high-confidence areas is more accurate.
[0079] The output fused depth map will provide reliable depth information for subsequent 3D point cloud construction and 3D scene synthesis, further improving the accuracy and continuity of 3D reconstruction.
[0080] Step S8: Combine the weighted fused depth map with the current frame's dual-view video image for a joint projection calculation to construct a dense 3D point cloud for the current frame. The core goal of this step is to convert the depth information of each pixel in the image into actual coordinates in 3D space, generating 3D point cloud data.
[0081] Depth map and image projection:
[0082] In the joint projection calculation, the depth value of each pixel in the depth map is paired with the corresponding pixel coordinates. Using the camera's intrinsic parameters (such as focal length and optical center) and extrinsic parameters (such as camera position and rotation angle), the depth information of these pixels is projected into three-dimensional space to obtain the corresponding three-dimensional coordinates of the pixel. The projection operation establishes a correspondence between the image's two-dimensional coordinate system and the three-dimensional world coordinate system, accurately mapping the spatial position of each pixel in the depth map to three-dimensional space.
[0083] Color assignment of point cloud:
[0084] After point cloud generation, the color information from the dual-view video images is typically used as the texture or color attribute of the point cloud, assigning RGB values to each 3D point cloud point. This ensures that the point cloud not only has spatial location information but also presents realistic color representation in 3D space. This color information comes from either the left or right view of the image, and the point cloud can be assigned color based on either a single image or a weighted average of the two images.
[0085] Spatial alignment of 3D scenes:
[0086] The generated 3D point cloud of the current frame needs to be spatially aligned with the 3D scene of the previous frame. This is achieved by calculating the relative pose (i.e., the change in camera position and attitude) between the current and previous frames. Relative pose can be calculated using image registration, visual odometry, or position and attitude information provided by external sensors such as an IMU. During this process, the 3D point cloud of the previous frame is transformed into the coordinate system of the current frame through a rigid transformation (rotation and translation), ensuring that the 3D point clouds of the previous and next frames are aligned in a unified 3D spatial coordinate system.
[0087] 3D scene update:
[0088] After aligning the current frame's 3D point cloud with the previous frame's scene, the point cloud of the current frame is merged with the previous frame's point cloud to update the overall 3D scene. This point cloud merging process employs an optimization strategy to ensure that the point cloud data within the scene does not overlap or become distorted. If duplicate points exist across multiple point clouds, the optimal depth value is selected based on weights (such as depth confidence or measurement accuracy) to enhance the accuracy of the final synthesized scene.
[0089] Ultimately, the output 3D scene will contain the combined 3D point cloud data of the current and previous frames. By continuously synthesizing the 3D point clouds of multiple frames, a dense 3D model of the entire scene is gradually constructed, ensuring the continuity and stability of the scene.
[0090] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A 3D scene synthesis method integrating depth estimation and dual-view video, characterized in that: include: Acquire a dual-view video image sequence captured by two synchronous camera devices with fixed view angles, including a synchronized left image sequence and a synchronized right image sequence; The left and right images of each frame are paired and input into the depth estimation model built based on the adaptive stereo matching network to extract image features, construct the disparity cost volume and perform depth regression to generate the initial depth map; Trigger regional occlusion detection according to the set number of frames or viewpoint motion monitoring, obtain the disparity matching results and inter-frame change characteristics of the dual-view video image of the current frame, and generate an occlusion area mask; The fused depth map in the previous frame is jointly spatially aligned with the dual-view video image of the current frame, and the occlusion area in the image of the current frame is located in combination with the occlusion area mask of the current frame; For the occluded area, the residual fusion strategy is used to perform depth completion operation based on the aligned fused depth map of the previous frame, the initial depth map of the current frame, and the texture features of the dual-view video image, and the updated depth map is output; Perform brightness detection on the dual-view video image of the current frame, select image enhancement or previous image superposition based on the detection result, output the processed dual-view video image, construct a depth confidence map, perform weighted fusion on the updated depth map, and output the fused depth map; The fused depth map and the dual-view video image are jointly projected to construct a dense 3D point cloud of the current frame, which is spatially aligned with the 3D scene of the previous frame to output the 3D scene of the current frame.
2. The method for synthesizing a three-dimensional scene by integrating depth estimation and dual-view video according to claim 1, characterized in that: Generating the initial depth map includes: Based on the geometric constraints of the internal and external parameters of the camera device, a feature matching algorithm is used to match the feature points of the left and right images of each frame. Based on the matching results of the feature points, the disparity cost between each pixel pair is calculated, and the initial depth information is generated according to the distribution of the disparity cost. The paired images and initial depth information are used as input, and feature extraction is performed on the input image through the depth estimation model to construct a disparity cost volume and obtain the optimized disparity distribution. Depth regression is performed on the disparity cost volume to generate an initial depth map.
3. The method for synthesizing a three-dimensional scene by integrating depth estimation and dual-view video according to claim 1, wherein: Generating an occlusion area mask includes: Perform disparity matching on the dual-view images, identify areas where the disparity matching degree in the left and right images in the current frame exceeds the set threshold and where the matching fails, and use them as the disparity matching results; obtain the inter-frame change features between the current frame and the previous frame, including image brightness difference values, pixel displacement values, and edge structure change values; jointly analyze the disparity matching results and the inter-frame change features through a rule library to determine the occlusion area and generate an occlusion area mask for marking.
4. The method for synthesizing a three-dimensional scene by fusing depth estimation with dual-view video according to claim 3, wherein: The trigger area occlusion detection includes: A frame detection threshold and a view angle movement detection threshold are set to control the triggering timing of occlusion detection; when the number of consecutive image frames reaches a preset threshold, or when the view angle of the camera device is detected to have moved beyond the preset angle range, the occlusion area detection process is triggered; the judgment of the view angle movement is based on analysis of image content changes, camera device position estimation or image registration results.
5. The method for synthesizing a three-dimensional scene by integrating depth estimation and dual-view video according to claim 1, wherein: The locating of the blocked area in the image of the current frame includes: The fused depth map generated by the previous frame is mapped to the current frame image coordinate system through a spatial transformation method based on the internal and external parameters of the camera; combined with the occlusion area mask, the exact position and boundary of the occlusion area in the dual-view video image of the current frame are determined; and image registration and pixel-level interpolation optimization methods are used to compensate for the occlusion deformation caused by perspective differences or camera movement; the spatial transformation includes three-dimensional reprojection of the depth map of the previous frame and coordinate alignment based on the position and posture parameters of the camera device in the current frame.
6. The method for synthesizing a three-dimensional scene by fusing depth estimation with dual-view video according to claim 1, characterized in that: The performing of the depth completion operation includes: In the determined occluded area, the aligned fused depth map of the previous frame, the initial depth map of the current frame, and the texture features of the dual-view image of the current frame are fused to construct multi-source residual image features, and pixel-by-pixel depth value optimization is performed to complete the missing depth information of the occluded area in the current frame.
7. The method for synthesizing a three-dimensional scene by fusing depth estimation with dual-view video according to claim 1, characterized in that: The brightness detection includes: Brightness detection is performed on the left and right images of the current frame respectively. Based on the absolute brightness distribution of the entire image and the brightness deviation of the local area relative to the image mean, the area with abnormal brightness in the image is jointly determined. When the brightness detection result meets the preset trigger condition, the image brightness is compensated by image enhancement or superposition of the previous image. The image enhancement is performed through a neural network model.
8. The method for synthesizing a three-dimensional scene by fusing depth estimation with dual-view video according to claim 1, wherein: The preceding image superposition includes: Select the qualified brightness detection area from the cached previous image, and fuse it with the corresponding area in the current frame image according to the set ratio based on the image expansion domain; The image expansion domain is obtained based on the dynamically updated image center point, and the expanded area is marked according to the preset grid rules. The dynamic update is triggered according to the set frame number or the perspective movement detection result; if the camera device supports infrared image acquisition or exposure control function, then after the brightness detection, the device is triggered to perform infrared image acquisition or exposure adjustment operation, obtain auxiliary images and participate in image fusion to replace the previous image superposition result.
9. The method for synthesizing a three-dimensional scene by integrating depth estimation and dual-view video according to claim 1, wherein: The weighted fusion of the updated depth map includes: Obtain the brightness distribution and texture clarity of the dual-view video image after brightness detection processing, and construct the depth confidence map of the image; jointly process the updated depth map and the depth confidence map, assign weighting coefficients to different pixel positions in the updated depth map through the confidence map, process the depth values in a pixel-weighted manner, and output a fused depth map.
Citation Information
Patent Citations
Rigid object virtual and real shielding method based on RGB-D three-dimensional reconstruction
CN117541755A
Depth estimation method based on multi-view self-supervised learning
CN118552596A