Three-dimensional scene reconstruction method and device based on monocular camera

Through the 2D segmentation and depth estimation model, a 3D point cloud map is generated, which solves the problem of high complexity in the reconstruction of three-dimensional scenes of monocular cameras and realizes efficient three-dimensional scene reconstruction of train inspection.

CN120580359APending Publication Date: 2025-09-02ZHENGZHOU THINK FREELY HI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510692235.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing three-dimensional scene reconstruction method of monocular cameras is highly complex and has problems with scale drift and dynamic scene sensitivity, making it difficult to directly apply to three-dimensional scene reconstruction during train inspection.

Method used

The 2D segmentation model and depth estimation model are used to segment the inspection video data captured by the monocular camera and dig deep feature. The 3D point cloud map is generated by fusion of 2D segmentation map and depth map, and the three-dimensional scene reconstruction is carried out using continuous frame 3D point cloud maps, and the fusion process is optimized by dynamic weighting and geometric correction technology.

Benefits of technology

The three-dimensional scene reconstruction process is simplified, the robustness and accuracy of reconstruction is improved, and the calculation complexity is reduced. It is suitable for three-dimensional scene reconstruction of train inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580359A_ABST
    Figure CN120580359A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional scene reconstruction method and device based on a monocular camera, and belongs to the technical field of three-dimensional reconstruction. According to the invention, image segmentation is carried out on each data frame in inspection video data through a 2D segmentation model to obtain a 2D segmentation image, depth feature mining is carried out on the inspection video data through a depth estimation model to obtain a two-dimensional single-channel depth image, and the obtained 2D segmentation image and the depth image are fused to generate a 3D point cloud image. And performing three-dimensional scene reconstruction by using the generated 3D point cloud image. Therefore, the method does not need a complex algorithm, only needs to carry out image segmentation and depth data extraction on the video data obtained by the monocular camera, can obtain the corresponding 3D point cloud image, carries out the three-dimensional scene reconstruction through the 3D point cloud, and greatly simplifies the complexity of carrying out the three-dimensional scene reconstruction through the monocular camera at present.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a three-dimensional scene reconstruction method and device based on a monocular camera, belonging to the technical field of three-dimensional reconstruction. Background Art

[0002] Train inspection is a crucial task for train operators and a crucial prerequisite for ensuring safe train operation. During train inspections, 3D scene reconstruction is typically required. This process can also be used for routine inspection video analysis to assess whether operators have completed various inspection checkpoints and standardize operational procedures. It can also be used for simulation training for train trainees during inspection training. Currently, however, monocular cameras are used for image capture during train inspections. However, the 2D images captured by monocular cameras cannot be directly used to reconstruct the 3D scene within the train compartment. Complex image conversion and processing are generally required, making the process cumbersome. The main existing methods for 3D reconstruction using monocular cameras include monocular SLAM and structure from motion (SFM). These two traditional methods have limitations (such as scale drift, sensitivity to dynamic scenes, and high computational complexity) that restrict their practical application. A method combining deep networks and segmentation networks can directly predict depth or segment dynamic regions through end-to-end learning, while also integrating semantic information, significantly improving the robustness and accuracy of reconstruction. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and device for reconstructing a three-dimensional scene based on a monocular camera, so as to solve the problem that the current three-dimensional scene reconstruction based on a monocular camera has a complicated process.

[0004] To solve the above technical problems, the present invention provides a three-dimensional scene reconstruction method based on a monocular camera, which includes:

[0005] 1) Obtain inspection video data captured by a monocular camera;

[0006] 2) Use the 2D segmentation model to perform image segmentation on each data frame in the inspection video data to obtain a 2D segmentation map containing the segmentation target; use the depth estimation model to perform depth feature mining on each data frame in the inspection video data captured by the monocular camera to obtain a two-dimensional single-channel depth map containing each pixel point;

[0007] 3) Fuse the obtained 2D segmentation map with the depth map to obtain a 3D point cloud map of each frame;

[0008] 4) Use continuous frame 3D point cloud images to reconstruct the three-dimensional scene.

[0009] Furthermore, in step 3), when fusing the obtained 2D segmentation map with the depth map, the image coordinates of the segmentation targets in the 2D segmentation map are multiplied pixel by pixel by the depth values ​​of the pixels corresponding to each segmentation target in the depth map to determine the 3D coordinates of each segmentation target.

[0010] Furthermore, the step 3) is performed by superposition in a dynamic weighted manner, and the weight of the depth data is determined according to the confidence score of the depth value of each pixel in the image frame output by the depth estimation model. The higher the confidence score of the pixel depth value, the greater the corresponding weight, and the lower the confidence score of the pixel depth value, the smaller the corresponding weight.

[0011] Furthermore, the step 3) is performed by dynamically weighting the superposition, where the weight of the depth data of the pixel point closer to the boundary of the segmentation target is smaller.

[0012] Furthermore, the step 3) is performed by dynamically weighting the superposition, reducing the weight of the image segmentation data in a low-light environment and reducing the weight of the image depth data in a strong light environment.

[0013] Furthermore, in step 3), when the obtained 2D segmentation map and the depth map are fused, the 2D segmentation map and the depth map are aligned, corresponding to the scene under the same perspective. If there is any misalignment between the two, correction is performed through geometric transformation or feature matching.

[0014] Furthermore, the process of reconstructing a three-dimensional scene using continuous frame 3D point cloud data in step 4) includes: matching the obtained 3D point cloud map of the current frame with the 3D point cloud map of the previous frame, fusing the matched result with the current global 3D point cloud map, and updating the global 3D point cloud map; the current global 3D point cloud map is obtained based on the fusion of 3D point cloud maps of frames before the current frame; during fusion, if there is a new target in the current frame, the new target is added to the current global 3D point cloud map for updating; if there is a reduction in the target in the current frame, the current global 3D point cloud map is not updated.

[0015] Furthermore, the depth estimation model adopts a ZeroDepth model or a ZoeDepth model.

[0016] Furthermore, each data frame in the inspection video data in step 2) refers to an extracted key frame, and the key frame is a frame with a large semantic change or a frame with a depth difference exceeding a set threshold.

[0017] The present invention also provides a three-dimensional scene reconstruction device based on a monocular camera, comprising a processor, wherein the processor is used to execute computer program instructions to implement the above-mentioned three-dimensional scene reconstruction method based on a monocular camera.

[0018] The beneficial effects of the present invention are as follows: As an improved invention, the present invention uses a 2D segmentation model to perform image segmentation on each data frame in the inspection video data to obtain a 2D segmentation map, uses a depth estimation model to perform deep feature mining on the inspection video data to obtain a two-dimensional single-channel depth map, fuses the obtained 2D segmentation map with the depth map to generate a 3D point cloud map, and uses the generated 3D point cloud map to reconstruct the three-dimensional scene. As can be seen, the present invention does not require complex algorithms, and only needs to perform image segmentation and depth data extraction on the video data obtained by the monocular camera to obtain the corresponding 3D point cloud map. Reconstructing the three-dimensional scene from the 3D point cloud greatly simplifies the complexity of current three-dimensional scene reconstruction using monocular cameras. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of a method for reconstructing a three-dimensional scene based on a monocular camera according to the present invention;

[0020] Figure 2 It is a schematic diagram of the process of reconstructing a three-dimensional scene using continuous frame 3D point cloud images of the present invention. DETAILED DESCRIPTION

[0021] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0022] The present invention performs 2D image segmentation and depth feature mining through data measured by a monocular camera, fuses the 2D segmentation map with the depth map to generate a 3D point cloud map, and reconstructs the corresponding three-dimensional scene based on the 3D point cloud map, thereby simplifying the reconstruction process.

[0023] Implementation method of three-dimensional scene reconstruction method based on monocular camera

[0024] The present invention first obtains the inspection video data shot by a monocular camera, and then uses a 2D segmentation model to perform image segmentation on each data frame in the inspection video data to obtain a 2D segmentation map containing the segmentation target; uses a depth estimation model to perform deep feature mining on the inspection video data shot by a monocular camera containing camera internal parameters to obtain a two-dimensional single-channel depth map containing each pixel point; then fuses the obtained 2D segmentation map with the depth map to obtain a 3D point cloud map of each frame; finally, uses the continuous 3D point cloud maps of each frame to reconstruct the three-dimensional scene. Its implementation process is as follows: Figure 1 As shown, the following is a detailed description.

[0025] 1. Obtain inspection video data captured by a monocular camera.

[0026] The present invention is aimed at the train inspection scenario, in which the inspection video data is captured by the inspection instrument carried by the inspector, wherein the inspection instrument carried by the inspector uses a monocular camera to capture the inspection video data. In addition to obtaining the inspection video data captured by the monocular camera, the present invention also needs to obtain the internal parameters of the monocular camera to facilitate subsequent in-depth analysis of the video data.

[0027] 2. Use the 2D segmentation model and depth estimation model to process the inspection video data to obtain a 2D segmentation map and depth map.

[0028] Since the video data captured by a monocular camera is continuous 2D data, the present invention uses a 2D segmentation model to perform target segmentation processing on each image frame in the inspection video data captured by the monocular camera. The 2D segmentation model here can adopt an image segmentation model based on deep learning, such as the FastSAM model, the MobileSAM model, etc., and each data frame in the inspection video data is sequentially input into the 2D segmentation model. Through the processing of the 2D segmentation model, a fine-grained 2D segmentation map, also called a segmentation mask map, is quickly segmented. The 2D segmentation map is a grayscale single-channel view that displays the target object category.

[0029] As a preferred embodiment, the 2D segmentation model of the present invention adopts the FastSAM model. The FastSAM model is a segmentation model with high computational efficiency and guaranteed accuracy. The core idea of ​​the FastSAM model is to decompose the segmentation task into two stages: 1) Full instance segmentation, which adopts the improved Yolov8-seg model based on the YOLACT method, and can quickly generate segmentation masks for all instances in the image; 2) Prompt-guided selection: According to the (points, boxes, text) provided by the user, the target area is selected from the full instance segmentation results. Through the FastSAM model, the present invention can quickly perform target segmentation on each frame image in the inspection video data. Specifically, the present invention uses the FastSAM model to process all frames, and can also extract image key frames according to different scene interval frames for segmentation and depth estimation processing. The key frames are selected in combination with segmentation or depth estimation. The selection strategy is:

[0030] 1) The keyframe selection method based on semantic segmentation is to select a keyframe if a significant semantic change is detected in the current frame (such as a scene change from inside the car to outside the car). This method can exclude frames with a high proportion of dynamic objects (such as people in the car), avoid interference, improve dynamic scene robustness, and reduce invalid keyframes. 2) The keyframe selection method based on depth estimation is to select a keyframe if the difference in depth map between the current frame and the previous keyframe (such as the mean squared error) exceeds a threshold, which can improve the robustness of low-texture scenes.

[0031] The inspection video data is processed using a depth estimation model to obtain a depth map. The purpose of the depth estimation model in the present invention is to obtain the depth value of the pixel points of each image frame in the inspection video data. Since the inspection video data is a 2D image, the internal parameters of the monocular camera during shooting are required to obtain accurate depth data from the 2D image. The depth estimation model of the present invention can adopt an image depth estimation model based on deep learning, and input the internal parameters of the monocular camera during shooting and each frame image in the inspection video data into the ZeroDepth model. The depth estimation model will output the depth map of each frame image.

[0032] As a preferred embodiment, the depth estimation model of the present invention adopts the ZeroDepth model, which is a monocular depth estimation model based on zero-sample scale perception. The model makes two key modifications to the standard architecture of monocular depth estimation to achieve robustness to geometric domain gaps. The model adopts the following two key technologies: 1) input-level geometric embedding, which combines monocular camera parameters and image features to enable the network to infer the physical size of objects and learn scale priors; 2) variational global latent variable representation, which decouples the encoding and decoding stages through a learned global latent variable representation, where the latent variable representation is variational and is sampled and decoded in a probabilistic manner to generate multiple predictions. This method enables the model to generate metrically accurate depth estimates between different data sets. The input of the ZeroDepth model includes a single image, camera intrinsic parameters, and geometric embedding (encoding of the camera intrinsic parameters). The single image is a single frame of inspection video data captured by the monocular camera in the present invention. The camera intrinsic parameters are the camera intrinsic parameters of the monocular camera used in the present invention, including the camera focal length and principal point coordinates. They are important information about the camera imaging geometry and help the model understand the actual size and distance of objects in the image. The geometric embedding is the pixel-level geometric features generated by normalizing and encoding the pixel line of sight directions in the image frame. The geometric features, combined with the image features, can provide richer contextual information for the model. In this embodiment, the camera intrinsic parameters of the monocular camera and the corresponding geometric features obtained by normalizing and encoding the pixel line of sight directions of each image frame of the video data are obtained. The camera intrinsic parameters, each image frame in the video data, and the geometric features of each image frame are input into the ZeroDepth model. The ZeroDepth model will then output a depth map corresponding to each image frame. The depth map includes the depth data of each pixel point.

[0033] As a preferred embodiment, the depth estimation model of the present invention may also adopt the ZoeDepth model. The ZoeDepth model is a multimodal monocular depth estimation network that combines relative depth and metric depth estimation. The input of the model is a single image.

[0034] The present invention uses the ZeroDepth model or the ZoeDepth model to process all frames, and can also extract image key frames according to different scene interval frames for depth estimation. The selection of key frames has been described before and will not be detailed here. The key frames used for image segmentation and depth estimation are the same.

[0035] 3. The obtained 2D segmentation map and depth map are superimposed and fused to obtain the corresponding 3D point cloud map.

[0036] Through step 2, the 2D segmentation map and 3D depth map corresponding to each image frame in the inspection video data can be obtained, and the 2D segmentation map and the 3D depth map are fused to determine the corresponding 3D point cloud map. The present invention fuses the obtained 2D segmentation map and the depth map by superimposing the segmentation target in the 2D segmentation map with the depth of the pixel points corresponding to each segmentation target in the depth map to obtain the depth data of each segmentation target. The map containing the depth data of each segmentation target is the 3D point cloud map. Since there is an alignment error problem between the segmentation edge and the depth edge of the same object, the segmentation map and the depth map need to be aligned when superimposed, and both correspond to the scene under the same perspective. If there is a slight alignment error, the present invention corrects it through geometric transformation or feature matching.

[0037] The present invention adopts a dynamic weighting method when superimposing the segmentation targets in the 2D segmentation map with the depths of the pixels corresponding to each segmentation target in the depth map. The dynamic weighting here refers to the dynamic change of the weight between the segmentation result of the 2D segmentation map and the depth data during superposition. The following three aspects are mainly considered to affect the change of the weight.

[0038] 1) Adjust the weight according to the confidence score of the depth value of each pixel in the image frame output by the depth estimation model: Determine the weight of the depth data based on the confidence score of the depth value of each pixel in the image frame output by the depth estimation model. The higher the confidence score of the depth value, the greater the corresponding weight, and the lower the confidence score of the depth value, the smaller the corresponding weight. Through this dynamic adjustment method, the influence of the high confidence area can be increased. 2) Adjust the weight according to the distance from the boundary pixel of the segmentation target: The closer the pixel is to the boundary of the segmentation target, the smaller the weight of the depth data. Since the depth estimation near the boundary of the segmentation target may be inaccurate, the present invention can appropriately reduce the weight of the area through this dynamic adjustment method. 3) Adjust the weight based on the lighting environment during shooting: Reduce the weight of the image segmentation data in low light environment, and reduce the weight of the image depth data in strong light.

[0039] Steps for overlaying the segmented target with the pixel points corresponding to the target in the depth map:

[0040] Step 1: Input a single RGB image, obtain the segmentation target object, such as the seat in the image, and obtain the segmentation mask map of the seat.

[0041] Step 2: The same RGB image is input into the depth estimation network to obtain the depth value of each pixel.

[0042] Step 3: Extract the depth information of the target object and multiply the segmentation mask with the depth map pixel by pixel to calculate the depth value of the target object (seat).

[0043] Step 4: Generate 3D point cloud. The calculation formula is as follows

[0044]

[0045] Z=z Formula (3)

[0046] Among them, (X, Y, Z) represents the 3D coordinates, (u, v) represents the image coordinates of each pixel, (c x ,c y ) is the principal point of the camera, (f x ,f y ) is the focal length of the camera, and z is the depth value of each pixel. Through the above calculation, the coordinates of all valid pixels are combined into a point cloud. During this calculation, corresponding weights can be assigned to the pixel's image coordinates and depth value.

[0047] 4. Use 3D point cloud images to reconstruct three-dimensional scenes.

[0048] Through the above process, the present invention can obtain a multi-frame 3D point cloud image, and the obtained multi-frame 3D point cloud image can be used to reconstruct a three-dimensional scene (three-dimensional map). The implementation process is as follows: Figure 2 shown.

[0049] The present invention is obtained by accumulating and reconstructing the obtained continuous multi-frame data. If the current frame is the t-th frame 3D point cloud map, the present invention now matches and fuses the t-th frame 3D point cloud map with the previous frame (t-1 frame) 3D point cloud map. Here, a 3D point cloud registration algorithm can be used for matching and fusion, such as the ICP registration algorithm, to register the t-th frame 3D point cloud map with the t-1-th frame 3D point cloud map. If the target in the t-th frame 3D point cloud map does not exist in the t-1 frame 3D point cloud map, it is considered that the target is a new target and is added to the global 3D point cloud map (the global 3D point cloud map here refers to the 3D point cloud map before the t-th frame). In the t+1 frame, the 3D point cloud map is updated. The updated global 3D point cloud map is the global 3D point cloud map corresponding to the t-1 frame. If the target in the 3D point cloud map of the t-1 frame does not exist in the 3D point cloud map of the t frame, and the target decreases, the current global 3D point cloud map is not updated. At the t+1 frame, the 3D point cloud map of the t+1 frame is aligned with the 3D point cloud map of the t frame. According to the alignment result, the global 3D point cloud map (the global 3D point cloud map here refers to the 3D point cloud map constructed before the t+1 frame, that is, the 3D point cloud map updated at the t frame) is updated to achieve the update of the t+1 frame global 3D point cloud map. Repeat this process until the last frame of 3D point cloud map. The global 3D point cloud map after the last frame of 3D point cloud map is the final global 3D point cloud map, that is, the final three-dimensional map.

[0050] Therefore, through the above process, the present invention can reconstruct the three-dimensional scene based on the inspection video data obtained by the monocular camera, and can provide reliable data support for daily inspection video analysis.

[0051] Implementation method of three-dimensional scene reconstruction device based on monocular camera

[0052] The monocular camera-based three-dimensional scene reconstruction device of the present invention includes a processor, wherein the processor is used to execute computer program instructions to implement the above-mentioned monocular camera-based three-dimensional scene reconstruction method. The specific implementation process of this method has been described in detail in the method implementation method and will not be repeated here.

Claims

1. A three-dimensional scene reconstruction method based on a monocular camera, characterized in that: The reconstruction method includes: 1) Obtain inspection video data captured by a monocular camera; 2) Use the 2D segmentation model to perform image segmentation on each data frame in the inspection video data to obtain a 2D segmentation map containing the segmentation target; use the depth estimation model to perform depth feature mining on each data frame in the inspection video data captured by the monocular camera to obtain a two-dimensional single-channel depth map containing each pixel point; 3) Fuse the obtained 2D segmentation map with the depth map to obtain a 3D point cloud map of each frame; 4) Use continuous frame 3D point cloud images to reconstruct the three-dimensional scene.

2. The method for reconstructing a three-dimensional scene based on a monocular camera according to claim 1, wherein: In step 3), when fusing the obtained 2D segmentation map with the depth map, the image coordinates of the segmentation targets in the 2D segmentation map are multiplied pixel by pixel by the depth values ​​of the pixels corresponding to each segmentation target in the depth map to determine the 3D coordinates of each segmentation target.

3. The method for 3D scene reconstruction based on a monocular camera according to claim 2, wherein: The step 3) adopts a dynamic weighted method to perform superposition, and the weight of the depth data is determined according to the confidence score of the depth value of each pixel in the image frame output by the depth estimation model. The higher the confidence score of the pixel depth value, the greater the corresponding weight, and the lower the confidence score of the pixel depth value, the smaller the corresponding weight.

4. The method for 3D scene reconstruction based on a monocular camera according to claim 2, wherein: The step 3) is performed by using a dynamic weighting method for superposition, where the weight of the depth data of the pixel point closer to the boundary of the segmentation target is smaller.

5. The method for reconstructing a three-dimensional scene based on a monocular camera according to claim 2, wherein: The step 3) performs superposition in a dynamic weighted manner, reducing the weight of the image segmentation data in a low-light environment and reducing the weight of the image depth data in a strong light environment.

6. The method for reconstructing a three-dimensional scene based on a monocular camera according to any one of claims 2 to 5, wherein: In step 3), when the obtained 2D segmentation map and the depth map are fused, the 2D segmentation map and the depth map are aligned, corresponding to the scene under the same perspective. If there is any misalignment between the two, correction is performed through geometric transformation or feature matching.

7. The method for 3D scene reconstruction based on a monocular camera according to claim 1, wherein: The process of reconstructing a three-dimensional scene using continuous frame 3D point cloud data in step 4) includes: matching the obtained 3D point cloud map of the current frame with the 3D point cloud map of the previous frame, fusing the matched result with the current global 3D point cloud map, and updating the global 3D point cloud map; the current global 3D point cloud map is obtained by fusing the 3D point cloud maps of the frames before the current frame; during fusion, if there is a new target in the current frame, the new target is added to the current global 3D point cloud map for updating; if there is a reduction in the number of targets in the current frame, the current global 3D point cloud map is not updated.

8. The method for reconstructing a three-dimensional scene based on a monocular camera according to any one of claims 2 to 5, wherein: The depth estimation model adopts the ZeroDepth model or the ZoeDepth model.

9. The method for reconstructing a three-dimensional scene based on a monocular camera according to any one of claims 2 to 5, wherein: In the step 2), each data frame in the inspection video data refers to an extracted key frame, and the key frame is a frame with a large semantic change or a frame with a depth difference exceeding a set threshold.

10. A three-dimensional scene reconstruction device based on a monocular camera, comprising a processor, characterized in that: The processor is used to execute computer program instructions to implement the three-dimensional scene reconstruction method based on a monocular camera according to any one of claims 1 to 9.