Three-dimensional map reconstruction method, device and equipment without dynamic object interference
Through the depth camera and Mask R-CNN network, dynamic objects are detected, combined with multi-view geometric verification and background repair technology, the problem of three-dimensional map construction accuracy and stability under dynamic object interference is solved, and a high-precision, three-dimensional map reconstruction without dynamic object interference is achieved.
Patent Information
- Application Number
- CN202510079057.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
When facing dynamic object interference, existing three-dimensional map reconstruction technology is difficult to build a high-precision, three-dimensional map without dynamic object interference, resulting in a decrease in map construction accuracy and stability.
Environmental data is collected through depth cameras, and dynamic objects are initially detected and pixel-by-pixel semantic segmentation is used to generate dynamic area masks. Combined with multi-view geometry verification technology, static feature points are screened, dynamic areas are eliminated, and background repair is carried out through multi-view information fusion and depth data interpolation, and the octree map is finally constructed.
It effectively solves the problems of mismatch, pose drift and background information loss caused by dynamic object interference in dynamic scenarios, and significantly improves the accuracy and quality of three-dimensional map reconstruction.
Smart Images

Figure CN119991986A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device and equipment for reconstructing a three-dimensional map without interference from dynamic objects. Background Art
[0002] After nearly three decades of development, Simultaneous Localization and Mapping (SLAM) technology has become a key technology in the research of robotics, automation and computer vision. SLAM can be seen everywhere in various fields, such as Micro Aerial Vehicle (MAV), Unmanned Ground Vehicle (UGV), autonomous driving, Virtual Reality (VR), and Augmented Reality (AR), all of which require SLAM technology to provide reliable positioning and mapping results.
[0003] The currently known simultaneous localization and mapping techniques all assume that the working scene is static and rigid. When these systems work in dynamic scenes, the erroneous data association caused by the static scene assumption will seriously weaken the accuracy and stability of the system. In general, the presence of dynamic objects in the scene divides all features into two categories: static features and dynamic features. How to detect and reject dynamic features is a key issue in research. Previous research work can be divided into three categories: geometric information method, semantic information method and method that integrates geometric information and semantic information.
[0004] The geometric information method detects dynamic objects through multi-view geometric constraints. The idea is that a static feature point in the current image must be located on the epipolar line corresponding to the same feature point in the previous image. If the distance between a feature point and the corresponding epipolar line exceeds the set empirical threshold, it is considered to be dynamic. Kundu et al.'s algorithm has the advantages of fast speed and strong scene generalization ability. However, it lacks a high-level understanding of the scene, so the empirical threshold of the algorithm is difficult to determine and the result is not accurate. The semantic information method obtains the semantic information of the dynamic area through deep learning technology, in which the YOLO target detection method is used to obtain the prior semantic information of dynamic objects in the working scene, and then the dynamic feature points are removed based on the semantic information to improve the tracking accuracy of the system. However, the way YOLO extracts semantic information through the bounding box will cause some static feature points to be mistakenly regarded as outliers and removed. In recent years, researchers have begun to explore the fusion method of geometric information and semantic information. For example, for RGB-D cameras, the semantic segmentation results of Mask R-CNN are currently used in combination with multi-view geometric constraints to detect dynamic objects and remove outliers. Although this method has made some progress in SLAM systems in dynamic scenes, it still has the following shortcomings in terms of dynamic obstacle interference: (1) In the process of matching a large number of dynamic features in the environment. Since dynamic features are prone to rapid changes, a large number of mismatches lead to an increase in data noise, which seriously affects the accuracy of map construction. Reducing the number of mismatches and improving the accuracy of feature matching, accurately identifying and matching dynamic feature points is one of the main difficulties. (2) Dynamic obstacles are constantly changing in the scene, and the motion trajectory of dynamic objects is unstable. It is necessary to ensure that the algorithm can accurately identify and track obstacles while having sufficient robustness to cope with the dynamic changes of obstacles. It is particularly difficult to design a dynamic target detection and tracking algorithm with strong adaptability and high real-time performance. (3) Background repair for the part occluded by dynamic obstacles and reconstruction of high-precision three-dimensional maps. Background repair technology not only needs to fill the texture of the occluded part. There are technical difficulties in synchronously reconstructing texture and depth information to maintain the authenticity of the occluded area. Background repair requires not only filling the texture details of the occluded part, but also requires restoring the depth information of the area while removing the object, thereby restoring the authenticity of the information of the occluded part, which has certain research difficulties.
[0005] In summary, the existing three-dimensional map reconstruction technology cannot well construct a high-precision, dynamic object-free three-dimensional map when faced with dynamic object interference, and cannot meet the urgent needs of various fields for high-quality three-dimensional maps. Summary of the invention
[0006] The embodiments of the present invention provide a method, device and equipment for reconstructing a three-dimensional map without dynamic object interference, which can effectively construct a high-precision three-dimensional map without dynamic object interference when facing dynamic object interference.
[0007] An embodiment of the present invention provides a method for reconstructing a three-dimensional map without interference from dynamic objects, comprising:
[0008] The environment data is collected by a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information;
[0009] Performing preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through the Mask R-CNN network to generate a dynamic region mask to separate the dynamic region from the static background; the dynamic region mask is used to mark the dynamic object region in the target image;
[0010] Performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy;
[0011] Combining the processed dynamic area mask, using a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask;
[0012] Based on the semantic segmentation mask, dynamic areas are eliminated, and the feature points of the target image that do not conform to the dynamic background are screened out in combination with the multi-view geometric verification technology, and the static parts are retained;
[0013] Using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation, background repair is performed on the target image's occluded area by the dynamic object to restore static background information;
[0014] The corresponding key frames of the repaired target image are selected to optimize the scene representation and the camera posture, and the static information of the key frames is used to construct an octree map, and the optimized scene representation and camera posture are integrated into the octree map to complete the construction of a three-dimensional map without interference from dynamic objects.
[0015] As an improvement of the above solution, after collecting environmental data through the depth camera to obtain the target image, the method also includes:
[0016] The target image is preprocessed to obtain a preprocessed target image; the preprocessing includes image graying, scale adjustment, denoising and smoothing operations.
[0017] As an improvement of the above solution, the target image is subjected to preliminary dynamic object detection and pixel-by-pixel semantic segmentation through the Mask R-CNN network to generate a dynamic region mask to separate the dynamic region from the static background, including:
[0018] Input the target image into the Mask R-CNN network, and generate candidate regions through the region proposal network; the candidate regions are regions in the image that may contain dynamic objects;
[0019] For each candidate region, bounding box regression, category classification and pixel-level segmentation are performed respectively through the branch structure of the Mask R-CNN network to generate a bounding box, category information and a dynamic region mask of the dynamic object; the bounding box includes the position coordinates of the dynamic object, the category information includes the category label of the dynamic object, and the dynamic region mask is used to mark the dynamic object region in the image;
[0020] According to the generated dynamic area mask, the dynamic area is separated from the static background; the dynamic object area marked by the dynamic area mask will be excluded from the static background.
[0021] As an improvement of the above solution, the segmentation accuracy processing is performed on the generated dynamic area mask to improve the segmentation accuracy, including:
[0022] Obtaining a dynamic region mask generated by performing preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network;
[0023] Performing morphological operations on the generated dynamic region mask, including erosion and dilation, to remove noise at the mask boundary and small misjudged regions; the morphological operations are used to improve the segmentation accuracy of the dynamic region mask;
[0024] The position of the bounding box of the dynamic object is updated by Kalman filtering or optical flow method to ensure accurate tracking of the position of the dynamic obstacle in consecutive frames; the position update is based on the bounding box information of the dynamic object to ensure the consistency of the position of the dynamic object in consecutive frames;
[0025] The processed dynamic area mask is used for subsequent dynamic object removal and background restoration; the dynamic object area marked by the processed dynamic area mask will be excluded from the three-dimensional map construction process and used to guide the background restoration of the area blocked by the dynamic object.
[0026] As an improvement of the above solution, the combined processed dynamic area mask uses a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask, including:
[0027] The processed dynamic area mask is used as input, and the target image is processed using a target detection algorithm to generate a bounding box and category information of the dynamic object; the bounding box includes the position coordinates of the dynamic object in the target image, and the category information includes the category label of the dynamic object;
[0028] The processed dynamic area mask and the generated bounding box and category information of the dynamic object are used as input, and the target image is subjected to pixel-by-pixel semantic segmentation operation through the Mask R-CNN network to generate a semantic segmentation mask of the dynamic object; the semantic segmentation mask is used to mark the dynamic object area in the image.
[0029] As an improvement of the above solution, the dynamic area is eliminated based on the semantic segmentation mask, and the feature points of the target image that do not conform to the dynamic background are screened in combination with the multi-view geometric verification technology, and the static part is retained, including:
[0030] Obtaining a generated semantic segmentation mask; the semantic segmentation mask is used to mark a dynamic object area in a target image;
[0031] According to the semantic segmentation mask, a dynamic area in the target image is eliminated; the dynamic area is a dynamic object area marked by the semantic segmentation mask;
[0032] Combined with multi-view geometric verification technology, feature points in the target image that do not conform to the dynamic background are screened out; the multi-view geometric verification technology calculates the projection error between the key frame and the current frame to screen out points that do not conform to the dynamic background characteristics;
[0033] The filtered static feature points are retained; the static feature points are used for subsequent three-dimensional map construction.
[0034] As an improvement of the above scheme, the method of using the multi-view view and the historical frames of the target image to perform background repair on the area of the target image blocked by the dynamic object through multi-view information fusion and depth data interpolation to restore the static background information includes:
[0035] Acquire a historical frame of the target image; the historical frame is an image frame captured before the target image, and contains static background information of the area blocked by the dynamic object;
[0036] Integrate the information of different perspectives in the multi-perspective view, perform fusion operation on the static part information of the target image under different perspectives, and obtain the static part information after multi-perspective fusion;
[0037] For the area blocked by the dynamic object in the target image, using a depth data interpolation method, taking the known depth information of the unblocked area as a reference, interpolating the depth data of the blocked area to obtain the depth information of the area;
[0038] According to the static part information after multi-view fusion and the depth information obtained by interpolating the depth data, the background repair operation is performed on the area blocked by the dynamic object, and the repaired information is used as the final static background information.
[0039] As an improvement of the above scheme, the corresponding key frame of the restored target image is selected to optimize the scene representation and camera posture, and the static information of the key frame is used to construct an octree map, and the optimized scene representation and camera posture are integrated into the octree map to complete the construction of a three-dimensional map without interference from dynamic objects, including:
[0040] Acquire a restored target image, wherein the restored target image includes static background information of a region blocked by a dynamic object, and the restored target image is obtained by performing background restoration on the region blocked by the dynamic object of the target image by using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation;
[0041] Selecting a corresponding key frame of the restored target image, wherein the key frame is a representative image frame in the restored target image and is screened from the restored target image by calculating inter-frame overlap and scene complexity;
[0042] According to the static feature points and depth information in the key frames, the geometric structure in the scene is optimized using multi-view geometric constraints to obtain an optimized scene representation;
[0043] Through multi-view geometric constraints and the projection relationship of feature points in different key frames, the rotation and translation parameters of the camera are optimized to obtain the optimized camera posture;
[0044] Convert the static information in different keyframes into nodes or units of the octree, present the three-dimensional space in a hierarchical manner, and construct an octree map;
[0045] The optimized scene representation and the information corresponding to the camera posture are integrated with the corresponding nodes or units in the octree map according to the structural rules of the octree to complete the construction of a three-dimensional map without interference from dynamic objects.
[0046] Another embodiment of the present invention provides a three-dimensional map reconstruction device without dynamic object interference, including:
[0047] An image acquisition module is used to acquire environmental data through a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information;
[0048] An image separation module, used to perform preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network, generate a dynamic region mask, and separate the dynamic region from the static background; the dynamic region mask is used to mark the dynamic object region in the target image;
[0049] A processing module, used for performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy;
[0050] A recognition module, used to combine the processed dynamic area mask, use the target detection algorithm to recognize the dynamic objects in the target image again and generate a semantic segmentation mask;
[0051] A screening module, used to remove dynamic areas based on the semantic segmentation mask, and screen feature points of the target image that do not conform to the dynamic background in combination with multi-view geometric verification technology, and retain static parts;
[0052] A restoration module, used to perform background restoration on the target image's occluded area by the dynamic object by using the multi-view views and the target image's historical frames, through multi-view information fusion and depth data interpolation, to restore static background information;
[0053] The reconstruction module is used to select the corresponding key frames of the restored target image to optimize the scene representation and camera posture, and use the static information of the key frames to build an octree map, integrate the optimized scene representation and camera posture into the octree map, and complete the three-dimensional map construction without dynamic object interference.
[0054] Another embodiment of the present invention provides a three-dimensional map reconstruction device without interference from dynamic objects, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the three-dimensional map reconstruction method without interference from dynamic objects described in the above-mentioned embodiment of the invention is implemented.
[0055] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0056] First, the depth camera is used to collect environmental data to obtain the target image, including continuous image frames or RGB-D images with depth information. Based on this, the Mask R-CNN network is used to perform preliminary detection and pixel-by-pixel semantic segmentation of dynamic objects, generate dynamic area masks, and achieve preliminary separation of dynamic areas from static backgrounds. This is because the powerful feature extraction and classification capabilities of the Mask R-CNN network can effectively distinguish different objects and provide important dynamic area information for subsequent processing. Then, the dynamic area mask is processed for segmentation accuracy. Methods such as morphological operations can be used to refine the boundaries of the dynamic area, improve segmentation accuracy, and ensure the accuracy of subsequent processing. Subsequently, combined with the processed dynamic area mask, the target detection algorithm is used to identify dynamic objects again and generate semantic segmentation masks to further refine the marking of dynamic objects. Thanks to the combined processed mask information, it can focus on possible dynamic areas and improve recognition accuracy. Based on the semantic segmentation mask, the dynamic area is eliminated, and the multi-view geometric verification technology is used to filter out feature points that do not conform to the dynamic background. The multi-view geometric verification technology uses the geometric relationship between multi-view images to ensure the accuracy of screening, thereby retaining the static part. Then, the multi-view view and historical frame are used to repair the area blocked by dynamic objects through multi-view information fusion and deep data interpolation. This is to utilize the complementarity of multi-view information and the advantages of deep data interpolation to restore more complete and accurate static background information. Finally, key frames are selected from the repaired target image, the scene representation and camera posture are optimized, and the octree map is constructed using the static information of the key frames. The optimized information is integrated into the octree. The octree structure is conducive to the hierarchical storage and representation of three-dimensional space, and the construction of a three-dimensional map without interference from dynamic objects is completed. The optimization of key frames and the use of octrees in the construction process can effectively integrate static information and improve the quality of map construction. From the above analysis, it can be seen that the embodiment of the present invention effectively solves the problems of mismatching, posture drift, and background information loss caused by dynamic object interference in dynamic scenes through technical means such as Mask R-CNN network, segmentation accuracy processing, multi-view geometry verification technology, multi-view information fusion, deep data interpolation, and octree map construction, and significantly improves the accuracy of three-dimensional map reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of a method for reconstructing a three-dimensional map without interference from dynamic objects provided by an embodiment of the present invention;
[0058] Figure 2 It is a technical roadmap provided by an embodiment of the present invention;
[0059] Figure 3 It is a technical idea diagram provided by an embodiment of the present invention;
[0060] Figure 4It is a technical research framework diagram provided by an embodiment of the present invention;
[0061] Figure 5 It is a schematic diagram of the principle of multi-view geometric detection of dynamic points provided by an embodiment of the present invention;
[0062] Figure 6 is a Mask R-CNN framework diagram for instance segmentation provided by an embodiment of the present invention;
[0063] Figure 7 This is a dynamic obstacle segmentation effect diagram provided by an embodiment of the present invention;
[0064] Figure 8 is a schematic diagram of a dynamic object culling algorithm flow chart provided by an embodiment of the present invention;
[0065] Fig. 9 This is a dynamic feature elimination effect diagram provided by an embodiment of the present invention;
[0066] Fig.10 This is a schematic diagram of the octree mapping principle provided by an embodiment of the present invention;
[0067] Fig.11 This is an octree map construction effect diagram provided by an embodiment of the present invention;
[0068] Fig.12 is a schematic diagram of a background repair process provided by an embodiment of the present invention;
[0069] Fig.13 This is a three-dimensional mapping effect diagram provided by an embodiment of the present invention;
[0070] Fig.14 It is a structural schematic diagram of a three-dimensional map reconstruction device without dynamic object interference provided by an embodiment of the present invention;
[0071] Fig.15 It is a structural schematic diagram of a three-dimensional map reconstruction device without dynamic object interference provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0073] See also Figure 1, is a flow chart of a method for reconstructing a three-dimensional map without interference from dynamic objects provided by an embodiment of the present invention. The method for reconstructing a three-dimensional map without interference from dynamic objects comprises:
[0074] The environment data is collected by a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information;
[0075] Performing preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through the Mask R-CNN network to generate a dynamic region mask to separate the dynamic region from the static background; the dynamic region mask is used to mark the dynamic object region in the target image;
[0076] Performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy;
[0077] Combining the processed dynamic area mask, using a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask;
[0078] Based on the semantic segmentation mask, dynamic areas are eliminated, and the feature points of the target image that do not conform to the dynamic background are screened out in combination with the multi-view geometric verification technology, and the static parts are retained;
[0079] Using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation, background repair is performed on the target image's occluded area by the dynamic object to restore static background information;
[0080] The corresponding key frames of the repaired target image are selected to optimize the scene representation and the camera posture, and the static information of the key frames is used to construct an octree map, and the optimized scene representation and camera posture are integrated into the octree map to complete the construction of a three-dimensional map without interference from dynamic objects.
[0081] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0082] First, the depth camera is used to collect environmental data to obtain the target image, including continuous image frames or RGB-D images with depth information. Based on this, the Mask R-CNN network is used to perform preliminary detection and pixel-by-pixel semantic segmentation of dynamic objects, generate dynamic area masks, and achieve preliminary separation of dynamic areas from static backgrounds. This is because the powerful feature extraction and classification capabilities of the Mask R-CNN network can effectively distinguish different objects and provide important dynamic area information for subsequent processing. Then, the dynamic area mask is processed for segmentation accuracy. Methods such as morphological operations can be used to refine the boundaries of the dynamic area, improve segmentation accuracy, and ensure the accuracy of subsequent processing. Subsequently, combined with the processed dynamic area mask, the target detection algorithm is used to identify dynamic objects again and generate semantic segmentation masks to further refine the marking of dynamic objects. Thanks to the combined processed mask information, it can focus on possible dynamic areas and improve recognition accuracy. Based on the semantic segmentation mask, the dynamic area is eliminated, and the multi-view geometric verification technology is used to filter out feature points that do not conform to the dynamic background. The multi-view geometric verification technology uses the geometric relationship between multi-view images to ensure the accuracy of screening, thereby retaining the static part. Then, the multi-view view and historical frame are used to repair the area blocked by dynamic objects through multi-view information fusion and deep data interpolation. This is to utilize the complementarity of multi-view information and the advantages of deep data interpolation to restore more complete and accurate static background information. Finally, key frames are selected from the repaired target image, the scene representation and camera posture are optimized, and the octree map is constructed using the static information of the key frames. The optimized information is integrated into the octree. The octree structure is conducive to the hierarchical storage and representation of three-dimensional space, and the construction of a three-dimensional map without interference from dynamic objects is completed. The optimization of key frames and the use of octrees in the construction process can effectively integrate static information and improve the quality of map construction. From the above analysis, it can be seen that the embodiment of the present invention effectively solves the problems of mismatching, posture drift, and background information loss caused by dynamic object interference in dynamic scenes through technical means such as MaskR-CNN network, segmentation accuracy processing, multi-view geometry verification technology, multi-view information fusion, deep data interpolation, and octree map construction, and significantly improves the accuracy of three-dimensional map reconstruction.
[0083] For example, sensors such as RGB cameras, depth cameras or lidar are used to collect environmental data and generate a series of continuous image frames or RGB-D images with depth information. These data will provide a basis for dynamic obstacle detection and segmentation.
[0084] As an improvement of the above embodiment, after collecting environmental data through the depth camera to obtain the target image, before the above, the method further includes:
[0085] The target image is preprocessed to obtain a preprocessed target image; the preprocessing includes image graying, scale adjustment, denoising and smoothing operations.
[0086] In this embodiment, in order to reduce the impact of image noise and illumination changes on subsequent detection, each frame of the image needs to be preprocessed. These operations usually include image grayscale, scale adjustment, denoising and smoothing, etc., to reduce the interference of noise and illumination changes on detection. The preprocessed image can more clearly reflect the shape and details of the target object. The preprocessed image can be expressed as: t '=Preprocess(I t ).
[0087] As an improvement of the above embodiment, the target image is subjected to preliminary dynamic object detection and pixel-by-pixel semantic segmentation by using a Mask R-CNN network to generate a dynamic region mask to separate the dynamic region from the static background, including:
[0088] Input the target image into the Mask R-CNN network, and generate candidate regions through the region proposal network; the candidate regions are regions in the image that may contain dynamic objects;
[0089] For each candidate region, bounding box regression, category classification and pixel-level segmentation are performed respectively through the branch structure of the Mask R-CNN network to generate a bounding box, category information and a dynamic region mask of the dynamic object; the bounding box includes the position coordinates of the dynamic object, the category information includes the category label of the dynamic object, and the dynamic region mask is used to mark the dynamic object region in the image;
[0090] According to the generated dynamic area mask, the dynamic area is separated from the static background; the dynamic object area marked by the dynamic area mask will be excluded from the static background.
[0091] In this embodiment, the powerful image analysis capability of the Mask R-CNN network is used to perform multi-dimensional processing on the target image, thereby accurately separating the dynamic area from the static background, laying the foundation for the subsequent construction of a high-precision three-dimensional map without interference from dynamic objects. Its technical effect is that through a series of orderly operations, it can accurately identify the area where the dynamic object is located in the complex target image, generate the corresponding bounding box, category information and dynamic area mask, realize the clear division of the dynamic area and the static background, effectively avoid the dynamic object information from mixing into the static background, and then ensure that the subsequent map construction link can be carried out based on accurate static background information, and improve the accuracy and quality of the entire three-dimensional map reconstruction.
[0092] Specifically, the first step is to input the target image into the Mask R-CNN network. As an advanced and mature deep learning architecture, the Region Proposal Network (RPN) inside this network will take the lead in playing a role. Based on information such as pixel features in the target image, the Region Proposal Network uses the feature extraction capability of the Convolutional Neural Network (CNN) to slide a window on the image and generate many candidate regions that may contain dynamic objects through a series of convolution, pooling and other operations. These candidate regions are essentially some local areas in the image that have been preliminarily screened and are considered to have a high probability of containing dynamic objects. Their role is to narrow the scope of subsequent processing and focus on parts of the image where dynamic objects are more likely to appear, avoiding indiscriminate and computationally intensive processing of the entire image, and improving processing efficiency and accuracy. In the second step, for each candidate region, the branch structure of the Mask R-CNN network is further used to perform key bounding box regression, category classification and pixel-level segmentation operations. In the bounding box regression operation, the specially designed regression module in the Mask R-CNN network is used. This module will be based on the rich feature information extracted from the candidate region, by learning the deviation relationship between the real boundary position of the dynamic object in the image and the boundary of the candidate region, and using appropriate loss functions (such as SmoothL1 loss function, etc.) to continuously optimize and adjust the coordinate parameters of the bounding box, and finally generate an accurate bounding box of the dynamic object. This bounding box contains the accurate position coordinates of the dynamic object in the target image. For example, with the upper left corner of the image as the origin, the horizontal right as the positive direction of the x-axis, and the vertical downward as the positive direction of the y-axis, the coordinate information of the bounding box (x1, y1, x2, y2) clearly defines the specific position range of the dynamic object on the two-dimensional image plane, providing a quantitative basis for the subsequent accurate judgment of the spatial position of the dynamic object. For the category classification operation, the classification branch of the Mask R-CNN network will match and judge the features in each candidate region based on the feature patterns of various objects learned in advance on a large-scale annotated dataset. It maps the feature vectors of the candidate area to the corresponding category space through structures such as multi-layer fully connected layers, and finally outputs the category information of the dynamic object, that is, to clarify the specific category label to which it belongs, such as pedestrians, vehicles, animals and other specific categories. This classification result is of great significance for subsequent targeted processing and analysis based on the characteristics of different types of dynamic objects, as well as distinguishing the impact of different dynamic objects in complex scenes. The pixel-level segmentation operation is achieved through the unique mask generation branch in the Mask R-CNN network. This branch will perform a detailed semantic analysis on each pixel in the candidate area, assign corresponding semantic values according to the degree of association between the pixel and the dynamic object, and finally generate a dynamic area mask after a complex calculation and mapping process.The mask marks the dynamic object area in the image with pixel-level accuracy. For example, in the mask image, the pixels belonging to the dynamic object are marked with a specific value (such as 1), and the pixels belonging to the static background are marked with another value (such as 0). In this way, the boundary between the dynamic area and the static background can be clearly demarcated with extremely high resolution, providing an accurate basis for the subsequent complete separation of the two. The last step is to separate the dynamic area from the static background based on the generated dynamic area mask. Specifically, at the level of the entire target image, according to the pixel values marked by the dynamic area mask, those parts marked as dynamic object areas (that is, the corresponding pixel values are specific values representing dynamic objects) are clearly separated from the overall image, so that they are completely isolated from the static background area. These dynamic object areas marked by the dynamic area mask will be strictly excluded from the static background, thereby ensuring the purity and integrity of the static background information. This provides accurate and dynamic-free static background basic data for a series of subsequent operations such as further dynamic object recognition based on the octree data structure, screening of feature points using multi-view geometry verification technology, and background repair, ensuring that the entire 3D map reconstruction process can proceed in the correct and precise direction.
[0093] As an improvement of the above embodiment, the step of performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy includes:
[0094] Obtaining a dynamic region mask generated by performing preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network;
[0095] Performing morphological operations on the generated dynamic region mask, including erosion and dilation, to remove noise at the mask boundary and small misjudged regions; the morphological operations are used to improve the segmentation accuracy of the dynamic region mask;
[0096] The position of the bounding box of the dynamic object is updated by Kalman filtering or optical flow method to ensure accurate tracking of the position of the dynamic obstacle in consecutive frames; the position update is based on the bounding box information of the dynamic object to ensure the consistency of the position of the dynamic object in consecutive frames;
[0097] The processed dynamic area mask is used for subsequent dynamic object removal and background restoration; the dynamic object area marked by the processed dynamic area mask will be excluded from the three-dimensional map construction process and used to guide the background restoration of the area blocked by the dynamic object.
[0098] In this embodiment, after obtaining the initial dynamic area mask, morphological operations are used to remove noise at the mask boundary and small areas of misjudgment. At the same time, the position of the dynamic object boundary box is updated with the help of Kalman filtering or optical flow method to ensure the accuracy of position tracking in continuous frames. Ultimately, the processed dynamic area mask can more accurately mark the dynamic object area, effectively reduce the errors caused by insufficient accuracy of the dynamic area mask in subsequent processing links, improve the accuracy of dynamic object removal and the effect of background restoration, and thus enhance the quality and accuracy of the entire three-dimensional map construction, so that it can better serve various application scenarios that rely on high-precision maps.
[0099] Specifically, we first need to obtain the dynamic region mask generated by preliminary detection of dynamic objects and pixel-by-pixel semantic segmentation of the target image through the Mask R-CNN network. This mask is the preliminary mark of the region where the dynamic object is located after analyzing the image based on the deep learning network in the previous step, but there may be certain accuracy problems, such as unclear boundaries and small-scale misjudgments, so further processing is required. Then, morphological operations are performed on the generated dynamic region mask, which mainly involve two basic operations: corrosion and dilation. The corrosion operation is to use a structural element (usually a matrix of a specific shape and size, such as a rectangle, circle, etc.) to perform convolution processing on the dynamic region mask. In this process, the structural element slides on the mask image. When part of the area covered by the structural element belongs to the background pixel, the corresponding central pixel of the area is judged as the background pixel, so that the boundary pixels of the dynamic region are gradually "eroded", so as to remove the boundary noise and some isolated, small-area areas that are misjudged as dynamic objects, making the boundary of the dynamic region clearer and more regular. The dilation operation is the opposite. It is also based on the sliding of the structural element on the mask image. However, as long as there are some pixels belonging to dynamic objects in the area covered by the structural element, the central pixel corresponding to the area is determined as a dynamic object pixel. In this way, some holes inside the dynamic area can be filled or the dynamic object pixels that may be excessively removed by the corrosion operation can be made up, so that the integrity of the dynamic area is improved, and the quality of the dynamic area mask is further optimized to improve its segmentation accuracy. Then, the position of the bounding box of the dynamic object is updated by Kalman filtering or optical flow method. Kalman filtering is an algorithm based on the linear system state equation, which makes the optimal estimation of the system state through the input and output observation data of the system. In this scenario, it will predict the position of the dynamic object bounding box in the current frame based on the position, speed and other state information of the dynamic object bounding box in the previous frame (this information can be obtained based on the previous processing steps), combined with the observation data of the current frame (such as the feature information detected in the image that may be related to the dynamic object), and then make corrections based on the new position information actually observed, and continuously iterate and update, so as to ensure that the position of the dynamic object in the continuous frames can be accurately tracked, ensure its position consistency, and avoid position deviation caused by factors such as object motion and image noise. The optical flow method mainly calculates the movement speed and direction of the pixel points based on the grayscale change of the pixel points in the image, and then infers the movement of the dynamic object. By analyzing the corresponding relationship between the pixels in adjacent frames, the accurate position change of the dynamic object bounding box in different frames is determined, and the position of the bounding box is updated, so that the position of the dynamic object in the continuous image frames can be accurately located, providing accurate dynamic object position information for subsequent processing. Finally, the dynamic area mask after the above processing is used for the subsequent dynamic object removal and background repair.The processed dynamic area mask has improved accuracy, and the marked dynamic object areas can be accurately identified. In the subsequent 3D map construction process, these accurately marked dynamic object areas will be strictly excluded to prevent them from interfering with static background information and map construction. At the same time, when performing background repair on areas occluded by dynamic objects, the dynamic area mask can play a guiding role. For example, according to the marked dynamic area range, the surrounding unobstructed static background information and other related data are reasonably used to restore the background of the occluded area through the corresponding repair algorithm, making the background repair work more accurate and effective, and laying a solid foundation for the final construction of a high-precision 3D map without interference from dynamic objects.
[0100] As an improvement of the above embodiment, the combined processed dynamic area mask uses a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask, including:
[0101] The processed dynamic area mask is used as input, and the target image is processed using a target detection algorithm to generate a bounding box and category information of the dynamic object; the bounding box includes the position coordinates of the dynamic object in the target image, and the category information includes the category label of the dynamic object;
[0102] The processed dynamic area mask and the generated bounding box and category information of the dynamic object are used as input, and the target image is subjected to pixel-by-pixel semantic segmentation operation through the Mask R-CNN network to generate a semantic segmentation mask of the dynamic object; the semantic segmentation mask is used to mark the dynamic object area in the image.
[0103] In this embodiment, by combining the processed dynamic area mask and the target detection algorithm, accurate recognition and semantic segmentation of dynamic objects in the target image are achieved, thereby generating a high-quality semantic segmentation mask. Specifically, the target image is first processed using a target detection algorithm (such as YOLO, FasterR-CNN or SSD) to generate a bounding box and category information of the dynamic object, and to clarify the position and type of the dynamic object; then, the processed dynamic area mask and the generated bounding box and category information are used as input, and a pixel-by-pixel semantic segmentation operation is performed through the Mask R-CNN network to generate a semantic segmentation mask for the dynamic object. The mask can accurately mark the dynamic object area in the image, ensuring that the edges of the dynamic object are clear and separated from the background. Through this technical solution, the system can effectively reduce the interference of dynamic objects on map construction, improve the accuracy and stability of the three-dimensional map, and provide a reliable foundation for subsequent static background repair and map construction.
[0104] Specifically, first, the processed dynamic area mask is used as input, and the target image is processed by the target detection algorithm to generate the bounding box and category information of the dynamic object. In this process, the target detection algorithm will extract and analyze the features of the target image based on the area information where the dynamic object may exist provided by the dynamic area mask. For feature extraction, a variety of feature descriptors may be used, such as gradient-based features, texture features, or color histograms. By extracting these features in the area that may contain dynamic objects, combined with machine learning or deep learning classification methods, it is determined to which dynamic object category the area belongs, and its bounding box position in the target image is determined. Specifically, the target detection algorithm will first use the dynamic area mask as prior information, narrow the detection range, and search for features only in the part marked as the dynamic area. For generating the bounding box, it will use a regression algorithm (such as linear regression or bounding box regression in deep learning) based on the extracted feature information to find the bounding box that best suits the dynamic object, accurately determine its position coordinates, and accurately delineate the range of the dynamic object. The generation of category information depends on the classification algorithm, which may be the use of support vector machines (SVM) or classifiers in deep learning (such as fully connected layers plus Softmax functions) to map the extracted features to different category labels, thereby determining the category of dynamic objects, such as pedestrians, vehicles, animals, etc. Then, the processed dynamic area mask and the generated bounding box and category information of the dynamic object are used as input, and the target image is subjected to pixel-by-pixel semantic segmentation through the Mask R-CNN network to generate a semantic segmentation mask for the dynamic object. In this link, the Mask R-CNN network will use its powerful convolutional neural network structure to first extract and fuse the input information. It will fuse features at different levels through multiple convolutional layers and pooling layers, so that the network can capture richer semantic information in the image. When performing pixel-by-pixel semantic segmentation, the Mask R-CNN network will use the bounding box and category information obtained previously to refine the segmentation task to the pixel level. For each pixel, the network will predict the probability of it belonging to a dynamic object, and assign the corresponding category label to the pixel belonging to the dynamic object, and finally generate a semantic segmentation mask. Here, the structure of a fully convolutional network (FCN) may be used to map the output of the network from the feature map back to the pixel size of the original image, and ensure that the semantic segmentation mask has the same resolution as the original image through upsampling and deconvolution operations. For example, in the branch structure of the Mask R-CNN network, one branch is responsible for generating bounding boxes and category information, and the other branch is responsible for generating pixel-by-pixel mask information.In this way, the final semantic segmentation mask can accurately mark the dynamic object area in the image and accurately mark each pixel as belonging to a dynamic object or a static background, providing a more precise basis for subsequent dynamic area elimination and precise extraction of the static background, and helping to achieve more accurate three-dimensional map construction without interference from dynamic objects.
[0105] As an improvement of the above embodiment, the method of removing the dynamic area based on the semantic segmentation mask and combining the multi-view geometric verification technology to screen the feature points of the target image that do not conform to the dynamic background and retain the static part includes:
[0106] Obtaining a generated semantic segmentation mask; the semantic segmentation mask is used to mark a dynamic object area in a target image;
[0107] According to the semantic segmentation mask, a dynamic area in the target image is eliminated; the dynamic area is a dynamic object area marked by the semantic segmentation mask;
[0108] Combined with multi-view geometric verification technology, feature points in the target image that do not conform to the dynamic background are screened out; the multi-view geometric verification technology calculates the projection error between the key frame and the current frame to screen out points that do not conform to the dynamic background characteristics;
[0109] The filtered static feature points are retained; the static feature points are used for subsequent three-dimensional map construction.
[0110] In this embodiment, semantic segmentation mask and multi-view geometric verification technology are used to accurately remove dynamic areas from the target image and filter out feature points that meet the characteristics of the static background to ensure that the subsequent three-dimensional map construction process is based only on pure static information. Among them, the dynamic area is marked by an accurate semantic segmentation mask to effectively remove it from the target image, and the multi-view geometric verification technology is used to further filter out feature points that may be misjudged as static but do not actually meet the dynamic background, thereby ensuring that the retained static part is highly accurate, providing a high-quality static information source for building a three-dimensional map without interference from dynamic objects, significantly improving the quality and accuracy of the final three-dimensional map, and meeting the demand for high-precision maps.
[0111] Specifically, first, obtaining the generated semantic segmentation mask is the starting point of this step. The semantic segmentation mask is obtained through a series of previous processes. It represents the dynamic object area in the target image in a specific pixel marking method. For example, different pixel values are used to distinguish dynamic and static areas. The pixels in the dynamic area may be marked as 1, while the pixels in the static area may be marked as 0. Next, the dynamic area in the target image is eliminated according to the semantic segmentation mask. In the specific implementation, each pixel of the target image is traversed. When it is detected that the pixel is marked as a dynamic area in the semantic segmentation mask (such as a pixel value of 1), the pixel is removed from the target image or set to a background pixel value, so as to eliminate the dynamic area. This may involve a pixel-level screening and processing algorithm to ensure that the dynamic area is completely excluded to avoid interference with subsequent processing. Then, the feature points in the target image that do not conform to the dynamic background are screened in combination with the multi-view geometry verification technology. In this link, the principle of multi-view geometry is applied to the key frame and the current frame by calculating the projection matrix between them. Specifically, the projection formula based on the camera intrinsic and extrinsic parameters is used to project the feature points in the key frame to the current frame, and then the projection error is calculated. The projection formula here can be x′=K[R|t]X, where K is the camera intrinsic parameter matrix, [R|t] is the extrinsic parameter matrix, X is the coordinates of the feature point in three-dimensional space, and x′ is the two-dimensional coordinate projected to the current frame. After calculating the projection error, a reasonable error threshold is set. For feature points whose projection error exceeds the threshold, they are judged as feature points that do not conform to the dynamic background. When calculating the projection error, the Euclidean distance formula is used To calculate the distance between the projection point and the actual point, use it as an error metric to filter out points that do not conform to the dynamic background features. Finally, retain the filtered static feature points. These static feature points have been strictly screened to exclude dynamic areas and feature points that do not conform to the dynamic background, ensuring that they can truly reflect the static background information and provide a reliable foundation for subsequent three-dimensional map construction. When storing and managing these static feature points, data structures such as arrays or linked lists may be used to store them in order so that subsequent algorithms can easily access and use this information, thereby providing accurate static information support for building high-quality three-dimensional maps. Through this series of steps, the static part of the target image is accurately extracted, avoiding the adverse effects of dynamic objects and their possible misjudged feature points on subsequent map construction, which helps to ultimately achieve high-precision three-dimensional map construction without interference from dynamic objects.
[0112] As an improvement of the above embodiment, the method of using the multi-view view and the historical frames of the target image to perform background repair on the area of the target image blocked by the dynamic object through multi-view information fusion and depth data interpolation to restore the static background information includes:
[0113] Acquire a historical frame of the target image; the historical frame is an image frame captured before the target image, and contains static background information of the area blocked by the dynamic object;
[0114] Integrate the information of different perspectives in the multi-perspective view, perform fusion operation on the static part information of the target image under different perspectives, and obtain the static part information after multi-perspective fusion;
[0115] For the area blocked by the dynamic object in the target image, using a depth data interpolation method, taking the known depth information of the unblocked area as a reference, interpolating the depth data of the blocked area to obtain the depth information of the area;
[0116] According to the static part information after multi-view fusion and the depth information obtained by interpolating the depth data, the background repair operation is performed on the area blocked by the dynamic object, and the repaired information is used as the final static background information.
[0117] In this embodiment, the multi-perspective view and the historical frames of the target image are comprehensively utilized, and the background repair problem of the area occluded by the dynamic object in the target image is solved by means of multi-perspective information fusion and depth data interpolation, aiming to restore complete and accurate static background information, and provide a high-quality static background foundation for the subsequent construction of a three-dimensional map without interference from dynamic objects. This embodiment can make full use of the information advantages of multi-perspective and historical frames, and fill in the missing static background information that is occluded by dynamic objects through fusion and interpolation operations, significantly improving the integrity and accuracy of the static background information, thereby improving the reconstruction quality of the entire three-dimensional map, and ensuring that it shows better results in actual application scenarios.
[0118] Specifically, first of all, obtaining the historical frames of the target image is a basic step. These historical frames are image frames collected before the target image. They contain important information, especially static background information that may not be blocked by dynamic objects in the previous frames. This information is very critical for the background repair of the current area blocked by dynamic objects. In actual operation, the historical frames may be stored in a data storage system, such as using a database or file system, and they are arranged in order according to timestamps or other identifiers to facilitate subsequent rapid search and call as needed. Then, the information of different perspectives in the multi-perspective view is integrated, and the static part information of the target image under different perspectives is fused to obtain the static part information after multi-perspective fusion. In this process, image registration technology is used to ensure that images of different perspectives can be accurately aligned. For example, a feature-based image registration method is used to extract feature points (such as SIFT, SURF or ORB feature points) and use feature matching algorithms (such as RANSAC algorithms) to find the correspondence between images of different perspectives, thereby achieving accurate image alignment. For the fusion of static partial information, the weighted average method can be used to assign different weights according to the reliability or importance of different perspectives, and weighted sum the static partial information under each perspective. For each pixel position, the fused pixel value can be expressed as Where P i is the pixel value of the i-th viewing angle, w i is the corresponding weight, and In this way, the advantages of different perspectives can be combined to improve the accuracy and robustness of the static part information. Next, for the area in the target image that is occluded by the dynamic object, the depth information is processed using the depth data interpolation method. Based on the known depth information of the unoccluded area, a variety of algorithms can be used to interpolate the depth information of the occluded area, such as bilinear interpolation, bicubic interpolation, or an interpolation method based on a spline function. Taking bilinear interpolation as an example, for a certain pixel point (x, y) in the occluded area, its depth value D(x, y) can be calculated by the surrounding four pixels with known depth values (x1, y1), (x1, y2), (x2, y1) and (x2, y2), and the calculation formula refers to the existing method. In this way, the depth information of the surrounding unoccluded area is used to reasonably estimate the depth of the occluded area, providing information support in the depth dimension for subsequent background restoration. Finally, based on the static part information after multi-perspective fusion and the depth information obtained by depth data interpolation, the background restoration operation is performed on the area occluded by the dynamic object. Here, a background repair algorithm based on texture synthesis may be used, and the fused static part information is used as texture information. Combined with the depth information, the pixel values of the occluded area are restored by copying, filling or deforming the texture. For example, the relative position relationship between the occluded area and the surrounding unoccluded area is determined based on the depth information, and the appropriate texture block is selected from the static part information after multi-view fusion based on this relationship, and it is copied or deformed to fill the occluded area, so that the repaired information and the surrounding static background are naturally transitioned, and finally the repaired information is used as the final static background information. In this way, a complete and accurate static background can be provided for the subsequent three-dimensional map construction, avoiding the impact of the dynamic object occlusion area, and ensuring the high-quality construction of the three-dimensional map.
[0119] As an improvement of the above embodiment, the corresponding key frame of the restored target image is selected to optimize the scene representation and the camera posture, and the static information of the key frame is used to construct an octree map, and the optimized scene representation and the camera posture are integrated into the octree map to complete the construction of a three-dimensional map without interference from dynamic objects, including:
[0120] Acquire a restored target image, wherein the restored target image includes static background information of a region blocked by a dynamic object, and the restored target image is obtained by performing background restoration on the region blocked by the dynamic object of the target image by using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation;
[0121] Selecting a corresponding key frame of the restored target image, wherein the key frame is a representative image frame in the restored target image and is screened from the restored target image by calculating inter-frame overlap and scene complexity;
[0122] According to the static feature points and depth information in the key frames, the geometric structure in the scene is optimized using multi-view geometric constraints to obtain an optimized scene representation;
[0123] Through multi-view geometric constraints and the projection relationship of feature points in different key frames, the rotation and translation parameters of the camera are optimized to obtain the optimized camera posture;
[0124] Convert the static information in different keyframes into nodes or units of the octree, present the three-dimensional space in a hierarchical manner, and construct an octree map;
[0125] The optimized scene representation and the information corresponding to the camera posture are integrated with the corresponding nodes or units in the octree map according to the structural rules of the octree to complete the construction of a three-dimensional map without interference from dynamic objects.
[0126] In this embodiment, by selecting key frames from the repaired target image, and then using the information in the key frames, the scene representation and camera posture are optimized through multi-view geometric constraints, and the optimized information is integrated into the octree map, so as to complete the construction of a three-dimensional map without interference from dynamic objects. This embodiment can achieve accurate representation of the three-dimensional scene by carefully selecting key frames and optimizing the scene and camera posture on the basis of the completed background repair, and finally present the static information in the form of an octree map, which can not only efficiently store and visualize three-dimensional information, but also eliminate the interference of dynamic objects, provide high-quality maps for the application of three-dimensional maps in the fields of autonomous driving, virtual reality, etc., and improve the accuracy and reliability of scene reconstruction.
[0127] Specifically, first, the restored target image is obtained. This step is based on the previous background restoration work. The background of the area occluded by the dynamic object is restored by using multi-view views and historical frames of the target image through multi-view information fusion and depth data interpolation technology. In the specific implementation, for multi-view information fusion, a feature point matching method may be used, such as using SIFT or SURF feature point algorithms to find matching feature points between images of different perspectives, and fuse the information of different perspectives based on these matching points. Depth data interpolation uses known depth information. For the occluded area, according to the depth information of the surrounding unoccluded areas, algorithms such as bilinear interpolation or bicubic interpolation are used to fill the depth data, and finally obtain the restored target image containing complete static background information. Next, the corresponding key frame of the restored target image is selected. Here, representative key frames are selected by calculating the overlap between frames and the complexity of the scene. For the calculation of the overlap between frames, the result of feature point matching can be used to measure by calculating the number of matching points or the proportion of matching points to the total feature points. The evaluation of scene complexity can be based on the texture complexity, depth change, edge information, etc. of the image. For example, the gradient amplitude and direction are calculated using a gradient operator (such as the Sobel operator) and used as an indicator for evaluating scene complexity. Taking these factors into consideration, key frames are selected from the restored target image to avoid redundant information and ensure that the key frames can effectively reflect the key information of the scene. Then, based on the static feature points and depth information in the key frames, the geometric structure in the scene is optimized using multi-view geometric constraints to obtain an optimized scene representation. In this process, multi-view geometric constraints will reconstruct the three-dimensional points in the scene based on the geometric relationship between different key frames and the position and depth information of the feature points in different key frames through the triangulation principle, thereby optimizing the geometric structure of the scene. For example, using the corresponding feature points in multiple sets of key frames, according to their projection relationship at different perspectives, the coordinates of the points in the three-dimensional space are calculated through the triangulation formula, and then the geometric elements in the scene, such as planes and curves, are adjusted based on these three-dimensional points to achieve the purpose of optimizing the scene representation. For the optimization of camera posture, the rotation and translation parameters of the camera are optimized through multi-view geometric constraints and the projection relationship of feature points in different key frames to obtain the optimized camera posture. In the specific implementation, the projection equation is established according to the projection position relationship of the feature points in different key frames, such as using the essential matrix or homography matrix, and then the rotation and translation parameters of the camera are obtained by solving these equations. Optimization algorithms such as nonlinear least squares can be used to minimize the error between the actual projection point and the projection point calculated according to the current camera posture, and the camera posture is continuously optimized iteratively. The static information in different key frames is converted into nodes or units of the octree, and the three-dimensional space is presented in a hierarchical manner to construct an octree map.The specific operation is to recursively divide the three-dimensional space into eight subspaces. For the static information in each keyframe, it is assigned to the corresponding octree node or unit according to its three-dimensional spatial position. For a three-dimensional point (x, y, z), it will be determined according to its coordinates to which octree unit it belongs, and it will be stored in the corresponding unit. By continuously dividing the space, a hierarchical structure is formed. This structure can efficiently store and manage three-dimensional information, and can flexibly display maps of different precisions according to the needs of different levels of details. Finally, the information corresponding to the optimized scene representation and the camera posture is integrated with the corresponding nodes or units in the octree map according to the structural rules of the octree. For the optimized scene representation, the information of the geometric structure (such as point, line, and surface information) will be added to the corresponding nodes or units of the octree, and the camera posture information can be stored in the metadata of the octree or associated with the nodes of the octree through additional data structures. In this way, the scene representation and camera attitude information are integrated into the octree map, and finally the three-dimensional map construction without interference from dynamic objects is completed, ensuring that the entire three-dimensional map contains both accurate scene information and accurate camera perspective information, providing complete, accurate and dynamic interference-free map information for subsequent applications.
[0128] In order to facilitate understanding of the above embodiment, the following specific description is given here:
[0129] This embodiment proposes a 3D map reconstruction method that combines theoretical research and synthetic dataset verification for occluded background repair technology in dynamic scenes. Synthetic datasets are used to simulate a variety of typical dynamic scenes, and the performance of the proposed method is systematically verified and compared. The accuracy of dynamic object tracking and background repair is evaluated, and the method is continuously adjusted and improved to ensure high-precision 3D map reconstruction in complex dynamic scenes. The technical route of this embodiment is shown in Figure 2 .
[0130] In order to better eliminate the influence of dynamic obstacles and completely reconstruct the dynamic background in the dynamic scene, this embodiment takes advantage of deep learning in target detection, and uses convolutional neural network (CNN) to obtain prior information and depth information of static scenes to eliminate dynamic targets, thereby restoring the static background blocked by dynamic obstacles. In order to detect dynamic objects, the target detection algorithm is used in the experiment, and the convolutional neural network is used to perform semantic segmentation on the input image. The mask obtained from the semantic segmentation network can effectively mark the dynamic objects in the image, and guide the dynamic SLAM system as a stable and reliable prior constraint.
[0131] Taking the original RGB image as input, a dedicated dynamic processing process is used to remove dynamic objects, and the binary mask of potential dynamic or movable objects in the image is output. In each mapping iteration, keyframes are selected to optimize the scene representation and camera pose and dynamic objects are removed. For the removed dynamic targets, the occluded background is repaired using static information obtained from previous viewpoints to synthesize a realistic image without dynamic targets. The repaired image contains more scene information, making the map presentation more accurate and enhancing the stability of camera tracking. See the technical concept architecture for details. Figure 3 .
[0132] In view of the complexity and diversity of dynamic scenes, in order to effectively solve the core problem of building three-dimensional maps in dynamic scenes, this embodiment mainly focuses on the detection and segmentation of dynamic targets and the reconstruction of static backgrounds. In the process of building maps in dynamic environments, we usually face the following problems. Finally, we build a system that can efficiently and accurately generate three-dimensional maps in dynamic environments, providing a solid technical foundation for intelligent systems to perceive the environment. The main research of each part is as follows:
[0133] 1) Detection and segmentation of dynamic obstacles: In dynamic environments, traditional SLAM systems often fail to distinguish between dynamic and static feature points, resulting in mismatches and reduced positioning accuracy of the map. The motion interference of dynamic obstacles (such as pedestrians, vehicles, etc.) is the main cause of these problems. This embodiment introduces deep learning models (such as SegNet) and motion consistency detection technology to accurately identify dynamic targets, segment them, and remove dynamic feature points. Combined with semantic segmentation technology, the system can distinguish between static background and dynamic obstacles in real time, so that the feature points of dynamic objects are removed during the map construction process, effectively reducing the posture drift problem caused by interference from dynamic obstacles, and enhancing the robustness and adaptability of the system in complex dynamic environments.
[0134] 2) Static background reconstruction: The occlusion of dynamic obstacles will cause the loss of information in the background area, which will affect the integrity and coherence of the three-dimensional map. In order to restore the occluded static background information, this embodiment designs a background repair strategy that combines multi-perspective information and prior frame data. When processing scenes occluded by dynamic objects, the system makes full use of the information fusion in the multi-perspective images and the texture and depth data in the historical frames to complete and restore the background area with high precision. By performing state tracking and depth reasoning on the occluded area of the dynamic target, the accurate restoration of the background texture and depth features of the occluded area is ensured, thereby significantly improving the detail expression and overall accuracy of the three-dimensional map, and providing strong technical support for the construction of high-quality maps in complex dynamic scenes, thereby restoring a static three-dimensional map that can significantly reduce the interference of dynamic obstacles.
[0135] The main framework of this embodiment combines the object detection and instance segmentation functions of Mask R-CNN to achieve high-precision 3D map construction in dynamic environments. Through the deep neural network structure of Mask R-CNN, the system is first able to detect and segment dynamic objects in the scene. Its region proposal network generates candidate regions, while the subsequent branch structure is responsible for performing bounding box regression, category classification, and pixel-level segmentation, thereby accurately identifying dynamic objects in complex scenes and generating detailed segmentation masks to ensure that the edges of dynamic objects are clear and cleanly separated from the background.
[0136] After obtaining the segmentation mask of the dynamic object, the system further screens and removes the dynamic features in the scene by combining the multi-view geometric features, and filters out the objects that do not conform to the static background features, thereby significantly improving the stability of the 3D reconstruction process. In this process, the mask and feature attributes of Mask R-CNN greatly improve the accuracy of scene parsing, ensuring that interfering features will not affect map construction.
[0137] After the dynamic objects are removed, the system needs to repair the occluded background area to ensure the integrity of the 3D map. With the help of the depth camera data and multi-view information, the system uses semantic segmentation and static information in historical frames to complete the background. Through the feature synthesis technology of deep learning, the true restoration of the background after the dynamic objects are removed can be ensured, providing a stable and accurate foundation for subsequent 3D reconstruction. The framework diagram of the main research can be seen in Figure 4 .
[0138] Step 1: Dynamic content recognition and segmentation based on Mask R-CNN and multi-view geometry
[0139] In the case of RGB-D, multi-view geometry is used to improve dynamic content segmentation from two aspects. To this end, it is necessary to know the pose of the camera, for which a low-cost tracking module has been implemented to locate the camera in the scene map that has been created. The general dynamic and static obstacle detection and segmentation methods mainly have the following steps: data acquisition and preprocessing, dynamic obstacle detection, and dynamic object recognition.
[0140] Use sensors such as RGB cameras, depth cameras or lidar to collect environmental data and generate a series of continuous image frames I t Or RGB-D images with depth information. These data will provide a basis for dynamic obstacle detection and segmentation. In order to reduce the impact of image noise and illumination changes on subsequent detection, each frame of the image needs to be preprocessed. These operations usually include image grayscale, scale adjustment, denoising and smoothing to reduce the interference of noise and illumination changes on detection. The preprocessed image can more clearly reflect the shape and details of the target object. The preprocessed image can be expressed as:
[0141] I t '=Preprocess(I t )(1)
[0142] At the same time, when performing obstacle detection, a deep learning-based target detection network such as YOLO, FasterR-CNN or SSD is used to perform preliminary recognition of dynamic objects. Different models have their own advantages: YOLO is faster and suitable for scenes with high frame rate requirements, while FasterR-CNN has higher accuracy and is suitable for complex environments. det , the bounding box B of the dynamic object can be extracted from the image t and category information C t Among them, B t represents the bounding box of the detected dynamic object, including the position coordinates (x, y, z, h), and C t Indicates the category information of dynamic objects, such as vehicles or pedestrians. That is:
[0143] B t , C t =f det (I′ t ) (2)
[0144] For detecting dynamic objects, Mask R-CNN is used to obtain pixel-wise semantic segmentation of the image. This is the state-of-the-art for object instance segmentation. Mask R-CNN can obtain both pixel-wise semantic segmentation and instance labels. But instance labels may be used in future work to track different moving objects. The input to Mask R-CNN is a raw RGB image. The idea is to segment those classes that are likely to be dynamic or movable (people, bicycles, cars, motorcycles, airplanes, buses, trains, trucks, boats, birds, cats, dogs, horses, sheep, cows, elephants, bears, zebras, and giraffes). It is believed that for most environments, the dynamic objects that may appear are included in this list. If additional classes are needed, the network trained on MSCOCO can be fine-tuned with new training data.
[0145] After obtaining dynamic obstacles through target detection, semantic dynamic segmentation is further used to segment dynamic objects at the pixel level. Nowadays, segmentation methods include DeepLab, U-Net or MaskR-CNN network models. Among them, Mask R-CNN can generate accurate object masks after target detection. seg Generate a mask M from the dynamic region extracted from the image t ,Right now:
[0146] M t =f seg (It ,B t ) (3)
[0147] This mask M t The dynamic object areas in the image are marked to ensure that these areas are not mistakenly included in the map background. After the mask is generated, it is used to separate the dynamic and static areas in the image. The mask M of the dynamic area t The information is passed to the background repair module for subsequent occluded background repair, while the static area is directly used for map construction. At the same time, by accumulating the static area information of the previous frame (i.e. building a static prior information library), static and dynamic obstacles can be better distinguished during the detection process.
[0148] The generated mask is processed to improve the segmentation accuracy. For example, the mask is refined through morphological operations (such as corrosion and dilation) to remove boundary noise and small areas of misjudgment, thereby improving the segmentation effect of dynamic obstacles. According to the movement of the target and the update of the mask, the Kalman filter or optical flow method can be used to update the position of the bounding box of the dynamic object to ensure accurate tracking of the position of the dynamic obstacle in continuous frames. The state update and measurement update formulas of the Kalman filter are:
[0149]
[0150]
[0151] where x t is the state vector (including position and velocity), A is the state transfer matrix, B is the control input matrix, H is the measurement matrix, and w t and v t denote process noise and measurement noise, respectively. By using Mask R-CNN, most dynamic objects can be segmented without tracking and mapping. However, some objects cannot be detected by this method because they are not dynamic a priori but movable. Figure 5 Schematic diagram of the principle of multi-view geometry detection of dynamic points:
[0152] When detecting motion between keyframes (KF), it is necessary to select several KFs with the highest overlap with the current frame (CF) from the KF database according to the current frame (CF). The upper limit of the KF database is generally set to 20. The larger the database, the more difficult it is to initialize the system and affect the speed of frame search; the number of overlapping KFs selected will also affect the running speed of the system and the accuracy of dynamic object detection. Calculate the projection of each key point x from the previous key frame to the current frame, and obtain the key frame x' and their projection depth z1. For each key point, its corresponding 3D point is X, and calculate the angle between the back projection of x and x', that is, the disparity angle a. If this angle is greater than 30°, the point may be occluded and will be ignored from then on. In the TUM dataset, static objects with a disparity greater than 30° are considered dynamic because of their viewpoint differences. Taking into account the reprojection error, obtain the depth of the remaining key points in the current frame z' (directly from the depth measurement) and compare them with z1. If the difference Δz=z1-z' exceeds the threshold Γz, the key point x' is considered to belong to a dynamic object.
[0153] Step 2: Introduction to Mask R-CNN segmentation algorithm
[0154] Mask R-CNN is a deep learning model that integrates object detection, instance segmentation, and bounding box localization, and is suitable for handling object recognition and segmentation tasks in complex dynamic environments. It extends Faster R-CNN and adds a branch on each region of interest (RoI) to predict the segmentation mask (segmentationMask) based on the existing classification and bounding box regression branches. The mask branch is a small FCN applied to each RoI to predict the segmentation mask in a pixel-topixel manner. Given the Faster R-CNN framework, Mask R-CNN is easy to implement and train, which facilitates a wide range of flexible architecture designs. In addition, the mask branch only adds a small computational overhead, enabling fast systems and fast experiments. Its structure is based on a two-stage framework design, first locating objects and then performing pixel-level segmentation. In the first stage, Mask R-CNN generates candidate regions through a region proposal network (RPN) to mark potential object locations. In the second stage, the model further extracts features from these regions and performs bounding box regression, category classification, and pixel-level segmentation respectively through a branching structure. Compared with other detection models, Mask R-CNN has more sophisticated pixel-level segmentation capabilities and is suitable for accurately distinguishing between background and dynamic objects. Figure 6 shown.
[0155] It plays an important role in the process of dynamic object removal and background restoration. First, the segmentation branch of Mask R-CNN can accurately identify the contours of dynamic objects, thereby generating pixel-level masks to separate these dynamic objects from the background. Secondly, this fine segmentation mask provides a basis for subsequent background restoration, so that after removing dynamic objects, the system can accurately restore the occluded background information. Specifically, the dynamic background reconstruction in the study uses the segmentation results of Mask R-CNN for feature extraction and screening, and combines semantic segmentation with multi-view geometric information to ensure the integrity and authenticity of the background after dynamic objects are removed. The dynamic object segmentation effect diagram is shown below. Figure 7 As shown, the red area is the segmentation effect using the Mask R-CNN network.
[0156] Step 3: Dynamic obstacle removal based on semantic segmentation method
[0157] The above method can better segment moving objects. This embodiment mainly uses a semantic segmentation network to detect and segment dynamic obstacles. The semantic segmentation network can better eliminate a priori dynamic objects, but the effect of distinguishing non-a priori dynamic objects in the scene is limited. A multi-view geometry algorithm is used for further processing. The proposed dynamic object elimination algorithm process is as follows: Figure 8 As shown. Dynamic points are detected according to the motion relationship between key frames KF. After the dynamic points are detected, they are divided into dynamic points with semantic information and dynamic points without semantic information by judging whether they have semantic labels. Semantic contour search is performed on the semantic map for feature points with semantic information, and regional growth is performed on the depth map for dynamic points without semantic information. Semantic information is fully utilized to reduce the number of regional growth seed points and improve the operating efficiency of the system. The dynamic object mask without semantics and the dynamic object mask with semantic information are fused to obtain a complete dynamic object mask.
[0158] The RGB-D camera is used to obtain RGB images and depth images. The RGB images are processed by the semantic segmentation network to obtain pixel-level semantic information. The semantic information is used to remove the feature points of the prior dynamic objects in the image. The multi-view geometry algorithm is used to further detect the feature points corresponding to the non-prior dynamic objects. The detection results of the multi-view geometry and the semantic segmentation network are cross-validated to obtain the complete dynamic area. After filtering out the feature points in the dynamic area, the tracking thread is entered to obtain a more accurate pose. The dynamic feature removal effect is as follows: Fig. 9 As shown, the red box marks the detected dynamic features.
[0159] Step 4: Static background repair 3D map construction based on octree map dynamic environment
[0160] Octree is a data structure with eight child nodes, and its name comes from its unique way of dividing space. The reason for dividing the space into eight sub-areas can be figuratively understood as cutting a cube once on three orthogonal planes, thereby dividing the entire cube into eight small cubes of equal volume. It is widely used in 3D mapping, environmental modeling, and spatial retrieval. Octree-based 3D map construction and static background restoration is an innovative method that combines space segmentation and dynamic scene processing, which can achieve efficient map construction in complex dynamic environments. Fig.10 A diagram of the main principles for building a graph for an octree.
[0161] The octree recursively divides the three-dimensional space into small cubic nodes to form a hierarchical structure. This method can not only dynamically adjust the resolution of the map to adapt to the complexity of the scene, but also reduce storage requirements and support efficient processing of sparse scenes. At the same time, it is convenient to quickly retrieve and locate specific spatial areas to meet the needs of real-time map construction. In dynamic scenes, dynamic objects (such as pedestrians, vehicles, etc.) interfere with the accuracy and integrity of map construction. Through target detection algorithms (such as Mask R-CNN or YOLO), dynamic objects can be identified and semantic segmentation masks can be generated to exclude dynamic areas from map construction. Combined with multi-view geometric verification technology, feature points can be further screened to ensure that only static parts are retained, laying the foundation for static background repair. The specific mapping effect of the octree is as follows: Fig.11 shown.
[0162] Static background restoration effectively reconstructs the area occluded by dynamic objects by fusing historical keyframes and multi-view information. Using the multi-frame data stored in the octree, the occluded area can be restored to a static background that conforms to the scene logic. For the missing depth data, interpolation or depth compensation methods are used to fill it, and combined with the classification results of the semantic segmentation network, realistic restoration images are generated based on semantic features. This restoration method enhances the accuracy of the map and the stability of camera tracking, ensuring the applicability of the map in dynamic environments.
[0163] In specific implementation, the RGB and depth channels of the previous key frame are projected to the dynamic area of the current frame based on the known position information of the previous and next frames. In this way, the static background information can be mapped to the current frame to fill the blank area left after the dynamic objects are removed. In this process, some areas are left blank due to lack of correspondence, while other areas may not be drawn because the required part of the scene has not yet appeared in the key frame, or has appeared but lacks effective depth information. During the background repair process, the system will use the depth camera to capture dynamic frame data from a global perspective for each dynamic object that has been removed. In order to further optimize the coherence of the repair, multi-perspective cross-frame information transmission is used to obtain the motion trend of the scene. By combining Mask R-CNN, multi-perspective information fusion and historical information in the prior frame, accurate restoration of the dynamically occluded area is achieved, and a high-precision three-dimensional map without dynamic objects is generated, ensuring stability and accuracy in dynamic environments. This process removes moving content from the generated map, presenting a natural, realistic image effect without dynamic interference. The static background repair process is as follows Fig.12 shown.
[0164] This embodiment can not only enhance the visual quality of the synthesized image, but also has important significance for application scenarios such as virtual reality and augmented reality, as well as tasks such as relocalization and camera tracking after map creation. In specific implementation, the RGB and depth channels of the previous key frame are projected to the dynamic area of the current frame based on the known position information of the previous and next frames. In this way, the static background information can be mapped to the current frame, thereby filling the blank area left after the dynamic objects are removed. However, in this process, some areas are left blank due to lack of correspondence, while other areas may not be drawn because the required part of the scene has not yet appeared in the key frame, or has appeared but lacks valid depth information.
[0165] In order to deal with these challenges, a method combining deep learning and motion consistency is used to better restore the content of blank areas. This method not only improves the integrity and realism of the synthetic image, but also provides more stable static background information for the subsequent dynamic SLAM system. It can clearly show the successful segmentation and deletion of dynamic content, and how most of the segmented areas are correctly integrated into the static background information. Through this background repair process, high-quality images without dynamic interference can be generated, providing more reliable visual data support for various applications. The mask M of the dynamic object is tIt is used to remove dynamic interference and only retain data in static areas to build maps. In the process of map construction, the selection of key frames is crucial. By selecting appropriate key frames, the scene representation and camera posture can be optimized, and the accuracy and consistency of the map can be improved. At the same time, combined with depth information and image data, point cloud fusion technology is used to merge static information in different frames into a complete three-dimensional map. The mapping effect of this embodiment is shown as follows Fig.13 shown.
[0166] The solution of this embodiment has the possibility of practical operation, which can not only promote the development of intelligent robot technology, but also demonstrate its value and potential in practical applications. By optimizing the dynamic target detection algorithm, it is planned to combine the powerful recognition ability of deep learning and the stability of traditional image processing technology to achieve accurate recognition and tracking of dynamic targets. The innovation of this algorithm lies in its ability to effectively separate dynamic elements from images through multiple steps, such as motion detection, depth analysis, classification and labeling, and tracking verification, thereby reducing errors in the process of building three-dimensional maps.
[0167] See also Fig.14 , is a schematic diagram of a structure of a three-dimensional map reconstruction device without dynamic object interference provided by an embodiment of the present invention. The three-dimensional map reconstruction device without dynamic object interference includes:
[0168] The image acquisition module 10 is used to acquire environmental data through a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information;
[0169] An image separation module 11 is used to perform preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network, generate a dynamic region mask, and separate the dynamic region from the static background; the dynamic region mask is used to mark the dynamic object region in the target image;
[0170] A processing module 12, configured to perform segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy;
[0171] The recognition module 13 is used to combine the processed dynamic area mask and use the target detection algorithm to recognize the dynamic objects in the target image again and generate a semantic segmentation mask;
[0172] A screening module 14 is used to remove dynamic areas based on the semantic segmentation mask, and to screen feature points of the target image that do not conform to the dynamic background in combination with a multi-view geometric verification technology, and retain static parts;
[0173] A restoration module 15 is used to perform background restoration on the target image's occluded area by the dynamic object by using the multi-view views and the historical frames of the target image, through multi-view information fusion and depth data interpolation, to restore the static background information;
[0174] The reconstruction module 16 is used to select the corresponding key frames of the restored target image to optimize the scene representation and camera posture, and use the static information of the key frames to build an octree map, integrate the optimized scene representation and camera posture into the octree map, and complete the three-dimensional map construction without dynamic object interference.
[0175] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0176] First, the depth camera is used to collect environmental data to obtain the target image, including continuous image frames or RGB-D images with depth information. Based on this, the Mask R-CNN network is used to perform preliminary detection and pixel-by-pixel semantic segmentation of dynamic objects, generate dynamic area masks, and achieve preliminary separation of dynamic areas from static backgrounds. This is because the powerful feature extraction and classification capabilities of the Mask R-CNN network can effectively distinguish different objects and provide important dynamic area information for subsequent processing. Then, the dynamic area mask is processed for segmentation accuracy. Methods such as morphological operations can be used to refine the boundaries of the dynamic area, improve segmentation accuracy, and ensure the accuracy of subsequent processing. Subsequently, combined with the processed dynamic area mask, the target detection algorithm is used to identify dynamic objects again and generate semantic segmentation masks to further refine the marking of dynamic objects. Thanks to the combined processed mask information, it can focus on possible dynamic areas and improve recognition accuracy. Based on the semantic segmentation mask, the dynamic area is eliminated, and the multi-view geometric verification technology is used to filter out feature points that do not conform to the dynamic background. The multi-view geometric verification technology uses the geometric relationship between multi-view images to ensure the accuracy of screening, thereby retaining the static part. Then, the multi-view view and historical frame are used to repair the area blocked by dynamic objects through multi-view information fusion and deep data interpolation. This is to utilize the complementarity of multi-view information and the advantages of deep data interpolation to restore more complete and accurate static background information. Finally, key frames are selected from the repaired target image, the scene representation and camera posture are optimized, and the octree map is constructed using the static information of the key frames. The optimized information is integrated into the octree. The octree structure is conducive to the hierarchical storage and representation of three-dimensional space, and the construction of a three-dimensional map without interference from dynamic objects is completed. The optimization of key frames and the use of octrees in the construction process can effectively integrate static information and improve the quality of map construction. From the above analysis, it can be seen that the embodiment of the present invention effectively solves the problems of mismatching, posture drift, and background information loss caused by dynamic object interference in dynamic scenes through technical means such as Mask R-CNN network, segmentation accuracy processing, multi-view geometry verification technology, multi-view information fusion, deep data interpolation, and octree map construction, and significantly improves the accuracy of three-dimensional map reconstruction.
[0177] It is understandable that the embodiment of the three-dimensional map reconstruction device without dynamic object interference may correspond to the relevant contents of the above-mentioned embodiment of the three-dimensional map reconstruction method without dynamic object interference, and will not be described in detail here.
[0178] See also Fig.15, is a schematic diagram of a three-dimensional map reconstruction device without dynamic object interference provided by an embodiment of the present invention. The three-dimensional map reconstruction device without dynamic object interference of this embodiment includes: a processor 100, a memory 101, and a computer program stored in the memory 101 and executable on the processor 100, such as a three-dimensional map reconstruction program without dynamic object interference. When the processor 100 executes the computer program, the steps in the above-mentioned three-dimensional map reconstruction method without dynamic object interference are implemented. Alternatively, when the processor 100 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0179] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the three-dimensional map reconstruction device without dynamic object interference.
[0180] The three-dimensional map reconstruction device without dynamic object interference can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The three-dimensional map reconstruction device without dynamic object interference may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the schematic diagram is only an example of a three-dimensional map reconstruction device without dynamic object interference, and does not constitute a limitation on the three-dimensional map reconstruction device without dynamic object interference. It can include more or less components than shown in the figure, or combine certain components, or different components. For example, the three-dimensional map reconstruction device without dynamic object interference can also include input and output devices, network access devices, buses, etc.
[0181] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the three-dimensional map reconstruction device without dynamic object interference, and uses various interfaces and lines to connect various parts of the three-dimensional map reconstruction device without dynamic object interference.
[0182] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the three-dimensional map reconstruction device without dynamic object interference by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0183] Wherein, if the module / unit integrated in the three-dimensional map reconstruction device without dynamic object interference is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media does not include electrical carrier signals and telecommunication signals.
[0184] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.
[0185] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A three-dimensional map reconstruction method without dynamic object interference, characterized in that: include: The environment data is collected by a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information; Perform preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through the Mask R-CNN network to generate a dynamic region mask to separate the dynamic region from the static background; The dynamic area mask is used to mark the dynamic object area in the target image; Performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy; Combining the processed dynamic area mask, using a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask; Based on the semantic segmentation mask, dynamic areas are eliminated, and the feature points of the target image that do not conform to the dynamic background are screened out in combination with the multi-view geometric verification technology, and the static parts are retained; Using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation, background repair is performed on the target image's occluded area by the dynamic object to restore static background information; The corresponding key frames of the repaired target image are selected to optimize the scene representation and the camera posture, and the static information of the key frames is used to construct an octree map, and the optimized scene representation and camera posture are integrated into the octree map to complete the construction of a three-dimensional map without interference from dynamic objects.
2. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: After the environmental data is collected by the depth camera to obtain the target image, the method also includes: The target image is preprocessed to obtain a preprocessed target image; the preprocessing includes image graying, scale adjustment, denoising and smoothing operations.
3. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The Mask R-CNN network is used to perform preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image to generate a dynamic region mask to separate the dynamic region from the static background, including: Input the target image into the Mask R-CNN network, and generate candidate regions through the region proposal network; the candidate regions are regions in the image that may contain dynamic objects; For each candidate region, bounding box regression, category classification and pixel-level segmentation are performed respectively through the branch structure of the Mask R-CNN network to generate a bounding box, category information and a dynamic region mask of the dynamic object; the bounding box includes the position coordinates of the dynamic object, the category information includes the category label of the dynamic object, and the dynamic region mask is used to mark the dynamic object region in the image; According to the generated dynamic area mask, the dynamic area is separated from the static background; the dynamic object area marked by the dynamic area mask will be excluded from the static background.
4. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The performing segmentation accuracy processing on the generated dynamic area mask to improve the segmentation accuracy includes: Obtaining a dynamic region mask generated by performing preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network; Performing morphological operations on the generated dynamic region mask, including erosion and dilation, to remove noise at the mask boundary and small misjudged regions; the morphological operations are used to improve the segmentation accuracy of the dynamic region mask; The position of the bounding box of the dynamic object is updated by Kalman filtering or optical flow method to ensure accurate tracking of the position of the dynamic obstacle in consecutive frames; the position update is based on the bounding box information of the dynamic object to ensure the consistency of the position of the dynamic object in consecutive frames; The processed dynamic area mask is used for subsequent dynamic object removal and background restoration; the dynamic object area marked by the processed dynamic area mask will be excluded from the three-dimensional map construction process and used to guide the background restoration of the area blocked by the dynamic object.
5. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The combined processed dynamic area mask uses a target detection algorithm to identify dynamic objects in the target image again and generate a semantic segmentation mask, including: The processed dynamic area mask is used as input, and the target image is processed using a target detection algorithm to generate a bounding box and category information of the dynamic object; the bounding box includes the position coordinates of the dynamic object in the target image, and the category information includes the category label of the dynamic object; Taking the processed dynamic area mask and the generated bounding box and category information of the dynamic object as input, the Mask R-CNN network performs pixel-by-pixel semantic segmentation operation on the target image to generate a semantic segmentation mask of the dynamic object; the semantic segmentation mask is used to mark the dynamic object area in the image.
6. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The method of eliminating the dynamic area based on the semantic segmentation mask and screening the feature points of the target image that do not conform to the dynamic background in combination with the multi-view geometric verification technology to retain the static part includes: Obtaining a generated semantic segmentation mask; the semantic segmentation mask is used to mark a dynamic object area in a target image; According to the semantic segmentation mask, a dynamic area in the target image is eliminated; the dynamic area is a dynamic object area marked by the semantic segmentation mask; Combined with multi-view geometric verification technology, feature points in the target image that do not conform to the dynamic background are screened out; the multi-view geometric verification technology calculates the projection error between the key frame and the current frame to screen out points that do not conform to the dynamic background characteristics; The filtered static feature points are retained; the static feature points are used for subsequent three-dimensional map construction.
7. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The method utilizes the multi-view views and the historical frames of the target image to perform background repair on the region of the target image blocked by the dynamic object through multi-view information fusion and depth data interpolation to restore the static background information, including: Acquire a historical frame of the target image; the historical frame is an image frame captured before the target image, and contains static background information of the area blocked by the dynamic object; Integrate the information of different perspectives in the multi-perspective view, perform fusion operation on the static part information of the target image under different perspectives, and obtain the static part information after multi-perspective fusion; For the area blocked by the dynamic object in the target image, using a depth data interpolation method, taking the known depth information of the unblocked area as a reference, interpolating the depth data of the blocked area to obtain the depth information of the area; According to the static part information after multi-view fusion and the depth information obtained by interpolating the depth data, the background repair operation is performed on the area blocked by the dynamic object, and the repaired information is used as the final static background information.
8. The method for reconstructing a three-dimensional map without dynamic object interference as claimed in claim 1, characterized in that: The method comprises: selecting the corresponding key frame of the restored target image to optimize the scene representation and the camera posture, and constructing an octree map using the static information of the key frame, integrating the optimized scene representation and the camera posture into the octree map, and completing the construction of a three-dimensional map without interference from dynamic objects, including: Acquire a restored target image, wherein the restored target image includes static background information of a region blocked by a dynamic object, and the restored target image is obtained by performing background restoration on the region blocked by the dynamic object of the target image by using multi-view views and historical frames of the target image, through multi-view information fusion and depth data interpolation; Selecting a corresponding key frame of the restored target image, wherein the key frame is a representative image frame in the restored target image and is screened from the restored target image by calculating inter-frame overlap and scene complexity; According to the static feature points and depth information in the key frames, the geometric structure in the scene is optimized using multi-view geometric constraints to obtain an optimized scene representation; Through multi-view geometric constraints and the projection relationship of feature points in different key frames, the rotation and translation parameters of the camera are optimized to obtain the optimized camera posture; Convert the static information in different keyframes into nodes or units of the octree, present the three-dimensional space in a hierarchical manner, and construct an octree map; The optimized scene representation and the information corresponding to the camera posture are integrated with the corresponding nodes or units in the octree map according to the structural rules of the octree to complete the construction of a three-dimensional map without interference from dynamic objects.
9. A three-dimensional map reconstruction device without dynamic object interference, characterized in that: include: An image acquisition module is used to acquire environmental data through a depth camera to obtain a target image; the target image includes a continuous image frame or an RGB-D image with depth information; An image separation module is used to perform preliminary dynamic object detection and pixel-by-pixel semantic segmentation on the target image through a Mask R-CNN network, generate a dynamic region mask, and separate the dynamic region from the static background; The dynamic area mask is used to mark the dynamic object area in the target image; A processing module, used for performing segmentation accuracy processing on the generated dynamic area mask to improve segmentation accuracy; A recognition module, used to combine the processed dynamic area mask, use the target detection algorithm to recognize the dynamic objects in the target image again and generate a semantic segmentation mask; A screening module, used to remove dynamic areas based on the semantic segmentation mask, and screen feature points of the target image that do not conform to the dynamic background in combination with multi-view geometric verification technology, and retain static parts; A restoration module, used to perform background restoration on the target image's occluded area by the dynamic object by using the multi-view views and the target image's historical frames, through multi-view information fusion and depth data interpolation, to restore static background information; The reconstruction module is used to select the corresponding key frames of the restored target image to optimize the scene representation and camera posture, and use the static information of the key frames to build an octree map, integrate the optimized scene representation and camera posture into the octree map, and complete the three-dimensional map construction without dynamic object interference.
10. A three-dimensional map reconstruction device without dynamic object interference, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for reconstructing a three-dimensional map without interference from dynamic objects as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Visual SLAM method based on semantic segmentation of deep learning
CN112132897A
Monocular vision-based positioning and map construction method
CN113298904A
Unmanned aerial vehicle low-altitude monitoring method
CN118504925A
Cited By
Super-resolution high-precision map construction method and device, storage medium and program product
CN120521628A
Photovoltaic module defect detection method and system based on photoluminescence technology
CN120823204A
Slam reality three-dimensional space scanning imaging system based on quadruped robot
CN120991829A
Dynamic shelter restoration method and system based on continuous streetscape panoramic image
CN122023201A
Dynamic target three-dimensional positioning and tracking method applied to earth surface fitting and dynamic separation of unmanned aerial vehicle
CN122023532A