Indoor dynamic environment-oriented high-precision visual localization and high-fidelity object-level map construction method and device and program product

By integrating semantic segmentation and depth clustering in VSLAM technology, visual sensor data in indoor dynamic environments is processed, and the problem of insufficient positioning accuracy and map reconstruction quality in the existing technology is solved, and high-precision visual positioning and high-fidelity map construction are realized.

CN119991802AActive Publication Date: 2025-05-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510058059.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing VSLAM technology lacks positioning accuracy and map reconstruction quality in dynamic environments, and it is particularly difficult to deal with passive dynamic objects and eliminate artifacts.

Method used

By collecting color images and depth images, using semantic segmentation large models for segmentation and feature extraction, combining depth clustering and feature encoding, active dynamic object masks and semantic images are obtained, and robust non-active dynamic points are obtained. Based on this information, we optimize the pose, calculate dynamic semantic features, generate the final dynamic object mask and static points, and use the differentiable rendering characteristics of the Gaussian ellipsoid for backpropagation to achieve high-precision visual positioning and high-fidelity map construction.

Benefits of technology

In indoor dynamic environments, high-precision visual positioning and high-fidelity object-level map construction are realized, and dynamic objects are accurately deleted, improving positioning accuracy and map reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991802A_ABST
    Figure CN119991802A_ABST
Patent Text Reader

Abstract

The invention provides an indoor dynamic environment-oriented high-precision visual localization and high-fidelity object-level map construction method and device and a program product, and the method comprises the steps: carrying out the segmentation and feature extraction of an image based on a semantic segmentation large model, generating an initiative dynamic object extension mask, obtaining robust non-initiative dynamic points, and carrying out the feature extraction of the robust non-initiative dynamic points; the method comprises the following steps: preliminarily optimizing a pose in combination with a Gaussian map, calculating a center distance between a robust non-active dynamic point corresponding to an object and a Gaussian ellipsoid, obtaining a final dynamic object mask in combination with a semantic loss variable quantity of a corresponding pixel region in a key frame sequence, and optimizing the loss of a back propagation static pixel region to obtain a final dynamic object mask. And densifying the Gaussian ellipsoid in the multi-view intersection area to obtain a high-precision visual positioning result and a high-fidelity object-level map. According to the method and the device, the dynamic object can be accurately and comprehensively deleted in any indoor dynamic environment, high-precision visual positioning is realized, and a high-fidelity object-level map is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision and artificial intelligence technology, and in particular to a method, device and program product for high-precision visual positioning and high-fidelity object-level map construction for indoor dynamic environments. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.

[0003] Visual Simultaneous Localization and Mapping (VSLAM) is a task that uses visual sensor data to reconstruct a 3D map in an unknown environment and estimate the camera pose at the same time. Traditional VSLAM methods have significant limitations in dynamic environments, especially when dealing with the impact of dynamic objects on camera pose estimation. In addition, existing methods are difficult to reconstruct high-fidelity object-level maps, which to a certain extent restricts the practical application of VSLAM technology. With the continuous development of 3D reconstruction technology, the reconstruction of high-fidelity maps has gradually become a research hotspot. 3D Gaussian Splatting for Real-Time Radiance Field Rendering (3DGS) technology reconstructs high-fidelity scene maps through input images and has powerful real-time rendering capabilities. Its efficient rendering performance enables 3DGS to be seamlessly integrated with the mapping backend of VSLAM, providing innovative solutions for cutting-edge fields such as autonomous driving and virtual reality.

[0004] At present, the VSLAM method based on 3DGS has become a research hotspot, but the robustness of most current systems in dynamic environments is still insufficient. In order to solve this problem, some researchers have tried to combine multiple methods to deal with dynamic objects in the scene to improve positioning accuracy and system robustness. However, these methods have not fully considered the edge noise between color images and depth images, failed to effectively deal with passive dynamic objects in the scene, and failed to effectively eliminate artifacts generated by dynamic objects. There is still a lot of room for improvement in positioning accuracy and map reconstruction quality, and further optimization and improvement are urgently needed. Summary of the invention

[0005] In view of this, the purpose of the present disclosure is to propose a method, device and program product for high-precision visual positioning and high-fidelity object-level map construction for indoor dynamic environments, which can at least solve some technical problems in the prior art to a certain extent.

[0006] Based on the above purpose, the first aspect of the exemplary embodiment of the present disclosure provides a method for high-precision visual positioning and high-fidelity object-level map construction for indoor dynamic environments, the method comprising:

[0007] Collect visual sensor data, including color images and depth images, segment and extract features from the color images based on a semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust inactive dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images;

[0008] A preliminary optimized pose is obtained based on the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map, and a distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is calculated based on the preliminary optimized pose to obtain a latent dynamic semantic feature, and a passive dynamic semantic feature is obtained based on the semantic loss change of the pixel area corresponding to the latent dynamic semantic feature in the key frame sequence, and a final dynamic object mask is obtained based on the passive dynamic semantic feature and the active dynamic object expansion mask, and a static point is obtained based on the passive dynamic semantic feature and the robust non-active dynamic point;

[0009] Based on the final dynamic object mask, the loss of the input image and the rendered image in the static pixel area is obtained, and the loss is back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result;

[0010] Inserting the static point into the Gaussian map based on the high-precision visual positioning result, optimizing the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss, and obtaining a high-fidelity map;

[0011] A high-loss object is obtained based on the loss, and a ray intersection area of ​​the high-loss object in a key frame sequence is calculated to obtain a multi-perspective ray intersection area. A Gaussian ellipsoid is densified based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

[0012] Based on the same inventive concept, the second aspect of the exemplary embodiment of the present disclosure provides a high-precision visual positioning and high-fidelity object-level map construction device for indoor dynamic environments, including:

[0013] A visual sensor data preprocessing module for indoor dynamic environments, configured to collect visual sensor data, including color images and depth images, segment and extract features from the color images based on a semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust inactive dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images;

[0014] A dynamic object detection and deletion module is configured to obtain a preliminary optimized pose based on robust non-active dynamic points and a Gaussian ellipsoid in a Gaussian map, calculate the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtain a latent dynamic semantic feature, obtain a passive dynamic semantic feature based on a semantic loss change of a pixel area corresponding to the latent dynamic semantic feature in a key frame sequence, obtain a final dynamic object mask based on the passive dynamic semantic feature and an active dynamic object expansion mask, and obtain a static point based on the passive dynamic semantic feature and the robust non-active dynamic point;

[0015] A high-precision visual positioning module is configured to obtain the loss of the input image and the rendered image in the static pixel area based on the final dynamic object mask, and to back-propagate the loss using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result;

[0016] A high-fidelity map optimization module, configured to insert the static point into a Gaussian map based on the high-precision visual positioning result, and optimize the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss to obtain a high-fidelity map;

[0017] A high-fidelity object-level map densification module is configured to obtain a high-loss object based on the loss, calculate the ray intersection area of ​​the high-loss object in a key frame sequence, obtain a multi-perspective ray intersection area, and densify a Gaussian ellipsoid based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

[0018] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer is enabled to execute the method as described in the first aspect.

[0019] From the above, it can be seen that the embodiments of the present disclosure provide for collecting visual sensor data, including color images and depth images, segmenting and extracting features from the color images based on a large semantic segmentation model, obtaining active dynamic object extension masks and semantic images through deep clustering and feature encoding, obtaining robust non-active dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images, obtaining preliminary optimized poses based on the robust non-active dynamic points and the Gaussian ellipsoid in the Gaussian map, calculating the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtaining potential dynamic semantic features, and obtaining passive dynamic semantic features based on the semantic loss change of the pixel area corresponding to the potential dynamic semantic features in the key frame sequence. , based on the passive dynamic semantic features and the active dynamic object expansion mask, a final dynamic object mask is obtained, based on the passive dynamic semantic features and the robust non-active dynamic points, a static point is obtained, based on the final dynamic object mask, the loss of the input image and the rendered image in the static pixel area is obtained, the loss is back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result, based on the high-precision visual positioning result, the static point is inserted into the Gaussian map, based on the loss, the parameters of the Gaussian ellipsoid in the Gaussian map are optimized to obtain a high-fidelity map, based on the loss, a high-loss object is obtained, the ray intersection area of ​​the high-loss object in the key frame sequence is calculated, and a multi-view ray intersection area is obtained, and the Gaussian ellipsoid is densified based on the multi-view ray intersection area to obtain a high-fidelity object-level map. The present invention discloses a method for high-precision visual positioning and high-fidelity object-level map construction for indoor dynamic environments. The method solves the limitations of traditional VSLAM technology through steps such as visual sensor data preprocessing, dynamic object detection and deletion, high-precision visual positioning, high-fidelity map optimization and high-fidelity object-level map densification for indoor dynamic environments. The method can accurately and comprehensively delete dynamic objects in any indoor dynamic environment, achieve high-precision visual positioning, and construct a high-fidelity object-level map. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 A schematic diagram of a process for a high-precision visual positioning and high-fidelity object-level map construction method for an indoor dynamic environment provided by an embodiment of the present disclosure;

[0022] Figure 2 A schematic diagram of the structure of a high-precision visual positioning and high-fidelity object-level map building device for indoor dynamic environments provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, the principles and spirit of the present disclosure will be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0024] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0025] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The article "one" or "an" before an element does not exclude the presence of multiple such elements.

[0026] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.

[0027] As described in the background technology, existing visual positioning technologies often face many challenges in dynamic environments. Accurately and comprehensively identifying and deleting dynamic objects in images is a key step to improve the accuracy of visual positioning. Existing methods usually delete active dynamic objects based on active dynamic masks segmented from color images, and delete passive dynamic objects based on depth masks obtained from multi-frame depth reprojection errors. When the depth image accuracy is low, the method is prone to large errors. In a dynamic environment, large movements will cause the edges of dynamic objects in the color image and the depth image to be unable to be strictly aligned, and due to the time synchronization problem between the color image and the depth image, it is difficult for existing methods to delete all pixel areas containing dynamic objects, which in turn affects the accuracy of visual positioning and map construction tasks. In addition, existing visual synchronous positioning and map construction technologies often use point clouds to reconstruct scenes. Point clouds lack size and direction attributes, resulting in a large number of holes in the map, which seriously affects the visual effect.

[0028] In order to solve the above problems, the present disclosure provides a method, device and program product solution for high-precision visual positioning and high-fidelity object-level map construction for indoor dynamic environments, specifically including:

[0029] Visual sensor data are collected, including color images and depth images. The color images are segmented and feature extracted based on a semantic segmentation model. Active dynamic object extension masks and semantic images are obtained through deep clustering and feature encoding. Robust non-active dynamic points are obtained based on the active dynamic object extension masks, semantic images, color images and depth images. Preliminary optimized poses are obtained based on the robust non-active dynamic points and the Gaussian ellipsoid in the Gaussian map. The distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is calculated based on the preliminary optimized pose to obtain potential dynamic semantic features. Passive dynamic semantic features are obtained based on the semantic loss change of the potential dynamic semantic features in the corresponding pixel area in the key frame sequence. The semantic features and the active dynamic object expansion mask are used to obtain the final dynamic object mask. Based on the passive dynamic semantic features and the robust non-active dynamic points, static points are obtained. Based on the final dynamic object mask, the loss of the input image and the rendered image in the static pixel area is obtained. The loss is back-propagated using the Gaussian ellipsoid differentiable rendering characteristics to obtain a high-precision visual positioning result. Based on the high-precision visual positioning result, the static point is inserted into the Gaussian map. Based on the loss, the parameters of the Gaussian ellipsoid in the Gaussian map are optimized to obtain a high-fidelity map. Based on the loss, a high-loss object is obtained. The ray intersection area of ​​the high-loss object in the key frame sequence is calculated to obtain a multi-view ray intersection area. The Gaussian ellipsoid is densified based on the multi-view ray intersection area to obtain a high-fidelity object-level map. The advantage of this solution is that object-level deletion is performed on active dynamic objects and passive dynamic objects in indoor dynamic environments respectively, dynamic objects are deleted accurately and comprehensively, and high-precision visual positioning is achieved based on the Gaussian ellipsoid differentiable rendering characteristics to construct a high-fidelity object-level map.

[0030] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0031] refer to Figure 1 , a high-precision visual positioning and high-fidelity object-level map construction method for indoor dynamic environments, the method comprising the following steps:

[0032] Step S110, collect visual sensor data, including color images and depth images, segment and extract features of the color images based on the semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust non-active dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images.

[0033] Next, we will introduce how to generate object-level masks and semantic images:

[0034] In this exemplary embodiment, the object-level mask and the semantic image are obtained by the following method, including:

[0035] Use the Grounding DINO model to detect instance objects in the current color image, use SAM2 to obtain object-level masks, and encode the high-dimensional semantic features of each object based on a predefined encoding table to obtain a semantic image.

[0036] In specific implementation, object-level masks are generated, including:

[0037] In a color image, a text prompt is input into the Grounding DINO model, the object specified by the text prompt is detected, and the bounding box of the object in the color image is obtained. The bounding box is used as a prompt for SAM2 to generate an object-level detection mask for the current frame. The mask transfer function of SAM2 is used to continuously track the object in the current color image to obtain an object-level transfer mask for subsequent frames. During the object tracking process, when the number of objects in the color image is reduced to three-quarters of the original number of objects, the objects in the color image are re-detected to obtain the object bounding box in the detection frame. The object bounding box is used as a prompt for SAM2 to generate an object-level detection mask again. The object-level detection mask and the object-level transfer mask of the color image are associated to obtain the object-level mask that is ultimately used for visual positioning and high-fidelity object-level map construction.

[0038] In specific implementation, generating a semantic image includes:

[0039] Based on the predefined encoding table, the high-dimensional feature vector output by the Grounding DINO model is encoded to assign a low-dimensional semantic feature to each object. Combined with the object-level mask, a semantic image is obtained for visual localization and high-fidelity object-level map construction.

[0040] In the above exemplary embodiments, the method of obtaining the object-level mask and the semantic image is introduced. Next, the method of extending the active dynamic object mask will be introduced:

[0041] In this exemplary embodiment, the active dynamic object extension mask is obtained by the following method, including:

[0042] An active dynamic object mask is obtained based on the object level mask and the semantic image, and the active dynamic object mask is expanded based on the depth image to obtain an active dynamic object expansion mask.

[0043] In specific implementation, the active dynamic object mask is obtained, including:

[0044] The semantic features corresponding to the object with active motion capability are queried, the pixel area corresponding to the semantic features in the semantic image is acquired, and the active dynamic object mask is obtained.

[0045] In specific implementation, the active dynamic object mask is extended to include:

[0046] Use the K-Means clustering algorithm to cluster the depth image to obtain a depth clustering result. Query the pixel coordinates whose values ​​are true in the active dynamic object mask, count the depth category labels of the pixel coordinates in the depth clustering result, and obtain the active dynamic depth category. Determine whether the category of each pixel in the depth image belongs to the active dynamic depth category, and obtain a binary result for each pixel. Merge the binary results to obtain a depth mask. Perform morphological expansion on the active dynamic object mask to obtain an expansion mask. Take the intersection area of ​​the expansion mask and the depth mask to obtain the active dynamic object expansion mask.

[0047] In the above exemplary embodiment, a method for obtaining an active dynamic object expansion mask is introduced. Next, a method for obtaining a robust inactive dynamic point is introduced:

[0048] In this exemplary embodiment, the robust non-active dynamic point is obtained by the following method, including:

[0049] Based on the active dynamic object expansion mask, the depth image and the visual sensor internal parameters, the inactive dynamic points are obtained. Based on the inactive dynamic point coordinates and semantic features thereof, the edge noise in the inactive dynamic points is filtered to obtain robust inactive dynamic points.

[0050] During the specific implementation, non-active dynamic points are obtained, including:

[0051] The active dynamic object expansion mask is applied to the semantic image, color image and depth image, and the inactive dynamic pixel area is retained. The inactive dynamic pixel area in the semantic image, color image and depth image is associated. Using the pinhole camera model, the pixel coordinates are back-projected based on the depth image and the visual sensor intrinsic reference to obtain the inactive dynamic point in the camera coordinate system, which can be expressed as:

[0052]

[0053] Among them, K represents the intrinsic parameter of the visual sensor, u and v represent the pixel coordinates, D represents the value of the pixel (u, v) in the depth image, and X, Y, and Z represent the three-dimensional coordinates in the camera coordinate system.

[0054] During the specific implementation, robust non-active dynamic points are obtained, including:

[0055] Normalize the spatial coordinates and semantic features of the non-active dynamic points. Concatenate the normalized spatial coordinates and the normalized semantic features to obtain a high-dimensional feature vector. Use the DBSCAN clustering algorithm to cluster the high-dimensional feature vector to obtain a density clustering result. Count the information of each cluster in the density clustering result to obtain the number of points contained in the cluster and the aspect ratio of the cluster. Query the number of objects in the current color image, sort the clusters according to the number of points contained in the clusters, obtain the cluster with the largest number of objects, and obtain the cluster screening threshold. Delete the clusters whose number of points is less than the cluster screening threshold, and delete the clusters whose aspect ratio is greater than the specified aspect ratio threshold to obtain robust non-active dynamic points.

[0056] Step S120, obtaining a preliminary optimized pose based on the robust non-active dynamic points and the Gaussian ellipsoid in the Gaussian map, calculating the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtaining a latent dynamic semantic feature, obtaining a passive dynamic semantic feature based on the semantic loss change of the pixel area corresponding to the latent dynamic semantic feature in the key frame sequence, obtaining a final dynamic object mask based on the passive dynamic semantic feature and the active dynamic object expansion mask, and obtaining a static point based on the passive dynamic semantic feature and the robust non-active dynamic point.

[0057] Next, we will introduce the method of preliminary optimization of posture:

[0058] In this exemplary embodiment, the preliminary optimized posture is obtained by the following method, including:

[0059] The covariance of each point in the robust non-active dynamic point is calculated, the central coordinates and covariance matrix of each Gaussian ellipsoid in the Gaussian map are restored, the correspondence between the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map is obtained by using the nearest neighbor search, the objective function is constructed, and the maximum likelihood estimation is used to obtain the optimal transformation of the current sample, that is, the preliminary optimized posture.

[0060] In a specific implementation, calculating the covariance of the robust non-active dynamic point includes:

[0061] The K nearest neighbors of each point in the robust non-active dynamic point are calculated, and the covariance matrix of each point is calculated using the K nearest neighbors to obtain the covariance of the robust non-active dynamic point, which can be expressed as:

[0062]

[0063] where Σ f represents the covariance of the robust inactive dynamic point, σ ab represents the covariance between the a and b dimensions in the robust non-active dynamic point, which can be expressed as:

[0064]

[0065] where x ka represents the coordinate value of the kth point in dimension a, Represents the average coordinate value in dimension a, i.e., the centroid, N K Indicates that the covariance matrix of each point is calculated using K nearest neighbors.

[0066] In the specific implementation, the center coordinates and covariance matrix of each Gaussian ellipsoid in the Gaussian map are restored, including:

[0067] The center coordinates, axial radius and rotation of each Gaussian ellipsoid in the Gaussian map are taken out, and the covariance matrix of each Gaussian ellipsoid is restored by the axial radius and the rotation, which can be expressed as:

[0068] Σ m =RSS T R T

[0069] Where R represents the rotation property of the Gaussian ellipsoid, S represents the scale property of the Gaussian ellipsoid, Σ m Represents the covariance matrix of the Gaussian ellipsoid.

[0070] In specific implementation, the nearest neighbor search is used to obtain the correspondence between the robust non-active dynamic points and the Gaussian ellipsoid in the Gaussian map, and the matching pairs with the same subscript are obtained.

[0071] During the specific implementation, the objective function is constructed, including:

[0072] Define each point as an independent Gaussian distribution, which can be expressed as:

[0073]

[0074] in represents the coordinate observation value of point i, represents the covariance matrix of point i, represents the Gaussian distribution of point i.

[0075] Define each Gaussian ellipsoid as an independent Gaussian distribution, which can be expressed as:

[0076]

[0077] in represents the coordinate observation value of Gaussian ellipsoid i, represents the covariance matrix of Gaussian ellipsoid i, represents the Gaussian distribution of Gaussian ellipsoid i.

[0078] Based on the matching pairs with the same subscript, an error term is defined, which can be expressed as:

[0079]

[0080] in, represents the Gaussian distribution of Gaussian ellipsoid i, represents the Gaussian distribution of point i, and T represents the pose transformation between the robust non-active dynamic point and the Gaussian map.

[0081] Assume that there is an optimal pose transformation T * , obviously:

[0082]

[0083] because and are independent Gaussian random variables, and their linear combinations is also a Gaussian random variable and can be expressed as:

[0084]

[0085] Where T * represents the optimal pose transformation assumed to exist, Represents the distribution of the Gaussian random variable corresponding to the error term.

[0086] In specific implementation, the maximum likelihood estimation is used to obtain the best transformation of the current sample, including:

[0087] The optimal pose transformation T * It can be seen as the need to The distribution parameters estimated in the probability distribution of . Using the maximum likelihood estimation method, we find the probability that can maximize the current sample Probability T * , which can be expressed as:

[0088]

[0089] in Represents finding a pose transformation that maximizes the likelihood function. represents the probability density function of the Gaussian random variable corresponding to the error term, Represents finding a pose transformation that minimizes the objective function, T** Represents the optimal pose transformation of the current sample, that is, the initially optimized pose.

[0090] In the above exemplary embodiment, a method for preliminarily optimizing the posture is introduced. Next, a method for obtaining potential dynamic semantic features will be introduced:

[0091] In this exemplary embodiment, the potential dynamic semantic features are obtained by the following method, including:

[0092] Based on the preliminary optimized pose, the robust inactive dynamic point is converted to a map coordinate system, and the distance between the center of the robust inactive dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is compared to obtain potential dynamic semantic features.

[0093] In a specific implementation, the robust non-active dynamic point is converted into a map coordinate system based on the preliminary optimized pose, including:

[0094] The robust inactive dynamic points in the current camera coordinate system are converted to the map coordinate system using the preliminary optimized pose to obtain the robust inactive dynamic points in the map coordinate system.

[0095] In a specific implementation, the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is compared to obtain potential dynamic semantic features, including:

[0096] For each semantic feature of an object, the center of the corresponding point in the robust non-active dynamic point in the map coordinate system and the center of the corresponding Gaussian ellipsoid in the current Gaussian map are calculated, and the distance between the two centers is calculated. The semantic features corresponding to the objects whose distance exceeds the threshold are marked to obtain the latent dynamic semantic features.

[0097] In the above exemplary embodiment, a method for obtaining potential dynamic semantic features is introduced. Next, a method for obtaining passive dynamic semantic features will be introduced:

[0098] In this exemplary embodiment, the passive dynamic semantic feature is obtained by the following method, including:

[0099] Based on the semantic image of the key frame sequence, the pixel area corresponding to the potential dynamic semantic feature is obtained, and the semantic loss change of the pixel area between frames is calculated to obtain the passive dynamic semantic feature.

[0100] In a specific implementation, obtaining the pixel area corresponding to the potential dynamic semantic feature includes:

[0101] The semantic image of each key frame in the key frame sequence is obtained, and the pixel area whose value is the potential dynamic semantic feature in the semantic image is merged to obtain the pixel area corresponding to the potential dynamic semantic feature in the key frame sequence.

[0102] During the specific implementation, the passive dynamic semantic features are obtained, including:

[0103] For each pixel region, the semantic image of the current Gaussian map under the perspective of the current frame and the key frame is rendered, and the loss between the rendered semantic image and the input semantic image in the pixel region is calculated to obtain the semantic loss of the current frame and the key frame, which can be expressed as:

[0104]

[0105] Where S represents the input semantic image, represents the rendered semantic image, N s Indicates the number of valid pixels in the semantic image, represents the l1 loss between each valid pixel j of the rendered semantic image and the input semantic image, represents the semantic loss.

[0106] The variation range of the semantic losses is calculated, the pixel regions whose variation range is greater than a threshold are counted, and the potential dynamic semantic features corresponding to the pixel regions are queried to obtain passive dynamic semantic features.

[0107] In the above exemplary embodiment, a method for obtaining passive dynamic semantic features is introduced. Next, a method for obtaining the final dynamic object mask will be introduced:

[0108] In this exemplary embodiment, the final dynamic object mask is obtained by the following method, including:

[0109] Query the pixel area whose value is the passive dynamic semantic feature in the semantic image of the current frame, merge the pixel area to obtain a passive dynamic object mask, merge the passive dynamic object mask and the active dynamic object extension mask to obtain a final dynamic object mask.

[0110] In specific implementation, a passive dynamic object mask is obtained, including:

[0111] Determine whether each pixel value in the current frame semantic image belongs to the passive dynamic semantic feature, and obtain a binary result of each pixel. Merge the binary results to obtain a passive dynamic object mask.

[0112] In specific implementation, the final dynamic object mask is obtained, including:

[0113] The passive dynamic object mask and the active dynamic object extension mask are merged pixel by pixel to obtain a final dynamic object mask.

[0114] Step S130: obtaining the loss of the input image and the rendered image in the static pixel area based on the final dynamic object mask, and back-propagating the loss using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result.

[0115] Next, we will introduce the high-precision visual positioning method:

[0116] In this exemplary embodiment, a high-precision visual positioning result is obtained by the following method, including:

[0117] Based on the final dynamic object mask, the static pixel areas of the input image and the rendered image are retained, the color, depth and semantic losses of the static pixel areas are calculated, and the losses are back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result.

[0118] In a specific implementation, the static pixel areas of the input image and the rendered image are obtained, including:

[0119] The final dynamic object mask is applied to the input image and the rendered image, the dynamic pixel area value is set to zero, and the static pixel area value remains unchanged, so as to obtain the static pixel area of ​​the input image and the rendered image.

[0120] In a specific implementation, calculating the color, depth and semantic loss of the static pixel area includes:

[0121] The color, depth and semantic losses of the static pixel area are calculated, which can be expressed as:

[0122]

[0123] Where C represents the input color image, Indicates rendering of color images, N r Indicates the number of valid pixels in a color image. represents the l1 loss between each valid pixel j of the rendered color image and the input color image, represents the pixel value of the final dynamic object mask at pixel j, Indicates the inverse operation. represents the color loss.

[0124]

[0125] Where D represents the input depth image, Represents the rendered depth image, N d Indicates the number of valid pixels in the depth image, represents the l1 loss between each valid pixel j of the rendered depth image and the input depth image, represents the pixel value of the final dynamic object mask at pixel j, Indicates the inverse operation. represents the depth loss.

[0126]

[0127] Where S represents the input semantic image, represents the rendered semantic image, N s Indicates the number of valid pixels in the semantic image, represents the l1 loss between each valid pixel j of the rendered semantic image and the input semantic image, represents the pixel value of the final dynamic object mask at pixel j, Indicates the inverse operation. represents the semantic loss.

[0128] During the specific implementation, high-precision visual positioning results are obtained, including:

[0129] A loss function consisting of the color, depth and semantic losses of the static pixel area is constructed, and the loss function is minimized to obtain a high-precision visual positioning result, which can be expressed as:

[0130]

[0131] in represents the color loss, λ1 represents the weight of the color loss, represents the depth loss, λ2 represents the weight of the depth loss, represents the semantic loss, λ3 represents the weight of the semantic loss, Represents a loss function obtained by weighted addition of the color loss, the depth loss, and the semantic loss.

[0132] Step S140: inserting the static point into a Gaussian map based on the high-precision visual positioning result, optimizing the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss, and obtaining a high-fidelity map.

[0133] Next, a method for inserting the static points into the Gaussian map is introduced:

[0134] In this exemplary embodiment, the static point is inserted into the Gaussian map by the following method, including:

[0135] Based on the passive dynamic semantic features, the passive dynamic points in the robust inactive dynamic points are deleted to obtain static points, and the static points are inserted into the Gaussian map based on the high-precision visual positioning results.

[0136] During the specific implementation, the static points are obtained, including:

[0137] The points containing the passive dynamic semantic features in the robust non-active dynamic points are queried to obtain passive dynamic points, and the passive dynamic points are deleted to obtain static points.

[0138] In a specific implementation, inserting the static point into the Gaussian map includes:

[0139] Based on the high-precision visual positioning result, the static point in the current camera coordinate system is converted to the map coordinate system, and the static point in the map coordinate system is inserted into the Gaussian map.

[0140] In the above exemplary embodiment, a method of inserting the static point into the Gaussian map is introduced. Next, a method of optimizing the high-fidelity map is introduced:

[0141] In this exemplary embodiment, the high-fidelity map is optimized by the following method, including:

[0142] The parameters of the Gaussian ellipsoid in the Gaussian map are optimized based on the loss to obtain a high-fidelity map.

[0143] During the specific implementation, a high-fidelity map is obtained, including:

[0144] A loss function consisting of color, depth and semantic losses of the static pixel area is constructed, the loss function is minimized, and the high-fidelity map is optimized.

[0145] Step S150: obtaining a high-loss object based on the loss, calculating the ray intersection area of ​​the high-loss object in the key frame sequence, obtaining a multi-perspective ray intersection area, and densifying a Gaussian ellipsoid based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

[0146] Next, we will introduce the method of obtaining the light intersection area:

[0147] In this exemplary embodiment, the multi-view light intersection area is obtained by the following method, including:

[0148] Based on the loss, a high-loss object in the current frame is queried, and a ray intersection area of ​​the high-loss object in a key frame sequence is calculated to obtain a multi-view ray intersection area.

[0149] In specific implementation, the multi-view light intersection area is obtained, including:

[0150] Based on the loss, query the high-loss object in the current frame image to obtain the semantic features of the high-loss object. Obtain the pixel area corresponding to the semantic feature in the key frame sequence, and generate the bounding box of the high-loss object based on the pixel area in each key frame. Based on the high-precision visual positioning result, back-project the light on the bounding box, calculate the light intersection area of ​​the high-loss object under the perspective of the current frame and the key frame sequence, and obtain the multi-perspective light intersection area.

[0151] In the above exemplary embodiment, a method for obtaining a multi-view light intersection area is introduced. Next, a method for obtaining a high-fidelity object-level map will be introduced:

[0152] In this exemplary embodiment, a high-fidelity object-level map is obtained by the following method, including:

[0153] Based on the multi-view light intersection region, the Gaussian ellipsoid inside the multi-view light intersection region is densified to obtain a high-fidelity object-level map.

[0154] During the specific implementation, a high-fidelity object-level map is obtained, including:

[0155] Query the Gaussian ellipsoid whose center is inside the multi-view light intersection area, screen the Gaussian ellipsoids with gradient greater than a threshold, clone the Gaussian ellipsoids with gradient greater than a threshold and scale less than a threshold, split the Gaussian ellipsoids with gradient greater than a threshold and scale greater than a threshold, and obtain a high-fidelity object-level map.

[0156] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a high-precision visual positioning and high-fidelity object-level map construction device for indoor dynamic environments.

[0157] refer to Figure 2 The high-precision visual positioning and high-fidelity object-level map construction device for indoor dynamic environments includes:

[0158] The visual sensor data preprocessing module 210 for indoor dynamic environment is configured to collect visual sensor data, including color images and depth images, segment and extract features of the color images based on the semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust non-active dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images;

[0159] The dynamic object detection and deletion module 220 is configured to obtain a preliminary optimized pose based on the robust non-active dynamic points and the Gaussian ellipsoid in the Gaussian map, calculate the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtain a potential dynamic semantic feature, obtain a passive dynamic semantic feature based on the semantic loss change of the pixel area corresponding to the potential dynamic semantic feature in the key frame sequence, obtain a final dynamic object mask based on the passive dynamic semantic feature and the active dynamic object expansion mask, and obtain a static point based on the passive dynamic semantic feature and the robust non-active dynamic point;

[0160] The high-precision visual positioning module 230 is configured to obtain the loss of the input image and the rendered image in the static pixel area based on the final dynamic object mask, and back-propagate the loss using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result;

[0161] A high-fidelity map optimization module 240 is configured to insert the static point into a Gaussian map based on the high-precision visual positioning result, and optimize the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss to obtain a high-fidelity map;

[0162] The high-fidelity object-level map densification module 250 is configured to obtain a high-loss object based on the loss, calculate the ray intersection area of ​​the high-loss object in the key frame sequence, obtain a multi-perspective ray intersection area, and densify a Gaussian ellipsoid based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

[0163] In this exemplary embodiment, the visual sensor data preprocessing module 210 for indoor dynamic environment is specifically configured as follows:

[0164] The Grounding DINO model is used to detect instance objects in the current color image, and the object-level mask is obtained using SAM2. The high-dimensional semantic features of each object are encoded based on a predefined encoding table to obtain a semantic image. Based on the object-level mask and the semantic image, an active dynamic object mask is obtained. The active dynamic object mask is expanded based on the depth image to obtain an active dynamic object extended mask. Based on the active dynamic object extended mask, the depth image and the visual sensor internal parameters, inactive dynamic points are obtained. Based on the coordinates of the inactive dynamic points and their semantic features, the edge noise in the inactive dynamic points is filtered to obtain robust inactive dynamic points.

[0165] In this exemplary embodiment, the dynamic object detection and deletion module 220 is specifically configured as follows:

[0166] The covariance of each point in the robust non-active dynamic point is calculated, and the central coordinates and covariance matrix of each Gaussian ellipsoid in the Gaussian map are restored. The corresponding relationship between the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map is obtained by using the nearest neighbor search. The objective function is constructed, and the maximum likelihood estimation is used to obtain the optimal transformation of the current sample, that is, the preliminary optimized pose. Based on the preliminary optimized pose, the robust non-active dynamic point is converted to the map coordinate system, and the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is compared to obtain the potential dynamic semantic feature. Based on the semantic image of the key frame sequence, the pixel area corresponding to the potential dynamic semantic feature is obtained, and the semantic loss change of the pixel area between frames is calculated to obtain the passive dynamic semantic feature. The pixel area whose value is the passive dynamic semantic feature in the semantic image of the current frame is queried, and the pixel area is merged to obtain the passive dynamic object mask. The passive dynamic object mask and the active dynamic object extension mask are merged to obtain the final dynamic object mask.

[0167] In this exemplary embodiment, the high-precision visual positioning module 230 is specifically configured as follows:

[0168] Based on the final dynamic object mask, the static pixel areas of the input image and the rendered image are retained, the color, depth and semantic losses of the static pixel areas are calculated, and the losses are back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result.

[0169] In this exemplary embodiment, the high-fidelity map optimization module 240 is specifically configured as follows:

[0170] Based on the passive dynamic semantic features, the passive dynamic points in the robust non-active dynamic points are deleted to obtain static points. Based on the high-precision visual positioning results, the static points are inserted into the Gaussian map. Based on the loss, the parameters of the Gaussian ellipsoid in the Gaussian map are optimized to obtain a high-fidelity map.

[0171] In this exemplary embodiment, the high-fidelity object-level map densification module 250 is specifically configured as follows:

[0172] Based on the loss, high-loss objects in the current frame are queried, and the ray intersection area of ​​the high-loss object in the key frame sequence is calculated to obtain a multi-perspective ray intersection area. Based on the multi-perspective ray intersection area, the Gaussian ellipsoid inside the multi-perspective ray intersection area is densified to obtain a high-fidelity object-level map.

[0173] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0174] The device of the above embodiment is used to implement the corresponding high-precision visual positioning and high-fidelity object-level map construction method for indoor dynamic environments in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0175] Based on the same inventive concept, corresponding to the method for building a high-precision visual positioning and high-fidelity object-level map for indoor dynamic environments described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor executes the method for building a high-precision visual positioning and high-fidelity object-level map for indoor dynamic environments. Corresponding to the execution subject corresponding to each step in each embodiment of the method for building a high-precision visual positioning and high-fidelity object-level map for indoor dynamic environments, the processor that executes the corresponding step may belong to the corresponding execution subject.

[0176] The computer program product of the above-mentioned embodiment is used to enable the computer and / or the processor to execute the high-precision visual positioning and high-fidelity object-level map construction method for indoor dynamic environments as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0177] In addition, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.

[0178] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0179] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0180] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0181] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.

[0182] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims. The scope of the attached claims conforms to the broadest interpretation, thereby including all such modifications and equivalent structures and functions.

Claims

1. A high-precision visual positioning and high-fidelity object-level map construction method for indoor dynamic environments, characterized in that: include: Collect visual sensor data, including color images and depth images, segment and extract features from the color images based on a semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust inactive dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images; A preliminary optimized pose is obtained based on the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map, and a distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is calculated based on the preliminary optimized pose to obtain a potential dynamic semantic feature, and a passive dynamic semantic feature is obtained based on the semantic loss change of the pixel area corresponding to the potential dynamic semantic feature in the key frame sequence, and a final dynamic object mask is obtained based on the passive dynamic semantic feature and the active dynamic object expansion mask, and a static point is obtained based on the passive dynamic semantic feature and the robust non-active dynamic point; Based on the final dynamic object mask, the loss of the input image and the rendered image in the static pixel area is obtained, and the loss is back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result; Inserting the static point into the Gaussian map based on the high-precision visual positioning result, optimizing the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss, and obtaining a high-fidelity map; A high-loss object is obtained based on the loss, and a ray intersection area of ​​the high-loss object in a key frame sequence is calculated to obtain a multi-perspective ray intersection area. A Gaussian ellipsoid is densified based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

2. The method according to claim 1, characterized in that: The visual sensor data is collected, including color images and depth images, the color images are segmented and feature extracted based on the semantic segmentation model, active dynamic object extension masks and semantic images are obtained through deep clustering and feature encoding, and robust non-active dynamic points are obtained based on the active dynamic object extension masks, semantic images, color images and depth images, including: Use the Grounding DINO model to detect instance objects in the current color image, use SAM2 to obtain object-level masks, and encode the high-dimensional semantic features of each object based on a predefined encoding table to obtain a semantic image; Based on the object-level mask and the semantic image, an active dynamic object mask is obtained, and based on the depth image, the active dynamic object mask is extended to obtain an active dynamic object extended mask; Based on the active dynamic object expansion mask, the depth image and the visual sensor internal parameters, the inactive dynamic points are obtained. Based on the inactive dynamic point coordinates and semantic features thereof, the edge noise in the inactive dynamic points is filtered to obtain robust inactive dynamic points.

3. The method according to claim 1, characterized in that The method comprises: obtaining a preliminary optimized pose based on the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map, calculating the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtaining a potential dynamic semantic feature, obtaining a passive dynamic semantic feature based on the semantic loss change of the pixel area corresponding to the potential dynamic semantic feature in the key frame sequence, obtaining a final dynamic object mask based on the passive dynamic semantic feature and the active dynamic object expansion mask, and obtaining a static point based on the passive dynamic semantic feature and the robust non-active dynamic point, including: Calculate the covariance of each point in the robust non-active dynamic point, restore the center coordinates and covariance matrix of each Gaussian ellipsoid in the Gaussian map, use the nearest neighbor search to obtain the correspondence between the robust non-active dynamic point and the Gaussian ellipsoid in the Gaussian map, construct the objective function, and use the maximum likelihood estimation to obtain the optimal transformation of the current sample, that is, the preliminary optimized posture; Based on the preliminary optimized pose, the robust inactive dynamic point is converted into a map coordinate system, and the distance between the center of the robust inactive dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map is compared to obtain a potential dynamic semantic feature; Based on the semantic image of the key frame sequence, a pixel area corresponding to the potential dynamic semantic feature is obtained, and the semantic loss change of the pixel area between frames is calculated to obtain the passive dynamic semantic feature; Query the pixel area whose value is the passive dynamic semantic feature in the semantic image of the current frame, merge the pixel area to obtain a passive dynamic object mask, merge the passive dynamic object mask and the active dynamic object extension mask to obtain a final dynamic object mask.

4. The method according to claim 1, characterized in that: The loss of the input image and the rendered image in the static pixel area is obtained based on the final dynamic object mask, and the loss is back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result, including: Based on the final dynamic object mask, the static pixel areas of the input image and the rendered image are retained, the color, depth and semantic losses of the static pixel areas are calculated, and the losses are back-propagated using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result.

5. The method according to claim 1, characterized in that The method of inserting the static point into a Gaussian map based on the high-precision visual positioning result, and optimizing the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss to obtain a high-fidelity map includes: Based on the passive dynamic semantic features, the passive dynamic points in the robust inactive dynamic points are deleted to obtain static points, and the static points are inserted into the Gaussian map based on the high-precision visual positioning results; The parameters of the Gaussian ellipsoid in the Gaussian map are optimized based on the loss to obtain a high-fidelity map.

6. The method according to claim 1, characterized in that The method of obtaining a high-loss object based on the loss, calculating a ray intersection area of ​​the high-loss object in a key frame sequence, obtaining a multi-view ray intersection area, and densifying a Gaussian ellipsoid based on the multi-view ray intersection area to obtain a high-fidelity object-level map includes: Based on the loss, query the high-loss object in the current frame, calculate the light intersection area of ​​the high-loss object in the key frame sequence, and obtain the multi-view light intersection area; Based on the multi-view light intersection region, the Gaussian ellipsoid inside the multi-view light intersection region is densified to obtain a high-fidelity object-level map.

7. A high-precision visual positioning and high-fidelity object-level map construction device for indoor dynamic environments, characterized in that: include: A visual sensor data preprocessing module for indoor dynamic environments, configured to collect visual sensor data, including color images and depth images, segment and extract features from the color images based on a semantic segmentation model, obtain active dynamic object extension masks and semantic images through deep clustering and feature encoding, and obtain robust inactive dynamic points based on the active dynamic object extension masks, semantic images, color images and depth images; A dynamic object detection and deletion module is configured to obtain a preliminary optimized pose based on robust non-active dynamic points and a Gaussian ellipsoid in a Gaussian map, calculate the distance between the center of the robust non-active dynamic point corresponding to each object and the center of the Gaussian ellipsoid in the Gaussian map based on the preliminary optimized pose, obtain a latent dynamic semantic feature, obtain a passive dynamic semantic feature based on a semantic loss change of a pixel area corresponding to the latent dynamic semantic feature in a key frame sequence, obtain a final dynamic object mask based on the passive dynamic semantic feature and an active dynamic object expansion mask, and obtain a static point based on the passive dynamic semantic feature and the robust non-active dynamic point; A high-precision visual positioning module is configured to obtain the loss of the input image and the rendered image in the static pixel area based on the final dynamic object mask, and to back-propagate the loss using the differentiable rendering characteristics of the Gaussian ellipsoid to obtain a high-precision visual positioning result; A high-fidelity map optimization module, configured to insert the static point into a Gaussian map based on the high-precision visual positioning result, and optimize the parameters of the Gaussian ellipsoid in the Gaussian map based on the loss to obtain a high-fidelity map; A high-fidelity object-level map densification module is configured to obtain a high-loss object based on the loss, calculate the ray intersection area of ​​the high-loss object in a key frame sequence, obtain a multi-perspective ray intersection area, and densify a Gaussian ellipsoid based on the multi-perspective ray intersection area to obtain a high-fidelity object-level map.

8. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Visual positioning and static map construction method and system in dynamic environment

    CN112991447A

  • Robustness dynamic vision synchronous positioning and mapping method and system

    CN117557748A

  • Method and system for reconstructing dense Gaussian map in dynamic environment

    CN118071873A

  • Indoor complex scene high-fidelity real-time rendering method based on three-dimensional Gaussian representation

    CN118096988A

  • Large-scale three-dimensional scene real-time reconstruction method based on Gaussian expression

    CN118314280A