Three-dimensional Gaussian environment sensing and positioning method and device for processing dynamic interference based on motion probability

The method improves three-dimensional Gaussian environment perception and positioning by deploying a YOLOv7 model on mobile devices to classify Gaussian points dynamically, optimizing rendering losses for enhanced pose estimation and reduced scene artifacts.

CN120318331APending Publication Date: 2025-07-15HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510462019.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing three-dimensional Gaussian environment perception and positioning systems have limited accuracy in dynamic environments and cannot be directly deployed to mobile platforms. The masks relying on optical flow estimation lead to reduced artifacts and tracking accuracy.

Method used

Using a motion probability processing method, the motion probability of Gaussian points is fused through semantic detection and geometric constraints, combined with Gaussian pyramid network and polar line constraints, the rendering loss function is constructed to optimize the pose estimation, reduce artifacts and improve accuracy.

Benefits of technology

Implement efficient dynamic interference processing on the mobile terminal, improve pose estimation accuracy, reduce artifacts, enhance system adaptability, and optimize rendering effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318331A_ABST
    Figure CN120318331A_ABST
Patent Text Reader

Abstract

The invention relates to a three-dimensional Gaussian environment sensing and positioning method for processing dynamic interference based on motion probability. The method comprises the following steps: obtaining the motion probability of a three-dimensional Gaussian point; obtaining dynamic and static attributes of the Gaussian point according to the initial motion probability of the Gaussian point; the final Gaussian point motion probability is obtained; constructing loss based on motion probability and loss based on edge distortion to obtain an overall rendering loss function; and optimizing parameters in the overall rendering loss function and the motion probability of the final Gaussian point. According to the method, a semantic box is provided by adopting a YOLOv7 model deployed by a mobile terminal, the motion probability of Gaussian points is obtained by performing variance weighted fusion on instance segmentation and a geometric method, and the Gaussian points are classified into dynamic or static states, so that the pose estimation precision is improved, and artifacts in a scene map are reduced; according to the method, the map is densified, and static Gaussian points which are wrongly marked as dynamic and corresponding feature points are recovered, so that simultaneous optimization of tracking and rendering information is realized, and redundant decoupling processing is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional Gaussian environment perception and positioning, and in particular to a three-dimensional Gaussian environment perception and positioning method and device for processing dynamic interference based on motion probability. Background Art

[0002] Three-dimensional environment perception and positioning is a core technology for autonomous systems, capable of constructing accurate maps and performing self-positioning in dynamic environments. Three-dimensional Gaussian splatting has demonstrated powerful real-time performance and high-fidelity texture representation in novel view synthesis, which has promoted the emergence of three-dimensional Gaussian environment perception and positioning systems. However, traditional three-dimensional Gaussian environment perception and positioning systems usually assume that the scene is static, which performs well in virtual environments but leads to significant pose drift and dynamic artifacts when facing dynamic interference in real scenes.

[0003] With the rapid development of deep learning, some existing three-dimensional Gaussian environment perception and positioning systems have adopted adaptive strategies. The core method mainly obtains the mask of dynamic objects by using a deep learning-based optical flow estimation system, and then adopts a divide-and-conquer strategy to adjust the rendering function, introducing a variety of complex rendering losses to remove dynamic elements. These methods usually face two main problems: First, since the current optical flow estimation systems generally rely on additional computing power support and cannot be directly deployed to mobile platforms, which limits the real-time performance of the system; Second, in order not to introduce additional information, existing end-to-end systems rely on the mask of optical flow estimation for the removal and tracking correction of dynamic objects. However, due to the accuracy of the optical flow mask being restricted by illumination and the state of moving objects in the actual environment, the pixel matching accuracy is limited, which leads to a decrease in the actual rendering artifacts and the tracking accuracy.

[0004] Therefore, there is a need for an efficient and accurate three-dimensional Gaussian environment perception and positioning method for processing dynamic interference. Summary of the Invention

[0005] To solve the problems that the existing system relies on additional computing power support, cannot be directly deployed to mobile platforms, and has limited accuracy in dynamic environments, the primary objective of the present invention is to provide a three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability that can be deployed to mobile devices, improve the pose estimation accuracy, and achieve effective artifact elimination.

[0006] To achieve the above objective, the present invention adopts the following technical solutions: A three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability, the method includes the following steps in sequence:

[0007] (1) Obtain the motion probability of 3D Gaussian points: Pass multiple input RGB images through the object detection network to obtain the semantic motion probability of the feature points in the image frame. Obtain the geometric motion probability of the feature points through multi-view geometric constraints. Fuse the semantic motion probability and the geometric motion probability through variance weighting to obtain the observed preliminary motion probability of the Gaussian points;

[0008] (2) Obtain the dynamic and static attributes of Gaussian points according to the preliminary motion probability of Gaussian points: Based on the median of the preliminary motion probability of Gaussian points in the current image frame, set an adaptive threshold. Gaussian points exceeding the adaptive threshold are marked as dynamic, otherwise marked as static;

[0009] (3) Reverse map the marked Gaussian points to the front-end feature tracking system through the Gaussian pyramid network, further judge using the epipolar constraint based on feature points, and finally densely map back to Gaussian points to obtain the final motion probability of Gaussian points;

[0010] (4) Construct a loss based on motion probability and a loss based on edge distortion based on the final motion probability of Gaussian points to obtain the overall rendering loss function;

[0011] (5) Optimize the parameters in the overall rendering loss function and the final motion probability of Gaussian points.

[0012] Step (1) specifically includes the following steps:

[0013] (1a) Use an RGB-D camera to obtain the frame stream of the scene. The RGB-D camera is fixed on the mobile device. The mobile device uses a mobile robot or a drone. Load the YOLOv7 model for detecting semantic detection boxes on the mobile device, and assign the corresponding semantic motion probability MP i to the feature point p ins (p i );

[0014] (1b) The geometric constraint uses the reprojection error. Select the key frame with the largest number of co-visible views from the historical key frames. The historical key frames are RGB images saved at past time points. Calculate the reprojection error between the current frame Gaussian point and its corresponding feature point p i :

[0015]

[0016] where π is the projection function, R is the rotation matrix, t is the translation vector, and both R and t are the camera pose T CW ; Inside the same instance area, calculate the variance σ of the reprojection error of the feature point e and the maximum error e max, and calculate the geometric motion probability MP of the feature point p within the region i of geo (p i ):

[0017]

[0018] where σ is a predefined variance threshold; n represents the number of feature points;

[0019] For points located in the outer region of the semantic detection box detected in step (1a):

[0020]

[0021] (1c) Use the variance-weighted fusion strategy to combine the obtained geometric motion probability with the semantic motion probability to obtain the motion probability M of the feature point p i : p M p

[0022] p = MP ins + H(MP geo - Mp ins )

[0023] where H is used to measure the contributions of MP ins and MP geo to M p , and the formula is:

[0024]

[0025] In the formula, is the confidence of instance segmentation, is the variance of the reprojection error; the motion probability of the initial Gaussian point depends on the information of the corresponding feature point observed in each image frame, so is expressed as a function of the motion probability M of the observed feature point p :

[0026]

[0027] where m represents the number of observed feature points, is the Gaussian point the motion probability of the observed feature point.

[0028] Step (2) specifically refers to: the motion probability of the initial Gaussian point forms a set Calculate the median of

[0029] The adaptive threshold τ is as follows:

[0030]

[0031] When the Gaussian points are marked as dynamic, otherwise they are static.

[0032] Step (3) specifically includes the following steps in sequence:

[0033] (3a) Construct a Gaussian pyramid network:

[0034] (3b) Reverse-map the marked Gaussian points through the constructed Gaussian pyramid network to the front-end feature tracking system, which is deployed on the mobile device, to obtain the feature points in the image frame corresponding to the Gaussian points. Further screen out the missing dynamic Gaussian points through the epipolar error of the feature points obtained by reverse mapping, and at the same time recover the static Gaussian points that are mislabeled as dynamic;

[0035] Assume that p1 and p2 are the matching feature points of the Gaussian points in the previous image frame and the current image frame respectively L1 and L2 are the epipolar lines in the corresponding frames respectively. Calculate L1 and L2 according to the fundamental matrix F:

[0036]

[0037] where A1, B1, C1, A2, B2, C2 are all the coefficients of the two-dimensional linear equation, K is the internal parameter matrix of the camera, t ∧ is the skew-symmetric matrix of the translation vector, (u, v) are the pixel coordinates of the feature point, and R is the rotation matrix;

[0038] The distance d i between the feature point p i and the corresponding epipolar line is calculated by the following formula:

[0039]

[0040] (3c) If d1 + d2 < ε, then the feature point p i will be densified, and the feature point p i is remapped back to the Gaussian point through the Gaussian function:

[0041]

[0042] where ∑ is the covariance matrix, o ∈ [0, 1] represents the opacity value, S is the scaling matrix, and R is the rotation matrix; d1 is the distance from the feature point p1 to the epipolar line L1, and d2 is the distance from the feature point p2 to the epipolar line L2; ε is the empirical threshold, with a value of 0.6;

[0043] Finally, for the Gaussian points remapped back, if di If it is less than the empirical threshold ε, the remapped Gaussian point is marked as static. If d i is greater than the empirical threshold ε, the remapped Gaussian point is marked as dynamic.

[0044] Step (4) specifically includes the following steps:

[0045] (4a) Construct the loss based on the motion probability: Use the photometric loss to measure the color error between the rendered image and the real input image, and combine it with the motion probability of the final Gaussian points to obtain the adjusted photometric loss L pho :

[0046]

[0047] where represents the photometric rendering of the Gaussian points using the camera pose T CW for the Gaussian points and represents the real color map corresponding to the given pose, is the motion probability of the final Gaussian points; the camera pose T CW adopts the transformation matrix;

[0048] Use the depth loss to constrain the depth of the 3D Gaussian points to obtain the adjusted depth loss L depth as:

[0049]

[0050] where represents the depth rasterization process, represents the real depth map corresponding to the given pose;

[0051] According to the position information of the Gaussian points themselves, introduce a penalty term based on the motion probability to constrain the influence of the truly dynamic Gaussian on the rendering, and obtain the loss L based on the motion probability MP :

[0052]

[0053] where represents the real 3D position of the Gaussian points, represents the estimated position;

[0054] (4b) Introduce the edge distortion loss to enhance the geometric consistency of data association between adjacent frames, that is, for the feature point p in the image frame i i , re-project it onto the frame j through the distortion operation:

[0055]

[0056] Among them, D and T ji represent depth information and transformation matrix respectively, is the homogeneous coordinate of p i ; for an edge set ε i , the loss L edge based on edge distortion is calculated as follows:

[0057]

[0058] Among them, ρ is a robust weight function used to reduce the influence of abnormal residuals;

[0059] (4c) Use different weights λ i to adjust the importance of motion probability in the camera tracking and mapping process, so as to obtain the overall rendering loss function L G :

[0060] L G = λ1·L pho + λ2·L depth + λ3·L MP + λ4·L edge

[0061] Among them, λ1 = 0.9, λ2 = 0.1, λ3 = 500, λ4 = 300.

[0062] Step (3a) specifically includes the following steps in sequence:

[0063] (3a1) Form a pyramid structure: from the bottom RGB image to the top RGB image, the image size decreases layer by layer to form a multi-resolution hierarchy;

[0064] (3a2) Smooth the RGB image with a 5×5 Gaussian blur kernel for each layer to eliminate high-frequency information and avoid aliasing problems during downsampling;

[0065] (3a3) After smoothing, perform subsampling on the image at every other point to achieve a halving of the image size reduction;

[0066] Assume that the original image is the image G0 at the 0th layer, then the image G i at the ith layer is generated by the following formula:

[0067] G i = Downsample(G i-1 * Gaussian Kernel)

[0068] Among them, * represents convolution, Gaussian Kernel is a 5×5 Gaussian blur kernel, and Downsample is downsampling.

[0069] Another object of the present invention is to provide an electronic device, including:

[0070] a processor; and

[0071] a memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability as described above.

[0072] The present invention also provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability as described above.

[0073] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention provides semantic boxes by using the YOLOv7 model that can be deployed on a mobile device, fuses instance segmentation and geometric methods through variance weighting to obtain the motion probability of Gaussian points, and accordingly classifies the Gaussian points as dynamic or static, thereby improving the pose estimation accuracy and reducing artifacts in the scene map; Second, the present invention uses the motion state verification based on epipolar geometry to densify the map, and restores the static Gaussian points and corresponding feature points that are mislabeled as dynamic, realizing the simultaneous optimization of tracking and rendering information and avoiding redundant decoupling processing; Third, by using the motion probability to adjust the photometric and depth rendering losses, and introducing a penalty term based on the motion probability to limit the influence of dynamic objects on rendering, and at the same time introducing an edge transformation loss to improve geometric consistency and reduce noise in the dynamic environment, effective artifact elimination can be achieved only by a rough estimated mask, and it does not depend on the accuracy of the mask, greatly enhancing the adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 is a flowchart of the method of the present invention;

[0075] Figure 2 is the positioning effect of a certain sequence in the dynamic real-world dataset of the present invention;

[0076] Figure 3 is the effect of restoring the static three-dimensional environment of a certain sequence in the dynamic real-world dataset of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0077] As Figure 1 shown, a three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability, the method includes the following steps in sequence:

[0078] (1) Obtain the motion probability of 3D Gaussian points: Pass multiple input RGB images through the object detection network to obtain the semantic motion probability of the feature points in the image frame. Obtain the geometric motion probability of the feature points through multi-view geometric constraints. Fuse the semantic motion probability and the geometric motion probability through variance weighting to obtain the observed preliminary motion probability of the Gaussian points;

[0079] (2) Obtain the dynamic and static attributes of the Gaussian points based on the preliminary motion probability of the Gaussian points: Based on the median of the preliminary motion probability of the Gaussian points in the current image frame, set an adaptive threshold. Gaussian points exceeding the adaptive threshold are marked as dynamic, otherwise they are marked as static;

[0080] (3) Reverse map the marked Gaussian points to the front-end feature tracking system through the Gaussian pyramid network. To prevent errors or omissions in marking dynamic Gaussian points during the marking process, after mapping, use the epipolar constraint based on feature points for further judgment, and finally densely map back to the Gaussian points to obtain the final motion probability of the Gaussian points;

[0081] (4) Construct a loss based on motion probability and a loss based on edge distortion based on the final motion probability of the Gaussian points to obtain the overall rendering loss function;

[0082] (5) Optimize the parameters in the overall rendering loss function and the final motion probability of the Gaussian points. Optimize the following parameters: p i 、C i 、camera pose T CW .

[0083] Step (1) specifically includes the following steps:

[0084] (1a) Use an RGB-D camera to obtain the frame stream of the scene. The RGB-D camera is fixed on the mobile device. The mobile device uses a mobile robot or a drone. Load the YOLOv7 model for detecting semantic detection boxes on the mobile device, and assign the corresponding semantic motion probability MP i to the feature point p ins (p i ); for example, if the feature point belongs to the area where "person" is located, the corresponding semantic motion probability is 0.8, "chair" is 0.5, small objects on the table are 0.3, and if it is the background, it is 0.2.

[0085] (1b) The geometric constraint uses the reprojection error. Select the key frame with the largest number of co-visible views from the historical key frames. The historical key frames are RGB images saved at past time points. Calculate the reprojection error between the current frame Gaussian point and its corresponding feature point p i :

[0086]

[0087] Among them, π is the projection function, R is the rotation matrix, t is the translation vector, and both R and t are the camera pose T CW ; Inside the same instance region, calculate the feature points variance σ of the reprojection error e and the maximum error e max , and calculate the geometric motion probability MP i of the feature points p geo (p i ):

[0088]

[0089] where σ is the predefined variance threshold; n represents the number of feature points; Since the projection relationship of pixels is fully considered, the obtained geometric motion probability can better compensate for the misrecognition of semantic information.

[0090] For the points located in the external region of the semantic detection box detected in step (1a):

[0091]

[0092] (1c) Use the variance weighted fusion strategy to combine the obtained geometric motion probability with the semantic motion probability to obtain the motion probability N i of the feature points p p :

[0093] M p = MP ins + H(MP geo - MP ins )

[0094] where H is used to measure the contributions of MP ins and MP geo to M p , and the formula is:[[]]

[0095]

[0096] In the formula,[[]] is the confidence of instance segmentation,[[]] is the variance of the reprojection error; The motion probability of the initial Gaussian points depends on the information of the corresponding feature points observed in each image frame, so is expressed as a function depending on the motion probability M p of the observed feature points:[[]]

[0097]

[0098] Among them, m represents the number of observed feature points, is the Gaussian point The motion probability of the observed feature points.

[0099] Step (2) specifically refers to: the motion probability of the preliminary Gaussian points form a set Calculate the median of

[0100] The adaptive threshold τ is:

[0101]

[0102] When , mark the Gaussian point as dynamic, otherwise as static.

[0103] Step (3) specifically includes the following steps in sequence:

[0104] (3a) Construct a Gaussian pyramid network:

[0105] (3b) Reverse-map the marked Gaussian points through the constructed Gaussian pyramid network to the front-end feature tracking system, which is deployed on the mobile device, to obtain the feature points in the image frame corresponding to the Gaussian points, and further screen out the missing dynamic Gaussian points through the epipolar error of the reverse-mapped feature points, and at the same time recover the static Gaussian points that are mislabeled as dynamic;

[0106] Assume that p1 and p2 are the matching feature points of the Gaussian points in the previous image frame and the current image frame respectively, and L1 and L2 are the epipolar lines in the corresponding frames. Calculate L1 and L2 according to the fundamental matrix F:

[0107]

[0108]

[0109] Among them, A1, B1, C1, A2, B2, C2 are all the coefficients of the two-dimensional linear equation, K is the internal parameter matrix of the camera, t ∧ is the skew-symmetric matrix of the translation vector, (u, v) is the pixel coordinates of the feature point, and R is the rotation matrix;

[0110] The distance d i between the feature point p i and the corresponding epipolar line is calculated by the following formula:

[0111]

[0112] Since dynamic points are not completely removed in the calculation of the fundamental matrix, the fundamental matrix may be inaccurate. To obtain a more accurate fundamental matrix, before the second fundamental matrix calculation, the distance from the matching points to the epipolar lines is used to exclude some points that do not meet the constraints.

[0113] (3c) If d1 + d2 < ε, then the feature point p i will be densified, and the feature point p i is remapped back to the Gaussian point through the Gaussian function:

[0114]

[0115] where ∑ is the covariance matrix, o ∈ [0, 1] represents the opacity value, S is the scaling matrix, R is the rotation matrix; d1 is the distance from the feature point p1 to the epipolar line L1, d2 is the distance from the feature point p2 to the epipolar line L2; ε is the empirical threshold, with a value of 0.6;

[0116] Finally, for the Gaussian points remapped back, if d i is less than the empirical threshold ε, then the remapped Gaussian point is marked as static, and if d i is greater than the empirical threshold ε, then the remapped Gaussian point is marked as dynamic.

[0117] Step (4) specifically includes the following steps:

[0118] (4a) Construct the loss based on the motion probability: Use the photometric loss to measure the color error between the rendered image and the real input image, and combine the motion probability of the final Gaussian points to obtain the adjusted photometric loss L pho :

[0119]

[0120] where represents photometric rendering of the Gaussian points using the camera pose T CW for and represents the real color map corresponding to the given pose, is the motion probability of the final Gaussian points; the camera pose T CW adopts the transformation matrix;

[0121] Use the depth loss to constrain the depth of the 3D Gaussian points to obtain the adjusted depth loss L depth which is:

[0122]

[0123] where represents the depth rasterization process, Represents the true depth map corresponding to a given pose;

[0124] According to the position information of the Gaussian points themselves, a penalty term based on motion probability is introduced to constrain the influence of the true dynamic Gaussian on rendering, and the loss L based on motion probability is obtained MP :

[0125]

[0126] Among them, Represents the true three-dimensional position of the Gaussian point, Represents the estimated position;

[0127] (4b) An edge distortion loss is introduced to enhance the geometric consistency of data association between adjacent frames. That is, for the feature point p in the image frame i i , it is reprojected onto the frame j through a warping operation:

[0128]

[0129] Among them, D and T ji Represent the depth information and the transformation matrix respectively, Is the homogeneous coordinate of p i ; for an edge set ε i , the loss L based on edge distortion edge The calculation formula is as follows:

[0130]

[0131] Among them, ρ is a robust weight function used to reduce the influence of abnormal residuals;

[0132] (4c) Different weights λ i Are used to adjust the importance of motion probability in the camera tracking and mapping process, so as to obtain the overall rendering loss function L G :

[0133] L G = λ1·L pho + λ2·L depth + λ3·L MP + λ4·L edge

[0134] Among them, λ1 = 0.9, λ2 = 0.1, λ3 = 500, λ4 = 300.

[0135] Step (3a) specifically includes the following steps in sequence:

[0136] (3a1) Form a pyramid structure: from the bottom RGB image to the top RGB image, the image size decreases layer by layer, forming a multi-resolution hierarchy;

[0137] (3a2) Smooth the RGB image with a 5×5 Gaussian blur kernel for each layer to eliminate high-frequency information and avoid aliasing problems during downsampling;

[0138] (3a3) After smoothing, perform point sampling on the image at intervals to halve the image size.

[0139] Assume that the original image is the image G0 of the 0th layer, then the generation formula of the image Gi of the ith layer is: i as follows:

[0140] Gi i = Downsample(Gi−1 * Gaussian Kernel) i-1 where * represents convolution, Gaussian Kernel is a 5×5 Gaussian blur kernel, and Downsample is downsampling.

[0141] As shown in, this dynamic real-world dataset is described as: two people walking through an office scene. The RGB-D camera was manually moved in three directions (x, y, z) while maintaining the same orientation. The movement trajectory of the camera was accurately estimated by the present invention. The red area (difference) represents the difference between the trajectory estimated (estimated) by the present invention and the ground truth trajectory. As shown in, the RGB image randomly captured in the dynamic scene is in front of the arrow, and the RGB image after being processed by the present invention is behind the arrow. It can be seen that the present invention has successfully filtered out the dynamic objects in the scene and restored the complete static environment.

[0142] As Figure 2 shown, this dynamic real-world dataset is described as: two people walking through an office scene. The RGB-D camera was manually moved in three directions (x, y, z) while maintaining the same orientation. The movement trajectory of the camera was accurately estimated by the present invention. The red area (difference) represents the difference between the trajectory estimated (estimated) by the present invention and the ground truth trajectory. As shown in, the RGB image randomly captured in the dynamic scene is in front of the arrow, and the RGB image after being processed by the present invention is behind the arrow. It can be seen that the present invention has successfully filtered out the dynamic objects in the scene and restored the complete static environment. Figure 3 As shown in, the RGB image randomly captured in the dynamic scene is in front of the arrow, and the RGB image after being processed by the present invention is behind the arrow. It can be seen that the present invention has successfully filtered out the dynamic objects in the scene and restored the complete static environment.

[0143] In summary, the present invention combines dynamic-static segmentation based on the motion probability of three-dimensional Gaussian points and a state verification mechanism based on epipolar geometry, and optimizes the Gaussian rendering module accordingly. Compared with the existing methods, the present invention achieves optimal tracking accuracy and more realistic scene rendering effects on the premise of ensuring deployability with low computing power on mobile devices, significantly improves the accuracy of pose estimation, and can reconstruct the static environment completely and clearly.

[0144] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A three-dimensional Gaussian environmental perception and positioning method based on motion probability for processing dynamic interference, characterized in that: The method includes the following steps in sequence: (1) Obtain the motion probability of three-dimensional Gaussian points: Pass the input multiple RGB images through the target detection network to obtain the semantic motion probability of the feature points in the image frame, obtain the geometric motion probability of the feature points through multi-view geometric constraints, and fuse the semantic motion probability and the geometric motion probability through variance weighting to obtain the observed preliminary motion probability of the Gaussian points; (2) Obtain the dynamic and static attributes of the Gaussian points according to the preliminary motion probability of the Gaussian points: Based on the median of the preliminary motion probability of the Gaussian points in the current image frame, set an adaptive threshold, and mark the Gaussian points exceeding the adaptive threshold as dynamic, otherwise mark them as static; (3) Inversely map the marked Gaussian points to the front-end feature tracking system through the Gaussian pyramid network, use the epipolar constraint based on the feature points for further judgment, and finally densely map back to the Gaussian points to obtain the final motion probability of the Gaussian points; (4) Construct a loss based on the motion probability and a loss based on edge distortion based on the final motion probability of the Gaussian points to obtain the overall rendering loss function; (5) Optimize the parameters in the overall rendering loss function and the final motion probability of the Gaussian points.

2. The three-dimensional Gaussian environment perception and positioning method based on motion probability for processing dynamic interference according to claim 1, wherein: Step (1) specifically includes the following steps: (1a) Obtain the frame stream of the scene using an RGB-D camera, which is fixed on a mobile device. The mobile device uses a mobile robot or a drone. Load the YOLOv7 model for detecting semantic detection boxes on the mobile device, and assign the corresponding semantic motion probability MP for the feature point p in the image frame i to ins (p i ); (1b) The geometric constraint uses the reprojection error, selects the key frame with the largest co-visibility amount from historical key frames, where the historical key frames are RGB images saved at past time points, and calculates the Gaussian points of the current frame and its corresponding feature point p i of the reprojection error: where π is the projection function, R is the rotation matrix, t is the translation vector, and both R and t are the camera pose T CW ; within the same instance region, calculate the feature points variance σ of the reprojection error e and the maximum error e max , and calculate the geometric motion probability MP i of the feature point p geo (p i ): where σ is a predefined variance threshold; n represents the number of feature points; For the points located in the external region of the semantic detection box detected in step (1a): (1c) Combine the obtained geometric motion probability and semantic motion probability using the variance-weighted fusion strategy to obtain the motion probability \(M\) of feature point \(p\). i p : M p = MP ins + H(MP geo - MP ins ) Among them, H is used to measure MP ins and MP geo 's contribution to M p The formula is: wherein, is the confidence of instance segmentation, is the variance of the reprojection error; the motion probability of the initial Gaussian points depends on the information of the corresponding feature points observed in each image frame, so is expressed as the motion probability M that depends on the observed feature points p function of: where m represents the number of observed feature points, is a Gaussian point the motion probability of the observed feature points.

3. The three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability according to claim 1, wherein: Step (2) specifically refers to: the motion probability of the initial Gaussian points to form a set to calculate the median of The adaptive threshold τ is: When the marked Gauss points are dynamic; otherwise, they are static.

4. The three-dimensional Gaussian environment perception and positioning method based on motion probability for processing dynamic interference according to claim 1, characterized in that: Step (3) specifically includes the following steps in sequence: (3a) Construct a Gaussian pyramid network: (3b) Inversely map the marked Gaussian points to the front-end feature tracking system through the constructed Gaussian pyramid network. The front-end feature tracking system is deployed on the mobile device to obtain the feature points in the image frame corresponding to the Gaussian points, and further screen out the missing dynamic Gaussian points through the epipolar error of the inversely mapped feature points, and at the same time recover the static Gaussian points mislabeled as dynamic; Suppose p1 and p2 are the matching feature points of the Gaussian points in the previous image frame and the current image frame respectively , and L1 and L2 are the epipolar lines in the corresponding frames respectively. L1 and L2 are calculated according to the fundamental matrix F: where A1, B1, C1, A2, B2, C2 are all coefficients of two-dimensional linear equations, K is the internal parameter matrix of the camera, t ∧ is the skew-symmetric matrix of the translation vector, (u, v) are the pixel coordinates of the feature point, and R is the rotation matrix; Feature point p i and the distance d between the corresponding epipolar line i is calculated by the following formula: (3c) If d1 + d2 < ε, then the feature point p i will be densified, and the feature point p i is remapped back to the Gaussian point through the Gaussian function: where ∑ is the covariance matrix, o ∈ [0,1] represents the opacity value, S is the scaling matrix, R is the rotation matrix; d1 is the distance from the feature point p1 to the epipolar line L1, and d2 is the distance from the feature point p2 to the epipolar line L2; ε is an empirical threshold with a value of 0.6; Finally, for the Gauss points remapped back, if d i is less than the empirical threshold ε, then the remapped Gauss point is marked as static, and if d i is greater than the empirical threshold ε, then the remapped Gauss point is marked as dynamic.

5. The three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability according to claim 1, characterized in that: Step (4) specifically includes the following steps: (4a) Construct the loss based on motion probability: Use the photometric loss to measure the color error between the rendered image and the real input image, and combine it with the motion probability of the final Gaussian points to obtain the adjusted photometric loss L pho : Among them, represents using the camera pose T CW to perform photometric rendering on the Gaussian points and represents the true color map corresponding to the given pose, is the motion probability of the final Gaussian points; the camera pose T CW adopts a transformation matrix. The depth of the 3D Gaussian points is constrained by the depth loss to obtain the adjusted depth loss L depth as follows: Among them, represents the depth rasterization process, represents the true depth map corresponding to the given pose; According to the position information of the Gaussian points themselves, a penalty term based on the motion probability is introduced to constrain the influence of the truly dynamic Gaussian on rendering, and the loss L based on the motion probability is obtained. MP : Among them, represents the true three-dimensional position of the Gauss point, represents the estimated position; (4b) Introduce an edge distortion loss to enhance the geometric consistency of data association between adjacent frames, that is, for the feature point p in image frame i i , re-project it onto frame j through a warping operation: Among them, D and T ji represent depth information and transformation matrix respectively, is the homogeneous coordinate of p i ; for an edge set ε i , the calculation formula of the loss L edge based on edge distortion is as follows: where ρ is a robust weight function used to reduce the influence of abnormal residuals; (4c) Use different weights λ i to adjust the importance of the motion probability during camera tracking and mapping, thereby obtaining the overall rendering loss function L G : L G = λ1·L pho + λ2·L depth + λ3·L MP + λ4·L edge where λ1 = 0.9, λ2 = 0.1, λ3 = 500, λ4 = 300.

6. The three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability according to claim 4, wherein: Step (3a) specifically includes the following steps in sequence: (3a1) Form a pyramid structure: From the bottom RGB image to the top RGB image, the image size decreases layer by layer to form a multi-resolution hierarchy; (3a2) Smooth the RGB image with a 5×5 Gaussian blur kernel for each layer to eliminate high-frequency information and avoid aliasing problems during downsampling; (3a3) After smoothing, perform subsampling on the image at every other point to reduce the image size by half; Assume that the original image is the image G0 of the 0th layer, then the generation formula of the image Gi of the ith layer is: i as follows: G i = Downsample(G i-1 * Gaussian Kernel) where * represents convolution, Gaussian Kernel is a 5×5 Gaussian blur kernel, and Downsample is downsampling.

7. An electronic device, including: A processor; And A memory in which computer program instructions are stored, and when the computer program instructions are run by the processor, the processor is caused to execute the three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability according to any one of claims 1-6.

8. A computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the three-dimensional Gaussian environment perception and positioning method for processing dynamic interference based on motion probability according to any one of claims 1-6.