An occlusion-oriented geometric constraint generated visual slam method
Patent Information
- Application Number
- CN202610871177.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]本发明的目的在于:提出一种面向遮挡的几何约束生成式视觉SLAM方法,解决现有技术在遭遇大面积动态遮挡时特征点匮乏导致跟踪丢失,以及传统图像修复技术破坏SLAM几何约束的缺陷
1.本发明提出的EnhancedGAN时空预测方案,通过光流扭曲、生成补全以及双注意力帧间匹配的核心技术特征,突破了现有被动剔除动态特征的技术瓶颈,在大面积持续遮挡场景中,可主动生成稳定特征维持系统跟踪连续性,大幅降低跟踪丢失概率。
Smart Images

Figure CN122799263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot navigation, and more particularly to an occlusion-oriented geometric constraint generative visual SLAM method. Background Technology
[0002] Simultaneous Localization and Mapping (SLAM) is a core underlying technology in mobile robotics, autonomous driving, AR / VR, and other fields. Classic visual SLAM systems such as ORB-SLAM3 exhibit extremely high localization accuracy in static, structured environments. However, their core algorithms heavily rely on static rigid body assumptions and multi-view geometric epipolar constraints. In real-world complex scenarios, when the field of view is obstructed by dynamic objects such as nearby buses or dense crowds over large areas for extended periods, the system faces severe challenges. Summary of the Invention
[0003] The purpose of this invention is to propose an occlusion-oriented geometric constraint generative visual SLAM method to solve the problems of existing technologies, such as the lack of feature points leading to tracking loss when encountering large-area dynamic occlusion, and the destruction of SLAM geometric constraints by traditional image inpainting techniques.
[0004] Specifically, this invention provides an occlusion-oriented geometric constraint-based generative visual SLAM method, device, and medium, the method comprising the following steps: S1. Obtain the input image of the current frame of the visual SLAM system, extract image feature points and perform feature matching, and count the number of valid matching feature points; if the number of valid matching feature points is not less than a preset judgment threshold, input the feature points into the visual SLAM system to perform conventional tracking and pose calculation; if the number of valid matching feature points is less than the preset judgment threshold, determine that there is severe occlusion in the current scene, trigger the Enhanced Generative Adversarial Network (EnhancedGAN), and execute step S2. S2. The occluded input image is fed into the semantic segmentation network and decoupled into a dynamic foreground region and a static background region. The generator of the EnhancedGAN is used to perform texture restoration and spatiotemporal prediction on the dynamic foreground region in combination with historical reference frames to generate a prediction frame image. S3. Input the predicted frame image and the historical benchmark reference frame into the feature generation network G3, extract the inter-frame correlation features through the self-attention mechanism and the cross-attention mechanism, and output the matching feature point pairs; wherein, the feature generation network G3 is obtained through adversarial training with the introduction of multi-view epipolar geometric constraint loss during the training phase, and the matching feature point pairs satisfy the geometric consistency of rigid body transformation. S4. The geometrically consistent matching feature point pairs output in step S3 are used as auxiliary observations and input into the tracking thread of the visual SLAM system to estimate the pose of the current frame. When the tracking is recovered and the key frame insertion condition is met, the verified feature observations are sent to the local mapping thread to perform local bundle adjustment optimization and loop closure detection, and output the global pose trajectory and map of the camera.
[0005] A storage medium storing instructions and data for implementing an occlusion-oriented geometric constraint generative visual SLAM method.
[0006] An occlusion-oriented geometric constraint generative visual SLAM device includes: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement an occlusion-oriented geometric constraint generative visual SLAM method.
[0007] The beneficial effects provided by this invention are: 1. The EnhancedGAN spatiotemporal prediction scheme proposed in this invention breaks through the technical bottleneck of passively removing dynamic features by using the core technical features of optical flow distortion, generation completion and dual attention inter-frame matching. In large-area continuous occlusion scenes, it can actively generate stable features to maintain the continuity of system tracking and greatly reduce the probability of tracking loss.
[0008] 2. Overcoming the limitation of GAN networks generating visually realistic but geometrically divergent images. This is achieved by transforming the fundamental solution basis of SLAM—the basis matrix and epipolar distance—into a penalty function of the adversarial network during the training phase. This ensures that the generated feature points have fundamental geometric consistency, significantly reducing absolute trajectory error.
[0009] 3. This invention achieves a lightweight design of semantic segmentation networks through a knowledge distillation strategy. While retaining 96.9% of the model performance, it reduces the number of model parameters by about 95%, achieving high real-time inference speed. It resolves the contradiction between accuracy and real-time performance in existing solutions and can be adapted to the deployment requirements of edge mobile robots. Attached Figure Description
[0010] Figure 1 This is a simplified flowchart of the method of the present invention; Figure 2 This is a diagram of the EnhancedGAN generator network architecture of the present invention; Figure 3 This is a diagram of the self-attention and cross-attention module architecture based on the baseline frame update in G3 of the present invention; Figure 4 This is a schematic diagram of the knowledge distillation process of the present invention; Figure 5This is a schematic diagram of the hardware device operation according to an embodiment of the present invention. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0012] Before formally describing the present invention, a general description of the solution of the present invention will be given first to facilitate understanding.
[0013] Example 1 Please refer to Figure 1 The present invention provides an occlusion-oriented geometric constraint generative visual SLAM method, comprising: S1. Obtain the input image of the current frame of the visual SLAM system, extract image feature points and perform feature matching, and count the number of valid matching feature points; if the number of valid matching feature points is not less than a preset judgment threshold, input the feature points into the visual SLAM system to perform conventional tracking and pose calculation; if the number of valid matching feature points is less than the preset judgment threshold, determine that there is severe occlusion in the current scene, trigger the Enhanced Generative Adversarial Network (EnhancedGAN), and execute step S2. It should be noted that the preset judgment threshold in step S1 is the minimum empirical value of the number of valid matching feature points, which is used to distinguish between normal tracking state and severe occlusion state. Specifically, this invention performs ORB feature point extraction on the input image, sets the threshold for the number of feature points extracted to 60, performs brute-force matching and RANSAC to remove false matches, and counts the number of valid matching feature points N. (1) If N≥60, the tracking status is determined to be normal. The effective feature points are input into the ORB-SLAM3 system to perform regular tracking thread, local BA optimization and loop closure detection. (2) If N<60, it is determined that there is a risk of occlusion or weak texture, triggering the EnhancedGAN module and executing step S2.
[0014] S2. The occluded input image is fed into the semantic segmentation network and decoupled into a dynamic foreground region and a static background region. The generator of the EnhancedGAN is used to perform texture restoration and spatiotemporal prediction on the dynamic foreground region in combination with historical reference frames to generate a prediction frame image. Please refer to Figure 2 and Figure 4 It should be noted that step S2 specifically includes: S21. Input the input image into a lightweight student semantic segmentation network obtained by compression through the knowledge distillation algorithm, output fine-grained semantic labels, and decouple them into dynamic foreground masks and static background masks. S22. The generator of the EnhancedGAN includes an optical flow prediction network G1, an image generation network G2, and a mask prediction network M, which collaboratively generate the predicted image of the t-th frame according to the following formula. : in, The spatially gated soft mask is the output of the mask prediction network M. The optical flow distortion image output by the optical flow prediction network G1. , Generate content for the foreground and background output by the G2 image generation network. is the static background area mask, and ⊙ represents element-wise multiplication; S23. Repeat step S22 until the number of generated prediction frame images reaches the preset maximum number of consecutively generated frames σ.
[0015] It should be noted that the maximum number of consecutively generated frames σ in step S2 is a preset hyperparameter used to control the duration of generative prediction and prevent error accumulation.
[0016] Specifically, the image that triggered EnhancedGAN was input into a lightweight segmentation network, which used PSPNet-R101 as the teacher model and was trained using the CWD+KD hybrid knowledge distillation algorithm. Through the semantic aggregation mapping function, the 22 fine-grained labels output by the network were decoupled into dynamic foreground and static background to generate a binary mask image, where the dynamic foreground region is the occluded region to be repaired and the static background region is the effective geometric constraint region.
[0017] The generator of EnhancedGAN consists of an optical flow prediction network G1, an image generation network G2, a mask prediction network M, and a feature generation network G3. It takes L=2 consecutive historical keyframes and their corresponding semantic segmentation mask frames as input and generates predicted frame images with repaired occluded areas frame by frame. Specifically, the optical flow prediction network G1 performs pixel-level optical flow distortion propagation on static background regions based on temporal motion inertia; the image generation network G2 performs texture generation and background completion on occluded and missing regions; and the mask prediction network M outputs a spatially gated soft mask. This achieves a smooth fusion of optical flow-distorted content and generated content, ultimately resulting in the predicted image for frame t. The generating expression is:
[0018] In the formula, The optical flow distortion image output by the optical flow prediction network G1. , Generate content for the foreground and background output by the G2 image generation network. ⊙ represents the binary mask for the static background region, and ⊙ represents the element-wise multiplication operation. The value range is [0,1].
[0019] S3. Input the predicted frame image and the historical benchmark reference frame into the feature generation network G3, extract the inter-frame correlation features through the self-attention mechanism and the cross-attention mechanism, and output the matching feature point pairs; wherein, the feature generation network G3 is obtained through adversarial training with the introduction of multi-view epipolar geometric constraint loss during the training phase, and the matching feature point pairs satisfy the geometric consistency of rigid body transformation. It should be noted that the feature generation network G3 in step S3 does not directly generate a complete image. Instead, it performs feature extraction, attention association, and matching confidence evaluation on the input predicted frame image and the historical benchmark reference frame, and outputs image pairs labeled with matching feature point pairs.
[0020] Step S3 further includes a similarity-based dynamic update mechanism for the baseline frame: S31. Set historical baseline reference frame The historical reference frame and the current prediction frame are divided into blocks, and feature maps are extracted through a local convolutional neural network. S32. Extract intra-frame features within each frame through a self-attention mechanism, extract inter-frame matching features through a cross-attention mechanism, and calculate the inter-frame region similarity probability matrix. S33. A preset similarity threshold is set. If the calculated average similarity probability is not lower than the similarity threshold, the current historical reference frame is used to match subsequent frames. If it is lower than the similarity threshold, the current predicted frame is updated to a new reference frame, the continuous generation count is reset, and subsequent frames are generated until the preset maximum number of consecutively generated frames is reached.
[0021] Please refer to Figure 3 Specifically, step S3 uses a feature generation network G3, which is an improvement on the lightweight LoFTR network, to complete feature extraction and matching between the reference frame and the predicted frame, as well as frame similarity judgment, and finally outputs image pairs with accurate matching. The specific implementation process is as follows: 1. Reference Frame Initialization: Set a fixed reference frame. The initial reference frame is determined according to the following rules: if the EnhancedGAN module is triggered for the first time in the current sequence, and GAN generation was not triggered in the previous frame, then the reference frame is... The original image is tracked normally in the previous frame. If GAN generation has been triggered in the previous frame, then the reference frame... Generate a valid frame for the last frame in the previous sequence. Once the reference frame is initialized, it remains fixed during continuous matching and does not switch as a single-frame prediction frame is generated.
[0022] 2. Dual-frame feature extraction and matching: Using a fixed reference frame... Compared with the currently generated predicted image of frame t The synchronous input feature generation network G3 is used. G3 is an improvement on the lightweight LoFTR architecture. First, a CNN backbone network with shared weights is used to extract multi-scale features from two frames of images. Then, a self-attention and cross-attention fusion module is used to complete feature matching. The self-attention module extracts intra-frame long-range dependency features from the reference frame and the prediction frame, respectively. The cross-attention module completes the feature association and matching pair solution between the two frames, resulting in... The expression is:
[0023] Final output The baseline and prediction frame image pairs are used to complete feature matching and annotation.
[0024] 3. Inter-frame similarity judgment and baseline frame update: Based on the feature matching results output by the G3 network attention module, an inter-frame similarity metric matrix is constructed, and the average matching confidence between two frames is calculated as the similarity score. A preset similarity threshold q=0.6 is used. During continuous generation, if the similarity score is greater than or equal to the threshold q, it indicates that the scene has not changed significantly, and the current baseline frame still has effective geometric constraints. The baseline frame is kept fixed, and the next frame prediction image is generated. Conversely, if the similarity score is less than or equal to the threshold q, it indicates that the scene has changed significantly due to camera movement, and the original baseline frame has lost its geometric constraint effectiveness. The latest effective prediction frame is updated as the new fixed baseline frame, the continuous generation count is reset, and the frame generation process continues.
[0025] 4. Termination Condition: The maximum number of consecutively generated frames is preset to σ=3. When the number of consecutively generated valid prediction frames reaches the upper limit of σ, the EnhancedGAN frame generation process is terminated, and the stable feature matching pairs output by the G3 network are sent to the ORB-SLAM3 backend for pose calculation. The completed features in the generated frames are only used as short-term tracking auxiliary features for pose stability estimation of the current frame; only when the feature is observed again in subsequent real observation frames and passes the reprojection error verification is it allowed to be converted into long-term map points. Generated features that fail verification will not participate in the permanent update of map points.
[0026] It should be noted that the feature generation network G3 adopts an adversarial training approach during the training phase. The EnhancedGAN includes a generator and a discriminator, and its composite objective loss function during the training phase... Represented as: in, This is the total adversarial loss, used to improve the quality of generated images and inter-frame continuity; Knowledge distillation loss is used to compress network models; This is the epipolar constraint loss, used to penalize generated feature points that do not conform to epipolar geometry; , , This represents the balancing weights for each type of loss.
[0027] The polar constraint loss The calculation formula is: in, The number of matching feature point pairs generated. The coordinates of the actual feature points in the historical reference frame. The generator predicts the coordinates of the corresponding generated feature points in the frame. The fundamental matrix is calculated based on the true pose of historical keyframes; the loss function forces the generated feature point pairs to satisfy... Multi-view geometric constraints.
[0028] Specifically, in traditional generative adversarial networks (GANs), the training objective focuses only on pixel-level differences in images. However, this method, during the training phase, forcibly incorporates the fundamental matrix F of the SLAM underlying solution into the gradient backpropagation network. The composite objective loss function is defined as:
[0029] Among them, total combat losses and knowledge distillation loss This is used to ensure the quality of a single frame image and to keep the model lightweight. This is a penalty term for polar geometry constraints.
[0030] The specific computational logic of epipolar geometric constraint loss is as follows: Assume that there are true static feature points in the historical reference frame. In the real three-dimensional physical world, the projection point of this point in the current prediction frame. They must fall precisely on the corresponding epipolar line. According to the principles of multi-view geometry, matching point pairs must satisfy... .
[0031] The epipolar geometry constraint loss function of this invention is designed as follows:
[0032] During training, when the generator predicts feature points Then, the system calculates the fundamental matrix F using the known true pose and obtains... The value of is the squared distance from the predicted feature point to the theoretical epipolar line. If the generative network only generates visually realistic pixels but with positional deviations, the epipolar line distance corresponding to that pixel is... The value will obviously not be 0, so the loss function... This will increase dramatically, imposing a huge gradient penalty on the generator.
[0033] S4. The geometrically consistent matching feature point pairs output in step S3 are used as auxiliary observations and input into the tracking thread of the visual SLAM system to estimate the pose of the current frame. When the tracking is recovered and the key frame insertion condition is met, the verified feature observations are sent to the local mapping thread to perform local bundle adjustment optimization and loop closure detection, and output the global pose trajectory and map of the camera.
[0034] Specifically, the stable matching feature points output in step S3 are input into the backend thread of the ORB-SLAM3 system to perform Local Bundle Adjustment (Local BA) to optimize the pose and map points, loop closure detection is performed using the bag-of-words model (DBoW2), global graph optimization is performed to eliminate cumulative drift, and finally the global pose trajectory and sparse feature map of the camera are output.
[0035] Example 2: This embodiment is built based on the ORB-SLAM3 monocular inertial navigation fusion mode, and the hardware platform is AMD R7-6800H CPU, 16GB RAM, and NVIDIA RTX 3090 GPU; the software environment is Ubuntu 20.04, ROS Noetic, Python 3.8, and PyTorch 1.10. The experimental dataset uses the EuRoC MAV public dataset. To verify the robustness of dynamic occlusion, this invention further constructs a dynamic occlusion enhancement test set based on EuRoC sequences, and simulates continuous occlusion scenarios by superimposing random dynamic occlusion masks.
[0036] In addition, this embodiment constructs a random occlusion degradation test set based on the EuRoC MAV dataset MH01-05 sequence. Specifically, random irregular polygons are generated in 30 frames of the original image sequence, with the occlusion area accounting for 30% of the image area. This is used to simulate situations such as field of view occlusion, expansion of weak texture regions, and insufficient effective feature points encountered during actual operation. All experiments are repeated 5 times, and the average result is taken as the final experimental data. When the system loses tracking for more than 10 consecutive frames in the test sequence, or the output trajectory length is less than 80% of the true trajectory length, or it cannot complete the timestamp alignment with the ground truth trajectory, the experiment is judged as Tracking Failed, and the ATE RMSE of the experiment is no longer calculated.
[0037] (1) Comparative experiment with existing mainstream visual SLAM systems The localization accuracy and robustness of this invention compared with ORB-SLAM3, Vins-mono, DynaSLAM, and RTAB-Map systems on the EuRoC MAV dataset are shown in Tables 1 and 2 below: Table 1. Evaluation of Absolute Trajectory Error (ATE)
[0038] Table 2 Evaluation of the number of successful tracking attempts
[0039] Experimental results show that in occluded scene sequences, the robustness and accuracy of the method of this invention are better than those of existing mainstream systems, while in simple structured scene sequences, the accuracy of the method of this invention is comparable to that of ORB-SLAM3, achieving a balance between positioning accuracy and real-time performance in complex scenes.
[0040] (2) Ablation experiment of core module For the EnhancedGAN module of this invention, an ablation experiment was conducted using the MH05 sequence superimposed in the above comparative experiment to verify the technical effect of each module. The results are shown in Table 3 below: Table 3. Baseline and performance evaluation of each module of the invention
[0041] Experimental results show that the EnhancedGAN module can effectively solve the feature loss problem in large-area occlusion scenarios and avoid tracking loss. After deep coupling of the module, the best balance between robustness and real-time performance is achieved in complex scenarios, verifying the effectiveness and inventiveness of the technical solution of this invention.
[0042] Example 3: Please see Figure 5 , Figure 5 This is a schematic diagram of the hardware device in operation according to an embodiment of the present invention. The hardware device specifically includes: an occlusion-oriented geometric constraint generative visual SLAM device 401, a processor 402, and a storage medium 403.
[0043] An occlusion-oriented geometric constraint generative visual SLAM device 401: The occlusion-oriented geometric constraint generative visual SLAM device 401 implements the occlusion-oriented geometric constraint generative visual SLAM method.
[0044] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the occlusion-oriented geometric constraint generative visual SLAM method.
[0045] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the occlusion-oriented geometric constraint generative visual SLAM method.
[0046] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An occlusion-oriented geometric constraint-based generative visual SLAM method, characterized in that: Includes the following steps: S1. Obtain the input image of the current frame of the visual SLAM system, extract image feature points and perform feature matching, and count the number of valid matching feature points; If the number of valid matching feature points is not less than the preset judgment threshold, the feature points are input into the visual SLAM system to perform regular tracking and pose calculation; if the number of valid matching feature points is less than the preset judgment threshold, it is determined that there is severe occlusion in the current scene, and the Enhanced Generative Adversarial Network (EnhancedGAN) is triggered to execute step S2. S2. The occluded input image is fed into the semantic segmentation network and decoupled into dynamic foreground region and static background region; The enhancedGAN generator, combined with historical reference frames, performs texture restoration and spatiotemporal prediction on the dynamic foreground region to generate a predicted frame image. S3. Input the predicted frame image and the historical benchmark reference frame into the feature generation network G3, extract the inter-frame correlation features through the self-attention mechanism and the cross-attention mechanism, and output the matching feature point pairs; wherein, the feature generation network G3 is obtained through adversarial training with the introduction of multi-view epipolar geometric constraint loss during the training phase, and the matching feature point pairs satisfy the geometric consistency of rigid body transformation. S4. The geometrically consistent matching feature point pairs output in step S3 are used as auxiliary observations and input into the tracking thread of the visual SLAM system to estimate the pose of the current frame. When the tracking is recovered and the key frame insertion condition is met, the verified feature observations are sent to the local mapping thread to perform local bundle adjustment optimization and loop closure detection, and output the global pose trajectory and map of the camera.
2. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 1, characterized in that, Step S2 specifically includes: S21. Input the input image into a lightweight student semantic segmentation network obtained by compression through the knowledge distillation algorithm, output fine-grained semantic labels, and decouple them into dynamic foreground masks and static background masks. S22. The generator of the EnhancedGAN includes an optical flow prediction network G1, an image generation network G2, and a mask prediction network M, which collaboratively generate the predicted image of the t-th frame according to the following formula. : in, The spatially gated soft mask is the output of the mask prediction network M. The optical flow distortion image output by the optical flow prediction network G1. , Generate content for the foreground and background output by the G2 image generation network. is the static background area mask, and ⊙ represents element-wise multiplication; S23. Repeat step S22 until the number of generated prediction frame images reaches the preset maximum number of consecutively generated frames σ.
3. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 1, characterized in that: In step S3, the feature generation network G3 does not directly generate a complete image. Instead, it performs feature extraction, attention association, and matching confidence evaluation on the input predicted frame image and the historical benchmark reference frame, and outputs image pairs labeled with matching feature point pairs.
4. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 3, characterized in that, Step S3 further includes a similarity-based dynamic update mechanism for the baseline frame: S31. Set historical baseline reference frame The historical reference frame and the current prediction frame are divided into blocks, and feature maps are extracted through a local convolutional neural network. S32. Extract intra-frame features within each frame through a self-attention mechanism, extract inter-frame matching features through a cross-attention mechanism, and calculate the inter-frame region similarity probability matrix. S33. A preset similarity threshold is set. If the calculated average similarity probability is not lower than the similarity threshold, the current historical reference frame is used to match subsequent frames. If it is lower than the similarity threshold, the current predicted frame is updated to a new reference frame, the continuous generation count is reset, and subsequent frames are generated until the preset maximum number of consecutively generated frames is reached.
5. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 1, characterized in that, The feature generation network G3 employs adversarial training during the training phase. The EnhancedGAN includes a generator and a discriminator, and its training phase uses a composite objective loss function. Represented as: in, This is the total adversarial loss, used to improve the quality of generated images and inter-frame continuity; Knowledge distillation loss is used to compress network models; This is the epipolar constraint loss, used to penalize generated feature points that do not conform to epipolar geometry; , , This represents the balancing weights for each type of loss.
6. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 5, characterized in that, The polar constraint loss The calculation formula is: in, The number of matching feature point pairs generated. The coordinates of the actual feature points in the historical reference frame. The generator predicts the coordinates of the corresponding generated feature points in the frame. The fundamental matrix is calculated based on the true pose of historical keyframes; the loss function forces the generated feature point pairs to satisfy... Multi-view geometric constraints.
7. The occlusion-oriented geometric constraint generative visual SLAM method as described in claim 1, characterized in that, In step S1, the preset judgment threshold is the minimum empirical value of the number of effective matching feature points, which is used to distinguish between normal tracking state and severe occlusion state; in step S2, the maximum number of consecutively generated frames σ is a preset hyperparameter, which is used to control the duration of generative prediction and prevent error accumulation.
8. A storage medium, characterized in that: The storage medium stores instructions and data for implementing the occlusion-oriented geometric constraint generative visual SLAM method according to any one of claims 1 to 7.
9. An occlusion-oriented geometric constraint generative visual SLAM device, characterized in that: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the occlusion-oriented geometric constraint generative visual SLAM method according to any one of claims 1 to 7.