Visual SLAM (Simultaneous Localization and Mapping) method, device and terminal in dynamic environment based on panoramic segmentation
By adopting a panoramic segmentation-based visual SLAM method in the SLAM system, the problem of unknown or unlabeled objects in the dynamic environment affecting SLAM accuracy is solved, and higher adaptability and accuracy are achieved.
Patent Information
- Application Number
- CN202411793763.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
AI Technical Summary
The accuracy of the SLAM system is significantly reduced when there are unknown or unlabeled dynamic objects.
The visual SLAM method in a dynamic environment based on panoramic segmentation is adopted. By receiving continuous frame images, filtering portraits, determining the pole lines corresponding to static points, obtaining matching points, and estimating SLAM state and updating maps are realized through static points.
It improves the adaptability and accuracy of SLAM system in dynamic environments, making it more practical and robust in complex practical scenarios.
Smart Images

Figure CN119941777A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of visual simultaneous positioning and map construction, and specifically to a visual SLAM method, device, and terminal in a dynamic environment based on panoramic segmentation. Background Art
[0002] Visual simultaneous localization and mapping (SLAM) is a core technology in fields such as autonomous robotics and unmanned driving, enabling robots to navigate and operate in unknown or unstructured environments.
[0003] In the real world, scenes often contain moving objects (such as pedestrians and vehicles), whose feature points affect the performance of SLAM systems. Most existing methods rely on deep learning models to detect and filter the feature points of dynamic objects.
[0004] However, the above methods usually require pre-training models, and the accuracy of the SLAM system will be significantly reduced in the presence of unknown or unlabeled dynamic objects. Summary of the Invention
[0005] The main purpose of this application is to provide a visual SLAM method, device, and terminal in a dynamic environment based on panoramic segmentation to solve the problem that the accuracy of the SLAM system will be significantly reduced when there are unknown or unmarked dynamic objects.
[0006] To achieve the above objectives, in a first aspect, the present application provides a visual SLAM method in a dynamic environment based on panoptic segmentation, comprising:
[0007] receiving a first image and a second image, wherein the first image and the second image are continuous frame images;
[0008] Filtering the people in the first image and the second image respectively to obtain a third image and a fourth image;
[0009] determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image;
[0010] Obtaining matching points, wherein the matching points are used to represent points corresponding to the object in the third image and the object in the fourth image;
[0011] If the matching point is a static point, the SLAM state estimation and map update are realized through the static point.
[0012] In one possible implementation, determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image includes:
[0013] extracting static points in the third image and static points in the fourth image respectively;
[0014] calculating a fundamental matrix between the third image and the fourth image based on the pixel coordinates of the static point in the third image and the pixel coordinates of the static point in the fourth image;
[0015] Based on a fundamental matrix between the third image and the fourth image, epipolar lines corresponding to the static points in the third image and epipolar lines corresponding to the static points in the fourth image are respectively calculated.
[0016] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, the epipolar lines corresponding to the static points in the third image are calculated using the following formulas:
[0017] L1=F*P1 j
[0018] Where F represents the basic matrix between the third image and the fourth image, P1 j represents a static point in the third image, and L1 represents an epipolar line corresponding to the static point in the third image.
[0019] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, a formula for calculating the epipolar line corresponding to the static point in the fourth image is as follows:
[0020] L2=F T *P2 j
[0021] Among them, F T represents the transpose of the fundamental matrix between the third and fourth images, P2 j represents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
[0022] In a possible implementation, obtaining matching points includes:
[0023] performing panorama separation on the third image and the fourth image respectively to obtain an object region and a background region of the third image and an object region and a background region of the fourth image;
[0024] extracting an object region in the third image and an object region in the fourth image;
[0025] Obtaining a first label and a first mask corresponding to the object in the object region in the third image, and a second label and a second mask corresponding to the object in the object region in the fourth image;
[0026] Based on the first mask and the second mask, detecting whether bounding boxes of the first label and the second label overlap, wherein the first label and the second label belong to the same label;
[0027] If there is overlap, get the matching points.
[0028] In a possible implementation, after obtaining the matching points, the following steps are further included:
[0029] Calculate the distance between the matching point and the corresponding epipolar line;
[0030] Based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it is determined whether the matching point is a static point.
[0031] In one possible implementation, the formula for calculating the distance between a matching point and its corresponding epipolar line is as follows:
[0032] Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image, F represents the basic matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
[0033] In one possible implementation, determining whether a matching point is a static point based on the distance between the matching point and the corresponding epipolar line and a preset threshold includes:
[0034] If the distance between the matching point and the corresponding epipolar line is less than the preset threshold, the matching point is a static point;
[0035] If the distance between the matching point and the corresponding epipolar line is greater than or equal to a preset threshold, the matching point is a dynamic point.
[0036] In a second aspect, an embodiment of the present invention provides a visual SLAM device in a dynamic environment based on panoramic segmentation, comprising:
[0037] A receiving module, configured to receive a first image and a second image, wherein the first image and the second image are continuous frame images;
[0038] a filtering module, configured to filter the persons in the first image and the second image respectively to obtain a third image and a fourth image;
[0039] an epipolar line calculation module, configured to determine, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image;
[0040] an acquisition module, configured to acquire matching points, wherein the matching points are used to represent points corresponding to the object in the third image and the object in the fourth image;
[0041] The visual SLAM implementation module is used to estimate the SLAM state and update the map through the static points if the matching points are static points.
[0042] In a third aspect, an embodiment of the present invention provides a terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above-mentioned visual SLAM methods in a dynamic environment based on panoramic segmentation are implemented.
[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above visual SLAM methods in a dynamic environment based on panoramic segmentation.
[0044] The embodiment of the present invention provides a visual SLAM method, device, and terminal in a dynamic environment based on panoramic segmentation, including: first receiving a first image and a second image, wherein the first image and the second image are continuous frame images, then filtering the people in the first image and the second image respectively to obtain a third image and a fourth image, and then based on the third image and the fourth image, determining the polar lines corresponding to the static points in the third image and the polar lines corresponding to the static points in the fourth image, and finally obtaining matching points, wherein the matching points are used to characterize the points corresponding to the objects in the third image and the objects in the fourth image, and if the matching points are static points, the static points are used to realize the estimation of the SLAM state and the update of the map. This application helps to improve the adaptability and accuracy of the SLAM system in a dynamic environment, making it more practical and robust in complex actual scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings that constitute part of this application are used to provide a further understanding of this application and make other features, objects and advantages of this application more apparent. The illustrative embodiment drawings of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0046] Figure 1 This is a flowchart of an implementation of a visual SLAM method in a dynamic environment based on panoptic segmentation provided by an embodiment of the present invention;
[0047] Figure 2 1 is a schematic structural diagram of a visual SLAM device in a dynamic environment based on panoramic segmentation provided by an embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0050] The terms "first," "second," "third," "fourth," and so forth (if any) in the description and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present invention described herein can be practiced in sequences other than those illustrated or described herein.
[0051] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0052] It should be understood that in the present invention, "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.
[0053] It should be understood that in the present invention, "multiple" refers to two or more. "And / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "Contains A, B and C", "Contains A, B, C" means that A, B, and C are all included, "Contains A, B or C" means that one of A, B, and C is included, and "Contains A, B and / or C" means that any one, any two, or any three of A, B, and C are included.
[0054] It should be understood that, in the present invention, "B corresponding to A," "B corresponding to A," "A corresponds to B," or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.
[0055] Depending on the context, "if" as used herein may be interpreted as "when" or "when" or "in response to determining" or "in response to detecting."
[0056] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0057] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below with reference to the accompanying drawings.
[0058] In one embodiment, Figure 1 As shown, a visual SLAM method in a dynamic environment based on panoramic segmentation is provided, comprising the following steps:
[0059] Step S101: receiving a first image and a second image.
[0060] The first image and the second image are continuous frame images.
[0061] The first image and the second image contain objects (Things), such as people, animals, vehicles, etc., and background (Stuff), such as the ground, sky, wall, etc.
[0062] The object contains both static and dynamic points, while the background contains only static points.
[0063] Step S102: filtering the people in the first image and the second image respectively to obtain a third image and a fourth image.
[0064] Since people are usually highly dynamic objects and rarely remain still for a long time, only people in the images are filtered for the first image and the second image.
[0065] Step S103: Based on the third image and the fourth image, determine the epipolar line corresponding to the static point in the third image and the epipolar line corresponding to the static point in the fourth image.
[0066] In order to determine the epipolar lines corresponding to static points in the third image and the epipolar lines corresponding to static points in the fourth image based on the third image and the fourth image, it is necessary to first extract the static points in the third image and the static points in the fourth image respectively, and then calculate the basic matrix between the third image and the fourth image based on the pixel coordinates of the static points in the third image and the pixel coordinates of the static points in the fourth image, and then calculate the epipolar lines corresponding to the static points in the third image and the epipolar lines corresponding to the static points in the fourth image respectively based on the basic matrix between the third image and the fourth image.
[0067] This method uses the RANSAC algorithm to calculate the fundamental matrix F between the third image and the fourth image as follows:
[0068] For two images of the same scene, such as a pair of matching points P1 in the third image and the fourth image j =[u1,v1,1] T and P2 j =[u2,v2,1] T
[0069] Among them, P1 j and P2 j Represent the pixel coordinates of the matching points of the third image and the fourth image respectively.
[0070] The fundamental matrix F between the third image and the fourth image can be calculated by the following constraints:
[0071] P2 j T *F*P1 j =0
[0072] The basic matrix F is a 3*3 matrix used to describe the geometric relationship between the third image and the fourth image.
[0073] After the fundamental matrix between the third image and the fourth image is calculated, the epipolar lines corresponding to the static points in the third image and the epipolar lines corresponding to the static points in the fourth image need to be calculated based on the fundamental matrix between the third image and the fourth image.
[0074] Specifically, based on the fundamental matrix between the third image and the fourth image, the calculation formulas for calculating the epipolar lines corresponding to the static points in the third image are as follows:
[0075] L1=F*P1 j
[0076] Where F represents the basic matrix between the third image and the fourth image, P1 j represents a static point in the third image, and L1 represents an epipolar line corresponding to the static point in the third image.
[0077] The calculation formula for calculating the epipolar line corresponding to the static point in the fourth image based on the basic matrix between the third image and the fourth image is as follows:
[0078] L2=F T *P2 j
[0079] Among them, F T represents the transpose of the fundamental matrix between the third and fourth images, P2 jrepresents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
[0080] Step S104: Acquire matching points.
[0081] The matching points are used to represent corresponding points between the object in the third image and the object in the fourth image.
[0082] To obtain matching points, it is necessary to first perform panoramic separation on the third image and the fourth image respectively to obtain the object area and background area of the third image and the object area and background area of the fourth image, then extract the object area in the third image and the object area in the fourth image, and then obtain the first label and the first mask corresponding to the object in the object area in the third image, as well as the second label and the second mask corresponding to the object in the object area in the fourth image. Finally, based on the first mask and the second mask, detect whether the bounding boxes of the first label and the second label overlap, where the first label and the second label belong to the same label. If they overlap, the matching point is obtained.
[0083] Since both the third image and the fourth image include objects and backgrounds, the area corresponding to the object is the object area, and the area corresponding to the background is the background area. Panoramic separation can be used to obtain the object area and background area of the third image and the object area and background area of the fourth image.
[0084] Since the object area in the image may contain different objects, this application marks different objects with different labels, such as the label of the car in the object area is 1, the label of the animal in the object area is 2, and so on.
[0085] Based on the first mask and the second mask, it is possible to detect whether the bounding boxes of the first label and the second label overlap. The specific calculation formula is as follows:
[0086] IoU=|C k ∩C k-1 | / |C k ∪C k-1 |
[0087] Among them, C k-1 Represents the first mask, C k Represents the second mask, IoU represents the intersection over union, and the intersection over union is used to determine whether the bounding boxes of the first label and the second label overlap.
[0088] If there is no overlap, it is known that an object in the fourth image does not match any object in the third image and is considered a new object, and features related to its mask are filtered. If there is overlap, it is known that an object in the fourth image matches an object in the third image and matching points are obtained.
[0089] After obtaining the matching point, the distance between the matching point and the corresponding epipolar line needs to be calculated, and then based on the distance between the matching point and the corresponding epipolar line and the preset threshold, it is determined whether the matching point is a static point.
[0090] The formula for calculating the distance between the matching point and the corresponding epipolar line is as follows:
[0091] Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image,
[0092] F represents the fundamental matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
[0093] Next, based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it can be determined whether the matching point is a static point. Specifically, if the distance between the matching point and the corresponding epipolar line is less than the preset threshold, the matching point is a static point; if the distance between the matching point and the corresponding epipolar line is greater than or equal to the preset threshold, the matching point is a dynamic point. The preset threshold can be set according to the situation and is not specifically limited here.
[0094] Step S105: If the matching point is a static point, the SLAM state estimation and map update are realized through the static point.
[0095] An embodiment of the present invention provides a visual SLAM method in a dynamic environment based on panoramic segmentation, comprising: first receiving a first image and a second image, wherein the first image and the second image are continuous frame images, then filtering the people in the first image and the second image respectively to obtain a third image and a fourth image, and then based on the third image and the fourth image, determining the polar lines corresponding to the static points in the third image and the polar lines corresponding to the static points in the fourth image, and finally obtaining matching points, wherein the matching points are used to characterize the points corresponding to the objects in the third image and the objects in the fourth image, and if the matching points are static points, the static points are used to realize the estimation of the SLAM state and the update of the map. This application helps to improve the adaptability and accuracy of the SLAM system in a dynamic environment, making it more practical and robust in complex actual scenes.
[0096] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0097] The following are device embodiments of the present invention. For details not fully described therein, reference may be made to the corresponding method embodiments described above.
[0098] Figure 2 The structure diagram of a visual SLAM device in a dynamic environment based on panoptic segmentation provided by an embodiment of the present invention is shown. For the sake of convenience, only the parts related to the embodiment of the present invention are shown. A visual SLAM device in a dynamic environment based on panoptic segmentation includes a receiving module 201, a filtering module 202, an epipolar line calculation module 203, an acquisition module 204 and a visual SLAM implementation module 205, which are specifically as follows:
[0099] A receiving module 201 is configured to receive a first image and a second image, wherein the first image and the second image are continuous frame images;
[0100] A filtering module 202 is configured to filter the people in the first image and the second image respectively to obtain a third image and a fourth image;
[0101] an epipolar line calculation module 203 for determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image;
[0102] An acquisition module 204 is configured to acquire matching points, wherein the matching points are used to represent corresponding points between an object in the third image and an object in the fourth image;
[0103] The visual SLAM implementation module 205 is used to implement SLAM state estimation and map update through the static points if the matching points are static points.
[0104] In a possible implementation, the epipolar line calculation module 203 is further configured to extract static points in the third image and static points in the fourth image respectively;
[0105] calculating a fundamental matrix between the third image and the fourth image based on the pixel coordinates of the static point in the third image and the pixel coordinates of the static point in the fourth image;
[0106] Based on a fundamental matrix between the third image and the fourth image, epipolar lines corresponding to the static points in the third image and epipolar lines corresponding to the static points in the fourth image are respectively calculated.
[0107] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, the epipolar lines corresponding to the static points in the third image are calculated using the following formulas:
[0108] L1=F*P1 j
[0109] Where F represents the basic matrix between the third image and the fourth image, P1j represents a static point in the third image, and L1 represents an epipolar line corresponding to the static point in the third image.
[0110] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, a formula for calculating the epipolar line corresponding to the static point in the fourth image is as follows:
[0111] L2=F T *P2 j
[0112] Among them, F T represents the transpose of the fundamental matrix between the third and fourth images, P2 j represents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
[0113] In a possible implementation, the acquisition module 204 is further configured to perform panorama separation on the third image and the fourth image, respectively, to obtain the object region and the background region of the third image and the object region and the background region of the fourth image;
[0114] extracting an object region in the third image and an object region in the fourth image;
[0115] Obtaining a first label and a first mask corresponding to the object in the object region in the third image, and a second label and a second mask corresponding to the object in the object region in the fourth image;
[0116] Based on the first mask and the second mask, detecting whether bounding boxes of the first label and the second label overlap, wherein the first label and the second label belong to the same label;
[0117] If there is overlap, get the matching points.
[0118] In a possible implementation, after the acquisition module 204, the method further includes: a judgment module configured to calculate the distance between the matching point and the corresponding epipolar line;
[0119] Based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it is determined whether the matching point is a static point.
[0120] In one possible implementation, the formula for calculating the distance between a matching point and its corresponding epipolar line is as follows:
[0121] Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image,
[0122] F represents the fundamental matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
[0123] In a possible implementation, the judgment module is further configured to determine that the matching point is a static point if the distance between the matching point and the corresponding epipolar line is less than a preset threshold;
[0124] If the distance between the matching point and the corresponding epipolar line is greater than or equal to a preset threshold, the matching point is a dynamic point.
[0125] An embodiment of the present invention provides a visual SLAM device in a dynamic environment based on panoramic segmentation, which is specifically used to: first receive a first image and a second image, wherein the first image and the second image are continuous frame images, then filter the people in the first image and the second image respectively to obtain a third image and a fourth image, and then determine the polar lines corresponding to the static points in the third image and the polar lines corresponding to the static points in the fourth image based on the third image and the fourth image, and finally obtain matching points, wherein the matching points are used to represent the points corresponding to the objects in the third image and the objects in the fourth image. If the matching points are static points, the static points are used to realize the estimation of the SLAM state and the update of the map. This application helps to improve the adaptability and accuracy of the SLAM system in a dynamic environment, making it more practical and robust in complex actual scenes.
[0126] Figure 3 Schematic diagram of a terminal provided by an embodiment of the present invention. Figure 3 As shown, the terminal 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and capable of running on the processor 301. When the processor 301 executes the computer program 303, the steps in the above-mentioned embodiments of the visual SLAM method in a dynamic environment based on panoptic segmentation are implemented, such as Figure 1 Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned embodiments of the visual SLAM device in a dynamic environment based on panoptic segmentation are realized, such as Figure 2 Functionality of modules / units 201-205 is shown.
[0127] The present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a processor, it is used to implement a visual SLAM method in a dynamic environment based on panoptic segmentation provided by the various embodiments described above, including:
[0128] receiving a first image and a second image, wherein the first image and the second image are continuous frame images;
[0129] Filtering the people in the first image and the second image respectively to obtain a third image and a fourth image;
[0130] determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image;
[0131] Obtaining matching points, wherein the matching points are used to represent points corresponding to the object in the third image and the object in the fourth image;
[0132] If the matching point is a static point, the SLAM state estimation and map update are realized through the static point.
[0133] In one possible implementation, determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image includes:
[0134] extracting static points in the third image and static points in the fourth image respectively;
[0135] calculating a fundamental matrix between the third image and the fourth image based on the pixel coordinates of the static point in the third image and the pixel coordinates of the static point in the fourth image;
[0136] Based on a fundamental matrix between the third image and the fourth image, epipolar lines corresponding to the static points in the third image and epipolar lines corresponding to the static points in the fourth image are respectively calculated.
[0137] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, the epipolar lines corresponding to the static points in the third image are calculated using the following formulas:
[0138] L1=F*P1 j
[0139] Where F represents the basic matrix between the third image and the fourth image, P1 j represents a static point in the third image, and L1 represents an epipolar line corresponding to the static point in the third image.
[0140] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, a formula for calculating the epipolar line corresponding to the static point in the fourth image is as follows:
[0141] L2=F T *P2 j
[0142] Among them, F T represents the transpose of the fundamental matrix between the third and fourth images, P2 j represents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
[0143] In a possible implementation, obtaining matching points includes:
[0144] performing panorama separation on the third image and the fourth image respectively to obtain an object region and a background region of the third image and an object region and a background region of the fourth image;
[0145] extracting an object region in the third image and an object region in the fourth image;
[0146] Obtaining a first label and a first mask corresponding to the object in the object region in the third image, and a second label and a second mask corresponding to the object in the object region in the fourth image;
[0147] Based on the first mask and the second mask, detecting whether bounding boxes of the first label and the second label overlap, wherein the first label and the second label belong to the same label;
[0148] If there is overlap, get the matching points.
[0149] In a possible implementation, after obtaining the matching points, the following steps are further included:
[0150] Calculate the distance between the matching point and the corresponding epipolar line;
[0151] Based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it is determined whether the matching point is a static point.
[0152] In one possible implementation, the formula for calculating the distance between a matching point and its corresponding epipolar line is as follows:
[0153] Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image,
[0154] F represents the fundamental matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
[0155] In one possible implementation, determining whether a matching point is a static point based on the distance between the matching point and the corresponding epipolar line and a preset threshold includes:
[0156] If the distance between the matching point and the corresponding epipolar line is less than the preset threshold, the matching point is a static point;
[0157] If the distance between the matching point and the corresponding epipolar line is greater than or equal to a preset threshold, the matching point is a dynamic point.
[0158] Among them, the readable storage medium can be a computer storage medium or a communication medium. Communication media include any medium that facilitates the transmission of computer programs from one place to another. Computer storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0159] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of a device can read the execution instructions from the readable storage medium, and at least one processor executes the execution instructions so that the device implements a visual SLAM method in a dynamic environment based on panoptic segmentation provided by various embodiments described above, including:
[0160] receiving a first image and a second image, wherein the first image and the second image are continuous frame images;
[0161] Filtering the people in the first image and the second image respectively to obtain a third image and a fourth image;
[0162] determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image;
[0163] Obtaining matching points, wherein the matching points are used to represent points corresponding to the object in the third image and the object in the fourth image;
[0164] If the matching point is a static point, the SLAM state estimation and map update are realized through the static point.
[0165] In one possible implementation, determining, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image includes:
[0166] extracting static points in the third image and static points in the fourth image respectively;
[0167] calculating a fundamental matrix between the third image and the fourth image based on the pixel coordinates of the static point in the third image and the pixel coordinates of the static point in the fourth image;
[0168] Based on a fundamental matrix between the third image and the fourth image, epipolar lines corresponding to the static points in the third image and epipolar lines corresponding to the static points in the fourth image are respectively calculated.
[0169] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, the epipolar lines corresponding to the static points in the third image are calculated using the following formulas:
[0170] L1=F*P1 j
[0171] Where F represents the basic matrix between the third image and the fourth image, P1 j represents a static point in the third image, and L1 represents an epipolar line corresponding to the static point in the third image.
[0172] In one possible implementation, based on the fundamental matrix between the third image and the fourth image, a formula for calculating the epipolar line corresponding to the static point in the fourth image is as follows:
[0173] L2=F T *P2 j
[0174] Among them, F T represents the transpose of the fundamental matrix between the third and fourth images, P2 j represents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
[0175] In a possible implementation, obtaining matching points includes:
[0176] performing panorama separation on the third image and the fourth image respectively to obtain an object region and a background region of the third image and an object region and a background region of the fourth image;
[0177] extracting an object region in the third image and an object region in the fourth image;
[0178] Obtaining a first label and a first mask corresponding to the object in the object region in the third image, and a second label and a second mask corresponding to the object in the object region in the fourth image;
[0179] Based on the first mask and the second mask, detecting whether bounding boxes of the first label and the second label overlap, wherein the first label and the second label belong to the same label;
[0180] If there is overlap, get the matching points.
[0181] In a possible implementation, after obtaining the matching points, the following steps are further included:
[0182] Calculate the distance between the matching point and the corresponding epipolar line;
[0183] Based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it is determined whether the matching point is a static point.
[0184] In one possible implementation, the formula for calculating the distance between a matching point and its corresponding epipolar line is as follows:
[0185] Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image,
[0186] F represents the fundamental matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
[0187] In one possible implementation, determining whether a matching point is a static point based on the distance between the matching point and the corresponding epipolar line and a preset threshold includes:
[0188] If the distance between the matching point and the corresponding epipolar line is less than the preset threshold, the matching point is a static point;
[0189] If the distance between the matching point and the corresponding epipolar line is greater than or equal to a preset threshold, the matching point is a dynamic point.
[0190] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0191] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A visual SLAM method in a dynamic environment based on panoptic segmentation, characterized in that: include: Receiving a first image and a second image, wherein the first image and the second image are continuous frame images; filtering the persons in the first image and the second image respectively to obtain a third image and a fourth image; Based on the third image and the fourth image, determining an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image; Acquire matching points, wherein the matching points are used to represent points corresponding to the object in the third image and the object in the fourth image; If the matching point is a static point, the SLAM state estimation and map update are realized through the static point.
2. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 1, characterized in that: The determining, based on the third image and the fourth image, epipolar lines corresponding to static points in the third image and epipolar lines corresponding to static points in the fourth image includes: extracting static points in the third image and static points in the fourth image respectively; Calculating a basic matrix between the third image and the fourth image based on pixel coordinates of the static point in the third image and pixel coordinates of the static point in the fourth image; Based on a basic matrix between the third image and the fourth image, epipolar lines corresponding to static points in the third image and epipolar lines corresponding to static points in the fourth image are calculated respectively.
3. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 2, characterized in that: The calculation formulas for respectively calculating the epipolar lines corresponding to the static points in the third image based on the basic matrix between the third image and the fourth image are as follows: L1=F*P1 j Wherein, F represents the basic matrix between the third image and the fourth image, P1 j represents a static point in the third image, and L1 represents the epipolar line corresponding to the static point in the third image.
4. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 2, characterized in that: The calculation formula for calculating the epipolar line corresponding to the static point in the fourth image based on the basic matrix between the third image and the fourth image is as follows: L2=F T *P2 j Among them, F T represents the transpose of the fundamental matrix between the third image and the fourth image, P2 j represents a static point in the fourth image, and L2 represents the epipolar line corresponding to the static point in the fourth image.
5. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 1, characterized in that: The obtaining of matching points comprises: Performing panoramic separation on the third image and the fourth image respectively to obtain an object region and a background region of the third image and an object region and a background region of the fourth image; extracting the object region in the third image and the object region in the fourth image; Acquire a first label and a first mask corresponding to the object in the object area in the third image, and a second label and a second mask corresponding to the object in the object area in the fourth image; Based on the first mask and the second mask, detecting whether the bounding boxes of the first label and the second label overlap, wherein the first label and the second label belong to the same label; If there is overlap, get the matching points.
6. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 5, characterized in that: After obtaining the matching points, the method further includes: Calculating the distance between the matching point and the corresponding epipolar line; Based on the distance between the matching point and the corresponding epipolar line and a preset threshold, it is determined whether the matching point is a static point.
7. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 6, characterized in that: The formula for calculating the distance between the matching point and the corresponding epipolar line is as follows: Among them, P1 j represents a static point in the third image, represents the transpose of the static point in the fourth image, F represents the basic matrix between the third image and the fourth image, L1 represents the epipolar line corresponding to the static point in the third image, L2 represents the epipolar line corresponding to the static point in the fourth image, D j Represents the distance between the matching point and the corresponding epipolar line.
8. A visual SLAM method in a dynamic environment based on panoptic segmentation according to claim 6, characterized in that: The determining whether the matching point is a static point based on the distance between the matching point and the corresponding epipolar line and a preset threshold value includes: If the distance between the matching point and the corresponding epipolar line is less than the preset threshold, the matching point is a static point; If the distance between the matching point and the corresponding epipolar line is greater than or equal to the preset threshold, the matching point is a dynamic point.
9. A visual SLAM device in a dynamic environment based on panoramic segmentation, characterized in that: include: A receiving module, configured to receive a first image and a second image, wherein the first image and the second image are continuous frame images; A filtering module, used to filter the people in the first image and the second image respectively to obtain a third image and a fourth image; an epipolar line calculation module, configured to determine, based on the third image and the fourth image, an epipolar line corresponding to a static point in the third image and an epipolar line corresponding to a static point in the fourth image; An acquisition module, used for acquiring matching points, wherein the matching points are used for representing points corresponding to the object in the third image and the object in the fourth image; The visual SLAM implementation module is used to estimate the SLAM state and update the map through the static point if the matching point is a static point.
10. A terminal, characterized in that: comprising a memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors, and the instructions are executed by the one or more processors so that the one or more processors implement the visual SLAM method in a dynamic environment based on panoramic segmentation as described in any one of claims 1 to 8.