Visual SLAM method, device, system, electronic device and storage medium

By identifying and processing panoramic bird's-eye view images in autonomous parking scenes, using semantic information for non-rigid fitting and pose correction, the problem of insufficient accuracy and stability in traditional visual SLAM methods in autonomous parking scenes is solved, and more efficient positioning and mapping are achieved.

CN118887644BActive Publication Date: 2025-05-13INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410801779.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-05-13
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

Traditional visual SLAM methods are affected by factors such as lighting, signal occlusion and time changes in autonomous parking scenarios, resulting in insufficient accuracy and stability.

Method used

By obtaining the original panoramic view of bird's-eye view of the current frame of the target moving object, semantic information is identified and extracted, and non-rigid fitting and semantic matching is used to obtain vectorized instances of the perceived target, and pose correction and keyframe judgment are performed based on this, and the global vector map is updated.

Benefits of technology

It improves the accuracy and stability of the visual SLAM method in autonomous parking scenarios, provides a more accurate data basis, and enhances user perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887644B_ABST
    Figure CN118887644B_ABST
Patent Text Reader

Abstract

The present invention provides a visual SLAM method, device, system, electronic device and storage medium, the method comprising: identifying a perceived target in an original panoramic bird's-eye view image of a current frame of a target moving body, obtaining a target panoramic bird's-eye view image of the current frame of the target moving body, performing semantic matching and non-rigid fitting on the target panoramic bird's-eye view image, obtaining a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image, performing posture correction on the vectorized instance based on the original posture data of the current frame of the target moving body, judging whether the current frame is a key frame based on the vectorized instance after posture correction, and updating the global vector map of the target moving body based on the above-mentioned vectorized instance when the current frame is a key frame. The visual SLAM method, device, system, electronic device and storage medium provided by the present invention can improve the accuracy and stability of the visual SLAM method in an autonomous parking scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a visual SLAM method, device, system, electronic equipment and storage medium. Background Art

[0002] With the rapid development of big data technology and artificial intelligence technology in recent years, driverless driving has become an important development trend in the automotive industry and even the entire transportation field. As an important part of driverless driving, the development of autonomous parking technology can greatly improve parking safety, improve parking efficiency and enhance driving comfort.

[0003] SLAM (Simultaneous Localization and Mapping) is a simultaneous positioning and mapping technology that is used to locate mobile objects in unknown environments in real time and simultaneously build an environmental map of the mobile object.

[0004] However, the traditional visual SLAM method in the related art is greatly affected by external factors such as lighting conditions, signal occlusion and time changes in the autonomous parking scene, resulting in insufficient accuracy and stability of the above-mentioned traditional visual SLAM method in the autonomous parking scene. Therefore, how to improve the accuracy and stability of the visual SLAM method in the autonomous parking scene is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0005] The present invention provides a visual SLAM method, device, system, electronic device and storage medium, which are used to solve the defect of insufficient accuracy and stability of traditional visual SLAM methods in autonomous parking scenarios in the prior art, and to improve the accuracy and stability of the visual SLAM method in autonomous parking scenarios.

[0006] The present invention provides a visual SLAM method, comprising the following steps.

[0007] Step S1, obtaining an original panoramic bird's-eye view image of a current frame of a target moving object and original position and posture data of the current frame of the target moving object.

[0008] Step S2, identifying the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information as the target panoramic bird's-eye view image of the current frame of the target mobile body, and the type of the perceived target includes at least one of a ground sign, a traffic sign and an obstacle.

[0009] Step S3, based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target mobile body, determine the actual outline of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, and based on a predefined semantic template, perform non-rigid fitting and semantic matching on the actual outline of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, wherein the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes multiple discrete semantic points, and a combination of each of the semantic points constitutes a standard outline of the perceived target.

[0010] Step S4, based on the original posture data of the current frame of the target mobile body, after correcting the posture of the vectorized instance of the perceived target identified in the target panoramic view bird's-eye view image of the current frame of the target mobile body, determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target posture identified in the target panoramic view bird's-eye view image of the current frame of the target mobile body.

[0011] Step S5, when it is determined that the current frame is not the key frame, the target mobile body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by the first preset time length, repeat steps S1 to S4 in the next frame of the current frame until it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length. When it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, the panoramic vector of the target mobile body is updated to obtain an updated global vector map of the target mobile body.

[0012] According to a visual SLAM method provided by the present invention, based on a predefined semantic template, non-rigid fitting and semantic matching are performed on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, including:

[0013] In the case where the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body is obscured, missing or deformed, the actual contour of any perceived target is repaired by non-rigid fitting; and it is determined whether the actual contour of each perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be successfully matched with the standard contour of each perceived target in the semantic template;

[0014] When the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body successfully matches the standard contour of any perceived target in the semantic template, each semantic point in the semantic point set corresponding to the any perceived target in the semantic template is used to replace each pixel point on the actual contour of the any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body; when the similarity between the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body and the standard contour of each perceived target in the semantic target fails to match successfully, the identified any perceived target is removed from the target panoramic bird's-eye view image of the current frame of the target moving body;

[0015] The type information, shape information and position information of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body are obtained as a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body.

[0016] According to a visual SLAM method provided by the present invention, the method of judging whether the current frame is a key frame based on the vectorized instance of the corrected perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body comprises: pairing the vectorized instance of the corrected perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body with the map vectorized instance in the environment map of the current frame of each target mobile body, according to the type of the perceived target;

[0017] Using a nearest point iteration algorithm, after aligning the vectorized instance of the corrected perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair with the map vectorized instance in the environment map of the current frame of the target mobile body in each vectorized instance pair, obtain a posture change relationship between the aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair and the map vectorized instance in the environment map of the current frame of the target mobile body;

[0018] Obtaining target pose data of the target mobile body in the current frame based on a pose change relationship between the aligned vectorized instances of the perceived targets identified in the target panoramic bird's-eye view image of the target mobile body in the current frame in each of the vectorized instance pairs and the map vectorized instances in the environment map of the target mobile body in the current frame;

[0019] Based on the target pose data of the current frame of the target moving body and the target pose data of each of the key frames of the target moving body, respectively calculating the matching degree between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of each of the key frames of the target moving body, and the difference degree between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of the previous key frame of the current frame of the target moving body;

[0020] When the degree of match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of each of the key frames of the target moving body is greater than a matching degree threshold, and the degree of difference in match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of the previous key frame of the current frame of the target moving body is less than a difference degree threshold, the current frame is determined as a key frame.

[0021] According to a visual SLAM method provided by the present invention, the step of obtaining an original panoramic bird's-eye view image of a current frame of a target moving body and original position and posture data of the current frame of the target moving body includes:

[0022] Based on the environment image of the target mobile body in the current frame, an original panoramic bird's-eye view image of the target mobile body in the current frame is generated, and based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located, the original posture data of the target mobile body in the current frame is acquired.

[0023] According to a visual SLAM method provided by the present invention, the original wheel speed meter data of the target mobile body collected by the wheel speed meter in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located are obtained, and the original posture data of the target mobile body in the current frame is obtained, including:

[0024] Based on the original wheel speed meter data of the target mobile body collected by the wheel speed meter in the first time window where the current frame is located, the first position data of the target mobile body in the current frame is obtained; based on the original inertial measurement unit data collected by the inertial measurement unit in the second time window where the current frame is located, the second position data of the target mobile body in the current frame is obtained;

[0025] The extended Kalman filter is used to fuse and correct the first pose data and the second pose data of the current frame of the target moving body to obtain the original pose data of the current frame of the target moving body.

[0026] According to a visual SLAM method provided by the present invention, the method of identifying a perceived target in an original panoramic bird's-eye view image of a current frame of the target moving body, and obtaining an original panoramic bird's-eye view image of the current frame of the target moving body carrying semantic information as a target panoramic bird's-eye view image of the current frame of the target moving body, comprises:

[0027] Inputting the original panoramic bird's-eye view image of the current frame of the target moving object into a perception target recognition model, obtaining the original panoramic bird's-eye view image of the current frame of the target moving object carrying semantic information output by the perception target recognition model as the target panoramic bird's-eye view image of the current frame of the target moving object;

[0028] Among them, the semantic information carried by each pixel point in the target panoramic bird's-eye view image of the current frame of the target moving body is used to indicate that each pixel point does not correspond to the perceived target or the type of the perceived target corresponding to each pixel point; the perceived target recognition model is constructed based on a semantic segmentation model, and is trained based on sample panoramic bird's-eye view images and the sample panoramic bird's-eye view images carrying semantic information. The semantic information carried by each sample pixel point in the sample panoramic bird's-eye view image carrying semantic information is used to indicate that each sample pixel point does not correspond to the perceived target or the type of the perceived target corresponding to each sample pixel point.

[0029] According to a visual SLAM method provided by the present invention, after obtaining the updated global vector map of the target moving body, the method further includes:

[0030] When the target moving body does not move within a second preset time period after the current frame, a global vector map of the target moving body is loop-closed detected and optimized by using a graph optimization method to obtain an optimized global vector map of the target moving body.

[0031] The present invention also provides a visual SLAM device, comprising the following modules.

[0032] A data acquisition module is used to obtain an original panoramic bird's-eye view image of a target moving object in a current frame and original position and posture data of the target moving object in a current frame;

[0033] a target recognition module, configured to recognize a perceived target in an original panoramic bird's-eye view image of a current frame of the target moving body, and obtain an original panoramic bird's-eye view image of a current frame of the target moving body carrying semantic information as a target panoramic bird's-eye view image of a current frame of the target moving body, wherein the type of the perceived target includes at least one of a ground sign, a traffic sign, and an obstacle;

[0034] A semantic matching module is used to determine the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, and to perform non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on a predefined semantic template, so as to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, wherein the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes a plurality of discrete semantic points, and a combination of the semantic points constitutes a standard contour of the perceived target;

[0035] A posture correction module is used to perform posture correction on a vectorized instance of a perceived target identified in a target panoramic bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body, and then determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body;

[0036] A map construction module is used to repeatedly execute steps S1 to S4 in the next frame of the current frame when it is determined that the current frame is not the key frame, the target mobile body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by the first preset time length, until it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length. When it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length, based on the vectorized instance after the correction of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, the panoramic vector of the target mobile body is updated to obtain an updated global vector map of the target mobile body.

[0037] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the visual SLAM methods described above is implemented.

[0038] The present invention also provides a visual SLAM system, comprising: the electronic device as described above, multiple image sensors, a wheel speed meter and an inertial measurement unit; each of the image sensor, the wheel speed meter and the inertial measurement unit are electrically connected to the electronic device respectively; each of the image sensors is arranged on the target mobile body, and is used to obtain the environmental image of the current frame of the environment around the target mobile body, and send the environmental image to the electronic device; the wheel speed meter is arranged on the target mobile body, and is used to obtain the original wheel speed meter data of the target mobile body, and send the obtained original wheel speed meter data to the electronic device; the inertial measurement unit is arranged on the target mobile body, and is used to obtain the original inertial measurement unit data of the target mobile body, and send the obtained original inertial measurement unit data of the target mobile body to the electronic device, so that the electronic device can obtain the original posture data of the current frame of the target mobile body based on the original wheel speed meter data and the original inertial measurement unit data.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the visual SLAM method as described above is implemented.

[0040] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the visual SLAM method described above is implemented.

[0041] The visual SLAM method, device, system, electronic device and storage medium provided by the present invention identify the perceived target in the original panoramic surround bird's-eye view image of the current frame of the target mobile body, obtain the original panoramic surround bird's-eye view image of the current frame of the target mobile body carrying semantic information, and then determine the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target mobile body based on the semantic information in the target panoramic surround bird's-eye view image of the current frame of the target mobile body, perform non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target mobile body based on a predefined semantic template, obtain a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target mobile body, perform posture correction on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target mobile body based on the original posture data of the current frame of the target mobile body, and then perform posture correction on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target mobile body based on the original posture data of the current frame of the target mobile body. The method further comprises the following steps: determining whether the current frame is a key frame based on the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image; when it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, repeating the above steps in the next frame of the current frame until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length, based on the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target moving body, updating the panoramic vector of the target moving body, obtaining an updated global vector map of the target moving body, and being able to more accurately obtain the global vector map of the moving body using the visual SLAM method in various scenarios, improving the accuracy and stability of the visual SLAM method in the autonomous parking scenario, providing a more accurate data basis for the autonomous parking technology, and improving user perception. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 This is one of the flow charts of the visual SLAM method provided by the present invention.

[0044] Figure 2 This is the second flow chart of the visual SLAM method provided by the present invention.

[0045] Figure 3 It is a structural schematic diagram of the visual SLAM device provided by the present invention.

[0046] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] In the description of the invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0049] In the description of the present application, the terms "first", "second", etc. are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually a class, and the number of objects is not limited. For example, the first object can be one or more. In addition, in the description of the present application, "and / or" represents at least one of the connected objects, and the character " / " generally represents that the front and back associated objects are in an "or" relationship.

[0050] It should be noted that with the rapid development of the automobile industry, traffic congestion and safe driving issues have attracted more and more attention from residents and researchers. With the rapid development of big data technology and artificial intelligence technology in recent years, driverless driving has become an important development trend in the automobile industry and even the entire transportation field. As an important part of driverless driving, the development of autonomous parking technology can greatly improve parking safety, improve parking efficiency and enhance driving comfort.

[0051] High-precision and high-frequency positioning of unmanned vehicles is an important prerequisite for the application of unmanned driving technology. The SLAM method is a method that can achieve simultaneous positioning and mapping. It is used to locate mobile objects in unknown environments in real time and build environmental maps of mobile objects at the same time. In the SLAM method, sensors are usually used to perceive the surrounding environment, and the positioning and mapping of mobile objects are achieved through the perception and movement of the mobile object itself, without relying on pre-established maps.

[0052] Autonomous parking scenarios may include indoor autonomous parking scenarios, outdoor autonomous parking scenarios, and underground garage autonomous parking scenarios, etc. However, factors such as lighting conditions, signal shielding, and time changes in indoor autonomous parking scenarios, outdoor autonomous parking scenarios, and underground garage autonomous parking scenarios are quite different.

[0053] SLAM methods in related technologies usually include laser SLAM methods and visual SLAM methods. However, in related technologies, traditional visual SLAM is greatly affected by external factors such as illumination, occlusion, and time changes in autonomous parking scenarios, resulting in insufficient accuracy and stability of the above-mentioned traditional visual SLAM methods in autonomous parking scenarios. Therefore, the accuracy of traditional visual SLAM in autonomous parking scenarios is difficult to meet the needs of autonomous parking. Although the accuracy of the laser SLAM method is very high, from the perspective of industrialization, the high equipment cost of laser radar makes it difficult for the laser SLAM method to be widely used in large-scale industrial production.

[0054] In addition, the features used by traditional visual SLAM methods in related technologies include ORB (Oriented FAST and Rotated BRIEF) features, SIFT (Scale-Invariant Feature Transform) features, etc. These features have great limitations in autonomous parking scenarios, such as lack of environmental texture and large temporal variability.

[0055] Therefore, how to improve the accuracy and stability of the SLAM method without using a lidar is a technical problem that needs to be solved urgently in this field.

[0056] In this regard, the present invention provides a visual SLAM method. The visual SLAM method provided by the present invention can greatly improve the accuracy of the map obtained by fusing sensor data of different modalities, and on the other hand, the visual SLAM method provided by the present invention can improve the use of visual information, convert the pictures collected by the camera into a panoramic bird's-eye view, and then perform semantic segmentation on the panoramic bird's-eye view, so that a map can be constructed according to the semantic features of the panoramic bird's-eye view.

[0057] Combine the following Figure 1-Figure 2The visual SLAM method of the present invention is described.

[0058] Figure 1 It is one of the flow charts of the visual SLAM method provided by the present invention, such as Figure 1 As shown, the method includes the following steps: Step S1, obtaining an original panoramic bird's-eye view image of a current frame of a target moving object and original position and posture data of a current frame of a target moving object.

[0059] It should be noted that the execution subject of the embodiment of the present invention is a visual SLAM device. The visual SLAM device can be a vehicle-mounted computer of the target moving body; the visual SLAM device can also be other electronic devices such as a user terminal.

[0060] Specifically, the target moving body in the embodiment of the present invention is an object for map construction by the visual SLAM method provided by the present invention. Based on the visual SLAM method provided by the present invention, an environment map of the target moving body can be constructed.

[0061] It is understandable that the target moving object in the embodiment of the present invention may be determined based on actual needs. The target moving object is not specifically limited in the embodiment of the present invention.

[0062] It should be noted that the mobile body in the embodiment of the present invention may include movable objects such as vehicles, robots, and drones. The specific type of the mobile body is not limited in the embodiment of the present invention.

[0063] Optionally, the moving object in the embodiment of the present invention may be a vehicle. The following takes the target moving object as a vehicle as an example to illustrate the visual SLAM method provided by the present invention.

[0064] In the embodiments of the present invention, the original panoramic bird's-eye view image of the current frame of the target moving body and the position and posture data of the current frame of the target moving body can be obtained in a variety of ways. For example, based on the image captured by the image sensor and the data captured by other sensors, the original panoramic bird's-eye view image of the current frame of the target moving body and the position and posture data of the current frame of the target moving body can be obtained. In the embodiments of the present invention, the specific method of obtaining the original panoramic bird's-eye view image of the current frame of the target moving body and the position and posture data of the current frame of the target moving body is not limited.

[0065] Figure 2 This is the second flow chart of the visual SLAM method provided by the present invention. Figure 2As shown, as an optional embodiment, the original panoramic bird's-eye view image of the current frame of the target mobile body and the original position and posture data of the current frame of the target mobile body are obtained, including: based on the environment image of the current frame of the target mobile body, the original panoramic bird's-eye view image of the current frame of the target mobile body is generated, based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located, the original position and posture data of the target mobile body in the current frame is obtained.

[0066] Specifically, a plurality of image sensors are arranged around the body of the target moving body. The above image sensors can be used to collect a plurality of environmental images including the surrounding environment of the target moving body, and the above environmental images include the environmental information of all angles around the target moving body.

[0067] After acquiring multiple environmental images of the current frame including the surrounding environment of the target mobile body collected by the above-mentioned image sensors, the above-mentioned environmental images can be dedistorted and corrected. After obtaining the environmental images after distortion correction, the transformation matrix obtained by image calibration can be used to project the environmental images after distortion correction from a bird's-eye view, so as to obtain the environmental images from a bird's-eye view.

[0068] By using the fuzzy weight mask set obtained by image calibration to fuzzy each environment image under the bird's-eye view, each weighted environment image can be obtained.

[0069] By using the transformation matrix obtained by image stitching calibration, the weighted environment images are stitched together to obtain the original panoramic bird's-eye view image of the current frame of the target moving object.

[0070] It should be noted that the target moving body in the embodiment of the present invention is provided with a wheel speed meter and an inertial measurement unit.

[0071] The wheel speed meter is a device that can obtain the rotation speed data of the wheels of a moving object. The original wheel speed meter data of the target moving object collected by the wheel speed meter in the embodiment of the present invention may include the rotation speed data of the target moving object vehicle.

[0072] An inertial measurement unit (IMU) is a device that integrates multiple inertial sensors and is used to measure and track the acceleration, angular velocity, and direction of a moving object. In an embodiment of the present invention, the raw inertial measurement unit data of the target moving object collected by the inertial measurement unit may include the acceleration, angular velocity, and direction of the target moving object.

[0073] It should be noted that since the data collection frequency of the wheel speed meter and the inertial measurement unit is much higher than that of the image data, in the embodiment of the present invention, without affecting the accuracy of the data, the first time length can be predefined as the time length of the first time window, and the second time length can be predefined as the time length of the second time window based on prior knowledge and / or actual conditions, so that the original wheel speed meter data of the target mobile body collected by the wheel speed meter in the first time window where the current frame is located, the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located, and the original panoramic bird's-eye view image of the target mobile body in the current frame can be marked as synchronous data.

[0074] Optionally, in an embodiment of the present invention, starting from the starting frame at which the wheel speed meter starts to collect the original wheel speed meter data of the target mobile body, a first time window can be determined every time a first time length has passed, and then the first time window where the current frame is located can be determined; similarly, in an embodiment of the present invention, starting from the starting frame at which the inertial measurement unit starts to collect the original inertial measurement unit data of the target mobile body, a second time window can be determined every time a second time length has passed, and then the second time window where the current frame is located can be determined.

[0075] Optionally, in an embodiment of the present invention, a time window with the current frame as the time midpoint and a time window with a first duration can be determined as the first time window where the current frame is located, and a time window with the current frame as the time midpoint and a time window with a second duration can be determined as the second time window where the current frame is located.

[0076] Based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the wheel speedometer in the second time window where the current frame is located, the original posture data used to describe the posture of the target mobile body in the current frame can be obtained through numerical calculation, mathematical statistics or deep learning technology.

[0077] As an optional embodiment, based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located, the original posture data of the target mobile body in the current frame is obtained, including: based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located, obtaining the first posture data of the target mobile body in the current frame; based on the original inertial measurement unit data collected by the inertial measurement unit in the second time window where the current frame is located, obtaining the second posture data of the target mobile body in the current frame.

[0078] The extended Kalman filter is used to fuse and correct the first pose data and the second pose data of the current frame of the target moving body to obtain the original pose data of the current frame of the target moving body.

[0079] Based on the original wheel speed meter data of the target mobile object collected by the wheel speed meter in the first time window where the current frame is located, the first position data of the target mobile object in the current frame can be obtained through numerical calculation, mathematical statistics or deep learning technology, which is used to describe the position and posture of the target mobile object in the current frame.

[0080] Based on the original inertial measurement unit data of the target mobile body collected by the wheel speed meter in the second time window where the current frame is located, the second posture data of the target mobile body in the current frame for describing the posture of the target mobile body in the current frame can be obtained through numerical calculation, mathematical statistics or deep learning technology.

[0081] After the first pose data and the second pose data are fused and corrected by using an extended Kalman filter, the pose data obtained after fusion and correction can be determined as the original pose data of the target moving body in the current frame.

[0082] It should be noted that the Extended Kalman Filter (EKF) is a commonly used state estimation algorithm for estimation and control in nonlinear systems. The Extended Kalman Filter is an extension of the Kalman Filter and can effectively handle the state estimation problem of nonlinear systems. The basic idea of ​​the Extended Kalman Filter is to model the nonlinear system as a linearized system and then use the Kalman Filter to estimate the linear system.

[0083] The state vector s of the known target moving body at the current frame t t , the state vector prediction value State transition matrix F t , the state covariance matrix P t , the state covariance matrix predicted value Process noise matrix Q t , the Kalman gain corresponding to the wheel speed meter Kalman gain corresponding to the inertial measurement unit Measurement matrix corresponding to the wheel speed meter The measurement matrix corresponding to the inertial measurement unit The measurement noise matrix corresponding to the wheel speed meter The measurement noise matrix corresponding to the inertial measurement unit In the prediction stage, we have:

[0084]

[0085] In the status update phase:

[0086]

[0087]

[0088] Based on the above formula, the first pose data is obtained respectively and the second pose data After that, the original pose data of the target moving body can be calculated by the following formula:

[0089]

[0090] Step S2, identifying the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information as the target panoramic bird's-eye view image of the current frame of the target mobile body, and the type of the perceived target includes at least one of a ground sign, a traffic sign and an obstacle.

[0091] It should be noted that the ground signs in the embodiments of the present invention may include but are not limited to lane signs, parking space signs, safety signs, guide signs, and prohibition signs, etc. Among them, the types of lane signs include left turn lane signs, right turn lane signs, straight lane signs, and U-turn lane signs, etc.

[0092] The traffic signs in the embodiments of the present invention are facilities used to provide road users with guidance, warnings, prohibitions or instructions about road traffic, including but not limited to warning signs, prohibition signs, instruction signs, guide signs, road construction safety signs and auxiliary signs.

[0093] Specifically, in an embodiment of the present invention, based on a deep learning approach, the perceived targets in the original panoramic bird's-eye view of the current frame of the target mobile body can be identified, and the original panoramic bird's-eye view of the current frame of the target mobile body carrying semantic information can be obtained as the target panoramic bird's-eye view image of the current frame of the target mobile body.

[0094] As an optional embodiment, the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body is identified, and the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information is obtained as the target panoramic bird's-eye view image of the current frame of the target mobile body, including: inputting the original panoramic bird's-eye view image of the current frame of the target mobile body into a perceived target recognition model, and obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information output by the perceived target recognition model as the target panoramic bird's-eye view image of the current frame of the target mobile body.

[0095] Among them, the semantic information carried by each pixel in the target panoramic bird's-eye view image of the current frame of the target moving body is used to indicate that each pixel does not correspond to a perceived target or the type of perceived target corresponding to each pixel.

[0096] The perception target recognition model is built based on the semantic segmentation model and is trained based on sample panoramic bird's-eye view images and sample panoramic bird's-eye view images carrying semantic information. The semantic information carried by each sample pixel in the sample panoramic bird's-eye view images carrying semantic information is used to indicate that each sample pixel does not correspond to a perception target or the type of perception target corresponding to each sample pixel.

[0097] Specifically, after obtaining the original panoramic bird's-eye view image of the current frame of the target moving object, the original panoramic bird's-eye view image of the current frame of the target moving object can be input into the perception target recognition model, and the above target recognition model recognizes the perception target in the original panoramic bird's-eye view image of the current frame of the target moving object.

[0098] If the target recognition model recognizes that any pixel point in the original panoramic bird's-eye view image of the current frame of the target moving object does not correspond to the above-mentioned perceived target, the above-mentioned pixel point can be annotated with semantic information indicating that the above-mentioned pixel point does not correspond to the above-mentioned perceived target. If the target recognition model recognizes that any pixel point in the original panoramic bird's-eye view image of the current frame of the target moving object corresponds to a certain type of perceived target, the above-mentioned pixel point can be annotated with semantic information indicating the type of perceived target corresponding to the above-mentioned pixel point. Then, the original panoramic bird's-eye view image of the current frame of the target moving object in which each pixel point output by the above-mentioned target recognition model carries semantic information can be obtained as the target panoramic bird's-eye view image of the current frame of the target moving object.

[0099] Step S3, based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, determine the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, and based on the predefined semantic template, perform semantic matching and non-rigid fitting on the actual contour to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the semantic template includes a semantic point set corresponding to the perceived target, the semantic point set corresponding to the perceived target includes multiple discrete semantic points, and the combination of each semantic point constitutes a standard contour of the perceived target.

[0100] Specifically, based on the semantic information of each pixel point in the target panoramic view bird's-eye view image of the current frame of the target mobile body, the pixel points corresponding to the perceived target identified in the target panoramic view bird's-eye view image of the current frame of the target mobile body can be determined, and then the pixel points located on the boundary of the identified perceived target in the target panoramic view bird's-eye view image of the current frame of the target mobile body can be determined as the actual outline of the perceived target.

[0101] It should be noted that after determining the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the type information and identification information of the perceived target can be annotated for the actual contour of the perceived target identified above.

[0102] It should be noted that the semantic point set corresponding to the perception target in the embodiment of the present invention is a semantic point set obtained by abstractly symbolizing the standard outline of the perception target, and the combination of discrete semantic points in the semantic point set corresponding to the perception target can constitute the standard outline of the above-mentioned perception target.

[0103] For example, if the type of the perceived target includes lane markings, and the types of the lane markings include straight lane markings, left turn lane markings, right turn lane markings and U-turn lane markings, and the shape of the straight lane marking is a straight arrow, the shape of the left turn lane marking is a left turn arrow, the shape of the right turn lane marking is a right turn arrow, and the shape of the U-turn lane marking is a U-turn arrow, then the combination of the semantic points in the semantic point set corresponding to the straight lane marking constitutes the standard outline of the straight arrow, the combination of the semantic points in the semantic point set corresponding to the left turn lane marking constitutes the standard outline of the left turn arrow, the combination of the semantic points in the semantic point set corresponding to the right turn lane marking constitutes the standard outline of the right turn arrow, and the combination of the semantic points in the semantic point set corresponding to the U-turn lane marking constitutes the standard outline of the U-turn arrow.

[0104] For another example, if the type of perceived target includes a parking space sign, and the shape of the parking space sign includes a rectangle formed by parking space lines for indicating the parking area, then the combination of the semantic points in the semantic point set corresponding to the parking space sign constitutes the standard outline of the rectangle formed by the parking space lines.

[0105] It should be noted that the semantic template in the embodiment of the present invention is predefined based on prior knowledge and / or actual conditions. The semantic template is not specifically limited in the embodiment of the present invention.

[0106] After determining the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body, semantic matching and two-dimensional non-rigid fitting can be performed on the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body through numerical calculation based on a predefined semantic template to obtain a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body.

[0107] It should be noted that two-dimensional non-rigid fitting refers to the process of fitting a non-rigid shape or curve on a two-dimensional plane. Non-rigid fitting means that the fitted shape is not rigid, that is, it can be deformed or distorted to a certain extent.

[0108] As an optional embodiment, based on a predefined semantic template, non-rigid fitting and semantic matching are performed on the actual contour of the perceived target identified in the target panoramic view bird's-eye view image of the current frame of the target moving body to obtain a vectorized instance of the perceived target identified in the target panoramic view bird's-eye view image of the current frame of the target moving body, including: in the event that the actual contour of any perceived target identified in the target panoramic view bird's-eye view image of the current frame of the target moving body is obscured, missing or deformed, the actual contour of any perceived target is repaired by non-rigid fitting.

[0109] Specifically, if the actual contour of the i-th perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body is shielded, missing or deformed, a two-dimensional non-rigid fitting can be performed on the actual contour of the i-th perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, and a repair operation such as deformation and filling can be performed on the actual contour of the i-th perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body to obtain the actual contour of the i-th perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body without shielding, missing or deformation. Wherein, i represents a positive integer greater than zero.

[0110] It is determined whether the actual contour of each perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be successfully matched with the standard contour of each perceived target in the semantic template.

[0111] Specifically, after repairing the actual contours of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body through non-rigid fitting, it is possible to determine by numerical calculation whether the actual contours of each perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be successfully matched with the standard contours of each perceived target in the semantic template.

[0112] The point P in the target point set is changed to obtain a new point P′, and the corresponding relationship between the two points is P′=R×P+T; where R represents the rotation matrix and T represents the translation vector T.

[0113] The set of all pixel points on the actual contours of all perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body is taken as the target point set. For each point P in the target point set, a reference point Q is found in the semantic template so that the distance s from P to Q = ||Q j -(Rj *P j +T j )|| 2 Shortest.

[0114] Iterate this process and finally get a set of solutions so that the sum of the squares of the distances from all points in the target point set to the corresponding reference points is minimum, where m represents the number of midpoints on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body.

[0115] At the same time, considering that the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body include several types, based on the sum of the actual contours σ of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the matching value of the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be calculated. n represents the number of perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body, and should ensure the matching value of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target moving body. The total is minimal.

[0116] Considering that in actual situations, ground signs also have corresponding constraints, for example, ground signs in adjacent lanes should remain parallel, adjacent parking space lines should remain parallel, etc. In order to ensure that the semantic template meets the actual constraints, after the above parallel constraints are successfully matched with the semantic template, the angle between the corresponding points can be obtained. For example, after the actual contours of the two adjacent parking space signs identified in the target panoramic bird's-eye view image of the current frame of the target moving body are successfully matched with the standard contours of the parking space signs in the semantic template, the straight line obtained by connecting two of the four corner points in the standard contour of the parking space sign in the semantic target can be selected, and the straight line obtained by connecting two of the four corner points in the actual contours of the two adjacent parking space signs identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be used to calculate the angle θ. If the obtained angle is 0 or close to 0, it proves that the similarity coefficient is high. Therefore, the parallel constraint Loss in the semantic template is defined parallel =1-cosθ. In summary, we can get the objective function L of semantic matching:

[0117]

[0118] By solving the objective function and obtaining a solution that satisfies the constraints, it is possible to determine whether the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be successfully matched with the standard contour of each perceived target in the semantic template.

[0119] When the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body successfully matches the standard contour of any perceived target in the semantic template, each semantic point in the semantic point set corresponding to any perceived target in the semantic template is used to replace each pixel point on the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body; when the similarity between the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body and the standard contour of each perceived target in the semantic target fails to match successfully, any perceived target identified is eliminated from the target panoramic bird's-eye view image of the current frame of the target moving body.

[0120] The type information, shape information and position information of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body are obtained as a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body.

[0121] Step S4, based on the original posture data of the current frame of the target moving body, after performing posture correction on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body.

[0122] Specifically, after obtaining the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the pose of the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be corrected by numerical calculation based on the original pose data of the target moving body in the current frame, and then based on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body after the pose correction, it can be determined whether the current frame is a key frame.

[0123] As an optional embodiment, determining whether the current frame is a key frame based on the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body includes: pairing the corrected vectorized instances of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body with the map vectorized instances in the environment map of the current frame of each target mobile body according to the type of the perceived target.

[0124] It can be understood that each key frame is a historical frame before the current frame.

[0125] Specifically, for the environmental map of any key frame target mobile body, after projecting the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body into the environmental map of the above key frame target mobile body, a map vectorized instance adjacent to the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body can be determined in the environmental map of the above key frame target mobile body.

[0126] After determining in the environment map of the above-mentioned key frame target mobile body a map vectorization instance adjacent to the vectorization instance after the correction of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, based on the type of the perceived target, the map vectorization instances in the environment map of the above-mentioned key frame target mobile body that are adjacent to and of the same type as the vectorization instance after the correction of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body can be determined as a vectorization instance pair.

[0127] By using the nearest point iteration algorithm, the vectorized instance after the correction of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance is aligned with the map vectorized instance in the environment map of the current frame of the target mobile body in each vectorized instance, and then the pose change relationship between the aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance and the map vectorized instance in the environment map of the current frame of the target mobile body is obtained.

[0128] It should be noted that the environmental map of the current frame of the target mobile body in the embodiment of the present invention refers to the map representation of the local environmental information perceived by the target mobile body in the current frame, including information such as surrounding objects, obstacles, feature points and local terrain. The global vector map of the target mobile body in the embodiment of the present invention refers to the complete map representation obtained after the target mobile body models and understands the entire environment, including the topological structure, landmark objects, large-scale terrain information, etc. of the entire environment. The global vector map can be used for long-term positioning, navigation and mission planning, as well as an overall understanding of the environment. In visual SLAM, the global vector map is usually continuously updated and optimized with the increase of exploration and observation to improve the ability to recognize and understand the environment.

[0129] In the embodiment of the present invention, the environment map of the target moving body key frame can be constructed based on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the target moving body key frame. The map vectorized instance in the environment map of the key frame target moving body can be determined based on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the target moving body key frame.

[0130] For each vectorized instance pair, the nearest point iteration algorithm can be used to gradually rotate and translate the vectorized instance after the correction of the perceived target pose identified in the target panoramic view bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair, until the vectorized instance after the correction of the perceived target pose identified in the target panoramic view bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair is aligned with the map vectorized instance in the environment map of the current frame of the target mobile body in each vectorized instance pair.

[0131] It should be noted that the Closest Point Iteration algorithm is a numerical method for solving nonlinear equations. It is usually used to solve distance field problems or the closest point projection problem in surface reconstruction. This algorithm is often used in fields such as computer graphics, computer vision, and geometry processing.

[0132] After the corrected vectorized instance of the perceived target pose identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance is aligned with the map vectorized instance in the environment map of the current frame of the target mobile body in each vectorized instance, the pose change relationship between the aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance and the map vectorized instance in the environment map of the current frame of the target mobile body can be obtained.

[0133] Based on the pose change relationship between the aligned vectorized instances of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair and the map vectorized instances in the environment map of the current frame of the target mobile body, the target pose data of the current frame of the target mobile body is obtained.

[0134] Specifically, after weighted summation of the posture change relationship between the vectorized instances after the alignment of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance and the map vectorized instances in the environment map of the current frame of the target mobile body, the result of the above weighted summation can be determined as the target posture data of the current frame of the target mobile body.

[0135] Based on the target pose data of the current frame of the target mobile body and the target pose data of each key frame of the target mobile body, the matching degree between the target panoramic bird's-eye view image of the current frame of the target mobile body and the environment map of each key frame of the target mobile body, as well as the difference between the target panoramic bird's-eye view image of the current frame of the target mobile body and the environment map of the previous key frame of the current frame of the target mobile body are calculated respectively.

[0136] Specifically, based on the target pose data of the target mobile body in the current frame and the target pose data of each target mobile body in the current frame, the matching degree between the target panoramic bird's-eye view image of the target mobile body in the current frame and the environmental map of each key frame of the target mobile body, as well as the difference between the target panoramic bird's-eye view image of the target mobile body in the current frame and the environmental map of the target mobile body in the previous key frame of the current frame can be calculated respectively by numerical calculation.

[0137] It should be noted that the matching degree in the embodiment of the present invention can be used to describe the matching degree between two data. The higher the matching degree between the two data, the higher the matching degree between the two data. The difference degree in the embodiment of the present invention can be used to describe the difference degree between the two data. The higher the difference degree between the two data, the higher the difference degree between the two data.

[0138] When the degree of match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of each key frame of the target moving body is greater than the matching degree threshold, and the degree of difference in match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of the previous key frame of the current frame of the target moving body is less than the difference threshold, the current frame is determined as a key frame.

[0139] It should be noted that the matching degree threshold and the difference degree threshold in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The matching degree threshold and the difference degree threshold are not specifically limited in the embodiment of the present invention.

[0140] Step S5, when it is determined that the current frame is not a key frame, the target moving body of the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, repeat steps S1 to S4 in the next frame of the current frame until it is determined that the current frame is a key frame, the target moving body of the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length. When it is determined that the current frame is a key frame, the target moving body of the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body.

[0141] It should be noted that the first preset duration in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The first preset duration in the embodiment of the present invention is not specifically limited.

[0142] When the target mobile body in the next frame of the current frame stops moving or the current frame is not a key frame but the current frame is separated from the previous key frame by a first preset time length, the vectorized instances after alignment of the perceived targets identified in the target panoramic bird's-eye view image of the current frame of the target mobile body are fused with the map vectorized instances in the environment map of the target mobile body in the previous key frame of the current frame, and the map vectorized instance set in the environment map of the target mobile body in the previous key frame of the current frame is updated.

[0143] Specifically, if any aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body and any map vectorized instance in the environment map of the target mobile body in the previous frame of the current frame point to the same real-world marking line, then the similarity between the above-mentioned aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body and the above-mentioned map vectorized instance in the environment map of the target mobile body in the previous key frame of the current frame is used to determine whether to use the above-mentioned aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body to replace the above-mentioned map vectorized instance in the environment map of the target mobile body in the previous key frame of the current frame.

[0144] After the fusion is completed, the marking line at each position in the real world will only save one vectorized instance in the environment map of the target moving body in the previous key frame of the current frame, thereby realizing the fusion of the vectorized instances.

[0145] If any aligned vectorized instance of a perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body does not point to the same real-world marking line as any map vectorized instance in the environment map of the target mobile body in the previous frame of the current frame, then the aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body is added to the environment map of the target mobile body in the previous key frame of the current frame, thereby achieving an update of the set of map vectorized instances in the environment map of the target mobile body in the previous key frame of the current frame.

[0146] It should be noted that, in the embodiment of the present invention, a KD tree can be constructed based on the semantic information carried in the target panoramic bird's-eye view image of the key frame to quickly locate the key frame. Among them, the KD tree (K-Dimensional Tree) is a multi-dimensional space data structure used to segment and organize high-dimensional data to support efficient nearest neighbor search. It is a binary tree, each node represents a super rectangular area, and the data space is divided into multiple subspaces by selecting a dimension for division at each node in turn.

[0147] As an optional embodiment, after generating the global vector map of the target mobile body, the method further includes: when the target mobile body does not move within a second preset time length after the current frame, using a graph optimization method, performing loop detection and optimization on the global vector map of the target mobile body to obtain an optimized global vector map of the target mobile body.

[0148] It should be noted that the second preset duration in the embodiment of the present invention may be determined based on prior knowledge and / or actual conditions. The second preset duration in the embodiment of the present invention is not specifically limited.

[0149] Specifically, when the target moving body does not move within the second preset time period after the current frame, it can be explained that the target moving body remains stationary for a long time, and the global vector map of the target moving body can be optimized and then output.

[0150] Based on the target pose data of each target moving body in the current frame, a pose graph of each target moving body in the current frame may be constructed.

[0151] A pose graph network is constructed by taking each key frame as a node of the pose graph and the pose transformation relationship between any two pose graphs as the edge of the pose graph.

[0152] Select loop detection frames and verification frames.

[0153] Let the current frame be t, and all frames between frame tx and frame t are determined as loop detection frames. In practice, the current frame and all frames within one second before the current frame are generally used as loop detection frames. After obtaining the loop detection frames, potential loop frames are selected from each key frame as verification frames based on the similarity of posture and time series.

[0154] The similarity between the pose graph of the target moving body in each loop detection frame and the pose graph of the target moving body in each verification frame is calculated. The similarity includes temporal similarity, spatial similarity, and semantic category similarity.

[0155] If the similarity between the pose graph of the target moving body in any loop detection frame and the pose graph of the target moving body in any verification frame exceeds the similarity threshold, it is considered that the above loop detection frame has a loop, and the loop information is fed back to the mapping part to optimize the pose and global vector map to reduce the cumulative error of the system. If the similarity between the pose graph of the target moving body in any loop detection frame and the pose graph of the target moving body in any verification frame does not exceed the similarity threshold, it is determined that the above loop detection frame has not a loop, and the target pose data of the above loop detection frame to compensate for the moving body is extracted, the pose graph of the target moving body in the above loop detection frame is constructed, and the nodes and edges corresponding to the pose graph of the target moving body in the above loop detection frame are determined, and then the pose graph of the above loop detection frame can be added to the above pose graph network.

[0156] It should be noted that each time the pose graph network is updated, a global BA pose optimization can be performed until an optimized global vector map of the target moving body is generated. Among them, BA (Bundle Adjustment) pose optimization is a technology used to improve the camera pose and three-dimensional scene structure, which is commonly used in photogrammetry, computer vision and robotics. Its goal is to optimize the camera pose and scene structure by minimizing the reprojection error, so as to minimize the error between the observed image feature points and their corresponding points in the three-dimensional scene.

[0157] The embodiment of the present invention obtains the original panoramic surround bird's-eye view image of the current frame of the target moving body carrying semantic information by identifying the perceived target in the original panoramic surround bird's-eye view image of the current frame of the target moving body, and then determines the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body based on the semantic information in the target panoramic surround bird's-eye view image of the current frame of the target moving body, performs non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body based on a predefined semantic template, obtains the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body, performs posture correction on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body, and then performs posture correction on the vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body. The vectorized instance after the correction of the recognized perceived target posture determines whether the current frame is a key frame. When it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, the above steps are repeated in the next frame of the current frame until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, but the current frame is separated from the previous key frame by a first preset time length, based on the vectorized instance after the correction of the perceived target posture recognized in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body, which can more accurately obtain the global vector map of the moving body using the visual SLAM method in various scenarios, can improve the accuracy and stability of the visual SLAM method in the autonomous parking scenario, can provide a more accurate data basis for the autonomous parking technology, and can improve user perception.

[0158] The visual SLAM method provided by the present invention can significantly improve the accuracy of positioning and mapping by fusing multi-source sensor data to optimize pose estimation, reduce equipment investment costs, and obtain high-quality global vector maps in autonomous parking scenarios. The visual SLAM method provided by the present invention can effectively improve the stability of mapping and reduce the probability of mismatching by fitting the semantic segmentation results into vectorized instances and then performing map matching based on the vectorized instances.

[0159] Figure 3 Schematic diagram of the structure of the visual SLAM device provided by the present invention. Figure 3 The visual SLAM device provided by the present invention is described, and the visual SLAM device described below and the visual SLAM method described above can be referred to each other. Figure 3 As shown, the device includes: a data acquisition module 301, a target recognition module 302, a semantic matching module 303, a posture correction module 304 and a map construction module 304.

[0160] The data acquisition module 301 is used to obtain the original panoramic bird's-eye view image of the target moving object in the current frame and the original position and posture data of the target moving object in the current frame.

[0161] The target recognition module 302 is used to identify the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, and obtain the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information as the target panoramic bird's-eye view image of the current frame of the target mobile body. The type of the perceived target includes at least one of a ground sign, a traffic sign and an obstacle.

[0162] The semantic matching module 303 is used to determine the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, and to perform non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on a predefined semantic template, so as to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, wherein the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes multiple discrete semantic points, and a combination of the semantic points constitutes a standard contour of the perceived target.

[0163] The posture correction module 304 is used to perform posture correction on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body, and then determine whether the current frame is a key frame based on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body after posture correction.

[0164] The map construction module 304 is used to repeatedly execute steps S1 to S4 in the next frame of the current frame when it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by the first preset time length. When it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by the first preset time length, based on the vectorized instance after the correction of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body.

[0165] Specifically, the data acquisition module 301 , the target recognition module 302 , the semantic matching module 303 , the posture correction module 304 and the map construction module 304 are electrically connected.

[0166] The visual SLAM device in the embodiment of the present invention can use the visual SLAM method in various scenarios to more accurately obtain the global vector map of the moving body, can improve the accuracy and stability of the visual SLAM method in the autonomous parking scenario, can provide a more accurate data basis for autonomous parking technology, and can improve user perception.

[0167] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the visual SLAM method, which includes: step S1, obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body and the original posture data of the current frame of the target mobile body; step S2, identifying the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information, as the target panoramic bird's-eye view image of the current frame of the target mobile body, and the type of the perceived target includes at least one of ground signs, traffic signs and obstacles; step S3, based on the target The semantic information in the target panoramic surround bird's-eye view image of the current frame of the target moving body is determined, and the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body is determined. Based on the predefined semantic template, the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body is non-rigidly fitted and semantically matched to obtain a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body. The semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes multiple discrete semantic points. The combination of each semantic point constitutes a perceived The standard outline of the target; step S4, based on the original posture data of the current frame of the target moving body, after performing posture correction on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body; step S5, if it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, repeat steps S1 to S5 in the next frame of the current frame 4. Until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length. When it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length, based on the vectorized instance after correction of the perceived target posture recognized in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body.

[0168] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0169] Based on the contents of the above embodiments, a visual SLAM system includes: the electronic device as described above, a plurality of image sensors, a wheel speed meter and an inertial measurement unit; each of the image sensor, the wheel speed meter and the inertial measurement unit is electrically connected to the electronic device respectively;

[0170] Each of the image sensors is arranged on the target mobile body, and is used to obtain an environmental image of a current frame of the surrounding environment of the target mobile body, and send the environmental image to the electronic device;

[0171] The wheel speed meter is arranged on the target mobile body, and is used to obtain the original wheel speed meter data of the target mobile body, and send the obtained original wheel speed meter data to the electronic device;

[0172] The inertial measurement unit is arranged on the target mobile body, and is used to obtain original inertial measurement unit data of the target mobile body, and send the obtained original inertial measurement unit data of the target mobile body to the electronic device, so that the electronic device can obtain original posture data of the current frame of the target mobile body based on the original wheel speed meter data and the original inertial measurement unit data.

[0173] It should be noted that the visual SLAM system in the embodiment of the present invention can be used to execute the visual SLAM method in the above embodiments. The specific steps of the visual SLAM system in the embodiment of the present invention to execute the visual SLAM method can refer to the contents of the above embodiments, which will not be repeated in the embodiment of the present invention.

[0174] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the visual SLAM method provided by the above-mentioned methods, and the method includes: step S1, obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body and the original posture data of the current frame of the target mobile body; step S2, identifying the perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, and obtaining the original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information as the original posture data of the current frame of the target mobile body. The target panoramic bird's-eye view image, the type of the perceived target includes at least one of a ground sign, a traffic sign and an obstacle; step S3, based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, determining the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, and performing non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on a predefined semantic template, to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the semantic template including a semantic point set corresponding to the perceived target, and the perceived target is detected by the vectorization method. The semantic point set corresponding to the known target includes multiple discrete semantic points, and the combination of each semantic point constitutes the standard outline of the perceived target; step S4, based on the original posture data of the current frame of the target moving body, after the posture of the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body is corrected, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body, determine whether the current frame is a key frame; step S5, if it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, in the current frame Repeat steps S1 to S4 for the next frame until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length. When it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body.

[0175] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the visual SLAM method provided by the above-mentioned methods, the method comprising: step S1, obtaining an original panoramic bird's-eye view image of a current frame of a target mobile body and original posture data of the current frame of the target mobile body; step S2, identifying a perceived target in the original panoramic bird's-eye view image of the current frame of the target mobile body, obtaining an original panoramic bird's-eye view image of the current frame of the target mobile body carrying semantic information, as the target panoramic bird's-eye view image of the current frame of the target mobile body, the types of perceived targets include ground At least one of a surface sign, a traffic sign and an obstacle; step S3, based on the semantic information in the target panoramic surround bird's-eye view image of the current frame of the target moving body, determining the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body, and based on a predefined semantic template, performing non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body to obtain a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body, the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes multiple discrete semantic points, and the combination of each semantic point constitutes the standard outline of the perceived target; step S4, based on the original posture data of the current frame of the target moving body, after performing posture correction on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body; step S5, if it is determined that the current frame is not a key frame, the target moving body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by a first preset time length, repeat the execution in the next frame of the current frame Perform steps S1 to S4 until it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length. When it is determined that the current frame is a key frame, the target moving body in the next frame of the current frame stops moving, or although the current frame is determined not to be a key frame, the current frame is separated from the previous key frame by a first preset time length, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body, the panoramic vector of the target moving body is updated to obtain an updated global vector map of the target moving body.

[0176] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0177] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual SLAM method, characterized in that: include: Step S1, obtaining an original panoramic bird's-eye view image of a current frame of a target moving object and original position and posture data of the current frame of the target moving object; Step S2, identifying a perceived target in the original panoramic bird's-eye view image of the current frame of the target moving body, acquiring the original panoramic bird's-eye view image of the current frame of the target moving body carrying semantic information as the target panoramic bird's-eye view image of the current frame of the target moving body, wherein the type of the perceived target includes at least one of a ground sign, a traffic sign, and an obstacle; Step S3, based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, determine the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, and based on a predefined semantic template, perform non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, wherein the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes multiple discrete semantic points, and a combination of each of the semantic points constitutes a standard contour of the perceived target; Step S4, after performing posture correction on the vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body, judging whether the current frame is a key frame based on the vectorized instance of the perceived target posture corrected in the target panoramic bird's-eye view image of the current frame of the target moving body; Step S5, when it is determined that the current frame is not the key frame, the target mobile body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by the first preset time length, repeat steps S1 to S4 in the next frame of the current frame until it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length. When it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length, based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, the panoramic vector of the target mobile body is updated to obtain an updated global vector map of the target mobile body.

2. The visual SLAM method according to claim 1, wherein The method of performing non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the predefined semantic template to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body comprises: In the case where the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body is shielded, missing or deformed, repairing the actual contour of any perceived target by non-rigid fitting; Determine whether the actual contour of each perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body can be successfully matched with the standard contour of each perceived target in the semantic template; When the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body successfully matches the standard contour of any perceived target in the semantic template, each semantic point in the semantic point set corresponding to the any perceived target in the semantic template is used to replace each pixel point on the actual contour of the any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body; when the similarity between the actual contour of any perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body and the standard contour of each perceived target in the semantic target fails to match successfully, the identified any perceived target is removed from the target panoramic bird's-eye view image of the current frame of the target moving body; The type information, shape information and position information of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body are obtained as a vectorized instance of the perceived target identified in the target panoramic surround bird's-eye view image of the current frame of the target moving body.

3. The visual SLAM method according to claim 1, wherein The determining whether the current frame is a key frame based on the corrected vectorized instance of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target moving body comprises: According to the type of the perceived target, pairing respectively the vectorized instances of the perceived target posture correction identified in the target panoramic bird's-eye view image of the current frame of the target moving body with the map vectorized instances in the environment map of the current frame of each of the target moving bodies; Using a nearest point iteration algorithm, after aligning the vectorized instance of the corrected perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair with the map vectorized instance in the environment map of the current frame of the target mobile body in each vectorized instance pair, obtain a posture change relationship between the aligned vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target mobile body in each vectorized instance pair and the map vectorized instance in the environment map of the current frame of the target mobile body; Obtaining target pose data of the target mobile body in the current frame based on a pose change relationship between the aligned vectorized instances of the perceived targets identified in the target panoramic bird's-eye view image of the target mobile body in the current frame in each of the vectorized instance pairs and the map vectorized instances in the environment map of the target mobile body in the current frame; Based on the target pose data of the current frame of the target moving body and the target pose data of each of the key frames of the target moving body, respectively calculating the matching degree between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of each of the key frames of the target moving body, and the difference degree between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of the previous key frame of the current frame of the target moving body; When the degree of match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of each of the key frames of the target moving body is greater than a matching degree threshold, and the degree of difference in match between the target panoramic bird's-eye view image of the current frame of the target moving body and the environment map of the previous key frame of the current frame of the target moving body is less than a difference degree threshold, the current frame is determined as a key frame.

4. The visual SLAM method according to claim 1, wherein: The step of acquiring an original panoramic bird's-eye view image of a current frame of a target moving object and original position and posture data of the current frame of the target moving object includes: Based on the environment image of the target mobile body in the current frame, an original panoramic bird's-eye view image of the target mobile body in the current frame is generated, and based on the original wheel speedometer data of the target mobile body collected by the wheel speedometer in the first time window where the current frame is located and the original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located, the original posture data of the target mobile body in the current frame is acquired.

5. The visual SLAM method according to claim 4, wherein: The method of acquiring original position data of the target mobile body in the current frame based on original wheel speed meter data of the target mobile body collected by the wheel speed meter in the first time window where the current frame is located and original inertial measurement unit data of the target mobile body collected by the inertial measurement unit in the second time window where the current frame is located comprises: Based on the original wheel speed meter data of the target mobile body collected by the wheel speed meter in the first time window where the current frame is located, the first position data of the target mobile body in the current frame is obtained; based on the original inertial measurement unit data collected by the inertial measurement unit in the second time window where the current frame is located, the second position data of the target mobile body in the current frame is obtained; The extended Kalman filter is used to fuse and correct the first pose data and the second pose data of the current frame of the target moving body to obtain the original pose data of the current frame of the target moving body.

6. The visual SLAM method according to claim 1, characterized in that: The step of identifying a perceived target in an original panoramic bird's-eye view image of a current frame of the target moving object and acquiring an original panoramic bird's-eye view image of a current frame of the target moving object carrying semantic information as a target panoramic bird's-eye view image of a current frame of the target moving object includes: Inputting the original panoramic bird's-eye view image of the current frame of the target moving object into a perception target recognition model, obtaining the original panoramic bird's-eye view image of the current frame of the target moving object carrying semantic information output by the perception target recognition model as the target panoramic bird's-eye view image of the current frame of the target moving object; The semantic information carried by each pixel in the target panoramic bird's-eye view image of the current frame of the target moving object is used to indicate that each pixel does not correspond to the perceived target or the type of the perceived target corresponding to each pixel; The perceived target recognition model is constructed based on a semantic segmentation model and is trained based on a sample panoramic view bird's-eye view image and the sample panoramic view bird's-eye view image carrying semantic information. The semantic information carried by each sample pixel in the sample panoramic view bird's-eye view image carrying semantic information is used to indicate that each sample pixel does not correspond to the perceived target or the type of the perceived target to which each sample pixel corresponds.

7. The visual SLAM method according to any one of claims 1 to 6, characterized in that: After obtaining the updated global vector map of the target moving object, the method further includes: When the target moving body does not move within a second preset time period after the current frame, a global vector map of the target moving body is loop-closed detected and optimized by using a graph optimization method to obtain an optimized global vector map of the target moving body.

8. A visual SLAM device, characterized in that: include: A data acquisition module is used to obtain an original panoramic bird's-eye view image of a target moving object in a current frame and original position and posture data of the target moving object in a current frame; a target recognition module, configured to recognize a perceived target in an original panoramic bird's-eye view image of a current frame of the target moving body, and obtain an original panoramic bird's-eye view image of a current frame of the target moving body carrying semantic information as a target panoramic bird's-eye view image of a current frame of the target moving body, wherein the type of the perceived target includes at least one of a ground sign, a traffic sign, and an obstacle; A semantic matching module is used to determine the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on the semantic information in the target panoramic bird's-eye view image of the current frame of the target moving body, and to perform non-rigid fitting and semantic matching on the actual contour of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body based on a predefined semantic template, so as to obtain a vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body, wherein the semantic template includes a semantic point set corresponding to the perceived target, and the semantic point set corresponding to the perceived target includes a plurality of discrete semantic points, and a combination of the semantic points constitutes a standard contour of the perceived target; A posture correction module is used to perform posture correction on a vectorized instance of a perceived target identified in a target panoramic bird's-eye view image of the current frame of the target moving body based on the original posture data of the current frame of the target moving body, and then determine whether the current frame is a key frame based on the corrected vectorized instance of the perceived target identified in the target panoramic bird's-eye view image of the current frame of the target moving body; A map construction module is used to repeatedly execute steps S1 to S4 in the next frame of the current frame when it is determined that the current frame is not the key frame, the target mobile body in the next frame of the current frame has not stopped moving, and the current frame is separated from the previous key frame by the first preset time length, until it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length. When it is determined that the current frame is the key frame, the target mobile body in the next frame of the current frame stops moving, or although the current frame is determined not to be the key frame, the current frame is separated from the previous key frame by the first preset time length, based on the vectorized instance after the correction of the perceived target posture identified in the target panoramic bird's-eye view image of the current frame of the target mobile body, the panoramic vector of the target mobile body is updated to obtain an updated global vector map of the target mobile body.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the visual SLAM method according to any one of claims 1 to 7 is implemented.

10. A visual SLAM system, characterized in that: include: The electronic device according to claim 9, a plurality of image sensors, a wheel speed meter and an inertial measurement unit; each of the image sensor, the wheel speed meter and the inertial measurement unit is electrically connected to the electronic device respectively; Each of the image sensors is arranged on the target mobile body, and is used to obtain an environmental image of a current frame of the surrounding environment of the target mobile body, and send the environmental image to the electronic device; The wheel speed meter is arranged on the target mobile body, and is used to obtain the original wheel speed meter data of the target mobile body, and send the obtained original wheel speed meter data to the electronic device; The inertial measurement unit is arranged on the target mobile body, and is used to obtain original inertial measurement unit data of the target mobile body, and send the obtained original inertial measurement unit data of the target mobile body to the electronic device, so that the electronic device can obtain original posture data of the current frame of the target mobile body based on the original wheel speed meter data and the original inertial measurement unit data.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the visual SLAM method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Identification-assisted visual inertia augmented reality registration method

    CN111784775A

  • Top view-based parking lot vehicle self-positioning and map construction method

    CN111862673A