Dynamic object recognition and obstacle avoidance navigation method and system based on visual SLAM
Through the visual SLAM system combined with laser sensors and vision sensors, deep learning technology is used to detect and track dynamic objects, solving the accuracy and reliability of SLAM technology in positioning and avoiding obstacles in dynamic environments, and achieving more efficient environmental perception and navigation.
Patent Information
- Application Number
- CN202310652810.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-06-02
AI Technical Summary
Existing SLAM technologies are poor in accuracy and reliability in positioning and obstacle avoidance, especially in environments containing dynamic objects.
A system based on vision SLAM is adopted, combining laser sensors and vision sensors to generate laser point cloud maps and images, and dynamic objects are detected and tracked through deep learning technology to realize environmental modeling and real-time perception and tracking of dynamic objects.
It improves the accuracy and reliability of positioning and obstacle avoidance, can sense and track dynamic objects in real time, adapt to complex environments, and enhances the stability and efficiency of robot navigation.
Smart Images

Figure CN116558526B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of navigation technology, and more specifically, to a dynamic object recognition and obstacle avoidance navigation method and system based on visual SLAM. Background Art
[0002] SLAM (Simultaneous Localization and Mapping) technology is a core technology for robotic navigation and autonomous control. It enables positioning and path planning by building a map of the environment. It is one of the most important functions in applications such as robotics and autonomous driving, playing a particularly crucial role in unknown and unstructured environments.
[0003] SLAM technology involves creating an environmental map and acquiring the robot's pose. However, traditional SLAM methods only consider mapping and localization in static environments and are unable to handle dynamic objects within them. Furthermore, many visual SLAM research proposals fail to consider large, dense, and dynamic indoor environments. To achieve a more accurate and comprehensive understanding of the environment, research on dynamic SLAM systems has generated widespread interest. These systems are capable of handling dynamic objects within the environment, enabling robots to better understand their environment and achieve more efficient and precise navigation and control. However, existing SLAM technology still suffers from technical challenges such as poor accuracy and reliability in positioning and obstacle avoidance.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present invention provide a dynamic object recognition and obstacle avoidance navigation method and system based on visual SLAM, so as to at least solve the technical problem of poor accuracy and reliability of positioning and obstacle avoidance in existing SLAM technology.
[0006] According to one aspect of an embodiment of the present invention, a dynamic object recognition and obstacle avoidance navigation system based on visual SLAM is provided, comprising: a sensor module configured to collect sensor data from an environment, wherein the sensor module comprises a laser sensor and a visual sensor, the laser sensor is configured to generate a laser point cloud map related to the environment, and the visual sensor is configured to collect images related to the environment; a processor module configured to convert the sensor data into an environment map, and perform feature extraction on the image to obtain feature point distribution information of dynamic objects in the image; and a dynamic object module configured to detect and track dynamic objects in the environment based on the environment map.
[0007] According to another aspect of an embodiment of the present invention, a dynamic object recognition and obstacle avoidance navigation method based on visual SLAM is also provided, including: collecting sensor data from an environment, wherein the sensor module includes a laser sensor and a visual sensor, the laser sensor is configured to generate a laser point cloud map related to the environment, and the visual sensor is configured to collect images related to the environment; converting the sensor data into robot posture information and an environment map, and performing feature extraction on the image to obtain feature point distribution information of dynamic objects in the image; detecting and tracking dynamic objects in the environment based on the robot posture information and the environment map.
[0008] In an embodiment of the present invention, a dynamic object recognition and obstacle avoidance navigation method based on visual SLAM is characterized in that it includes: collecting sensor data from the environment, wherein the sensor module includes a laser sensor and a visual sensor, the laser sensor is configured to generate a laser point cloud map related to the environment, and the visual sensor is configured to collect images related to the environment; converting the sensor data into robot posture information and an environmental map, and performing feature extraction on the image to obtain feature point distribution information of dynamic objects in the image; detecting and tracking dynamic objects in the environment based on the robot posture information and the environmental map, thereby solving the technical problem of poor accuracy and reliability in positioning and obstacle avoidance in existing SLAM technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0010] Figure 1 1 is a structural diagram of a dynamic object recognition and obstacle avoidance navigation system based on visual SLAM according to an embodiment of the present application;
[0011] Figure 2 1 is a structural diagram of another dynamic object recognition and obstacle avoidance navigation system based on visual SLAM according to an embodiment of the present application;
[0012] Figure 3 is a comparison of the RPE graphs of ORB-SLAM2, DynaSLAM, and this system according to an embodiment of the present application;
[0013] Figure 4 This is a comparison of the RPE curve errors of the ORB-SLAM2 system (a) and the present system (b) for the selected crowd1 sequence;
[0014] Figure 5This is a comparison of the RPE curve errors of the ORB-SLAM2 system (a) and the present system (b) for the selected crowd3 sequence;
[0015] Figure 6 This is a flowchart of a dynamic object recognition and obstacle avoidance navigation method based on visual SLAM disclosed in an embodiment of the present application;
[0016] Figure 7 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0017] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0018] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0019] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to actual proportional relationships. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values should be interpreted as being merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.
[0020] Overview
[0021] The main technical solutions of this application are: dynamic object detection and tracking based on visual SLAM and deep learning: this solution uses visual SLAM technology to achieve environment modeling and positioning, and at the same time uses deep learning technology to achieve real-time detection and tracking of dynamic objects; dynamic SLAM system based on RGB-D camera: this solution uses RGB-D camera sensor technology to achieve environment modeling and real-time perception and tracking of dynamic objects.
[0022] The main technical features of this application are: real-time perception and tracking of dynamic objects: the dynamic SLAM system can perceive and track dynamic objects in the environment in real time, thereby improving the accuracy and reliability of positioning and obstacle avoidance; robustness and stability: the dynamic SLAM system has strong robustness and stability, and can cope with complex situations in different environments, such as lighting changes, occlusion and noise; multi-sensor fusion: the dynamic SLAM system can improve the accuracy and efficiency of positioning and tracking through the fusion of multiple sensors, such as lidar, cameras, inertial measurement units, etc.
[0023] The main technical effects of this application are: real-time perception and tracking of dynamic objects, improving the accuracy and reliability of positioning and obstacle avoidance; having strong robustness and stability, and being able to cope with complex situations in different environments; multi-sensor fusion, improving the accuracy and efficiency of positioning and tracking.
[0024] Example 1
[0025] The present application embodiment provides another dynamic object recognition and obstacle avoidance navigation system based on visual SLAM, such as Figure 1 As shown, it includes: a sensor module 10, a processor module 12 and a dynamic object module 14.
[0026] Sensor module 10 collects sensor data from the environment. This module includes a laser sensor and a visual sensor, which are used to collect sensor data from the environment. The laser sensor can generate a laser point cloud map, and the visual sensor provides RGB image data. This data is transmitted to dynamic object module 14, which detects people and other objects via the YoloV4-tiny detector.
[0027] The processor module 12 converts sensor data into robot pose information and an environmental map. Specifically, it uses the laser point cloud and image data to perform feature extraction and matching, estimating the robot's pose and building an environmental map. Simultaneously, this module also updates the map and corrects the robot's pose.
[0028] The dynamic object module 14 detects and tracks dynamic objects in the environment and integrates them with the environment map to generate a map containing dynamic information. It first collects sensor data from the environment and converts it into robot posture information and an environment map. Next, it detects and tracks dynamic objects in the environment and integrates them with the environment map to generate a map containing dynamic information. Based on this dynamic information and the environment map, the robot's movement path and target are determined, enabling navigation and control. This process enables the robot to process dynamic objects in the environment, perceive the environment more accurately, and better adapt to environmental changes.
[0029] Specifically, the dynamic object module 14 detects and tracks dynamic objects in the environment. First, the RGB-D image output by the visual sensor is obtained, and the feature point distribution information of the crowd and other objects in the image is obtained through the processor feature extraction. At the same time, the detection frame information of the person in the image is obtained through the improved YC target detection thread of yolov4-tiny. Then, the information of the detection frame is passed to the Tranking thread and combined with the feature point distribution information. The dynamic and static features are distinguished by the variance fitting feature point filtering method, and the features belonging to dynamic objects are regarded as abnormal blocks and eliminated. The static features are then integrated with the environmental map to generate a map with dynamic information. Among them, the target detection uses the yolov4-tiny network, which is trained on the crowdhuman dataset so that it can quickly detect individuals and crowds in the environment.
[0030] The specific implementation steps of the variance fitting feature point filtering method are as follows: first, traverse each detection frame, extract the crowd information in the frame, then extract the feature points in each frame and calculate the depth value of each feature point, filter the depth value, calculate the variance of each pair of adjacent feature points corresponding to the useful depth value, filter the feature points with the calculated variance within a certain range, and the feature points left after filtering are the static points we need.
[0031] In addition, an adaptive key point determination strategy is designed in the key frame. The implementation method is divided into two steps. First, the corresponding pixel points of the two selected key frames are matched. When it is found that there are extra points in the previous frame that are not matched compared to the next frame, these points are regarded as dynamic points. After that, if all points can be matched, the epipolar line is calculated for each point on the two key frames. When the feature point moves, its projected position in the next key frame will also change. When the distance from the projected point to the epipolar line exceeds a certain threshold, it is regarded as a dynamic point. The algorithm of the adaptive threshold th is as follows:
[0032] 1) Get the depth d of the key point in the current frame and the depth d of the corresponding key point in the reference key frame ref .
[0033] 2) Calculate the depth ratio
[0034] 3) Calculate the camera intrinsic parameter f x and f y The maximum value of max(f x ,f y ).
[0035] 4) Calculate the adaptive threshold
[0036] The principle of adaptive thresholding is that as the distance between the camera and the object changes, the same pixel error will result in different depth errors. The size of the adaptive threshold depends on the depth of the current keypoint in the camera coordinate system and the distance between the two keyframes. When the epipolar distance between the two keyframes is large, the adaptive threshold is also increased, which can tolerate a larger epipolar distance and reduce errors. Conversely, when the distance between the two keyframes is small, the adaptive threshold is also reduced, which can improve the detection ability of keypoints that move across the object between the two keyframes.
[0037] This embodiment has the following beneficial effects: by using visual sensors, static and dynamic information in the environment can be fully perceived to achieve more accurate SLAM; by detecting and tracking dynamic objects and integrating them with the environmental map, a map can be generated to better reflect the characteristics and changes of the environment; the system has good real-time and robustness and is suitable for various environments and application scenarios, such as autonomous driving cars, robot patrols, smart logistics, etc.
[0038] Example 2
[0039] The present application embodiment provides another dynamic object recognition and obstacle avoidance navigation system based on visual SLAM, such as Figure 2 As shown, the system includes: a sensor module, a processor module and a dynamic object module, wherein the dynamic object module includes a dynamic object detection thread 142 , a tracking thread 144 , a local mapping thread 146 and a loop detection thread 148 .
[0040] The dynamic object detection thread 142 takes an RGB-D image as input. First, through feature extraction, it obtains the distribution information of the feature points of people and other objects in the image. At the same time, it adds an improved yolov4-tiny object detection thread to the original thread to obtain the detection box information of people in the image.
[0041] The detection frame information is then passed to the tracking thread 144 and combined with the feature point distribution information. Dynamic features are identified and removed using the variance fitting feature point removal method. The resulting static features are then integrated with the environment map to generate a map with dynamic information.
[0042] The local mapping thread 146 retrieves keyframe information from the local map and filters out map points. It then uses local bundle adjustment (BA) to further adjust the pose and improve the map points. The keyframes are then re-filtered.
[0043] The loop detection thread 148 includes two parts: loop detection and loop correction, which mainly detect and correct key frame information. Finally, the pose is optimized through global BA.
[0044] From the feature detection comparison under the fr3_walking_xyz sequence, it can be seen that the system has a good effect in eliminating the characteristics of dynamic crowds.
[0045] exist Figure 3 In this example, the high-dynamic sequence fr3_walking_static is selected to compare the RPE graphs of ORB-SLAM2, DynaSLAM, and this system. (a) shows the ORB-SLAM2 system, (b) shows the DynaSLAM system, and (c) shows the CP-SLAM system. In the RPE curve, the horizontal axis represents time in seconds, while the vertical axis represents the pose error in meters. Higher curves indicate greater error. This demonstrates the high precision and robustness of this system.
[0046] exist Figure 4 and Figure 5 In the figure, the ORB-SLAM2 system (a) and the present system (b) are used to compare the RPE curve errors of the selected crowd1 and crowd3 sequences.
[0047] Table 1 compares this system with ORB-SLAM2 for experimental analysis. The RMSE and mean values of the two systems under ATE for multiple sequences in the dataset are compared. Data marked with bolder colors indicates improved accuracy, and this applies to all the following. The fr3_walking_rpy sequence was selected as a high-dynamic sequence because the people in it are constantly moving in large movements. In this high-dynamic sequence, the RMSE of this system improved by 93.60%, and the mean value increased by 94.33%. The average RMSE improvement in the high-dynamic walking sequence reached 95.93%. This demonstrates that this system can achieve good accuracy in dynamic environments. Table 2 compares some of the most outstanding dynamic SALM systems in recent years.
[0048]
[0049] Table 1
[0050]
[0051] Table 2
[0052] The embodiments of the present application have the following beneficial effects:
[0053] 1. Through a dynamic SLAM system that can handle dynamic objects in the environment, the robot can better perceive the environment, thereby achieving more efficient and precise navigation and control.
[0054] 2. By integrating dynamic object information with the environment map, a map with dynamic information can be generated to better reflect the characteristics and changes of the environment.
[0055] 3. It can adapt to various environments and application scenarios and has broad application prospects.
[0056] Example 3
[0057] This embodiment of the present application provides another dynamic object recognition and obstacle avoidance navigation system based on visual SLAM. The structure of this system is similar to that of Example 1. The difference lies in the specific implementation of the variance fitting and feature point filtering method in the tracking thread of the dynamic object module. The following will focus on the differences, and other functions will not be repeated.
[0058] In this embodiment, the method for filtering out feature points by variance fitting implemented by the tracking thread may include the following steps:
[0059] 1) Traverse each detection box.
[0060] For each detection frame, extract the feature points within it. Feature point extraction algorithms (e.g., SIFT, SURF, ORB, etc.) can be used to detect and extract feature points within the frame. These feature points can be keypoints or corner points in the image. Calculate the depth value of each feature point. By combining data from the laser sensor and the visual sensor, the depth information of the feature points can be obtained. Depth can represent the distance between feature points in three-dimensional space.
[0061] 2) Filter the depth value.
[0062] Filter the depth values of feature points to retain those with useful depth values. You can set a threshold to retain only those with depth values within the threshold range, while excluding those with abnormal or unreliable depth values. This can filter out inaccurate feature points caused by outliers or incorrect depth estimation.
[0063] 3) Calculate the variance of two adjacent feature points.
[0064] The variance of the retained feature points is calculated for each pair. A specific feature descriptor algorithm (for example, SIFT, SURF, etc.) can be selected to calculate the descriptor vector of the feature point.
[0065] For each pair of adjacent feature points, the variance between their descriptor vectors is calculated. The variance can measure the degree of difference between feature points, thereby judging their stability and consistency.
[0066] Specifically, the following sub-steps may be included:
[0067] 3.1) Descriptor vector calculation.
[0068] The specific formulas for generating descriptor vectors (such as SIFT and SURF) may vary depending on the method selected. For example, in the SIFT algorithm, descriptor vector calculation involves complex steps such as Gaussian filtering, gradient calculation, and direction assignment, which will not be described here.
[0069] 3.2) Descriptor difference calculation.
[0070] Descriptor difference calculation involves the comparison and difference measurement between two descriptor vectors. Depending on the difference measurement method chosen, the following formula can be used:
[0071]
[0072] Among them, v1 and v2 represent two descriptor vectors, and v1i and v2i represent the corresponding elements of v1 and v2 respectively.
[0073] 3.3) Convert the descriptor difference into a numerical value.
[0074] The descriptor difference can be converted into a numerical value according to the selected difference measurement method. Taking the Euclidean distance as an example, the difference value can be directly used as a numerical value.
[0075] 3.4) Calculation of variance between feature points.
[0076] The variance calculation between feature points involves statistical analysis of the descriptor difference values. The specific variance calculation formula can be as follows:
[0077]
[0078] Among them, μ k Represents the variance of the feature points, x i represents the descriptor difference value, represents the average value of the descriptor difference value, n represents the number of feature point pairs, k represents the order of the central moment, w i Indicates the weight of each distribution, var i Indicates the variance of each distribution, mean i represents the mean of each distribution, represents the average value of the entire data, and m represents the number of distributions.
[0079] By calculating the variance between feature points using the above method, we can accurately identify and distinguish static points from dynamic points. Static points represent relatively fixed structures in a scene, while dynamic points represent parts of the scene that are changing. This is very useful for tasks such as motion tracking, scene analysis, and dynamic object detection.
[0080] 4) Separate static points and dynamic points.
[0081] Based on the calculated variance values of adjacent feature points, feature points with variances within a certain range are considered static points, assuming they come from static objects in the environment. Feature points corresponding to variances exceeding a threshold are considered dynamic points, assuming they may come from moving objects or dynamic changes in the environment.
[0082] The variance fitting and feature point filtering method in this embodiment can more accurately distinguish static points from dynamic points, thereby providing a more reliable static environment map to support the functions of the dynamic object recognition and obstacle avoidance navigation system based on visual SLAM.
[0083] Example 4
[0084] The present invention provides a method for dynamic object recognition and obstacle avoidance navigation based on visual SLAM. Figure 6 As shown, the following steps are included:
[0085] Step S702: collecting sensor data from the environment, wherein the sensor module includes a laser sensor and a visual sensor, the laser sensor is configured to generate a laser point cloud map related to the environment, and the visual sensor is configured to collect images related to the environment.
[0086] Step S704: converting the sensor data into robot posture information and an environment map, and performing feature extraction on the image to obtain feature point distribution information of dynamic objects in the image;
[0087] Step S706: Detect and track dynamic objects in the environment based on the robot posture information and the environment map.
[0088] The dynamic object recognition and obstacle avoidance navigation method in visual SLAM provided in this embodiment and the dynamic object recognition and obstacle avoidance navigation device embodiment in visual SLAM have the same concept. The specific implementation process is detailed in the device embodiment and will not be repeated here.
[0089] Example 5
[0090] Figure 7 Schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present disclosure is shown. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0091] like Figure 7As shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0092] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0093] In particular, according to an embodiment of the present disclosure, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the method and apparatus of the present application are executed. In some embodiments, the electronic device may further include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.
[0094] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0096] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0097] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device.
[0098] The computer-readable medium carries one or more programs, and when the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments. For example, the electronic device can implement each step of the above method embodiments.
[0099] Industrial Applicability
[0100] Application prospects and advantages of dynamic SLAM systems in industry:
[0101] 1. Autonomous driving cars: Dynamic SLAM systems can achieve more accurate positioning and obstacle avoidance by perceiving and tracking dynamic objects in the environment in real time, thereby improving the safety and reliability of autonomous driving cars.
[0102] 2. Robot patrol: Dynamic SLAM system can help robots plan and navigate paths in complex indoor and outdoor environments, while detecting and tracking moving objects, improving the intelligence and autonomy of robots.
[0103] 3. Intelligent Logistics: Dynamic SLAM systems can perceive and track moving objects in real time, such as transport vehicles and items in logistics, thereby improving the efficiency and safety of logistics operations.
[0104] 4. Industrial Automation: Dynamic SLAM systems can perform robot navigation and positioning in industrial scenarios, while detecting and tracking the position and status of workpieces in real time, improving the efficiency and accuracy of industrial automation.
[0105] In short, with the rapid development of autonomous driving, intelligent logistics, industrial automation, and other fields, dynamic SLAM systems have broad market prospects and commercial application potential. According to market research reports, the global dynamic SLAM market size will continue to grow in the next few years, with autonomous vehicles and robotic patrols being the primary application scenarios. Furthermore, with the popularization and acceleration of industrial automation, dynamic SLAM systems will also be widely used in industries such as manufacturing and logistics.
[0106] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0107] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed terminal device can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0109] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0110] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0111] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A dynamic object recognition and obstacle avoidance navigation system based on visual SLAM, characterized in that: include: a sensor module configured to collect sensor data from an environment, wherein the sensor module includes a laser sensor and a visual sensor, the laser sensor being configured to generate a laser point cloud map related to the environment, and the visual sensor being configured to collect images related to the environment; a processor module configured to convert the sensor data into an environment map, and perform feature extraction on the image to obtain feature point distribution information of dynamic objects in the image; a dynamic object module, configured to detect and track dynamic objects in the environment based on the environment map; The dynamic object module further includes: a dynamic object detection thread configured to detect detection frame information of the dynamic object in the image; a tracking thread configured to combine the detection frame information with the feature point distribution information, and distinguish static feature points from dynamic feature points based on the combination result, and output the distinguished feature points; a local mapping thread configured to convert key points in the distinguished feature points into map points, wherein the map points are information distribution for realizing a map; and a loop detection thread configured to detect and track the dynamic object in the environment based on the map points. The local mapping thread is further configured to: match corresponding pixels of the two selected key frames; if there are extra unmatched pixels in the current key frame compared to the subsequent key frame, the unmatched pixels are regarded as dynamic points; if there are no extra unmatched pixels in the current key frame compared to the subsequent key frame, the epipolar lines of each pixel on the two key frames are calculated; when the pixel in the previous key frame moves, the position of the projection point of the pixel in the subsequent key frame will also change; when the distance from the projection point to the epipolar line exceeds a preset adaptive threshold, the pixel is regarded as a dynamic point; The adaptive threshold is obtained by: obtaining the depth d of the key point in the current frame and the depth d of the corresponding key point in the reference key frame. ref Based on the depth d of the key point in the current frame and the depth d of the key point corresponding to the reference key frame ref , to calculate the depth ratio; calculate the larger value of the camera intrinsic parameter; calculate the adaptive threshold based on the larger value of the camera intrinsic parameter and the depth ratio; The step of calculating the adaptive threshold based on the larger value of the camera intrinsic parameter and the depth ratio includes calculating the adaptive threshold based on the following formula: Wherein, th represents the adaptive threshold, r represents the depth ratio, fx and fy represent the first focal length and the second focal length in the camera intrinsic parameters.
2. The system according to claim 1, wherein: The tracking thread is also configured to: The dynamic feature points and the static feature points are distinguished by using a variance fitting and filtering feature point method, and the dynamic feature points belonging to the dynamic object are regarded as abnormal blocks and removed; The remaining static feature points are integrated with the environment map to generate a map with the static features.
3. The system according to claim 2, characterized in that The tracking thread is also configured to: traversing each detection frame in the image based on the detection frame information, and extracting dynamic object information in each detection frame; Extracting feature points from each frame and calculating a depth value of each feature point, and screening the depth values to obtain a useful depth value; The variances of two adjacent feature points corresponding to the useful depth values are calculated, and the feature points whose calculated variances are within a preset range are filtered, and the feature points remaining after filtering are used as static points.
4. The system according to claim 1, wherein: The size of the adaptive threshold depends on the depth of the key point in the current frame in the camera coordinate system and the distance between the two key frames. When the distance between the two key frames is larger, the adaptive threshold also becomes larger. Conversely, when the distance between the two key frames is smaller, the adaptive threshold also becomes smaller.
5. A dynamic object recognition and obstacle avoidance navigation method based on visual SLAM, characterized in that: include: Collecting sensor data from an environment, wherein the sensor module includes a laser sensor and a visual sensor, the laser sensor is configured to generate a laser point cloud map related to the environment, and the visual sensor is configured to collect images related to the environment; Converting the sensor data into robot posture information and an environment map, and performing feature extraction on the image to obtain feature point distribution information of dynamic objects in the image; A dynamic object module detects and tracks dynamic objects in the environment based on the robot posture information and the environment map; The dynamic object module includes: a dynamic object detection thread configured to detect detection frame information of the dynamic object in the image; a tracking thread configured to combine the detection frame information with the feature point distribution information, and distinguish static feature points from dynamic feature points based on the combination result, and output the distinguished feature points; a local mapping thread configured to convert key points in the distinguished feature points into map points, wherein the map points are information distribution for realizing a map; and a loop detection thread configured to detect and track the dynamic object in the environment based on the map points. The local mapping thread is further configured to: match corresponding pixels of the two selected key frames; if there are extra unmatched pixels in the current key frame compared to the subsequent key frame, the unmatched pixels are regarded as dynamic points; if there are no extra unmatched pixels in the current key frame compared to the subsequent key frame, the epipolar lines of each pixel on the two key frames are calculated; when the pixel in the previous key frame moves, the position of the projection point of the pixel in the subsequent key frame will also change; when the distance from the projection point to the epipolar line exceeds a preset adaptive threshold, the pixel is regarded as a dynamic point; The adaptive threshold is obtained by: obtaining the depth d of the key point in the current frame and the depth d of the corresponding key point in the reference key frame. ref Based on the depth d of the key point in the current frame and the depth d of the key point corresponding to the reference key frame ref , to calculate the depth ratio; calculate the larger value of the camera intrinsic parameter; calculate the adaptive threshold based on the larger value of the camera intrinsic parameter and the depth ratio; The step of calculating the adaptive threshold based on the larger value of the camera intrinsic parameter and the depth ratio includes calculating the adaptive threshold based on the following formula: Wherein, th represents the adaptive threshold, r represents the depth ratio, fx and fy represent the first focal length and the second focal length in the camera intrinsic parameters.
6. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed, the computer is caused to execute the method according to claim 5 .
Citation Information
Patent Citations
Visual SLAM method and device, terminal equipment and storage medium
CN113628334A
Semantic vision SLAM positioning method based on target detection in indoor dynamic scene
CN114677323A