Object Detection Method for Complex Environments Based on Visual Map Reconstruction
Through the method based on visual map reconstruction, feature points are extracted using monocular cameras and ORB algorithms, and bounding box identification is combined with convolutional neural networks, which solves the problem of insufficient object detection efficiency, adaptability and accuracy in complex environments in the existing technology, and achieves more efficient and accurate object detection effects.
Patent Information
- Application Number
- CN202310831744.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-07-07
AI Technical Summary
The prior art is difficult to achieve efficient, adaptable and accurate object detection in real complex environments, especially in SLAM algorithms, the real-time and applicability of graph construction are insufficient.
Using a method based on visual map reconstruction, the video stream is obtained through a monocular camera, converted into a grayscale map, and feature points are extracted using the ORB algorithm, keyframes are obtained and local maps are constructed, and bounding box recognition and object detection are combined with a convolutional neural network.
The efficiency, adaptability and accuracy of object detection are improved, and the target recognition tasks can be better adapted to complex environments.
Smart Images

Figure CN116758460B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image recognition technology, and in particular, to an object detection method for complex environments based on visual map reconstruction. Background Art
[0002] Currently, simultaneous localization and mapping is a fundamental problem for mobile robots. SLAM algorithms assume a static scene, which limits their applicability in real-world environments. The content in Visual SLAM remains an open question.
[0003] SLAM (Simultaneous Localization and Mapping) is a key technology used in fields such as mobile robots and autonomous driving vehicles to simultaneously achieve localization and mapping. Its development stems from the need for mobile robots to localize and navigate in unknown environments. In the past few decades, the SLAM field has experienced rapid development and progress, mainly due to the emergence of many innovative algorithms and methods in the SLAM field. Among them, filters (such as extended Kalman filters, unscented Kalman filters), optimization algorithms (such as graph optimization), particle filters, etc. are commonly used technical means. In addition, some methods based on feature extraction, deep learning, and machine learning have been applied to SLAM tasks, such as feature-point-based visual SLAM, semantic SLAM, etc. The continuous improvement and innovation of these algorithms have promoted the development of SLAM technology. However, these existing methods ignore or fail to simultaneously solve two basic problems: the real-time nature of mapping in the real environment and the applicability in real complex environments.
[0004] Currently, existing solutions usually rely on pure geometric methods. Learning techniques can improve SLAM solutions in environments with prior dynamic objects, providing more scene information. However, most solutions are not ready to handle real complex scenarios.
[0005] It can be seen that there is an urgent need for an object detection method for complex environments based on visual map reconstruction with high detection efficiency, adaptability, and accuracy. Summary of the Invention
[0006] In view of this, embodiments of the present disclosure provide an object detection method for complex environments based on visual map reconstruction, which at least partially solves the problems of poor detection efficiency, adaptability, and accuracy in the prior art.
[0007] Embodiments of the present disclosure provide an object detection method for complex environments based on visual map reconstruction, including:
[0008] Step 1: Receive the video stream of the target area obtained by the monocular camera, and convert the RGB image to be mapped in the video stream into a grayscale image;
[0009] Step 2: Use the ORB algorithm to extract feature points in the grayscale image;
[0010] Step 3: Obtain key frames according to the feature points and the preset selection strategy;
[0011] Step 4: Insert the key frames into the local map for mapping;
[0012] Step 5: Filter and update the feature points of the key frames after mapping;
[0013] Step 6: Use a convolutional neural network to divide the filtered and updated key frames into grids of different sizes, and each grid predicts a set of bounding boxes and corresponding class probabilities;
[0014] Step 7: Use the detection model to predict the probability that each bounding box belongs to the human category to obtain the target detection result;
[0015] Step 8: After completing the local mapping of a single key frame, remove the local key frame, and at the same time return to Step 3 to update the key frame until all key frames are recognized to obtain the target recognition result corresponding to the target area.
[0016] According to a specific implementation manner of the embodiment of the present disclosure, Step 1 specifically includes:
[0017] Receive the video stream of the target area obtained by the monocular camera;
[0018] Perform distortion removal processing on the RGB image to be mapped in the video stream;
[0019] Convert the RGB image after distortion removal processing into a grayscale image.
[0020] According to a specific implementation manner of the embodiment of the present disclosure, Step 2 specifically includes:
[0021] Step 2.1: Grid division and FAST corner point extraction. Perform grid division on each image, divide the image into multiple grids, and extract feature points by FAST corner points within each grid;
[0022] Step 2.2: Feature point homogenization. Homogenize the extracted feature points according to the number of feature points pre-allocated for each layer of the pyramid;
[0023] Step 2.3: Set the quadtree splitting stop condition. Use the quadtree method to process the feature points. Among them, the stop condition for node splitting is that the number of split nodes is greater than or equal to the required number of feature points, or there is only one feature point in each split node;
[0024] Step 2.4, calculate the direction of each feature point by the gray centroid method.
[0025] According to a specific implementation manner of the embodiment of the present disclosure, the step 3 specifically includes:
[0026] Select any frame in the video stream as the initial key frame, calculate its feature points and descriptors, and perform initial camera positioning and mapping initialization by matching the feature points between the current frame and the initial key frame;
[0027] Select new key frames according to a preset strategy, extract their feature points, and calculate descriptors.
[0028] According to a specific implementation manner of the embodiment of the present disclosure, the step 5 specifically includes:
[0029] Step 5.1, use a motion detection algorithm to detect dynamic regions or objects in the key frame;
[0030] Step 5.2, use a feature point extraction algorithm to extract feature points on the entire key frame;
[0031] Step 5.3, match the extracted feature points with the dynamic region, and determine whether the feature points belong to the dynamic region according to the matching degree;
[0032] Step 5.4, according to the matching result, eliminate the feature points belonging to the dynamic region, and retain stable static feature points;
[0033] Step 5.5, calculate the ratio of the area of the dynamic region to the key frame, and update the number of feature points according to the ratio value.
[0034] According to a specific implementation manner of the embodiment of the present disclosure, the step 7 specifically includes:
[0035] For each scale of the feature map, screen according to the confidence score of the predicted bounding box, use the non-maximum suppression algorithm to eliminate highly overlapping bounding boxes, and retain the bounding box with the highest confidence;
[0036] Screen the retained bounding boxes according to the confidence threshold and the class threshold, and at the same time, output the detection result of the crowd according to the position and class information of the bounding box, wherein the target detection result includes the bounding box coordinates, the class label, and the confidence score.
[0037] According to a specific implementation manner of the embodiment of the present disclosure, the detection model is the Throng-Human YOLOv3Tiny model.
[0038] According to a specific implementation manner of the embodiment of the present disclosure, after the step 7, the method further includes:
[0039] Using the labeled bounding boxes and class information, calculate the loss function between the predicted bounding boxes and the ground truth bounding boxes, and update the network parameters through the backpropagation algorithm to optimize the performance and accuracy of the model;
[0040] Use the test set to evaluate the model, obtain the evaluation results, and accordingly optimize and improve the model.
[0041] The object detection solution for complex environments based on visual map reconstruction in the embodiments of the present disclosure includes: Step 1, receive the video stream of the target area acquired by the monocular camera, and convert the RGB image to be mapped in the video stream into a grayscale image; Step 2, use the ORB algorithm to extract the feature points in the grayscale image; Step 3, obtain the key frames according to the feature points and the preset selection strategy; Step 4, insert the key frames into the local map for mapping; Step 5, perform feature point filtering and updating on the key frames after mapping; Step 6, use a convolutional neural network to divide the filtered and updated key frames into grids of different sizes, and each grid predicts a set of bounding boxes and corresponding class probabilities; Step 7, use the detection model to predict the probability that each bounding box belongs to the human category to obtain the object detection result; Step 8, after completing the local mapping of a single key frame, remove the local key frame, and at the same time return to Step 3 to update the key frame until all key frames are recognized to obtain the object recognition result corresponding to the target area.
[0042] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, feature points in the image are extracted, key frames are selected accordingly and then local mapping is performed, and then feature point filtering and updating are performed on the key frames after mapping for bounding box recognition, and object detection is completed using the recognized bounding boxes, improving the detection efficiency, adaptability and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of a method for object detection for complex environments based on visual map reconstruction provided by the embodiments of the present disclosure;
[0045] Figure 2 It is a schematic diagram of a feature point filtering algorithm provided by the embodiments of the present disclosure;
[0046] Figure 3Schematic diagram of the specific implementation process of an object detection method for complex environments based on visual map reconstruction provided by an embodiment of the present disclosure. Detailed implementation manners
[0047] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0048] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without making creative efforts belong to the scope of protection of the present disclosure.
[0049] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement a device and / or practice a method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.
[0050] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0051] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0052] The embodiments of the present disclosure provide an object detection method for complex environments based on visual map reconstruction. The method can be applied to the object recognition process in scenarios such as mobile robots and autonomous driving vehicles.
[0053] See Figure 1 , which is a schematic flowchart of a target detection method for complex environments based on visual map reconstruction provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:
[0054] Step 1, receive the video stream of the target area obtained by a monocular camera, and convert the RGB image to be mapped in the video stream into a grayscale image;
[0055] Further, the specific steps of Step 1 include:
[0056] Receive the video stream of the target area obtained by a monocular camera;
[0057] Perform distortion removal processing on the RGB image to be mapped in the video stream;
[0058] Convert the RGB image after distortion removal processing into a grayscale image.
[0059] Specifically, when implemented, the target detection method for complex environments based on visual map reconstruction can be implemented in cooperation with software. In this embodiment, taking the target detection method for complex environments based on visual map reconstruction implemented in cooperation with Throng-slam as an example, Throng-slam supports monocular, stereo, and RGBD cameras at the same time, is applicable to complex environments, and has higher accuracy. First, it can receive the video stream of the target area obtained by a monocular camera, perform distortion removal processing on the RGB image to be mapped in the video stream, and then convert the RGB image after distortion removal processing into a grayscale image.
[0060] Step 2, use the ORB algorithm to extract feature points from the grayscale image;
[0061] Based on the above embodiment, the specific steps of Step 2 include:
[0062] Step 2.1, Grid division and FAST corner point extraction: Perform grid division on each image, divide the image into multiple grids, and extract feature points using FAST corner points within each grid;
[0063] Step 2.2, Feature point homogenization: Homogenize the extracted feature points according to the number of feature points pre-allocated for each layer of the pyramid;
[0064] Step 2.3, Set the quadtree splitting stop condition: Use the quadtree method to process the feature points. Among them, the stop condition for node splitting is that the number of split nodes is greater than or equal to the required number of feature points, or there is only one feature point within each split node;
[0065] Step 2.4, Calculate the direction of each feature point through the gray centroid method.
[0066] In specific implementation, the obtained RGB image is subjected to grayscale conversion, and the extraction of ORB feature points is based on the grayscale image. The feature points are extracted by traversing the images of all pyramid layers. Each image is meshed, and FAST corner points are extracted within the grid, which can improve the efficiency of feature point extraction. The feature points are homogenized and eliminated according to the number of feature points pre-allocated for each layer of the pyramid. The quadtree method is used, and the node splitting stop condition is that the number of split nodes is greater than or equal to the required number of feature points, or there is only 1 feature point in each split node. Finally, the direction of each feature point is calculated by the grayscale centroid method.
[0067] Select image A, and take the moment m of image A ab Make the following definitions:
[0068] m ab = ∑ x,y∈A x a y b K(x, y), a, b = {0, 1}
[0069] where (x, y) are the pixel coordinates of the image, and K(x, y) is the grayscale value of the image.
[0070] Step 3, obtain key frames according to the feature points and a preset selection strategy;
[0071] Furthermore, the specific content of step 3 includes:
[0072] Select any frame in the video stream as the initial key frame, and calculate its feature points and descriptors. By matching the feature points between the current frame and the initial key frame, initial camera positioning and mapping initialization are performed;
[0073] Select new key frames according to the preset strategy, extract their feature points, and calculate the descriptors.
[0074] In specific implementation, one frame can be selected as the initial key frame, and its feature points and descriptors are calculated. By matching the feature points between the current frame and the initial key frame, initial camera positioning and mapping initialization are performed, and then new key frames are selected according to a certain strategy. For example, the distance threshold between key frames can be set or selected according to the amount of motion. For the selected key frames, their feature points are extracted and the descriptors are calculated. The process of extracting their feature points and calculating the descriptors can use the ORB feature point extraction algorithm to extract feature points in the image sequence and calculate the ORB descriptors of each feature point, which represent the local features around the feature points. The role of key frames is to identify the significant key frames in the image for feature matching and map construction, so as to achieve camera positioning and the establishment of an environmental map. Key frames have uniqueness and distinguishability and are used to track the camera pose and extract scene structure information.
[0075] Step 4, insert the key frame into the local map for mapping;
[0076] In specific implementation, after obtaining the key frame, the position of the target area corresponding to the key frame can be used as the local map, and the key frame is inserted into the local map for mapping.
[0077] Step 5, perform feature point filtering and updating on the key frame after mapping;
[0078] Based on the above embodiments, step 5 specifically includes:
[0079] Step 5.1, use a motion detection algorithm to detect dynamic regions or objects in the key frame;
[0080] Step 5.2, use a feature point extraction algorithm to extract feature points on the entire key frame;
[0081] Step 5.3, match the extracted feature points with the dynamic region, and determine whether the feature points belong to the dynamic region according to the matching degree;
[0082] Step 5.4, according to the matching result, eliminate the feature points belonging to the dynamic region, and retain stable static feature points;
[0083] Step 5.5, calculate the proportion value of the area of the dynamic region in the key frame, and update the number of feature points according to the proportion value.
[0084] In specific implementation, a motion detection algorithm can be first used to detect dynamic regions or objects in the image, and then a feature point extraction algorithm such as FAST, SIFT, SURF, etc. is used to extract feature points on the entire image. The extracted feature points are matched with the dynamic region, and according to the matching degree, it is determined whether the feature points belong to the dynamic region. According to the matching result, a feature point filtering algorithm is used to eliminate the feature points belonging to the dynamic region, and only stable static feature points are retained. The feature point filtering algorithm is as Figure 2 shown, where the x and y of P Ti represent the upper left coordinates of the target bounding box, and w and h represent the width and height of the bounding box respectively. Different from other algorithms, this algorithm does not need to use the static or dynamic region information of the frame as a mask. It directly uses the target bounding box for filtering, so it does not depend on the size of the image. This filtering algorithm will erase all feature points in the area where there are people.
[0085] The advantage of this feature point filtering algorithm lies in its simplicity and efficiency. By directly using the target bounding box to eliminate the feature points in the person area, we can quickly and accurately remove outliers, improving the accuracy and efficiency of subsequent processing.
[0086] The key to this method lies in the accuracy of object detection and the robustness of feature points. When the object detection algorithm can accurately identify people and provide accurate bounding boxes, the feature point filtering algorithm can better eliminate outliers. At the same time, the robustness of feature points also plays an important role in the effectiveness of the algorithm, ensuring the accuracy and stability of feature points under different poses, illuminations, and occlusions.
[0087] Through the feature point filtering algorithm, we can eliminate the feature points related to people in the image according to the object bounding box, thus improving the effect of outlier elimination. This method is simple and efficient, and does not depend on the image size, and can quickly and accurately remove outliers in the human area.
[0088] Considering the need to solve the problem of calculating the similarity of feature points. In the co-occurrence matrix, the row vector corresponding to each target point can actually be regarded as a vector of a target point. The most common method is to use cosine similarity to measure the vector angle between target vector a and target vector b. The smaller the angle, the greater the cosine similarity, and the more similar the two target points. Its definition is as follows:
[0089]
[0090] At the same time, this algorithm does not need to use the static or dynamic region information of the frame as a mask. It directly uses the object bounding box for filtering, so it does not depend on the image size. This filtering algorithm will erase all feature points in the area where there are people. By directly using the object bounding box to eliminate the feature points in the human area, outliers can be quickly and accurately removed.
[0091] Secondly, considering that when a large number of feature points are filtered in a certain frame, the amount of information available to the system may not be sufficient for tracking. This problem may have two situations: one is that there may be too many people in the scene, and the other is that a person is too close to the camera and occupies most of the image. These two situations may occur simultaneously or independently. Even using the ORB-SLAM2 algorithm specifically designed to recover from lost trajectories, a crowded scene may still hinder the SLAM process. The present invention proposes a module to check the filtered area and update the number of detected ORB feature points, rather than setting a static high value. To improve performance, we use a method of dynamically adjusting the number of feature points. The number of feature points starts from a given initial value and increases accordingly according to the size of the filtered area. The specific rules are as follows:
[0092] When the filtered area reaches 30%, 300 feature points are added.
[0093] When the filtered area reaches 60%, 500 feature points are added.
[0094] When the filtering area reaches 90%, 700 feature points are added.
[0095] When the filtering area is greater than 95%, 1200 feature points are added.
[0096] The advantage of this method is that it dynamically adjusts the number of feature points according to the filtering area in the scene. By increasing the number of feature points, we can improve the performance of the system. However, in a scene where there are no people, simply defining a fixed high value as the number of feature points may reduce the tracking speed. This method helps to improve performance because extracting more feature points means an increase in the amount of computation, while in a scene without people, simply defining a high static value will reduce the tracking speed. Through this method of updating feature points, we can better adapt to the changes in the density of feature points in different scenes, improve the performance of the Throng-slam system, and maintain a high tracking speed in crowded scenes.
[0097] Step 6: Use a convolutional neural network to divide the key frames after filtering update into grids of different sizes, and each grid predicts a set of bounding boxes and corresponding class probabilities;
[0098] In specific implementation, the object detection method for complex environments based on visual map reconstruction in cooperation with Throng-slam may include a human detection thread, and the positioning of the crowd is achieved by predicting the object bounding boxes in the key frames through the human detection thread. Using a convolutional neural network, the input image is divided into grids of different sizes, and each grid predicts a set of bounding boxes and corresponding class probabilities, so as to achieve the rapid detection and positioning of multiple objects in the image.
[0099] Step 7: Use the detection model to predict the probability that each bounding box belongs to the human category to obtain the object detection result;
[0100] Based on the above embodiments, step 7 specifically includes:
[0101] For each scale of feature map, screen according to the confidence score of the predicted bounding box, use the non-maximum suppression algorithm to remove the highly overlapping bounding boxes, and retain the bounding box with the highest confidence;
[0102] Screen the retained bounding boxes according to the confidence threshold and class threshold, and at the same time, output the detection result of the crowd according to the position and class information of the bounding box, where the object detection result includes the bounding box coordinates, class label, and confidence score.
[0103] Optionally, the detection model is the Throng-Human YOLOv3 Tiny model.
[0104] Based on the above embodiments, after step 7, the method further includes:
[0105] Using the labeled bounding boxes and class information, calculate the loss function between the predicted bounding boxes and the ground truth bounding boxes, and update the network parameters through the backpropagation algorithm to optimize the performance and accuracy of the model;
[0106] Use the test set to evaluate the model, obtain the evaluation results, and accordingly tune and improve the model.
[0107] In specific implementation, the present disclosure uses the YOLOv3 object detection algorithm to implement crowd detection. This algorithm can provide the detected object classes, 2D bounding boxes at the corresponding positions, and the confidence of each box. To improve the detection speed, we adopted the YOLOv3 Tiny version, which has fewer layers and filters, and the inference speed is increased by 10 times.
[0108] However, YOLOv3 Tiny may have a slight decrease in accuracy because the standard YOLOv3 includes 80 different object classes, while in the present invention we only focus on crowd detection. To improve the accuracy of crowd detection, we propose the Throng-Human YOLOv3 Tiny model, which can more accurately identify crowd targets in crowded environments.
[0109] By using the Throng-Human YOLOv3 Tiny model, we can achieve efficient and accurate crowd detection in the Throng-SLAM system. Such a detection function can provide key information for the system to help identify and track crowd targets, thereby improving the performance of localization and mapping.
[0110] In summary, we utilize the YOLOv3 object detection algorithm to achieve the function of crowd detection, and improve the detection accuracy by introducing the Throng-Human YOLOv3 Tiny model. This function is of great significance for the application of Throng-SLAM in crowded environments.
[0111] The detection accuracy is represented by the following formula
[0112]
[0113] where the detection accuracy represents the ratio between the detection correctness and the total number of detections. The total number of training is the sum of the correct results (TP) and the wrong results (FP).
[0114] The multi-object detection accuracy is represented by the following formula
[0115]
[0116] where N frames is the total number of detection frames, and N mapped (t) is the number of mapped objects in frame t, and the overlap ratio is the sum of the intersections and unions of each object in each frame.
[0117] At the same time, the test set can also be used to evaluate the model and calculate metrics such as the accuracy of the prediction results. According to the evaluation results, the model is tuned and improved, such as adjusting hyperparameters, increasing the amount of training data, or modifying the network architecture.
[0118] Step 8, after completing the local mapping of a single key frame, remove the local key frame, and at the same time return to step 3 to update the key frame until all key frames are recognized, and obtain the target recognition result corresponding to the target area.
[0119] Specifically, when implementing, after completing the local mapping of a single key frame, remove the local key frame, and at the same time return to step 3 to update the key frame until all key frames are recognized, and obtain the target recognition result corresponding to the target area.
[0120] The object detection method for complex environments based on visual map reconstruction provided in this embodiment extracts feature points in the image, selects key frames based on this, performs local mapping, then filters and updates the feature points of the mapped key frames for bounding box recognition, and uses the recognized bounding boxes to complete object detection, improving the detection efficiency, adaptability, and accuracy.
[0121] The following will illustrate this solution through a specific embodiment, as Figure 3 shown:
[0122] Step 1: Preprocess the RGB image and extract ORB feature points
[0123] The obtained RGB image is grayscale-converted, and the extraction of ORB feature points is based on the grayscale image. Traverse the images of all pyramid layers to extract feature points, grid each image, and extract FAST corner points within the grid, which can improve the efficiency of feature point extraction. The feature points are homogenized and removed according to the number of feature points pre-allocated for each layer of the pyramid. Use the quadtree method, and the node splitting stop condition is that the number of split nodes >= the required number of feature points, or there is only 1 feature point in each split node. Finally, calculate the direction of each feature point through the gray centroid method.
[0124] Select image A, and define the moment m of image A ab as follows:
[0125] m ab = ∑ x,y∈A x a y bK(x,y), a, b = {0, 1}
[0126] where (x, y) are the pixel coordinates of the image, and K(x, y) is the grayscale value of the image.
[0127] Step 2: Update Feature Points
[0128] The present invention proposes a module to check the filtering area and update the number of detected ORB feature points, rather than setting a static high value. To improve performance, we use a method to dynamically adjust the number of feature points. The number of feature points starts from a given initial value and increases accordingly according to the size of the filtering area. The specific rules are as follows:
[0129] When the filtering area reaches 30%, add 300 feature points.
[0130] When the filtering area reaches 60%, add 500 feature points.
[0131] When the filtering area reaches 90%, add 700 feature points.
[0132] When the filtering area is greater than 95%, add 1200 feature points.
[0133] The advantage of this method is that it dynamically adjusts the number of feature points according to the filtering area in the scene. By increasing the number of feature points, we can improve the performance of the system. However, in a scene without people, simply defining a fixed high value as the number of feature points may reduce the tracking speed.
[0134] Step 3: Eliminate invalid key frames through the feature point filtering algorithm
[0135] First, we need to solve the problem of calculating the similarity of feature points. In the co-occurrence matrix, the row vector corresponding to each target point can actually be regarded as a vector of a target point. The most common method is to use cosine similarity to measure the vector angle between target vector a and target vector b. The smaller the angle, the greater the cosine similarity, and the more similar the two target points. Its definition is as follows:
[0136]
[0137] At the same time, this algorithm does not need to use the static or dynamic region information of the frame as a mask. It directly uses the target bounding box for filtering, so it does not depend on the size of the image. This filtering algorithm will erase all feature points in the area where people exist. By directly using the target bounding box to eliminate the feature points in the human area, outliers can be removed quickly and accurately.
[0138] Step 4: Start four threads of Throng-slam
[0139] (1) Start the tracking thread:
[0140] Figure 2 The tracking thread of Throng-slam is shown. The tracking thread mainly consists of two steps, including key-frame tracking and local map tracking.
[0141] Step 1: Based on the map points of the current frame obtained previously, we search for the first-level co-visible key frames that can observe the current frame. Then, we find the second-level co-visible key frames, sub-key frames, and parent key frames of these first-level co-visible key frames and use them as part of the local key frames. At the same time, we extract all the map points in the local key frames to construct a local map point set.
[0142] Step 2: Project the local map points onto the current frame and remove the invalid map points outside the field of view of the current frame. The remaining valid local map points will be matched with the current frame, and the matching result will be optimized (only optimize the pose).
[0143] (2) Start the local mapping thread: When the local mapping thread detects that there are new key frames inserted into the list, it starts to process these newly inserted key frames. When there are no new key frames inserted, the local mapping thread will be in an idle state.
[0144] The purpose of starting the local mapping thread is to process the local map around the newly created key frames, which mainly includes co-visible key frames and map points. Compared with the global map, the map around the newly created key frames is called the local map. Therefore, the task of the local mapping thread is to perform local mapping for the newly created key frames, that is, to create or update the co-visible key frames and map point information of this key frame. Once this information is determined, the local map can be successfully created.
[0145] (3) Start the loop closure detection thread: Loop closure detection plays a very important role in Throng-slam. In Throng-slam, the loop closure thread mainly includes the following three steps:
[0146] Step 1: Loop closure detection
[0147] This step is used to detect whether the robot has returned to the starting position and obtain possible loop closure candidate key frames.
[0148] Step 2: Loop closure correction
[0149] After successful loop closure detection, use the detected loop closure relationship to correct the poses of all key frames and map points. By spreading the loop closure error to each frame and its adjacent frames, the correction of the entire trajectory is achieved.
[0150] Step 3: Global optimization
[0151] After the closed-loop correction is completed, the poses of all key frames and the map points are globally optimized. This step aims to further optimize the accuracy and consistency of the entire SLAM system.
[0152] (4) Start the human detection thread: Implement crowd detection in the Throng-SLAM system by using the Throng-Human YOLO Tiny model. Its bounding box prediction formula mainly adopts the method of logistic regression:
[0153] A x = σ(t x ) + K x
[0154] A y = σ(t y ) + K y
[0155]
[0156]
[0157] where K x , K y represent the offsets of the cells, (x, y) are the coordinates of the upper left corner of the bounding box, P h , P w are the preset bounding box sizes, t x , t y , t h , t w are the four prediction values of the convolutional model.
[0158] Step 5: Evaluate the global consistency of the estimated trajectory through the absolute trajectory error, and compare the absolute distance between the translational components of the estimated trajectory and the ground truth trajectory.
[0159] STR t = R t -1 TQ t
[0160] where R is the estimated trajectory, Q represents the ground truth, and T is the transformation that aligns the two trajectories.
[0161] The units involved in the embodiments described in this disclosure can be implemented in software or in hardware.
[0162] It should be understood that each part of this disclosure can be implemented by hardware, software, firmware, or a combination thereof.
[0163] As described above, it is only the specific implementation manner of the present disclosure. However, the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present disclosure should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.
Claims
1. A target detection method for complex environments based on visual map reconstruction, characterized in that Including: Step 1: Receive the video stream of the target area obtained by the monocular camera, and convert the RGB image to be mapped in the video stream into a grayscale image; Step 2: Use the ORB algorithm to extract feature points in the grayscale image; Step 3: Obtain key frames according to the feature points and a preset selection strategy; Step 4: Insert the key frames into the local map for mapping; Step 5: Filter and update the feature points of the key frames after mapping; Step 6: Use a convolutional neural network to divide the filtered and updated key frames into grids of different sizes, and each grid predicts a set of bounding boxes and corresponding class probabilities; Step 7: Use the detection model to predict the probability that each bounding box belongs to the human body category to obtain the target detection result; Step 8: After completing the local mapping of a single key frame, remove the local key frame, and at the same time return to Step 3 to update the key frame until all key frames are recognized to obtain the target recognition result corresponding to the target area.
2. The method according to claim 1, wherein , The specific content of Step 1 includes: Receive the video stream of the target area obtained by the monocular camera; Perform distortion removal processing on the RGB image to be mapped in the video stream; Convert the RGB image after distortion removal processing into a grayscale image.
3. The method according to claim 1, wherein , The specific content of Step 2 includes: Step 2.1: Grid division and FAST corner point extraction. Perform grid division on each image, divide the image into multiple grids, and extract feature points by FAST corner points within each grid; Step 2.2: Feature point homogenization. Homogenize the extracted feature points according to the number of feature points pre-allocated for each layer of the pyramid; Step 2.3: Set the quadtree splitting stop condition. Use the quadtree method to process the feature points. Among them, the stop condition for node splitting is that the number of split nodes is greater than or equal to the required number of feature points, or there is only one feature point in each split node; Step 2.4: Calculate the direction of each feature point by the gray centroid method.
4. The method according to claim 1, wherein , The specific content of Step 3 includes: Select any frame in the video stream as the initial key frame, and calculate its feature points and descriptors. Through matching the feature points between the current frame and the initial key frame, perform initial camera positioning and mapping initialization; Select new key frames according to the preset strategy, extract their feature points, and calculate descriptors.
5. The method according to claim 1, characterized in that , The specific content of Step 5 includes: Step 5.1: Use a motion detection algorithm to detect dynamic regions or objects in the key frames; Step 5.2: Use a feature point extraction algorithm to extract feature points on the entire key frame; Step 5.3: Match the extracted feature points with the dynamic regions, and determine whether the feature points belong to the dynamic regions according to the matching degree; Step 5.4: According to the matching result, remove the feature points belonging to the dynamic regions and retain the stable static feature points; Step 5.5: Calculate the proportion value of the area of the dynamic region in the key frame, and update the number of feature points according to the proportion value.
6. The method according to claim 1, wherein The specific content of Step 7 includes: For the feature maps of each scale, screen according to the confidence scores of the predicted bounding boxes, use the non-maximum suppression algorithm to remove the highly overlapping bounding boxes, and retain the bounding box with the highest confidence. Filter the remaining bounding boxes according to the confidence threshold and the class threshold, and at the same time, according to the position and class information of the bounding boxes, output the detection results of the crowd, where the object detection results include the bounding box coordinates, class labels, and confidence scores.
7. The method according to claim 1, wherein , the detection model is the Throng-HumanYOLOv3Tiny model.
8. The method according to claim 7, wherein , after step 7, the method further includes: Use the labeled bounding boxes and class information to calculate the loss function between the predicted bounding boxes and the ground truth bounding boxes, and update the network parameters through the backpropagation algorithm to optimize the performance and accuracy of the model; Use the test set to evaluate the model, obtain the evaluation results and accordingly optimize and improve the model.
Citation Information
Patent Citations
Visual SLAM method for target detection based on deep learning
CN112884835A
Semantic vision SLAM positioning method based on target detection in indoor dynamic scene
CN114677323A