A visual-based real-time three-dimensional positioning and counting method for tree crowns

By using a vision-based real-time 3D tree canopy localization and counting method, which utilizes 3D environment maps and camera image sequences, the problem of low efficiency and insufficient accuracy in existing tree canopy localization and counting technologies is solved. This method achieves real-time and accurate tree canopy localization and counting, and is suitable for landscaping and forestry management.

CN117115219BActive Publication Date: 2026-03-24UNIV OF SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing tree canopy location and counting technologies suffer from problems such as low efficiency, high subjectivity, insufficient accuracy, complex operation, high cost, and difficulty in large-scale application. In particular, existing methods are difficult to achieve accurate tree canopy location and counting in garden management and forest resource protection.

Method used

A vision-based real-time 3D localization and counting method for tree canopies is adopted. By introducing a 3D environment map and using a camera to acquire image sequences, the method combines target tracking, feature tracking, local mapping, loop closure, and global map optimization to avoid duplicate counting and optimize the object radius, thereby achieving real-time 3D localization and counting of tree canopies.

Benefits of technology

It enables real-time three-dimensional positioning and counting of tree canopies, avoiding duplicate counting, improving counting accuracy and efficiency, reducing equipment costs, and is suitable for platforms such as drones, cars, and smartphones, as well as for garden management and forestry management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115219B_ABST
    Figure CN117115219B_ABST
Patent Text Reader

Abstract

The application provides a visual-based real-time three-dimensional positioning and counting method for tree crowns, which comprises the following steps: sensor input, wherein the sensor input information is a picture sequence of a camera; target tracking, which is used for obtaining the number of all targets in the picture sequence and the two-dimensional coordinates of all targets in each picture; feature tracking, which is used for extracting geometric information in the picture sequence and estimating the initial pose of a motion carrier, and key frames are extracted from common frames used for feature tracking according to a rule; local mapping, which is used for receiving the key frames and optimizing the initial value of the motion of the motion carrier and an environment map; loop closure, which is used for identifying a loop and triggering global optimization; and global map optimization, which is used for avoiding repeated counting of the same target caused by different target numbers. The application has the advantages of multiple functions, real-time, accuracy and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of target tracking, object positioning, three-dimensional map construction, and intelligent agriculture and forestry, and is mainly used for positioning, counting, and volume estimation of tree crowns in gardens, and particularly relates to a visual-based real-time three-dimensional positioning and counting method for tree crowns. BACKGROUND

[0002] Existing tree crown positioning and counting are mainly used for garden management and forest resource protection. Existing tree crown positioning and counting techniques and methods have the following disadvantages. One, manual field investigation: low efficiency, strong subjectivity, and high risk. Among them, a related paper on tree counting and measurement method under the condition of shrub and grass sheltering has relevant records. Two, satellite remote sensing: this method has the following disadvantages: wide range, difficulty in distinguishing tree crowns and understory vegetation, no accurate results, difficulty in data acquisition and updating, a paper on tree species classification and counting method and device based on high-resolution satellite remote sensing image has relevant records. Three, biochemical statistics: the same as above, and biochemical field skills are required, the biochemical model established must be consistent with the actual situation, the operation is complex, and the universality is weak (example, a method for estimating hickory yield based on remote sensing and statistical data: class spectral angle index map, standard yield model (tree age), suitability evaluation index (climate, topography, and soil), and growth factor support (crown width, diameter at breast height, and tree height)). The sensors used for tree crown counting are cameras or radars. Both of them have the characteristic that the accuracy is negatively correlated with the distance from the tree crown, so in order to obtain accurate positioning and counting results, the sensors are often located near the tree crown, such as handheld, small car, or unmanned aerial vehicle, and the present method does not limit the type of carrier.

[0003] Radar returns environmental information in the form of point cloud, without color, so it is difficult to separate tree crown and understory vegetation, and it is expensive, fragile, and not convenient for large-scale agricultural application (example 1, information acquisition and navigation method of a fully autonomous fruit tree information acquisition robot: multiple sensors are carried on a four-wheel vehicle platform for positioning, sensing, and navigation, but only radar is used to fit the coordinates of fruit trees, and this method needs to exist auxiliary positioning fruit trees for positioning ordinary tree crowns, so it cannot be universally applied to general orchards. Example 2, ground-based laser point cloud single tree crown volume extraction method based on spherical coordinate integral: laser radar acquires point cloud, separates tree crown and trunk, and estimates tree crown volume, which can only be used for single tree). Inexpensive cameras return environmental information in the form of pictures, which are large in information quantity, consistent with the input of human eyes, and convenient for human interaction. The present method uses camera input. The process of tree crown counting in the form of pictures usually has two steps of identification and statistics.

[0004] There are color space morphology-based recognition and neural network-based recognition in recognition. There are connected component-based statistics and tracking-based statistics in statistics. In the overhead view of the UAV, the boundary between trees can be clearly seen, and in the tracking-based statistics, a video or picture sequence with more than 50% overlap between adjacent frames is usually used. In recognition, the adjustable network structure and the large number of parameters of the neural network mean that it has better performance in complex scenes, while the color space morphology-based method is more friendly to embedded devices in terms of computing resources (for example, a tree fruit counting method based on YCbCr color space). This method does not limit the recognition method.

[0005] Connected component-based statistics requires a photo that covers all the tree crowns in the garden. In a narrow green belt or a wide cultivation forest scene, the panorama is usually generated by stitching or three-dimensional modeling. If the above operations are not performed, only the number of tree crowns within the camera's field of view can be counted, not the number of tree crowns in the entire garden. Simply adding the counting results from different angles will likely result in repeated counting and missed counting (for example, a fruit tree crown recognition and counting method, system, electronic device, and storage medium: an improved YOLOv4 network model uses a single image and two-dimensional coordinates). The distortion of the lens of a cheap camera to the light field is not linear, which causes the tree crowns at the edge of the field of view to deform, making the trees near the stitching line disappear or overlap (for example, a peach orchard fruit status recognition and counting method and system: two-view video stitching, recognition, and tracking).

[0006] However, it requires more time and resources. (Example 1: A method for identifying dead trees in aerial photos using deep learning is described in the patent "Dead Tree Positioning Method and System". This method identifies the two-dimensional coordinates of dead trees in the image through a neural network, and then inputs them into the DOM to obtain the three-dimensional coordinates in the geodetic coordinate system. This method completes the acquisition of tree crown coordinates in two steps, and the DOM in the second step is completed by time-consuming and resource-consuming three-dimensional reconstruction, so it cannot output in real time. Example 2: A tree counting method, system, storage medium, computer device, and terminal: DOM gets the shadow background, DSM gets the light background, and the background is removed before single-tree counting. Example 3: A method and system for identifying and counting fruit trees using UAV images: DOM calculates the vegetation index, identifies the vegetation area and background fruit tree radius, substitutes the preliminary identified fruit tree coordinates and diameter into DSM, and uses the minimum value of the elevation near the fruit tree coordinates and the maximum value of the fruit tree elevation to obtain the fruit tree height.

[0007] The tracking-based statistics is to obtain the two-dimensional coordinates of the tree crown in the continuous pictures, and the tracking of the tree crown is stopped when the tree crown moves out of the field of view or is blocked. Therefore, when the same tree crown appears in the field of view for multiple times but discontinuously, repeated counting is caused (for example, a fruit orchard fruit recognition and yield statistics system and method: acquisition, recognition, tracking, counting). The present method limits the statistics method to be tracking-based, and the innovation point is to introduce a three-dimensional environment map to avoid repeated counting of the same tree crown at different times. The first difference from the “fruit real-time counting method based on visual slam and target detection” is that the initial three-dimensional position of the object in the present method is obtained by triangulation, instead of being directly output by the point cloud of the camera; the second difference is that the object radius in the present method is continuously optimized as an optimization variable through the constraint of multiple observations, instead of being selected as the best one in each iteration of RANSAC; and the third difference is that the target matching (similarity of the target in the two-dimensional image + overlap of the target in the two-dimensional image) and the object fusion (distance of the object in the three-dimensional map) are used to avoid repeated counting in the present method, instead of only the repetition rate in adjacent frames. SUMMARY

[0008] Therefore, the present application proposes a visual-based real-time three-dimensional positioning and counting method for tree crowns. The present method limits the statistics method to be tracking-based, and introduces a three-dimensional environment map to avoid repeated counting of the same tree crown at different times. The initial three-dimensional coordinates of the object in the present method are obtained by triangulation; the object radius in the present method is continuously optimized as an optimization variable through the constraint of multiple observations; and the target matching (similarity of the target in the two-dimensional image + overlap of the target in the two-dimensional image) and the object fusion (distance of the object in the three-dimensional map) are used to avoid repeated counting in the present method.

[0009] In the fruit tree plant protection, the crown positioning of the present method can realize targeted spraying, and the results of the crown radius and counting in the present method can be used to estimate the amount of medicine, so as to realize the saving of the amount of medicine and the reduction of the pollution of the medicine to the environment; in the forestry management, the tree forest map constructed by the present method can be used to record the growth state of the tree crown, evaluate the urban green index, and the positioning results are used to avoid repeated recording of the same tree.

[0010] The present application proposes a visual-based real-time three-dimensional positioning and counting method for tree crowns, which comprises the following steps:

[0011] Step one, sensor input, the sensor input information is a picture sequence, and the sensor is located on a moving carrier;

[0012] Step two, target tracking, used to obtain the number of all targets in the picture sequence and the two-dimensional coordinates of all targets in each frame of picture;

[0013] Step three, feature tracking, is used to realize the extraction of geometric information and the initial pose estimation of the motion carrier in the picture sequence; key frames are extracted from the picture sequence and an environment map is created;

[0014] Step four, local mapping, is used to receive key frames and optimize the initial value of the motion of the motion carrier and the environment map therefrom;

[0015] Step five, loop closure, is used to identify loops and trigger global optimization;

[0016] Step six, global map optimization, is used to avoid the same target being counted repeatedly because it has different target numbers.

[0017] Further, the specific implementation method of step two is to perform target recognition and tracking after obtaining a frame of picture from the camera; if the target in the current frame finds a corresponding target in the previous frame, the same target number is assigned; if the target in the current frame cannot find a corresponding target in the previous picture, a new target number is assigned, and the two-dimensional coordinates and length and width of the recognition box of the target are output at the same time.

[0018] Further, the specific implementation method of step two is that,

[0019] The corner points in each frame of picture are extracted and their descriptors are obtained to obtain geometric information, similar corner points in two frames of picture are associated, and the initial value of the motion of the motion carrier between the two frames of picture is calculated by means of the two-dimensional coordinates of a plurality of pairs of matched corner points; the environment map includes a map point set, a key frame set, and an object set, the key frame set stores all key frames, the key frames are added to the key frame set after being created; the map point set stores all map points, the map points are added to the map point set after the three-dimensional coordinates of the corner points are obtained by triangulation of the corner points; the object set stores all objects, the objects are added to the object set after the three-dimensional coordinates of the targets are obtained by triangulation of the targets.

[0020] Further, the specific implementation method of step three is that,

[0021] Receiving key frames, and optimizing initial values of motion of the motion carrier and an environment map from the key frames; triangulating corners and objects in two pictures under different perspectives to obtain three-dimensional coordinates of the corners and the objects, then adding the former into a map point set and adding the latter into an object set; when two map points have small re-projection distances and similar descriptors, or two objects have consistent numbers, a fusion module is triggered, that is, a newly created object is selected to replace an old object and is added into the environment map; optimizing the map points and the objects observed in a current picture, optimization variables including a motion carrier pose, map point coordinates, object coordinates and a radius, constraint factors including camera observations on the map point coordinates, camera observations on the object coordinates and the radius; after optimization, it is determined whether to add a key frame set or discard the current frame according to a number of key frames in the environment map, relative motion of two frames or a map point repetition rate between the two frames.

[0022] Further, the specific implementation method of step four is,

[0023] Comparing the current frame with all key frames in the key frame set, selecting a plurality of most similar frames and calculating similarity of the most similar frames and the current frame, triggering loop closure when the similarity exceeds a threshold, otherwise waiting for a new key frame to detect again, when the loop closure is triggered, performing feature matching between the current frame and the similar frames, then estimating a similarity transformation between the two frames, and applying the similarity transformation to three-dimensional coordinates of the map points and the objects that can be observed in the current frame, and finally triggering a global map optimization step.

[0024] Further, the specific implementation method of step five is,

[0025] Firstly, an optimization problem including all variables in the environment map is constructed, optimization variables including a motion carrier pose, map point coordinates, object coordinates and a radius, constraint factors including two-dimensional camera observations on the map point coordinates, three-dimensional camera observations on the object coordinates and the radius, after optimization, all objects are traversed, when radii of two objects are both greater than a distance between the object center coordinates, object fusion is performed, and coordinates and a radius of a new object are the mean values of coordinates and a radius of old objects.

[0026] The present application has the following beneficial technical effects:

[0027] Multiple functions: counting while obtaining three-dimensional coordinates and radii of tree crowns;

[0028] Real-time: no need for three-dimensional mapping, positioning and counting while moving;

[0029] Accuracy: counting in a three-dimensional environment map can avoid repeated and missed counting;

[0030] Affordability: as long as a camera is mounted on a drone or a small vehicle, it can also be deployed to a smart phone or a tablet in the form of an APP. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 A schematic diagram illustrating the acquisition of target depth using triangulation.

[0032] Figure 2 This is a flowchart of the present invention. Detailed Implementation

[0033] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0034] The ultimate goal of this method is to obtain the coordinates (3D variable, in meters), radius (1D variable, in meters), and quantity (1D variable, in trees) of an object, typically a tree canopy. Therefore, as... Figure 2 As shown, a data structure called an environment map is maintained, which has two substructures: a common view and an octree. The environment map stores map points (3D variables, N-H coordinates), keyframes (3D variables N-H coordinates and N-H pose, the former two referred to as pose), and objects (3D variables N-H coordinates and one-dimensional variable radius). The set of the first two, representing geometric information, is maintained by the common view, while the set of objects, representing semantic information, is typically maintained by an octree, but can also be other data structures such as raster maps or KD trees. To maintain the environment map, the method consists of seven steps, arranged in the order of data flow: sensor input, target tracking, feature tracking, local mapping, loop closure, global map optimization, and human-computer interaction. Human-computer interaction is mainly used to display the sensor input data and the state of the environment map; it can be omitted in non-interactive scenarios. The method described here can be deployed on specific computing platforms and placed with sensors on drones, vehicles, and handheld platforms, or it can be used as an application on smartphones or tablets.

[0035] The sensor type is a camera. The camera can be monocular or binocular, typically a monocular global shutter camera, and the input format is an image, typically a color image, at 640P resolution.

[0036] The target tracking goal is to obtain the number of all targets in the picture sequence and their two-dimensional coordinates appearing in each frame of picture. The network training module needs to be carried out in advance, using the labeled target data set to adjust the weight of the neural network to a smaller error stage, and the typical labeling tool is labelimg, and the typical training tool is GPU. The target recognition module is to obtain the two-dimensional coordinates of all three-dimensional objects projected into each frame of picture as targets, and the typical tool is YOLOV5. The target matching module is to assign the same target number to the same object appearing in the picture sequence, which is convenient for target fusion in local mapping, and the typical tool is DeepSort. In practice, in order to realize target tracking, it can be divided into two steps, that is, first realize target recognition, and then realize target tracking, or directly track the target sequence. Among them, the target recognition can be completed by the method of color space morphology, or the neural network output recognition result. The flow chart shows a typical two-step process of using neural network recognition and tracking.

[0037] Feature tracking is composed of feature extraction, feature matching and pose estimation modules to realize the extraction of geometric information in the picture and the initial pose estimation of the moving carrier. The purpose of the feature extraction module is to extract some representative corner points in each frame of picture and obtain their descriptors to obtain geometric information. There are many feature extraction and description algorithms to choose from, such as classic SIFT, SURF, FAST, SuperPoint based on neural network, typical Oriented FAST and Rotated BRIEF, and ORB. The purpose of feature matching is to associate similar corner points in two frames of picture, and the typical method is the descriptor matching method based on Hamming distance. The method is as follows: the Hamming distance between two equal-length strings is the number of different characters in the corresponding positions of the two strings, in other words, it is the number of characters that need to be replaced to transform one string into another. The typical descriptor of ORB is a string of 128-bit binary string. The purpose of pose estimation is to calculate the initial value of the carrier between two frames of picture by means of a plurality of pairs of matched corner points, and the typical method is epipolar geometry, which simultaneously calculates the essential matrix and the homography matrix and selects the optimal value from them.

[0038] The local mapping aims to receive keyframes and optimize the initial value of the motion of the carrier and the environment map from them. In the creation module of the map points and objects, first, the corner points and targets in the two frames of pictures under different perspectives are triangulated to obtain their three-dimensional coordinates, and then the former is added to the map point set and the latter is added to the object set. When the re-projection distance of two map points is small and the descriptor hamming distance is similar, or the numbers of two objects are consistent, the fusion module is triggered, that is, the newly created object is selected to replace the old one and added to the environment map. The local map optimization module needs to optimize the map points and objects observed in the current picture, and the optimization variables include the pose of the moving carrier, the coordinates of the map points, the coordinates and radius of the objects, and the constraint factors include the observation of the camera on the map point coordinates, the observation of the camera on the object coordinates and radius, and the optimization framework is typically G2O or Ceres. After optimization, according to the number of keyframes in the environment map, the relative motion of two frames or the repetition rate of map points between two frames, it is determined whether to add the keyframe set or discard the current frame.

[0039] Loop closure is to reduce the cumulative error and avoid repeated counting of the same object due to different target numbers. The loop detection module can detect by picture similarity, compare the current frame with similar pictures in the keyframe set, and trigger loop closure when the similarity exceeds the threshold, otherwise wait for new keyframes to detect again. When it is detected that there is a loop, feature matching is performed between the current frame and the similar frame, and the similarity transformation is calculated. The similarity transformation is applied to the three-dimensional coordinates of the map points and objects in the current frame, and finally the global map optimization is triggered.

[0040] Since global optimization of the environment map often consumes a lot of resources and time, its typical triggering mode is to detect loop closure, light system load, and long time without global optimization. The map optimization module is consistent with the local map optimization module in terms of optimization variable types, constraint factors and optimization framework, but the former includes all elements in the environment map, so it has a larger number of optimization variables. There may be errors in the environment map introduced by tracking interruption in the target tracking step, that is, the same object has different target numbers. Therefore, in order to avoid repeated counting, object fusion is needed when global map optimization is performed. The object fusion module compares the distance between the two object coordinates with a certain threshold, and the object is fused when the latter is larger. The typical threshold is the object radius, and the average of the two objects is the new value.

[0041] The above is the flow of the entire method, and the features of the method are first in the object creation module in the local mapping step, that is, the targets with the same number in the two frames of pictures under different perspectives are triangulated to obtain the three-dimensional coordinates of the objects.

[0042] As Figure 1As shown, the left and right frames are images img1 and img2, respectively, with their optical centers O1 and O2. The normalized coordinates of target p1 in img1 are X1, and the normalized coordinates of target p2 in img2 are X2. The two-dimensional pixel coordinates (u, v) obtained from target tracking and the corresponding normalized coordinates (X1, X2, X2) are... X X Y 1) There exists u = X X ×F X +C X v = X Y ×F Y +C Y The relationship between F X C X F Y C Y These are all internal parameters of the camera, obtained from the camera's factory calibration. From the feature tracking steps, we know that the rotational motion of img2 relative to img1 is R, and the translational motion is t. Therefore, s1×X1=s2×R×X2+t exists. 。 Where s1 and s2 are the distances from object P to O1 and O2, respectively. Solving the equation S2×(X1)×R×X2+(X1)×t=0 yields the distance between object P and O2. In the formula, (M) represents transforming vector M into an antisymmetric matrix. Finally, the three-dimensional coordinates of object P are obtained from the pose R and t of img2.

[0043] The second feature of this method lies in the object fusion module in the local mapping and global map optimization steps, which involves fusing the old 3D coordinates X of two objects with the same number or close proximity. n Weighted average to new coordinate X new , that is, X new =W1×X1+W2×X2+W3×X3+…+W n ×X n In this method, the typical weight W n This represents the number of times the object has been observed consecutively.

[0044] The last two features of this method lie in the map optimization module within the local mapping and global map optimization steps. The third feature is the addition of the object radius as an optimization variable. The target information obtained from target tracking in the image is u, v, w, and h, representing the x-coordinate, y-coordinate, width, and height of the object's projection onto the image, respectively. The equation for measuring the actual object's radius is... sqrt represents the square root, and s is the distance from object P to the optical center O. Therefore, a constraint factor on the object's radius can be added each time a target is tracked.

[0045] A specific embodiment is given as follows, the method is deployed to tablet in APP form to locate, count and estimate the volume of fruit tree crown.

[0046] The sensor input is a sequence of pictures from the rear camera. After a picture is acquired from the camera, the target tracking step is input to recognize and track the fruit tree crown. If the crown in the current frame is found in the previous frame, the same crown number is assigned. If the crown in the current frame is not found in the previous frame, a new crown number is assigned. The two-dimensional coordinates and the length and width of the recognition box of the crown are output at the same time as the crown number.

[0047] In the feature tracking step, the first 800 ORB corners and their corresponding descriptors in the current frame are extracted, and then matched with the descriptors of all corners in the previous frame according to the Hamming distance to obtain the two-dimensional coordinates of the same corner in the two frames. Then the initial estimate of the motion between the two frames is calculated by solving the essential matrix or the homography matrix.

[0048] In the local mapping step, the map points and the fruit tree crowns in the orchard map are created according to the two-dimensional coordinates of the matched corners or crowns in the two frames. Then the map points or crowns with small re-projection distance and similar descriptors or the same number are fused, and the old coordinate mean value is used as the new value to replace the map point set or the crown set. Then the G2O optimization framework is called to jointly optimize the variables that can be observed in the current frame in the environment map. The optimization variables include the tablet pose, the orchard map point coordinates, the crown coordinates and the radius, and the constraint factors include the two-dimensional observation of the map point coordinates by the rear camera, the three-dimensional observation of the crown coordinates and the radius by the rear camera, and the initial value of the crown radius is 1 meter. After optimization, if the number of map points observed by the current frame and the previous frame is less than 80% of the number of all map points observed by the previous frame, the key frame set is added, otherwise it is discarded.

[0049] The loop closure detection module compares the current frame with all key frames in the key frame set, selects the top 5 frames using the DBow method, and calculates their similarity with the current frame. When the similarity exceeds the threshold, the loop closure is triggered, otherwise the new key frame is detected again. When the loop closure is triggered, the feature matching is performed between the current frame and the similar frame in the feature tracking step, then the similarity transformation between the two frames is estimated, and the three-dimensional coordinates of the orchard map points and crowns observed in the current frame are applied, and finally the global map optimization step is triggered.

[0050] The global map optimization step firstly constructs an optimization problem containing all variables in the environment map. The optimization variables include the tablet pose, the orchard map point coordinates, the tree crown coordinates and the radius. The constraint factors include the 2D observation of the map point coordinates by the back camera, the 3D observation of the tree crown coordinates and the radius by the back camera. After the optimization, all the tree crowns are traversed. When the radius of two tree crowns are both greater than the distance between the two tree crown coordinates, target fusion is performed. The coordinate and the radius of the new tree crown are the mean values of the old tree crowns.

[0051] The human-computer interaction interface includes an orchard map interface and a camera interface. The orchard map interface displays the current tablet pose in the map, the coordinates of the orchard map points, the coordinates and size of the tree crowns, and the motion trajectory of the tablet. The camera interface displays the picture taken by the current back camera, and the recognition frame of the tree crown in the current frame obtained by target tracking.

[0052] Those skilled in the art can easily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A vision-based real-time three-dimensional localization and counting method for tree canopies, characterized in that, The method includes the following steps: Step 1: Sensor input. The sensor input information is a sequence of images, and the sensor is located on the moving platform. Step 2, target tracking, is used to obtain the ID of all targets in the image sequence and the two-dimensional coordinates of all targets in each frame of the image; Step 3: Feature tracking, used to extract geometric information from the image sequence and estimate the initial pose of the moving vehicle; extract keyframes from the image sequence and create an environment map; Step 4: Local mapping, used to receive keyframes and optimize the initial values ​​of the motion of the moving vehicle and the environmental map from them; Step 5: Loop Closure, used to identify loop closures and trigger global optimization; Step 6: Global map optimization to avoid duplicate counting of the same target due to different target numbers; The specific implementation method of step five is as follows: First, we construct an optimization problem that includes all variables in the environment map. The optimization variables include the pose of the moving vehicle, the coordinates of map points, the coordinates of the target, and the radius. The constraint factors include the two-dimensional observation of the map point coordinates by the camera and the three-dimensional observation of the target coordinates and radius by the camera. After optimization, we traverse all targets. When the radii of two targets are both greater than the distance between the center coordinates of the two targets, we perform target fusion. The coordinates and radius of the new target are the average of the old targets.

2. The method according to claim 1, characterized in that, The specific implementation method of step two is as follows: after acquiring a frame image from the camera, target recognition and tracking are performed; if the target in the current frame has a corresponding target in the previous frame, the same target number is assigned; if the target in the current frame does not have a corresponding target in the previous image, a new target number is assigned, and the two-dimensional coordinates and length and width of the target's recognition box are output at the same time as the target number.

3. The method according to claim 2, characterized in that, The specific implementation method of step two is as follows: Corner points are extracted from each frame of the image, and their descriptors are obtained to acquire geometric information. Similar corner points in two frames of the image are associated, and the initial motion values ​​of the moving vehicle between the two frames are calculated using the two-dimensional coordinates of multiple pairs of matched corner points. The environment map includes a map point set, a keyframe set, and an object set. The keyframe set stores all keyframes, and keyframes are added to the keyframe set after creation. The map point set stores all map points. Map points are added to the map point set after the three-dimensional coordinates of the map points corresponding to the corner points are obtained by triangulation of the corner points. The object collection stores all objects. After the object is triangulated from the target to obtain the three-dimensional coordinates of the object corresponding to the target, it is added to the object collection.

4. The method according to claim 3, characterized in that, The specific implementation method of step three is as follows: The system receives keyframes and optimizes the initial values ​​of the motion of the moving vehicle and the environment map from them. It performs triangulation on corner points and targets in two frames from different viewpoints to obtain their 3D coordinates. The former is then added to the map point set, and the latter to the object set. If the reprojection distance between two map points is small and their descriptors are similar, or if two objects have the same number, the fusion module is triggered, replacing the old one with the newly created one and adding it to the environment map. The system optimizes the map points and objects observed in the current image. Optimization variables include the moving vehicle pose, map point coordinates, object coordinates, and radius. Constraint factors include the camera's observation of map point coordinates, and the camera's observation of object coordinates and radius. After optimization, the system decides whether to add the current frame to the keyframe set or discard it based on the number of keyframes in the environment map, the relative motion between the two frames, or the map point repetition rate between the two frames.

5. The method according to claim 4, characterized in that, The specific implementation method for step four is as follows: The current frame is compared with all keyframes in the keyframe set. Multiple most similar frames are selected and their similarity to the current frame is calculated. When the similarity exceeds a threshold, loop closure is triggered. Otherwise, the system waits for a new keyframe to be detected again. When loop closure is triggered, feature matching is performed between the current frame and similar frames. Then, the similarity transformation between the two frames is estimated and applied to the 3D coordinates of map points and targets that can be observed in the current frame. Finally, the global map optimization step is triggered.

Citation Information

Patent Citations

  • Monocular vision SLAM (Simultaneous Localization and Mapping) method for dynamic environment

    CN115471748A