A method for dynamic environment perception based on a binocular camera

Through the dynamic environment perception method based on binocular camera, deep images are processed to generate sparse point clouds, cluster and association, and identify dynamic or static attributes of obstacles, solving the problem of obstacle perception in complex environments by drones, and achieving fast and accurate dynamic environment perception.

CN114387462BActive Publication Date: 2025-06-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111652247.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-13
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

The existing drone obstacle perception technology has harsh usage conditions and poor adaptability, and is not suitable for use in drones, especially in complex dynamic environments, and it is difficult to accurately identify dynamic and static obstacles.

Method used

Using a dynamic environment perception method based on a binocular camera, acquiring multi-frame depth images, processing them into sparse point clouds, clustering and association, identifying the dynamic or static attributes of obstacles, and dynamically updating the environment information.

Benefits of technology

Quickly and accurately detect and identify obstacles in complex environments, realize dynamic environmental perception, suitable for drones, without the need to be equipped with peripherals such as lidar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387462B_ABST
    Figure CN114387462B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic environment perception method based on a binocular camera, which includes steps of acquiring multiple frames of depth images captured by the binocular camera, processing each frame of the depth images to obtain multiple frames of sparse point clouds, clustering each frame of the sparse point clouds respectively to obtain multiple clustering clusters, associating the multiple clustering clusters representing the same obstacle in the sparse point clouds of different frames, and identifying whether the obstacle corresponding to the associated clustering cluster is a dynamic obstacle or a static obstacle. The present invention can quickly and preferably realize the dynamic environment perception of detecting, dividing the dynamic and static attributes, and tracking obstacles of various shapes in the environment under complex environments such as when the distance between obstacles is relatively close, and can make a relatively robust judgment on the dynamic and static attributes of obstacles in complex situations where obstacles are occluded and the camera itself moves. The unmanned aerial vehicle applying the present invention has advantages of small load, low power consumption, being light and easy to carry. The present invention is widely applied to the technical field of image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a dynamic environment perception method based on a binocular camera. Background Art

[0002] During the flight of a drone, it may encounter obstacles. In order to successfully complete the flight mission, it is necessary to perceive the obstacles and perform actions such as avoidance. In a real flight environment, the drone often faces unknown and complex dynamic obstacle scenarios, where the obstacles may be either dynamic obstacles or static obstacles. Since the coping measures of the drone for dynamic obstacles and static obstacles are different, it is required that the drone can accurately perceive whether the obstacle is a dynamic obstacle or a static obstacle. Currently, relatively advanced obstacle perception technologies include the obstacle perception technology that uses a high-precision lidar to generate point cloud data, and the technology that uses visual information from images to perceive obstacles. The obstacle perception technology that uses a high-precision lidar to generate point cloud data relies on the lidar to collect point cloud data, and the volume and mass of the lidar are relatively large, which is not convenient for the drone to carry and fly, thus limiting its application in drones. In the current technologies that use visual information from images to perceive obstacles, there are generally disadvantages such as harsh usage conditions and poor adaptability. For example, the technology that uses the Frame-Difference method to perceive dynamic obstacles can only be applied when the drone hovers stably, and the technology that uses neural networks to perceive obstacles only performs well in perceiving predefined types of obstacles, but has poor effects on general unknown types of obstacles. Summary of the Invention

[0003] Aiming at at least one technical problem such as the harsh usage conditions, poor adaptability, and inapplicability to drones of the current obstacle perception technologies, the purpose of the present invention is to provide a dynamic environment perception method based on a binocular camera, including:

[0004] Obtaining multiple frames of depth images captured by a binocular camera;

[0005] Processing each frame of the depth images to obtain multiple frames of sparse point clouds; wherein, the processing result of one frame of the depth images is to obtain a corresponding frame of sparse point cloud;

[0006] Clustering each frame of the sparse point clouds respectively to obtain multiple clustering clusters; wherein, the clustering result of one frame of the sparse point clouds is to obtain corresponding several clustering clusters, and different clustering clusters in one frame of the sparse point clouds respectively represent different obstacles;

[0007] Associating multiple clustering clusters representing the same obstacle in the sparse point clouds of different frames;

[0008] Identify whether the obstacle corresponding to the associated clustering cluster is a dynamic obstacle or a static obstacle.

[0009] Further, the binocular camera-based dynamic environment perception method further includes:

[0010] When the obstacle corresponding to the clustering cluster is a static obstacle, update the obstacle information on the occupancy map according to the clustering cluster;

[0011] When the obstacle corresponding to the clustering cluster is a dynamic obstacle, model the clustering cluster into an ellipsoid, and use a Kalman filter to track the modeled clustering cluster.

[0012] Further, the binocular camera-based dynamic environment perception method further includes:

[0013] When the modeled clustering cluster is not tracked within a continuous time period exceeding the threshold length, end the tracking of the clustering cluster and delete the data corresponding to the clustering cluster.

[0014] Further, the processing of each frame of the depth image to obtain multiple frames of sparse point clouds includes:

[0015] Obtain the extrinsic matrix T and the intrinsic matrix K of the binocular camera;

[0016] Obtain the pixel coordinates P of the depth image uv ;

[0017] Determine the original point cloud through the formula P w = T -1 K -1 P uv where P w is the world coordinate of the original point cloud;

[0018] Crop the original point cloud to obtain a dense point cloud;

[0019] Filter the dense point cloud using the voxel filtering method to obtain the sparse point cloud.

[0020] Further, the clustering of each frame of the sparse point cloud respectively to obtain multiple clustering clusters includes:

[0021] A1. Obtain a sample point set D = {x 1 , x 2 , ……, x m}, where x m represents the m-th point in the sparse point cloud, set the neighborhood distance threshold ε and the connectivity threshold δ, and initialize the core object set Initialize the number of clustering clusters k = 0, initialize the unvisited sample set F = D, and initialize the cluster partition

[0022] A2. For j = 1, 2, ……, m, find all core objects according to the following steps A2a - A2b:

[0023] A2a. Find the subset of sample points N j within the ε-neighborhood of the sample x ε (x j ); where x j is a sample point in the sample point set D;

[0024] A2b. If the number of sample points in the subset of sample points N ε (x j ) satisfies |N ε (x j )| ≥ MinPts, then calculate the number of connected components n of the point Q = N ε (x j ) ∪ x j ; if n < δ, add the sample x j to the core object set Ω through the formula Ω = Ω ∪ {x j};

[0025] A3. If the core object set then end the execution of steps A1 - A6, otherwise execute step A4;

[0026] A4. In the core object set Ω, randomly select a core object o, initialize the current cluster core object list Ω cur = {o}, initialize the class serial number k = k + 1, initialize the current cluster sample set C k = {o}, and update the unvisited sample set F - F - {o};

[0027] A5. If the current cluster core object queue then the current clustering cluster C k is generated, update the cluster partition C = {C 1 , C 2 , ……, C k}, update the core object set Ω = Ω - C k , and return to execute step A3, otherwise update the core object set Ω = Ω - C k ;

[0028] A6. Take out a core object o'' from the current cluster core object queue Ω cur , find the ε-neighborhood subset of sample points N ε (o''), let Δ = N ε (o'') ∩ F, and update the current cluster sample set C k = Ck ∪Δ, update the unvisited sample set F = F - Δ, and update Ω cur = Ω cur ∪(Δ ∩ Ω) - o'', and return to execute step A5.

[0029] Further, the associating of multiple clustering clusters representing the same obstacle in the sparse point clouds of different frames includes:

[0030] B1. Obtain m clustering clusters where t represents C t Each clustering cluster in is obtained by clustering the sparse point cloud with the acquisition time t, and the sparse point cloud with the acquisition time t is obtained by processing the depth image with the shooting time t. Predict the positions of all obstacles in C t at time t Set a distance threshold ε, and initialize the association set F = K t ;

[0031] B2. Calculate the centroid of each clustering cluster in C t to obtain the centroids of all clustering clusters in the current frame

[0032] B3. Initialize the set to be associated Ω = D t ;

[0033] B4. For each execute according to the following steps B4a - B4c:

[0034] B4a. Find the nearest neighbor k of in F j ;

[0035] B4b. Find the nearest neighbor d of k j in Ω j ;

[0036] B4c. If i.e., and k j are nearest neighbors to each other, through the formula associate to obstacle j; through the formula remove from the set to be associated Ω; through the formula F = F - {k j} remove k j from the association set F;

[0037] B5. If or End the execution of steps B1 - B6. Conversely, for each Execute the following steps B5a - B5b:

[0038] B5a. Find the nearest neighbor k in F j , calculate the distance δ between j k;

[0039] B5b. If δ < ε, through the formula associate with the obstacle j, and through the formula remove from the set Ω to be associated. Remove k from the associated set F through the formula F = F - {k j}; j

[0040] B6. If End the execution of steps B1 - B6. Conversely, for each Consider it as a newly emerged obstacle and establish an obstacle tracking history, which is represented as where Δt represents the time interval and n represents the label of the obstacle.

[0041] Furthermore, identifying whether the obstacles corresponding to the respective clustering clusters are dynamic obstacles or static obstacles includes:

[0042] C1. Obtain the depth image D taken at time t - Δt t-Δt and the depth image D taken at time t t , obtain the dense point cloud collected at time t - Δt where the dense point cloud collected at time t - Δt is obtained by processing the depth image taken at time t - Δt. Obtain the pose O of the UAV where the binocular camera is located at time t - Δt t-Δt and the pose O at time t t , obtain where p represents a point in the world coordinate system, l represents the label of the point, and obtain the discrimination parameters (β, β min , V min );

[0043] C2. Traverse any clustering cluster in C t where I = 1, 2, ……, m. If the clustering cluster corresponds to a newly identified obstacle, the clustering cluster ​​The corresponding obstacle is identified as an unknown obstacle, where the unknown obstacle is an obstacle that is neither a dynamic obstacle nor a static obstacle;

[0044] C3. Initialize variables vote = 0, dyn = 0, static = 0;

[0045] C4. Clustering Each point p in performs the following voting process:

[0046] C4a. According to the posture O t-Δt Determine whether point p is within the field of view of the drone at time t-Δt, and determine whether point p is blocked at time t-Δt; when point p is not within the field of view of the drone at time t-Δt, and point p is blocked at time t-Δt, operate the variable vote according to the formula vote=vote+1, otherwise execute step C4b;

[0047] C4b. In dense point cloud Find the nearest neighbor nn of point p; when Where kength(p,nn) is the distance between point p and its nearest neighbor nn, Δt represents the time interval, and the variable dyn is operated according to the formula dyn=dyn+1, otherwise, step C4c is executed;

[0048] C4c. Operate the variable static according to the formula static=static+1;

[0049] C5. When Clustering The corresponding obstacles are identified as dynamic obstacles, otherwise the clusters are The corresponding obstacle is identified as a static obstacle; Represents clusters The center of mass, express With O t The distance between range Indicates the maximum measurement distance of the binocular camera.

[0050] Further, the determining whether the point p is blocked at the time t-Δt includes:

[0051] D1. Get c t The centroid of t 、c t-Δt The centroid of t-Δt , and the judgment parameter β(Δt) related to time Δt;

[0052] D2. According to O t Put cent Project it onto the pixel coordinate system at time t to obtain q t ;

[0053] D3. According to O t-Δt Project cen t-Δt onto the pixel coordinate system at time t - Δt to obtain q t-Δt ;

[0054] D4. Calculate the average depth roundAvg(D t within the first range around q t [q t );

[0055] D5. Calculate the average depth roundAvg(D t-Δt within the first range around q t-Δt [q t-Δt );

[0056] D6. If roundAvg(D t [q t ) - roundAvg(D t-Δt [q t-Δt ) > β(Δt), it is determined that point p is self-occluded at time t - Δt, otherwise it is determined that point p is not self-occluded at time t - Δt;

[0057] D7. According to O t-Δt Project point p onto the camera coordinate system at time t - Δt to obtain the depth from the position of the binocular camera at time t - Δt to point p where |.| z represents the coordinate value in the z direction;

[0058] D8. According to O t-Δt Project point p onto the pixel coordinate system q at time t - Δt t-Δt ;

[0059] D9. Calculate the average depth roundAvg(D t-Δt within the second range around q t-Δt [q t-Δt );

[0060] D10. If it is determined that point p is occluded by other points at time t - Δt, otherwise it is determined that point p is not occluded by other points at time t - Δt;

[0061] D11. If it is determined in step D6 that the object where point p is located is not self-occluded at time t - Δt, and it is determined in step D10 that point p is occluded by other points at time t - Δt, then it is determined that point p is occluded at time t - Δt.

[0062] Further, identifying whether the obstacles corresponding to the respective clustering clusters are dynamic obstacles or static obstacles includes:

[0063] E1. Obtain the centroid cen t of c t , the centroid cen t-Δt of c t-Δt , the speed threshold parameter V min , and the compensated motion threshold parameter β;

[0064] E2. Determine whether the obstacle corresponding to the clustering cluster c t needs to perform self - motion compensation of the binocular camera. If so, jump to step E3. Otherwise, calculate the approximate speed v of the obstacle corresponding to the clustering cluster c t as v = Length(cen t , en t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to the clustering cluster c t as a dynamic obstacle. Conversely, identify the obstacle corresponding to the clustering cluster c t as a static obstacle;

[0065] E3. Calculate the compensation vector V t of the change in the centroid of the obstacle corresponding to the binocular camera motion relative to the clustering cluster c c , and calculate the actual movement vector V t of the centroid of the obstacle corresponding to the clustering cluster c a ; if difference(V c , V a ) < β, where difference() represents the modulus of the difference between two vectors, then determine that the object is a static obstacle and end the execution of steps E1 - E4. Conversely, jump to step E4;

[0066] E4. Calculate the speed v of the obstacle corresponding to the clustering cluster c t as v = Length(cen t , cen t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to the clustering cluster c t as a dynamic obstacle. Conversely, identify the obstacle corresponding to the clustering cluster c t as a static obstacle.

[0067] Further, identifying whether the obstacles corresponding to the respective clustering clusters are dynamic obstacles or static obstacles includes:

[0068] F1. Obtain the time scale t h, attenuation factor α and obstacle tracking history where Δt represents the time interval and k represents the sequence number;

[0069] F2. Initialize variables Weight = 1, DynWeight = 0, StaticWeight = 0, unknownWeight = 0;

[0070] F3. Set variable I, and I ranges from 0 to t h Traverse and execute the following steps F3a - F3d:

[0071] F3a. If the obstacle corresponding to cluster c t-i*Δt belongs to a dynamic obstacle, operate on variable DynWeight according to the formula DynWeight = DynWeight + Weight;

[0072] F3b. If the obstacle corresponding to cluster c t-i*Δt belongs to a static obstacle, operate on variable StaticWeight according to the formula StaticWeight = StaticWeight + Weight;

[0073] F3c. If the obstacle corresponding to cluster c t-i*Δt belongs to an unknown obstacle, operate on variable unknownWeight according to the formula unknownWeight = unknownWeight + Weight;

[0074] F3d. Operate on variable Weight according to the formula Weight = Weight * α;

[0075] F4. If DynWeight > StaticWeight and DynWeight > unknownWeight, identify the obstacle corresponding to cluster c t as belonging to a dynamic obstacle, otherwise execute step F5;

[0076] F5. If StaticWeight > DynWeight and StaticWeight > unknownWeight, identify the obstacle corresponding to cluster c t as belonging to a static obstacle, otherwise execute step F6;

[0077] F6. Identify the obstacle corresponding to cluster c t as belonging to an unknown obstacle.

[0078] The beneficial effects of the present invention are as follows: The dynamic environment perception method based on a binocular camera in the embodiment can, by processing the depth images collected by the binocular camera, quickly and preferably achieve dynamic environment perception of detecting, classifying the dynamic and static attributes of, and tracking obstacles of various shapes in the environment under complex conditions such as when the distance between obstacles is relatively close, when obstacles are occluded, and when the camera itself is in motion. It can make a relatively robust judgment on the dynamic and static attributes of obstacles in complex situations where obstacles are occluded and the camera itself is in motion. The dynamic environment perception method based on a binocular camera in the embodiment only requires a binocular camera to collect image data for processing, without the need to be equipped with external devices such as lidar. The unmanned aerial vehicle applying the dynamic environment perception method based on a binocular camera in the embodiment has the advantages of small load, low price, small volume, low power consumption, light weight and easy portability. Therefore, the dynamic environment perception method based on a binocular camera in the embodiment is a method suitable for application to unmanned aerial vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is a flowchart of the dynamic environment perception method based on a binocular camera in the embodiment;

[0080] Figure 2 is an algorithm framework diagram of the dynamic environment perception method based on a binocular camera in the embodiment;

[0081] Figure 3 is a schematic diagram of the first basic situation where self-motion compensation of the binocular camera is required in the embodiment;

[0082] Figure 4 is a schematic diagram of the second basic situation where self-motion compensation of the binocular camera is required in the embodiment;

[0083] Figure 5 and Figure 6 is a schematic diagram of the simulation perception effect obtained by simulating and running the dynamic environment perception method based on a binocular camera in the embodiment;

[0084] Figure 7 is a schematic diagram of the actual machine perception effect obtained by actually running the dynamic environment perception method based on a binocular camera in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] In this embodiment, the dynamic environment perception method based on a binocular camera can be executed by an unmanned aerial vehicle equipped with a binocular camera. Specifically, each step in the dynamic environment perception method based on a binocular camera can be executed by the CPU in the unmanned aerial vehicle.

[0086] Referring to Figure 1 , the dynamic environment perception method based on a binocular camera includes the following steps:

[0087] S1. Obtain multiple frames of depth images captured by the binocular camera;

[0088] S2. Process each frame of depth image to obtain multiple frames of sparse point clouds. Among them, the processing result of one frame of depth image is to obtain the corresponding one frame of sparse point cloud.

[0089] S3. Cluster each frame of sparse point clouds respectively to obtain multiple clusters. Among them, the clustering result of one frame of sparse point cloud is to obtain the corresponding several clusters, and different clusters in one frame of sparse point cloud represent different obstacles respectively.

[0090] S4. Associate multiple clusters representing the same obstacle in sparse point clouds of different frames.

[0091] S5. Identify whether the obstacle corresponding to the associated cluster belongs to a dynamic obstacle or a static obstacle.

[0092] S6. When the obstacle corresponding to the cluster belongs to a static obstacle, update the obstacle information on the occupancy map according to the cluster. When the obstacle corresponding to the cluster belongs to a dynamic obstacle, model the cluster into an ellipsoid and use the Kalman filter to track the modeled cluster. When the modeled cluster is not tracked within a continuous time period exceeding the threshold length, end the tracking of the cluster and delete the data corresponding to the cluster.

[0093] Among them, steps S1 and S2 belong to the point cloud generation steps, step S3 belongs to the point cloud clustering step, step S4 belongs to the obstacle association step, step S5 belongs to the obstacle attribute division step, and step S6 belongs to the dynamic environment information update step. The algorithm framework of steps S1 - S6 is as Figure 2 shown.

[0094] In step S1, multiple frames of depth images are captured by the binocular camera on the UAV. The capture times of two adjacent frames of depth images can be equal, that is, the time difference between the capture times of two adjacent frames of depth images can be a fixed time interval Δt. In this embodiment, based on the capture time t, the frame of depth image with the capture time t is called the current frame of depth image. Correspondingly, in step S2, the one frame of sparse point cloud obtained by processing the current frame of depth image is the current frame of sparse point cloud. Then, the frame of depth image with the capture time t - Δt is called the previous frame of depth image. Correspondingly, in step S2, the one frame of sparse point cloud obtained by processing the current frame of depth image is the previous frame of sparse point cloud.

[0095] In step S2, each frame of depth image is processed to obtain multiple frames of sparse point clouds. Specifically, for each frame of depth image processed, the corresponding one frame of sparse point cloud is obtained. Therefore, there is a one-to-one correspondence between multiple frames of depth images and the multiple frames of sparse point clouds obtained by processing them.

[0096] When performing step S2, that is, processing each frame of depth image to obtain multiple frames of sparse point clouds, the following steps can be specifically executed:

[0097] S201. Obtain the extrinsic parameter matrix T and the intrinsic parameter matrix K of the binocular camera;

[0098] S202. Obtain the pixel coordinates P of the depth image uv ;

[0099] S203. Determine the original point cloud through the formula P w = T -1 K -1 P uv ; where P w is the world coordinate of the original point cloud;

[0100] S204. Crop the original point cloud to obtain a dense point cloud;

[0101] S205. Filter the dense point cloud using the voxel filtering method to obtain a sparse point cloud.

[0102] Among them, steps S201 - S205 are for processing one frame of depth image to obtain the corresponding one frame of sparse point cloud. When there are multiple frames of depth images, steps S201 - S205 are respectively executed for each frame of depth image.

[0103] In step S203, T -1 represents the inverse matrix of the extrinsic parameter matrix T, and K -1 represents the inverse matrix of the intrinsic parameter matrix K. Through the formula P w = T -1 K -1 P uv , the coordinates P w of the original point cloud in the world coordinate system can be calculated. In this embodiment, the original point cloud can be represented by the world coordinate P w .

[0104] In step S204, the points with relatively low credibility in the original point cloud P w (such as the depth is less than D min or greater than D max , both D min and D max are preset thresholds), as well as the ground points (points with a height less than H min ) and ceiling points (points with a height greater than H max ) are cropped off. The points that are not cropped off (depth greater than D min and less than D max , height greater than H min and less than H max ) form a dense point cloud.

[0105] In step S205, the dense point cloud is filtered using the voxel filtering method, and a sparser point cloud than the dense point cloud, i.e., a sparse point cloud, can be obtained. In subsequent steps such as S3 - S6, the sparse point cloud rather than the dense point cloud is processed, which can reduce the amount of data to be processed and thus improve the processing speed.

[0106] When performing step S3, that is, clustering each frame of the sparse point cloud separately to obtain multiple clustering clusters, the following steps A1 - A6 are specifically executed:

[0107] A1. Obtain a sample point set D = {x 1 , x 2 , ……, x m}, where x m represents the m-th point in the sparse point cloud. Set the neighborhood distance threshold ε and the connectivity threshold δ, and initialize the core object set Initialize the number of clustering clusters k = 0, initialize the unvisited sample set F = D, and initialize the cluster partition

[0108] A2. For j = 1, 2, ……, m, find all core objects according to the following steps A2a - A2b:

[0109] A2a. Find the subset of sample points N j within the ε-neighborhood of the sample x ε (x j ); where x j is a sample point in the sample point set D;

[0110] A2b. If the number of sample points in the subset of sample points N ε (x j ) satisfies |N ε (x j )| ≥ MinPts, then calculate the number of connected components n of the point Q = N ε (x j ) ∪ x j ; if n < δ, add the sample x j to the core object set Ω through the formula Ω = Ω ∪ {x j};

[0111] A3. If the core object set then end the execution of steps A1 - A6, otherwise execute step A4;

[0112] A4. In the core object set Ω, randomly select a core object o, initialize the current cluster core object list Ω cur = {o}, initialize the class serial number k = k + 1, and initialize the current cluster sample set C k={o}, update the unvisited sample set F - F - {o};

[0113] A5. If the current cluster core object queue then the current clustering cluster C k is generated, update the cluster partition C = {C 1 , C 2 , ……, C k}, update the core object set Ω = Ω - C k , return to execute step A3, otherwise update the core object set Ω = Ω - C k ;

[0114] A6. Take out a core object o'' from the current cluster core object queue Ω cur and find the ε-neighborhood sub-sample point set N ε (o'') through the neighborhood distance threshold ε, let Δ = N ε (o'') ∩ F, update the current cluster sample set C k = C k ∪Δ, update the unvisited sample set F = F - Δ, update Ω cur = Ω cur ∪(Δ ∩ Ω) - o'', return to execute step A5.

[0115] Among them, steps A1 - A6 are for clustering a frame of sparse point cloud to obtain clustering clusters. When there are multiple frames of sparse point cloud, steps A1 - A6 are respectively executed for each frame of sparse point cloud.

[0116] The result of executing steps A1 - A6 is to cluster each point in a frame of sparse point cloud into n clustering clusters (for different frames of sparse point cloud, the specific value of n is different). Among the multiple clustering clusters obtained by clustering the same frame of sparse point cloud, each clustering cluster respectively represents an obstacle in the corresponding frame of depth image captured by the binocular camera. Therefore, each clustering cluster reflects the information of the obstacle in the sparse point cloud of its corresponding frame, and the clustering clusters obtained by clustering different frames of sparse point cloud may represent the same obstacle or different obstacles. Steps A1 - A6 also have a good clustering effect on obstacles closer to the binocular camera.

[0117] When executing step S4, that is, the step of associating multiple clustering clusters representing the same obstacle in different frames of sparse point cloud, the following steps B1 - B6 are specifically executed:

[0118] B1. Obtain m clustering clusters where t represents that each clustering cluster in C t is obtained by clustering the sparse point cloud with the acquisition time of t, and the sparse point cloud with the acquisition time of t is obtained by processing the depth image with the shooting time of t. Predict C through the Kalman filtert The positions of all obstacles at time t Set a distance threshold ε and initialize the association set F = K t ;

[0119] B2. Calculate C t For each cluster Centroid Obtain the centroids of all clusters in the current frame

[0120] B3. Initialize the set to be associated Ω = D t ;

[0121] B4. For each Execute according to the following steps B4a - B4c:

[0122] B4a. Find The nearest neighbor k in F j ;

[0123] B4b. Find the nearest neighbor d of k j In Ω j ;

[0124] B4c. If That is And k j Are nearest neighbors to each other, through the formula Associate To obstacle j; through the formula Remove From the set to be associated Ω; through the formula F = F - {k j} Remove k j From the association set F;

[0125] B5. If Or End the execution of steps B1 - B6. Otherwise, for each Execute the following steps B5a - B5b:

[0126] B5a. Find The nearest neighbor k in F j , calculate The distance δ between and k j ;

[0127] B5b. If δ < ε, through the formula Associate To obstacle j, through the formula Remove From the set to be associated Ω, through the formula F = F - {k j} Remove k jRemove the associated set F;

[0128] B6. If End the execution of steps B1 - B6. Conversely, for each Consider it as a newly emerged obstacle and establish an obstacle tracking history, which is represented as Where Δt represents the time interval, specifically, it can be the time interval between two frames of depth images collected by the binocular camera, and n represents the label of the obstacle.

[0129] Among them, steps B1 - B6 are for the sparse point cloud at the acquisition time t (i.e., the current frame sparse point cloud), and the current frame sparse point cloud is obtained by processing the depth image at the shooting time t (i.e., the current frame depth image).

[0130] Steps B1 - B6 are a clustering cluster association algorithm based on Euclidean distance. By executing steps B1 - B6, each clustering cluster in the current frame sparse point cloud can be associated with the clustering clusters in the sparse point clouds of the previous frames of the current frame sparse point cloud. If the clustering clusters in different frames of sparse point clouds are associated, it means that the obstacles represented by these clustering clusters are the same obstacle. From another perspective, for a clustering cluster in the current frame sparse point cloud, it itself represents the information such as the position of an obstacle in the current frame depth image. After associating this clustering cluster with the corresponding clustering clusters in the previous frames of sparse point clouds, the corresponding clustering clusters in the previous frames of sparse point clouds respectively represent the information such as the position of the same obstacle in the previous frames of depth images, which is equivalent to endowing this clustering cluster in the current frame sparse point cloud with historical information.

[0131] In this embodiment, when identifying whether the obstacles corresponding to each clustering cluster are dynamic obstacles or static obstacles, two different discrimination algorithms can be specifically used. The first is the depth - adaptive discrimination algorithm based on point voting, and the second is the camera self - motion compensation discrimination algorithm based on the object centroid. Finally, in order to improve the robustness of the discrimination, it is necessary to ensure the historical consistency of the discrimination results. Therefore, a discrimination algorithm with a greater weight for closer frames based on historical consistency is used to finally determine whether the obstacle is a dynamic obstacle or a static obstacle.

[0132] When applying the depth - adaptive discrimination algorithm based on point voting, when executing step S5, that is, the step of identifying whether the obstacles corresponding to each clustering cluster are dynamic obstacles or static obstacles, the following steps C1 - C5 are specifically executed:

[0133] C1. Obtain the depth image D at the shooting time t - Δt t-Δt and the depth image D at the shooting time t t , and obtain the dense point cloud at the acquisition time t - Δt The dense point cloud with the acquisition time of t - Δt is obtained from the depth image with the shooting time of t - Δt, and the pose O of the UAV where the binocular camera is located at the moment of t - Δt is obtained t-Δt and the pose O at the moment of t t , and obtain where p represents the point in the world coordinate system, l represents the label of the point, and the discrimination parameters (β, β min , V min ) are obtained;

[0134] C2. Traverse any clustering cluster in C t where I = 1, 2,..., m. If the clustering cluster corresponds to a newly recognized obstacle, the obstacle corresponding to the clustering cluster is recognized as belonging to an unknown obstacle, and the unknown obstacle is an obstacle that is neither determined to belong to a dynamic obstacle nor determined to belong to a static obstacle;

[0135] C3. Initialize the variables vote = 0, dyn = 0, static = 0;

[0136] C4. Perform the following voting process for each point p in the clustering cluster :

[0137] C4a. According to the pose O t-Δt judge whether the point p is within the field of view angle of the UAV at the moment of t - Δt, and judge whether the point p is occluded at the moment of t - Δt; when the point p is not within the field of view angle of the UAV at the moment of t - Δt and the point p is occluded at the moment of t - Δt, operate on the variable vote according to the formula vote = vote + 1, otherwise execute step C4b;

[0138] C4b. Search for the nearest neighbor nn of the point p in the dense point cloud ; when where length(p, nn) is the distance between the point p and the nearest neighbor nn, and Δt represents the time interval, operate on the variable dyn according to the formula dyn = dyn + 1, otherwise execute step C4c;

[0139] C4c. Operate on the variable static according to the formula static = static + 1;

[0140] C5. When the obstacle corresponding to the clustering cluster is recognized as belonging to a dynamic obstacle, otherwise the obstacle corresponding to the clustering cluster is recognized as belonging to a static obstacle; where represents the clustering cluster ​The centroid of denotes the distance between t and O, and sensing range denotes the maximum measurement distance of the binocular camera.

[0141] The principle of the depth adaptive discrimination algorithm based on point voting performed in steps C1 - C5 is as follows: Let all the points that make up the obstacle find its nearest neighbor in the previous frame of point cloud, so as to calculate the moving speed of this point. If the moving speed of this point is greater than the set threshold, then this point votes that the obstacle it is in is a dynamic obstacle, otherwise it votes that it is a static obstacle. And only those points that appear in the previous frame of FOV and are not blocked by other objects can participate in the voting (being blocked by itself is also allowed to vote). When the proportion of points voting for a dynamic obstacle exceeds the set threshold, then this obstacle is considered a dynamic obstacle, otherwise it is a static obstacle. This voting threshold will be adaptively adjusted according to the distance (depth) of the object from the camera. Because the farther the object is from the camera, the fewer the point clouds projected into the world coordinate system, the greater the possibility of being interfered by noise, and the greater the noise of the camera at greater depths. At this time, the voting threshold needs to be increased accordingly.

[0142] The depth adaptive discrimination algorithm based on point voting performed in steps C1 - C5 can be represented by the following pseudocode:

[0143] Input: The dense point cloud that has not been filtered in the previous frame The pose O of the previous frame of the UAV system t-Δt and the pose O of the current frame t ; All clustered obstacles in the current frame where (p represents the point in the world coordinate system), discrimination parameters (β, β min , V min ); The depth map D of the previous frame t-Δt and the depth map D of the current frame t .

[0144] Output: The attributes (dynamic / static / UNKNOWN) of all obstacles in the current frame.

[0145] Algorithm steps:

[0146]

[0147] "The point p is occluded at time t - Δt" in step C4a specifically means that the object where point p is located is not self-occluded at time t - Δt, and point p is occluded by other points at time t - Δt. That is to say, when both "the point p is not self-occluded at time t - Δt" and "it is determined that the point p is occluded by the points of other objects at time t - Δt" occur, it can be considered that "the point p is occluded at time t - Δt" in step C4a has occurred.

[0148] Based on the above principle, when performing the step of determining whether point p is occluded at time t - Δt in step C4a, the following steps can be specifically executed:

[0149] D1. Obtain the centroid cen of c t and the centroid cen of c t , as well as the judgment parameter β(Δt) related to the time Δt; t-Δt and the centroid cen of c t-Δt , and the judgment parameter β(Δt) related to the time Δt;

[0150] D2. Project cen t onto the pixel coordinate system at time t according to O t to obtain q t ;

[0151] D3. Project cen t-Δt onto the pixel coordinate system at time t - Δt according to O t-Δt to obtain q t-Δt ;

[0152] D4. Calculate the average depth roundAvg(D t [q t ) within the first range around q t ;

[0153] D5. Calculate the average depth roundAvg(D t-Δt ) within the first range around q t-Δt [q t-Δt ;

[0154] D6. If roundAvg(D t [q t ) - roundAvg(D t-Δt [q t-Δt ) > β(Δt), it is determined that point p is self-occluded at time t - Δt, otherwise it is determined that point p is not self-occluded at time t - Δt.

[0155] When some points of an object are occluded by other objects in the previous frame, their positions cannot be found in the point cloud of the previous frame. Therefore, the movement speed of these points cannot be calculated, and the dynamic or static attributes of the object cannot be determined based on these points. As a result, these points cannot participate in the voting. The principle for determining whether a point is occluded is mainly as follows: If there are points in the previous frame that are closer to the camera than point p (there are points with smaller depth at the same position in the previous frame), then point p may be occluded. Based on the above principle, the following steps D7 - D10 can be executed:

[0156] D7. According to O t-Δt Project point p onto the camera coordinate system at time t - Δt Obtain the depth from the position of the binocular camera at time t - Δt to point p where |.| z represents the coordinate value in the z - direction;

[0157] D8. According to O t-Δt Project point p onto the pixel coordinate system q at time t - Δt t-Δt ;

[0158] D9. Calculate the average depth roundAvg(D t-Δt within the second range around q t-Δt [q t-Δt );

[0159] D10. If Determine that point p is occluded by other points at time t - Δt, otherwise determine that point p is not occluded by other points at time t - Δt;

[0160] D11 If it is determined in step D6 that the object where point p is located does not undergo self - occlusion at time t - Δt, and it is determined in step D10 that point p is occluded by points of other objects at time t - Δt, then determine that point p is occluded at time t - Δt.

[0161] In the above steps D1 - D11, steps D1 - D6 can determine whether point p undergoes self - occlusion at time t - Δt, steps D7 - D10 can determine whether point p is occluded by other points at time t - Δt, and step D11 combines the judgment results of steps D1 - D6 and the judgment results of steps D7 - D10 to determine whether point p is occluded at time t - Δt.

[0162] When applying the camera self - motion compensation discrimination algorithm based on the centroid of an object, when executing step S5, that is, the step of identifying whether the obstacles corresponding to each clustering cluster are dynamic obstacles or static obstacles, the following steps E1 - E4 are specifically executed:

[0163] E1. Obtain the centroid cen t of c t 、ct-Δt The centroid cen t-Δt , the speed threshold parameter V min and the compensation motion threshold parameter β;

[0164] E2. Determine whether the obstacle corresponding to the clustering cluster c t needs to perform self-motion compensation for the binocular camera. If so, jump to step E3. Otherwise, calculate the approximate speed v of the obstacle corresponding to the clustering cluster c t = Length(cen t , cen t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to the clustering cluster c t as a dynamic obstacle. Otherwise, identify the obstacle corresponding to the clustering cluster c t as a static obstacle;

[0165] E3. Calculate the compensation vector V for the change in the centroid of the obstacle corresponding to the clustering cluster c t of the binocular camera, and calculate the actual moving vector V of the centroid of the obstacle corresponding to the clustering cluster c c ; if difference(V t , V a ) < β, where difference() represents the modulus of the difference between two vectors, then determine that the object is a static obstacle and end the execution of steps E1 - E4. Otherwise, jump to step E4; c , V a ) < β, where difference() represents the modulus of the difference between two vectors, then determine that the object is a static obstacle and end the execution of steps E1 - E4. Otherwise, jump to step E4;

[0166] E4. Calculate the speed v of the obstacle corresponding to the clustering cluster c t = Length(cen t , cen t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to the clustering cluster c t as a dynamic obstacle. Otherwise, identify the obstacle corresponding to the clustering cluster c t as a static obstacle.

[0167] The principle of the camera self-motion compensation discrimination algorithm based on the object's centroid executed by steps E1-E4 is: based on the change of the centroid of the same obstacle in the previous and next frames, the moving speed of the object can be calculated. If the speed is greater than the threshold, the object is considered to be moving, otherwise the object is considered to be stationary. At the same time, considering the impact of the relative motion of the object caused by the camera's own motion, the camera's motion is compensated to the change of the object's centroid, so as to better judge the dynamic and static properties of the object. When the object is always in the camera's FOV, the centroid position of the object in the previous and next frames can be calculated more accurately, and the camera's self-motion will not cause obvious interference to it. At this time, there is no need to compensate for the camera's self-motion. However, when the previous frame or the current frame of the object is on the boundary of the camera's FOV, a part of the object will appear / leave the camera's FOV due to the camera's self-motion, resulting in the displacement of the object's centroid. At this time, the displacement of the object's centroid needs to be compensated.

[0168] Figure 3 and Figure 4 These are the two basic situations where binocular camera self-motion compensation is required, and other situations can be expanded from these two situations. One of the basic situations that requires compensation is as follows Figure 3 As shown in the figure, when the object is at the boundary of FOV, if the binocular camera performs translational motion, a part of the object will appear / leave FOV, which will cause an offset when calculating the object's center of mass. At this time, it is necessary to compensate for the object's center of mass offset. The corresponding compensation vector v3 can be approximately calculated as: v3 = v2-v4, where v1 is the translation vector of the binocular camera, v2 is the projection vector of v1 on the object surface, and v4 is the vector from the object boundary to the FOV boundary in the current frame. Another basic situation that needs compensation is as follows Figure 4 As shown in the figure, when the object is at the boundary of the FOV and the camera rotates, a part of the object will also appear / leave the FOV, causing an offset when calculating the object's center of mass. At this time, it is also necessary to compensate for the offset of the object's center of mass. The corresponding compensation vector can be approximately calculated as: v5 = rot*vc-vc. Among them, rot is the rotation matrix of the camera, and vc is the center of mass of the object in the previous frame. Therefore, the compensation vector V for the change of the camera's movement to the object's center of mass is c It can be expressed as V c =v3+v5. Only the object's center of mass motion vector and V c Only when there is a clear difference and the object moves faster than the set threshold, it is considered a dynamic obstacle.

[0169] When the discrimination algorithm based on historical consistency with a greater weight for closer frames is applied, when executing step S5, that is, the step of identifying whether the obstacles corresponding to each cluster are dynamic obstacles or static obstacles, the following steps F1-F6 are specifically executed:

[0170] F1. Obtain the time scale t h , the attenuation factor α, and the obstacle tracking history where Δt represents the time interval and k represents the sequence number;

[0171] F2. Initialize the variables Weight = 1, DynWeight = 0, StaticWeight = 0, unknownWeight = 0;

[0172] F3. Set the variable I, and I traverses from 0 to t h and execute the following steps F3a - F3d:

[0173] F3a. If the obstacle corresponding to the cluster c t-i*Δt belongs to a dynamic obstacle, operate on the variable DynWeight according to the formula DynWeight = DynWeight + Weight;

[0174] F3b. If the obstacle corresponding to the cluster c t-i*Δt belongs to a static obstacle, operate on the variable StaticWeight according to the formula StaticWeight = StaticWeight + Weight;

[0175] F3c. If the obstacle corresponding to the cluster c t-i*Δt belongs to an unknown obstacle, operate on the variable unknownWeight according to the formula unknownWeight = unknownWeight + Weight;

[0176] F3d. Operate on the variable Weight according to the formula Weight = Weight * α;

[0177] F4. If DynWeight > StaticWeight and DynWeight > unknownWeight, identify the obstacle corresponding to the cluster c t as belonging to a dynamic obstacle, otherwise execute step F5;

[0178] F5. If StaticWeight > DynWeight and StaticWeight > unknownWeight, identify the obstacle corresponding to the cluster c t as belonging to a static obstacle, otherwise execute step F6;

[0179] F6. Identify the obstacle corresponding to the cluster c t as belonging to an unknown obstacle.

[0180] The discrimination algorithm with a greater weight for more recent frames based on historical consistency executed in steps F1 - F6 can be represented by the following pseudocode:

[0181]

[0182] The principle of the discrimination algorithm with a greater weight for more recent frames based on historical consistency executed in steps F1 - F6 is as follows: According to the determination results of obstacles in the recent several frames, a vote is carried out to finally determine the static or dynamic attribute of the obstacle. And the frame closer to the present has a higher voting weight, while the frame farther from the present has less influence on the recognition of the static or dynamic attribute of the current obstacle. Therefore, it has a high robustness in the discrimination of the static or dynamic attribute of the current obstacle.

[0183] In step S6, when it is recognized in step S5 that the obstacle corresponding to the cluster belongs to a static obstacle, the obstacle information on the occupancy map is updated according to the cluster; when it is recognized in step S5 that the obstacle corresponding to the cluster belongs to a dynamic obstacle, the cluster is modeled into an ellipsoid, and the Kalman filter is used to track the modeled cluster. If the modeled cluster is not tracked within a continuous time period exceeding the threshold length, the tracking of the cluster is ended, and the data corresponding to the cluster is deleted.

[0184] By executing step S6, the UAV can achieve the dynamic update of environmental information such as obstacles and realize dynamic environment perception.

[0185] When using a computer to run steps S1 - S6 in a simulation environment, the obtained simulation perception effect is as Figure 5 and Figure 6 shown. The obtained real - machine perception effect diagram of the UAV equipped with a binocular camera running steps S1 - S6 is as Figure 7 shown.

[0186] By executing steps S1 - S6, the drone can process the depth images collected by the binocular camera using its own CPU or a cloud server when only using the binocular camera, and can quickly and preferably achieve dynamic environment perception of detecting, classifying the dynamic and static attributes, and tracking variously shaped obstacles in the environment under complex conditions such as close distances between obstacles, occlusion of obstacles, and the movement of the camera itself. Specifically, in the clustering stage, i.e., step S3, the connected density clustering algorithm is used to cluster the sparse point cloud, which preferably solves the problem in the traditional DBSCAN algorithm that different objects are likely to be clustered into one class when the object distances are close; in the stage of determining the dynamic and static attributes of the obstacle, i.e., step S5, through two algorithms of the camera self - motion compensation discrimination method based on the object centroid and the depth - adaptive discrimination method based on point voting, and using the dynamic and static obstacle discrimination algorithm with a greater weight for the most recent frame based on historical consistency for the final determination, the dynamic and static attributes of the obstacle can be robustly judged under complex conditions where the obstacle is occluded and the camera itself is moving.

[0187] Since the drone only uses the binocular camera and its own CPU to execute the dynamic environment perception method in the embodiment, without the need to be equipped with external devices such as lidar or special computing devices, the drone has the advantages of small load, low price, small volume, low power consumption, light weight, and easy portability.

[0188] A computer program for executing the dynamic environment perception method based on the binocular camera in this embodiment can be written, and this computer program can be written into a computer device or a storage medium. When the computer program is read and run, the dynamic environment perception method in this embodiment is executed, thereby achieving the same technical effects as the dynamic environment perception method based on the binocular camera in the embodiment.

[0189] It should be noted that, unless otherwise specified, when a certain feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to another feature, or indirectly fixed or connected to another feature. In addition, the up, down, left, right, etc. descriptions used in this disclosure are only relative to the mutual positional relationship of the various components of this disclosure in the drawings. The singular forms "a", "the", and "said" used in this disclosure are also intended to include the plural forms unless the context clearly indicates otherwise. In addition, unless otherwise defined, all the technical and scientific terms used in this embodiment have the same meanings as those commonly understood by those skilled in the art of this technical field. The terms used in the description of this embodiment are only for describing specific embodiments, rather than for limiting the present invention. The term "and / or" used in this embodiment includes any combination of one or more of the related listed items.

[0190] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as", etc.) provided in this embodiment is only intended to better illustrate the embodiments of the present invention and will not impose a limitation on the scope of the present invention unless otherwise required.

[0191] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The methods can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a programmed application-specific integrated circuit.

[0192] In addition, the operations of the processes described in this embodiment can be performed in any suitable order, unless this embodiment otherwise indicates or is otherwise clearly inconsistent with the context. The processes described in this embodiment (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.

[0193] Further, the method can be implemented in any type of computing platform operatively connected, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it is readable by a programmable computer and, when read by the storage medium or device, can be used to configure and operate the computer to perform the processes described herein. Additionally, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention as described in this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the invention also includes the computer itself.

[0194] A computer program can be applied to input data to perform the functions described in this embodiment, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.

[0195] As described above, only the preferred embodiments of the present invention are given, and the present invention is not limited to the above-described embodiments. As long as it achieves the technical effects of the present invention by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation manners can have various different modifications and variations.

Claims

1. A method for dynamic environment perception based on a binocular camera, characterized in that, the method for dynamic environment perception based on a binocular camera includes: Obtaining multiple frames of depth images captured by the binocular camera; Processing each frame of the depth image to obtain multiple frames of sparse point clouds; wherein, the processing result of one frame of the depth image is to obtain a corresponding frame of sparse point cloud; Clustering each frame of the sparse point clouds respectively to obtain multiple clustering clusters; wherein, the clustering result of one frame of the sparse point cloud is to obtain corresponding several clustering clusters, and different clustering clusters in one frame of the sparse point cloud respectively represent different obstacles; Associating multiple clustering clusters representing the same obstacle in the sparse point clouds of different frames; Identifying whether the obstacle corresponding to the associated clustering cluster belongs to a dynamic obstacle or a static obstacle; The processing of each frame of the depth image to obtain multiple frames of sparse point clouds includes: Obtaining the external parameter matrix T and internal parameter matrix K of the binocular camera; Obtain the pixel coordinates P of the depth image uv ; Determine the original point cloud through the formula P w = T -1 K -1 P uv where P w is the world coordinate of the original point cloud; Cropping the original point cloud to obtain a dense point cloud; Filtering the dense point cloud using a voxel filtering method to obtain the sparse point cloud; The associating multiple clustering clusters representing the same obstacle in the sparse point clouds of different frames includes: B1. Obtain m clusters where t represents C t Each cluster in C is obtained by clustering the sparse point cloud at acquisition time t, where the sparse point cloud at acquisition time t is obtained by processing the depth image at shooting time t, and the positions of all obstacles in C at time t are predicted through a Kalman filter t Set a distance threshold ε and initialize the association set F = K t ;​ B2. Calculate C t for each cluster centroid to obtain the centroids of all clusters in the current frame B3. Initialize the set Ω to be associated as Ω = D t ; B4. For each perform according to the following steps B4a - B4c: B4a. Find the nearest neighbor k in F j ; B4b. Find k j Nearest neighbor d in Ω j ; B4c. If i.e. is the nearest neighbor of k j associate with to obstacle j through the formula remove from the set Ω to be associated; remove k j from the associated set F through the formula F = F - {k j}; B5. If or end the execution of steps B1 - B6; conversely, for each perform the following steps B5a - B5b: B5a. Locate the nearest neighbor k in F j , calculate the distance δ between j k; B5b. If δ < ε, through the formula associate with obstacle j, through the formula associate remove from the set Ω to be associated, through the formula F = F - {k j} to remove k j from the associated set F; B6. If the execution of steps B1 - B6 is ended, and conversely, for each obstacle considered to be newly emerged, an obstacle tracking history is established, and the obstacle tracking history is expressed as where Δt represents the time interval and n represents the label of the obstacle.

2. The method for dynamic environment perception based on a binocular camera according to claim 1, characterized in that, the method for dynamic environment perception based on a binocular camera further includes: When the obstacle corresponding to the clustering cluster belongs to a static obstacle, updating the obstacle information on the occupancy map according to the clustering cluster; When the obstacle corresponding to the clustering cluster belongs to a dynamic obstacle, modeling the clustering cluster into an ellipsoid and using a Kalman filter to track the modeled clustering cluster.

3. The method for dynamic environment perception based on a binocular camera according to claim 2, characterized in that, the method for dynamic environment perception based on a binocular camera further includes: When the modeled clustering cluster is not tracked within a continuous time period exceeding the threshold length, ending the tracking of the clustering cluster and deleting the data corresponding to the clustering cluster.

4. The method for dynamic environment perception based on a binocular camera according to claim 1, characterized in that, the clustering of each frame of the sparse point clouds respectively to obtain multiple clustering clusters includes: A1. Obtain a sample point set D = {x 1 , x 2 , ……, x m}, where x m represents the m-th point in the sparse point cloud. Set the neighborhood distance threshold ε and the connectivity threshold δ, and initialize the core object set Initialize the number of clustering clusters k = 0, initialize the unvisited sample set F = D, and initialize the cluster partition A2. For j = 1, 2, ……, m, find all core objects according to the following steps A2a - A2b: A2a. Find the sub-sample point set N within the ε-neighborhood of the sample x j ; where N is within the ε-neighborhood of the sample point x ε (x j ); where x j is a sample point in the sample point set D A2b. If the number of samples in the sub-sample point set N ε (x j ) satisfies |N ε (x j )|≥MinPts, then calculate the number of connected components n of the point Q = N ε (x j ) ∪ x j ; if n < δ, add the sample x j to the core object set Ω through the formula Ω = Ω ∪ {x j}; A3. If the set of core objects then terminate the execution of steps A1 - A6, otherwise execute step A4; A4. Randomly select a core object o from the set of core objects Ω, and initialize the current cluster core object list Ω cur = {o}, initialize the category serial number k = k + 1, and initialize the current cluster sample set C k = {o}, and update the set of unvisited samples f - F - {o}; A5. If the current cluster core object queue then the current clustering cluster C k is generated, update the cluster partition C = {C 1 , C 2 , ……, C k}, update the core object set Ω = Ω - C k , return to execute step A3, otherwise update the core object set Ω = Ω - C k ; A6. Take a core object o'' from the current cluster core object queue Ω cur and find the ε-neighborhood subsample point set N ε (o'') through the neighborhood distance threshold ε. Let Δ = N ε ∩ F of (o''), and update the current cluster sample set C k = C k ∪ Δ, update the unvisited sample set F = F - Δ, and update Ω cur = Ω cur ∪ (Δ ∩ Ω) - o'', and return to execute step A5.

5. The method for dynamic environment perception based on a binocular camera according to claim 1, characterized in that, the identifying whether the obstacle corresponding to each clustering cluster belongs to a dynamic obstacle or a static obstacle includes: C1. Obtain the depth image D captured at time t - Δt t-Δt and the depth image D captured at time t t , obtain the dense point cloud at the acquisition time t - Δt wherein the dense point cloud at the acquisition time t - Δt is obtained by processing the depth image captured at time t - Δt, and obtain the pose O of the UAV where the binocular camera is located at time t - Δt t-Δt and the pose O at time t t , obtain where p represents a point in the world coordinate system, l represents the label of the point, and obtain the discrimination parameters (β, β min , V mtn ); Traverse C t in any of the clustering clusters where I = 1, 2, ……, m. If the clustering cluster corresponds to a newly identified obstacle, identify the obstacle corresponding to the clustering cluster as belonging to an unknown obstacle, where the unknown obstacle is an obstacle that is not determined to belong to either a dynamic obstacle or a static obstacle; C3. Initializing variables vote = 0, dyn = 0, static = 0; C4. For each point p in the cluster perform the following voting process: C4a. According to the pose O t-Δt Determine whether the point p is within the field of view angle of the UAV at time t - Δt, and determine whether the point p is occluded at time t - Δt; when the point p is not within the field of view angle of the UAV at time t - Δt and the point p is occluded at time t - Δt, operate on the variable vote according to the formula vote = vote + 1, otherwise execute step C4b; C4b. Find the nearest neighbor nn of point p in the dense point cloud ; when where length(p, nn) is the distance between point p and the nearest neighbor nn, and Δt represents the time interval, operate on the variable dyn according to the formula dyn = static + 1, otherwise execute step C4c; C4c. Operating on the variable static according to the formula static = static + 1; C5. When the clustering cluster corresponding obstacle is recognized as belonging to a dynamic obstacle, and conversely the clustering cluster corresponding obstacle is recognized as belonging to a static obstacle; where represents the centroid of the clustering cluster , represents the distance between t and O range sensing represents the maximum measurement distance of the binocular camera.

6. The method for dynamic environment perception based on a binocular camera according to claim 5, characterized in that, the judging whether the point p is occluded at the moment t - Δt includes: D1. Obtain c t The centroid cen of t c t-Δt The centroid cen of t-Δt , and the judgment parameter β(Δt) related to the time Δt; D2. According to O t Project cen t onto the pixel coordinate system at time t to obtain q t ; D3. According to O t-Δt project cen t-Δt onto the pixel coordinate system at time t - Δt to obtain q t-Δt ; D4. Calculate q t The average depth roundAvg(D within the first range around t [q t ); D5. Calculate q t-Δt The average depth roundAvg(D within the first range around t-Δt [q t-Δt ) D6. If roundAvg(D t [q t ) - roundAvg(D t-Δt [q t-Δt ) > β(Δt), it is determined that point p has self-occluded at time t - Δt; otherwise, it is determined that point p has not self-occluded at time t - Δt; D7. According to O t-Δt Project the point p onto the camera coordinate system at time t - Δt Obtain the depth from the position of the binocular camera at time t - Δt to the point p where |.| z represents the coordinate value in the z direction; D8. According to O t-Δt Project the point p onto the pixel coordinate system q at time t - Δt t-Δt ; D9. Calculate q t-Δt The average depth roundAvg(D within the second range around t-Δt [q t-Δt ) D10. If it is determined that point p is occluded by other points at time t - Δt, otherwise it is determined that point p is not occluded by other points at time t - Δt; D11. If the object where the point p is located in step D6 does not undergo self-occlusion at time t - Δt, and the point p is occluded by the points of other objects at time t - Δt in step D10, then it is determined that the point p is occluded at time t - Δt.

7. The binocular camera-based dynamic environment perception method according to claim 1, characterized in that identifying whether the obstacles corresponding to the respective clustering clusters are dynamic obstacles or static obstacles includes: E1. Obtain c t 's centroid cen t , c t-Δt 's centroid cen t-Δt , speed threshold parameter V min and compensation motion threshold parameter β; E2. Determine the clustering cluster c t Whether the corresponding obstacle requires self-motion compensation of the binocular camera. If so, jump to step E3. Otherwise, calculate the approximate speed v of the obstacle corresponding to the clustering cluster c t v = Length(cen t , cen t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to the clustering cluster c t as a dynamic obstacle. Otherwise, identify the obstacle corresponding to the clustering cluster c t as a static obstacle; E3. Calculate the relative motion of the binocular camera with respect to the clustering cluster c t The compensation vector V for the change in the centroid of the corresponding obstacle c , calculate the clustering cluster c t The actual movement vector V of the centroid of the corresponding obstacle a ; if difference(V c , V a ) < β, where difference() represents the modulus of the difference between two vectors, then determine that the object is a static obstacle and end the execution of steps E1 - E4. Otherwise, jump to execute step E4; E4. Calculate the cluster c t The speed v of the corresponding obstacle is v = Length(cen t , cen t-Δt ) / Δt. If v > V min , identify the obstacle corresponding to cluster c t as a dynamic obstacle, otherwise, identify the obstacle corresponding to cluster c t as a static obstacle.

8. The binocular camera-based dynamic environment perception method according to claim 1, characterized in that identifying whether the obstacles corresponding to the respective clustering clusters are dynamic obstacles or static obstacles includes: F1. Obtain the time scale t h , the attenuation factor α, and the obstacle tracking history where Δt represents the time interval and k represents the sequence number; F2. Initialize variables Weight = 1, DynWeight = 0, StaticWeight = 0, unknownWeight = 0; Set variable I, where I ranges from 0 to t h Traverse and perform the following steps F3a - F3d: F3a. If the cluster c t-i*Δt corresponding obstacle belongs to a dynamic obstacle, operate on the variable DynWeight according to the formula DynWeight = DynWeight + Weight; F3b. If the obstacle corresponding to cluster c t-i*Δt belongs to a static obstacle, operate on the variable StaticWeight according to the formula StaticWeight = StaticWeight + Weight; F3c. If the cluster c t-i*Δt corresponding obstacle belongs to the unknown obstacle, operate on the variable unknownWeight according to the formula unknownWeight = unknownWeight + Weight; F3d. Operate on the variable Weight according to the formula Weight = Weight * α; F4. If DynWeight > StaticWeight and DynWeight > unknownWeight, identify the obstacle corresponding to cluster c as a dynamic obstacle; otherwise, execute step F5; t ​ F5. If StaticWeight > DynWeight and StaticWeight > unknownWeight, cluster c t The corresponding obstacle is identified as a static obstacle, otherwise step F6 is executed; F6. Identify the obstacle corresponding to the cluster c t as an unknown obstacle.

Citation Information

Patent Citations

  • Dynamic obstacle tracking method based on sparse laser radar data

    CN111337941A

  • Dynamic obstacle detection method based on stereoscopic vision

    CN113536959A