Binocular optical detection method for target motion trail in cluster distribution

By using three binocular cameras to acquire multi-view depth images and using similarity comparison networks to match images, the problem of target tracking error and trajectory tracking difficulties under high frame rate shooting is solved, and high-precision tracking of target motion trajectories in the cluster is achieved.

CN120070509APending Publication Date: 2025-05-30CHANGCHUN UNIV OF SCI & TECH +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510527887.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing binocular optical detection technology is difficult to accurately track targets in cluster distribution when shooting at high frame rates, and due to the irregularity of target motion, it is difficult to achieve continuous trajectory tracking of the same target.

Method used

Three binocular cameras are used to shoot targets in the cluster from different directions at the same frame frequency, obtain depth images from multiple perspectives, and image matching is performed through similarity comparison networks to ensure accurate tracking of any target.

Benefits of technology

Through multi-view angle matching and depth image reconstruction, high-precision tracking of target motion trajectories in the cluster is achieved, suitable for high frame rate shooting, reducing the reduction in three-dimensional reconstruction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070509A_ABST
    Figure CN120070509A_ABST
Patent Text Reader

Abstract

The invention discloses a binocular optical detection method for a target movement track in cluster distribution. Belongs to the technical field of target detection and particularly relates to the technical field of target detection in cluster distribution. The method solves the problem of tracking error caused by the fact that two cameras cannot accurately shoot targets at the same time during high-frame-rate shooting when binocular optical detection equipment is used, and solves the problem that the tracking error is low due to the irregularity of target movement in cluster distribution and the mutual shielding effect between cluster targets. Therefore, the problem of difficulty in continuous trajectory tracking of the same target is solved. For any target in the cluster, three-view-angle shooting is adopted, matching of the same target is carried out on continuous two times of shooting in three view angles through a self-designed image matching method, in the method, image similarity calculation is carried out through feature vectors, the calculated data size can be simplified, and the method is suitable for target matching under the condition of high-frame-rate shooting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and particularly relates to the technical field of target detection in cluster distribution. Background Art

[0002] In a population with a cluster distribution, the positions of each target are random and there is no fixed pattern. In the cluster, the movement directions of each target are generally the same, all moving towards the destination, and the movement trajectories of each target in the cluster are divergent.

[0003] The detection of the movement trajectories of targets in cluster distribution is specifically applied in various fields as follows: Military field: UAV swarm monitoring, missile trajectory early warning (such as infrared early warning satellites); Civil field: traffic flow monitoring (vehicle flow, pedestrian flow), wildlife migration research; Aerospace field: space debris tracking, low-earth orbit satellite formation management.

[0004] In the existing optical detection technology for the movement trajectories of targets in cluster distribution, there are two main problems. One is that when using a binocular optical detection device, during high-frame-rate shooting, there is a tracking error problem caused by the two cameras not being able to completely and accurately capture the target simultaneously. The other is that due to the irregular movement of targets in cluster distribution, it is difficult to continuously track the same target. Summary of the Invention

[0005] In order to solve the above two main technical problems, the present invention provides a binocular optical detection method for the movement trajectories of targets in cluster distribution, and the method includes the following steps: S1. Use three binocular cameras at the same frame rate From Start shooting the targets in the cluster from three different directions at the same time to obtain the depth images of any target in the cluster At the moment and At the moment corresponding to three different perspectives; S2. For any target in the cluster, match the depth images of the same target at At the moment and At the moment through the cluster images captured by the three binocular cameras to ensure correct tracking of any target; S3. For any target in the cluster, obtain its spatial three-dimensional coordinates corresponding to At the moment and At the moment through the depth images of three different perspectives corresponding to At the moment and At the moment respectively, so as to determine the movement trajectory of any target from At the moment to At the moment; S4. Replace With Return to step S1 and continue to execute until the detection of the target motion trajectory ends.

[0006] Further, step S2 is specifically as follows: S21. For any target , record the depth image of it taken by the first binocular camera at as the target . Compare the similarity of the target with all the target depth images taken by the corresponding binocular camera at in sequence, and extract the depth image with the highest similarity to the target from the depth images taken at as the target at ; S22. For any target , record the depth image of it taken by the second binocular camera at as the target . Compare the similarity of the target with all the target depth images taken by the corresponding binocular camera at in sequence, and extract the depth image with the highest similarity to the target from the depth images taken at as the target at ; S23. For any target , record the depth image of it taken by the third binocular camera at as the target . Compare the similarity of the target with all the target depth images taken by the corresponding binocular camera at in sequence, and extract the depth image with the highest similarity to the target from the depth images taken at as the target at .

[0007] Further, the similarity comparison is performed through a similarity comparison network, and the structure of the similarity comparison network is specifically as follows: The two input images to be compared respectively pass through two feature extraction networks with the same structure, and respectively obtain the feature vectors and the feature vector . Subtract the feature vector from the feature vector and take the absolute value to obtain the vector . The vector Input the fully connected layer, and map the output similarity value between 0 and 1 through the activation function. The larger the similarity value, the higher the similarity.

[0008] Furthermore, the feature extraction network sequentially includes a first convolutional layer, a first activation function layer, a second convolutional layer, a second activation function layer, and a max pooling layer from the input to the output.

[0009] Furthermore, when training the similarity comparison network, set labels for the images in the training set, where 1 represents the same image and 0 represents different images, and select the contrast loss function as the loss function.

[0010] Furthermore, step S3 is specifically as follows: S31. Obtain the depth image of the target captured by a binocular camera at a moment, and restore the depth image to a three-dimensional spatial image ; S32. In the same way as step S31, restore the depth images of the targets captured by the other two binocular cameras at moments to three-dimensional spatial images and ; S33. Perform spatial coordinate fusion on the three-dimensional spatial images , and to obtain the three-dimensional spatial coordinates of the target at a moment; S34. In the same way as steps S31 - S33, obtain the three-dimensional spatial coordinates of the same target at a moment, and determine the motion trajectory of any target from a moment to a moment.

[0011] Furthermore, step S31 is specifically as follows: S311. Set a mark for any pixel point in the depth image; S312. Input the depth image into the voxel feature extraction model to extract voxel features; S313. Restore the three-dimensional image corresponding to the depth image through the extracted voxel features. In the three-dimensional image, the marked pixel points are converted into marked point clouds in the three-dimensional image; S314. Convert the coordinates of the pixel points in S311 from the camera coordinate system to the world coordinate system, which is the spatial coordinate of the corresponding marked point cloud in the three-dimensional image. Through the corresponding positional relationship, convert the spatial positions of all the point clouds in the three-dimensional image to be represented by the world coordinate system to obtain the three-dimensional image spatial coordinates.

[0012] Further, step S313 is specifically as follows: The feature vectors of the entire space are obtained from the voxel features by using a spatial sampling interpolation method, where each sampling point corresponds to a feature vector, and a multi-layer perceptron is used to predict the directed segment distance field corresponding to each sampling point in the space, so as to obtain the directed truncation distance field of the entire space. According to the directed truncation distance field of the entire space, a three-dimensional image corresponding to the voxel features is reconstructed.

[0013] The beneficial effects of the method of the present invention are as follows: For any target in the cluster, through the self-designed image matching method, the images taken by three binocular cameras in two consecutive shootings are respectively matched for the same target, so as to ensure accurate tracking of the same target at any time. In the method, the image similarity is calculated through feature vectors, which can simplify the amount of data for calculation and speed up the matching speed, and is suitable for target matching in the case of high-frame-rate shooting.

[0014] In the traditional method of three-dimensional image reconstruction by binocular cameras, it is necessary to perform epipolar rectification on the left and right cameras in the binocular cameras: adjust the image planes of the left and right cameras to be coplanar and row-aligned, so that the matching points are located on the same horizontal line (epipolar line), which simplifies the subsequent matching; however, in the shooting of continuously moving objects in space, due to the use of high frame rate, real-time epipolar rectification cannot be provided for each frame of image, because it involves real-time processing of images and can only be initially corrected at the beginning of shooting, which will lead to a decrease in the accuracy of three-dimensional reconstruction. In the method adopted in the present invention, since it does not involve stereo matching of the left and right cameras (the next step of epipolar rectification), but first reconstructs the three-dimensional image and then performs one-by-one matching according to a marker point, high-precision three-dimensional reconstruction images can be obtained even without real-time epipolar rectification. Description of the Drawings

[0015] Figure 1 It is a similarity comparison network structure diagram in an embodiment of the present invention.

[0016] Figure 2 It is a schematic diagram of the three-view reconstruction of the target space position and trajectory prediction in an embodiment of the present invention. Detailed Embodiments

[0017] Next, the technical solutions of the present invention will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] This embodiment provides a binocular optical detection method for the target motion trajectory in cluster distribution, as Figure 2As shown, the method includes the following steps: S1. Use three binocular cameras at the same frame rate starting from to take pictures of the targets in the cluster in three different directions, and obtain the depth images of any target in the cluster at time and the depth images of three different perspectives corresponding to time

[0019] Since objects in the cluster in motion are prone to occlusion problems, three binocular cameras are used to take pictures of the objects from different angles to overcome the occlusion problem. Using three binocular cameras can ensure that even if one side is occluded, the target can still be photographed from the other two sides, and a series of subsequent methods are also applicable to the perspectives of two binocular cameras and one binocular camera, but the accuracy will decrease relatively compared to the perspective of three binocular cameras. After multiple experiments, taking pictures from different angles with three binocular cameras is currently the method that can achieve relatively high shooting accuracy with the least number of cameras.

[0020] S2. For any target in the cluster, match the depth images of the same target at time and time through the cluster images taken by three binocular cameras to ensure correct tracking of any target.

[0021] Since the objects in the cluster in motion are constantly changing, their positions in each frame of the image are also changing, which will lead to the problem of difficult tracking. Step S2 is to ensure the accuracy of tracking.

[0022] S21. For any target , record the depth image of it taken by the first binocular camera at time as target . Compare the similarity of target with all the target depth images taken by the corresponding binocular camera at time in sequence, and extract the depth image with the highest similarity to target from the depth images taken at time as target at ; S22. For any target , record the depth image of it taken by the second binocular camera at time as target . Compare the similarity of target with all the target depth images taken by the corresponding binocular camera at time in sequence, and extract the depth image with the highest similarity to target from the depth images taken at The depth image with the highest similarity is used as the target at time; ; S23. For any target , the depth image of it taken by the third binocular camera at time is denoted as target . For target and all the target depth images taken by the corresponding binocular camera at time are successively compared in similarity, and the depth image with the highest similarity to target in the depth image taken at time is extracted and used as the target at time .

[0023] Considering the occlusion effect in the cluster, this target at time and time may be occluded in one (or several) perspectives. Therefore, similarity comparison is required to determine whether this target at time and time exists in the depth image obtained by the detector.

[0024] Among them, the similarity comparison is carried out through a similarity comparison network, and the structure of the similarity comparison network is as Figure 1 shown. Specifically, the two input images to be compared respectively pass through two feature extraction networks with the same structure and respectively obtain feature vectors and feature vector . The absolute value is obtained by subtracting feature vector from feature vector to obtain vector . Vector is input into the fully connected layer, and the output similarity value is mapped between 0 and 1 through the activation function. The larger the similarity value, the higher the similarity. The activation function selects the sigmoid activation function.

[0025] The feature extraction network successively includes a first convolutional layer, a first activation function layer, a second convolutional layer, a second activation function layer and a max pooling layer from input to output. Among them, the activation function selects the ReLU activation function.

[0026] When training the similarity comparison network, labels are set for the images in the training set. 1 represents the same image, and 0 represents different images. The loss function selects the contrastive loss function.

[0027] S3. For any target in the cluster, obtain its three-dimensional spatial coordinates corresponding to the moments of and through depth images from three different perspectives, so as to determine the motion trajectory of any target from the moment of to the moment of . and through depth images from three different perspectives, and then determine the three-dimensional spatial coordinates corresponding to the moments of and respectively for any target, thereby determining the motion trajectory of any target from the moment of to the moment of . and respectively, so as to determine the motion trajectory of any target from the moment of to the moment of . to

[0028] S31. Obtain the depth image of the target captured by a binocular camera at the moment of , and restore the depth image to a three-dimensional spatial image ; and restore the depth image to a three-dimensional spatial image . ; S32. In the same way as in step S31, restore the depth images of the target captured by the other two binocular cameras at the moment of to three-dimensional spatial images and respectively; and restore the depth images of the target captured by the other two binocular cameras at the moment of to three-dimensional spatial images and respectively. and ; S33. Perform spatial coordinate fusion on the three-dimensional spatial images , , and to obtain the three-dimensional spatial coordinates of the target at the moment of ; , and to obtain the three-dimensional spatial coordinates of the target at the moment of . ; S34. In the same way as in steps S31 - S33, obtain the three-dimensional spatial coordinates of the same target at the moment of , and determine the motion trajectory of any target from the moment of to the moment of . and determine the motion trajectory of any target from the moment of to the moment of . to

[0029] Specifically, step S31 is as follows: S311. Set a mark for any pixel point in the depth image; S312. Input the depth image into a voxel feature extraction model to extract voxel features; The voxel feature extraction model can be an existing voxel feature extractor. The voxel feature extractor uses a voxel feature encoding (VFE) layer to extract voxel features. The VFE layer takes all points in the same voxel as input and uses a fully connected network (FCN) composed of a linear layer, a batch normalization (BatchNorm) layer, and a ReLU layer to extract point features. Then, element-wise max pooling is used to obtain the local aggregation features of each voxel. Finally, the obtained features are flattened, and these flattened features and point features are concatenated together.

[0030] S313. Restore the three-dimensional image corresponding to the depth image through the extracted voxel features. In the three-dimensional image, the marked pixel points are converted into marked point clouds in the three-dimensional image. The feature vectors of the entire space are obtained from the voxel features by using a spatial sampling interpolation method. Each sampling point corresponds to a feature vector. A multi-layer perceptron is used to predict the directed phase distance field corresponding to each sampling point in the space, and the directed truncated distance field of the entire space is obtained. According to the directed truncated distance field of the entire space, a three-dimensional image corresponding to the voxel features is reconstructed (three-dimensional reconstruction algorithm).

[0031] Common spatial sampling interpolation methods include nearest neighbor interpolation, linear interpolation, cubic spline interpolation, etc. In this embodiment, spatial bilinear interpolation is used.

[0032] According to the directed truncated distance field (TSDF) of the entire space, a three-dimensional image corresponding to the voxel features for three-dimensional reconstruction is obtained by using a differentiable rendering algorithm based on orthographic projection.

[0033] S314: Convert the coordinates of the pixel points in S311 from the camera coordinate system to the world coordinate system, which are the spatial coordinates of the corresponding marked point cloud in the three-dimensional image. Through the corresponding positional relationship, the spatial positions of all the point clouds in the three-dimensional image are converted to be represented by the world coordinate system, and the three-dimensional image spatial coordinates are obtained.

[0034] Since each pixel point in the depth image corresponds to the point cloud of the reconstructed three-dimensional image, it can be understood that each pixel has a corresponding relationship with each point cloud, and there is a definite positional relationship between each pixel and other pixels, and there is also a definite positional relationship between each point cloud and other point clouds. Then it can be inferred that knowing the spatial coordinates of one pixel in the depth image can know the spatial coordinates of other pixels, and knowing the spatial coordinates of one point cloud in the three-dimensional image can know the spatial coordinates of other point clouds. Therefore, at this time, because there is a known correspondence between a pixel and a point cloud, the spatial coordinates of this pixel are the spatial coordinates of the point cloud, and the spatial coordinates of all point clouds can be inferred from the spatial coordinates of one point cloud, that is, the three-dimensional image spatial coordinates.

[0035] The reason for determining the three-dimensional image spatial coordinates in this way is that when the method of the present invention uses the depth image to restore the three-dimensional image, it is based on a modeling method. Therefore, it is also necessary to map the corresponding point cloud in the model to the world coordinate system to obtain the true position of the corresponding target at this moment, so as to achieve trajectory tracking.

Claims

1. A binocular optical detection method for target motion trajectory in cluster distribution, characterized in that: The method comprises the following steps: S1, using three binocular cameras with the same frame rate from Shoot the targets in the cluster from three different directions at the same time, and obtain any target in the cluster Moment and Depth images of three different perspectives corresponding to each moment; S2. For any target in the cluster, the cluster images taken by three binocular cameras are used to identify the same target in Moment and Match the depth images at each moment to ensure that any target is tracked correctly; S3. For any target in the cluster, Moment and The depth images of three different perspectives corresponding to the time are obtained Moment and The three-dimensional coordinates of the space corresponding to each moment, so as to determine the Time has come The trajectory of movement at each moment; S4. Replace with Return to step S1 and continue executing until the target motion trajectory detection is completed.

2. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 1, characterized in that: Step S2 is specifically as follows: S21. For any target , the first binocular camera captures the The depth image at the moment is recorded as the target , for the target With the corresponding binocular camera All target depth images taken at the moment are compared for similarity in turn to extract The depth image captured at the moment and the target The depth image with the highest similarity is used as The goal of the moment ; S22. For any target , the second binocular camera captures the The depth image at the moment is recorded as the target , for the target With the corresponding binocular camera All target depth images taken at the moment are compared for similarity in turn to extract The depth image captured at the moment and the target The depth image with the highest similarity is used as The goal of the moment ; S23. For any target , the third binocular camera takes The depth image at the moment is recorded as the target , for the target With the corresponding binocular camera All target depth images taken at the moment are compared for similarity in turn to extract The depth image captured at the moment and the target The depth image with the highest similarity is The goal of the moment .

3. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 2, characterized in that: The similarity comparison is performed through a similarity comparison network. The structure of the similarity comparison network is as follows: two input images to be compared are respectively passed through two feature extraction networks with the same structure. , and obtain the feature vectors and the eigenvector , the feature vector and the eigenvector Subtract and find the absolute value to get the vector , the vector Input the fully connected layer and map the output similarity value between 0 and 1 through the activation function. The larger the similarity value, the higher the similarity.

4. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 3, characterized in that: Feature extraction network From input to output, it includes the first convolution layer, the first activation function layer, the second convolution layer, the second activation function layer and the maximum pooling layer.

5. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 4, characterized in that: When training the similarity comparison network, labels are set for the images in the training set, 1 represents the same image, 0 represents different images, and the loss function selects the comparison loss function.

6. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 5, characterized in that: Step S3 is specifically as follows: S31, obtain a target photographed by a binocular camera The depth image at the moment is restored to a spatial three-dimensional image ; S32, using the same method as step S31, the targets photographed by the other two binocular cameras are The depth images at each moment are restored to spatial three-dimensional images and ; S33, the spatial three-dimensional image , and Perform spatial coordinate fusion to obtain the target The three-dimensional spatial coordinates of the moment; S34, using the same method as steps S31-S33, obtain the same target The three-dimensional spatial coordinates at the time of any target Time has come The movement trajectory of time.

7. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 6, characterized in that: The step S31 is specifically as follows: S311, setting a mark for any pixel in the depth image; S312, inputting the depth image into a voxel feature extraction model to extract voxel features; S313, restoring the three-dimensional image corresponding to the depth image by using the extracted voxel features, and converting the marked pixel points in the three-dimensional image into a marked point cloud in the three-dimensional image; S314, converting the coordinates of the pixel points in S311 from the camera coordinate system to the world coordinate system, that is, the spatial coordinates of the corresponding marked point cloud in the three-dimensional image, and through the corresponding positional relationship, converting the spatial positions of all point clouds in the three-dimensional image to be represented by the world coordinate system to obtain the three-dimensional image spatial coordinates.

8. The binocular optical detection method for target motion trajectory in cluster distribution according to claim 7, characterized in that: Step S313 is specifically as follows: a feature vector of the entire space is obtained from the voxel features using a spatial sampling interpolation method, wherein each sampling point corresponds to a feature vector, and a multi-layer perceptron is used to predict the directed segment distance field corresponding to each sampling point in the space to obtain the directed truncated distance field of the entire space, and a three-dimensional image corresponding to the voxel features is three-dimensionally reconstructed based on the directed truncated distance field of the entire space.

Citation Information

Patent Citations

  • Target tracking method and target tracking device

    CN102867311A

  • Cross-shot multi-target tracking method and system

    CN108876821A

  • Mechanical arm and three-dimensional reconstruction method and device thereof

    CN111540045A

  • Multi-target tracking method based on graph network

    CN111881840A

  • Vehicle pavement feature recognition and three-dimensional reconstruction method based on binocular vision

    CN113129449A