A method and system for tracking spatial target motion by combining local and global information

By combining local and global information in a spatial target motion tracking method, and using binary description vectors and neural networks to calculate Euclidean distance, the problem of repeated textures affecting traditional methods is solved, and efficient target motion tracking is achieved.

CN116958202BActive Publication Date: 2025-10-31HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310945393.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-28
Publication Date
2025-10-31
Estimated Expiration
2043-07-28

AI Technical Summary

Technical Problem

Traditional feature location description processes only utilize information from the local area surrounding the feature point location, which is insufficient to handle a large number of repetitive textures in spatial targets, leading to target motion tracking failures, and also incurring high computational and storage requirements.

Method used

A spatial target motion tracking method that combines local and global information obtains binary description vectors of feature locations, calculates Euclidean distance using a pre-trained neural network, and selects the best matching relationship for tracking.

Benefits of technology

It improves computational efficiency, reduces storage consumption, effectively addresses the impact of repetitive textures, and reduces the computational load of the matching process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958202B_ABST
    Figure CN116958202B_ABST
Patent Text Reader

Abstract

This invention discloses a spatial target motion tracking method and system that combines local and global information, relating to the field of target tracking technology, to solve the problem of tracking failure caused by repetitive textures of spatial targets in existing technologies. The key technical points of this invention include: encoding the image feature positions of the spatial target into binary description vectors; calculating the Euclidean distance between multiple binary description vectors in all video frames; for each feature position in the previous video frame, determining the multiple feature positions with the smallest corresponding Euclidean distance in the subsequent video frame as preliminary matching relationships; using a neural network to calculate the Euclidean distance of the preliminary matching relationships for verification, retaining only the best matching relationship with the smallest distance; and solving the spatial target motion based on the best matching relationship of the feature positions to achieve spatial target tracking. This invention improves the calculation efficiency of the Euclidean distance between feature positions in images, reduces the computational load of the matching process, and has lower storage consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target tracking technology, specifically to a spatial target motion tracking method and system that combines local and global information. Background Technology

[0002] Optical-based target motion tracking can recover the motion of the tracked target, providing information support for tasks such as navigation approximation and target perception and recognition, and is one of the key problems in the field of computer vision. In target motion tracking, it is usually necessary to extract the feature positions in the target, analyze the changes in these feature positions between images, and solve for motion based on the projection process.

[0003] However, traditional feature location description processes only utilize information from the local region surrounding the feature point, making it difficult to handle a large number of repetitive textures in spatial targets, and posing a potential risk of failure in tracking the motion of spatial targets. Introducing global information through neural networks can construct different description vectors for similar textures at different locations, thus effectively addressing the impact of repetitive textures in spatial target tracking. However, this process has higher requirements for computation and storage space compared to binary description vectors. Summary of the Invention

[0004] In view of the above problems, this invention proposes a spatial target motion tracking method and system that combines local and global information to solve the problem of repetitive textures of spatial targets that are difficult to handle by existing methods, and improves computational efficiency while dealing with similar structures.

[0005] According to one aspect of the present invention, a spatial target motion tracking method combining local and global information is proposed, the method comprising the following steps:

[0006] Step 1: Acquire the image and video stream of the spatial target motion and sample it into a continuous video frame sequence;

[0007] Step 2: For each video frame, obtain the feature positions of spatial targets and encode the feature positions into binary description vectors;

[0008] Step 3: Calculate the Euclidean distance between multiple binary description vectors in all video frames. For each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as the preliminary matching relationship.

[0009] Step 4: Input the video frame sequence into the pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position;

[0010] Step 5: Calculate the Euclidean distance between the multidimensional vector corresponding to each feature position in the previous video frame and the multidimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame. Select the feature position with the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame.

[0011] Step 6: Solve the motion of the spatial target based on the best matching relationship of the feature positions in order to track the spatial target.

[0012] Furthermore, the specific steps of step two include:

[0013] Step 2: For each video frame, calculate the Harris feature points of the image;

[0014] Step 22: Set up the non-maximum suppression region, filter Harris feature points according to the non-maximum suppression region, and use the filtered Harris feature points as the spatial target feature locations.

[0015] Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region;

[0016] Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

[0017] Furthermore, the specific process of filtering Harris feature points based on the non-maximum suppression region in step 22 includes: taking the maximum value positions in sequence according to the Harris response amplitude, setting the response values ​​of all pixel positions in the non-maximum suppression region around the maximum value position to 0, and repeating the above process to filter Harris feature points.

[0018] Furthermore, the specific steps in steps two and three include:

[0019] Extract a circular region M centered at feature location O with radius r. O Calculate the circular region M O Pixel value centroid C:

[0020]

[0021] Where I(x,y) is the pixel value corresponding to pixel (x,y) in the region;

[0022] Connect the feature location O with the centroid C to construct the main direction of the region. For the circular region M O Calculate the value of any point P in the middle. and included angle ∠COP;

[0023] According to ∠COP, the circular region M O The system is divided into m quadrants, and within each quadrant, it is further divided into n intervals based on the distance from the pixel to the feature location O, thus dividing the feature location into multiple sub-regions.

[0024] Furthermore, the specific steps of step two four include:

[0025] For circular region M O Histogram equalization is performed on all pixels, and the equalized circular region M is calculated. O Average pixel value;

[0026] The proportional interval [0, 100%] is uniformly divided into p1 intervals. A corresponding binary code sequence B1 is assigned to each interval, where adjacent intervals differ by only one bit. The encoding length k1 of the binary code sequence B1 satisfies... The smallest integer solution;

[0027] Calculate the circular region M O The proportion of pixels in each sub-region that are higher than the average pixel value According to pixel ratio The corresponding proportional range determines the value of the binary code sequence B1 for that sub-region;

[0028] Calculate the sub-centroid of each sub-region

[0029]

[0030] according to get The angle between the angle along the clockwise direction and its nearest quadrant axis is determined based on the number of quadrants (m) in steps two and three, with the upper limit of this angle being... The included angle range The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2, with the following encoding rules: adjacent interval code sequences differ by only one bit, and the first and last interval codes also differ by only one bit; the encoding length k2 of the binary code sequence B2 satisfies... The smallest integer solution;

[0031] Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region;

[0032] By concatenating B1 and B2, a circular region M is obtained. O The encoding of each sub-region;

[0033] By concatenating the codes of all sub-regions in order, we obtain a (k1+k2)mn-bit binary code, which is the binary description vector.

[0034] Furthermore, the structure design of the neural network described in step four is as follows:

[0035] The neural network has a total of 8 layers. The first 7 layers all use the form of "2D convolution + ReLU activation function", and the last layer only uses 2D convolution. The output dimensions of each layer are 8, 16, 32, 32, 64, 64, 128 and 128 respectively. The third and sixth layers use a convolution kernel of size 5, a stride of 2 and edge padding of 1. The convolution kernel, stride kernel and edge padding of the remaining layers are set to 3, 1 and 1 respectively.

[0036] According to another aspect of the present invention, a space target motion tracking system combining local and global information is proposed, the system comprising:

[0037] The image sequence acquisition module is configured to acquire an image video stream of spatial target motion and sample it into a continuous video frame sequence.

[0038] The local information acquisition module is configured to acquire the feature positions of spatial targets in each video frame and encode the feature positions into binary description vectors; calculate the Euclidean distance between multiple binary description vectors in all video frames; and for each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as preliminary matching relationships.

[0039] The global information acquisition module is configured to input the video frame sequence into a pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position; calculate the Euclidean distance between the multi-dimensional vector corresponding to each feature position in the previous video frame and the multi-dimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame; and select the feature position corresponding to the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame.

[0040] The target tracking module is configured to solve the spatial target motion based on the best matching relationship of feature positions in order to track the spatial target.

[0041] Furthermore, the specific steps in the local information acquisition module for acquiring the feature positions of spatial targets in each video frame and encoding the feature positions into binary description vectors include:

[0042] Step 2: For each video frame, calculate the Harris feature points of the image;

[0043] Step 22: Set up the non-maximum suppression region, filter Harris feature points according to the non-maximum suppression region, and use the filtered Harris feature points as the spatial target feature locations.

[0044] Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region;

[0045] Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

[0046] Furthermore, the specific steps in the local information acquisition module for calculating the main direction of the region for each feature position of the spatial target, and dividing the feature position into multiple sub-regions according to the main direction of the region, include:

[0047] Extract a circular region M centered at feature location O with radius r. O Calculate the circular region M O Pixel value centroid C:

[0048]

[0049] Where I(x,y) is the pixel value corresponding to pixel (x,y) in the region;

[0050] Connect the feature location O with the centroid C to construct the main direction of the region. For the circular region M O Calculate the value of any point P in the middle. and included angle ∠COP;

[0051] According to ∠COP, the circular region M O The system is divided into m quadrants, and within each quadrant, it is further divided into n intervals based on the distance from the pixel to the feature location O, thus dividing the feature location into multiple sub-regions.

[0052] Furthermore, the specific steps in the local information acquisition module to encode each sub-region of the feature location and concatenate the encoded description vectors to obtain a binary description vector include:

[0053] For circular region M O Histogram equalization is performed on all pixels, and the equalized circular region M is calculated. O Average pixel value;

[0054] The proportional interval [0, 100%] is uniformly divided into p1 intervals. A corresponding binary code sequence B1 is assigned to each interval, where adjacent intervals differ by only one bit. The encoding length k1 of the binary code sequence B1 satisfies... The smallest integer solution;

[0055] Calculate the circular region M O The proportion of pixels in each sub-region that are higher than the average pixel value According to pixel ratio The corresponding proportional range determines the value of the binary code sequence B1 for that sub-region;

[0056] Calculate the sub-centroid of each sub-region

[0057]

[0058] according to get The angle between the angle along the clockwise direction and its nearest quadrant axis is determined based on the number of quadrants (m) in steps two and three, with the upper limit of this angle being... The included angle range The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2, with the following encoding rules: adjacent interval code sequences differ by only one bit, and the first and last interval codes also differ by only one bit; the encoding length k2 of the binary code sequence B2 satisfies... The smallest integer solution;

[0059] Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region;

[0060] By concatenating B1 and B2, a circular region M is obtained. O The encoding of each sub-region;

[0061] By concatenating the codes of all sub-regions in order, we obtain a (k1+k2)mn-bit binary code, which is the binary description vector.

[0062] Furthermore, the structural design of the neural network in the global information acquisition module is as follows:

[0063] The neural network has a total of 8 layers. The first 7 layers all use the form of "2D convolution + ReLU activation function", and the last layer only uses 2D convolution. The output dimensions of each layer are 8, 16, 32, 32, 64, 64, 128 and 128 respectively. The third and sixth layers use a convolution kernel of size 5, a stride of 2 and edge padding of 1. The convolution kernel, stride kernel and edge padding of the remaining layers are set to 3, 1 and 1 respectively.

[0064] The beneficial technical effects of this invention are:

[0065] This invention proposes a spatial target motion tracking method and system that combines local and global information. For each video frame, the feature positions of the spatial target are obtained and encoded into binary description vectors. The Euclidean distance between multiple binary description vectors in all video frames is calculated. For each feature position in the previous video frame, the multiple feature positions with the smallest Euclidean distance in the subsequent video frame are determined as preliminary matching relationships. The video frame sequence is input into a pre-trained neural network to obtain the multidimensional vector corresponding to each feature position. The Euclidean distance between the multidimensional vector corresponding to each feature position in the previous video frame and the multiple multidimensional vectors corresponding to multiple feature positions in the preliminary matching relationship in the subsequent video frame is calculated. The feature position with the smallest Euclidean distance in the subsequent video frame is selected as the optimal matching relationship for the feature position in the previous video frame. The spatial target motion is solved based on the optimal matching relationship of the feature positions to track the spatial target. Encoding feature locations into binary description vectors improves the computational efficiency of calculating the Euclidean distance between feature locations in images. The description vectors proposed in this invention are shorter and have lower storage consumption. After constructing multiple preliminary matching results for each feature location using binary description vectors, the Euclidean distance of the preliminary matching results is calculated using a neural network for verification. Only the matching relationship with the smallest distance is retained, reducing the computational load of the matching process. Attached Figure Description

[0066] The present invention can be better understood by referring to the description given below in conjunction with the accompanying drawings, which together with the following detailed description are included in and form part of this specification, and are used to further illustrate preferred embodiments of the invention and explain the principles and advantages of the invention.

[0067] Figure 1 This is a flowchart of a spatial target motion tracking method that combines local and global information according to an embodiment of the present invention;

[0068] Figure 2 This is a schematic diagram of the generation of sub-centroids of each region of the feature points in an embodiment of the present invention;

[0069] Figure 3 This is an example diagram illustrating the encoding order of the included angle information in the local information encoding of this invention.

[0070] Figure 4 This is a schematic diagram of global information generation in an embodiment of the present invention;

[0071] Figure 5 This is an example diagram of the feature position matching results between consecutive frames in an embodiment of the present invention;

[0072] Figure 6 This is a graph showing the relationship between the accuracy of consecutive inter-frame feature matching and the target motion angle in an embodiment of the present invention. Detailed Implementation

[0073] To enable those skilled in the art to better understand the present invention, exemplary embodiments or examples of the present invention will be described below in conjunction with the accompanying drawings. Obviously, the described embodiments or examples are merely some, not all, of the embodiments or examples of the present invention. All other embodiments or examples obtained by those skilled in the art based on the embodiments or examples of the present invention without inventive effort should fall within the scope of protection of the present invention.

[0074] This invention proposes a spatial target motion tracking method that combines local and global information, such as... Figure 1 As shown, the method includes the following steps:

[0075] Step 1: Acquire the image and video stream of the spatial target motion and sample it into a continuous video frame sequence;

[0076] Step 2: For each video frame, obtain the feature positions of spatial targets and encode the feature positions into binary description vectors;

[0077] Step 3: Calculate the Euclidean distance between multiple binary description vectors in all video frames. For each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as the preliminary matching relationship.

[0078] Step 4: Input the video frame sequence into the pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position;

[0079] Step 5: Calculate the Euclidean distance between the multidimensional vector corresponding to each feature position in the previous video frame and the multidimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame. Select the feature position with the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame.

[0080] Step 6: Solve the motion of the spatial target based on the best matching relationship of the feature positions in order to track the spatial target.

[0081] As a preferred embodiment of the present invention, the specific steps of step two are as follows:

[0082] Step 2: For each video frame, convert the image to grayscale and calculate the Harris feature points of the image.

[0083] Step 22: Set up a non-maximum suppression region and filter Harris feature points based on the non-maximum suppression region. Use the filtered feature points as the spatial target feature locations. For example, the process of filtering Harris feature points based on the non-maximum suppression region is as follows: take the maximum value position in sequence according to the Harris response amplitude, set the response of all pixels in the non-maximum suppression region around the position to 0, and repeat this process to filter Harris feature points and obtain the feature locations of spatial targets.

[0084] Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region;

[0085] According to an embodiment of the present invention, each image feature location is divided into multiple sub-regions based on the main direction to ensure that the subsequent binary description vector remains stable with respect to image rotation. The specific division process is as follows:

[0086] Extract a circular region M centered at feature location O with a radius r (e.g., 20 pixels). O For M O If histogram equalization is performed on all pixels to enhance texture contrast, then the grayscale image with a grayscale range of [0, 255] after equalization is represented as follows:

[0087]

[0088] Where, n j g is the number of pixels with grayscale value j in the region. k This represents the grayscale value of a pixel after equalization.

[0089] For the circular region M O The centroid C of the pixel values ​​in the calculation region:

[0090]

[0091] Where x, y are the coordinates of all pixels in the region, and I(x, y) is the pixel value corresponding to pixel (x, y) in the region.

[0092] Connect the feature location O with the centroid C to construct the main direction of the region. For M O Calculate the value at any position P in the matrix. and included angle:

[0093]

[0094] Furthermore, since the range of the arccos function is [0, π), in order to determine the quadrant to which P belongs, we will... Expanding to three dimensions, that is:

[0095]

[0096] calculate cross product remember The last dimension has a value of t. Then... and The included angle ∠COP is represented as:

[0097]

[0098] According to ∠COP, M O The system is divided into m quadrants. Within each quadrant, the system is further divided into n intervals based on the distance from the pixel to the feature location O, ensuring that each sub-region has an equal area. Each sub-region is denoted as S. i (i = 1, ..., mn). For example... Figure 2 As shown, it is divided into 4 quadrants, each quadrant is further divided into... There are a total of 3 intervals, which divide the area around the feature location O into a total of 12 sub-regions.

[0099] Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

[0100] According to an embodiment of the present invention, each feature position is encoded as a (k1+k2)mn-bit binary description vector, where k1 and k2 are the lengths of encodings B1 and B2, respectively. The specific process is as follows:

[0101] For feature location O, there are a total of mn sub-regions, each corresponding to B1 and B2 codes. First, calculate the B1 code for each sub-region: calculate the equalized region M. O The average pixel value; the ratio interval [0, 100%] is uniformly divided into p1 intervals, and a corresponding binary code sequence B1 is set for each interval. B1 satisfies that the code sequences of adjacent intervals differ by only one bit, and its encoding length k1 satisfies the following condition. Find the smallest integer solution; count the proportion of pixels in each sub-region that are higher than the average pixel value. According to pixel ratio The value of the binary code sequence B1 corresponding to the sub-region is determined by the proportion range to which it belongs.

[0102] Taking p1 = 8 as an example, we can find that the minimum value of k1 is 3. The corresponding codes for each interval are as follows:

[0103]

[0104]

[0105] Then, calculate the B2 code for each sub-region: calculate the sub-centroid of each sub-region:

[0106]

[0107] according to get The angle between the angle and its nearest quadrant axis along the clockwise direction; determine the upper limit of the angle based on the number of quadrants m divided in steps two and three. like Figure 3 As shown, the interval The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2. The encoding rule is as follows: in addition to ensuring that the code sequences of adjacent intervals differ by only one bit, the first and last intervals must also differ by only one bit. The encoding length k2 of B2 is... The smallest integer solution;

[0108] Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region.

[0109] Taking the number of quadrants m=4 and the number of intervals p2=8 as an example: Since there are 4 quadrants, the upper limit of the included angle... Divide the interval [0°, 90°] into 8 included angle intervals, and the corresponding codes for each interval are as follows:

[0110]

[0111] By concatenating the codes of all sub-regions in sequence, it can be represented as a (k1+k2)mn-bit binary code. Taking the above example, the region is divided into 4 quadrants, each quadrant is divided into 3 sub-regions, and the B1 and B2 codes of each region are both 3 bits. Then the binary description vector of this feature position has a total of (3+3)×4×3=72 bits.

[0112] As a preferred embodiment of the present invention, step three is as follows: For two consecutive frames in a video frame sequence, feature positions and corresponding binary description vectors are extracted according to the above steps, and preliminary matching relationships are constructed by calculating the Euclidean distance of all description vectors in different images. For each feature position in the previous frame, several (e.g., three) preliminary matching relationships with the smallest Euclidean distance in the subsequent frame are retained. That is, for two consecutive frames, for each feature position in the previous frame, the Euclidean distance between the feature vector corresponding to that position and all feature description vectors in the subsequent frame is calculated. For binary vectors, the Euclidean distance is calculated by performing a bitwise XOR operation and then summing the results. For each feature position, the three feature positions with the smallest Euclidean distance in the subsequent frame are retained.

[0113] As a preferred embodiment of the present invention, step four is as follows: The two frames of images are scaled to 640×640, and the images are input into the trained neural network. Interpolation is used to obtain the 128-dimensional vector corresponding to each feature position. For an image with an original size of (3, H, W), it is first scaled to (3, 640, 640) and input into the neural network, outputting a three-dimensional matrix of size (128, 160, 160). This matrix is ​​then upsampled to (128, H, W) through interpolation. The 128-dimensional information of the upsampled matrix is ​​used as the description vector for each pixel position in the image, such as... Figure 4 As shown.

[0114] The neural network is designed as follows: The neural network has a total of 8 layers. The first 7 layers all use a "2D convolution + ReLU activation function" approach, while the last layer uses only 2D convolution. The output dimensions of each layer are 8 / 16 / 32 / 32 / 64 / 64 / 128 / 128, respectively. The third and sixth layers use a convolution kernel of size 5, a stride of 2, and edge padding of 1. The kernel, stride, and edge padding settings for the remaining layers are 3, 1, and 1, respectively.

[0115] After obtaining a three-dimensional matrix of size (128, H, W), the following normalization operation is performed: the 128-dimensional description information of the three-dimensional matrix corresponding to H=i, W=j is denoted as n. ij Calculate n ij The 2-norm ||n ij ||2, that is, calculate n ij Take the square root of the sum of the squares of all elements in the matrix. Then use... Replace the original n ij This serves as global descriptive information for the image pixel coordinates (i, j).

[0116] In a preferred embodiment of the present invention, step five specifically includes: for any feature position coordinates (u, v) in the previous frame, the multiple (e.g., three) preliminary matching relationships obtained in step three are represented as (u1, v1), (u2, v2), and (u3, v3); the description vector n of the previous frame I1 image at position (u, v) is obtained through step four. uv The description vector of the subsequent frame image I2 at the corresponding matching feature position is represented as follows: Calculate n respectively uv The Euclidean distances to the three vectors mentioned above are used to determine the best matching relationship, with only the feature position corresponding to the smallest Euclidean distance being retained.

[0117] In a preferred embodiment of the present invention, step six specifically includes: recovering the essential matrix through the correspondence of feature positions, and obtaining relative motion information by decomposing the essential matrix. Specifically, it is assumed that a total of N sets of matching feature positions are selected from the two frames of images, denoted as follows: and Find the least-squares solution to the following overdetermined system of equations:

[0118]

[0119] Where the matrix Performing SVD decomposition on the matrix yields E = U∑V T Where U and V are orthogonal matrices.

[0120] Let the relative motion of the target be represented by the rotation matrix R and the translation matrix t = [t1, t2, t3]. T Based on the SVD decomposition results, we can obtain R = UWV T or UW T V T , or UZ T V T There are a total of 4 solutions, among which The sign of the projected depth of the spatial coordinate point is verified to obtain a unique solution.

[0121] The technical effects of the present invention were further verified through experiments.

[0122] The experiment combined two target models (box and satellite models) for image simulation and analysis, and the feature location matching results are as follows: Figure 5 As shown. Analysis was performed on 100 consecutive simulated images, each with a size of 640×480. Ten different initial attitudes were simulated using simulation software, with the pitch / yaw / roll angles of the model simultaneously increased by a step size of 1°. The average matching accuracy of feature points was calculated as follows: Figure 6 As shown. From Figure 5 and Figure 6 It can be seen that the present invention can effectively construct the correct matching relationship of similar texture features in an image, and also has a high matching accuracy for the increase in disparity caused by target rotation.

[0123] Another embodiment of the present invention proposes a spatial target motion tracking system that combines local and global information, the system comprising:

[0124] The image sequence acquisition module is configured to acquire an image video stream of spatial target motion and sample it into a continuous video frame sequence.

[0125] The local information acquisition module is configured to acquire the feature positions of spatial targets in each video frame and encode the feature positions into binary description vectors; calculate the Euclidean distance between multiple binary description vectors in all video frames; and for each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as preliminary matching relationships.

[0126] The global information acquisition module is configured to input the video frame sequence into a pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position; calculate the Euclidean distance between the multi-dimensional vector corresponding to each feature position in the previous video frame and the multi-dimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame; and select the feature position corresponding to the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame.

[0127] The target tracking module is configured to solve the spatial target motion based on the best matching relationship of feature positions in order to track the spatial target.

[0128] In a preferred embodiment of the present invention, the specific steps of the local information acquisition module for acquiring the feature positions of spatial targets in each video frame and encoding the feature positions into binary description vectors include:

[0129] Step 2: For each video frame, calculate the Harris feature points of the image;

[0130] Step 22: Set up the non-maximum suppression region, filter Harris feature points according to the non-maximum suppression region, and use the filtered Harris feature points as the spatial target feature locations.

[0131] Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region;

[0132] Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

[0133] In a preferred embodiment of the present invention, the specific steps of the local information acquisition module for calculating the main direction of the region for each feature position of the spatial target, and dividing the feature position into multiple sub-regions according to the main direction of the region, include:

[0134] Extract a circular region M centered at feature location O with radius r. O Calculate the circular region M O Pixel value centroid C:

[0135]

[0136] Where I(x,y) is the pixel value corresponding to pixel (x,y) in the region;

[0137] Connect the feature location O with the centroid C to construct the main direction of the region. For the circular region M O Calculate the value of any point P in the middle. and included angle ∠COP;

[0138] According to ∠COP, the circular region M O The system is divided into m quadrants, and within each quadrant, it is further divided into n intervals based on the distance from the pixel to the feature location O, thus dividing the feature location into multiple sub-regions.

[0139] In a preferred embodiment of the present invention, the local information acquisition module encodes each sub-region of the feature location and concatenates the encoded description vectors to obtain a binary description vector, including the following specific steps:

[0140] For circular region M O Histogram equalization is performed on all pixels, and the equalized circular region M is calculated. O Average pixel value;

[0141] The proportional interval [0, 100%] is uniformly divided into p1 intervals. A corresponding binary code sequence B1 is assigned to each interval, where adjacent intervals differ by only one bit. The encoding length k1 of the binary code sequence B1 satisfies... The smallest integer solution;

[0142] Calculate the circular region M O The proportion of pixels in each sub-region that are higher than the average pixel value According to pixel ratio The corresponding proportional range determines the value of the binary code sequence B1 for that sub-region;

[0143] Calculate the sub-centroid of each sub-region

[0144]

[0145] according to get The angle between the angle along the clockwise direction and its nearest quadrant axis is determined based on the number of quadrants (m) in steps two and three, with the upper limit of this angle being... The included angle range The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2, with the following encoding rules: adjacent interval code sequences differ by only one bit, and the first and last interval codes also differ by only one bit; the encoding length k2 of the binary code sequence B2 satisfies... The smallest integer solution;

[0146] Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region;

[0147] By concatenating B1 and B2, a circular region M is obtained. O The encoding of each sub-region;

[0148] By concatenating the codes of all sub-regions in order, we obtain a (k1+k2)mn-bit binary code, which is the binary description vector.

[0149] As a preferred embodiment of the present invention, the structure design of the neural network in the global information acquisition module is as follows: the neural network has a total of 8 layers, the first 7 layers all adopt the form of "two-dimensional convolution + ReLU activation function", and the last layer only uses two-dimensional convolution; the output dimensions of each layer are 8, 16, 32, 32, 64, 64, 128 and 128 respectively, the third and sixth layers use a convolution kernel of size 5, a stride of 2, and edge padding of 1, and the convolution kernel, stride kernel and edge padding of the remaining layers are all set to 3, 1 and 1.

[0150] The functionality of a spatial target motion tracking system that combines local and global information in this embodiment of the invention can be described by the aforementioned spatial target motion tracking method that combines local and global information. Therefore, for the parts not detailed in the system embodiment, please refer to the above method embodiment, and they will not be repeated here.

[0151] Although the invention has been described with respect to a limited number of embodiments, those skilled in the art will understand from the foregoing description that other embodiments are conceivable within the scope of the invention described herein. The disclosure of the invention is illustrative and not restrictive, and the scope of the invention is defined by the appended claims.

Claims

1. A spatial target motion tracking method combining local and global information, characterized in that, Includes the following steps: Step 1: Acquire the image and video stream of the spatial target motion and sample it into a continuous video frame sequence; Step 2: For each video frame, obtain the feature positions of spatial targets and encode the feature positions into binary description vectors; Step 3: Calculate the Euclidean distance between multiple binary description vectors in all video frames. For each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as the preliminary matching relationship. Step 4: Input the video frame sequence into the pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position; Step 5: Calculate the Euclidean distance between the multidimensional vector corresponding to each feature position in the previous video frame and the multidimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame. Select the feature position with the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame. Step 6: Solve the motion of the spatial target based on the best matching relationship of the feature positions in order to track the spatial target.

2. The spatial target motion tracking method combining local and global information according to claim 1, characterized in that, Step two includes the following steps: Step 2: For each video frame, calculate the Harris feature points of the image; Step 22: Set up the non-maximum suppression region, filter Harris feature points according to the non-maximum suppression region, and use the filtered Harris feature points as the spatial target feature locations. Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region; Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

3. The spatial target motion tracking method combining local and global information according to claim 2, characterized in that, Step 22 involves filtering Harris feature points based on non-maximum suppression regions. The specific process includes: taking the maximum value positions in sequence according to the Harris response amplitude, setting the response values ​​of all pixels in the non-maximum suppression regions around the maximum value positions to 0, and repeating the above process to filter Harris feature points.

4. The spatial target motion tracking method combining local and global information according to claim 2, characterized in that, The specific steps in steps two and three include: Extract a circular region M centered at feature location O with radius r. O Calculate the circular region M O Pixel value centroid C: Where I(x,y) is the pixel value corresponding to pixel (x,y) in the region; Connect the feature location O with the centroid C to construct the main direction of the region. For the circular region M O Calculate the value of any point P in the middle. and included angle ∠COP; Based on ∠COP, the circular region M is... O The system is divided into m quadrants, and within each quadrant, it is further divided into n intervals based on the distance from the pixel to the feature location O, thus dividing the feature location into multiple sub-regions.

5. A spatial target motion tracking method combining local and global information according to claim 4, characterized in that, The specific steps in step two and four include: For circular region M O Histogram equalization is performed on all pixels, and the equalized circular region M is calculated. O Average pixel value; The proportional interval [0, 100%] is uniformly divided into p1 intervals. A corresponding binary code sequence B1 is assigned to each interval, where adjacent intervals differ by only one bit. The encoding length k1 of the binary code sequence B1 satisfies... The smallest integer solution; Calculate the circular region M O The proportion of pixels in each sub-region that are higher than the average pixel value According to pixel ratio The corresponding proportional range determines the value of the binary code sequence B1 for that sub-region; Calculate the sub-centroid of each sub-region according to get The angle between the angle along the clockwise direction and its nearest quadrant axis is determined based on the number of quadrants (m) in steps two and three, with the upper limit of this angle being... The included angle range The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2, with the following encoding rules: adjacent interval code sequences differ by only one bit, and the first and last interval codes also differ by only one bit; the encoding length k2 of the binary code sequence B2 satisfies... The smallest integer solution; Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region; By concatenating B1 and B2, a circular region M is obtained. O The encoding of each sub-region; By concatenating the codes of all sub-regions in order, we obtain a (k1+k2)mn-bit binary code, which is the binary description vector.

6. The spatial target motion tracking method combining local and global information according to claim 1, characterized in that, The neural network structure design described in step four is as follows: The neural network has a total of 8 layers. The first 7 layers all use the form of "2D convolution + ReLU activation function", and the last layer only uses 2D convolution. The output dimensions of each layer are 8, 16, 32, 32, 64, 64, 128 and 128 respectively. The third and sixth layers use a convolution kernel of size 5, a stride of 2 and edge padding of 1. The convolution kernel, stride kernel and edge padding of the remaining layers are set to 3, 1 and 1 respectively.

7. A spatial target motion tracking system that combines local and global information, characterized in that, include: The image sequence acquisition module is configured to acquire an image video stream of spatial target motion and sample it into a continuous video frame sequence. The local information acquisition module is configured to acquire the feature positions of spatial targets in each video frame and encode the feature positions into binary description vectors; calculate the Euclidean distance between multiple binary description vectors in all video frames; and for each feature position in the previous video frame, determine the multiple feature positions with the smallest Euclidean distance in the next video frame as preliminary matching relationships. The global information acquisition module is configured to input the video frame sequence into a pre-trained neural network to obtain the multi-dimensional vector corresponding to each feature position; calculate the Euclidean distance between the multi-dimensional vector corresponding to each feature position in the previous video frame and the multi-dimensional vector corresponding to multiple feature positions in the preliminary matching relationship in the next video frame; and select the feature position corresponding to the smallest Euclidean distance in the next video frame as the best matching relationship of the feature position in the previous video frame. The target tracking module is configured to solve the spatial target motion based on the best matching relationship of feature positions in order to track the spatial target.

8. A spatial target motion tracking system combining local and global information according to claim 7, characterized in that, The specific steps in the local information acquisition module for each video frame to acquire the feature positions of spatial targets and encode the feature positions into binary description vectors include: Step 2: For each video frame, calculate the Harris feature points of the image; Step 22: Set up the non-maximum suppression region, filter Harris feature points according to the non-maximum suppression region, and use the filtered Harris feature points as the spatial target feature locations. Steps 2 and 3: For each feature location of the spatial target, calculate the main direction of the region, and divide the feature location into multiple sub-regions according to the main direction of the region; Step 24: Encode each sub-region of the feature location separately, and concatenate the encoded description vectors to obtain a binary description vector.

9. A spatial target motion tracking system combining local and global information according to claim 8, characterized in that, The specific steps in the local information acquisition module for calculating the main direction of the region for each feature position of the spatial target and dividing the feature position into multiple sub-regions according to the main direction of the region include: Extract a circular region M centered at feature location O with radius r. O Calculate the circular region M O Pixel value centroid C: Where I(x,y) is the pixel value corresponding to pixel (x,y) in the region; Connect the feature location O with the centroid C to construct the main direction of the region. For the circular region M O Calculate the value of any point P in the middle. and included angle ∠COP; According to ∠COP, the circular region M O The system is divided into m quadrants, and within each quadrant, it is further divided into n intervals based on the distance from the pixel to the feature location O, thus dividing the feature location into multiple sub-regions.

10. A spatial target motion tracking system combining local and global information according to claim 9, characterized in that, The specific steps in the local information acquisition module to encode each sub-region of the feature location and concatenate the encoded description vectors to obtain a binary description vector include: For circular region M O Histogram equalization is performed on all pixels, and the equalized circular region M is calculated. O Average pixel value; The proportional interval [0, 100%] is uniformly divided into p1 intervals. A corresponding binary code sequence B1 is assigned to each interval, where adjacent intervals differ by only one bit. The encoding length k1 of the binary code sequence B1 satisfies... The smallest integer solution; Calculate the circular region M O The proportion of pixels in each sub-region that are higher than the average pixel value According to pixel ratio The corresponding proportional range determines the value of the binary code sequence B1 for that sub-region; Calculate the sub-centroid of each sub-region according to get The angle between the angle along the clockwise direction and its nearest quadrant axis is determined based on the number of quadrants (m) in steps two and three, with the upper limit of this angle being... The included angle range The code is evenly divided into p2 intervals. Each included interval is encoded as a binary code sequence B2, with the following encoding rules: adjacent interval code sequences differ by only one bit, and the first and last interval codes also differ by only one bit; the encoding length k2 of the binary code sequence B2 satisfies... The smallest integer solution; Based on the sub-centroid of each sub-region corresponding The included angle interval determines the value of the binary code sequence B2 corresponding to the sub-region; By concatenating B1 and B2, a circular region M is obtained. O The encoding of each sub-region; By concatenating the codes of all sub-regions in order, we obtain a (k1+k2)mn-bit binary code, which is the binary description vector.

Citation Information

Patent Citations

  • Medical endoscope continuous frame image feature point matching method and device

    CN113538540A

  • Apparatus and method for processing video data

    WO2006055512A2