A method and system for tracking object motion trajectories based on visual features

By employing a visual feature-based object motion trajectory tracking method, which utilizes visual feature extraction and cross-attention mechanisms combined with 3D point cloud data processing, the problem of low accuracy and high computational complexity in existing object motion trajectory tracking technologies is solved, achieving higher tracking accuracy and robustness.

CN120411174BActive Publication Date: 2025-11-14BEIJING HANGYU CHUANGTONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510905350.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-11-14
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing object motion trajectory tracking methods have low accuracy and high computational complexity in multi-target or complex background environments, making it difficult to meet real-time requirements.

Method used

The object motion trajectory tracking method based on visual features extracts visual features from the target object search image information and template image information, extracts temporal features using cross-attention mechanism and Transformer encoder-decoder, and combines 3D point cloud data for inverse density sampling to generate the object motion trajectory vector.

Benefits of technology

It improves the accuracy and robustness of object motion trajectory tracking, effectively copes with occlusion and complex scenes, and maintains high tracking accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411174B_ABST
    Figure CN120411174B_ABST
Patent Text Reader

Abstract

This invention relates to the field of motion trajectory tracking technology, specifically to a method and system for object motion trajectory tracking based on visual features. The method includes: acquiring continuous target object search image information and target object template image information based on an object motion video stream; performing information interaction processing on template features and search features; obtaining candidate predicted target object bounding boxes based on temporal features and the information-interacted search features; determining whether the candidate predicted target object bounding boxes meet the selection accuracy requirements based on the average coverage of the candidate predicted target object bounding boxes; extracting 3D point cloud data of the target objects within the predicted target object bounding boxes; performing inverse density sampling processing on the 3D point cloud data to obtain a set of points to be processed; and generating an object motion trajectory vector based on the set of points to be processed. This method can improve the accuracy and robustness of object motion trajectory tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of motion trajectory tracking technology, specifically to a method and system for tracking object motion trajectories based on visual features. Background Technology

[0002] With the rapid development of computer vision technology, object motion trajectory tracking methods based on visual features have become key technologies in many fields, including autonomous driving, intelligent monitoring, virtual reality, and human-computer interaction.

[0003] Existing tracking methods and their limitations:

[0004] Frame differencing detects moving objects by comparing pixel differences between consecutive image frames. However, it is sensitive to changes in lighting and background interference, making it prone to false positives, especially in multi-object or complex background environments. Background subtraction detects moving objects by building a background model and comparing it with the current frame. This method works well in scenes where the camera is stationary and the background is stable, but it struggles to adapt to complex situations such as changes in lighting and dynamic background updates. Furthermore, background subtraction cannot effectively distinguish between moving objects and their shadows, leading to inaccurate detection results. Optical flow infers the trajectory of objects by analyzing the motion vectors of pixels in an image sequence. This method can handle problems such as changes in lighting and background interference, but its accuracy drops significantly when objects move rapidly or when motion blur is present. In addition, optical flow has high computational complexity, making it difficult to meet the real-time requirements of applications.

[0005] To address the above problems, this invention proposes a method and system for tracking object motion trajectories based on visual features. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for tracking object motion trajectories based on visual features: to solve the technical problems of low accuracy and high computational complexity in existing tracking methods in multi-target or complex background environments.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] On the one hand, there are methods for tracking object motion trajectories based on visual features, including:

[0009] Based on the acquisition of continuous target object search image information and target object template image information from the object motion video stream, visual features are extracted from the target object search image information and target object template image information to obtain template features and search features, and information interaction processing is performed on the template features and search features.

[0010] Temporal features are extracted based on search features and cross-attention mechanism. Based on the temporal features and the search features after information interaction processing, the bounding boxes of the candidate predicted target objects are obtained. Based on the average coverage of the bounding boxes of the candidate predicted target objects, it is determined whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements. If so, the bounding boxes of the candidate predicted target objects are marked as the predicted target object bounding boxes.

[0011] Extract the 3D point cloud data of the target object within the bounding box of the predicted target object, perform inverse density sampling on the 3D point cloud data to obtain the point set to be processed, and generate the object motion trajectory vector based on the point set to be processed.

[0012] Furthermore, visual feature extraction is performed on the target object search image information and the target object template image information to obtain template features and search features. Specifically, this includes the following processes:

[0013] Image information for target object search includes images containing both background and target objects. The target object template image information includes the image of the target object. ,in, Represents the set of real numbers. , Images ,image height, , Images ,image The width is 3, where 3 represents the three color channels of the image (RGB).

[0014] Image and images Each image is split into multiple non-overlapping image patch sequences, denoted as follows: , ,in, , These are the dimensions of the image patches, The size of each image patch is represented by [insert image patch size here]. All image patches are flattened, and the image embedding representation is obtained through a linear mapping function, denoted as [insert image embedding representation here]. , ,in, , The length of the embedded representation, , To define the dimension of the embedding representation, the embedding representations of the search image and template image are input into the Transformer encoding layer for feature extraction. The extracted search features and template features are... , .

[0015] Furthermore, the information interaction processing of template features and search features specifically includes the following steps:

[0016] The template features and search features are unified in dimensionality through a 1×1 convolution operation:

[0017] ; ;

[0018] in, , These are the feature vectors after unifying the dimensions of the search features and template features, respectively.

[0019] based on and Generate graph structure ,in, To search for a set of nodes in the image information representing the feature information of the target object, the node set... , For the number of nodes, Search for the edge set representing the feature information of the target object in the image information. The number of sides is ;

[0020] compute nodes The size of the d-th eigenvalue : ;

[0021] in, For nodes Features Represents the measure of node centrality;

[0022] based on The interaction probability of the attribute representation feature of the compute node : ;

[0023] in, , for The maximum value, for The average value, Hyperparameters for controlling the overall size of feature enhancement;

[0024] By importance Delete edges from a graph structure, where, This yields search features that have undergone information interaction processing.

[0025] Furthermore, the extraction of temporal features based on search features and cross-attention mechanisms specifically includes the following processes:

[0026] In the Transformer decoder, the search features are linearly projected onto the key and value vectors, representing respectively... , And introduce learnable time series information queries. By calculating the cross-attention between search features and learnable temporal information queries, temporal information is aggregated to obtain temporal features. :

[0027] ;in, express transpose, Representing time series features Dimensions The function normalizes the original attention weights into a probability distribution, so that these weights are interpreted as the relative importance of different positions.

[0028] Furthermore, obtaining the candidate predicted target object bounding box based on temporal features and search features processed through information interaction specifically includes the following process:

[0029] The search features F after information interaction processing and the learnable target query and time series characteristics The input is fed into a target decoder consisting of a cross-attention and a feedforward network, and the final target features are obtained by calculating the cross-attention. , means as follows:

[0030] ; ; ;

[0031] ;

[0032] ;in, Representation of features The dimension; Indicates a feedforward network. Presentation layer normalization processing.

[0033] Furthermore, calculating the average coverage of the bounding boxes of the candidate predicted target objects specifically includes the following process:

[0034] ;

[0035] This represents the bounding box of the candidate predicted target object in the i-th frame of the j-th motion video. This represents the ground truth bounding box of the i-th frame in the j-th motion video. Indicates the number of motion videos. The number of frames in each motion video, This is an indicator function; when the overlap rate between the candidate predicted object bounding box and the ground truth object bounding box is greater than a threshold... The value is 1 if it is true, and 0 otherwise.

[0036] Furthermore, determining whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements based on the average coverage of the bounding boxes specifically includes the following process:

[0037] Load the average coverage threshold and determine whether the average coverage of the bounding boxes of the candidate predicted target objects exceeds the average coverage threshold. If yes, the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements. If no, the bounding boxes of the candidate predicted target objects do not meet the selection accuracy requirements.

[0038] Furthermore, the inverse density sampling process for the 3D point cloud data to obtain the point set to be processed specifically includes the following steps:

[0039] Set the neighborhood hyperparameter K. For the input 3D point cloud P, calculate the distance between each point in the point cloud, and select K neighborhood points as local regions for each point in the point cloud based on the distance.

[0040] Calculate the average distance between each point in the 3D point cloud P and its neighboring points. This distance is negatively correlated with the local point distribution density. That is, the higher the local point density, the smaller the calculated average distance, and vice versa.

[0041] Based on the average distance point by point, the sampling probability is calculated, and NS points are sampled according to the calculated probability and added to the point set PS. The point set PS is denoted as the point set to be processed.

[0042] Furthermore, generating the object's motion trajectory vector based on the set of points to be processed specifically includes the following process:

[0043] An ultrasonic probe coordinate system is introduced. The coordinates of the points to be processed are determined by taking the top midpoint of the image frame containing the target object as the origin of the ultrasonic probe coordinate system. The unique coordinates of each point in the ultrasonic probe coordinate system are then obtained. The three-dimensional coordinates of the point in the coordinate system are obtained through coordinate system transformation. The coordinate transformation formula is as follows:

[0044] ;

[0045] in, The transformation matrix depends on the relative positions between the points in the set of points to be processed and can be preset freely. It generates the object's motion trajectory vector based on the three-dimensional coordinate values.

[0046] On the other hand, a visual feature-based object motion trajectory tracking system, applicable to any of the visual feature-based object motion trajectory tracking methods described above, includes:

[0047] The feature interaction module is used to acquire continuous target object search image information and target object template image information based on the object motion video stream, extract visual features from the target object search image information and target object template image information to obtain template features and search features, and perform information interaction processing on the template features and search features.

[0048] The bounding box determination module is used to extract temporal features based on search features and cross-attention mechanism, obtain candidate predicted target object bounding boxes based on temporal features and search features after information interaction processing, and determine whether candidate predicted target object bounding boxes meet the selection accuracy requirements based on the average coverage of candidate predicted target object bounding boxes. If so, the candidate predicted target object bounding boxes are marked as predicted target object bounding boxes.

[0049] The motion trajectory tracking module is used to extract the 3D point cloud data of the target object in the bounding box of the predicted target object, perform inverse density sampling on the 3D point cloud data to obtain the point set to be processed, and generate the motion trajectory vector of the object based on the point set to be processed.

[0050] Compared to existing solutions, the beneficial effects achieved by this invention are:

[0051] This invention acquires continuous target object search image information and target object template image information from object motion video streams, and performs information interaction processing on template features and search features. Based on temporal features and the information-interactive search features, candidate predicted target object bounding boxes are obtained. The average coverage of the candidate predicted target object bounding boxes is used to determine whether the candidate predicted target object bounding boxes meet the selection accuracy requirements. The 3D point cloud data of the target objects in the predicted target object bounding boxes is extracted, and inverse density sampling processing is performed on the 3D point cloud data to obtain a set of points to be processed. Based on the set of points to be processed, an object motion trajectory vector is generated, which can improve the accuracy and robustness of object motion trajectory tracking.

[0052] Furthermore, this method can effectively handle occlusion and complex scenes:

[0053] Multi-frame information fusion: The method analyzes image information from consecutive frames and can use multi-frame information to infer the position of the target object. Even if the target object is occluded in some frames, it can be completed using information from other frames, thus maintaining the continuity of tracking.

[0054] Feature richness: By comprehensively utilizing a variety of visual features, the method can maintain high tracking accuracy even when the appearance of the target object changes (such as rotation or scaling) or is partially occluded. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0056] Figure 1 This is a flowchart illustrating the first method for tracking object motion trajectories based on visual features according to an embodiment of the present invention.

[0057] Figure 2 This is a flowchart illustrating the second method for tracking object motion trajectories based on visual features according to an embodiment of the present invention.

[0058] Figure 3 This is a system block diagram of an object motion trajectory tracking system based on visual features according to an embodiment of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0061] This embodiment provides a method for tracking object motion trajectories based on visual features. Figure 1 This is a flowchart illustrating the first object motion trajectory tracking method based on visual features according to an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes the following steps:

[0062] Step S101: Acquire continuous target object search image information and target object template image information based on the object motion video stream;

[0063] Step S102: Extract visual features from the target object search image information and the target object template image information to obtain template features and search features, and perform information interaction processing on the template features and search features;

[0064] Step S103: Extract temporal features based on search features and cross-attention mechanism, and obtain the bounding boxes of candidate predicted target objects based on temporal features and search features after information interaction processing;

[0065] Step S104: Determine whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements based on the average coverage of the bounding boxes of the candidate predicted target objects. If so, mark the bounding boxes of the candidate predicted target objects as the bounding boxes of the predicted target objects.

[0066] Step S105: Extract the 3D point cloud data of the target object in the bounding box of the predicted target object, perform inverse density sampling on the 3D point cloud data to obtain the point set to be processed, and generate the object motion trajectory vector based on the point set to be processed.

[0067] In summary, this invention acquires continuous target object search image information and target object template image information from object motion video streams, and performs information interaction processing on template features and search features. Based on temporal features and the search features processed by information interaction, it obtains candidate predicted target object bounding boxes. Based on the average coverage of the candidate predicted target object bounding boxes, it determines whether the candidate predicted target object bounding boxes meet the selection accuracy requirements. It extracts the 3D point cloud data of the target objects in the predicted target object bounding boxes, performs inverse density sampling processing on the 3D point cloud data to obtain a set of points to be processed, and generates an object motion trajectory vector based on the set of points to be processed. This can improve the accuracy and robustness of object motion trajectory tracking.

[0068] In some embodiments, visual feature extraction is performed on the target object search image information and the target object template image information to obtain template features and search features, specifically including the following process:

[0069] Image information for target object search includes images containing both background and target objects. The target object template image information includes the image of the target object. ,in, Represents the set of real numbers. , Images ,image height, , Images ,image The width is 3, where 3 represents the three color channels of the image (RGB).

[0070] Image and images Each image is split into multiple non-overlapping image patch sequences, denoted as follows: , ,in, , These are the dimensions of the image patches, The size of each image patch is represented by [insert image patch size here]. All image patches are flattened, and the image embedding representation is obtained through a linear mapping function, denoted as [insert image embedding representation here]. , ,in, , The length of the embedded representation, , To define the dimension of the embedding representation, the embedding representations of the search image and template image are input into the Transformer encoding layer for feature extraction. The extracted search features and template features are... , .

[0071] It is worth noting that by flattening all image patches and obtaining the image embedding representation through a linear mapping function, the image data is transformed into a vector form suitable for feature extraction, providing a foundation for subsequent visual feature extraction. Furthermore, the linear mapping can help remove redundant information, reduce computational complexity, and suppress noise to some extent. The embedding representation is usually the input form of deep learning models (such as convolutional neural networks, Transformers, etc.), which facilitates more complex feature learning and pattern recognition by the model and can improve the computational efficiency of object motion trajectory tracking.

[0072] In some embodiments, the information interaction processing of template features and search features specifically includes the following process:

[0073] The template features and search features are unified in dimensionality through a 1×1 convolution operation:

[0074] ; ;

[0075] in, , These are the feature vectors after unifying the dimensions of the search features and template features, respectively.

[0076] based on and Generate graph structure ,in, To search for a set of nodes in the image information representing the feature information of the target object, the node set... , For the number of nodes, Search for the edge set representing the feature information of the target object in the image information. The number of sides is ;

[0077] compute nodes The size of the d-th eigenvalue : ;

[0078] in, For nodes Features Represents the measure of node centrality;

[0079] based on The interaction probability of the attribute representation feature of the compute node : ;

[0080] in, , for The maximum value, for The average value, Hyperparameters for controlling the overall size of feature enhancement;

[0081] By importance Delete edges from a graph structure, where, This yields search features that have undergone information interaction processing.

[0082] In summary, information interaction processing of template features and search features can reduce the impact of image background on visual feature extraction and improve the accuracy of target object feature extraction.

[0083] In some embodiments, extracting temporal features based on search features and cross-attention mechanisms specifically includes the following process:

[0084] In the Transformer decoder, the search features are linearly projected onto the key and value vectors, representing respectively... , And introduce learnable time series information queries. By calculating the cross-attention between search features and learnable temporal information queries, temporal information is aggregated to obtain temporal features. :

[0085] ;in, express transpose, Representing time series features Dimensions The function normalizes the original attention weights into a probability distribution, so that these weights are interpreted as the relative importance of different positions.

[0086] In some embodiments, obtaining the bounding box of the candidate predicted target object based on temporal features and search features processed by information interaction specifically includes the following process:

[0087] The search features F after information interaction processing and the learnable target query and time series characteristics The input is fed into a target decoder consisting of a cross-attention and a feedforward network, and the final target features are obtained by calculating the cross-attention. , means as follows:

[0088] ; ; ;

[0089] ;

[0090] ;in, Representation of features The dimension; Indicates a feedforward network. Presentation layer normalization processing.

[0091] In some embodiments, calculating the average coverage of the bounding boxes of the candidate predicted target objects specifically includes the following process:

[0092] ;

[0093] This represents the bounding box of the candidate predicted target object in the i-th frame of the j-th motion video. This represents the ground truth bounding box of the i-th frame in the j-th motion video. Indicates the number of motion videos. The number of frames in each motion video, This is an indicator function; when the overlap rate between the candidate predicted object bounding box and the ground truth object bounding box is greater than a threshold... The value is 1 if it is true, and 0 otherwise.

[0094] Furthermore, determining whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements based on the average coverage of the bounding boxes specifically includes the following process:

[0095] Load the average coverage threshold and determine whether the average coverage of the bounding boxes of the candidate predicted target objects exceeds the average coverage threshold. If yes, the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements. If no, the bounding boxes of the candidate predicted target objects do not meet the selection accuracy requirements.

[0096] In some embodiments, Figure 2 This is a flowchart illustrating the second method for tracking object motion trajectories based on visual features according to an embodiment of the present invention. Figure 2 As shown, the specific steps for obtaining the point set to be processed by inverse density sampling of 3D point cloud data include:

[0097] Step S201: Set the hyperparameter K of the neighborhood points. For the input 3D point cloud P, calculate the distance between each point in the point cloud, and take K neighborhood points as local regions for each point in the point cloud based on the distance.

[0098] Step S202: Calculate the average distance between each point in the three-dimensional point cloud P and its neighboring points. This distance is negatively correlated with the local point distribution density. That is, the higher the local point density, the smaller the calculated average distance, and vice versa.

[0099] Step S203: Calculate the sampling probability based on the average distance point by point, and sample NS points according to the calculated probability and add them to the point set PS. Record the point set PS as the point set to be processed.

[0100] In some embodiments, generating an object motion trajectory vector based on a set of points to be processed specifically includes the following process:

[0101] An ultrasonic probe coordinate system is introduced. The coordinates of the points to be processed are determined by taking the top midpoint of the image frame containing the target object as the origin of the ultrasonic probe coordinate system. The unique coordinates of each point in the ultrasonic probe coordinate system are then obtained. The three-dimensional coordinates of the point in the coordinate system are obtained through coordinate system transformation. The coordinate transformation formula is as follows:

[0102] ;

[0103] in, The transformation matrix depends on the relative positions between the points in the set of points to be processed and can be preset freely. It generates the object's motion trajectory vector based on the three-dimensional coordinate values.

[0104] In some embodiments, Figure 3 This is a system block diagram of an object motion trajectory tracking system based on visual features according to an embodiment of the present invention, such as... Figure 3 As shown, the system includes:

[0105] The feature interaction module is used to acquire continuous target object search image information and target object template image information based on the object motion video stream, extract visual features from the target object search image information and target object template image information to obtain template features and search features, and perform information interaction processing on the template features and search features.

[0106] The bounding box determination module is used to extract temporal features based on search features and cross-attention mechanism, obtain candidate predicted target object bounding boxes based on temporal features and search features after information interaction processing, and determine whether candidate predicted target object bounding boxes meet the selection accuracy requirements based on the average coverage of candidate predicted target object bounding boxes. If so, the candidate predicted target object bounding boxes are marked as predicted target object bounding boxes.

[0107] The motion trajectory tracking module is used to extract the 3D point cloud data of the target object in the bounding box of the predicted target object, perform inverse density sampling on the 3D point cloud data to obtain the point set to be processed, and generate the motion trajectory vector of the object based on the point set to be processed.

[0108] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0109] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0110] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0111] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0112] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for tracking object motion trajectories based on visual features, characterized in that, The methods include: Based on the acquisition of continuous target object search image information and target object template image information from the object motion video stream, visual features are extracted from the target object search image information and target object template image information to obtain template features and search features, and information interaction processing is performed on the template features and search features. The process of extracting visual features from the target object search image information and the target object template image information to obtain template features and search features includes the following steps: Image information for target object search includes images containing both background and target objects. The target object template image information includes the image of the target object. ,in, Represents the set of real numbers. , Images ,image height, , Images ,image The width is 3, where 3 represents the three color channels of the image (RGB). Image and images Each image is split into multiple non-overlapping image patch sequences, denoted as follows: , ,in, , These are the dimensions of the image patches, The size of each image patch is represented by [insert image patch size here]. All image patches are flattened, and the image embedding representation is obtained through a linear mapping function, denoted as [insert image embedding representation here]. , ,in, , The length of the embedded representation, , To define the dimension of the embedding representation, the embedding representations of the search image and template image are input into the Transformer encoding layer for feature extraction. The extracted search features and template features are... , ; The information interaction processing of template features and search features specifically includes the following steps: The template features and search features are unified in dimensionality through a 1×1 convolution operation: ; ; in, , These are the feature vectors after unifying the dimensions of the search features and template features, respectively. based on and Generate graph structure ,in, To search for a set of nodes in the image information representing the feature information of the target object, the node set... , For the number of nodes, Search for the edge set representing the feature information of the target object in the image information. The number of sides is ; compute nodes The size of the d-th eigenvalue : ; in, For nodes Features Represents the measure of node centrality; based on The interaction probability of the attribute representation feature of the compute node : ; in, , for The maximum value, for The average value, Hyperparameters for controlling the overall size of feature enhancement; By importance Delete edges from a graph structure, where, This yields the search features after information interaction processing; Temporal features are extracted based on search features and cross-attention mechanism. Based on the temporal features and the search features after information interaction processing, the bounding boxes of the candidate predicted target objects are obtained. Based on the average coverage of the bounding boxes of the candidate predicted target objects, it is determined whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements. If so, the bounding boxes of the candidate predicted target objects are marked as the predicted target object bounding boxes. The calculation of the average coverage of the bounding boxes of the candidate predicted target objects specifically includes the following process: ; This represents the bounding box of the candidate predicted target object in the i-th frame of the j-th motion video. This represents the ground truth bounding box of the i-th frame in the j-th motion video. Indicates the number of motion videos. The number of frames in each motion video, This is an indicator function; when the overlap rate between the candidate predicted object bounding box and the ground truth object bounding box is greater than a threshold... The value is 1 if it is true, and 0 otherwise. Determining whether the bounding boxes of the candidate predicted target objects meet the selection accuracy requirements based on the average coverage of the bounding boxes of the candidate predicted target objects includes the following process: Load the average coverage threshold and determine whether the average coverage of the bounding box of the candidate predicted target object exceeds the average coverage threshold. If it does, the bounding box of the candidate predicted target object meets the selection accuracy requirement. If not, the bounding box of the candidate predicted target object does not meet the selection accuracy requirement. Extract the 3D point cloud data of the target object within the bounding box of the predicted target object, perform inverse density sampling on the 3D point cloud data to obtain the point set to be processed, and generate the object motion trajectory vector based on the point set to be processed.

2. The object motion trajectory tracking method based on visual features according to claim 1, characterized in that, Extracting temporal features based on search features and cross-attention mechanism The process includes the following: In the Transformer decoder, the search features are linearly projected onto the key and value vectors, representing respectively... , And introduce learnable time series information queries. By calculating the cross-attention between search features and learnable temporal information queries, temporal information is aggregated to obtain temporal features. : ;in, express transpose, Representing time series features Dimensions The function normalizes the original attention weights into a probability distribution, so that these weights are interpreted as the relative importance of different positions.

3. The object motion trajectory tracking method based on visual features according to claim 1, characterized in that, The bounding boxes of candidate predicted target objects are obtained based on temporal features and search features processed by information interaction. The process includes the following: The search features F after information interaction processing and the learnable target query and time series characteristics The input is fed into a target decoder consisting of a cross-attention and a feedforward network, and the final target features are obtained by calculating the cross-attention. , means as follows: ; ; ; ; ;in, Representation of features The dimension; Indicates a feedforward network. Presentation layer normalization processing.

4. The object motion trajectory tracking method based on visual features according to claim 1, characterized in that, The process of performing inverse density sampling on 3D point cloud data to obtain the point set to be processed includes the following steps: Set the neighborhood hyperparameter K. For the input 3D point cloud P, calculate the distance between each point in the point cloud, and select K neighborhood points as local regions for each point in the point cloud based on the distance. Calculate the average distance between each point in the 3D point cloud P and its neighboring points. This distance is negatively correlated with the local point distribution density. That is, the higher the local point density, the smaller the calculated average distance, and vice versa. Based on the average distance point by point, the sampling probability is calculated, and NS points are sampled according to the calculated probability and added to the point set PS. The point set PS is denoted as the point set to be processed.

5. The object motion trajectory tracking method based on visual features according to claim 1, characterized in that, Specifically, generating object motion trajectory vectors based on the set of points to be processed. The process includes the following: An ultrasonic probe coordinate system is introduced. The coordinates of the points to be processed are determined by taking the top midpoint of the image frame containing the target object as the origin of the ultrasonic probe coordinate system. The unique coordinates of each point in the ultrasonic probe coordinate system are then obtained. The three-dimensional coordinates of the point in the coordinate system are obtained through coordinate system transformation. The coordinate transformation formula is as follows: ; in, The transformation matrix depends on the relative positions between the points in the set of points to be processed and can be preset freely. It generates the object's motion trajectory vector based on the three-dimensional coordinate values.

Citation Information

Patent Citations

  • Twin network tracking system and method based on space-time attention mechanism

    CN114707604A

  • Unmanned aerial vehicle target infrared tracking method and system for complex scene

    CN119810141A