Point tracking method, device, electronic device and storage medium

By extracting space-time dependency features and using adaptive sliding window strategies for iterative updates, the problem of insufficient point tracking accuracy and robustness in the existing technology is solved, and a more efficient and robust point tracking effect is achieved.

CN117745761BActive Publication Date: 2025-06-06SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311781116.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-06
Estimated Expiration
2043-12-22

AI Technical Summary

Technical Problem

The existing point tracking methods have shortcomings in terms of tracking accuracy and robustness in occlusion scenarios, resulting in low accuracy and low robustness of point tracking.

Method used

By determining the initial feature trajectory and position trajectory based on the current position information of the point to be tracked, the space-time dependency characteristics are extracted, and iterative updates are performed using the preset adaptive sliding window strategy to generate the target feature trajectory and the target position trajectory until the final position trajectory of the point to be tracked is determined.

Benefits of technology

It effectively improves the accuracy and efficiency of point tracking in the case of target occlusion, improves the robustness and performance of point tracking, and can better deal with long-term occlusion scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117745761B_ABST
    Figure CN117745761B_ABST
Patent Text Reader

Abstract

The present invention discloses a point tracking method, device, electronic device and storage medium. The method includes: determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked; determining the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; iteratively updating the initial feature trajectory and initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features and a preset number of iterations to generate a target feature trajectory and target position trajectory corresponding to the current video frame segment; determining the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence in turn, and then determining the final position trajectory of the point to be tracked corresponding to the video frame sequence. By extracting the spatiotemporal dependency features of the point to be tracked and adopting the preset adaptive sliding window strategy, the point tracking accuracy and robustness under target occlusion are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a point tracking method, device, electronic device and storage medium. Background Art

[0002] Point tracking is a key task in computer vision, which is used to track the position of objects, specific points or features over time in a video sequence. This technology is of great significance in multiple applications, including behavior recognition, future behavior prediction, intention inference, and robot navigation. In addition, it also solves the challenge of long-term motion estimation in real-world scenes, which also has great potential for video surveillance, behavior analysis, and video editing.

[0003] There are still some shortcomings in the existing point tracking methods: for example, the point tracking method based on optical flow, the point tracking method based on feature matching, etc., the above point tracking methods usually track the movement of points independently, and do not take into account the relationship between multiple points and multiple frames, resulting in low accuracy of point trajectory tracking; in addition, the above point tracking methods have low robustness when dealing with complex scenes such as occlusion. Summary of the invention

[0004] The present invention provides a point tracking method, device, electronic device and storage medium to solve the problem of low tracking accuracy in existing point tracking methods, especially low tracking robustness in occluded scenes.

[0005] According to one aspect of the present invention, a point tracking method is provided, the method comprising:

[0006] Determine an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked;

[0007] Determine the temporal and spatial dependency features corresponding to the point to be tracked in the current video frame segment;

[0008] Iteratively updating the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features, and a preset number of iterations to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust the preset window size corresponding to the current video frame segment according to the target feature trajectory;

[0009] Based on the target position trajectory, the target position information of the last frame image of the point to be tracked in the current video frame segment is determined, and the target position information is used as the current position information of the point to be tracked, and the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked is returned until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined.

[0010] According to another aspect of the present invention, there is provided a point tracking device, the device comprising:

[0011] A trajectory initialization module, used to determine an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked;

[0012] A feature determination module is used to determine the time-space dependent features corresponding to the point to be tracked in the current video frame segment;

[0013] A target trajectory determination module is used to iteratively update the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features, and a preset number of iterations to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust the preset window size corresponding to the current video frame segment according to the target feature trajectory;

[0014] The final trajectory determination module is used to determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the point tracking method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the point tracking method described in any embodiment of the present invention when executed.

[0020] The technical solution of the embodiment of the present invention is to determine the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked; determine the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; iteratively update the initial feature trajectory and initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features and a preset number of iterations to generate a target feature trajectory and target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust the preset window size corresponding to the current video frame segment according to the target feature trajectory; determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined. The embodiment of the present invention realizes information exchange between points to be tracked and captures complex dependencies across points by extracting the corresponding spatiotemporal dependency features of the points to be tracked in the current video frame segment, thereby effectively improving the point tracking accuracy and efficiency under target occlusion. At the same time, in the iterative update process of the trajectory, a preset adaptive sliding window strategy is adopted, which can dynamically adjust the sliding window size of the video frame segment, so that the technical solution can effectively cope with long-term occlusion scenarios, further improving the robustness and performance of point trajectory tracking.

[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a flow chart of a point tracking method provided according to Embodiment 1 of the present invention;

[0024] Figure 2is a flow chart of a point tracking method provided according to Embodiment 2 of the present invention;

[0025] Figure 3 is a flow chart of a point tracking method provided according to Embodiment 3 of the present invention;

[0026] Figure 4 is a flowchart of a point tracking method provided according to Embodiment 3 of the present invention;

[0027] Figure 5 is a schematic diagram of a local similarity measurement module provided according to Embodiment 3 of the present invention;

[0028] Figure 6 is a schematic diagram of a multi-point information interaction module provided according to Embodiment 3 of the present invention;

[0029] Figure 7 is a schematic diagram of a visual comparison result provided according to Embodiment 3 of the present invention;

[0030] Figure 8 is a schematic diagram of the structure of a point tracking device provided according to a fourth embodiment of the present invention;

[0031] Fig. 9 It is a schematic diagram of the structure of an electronic device for implementing the point tracking method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0034] Embodiment 1

[0035] Figure 1 A flowchart of a point tracking method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of performing point trajectory tracking on a point to be tracked in a video frame sequence. The method can be performed by a point tracking device. The point tracking device can be implemented in the form of hardware and / or software. The point tracking device can be configured in an electronic device with computing capabilities. Figure 1 As shown, a point tracking method provided in the first embodiment specifically includes the following steps:

[0036] S110 . Determine an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked.

[0037] The point to be tracked may refer to a target input point whose position trajectory needs to be tracked in a video frame sequence. The number of points to be tracked may be one or more. The specific number may be set according to actual needs. Preferably, one main tracking point and multiple auxiliary tracking points may be selected. The current video frame segment may refer to a video frame segment to be processed obtained by sliding a preset window size in a video frame sequence. The current position information may refer to the position coordinate information of the point to be tracked in the first frame image of the current video frame segment.

[0038] The initial feature trajectory may refer to a feature vector trajectory (sequence) obtained by initializing the feature map (feature map) corresponding to the previous video frame segment using the current position information of the point to be tracked.

[0039] The initial position trajectory may refer to a position trajectory (sequence) obtained by initializing the current video frame segment using the current position information of the point to be tracked. The position trajectory is a sequence composed of the position coordinates of the point to be tracked in each frame image in the video frame sequence.

[0040] In an embodiment of the present invention, one or more points to be tracked can be selected in the current video frame segment according to actual needs, and then the current position information of the points to be tracked can be used to initialize and obtain the initial feature trajectory and initial position trajectory of the points to be tracked corresponding to the current video frame segment. Specifically, the position coordinates of the points to be tracked in the first frame image of the current video frame segment can be determined according to the current position information of the points to be tracked, and then the above position coordinates can be tiled (copied) to each frame image of the current video frame segment to obtain the initial position trajectory of the points to be tracked corresponding to the current video frame segment; for the feature trajectory initialization process of the points to be tracked, an encoder (Encoder) such as a convolutional neural network (CNN) can be called first to extract feature maps corresponding to each frame image in the current video frame segment, and then a sampling algorithm such as bilinear sampling can be used to extract the corresponding feature vector in each feature map according to the position coordinates of the points to be tracked, and the initial feature trajectory of the points to be tracked corresponding to the current video frame segment can be obtained by tiling the feature vectors corresponding to the position coordinates to the feature map corresponding to each frame image.

[0041] S120: Determine the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment.

[0042] Among them, the spatiotemporal dependency features can be understood as features used to characterize the correlation between the points to be tracked in time and space. The spatiotemporal dependency features can be used to fully explore the dependencies between points to improve the point tracking performance under occlusion.

[0043] In an embodiment of the present invention, a pre-trained neural network model may be called to extract the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment, wherein the above-mentioned neural network model may at least include a network model based on a self-attention mechanism. In a specific embodiment, a Transformer decoder may be called to extract the spatiotemporal dependency features of the point to be tracked in the current video frame segment through a temporal self-attention mechanism and a spatial self-attention mechanism, and the temporal self-attention mechanism is used to capture the mutual dependence of the points to be tracked in the time dimension, and the spatial self-attention mechanism is used to capture the dependency relationship between each point to be tracked in the same frame image, so the combination of the temporal self-attention and spatial self-attention operations, that is, the spatiotemporal dependency features can capture the dependency relationship between cross-points.

[0044] It should be understood that in existing point tracking methods, each point to be tracked is usually tracked in isolation, and the mutual information between different point trajectories cannot be used. However, the embodiments of the present invention can capture and utilize the spatial and temporal correlations between multiple points by extracting the corresponding spatiotemporal dependency features of the point to be tracked in the current video frame segment, so as to fully explore the dependencies across points. This information exchange establishes mutual information exchange between point trajectories, which can effectively improve point tracking performance in the case of occlusion.

[0045] S130, iteratively updating the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features, and a preset number of iterations, to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment.

[0046] Among them, the preset adaptive sliding window strategy can be understood as a pre-configured strategy for adaptively and dynamically adjusting the preset window size corresponding to the current video frame segment. For example, after the initial feature trajectory and the initial position trajectory are iteratively updated for a preset number of times, the preset window size corresponding to the current video frame segment can be adaptively adjusted by predicting the occlusion of the point to be tracked (for example, the main tracking point) in the last frame image of the current video frame segment. If it is predicted that the point to be tracked is occluded in the last frame image, the preset window size corresponding to the current video frame segment is increased, and the trajectory is iteratively updated again; if it is predicted that the tracking point is not occluded, the initial feature trajectory and initial position trajectory after the current iterative update are used as the corresponding target feature trajectory and target position trajectory.

[0047] The target feature trajectory and the target position trajectory may respectively refer to the feature trajectory and the position trajectory after iteratively updating the initial feature trajectory and the initial position trajectory of the current video frame segment.

[0048] In an embodiment of the present invention, a group of continuous video frames can be processed at a time in a sliding window manner. Taking the current video frame segment as an example, the corresponding window size is a preset window size. In order to improve the performance and adaptability of trajectory updating, the preset window size corresponding to the current video frame segment can be adaptively adjusted using a preset adaptive sliding window strategy. Specifically, the initial feature trajectory and the initial position trajectory can be iteratively updated for a preset number of iterations based on the spatiotemporal dependency characteristics of the point to be tracked, and a target feature trajectory and a target position trajectory corresponding to the current video frame segment are generated. Then, the occlusion of the point to be tracked in the last frame image of the current video frame segment is predicted based on the target feature trajectory. If it is predicted that the point to be tracked is occluded, the preset window size corresponding to the current video frame segment is increased, and a new current video frame segment is re-determined based on the new preset window size, and the new target feature trajectory and target position trajectory are re-iterated and updated to determine the new target feature trajectory and target position trajectory; if it is predicted that the point to be tracked is not occluded, it is considered that the target feature trajectory and the target position trajectory corresponding to the current video frame segment are determined, and the target feature trajectory and the target position trajectory corresponding to the next video frame segment of the point to be tracked can be continued to be determined.

[0049] It should be understood that, compared with the existing point tracking methods that directly process the entire video frame sequence, or use a model / algorithm with a fixed sliding window size, the embodiments of the present invention adopt a preset adaptive sliding window strategy, which can dynamically adjust the window size according to the occlusion prediction situation. If it is predicted that the point to be tracked is occluded in the last row of images of the video frame segment, the window size is increased to obtain more context information, thereby improving the point tracking performance; conversely, if it is predicted that the point to be tracked is not occluded, a smaller window size can be used for processing, thereby improving the point tracking efficiency.

[0050] S140, determining the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and using the target position information as the current position information of the point to be tracked, and returning to the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined.

[0051] The target position information may refer to the position coordinate information of the point to be tracked in the last frame of the current video frame segment. The final position trajectory may refer to the position trajectory of the entire video frame sequence corresponding to the point to be tracked, and the final position trajectory may be obtained by splicing the target position trajectories corresponding to multiple video frame segments according to the time dimension.

[0052] In an embodiment of the present invention, after determining the target feature trajectory and the target position trajectory corresponding to the point to be tracked in the current video frame segment, the target position information of the last frame image of the point to be tracked in the current video frame segment can be determined based on the above position trajectory, and the target position information is used as the current position information of the next video frame segment in the video frame sequence, and referring to the above process of determining the target feature trajectory and the target position trajectory corresponding to the point to be tracked in the current video frame segment, the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence is determined in turn, and each target position trajectory is spliced ​​into a final position trajectory according to the time dimension.

[0053] The technical solution of the embodiment of the present invention is to determine the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked; determine the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; iteratively update the initial feature trajectory and initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features and a preset number of iterations to generate a target feature trajectory and target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust the preset window size corresponding to the current video frame segment according to the target feature trajectory; determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined. The embodiment of the present invention realizes information exchange between points to be tracked and captures complex dependencies across points by extracting the corresponding spatiotemporal dependency features of the points to be tracked in the current video frame segment, thereby effectively improving the point tracking accuracy and efficiency under target occlusion. At the same time, in the iterative update process of the trajectory, a preset adaptive sliding window strategy is adopted, which can dynamically adjust the sliding window size of the video frame segment, so that the technical solution can effectively cope with long-term occlusion scenarios, further improving the robustness and performance of point trajectory tracking.

[0054] Embodiment 2

[0055] Figure 2 This is a flowchart of a point tracking method provided in the second embodiment of the present invention, which is further optimized and expanded based on the above implementation, and can be combined with various optional technical solutions in the above implementation. Figure 2 As shown, a point tracking method provided in the second embodiment specifically includes the following steps:

[0056] S210, calling the convolutional neural network of the pre-trained point tracking model to extract the initial feature map corresponding to each frame image in the current video frame segment.

[0057] The point tracking model may refer to a pre-built model for realizing point trajectory tracking, and the point tracking model may include a trajectory initialization module, a local similarity measurement module and a multi-point information interaction module.

[0058] In an embodiment of the present invention, a pre-trained CNN network can be called to extract feature maps, i.e., initial feature maps, corresponding to each frame of the image from the current video frame segment. The specific model parameters of the CNN network, such as the number of layers, are not specifically limited in the embodiment of the present invention, and image feature extraction can be achieved.

[0059] S220 , determining the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment according to the current position information, and tiling the initial position coordinates to each frame image to obtain an initial position trajectory.

[0060] In an embodiment of the present invention, the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment can be determined based on the current position information of the point to be tracked, and then the above initial position coordinates are tiled (copied) to each frame image of the current video frame segment, thereby obtaining the initial position trajectory of the point to be tracked corresponding to the current video frame segment.

[0061] S230. In the initial feature map corresponding to the first frame image, a preset bilinear sampling algorithm is called to determine a feature vector corresponding to the initial position coordinate, and the feature vector is flattened to the initial feature map corresponding to each frame image to obtain an initial feature trajectory.

[0062] In an embodiment of the present invention, a preset bilinear sampling algorithm can be called to determine the feature vector corresponding to the initial position coordinates in the initial feature map corresponding to the first frame image of the current video frame fragment, and the feature trajectory can be initialized based on this sampling feature, that is, the above-mentioned feature vector is flattened in the time dimension to the initial feature map corresponding to each frame image, so as to obtain the initial feature trajectory of the current video frame fragment corresponding to the point to be tracked.

[0063] S240, determining the inner product result between the feature vector of the point to be tracked in the initial feature trajectory and the initial feature map of the corresponding time step, and extracting multi-scale similarity features from the inner product result to construct a correlation matrix corresponding to the current video frame segment.

[0064] Among them, the multi-scale similarity feature can be understood as the multi-scale similarity feature extracted by using the spatial pyramid algorithm based on the position coordinates of the point to be tracked in the inner product result.

[0065] In an embodiment of the present invention, during the iterative update process of the initial feature trajectory and the initial position trajectory corresponding to the current video frame segment, the inner product result between the feature vector of the point to be tracked in the initial feature trajectory and the initial feature map of the corresponding time step (i.e., the corresponding frame) can be determined, and the spatial pyramid algorithm is called to extract the multi-scale similarity feature from the above inner product result, and then the multi-scale similarity feature corresponding to the point to be tracked is spliced ​​and tiled, so as to obtain the correlation matrix of the point to be tracked corresponding to the current video frame segment. In addition, the above inner product result can also be called a similarity map, which is used to characterize the degree of matching between the position trajectory and its associated features and the initial feature map. A large positive value in the similarity map indicates that the feature of the point to be tracked is highly similar to the convolution feature of the corresponding position.

[0066] S250, concatenating the correlation matrix, the initial feature trajectory, and the initial position trajectory into fused features, determining the temporal self-attention and spatial self-attention corresponding to the fused features, and taking the sum of the temporal self-attention and the spatial self-attention as the spatiotemporal dependency feature.

[0067] Among them, the temporal self-attention of the fused features can be used to characterize the mutual dependence of the points to be tracked in the temporal dimension. The spatial self-attention of the fused features can be used to characterize the dependency between the points to be tracked in the same frame image.

[0068] In an embodiment of the present invention, in order to capture and utilize the temporal and spatial correlations between multiple points to be tracked, the Transformer decoder can be called to obtain the temporal self-attention and spatial self-attention corresponding to the fusion features respectively, and then the sum of the temporal self-attention and the spatial self-attention is used as the spatiotemporal dependency feature. The spatiotemporal dependency feature is used to capture the dependency relationship between points to improve the point tracking performance under occlusion.

[0069] S260, using the spatiotemporal dependent features as the input of the feedforward neural network layer, generating trajectory update amounts corresponding to the initial feature trajectory and the initial position trajectory at the current iteration number, and then obtaining the updated initial feature trajectory and initial position trajectory at the current iteration number based on the trajectory update amounts.

[0070] In the embodiment of the present invention, the captured spatiotemporal dependency features can be input into the feedforward neural network (FNN) layer of the point tracking model for processing to generate the initial feature trajectory F under the current iteration number K. K and the initial position trajectory X K The corresponding trajectory updates ΔF and ΔX are then applied to the current state by addition, that is, F K =F K-1 +ΔF,X K =X K-1+ΔX, where F K-1 and X K-1 They respectively represent the initial feature trajectory and initial position trajectory obtained after the last iterative update.

[0071] S270. After performing iterative updates for a preset number of iterations, a target feature trajectory and a target position trajectory are obtained.

[0072] In an embodiment of the present invention, referring to steps S240 to S260, after iteratively updating the initial feature trajectory and initial position trajectory corresponding to the current video frame segment for a preset number of iterations, the target feature trajectory and target position trajectory of the point to be tracked corresponding to the current video frame segment are obtained.

[0073] S280, input the target feature trajectory into the linear layer and the Sigmoid activation layer in sequence, and determine the position occlusion state of the last frame image corresponding to the point to be tracked under the current preset window size.

[0074] The position occlusion state may refer to the occlusion prediction of the point to be tracked in the last frame image of the current video frame segment. The position occlusion state may be represented in the form of visibility score, occlusion prediction probability, etc.

[0075] In an embodiment of the present invention, after obtaining the target feature trajectory, the target feature trajectory can be sent to the linear layer and the Sigmoid activation layer of the model to be tracked to obtain the position occlusion state of the last frame image corresponding to the point to be tracked under the current preset window size.

[0076] S290: When the position occlusion state satisfies the preset window update condition, a new preset window size is determined according to a sliding window update formula corresponding to the preset adaptive sliding window strategy, and a new current video frame segment is re-determined according to the new preset window size.

[0077] Among them, the preset window update condition may refer to a pre-configured condition for determining whether it is necessary to adaptively and dynamically adjust the preset window size corresponding to the current video frame segment. The preset window update condition may include: the visibility score is less than the preset score threshold, the occlusion prediction probability is greater than the preset prediction probability threshold, etc.

[0078] In an embodiment of the present invention, after determining the position occlusion state of the point to be tracked corresponding to the last frame of the image, it can be determined whether it satisfies the pre-configured preset window update condition. When the condition is met, the sliding window update formula is called to determine a new preset window size, and the window size is used to slide again in the video frame sequence to obtain a new current video frame segment.

[0079] Further, based on the above-mentioned embodiment of the invention, the sliding window update formula can be expressed as: W' = W + 2n ; Wherein, W represents the preset window size; W' represents the new preset window size, that is, the re-determined window size; n represents the number of times the position occlusion state meets the preset window update condition during the feature trajectory update process for the same video frame segment.

[0080] S2100: Re-determine a new target feature trajectory and a new target position trajectory corresponding to the new current video frame segment, until the new target feature trajectory corresponds to a new position occlusion state that satisfies a preset window update condition.

[0081] In an embodiment of the present invention, the process of determining the target feature trajectory and the target position trajectory corresponding to the point to be tracked in the current video frame segment can be referred to, and the new target feature trajectory and the new target position trajectory corresponding to the new current video frame segment can be re-determined, and the new position occlusion state corresponding to the new target feature trajectory can be re-determined. If the new position occlusion state does not meet the preset window update condition, the current feature trajectory and position trajectory are used as the target feature trajectory and target position trajectory corresponding to the current video frame segment; if the new position occlusion state still meets the above-mentioned preset window update condition, the preset window size corresponding to the current video frame segment continues to be adjusted until the new position occlusion state corresponding to the new target feature trajectory does not meet the preset window update condition.

[0082] S2110: Slide in the video frame sequence using a preset window size to obtain a next set of video frame segments, and determine the initial position coordinates of the point to be tracked in the first frame image of the next set of video frame segments according to the target position information.

[0083] In an embodiment of the present invention, a preset window size can be used to continue sliding in the video frame sequence to obtain the next group of video frame fragments, and the position coordinates of the point to be tracked corresponding to the last frame image in the target position trajectory of the current video frame fragment (i.e., the previous group of video frame fragments) are used as the initial position coordinates in the first frame image of the next group of video frame fragments.

[0084] S2120, returning to the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment, and obtaining the target position trajectory of the point to be tracked corresponding to the next set of video frame segments.

[0085] In an embodiment of the present invention, the target feature trajectory and target position trajectory corresponding to the point to be tracked in the current video frame segment can be determined by referring to the above process, and the target position trajectory of the point to be tracked corresponding to the next set of video frame segments can be continued to be determined. The specific process will not be repeated here.

[0086] S2130 , sequentially determining the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence, and splicing the target position trajectories into a final position trajectory according to the time dimension.

[0087] In an embodiment of the present invention, the target position trajectory of the point to be tracked corresponding to each video frame segment can be determined in sequence in the video frame sequence, and the target position trajectories are spliced ​​into a final position trajectory according to the time dimension, thereby realizing trajectory tracking of the point to be tracked.

[0088] Furthermore, based on the above-mentioned embodiments of the invention, during the training process of the point tracking model, a time contrast loss mechanism may be used to optimize the feature trajectory, wherein the time contrast loss mechanism is implemented based on positive sample pairs and negative sample pairs.

[0089] Specifically, in order to further improve the performance of the point tracking model, especially to maintain the consistency of target features when processing time series data, the embodiment of the present invention adopts a time contrast loss mechanism. The time contrast loss strengthens the model's learning of the continuity of the target in the time dimension by minimizing the distance between the feature representations of the same target at different time steps and maximizing the distance between different target or background features.

[0090] The technical solution of the embodiment of the present invention extracts the initial feature map corresponding to each frame image in the current video frame segment by calling the convolutional neural network of the pre-trained point tracking model; determines the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment according to the current position information, and flattens the initial position coordinates to each frame image to obtain an initial position trajectory; in the initial feature map corresponding to the first frame image, calls a preset bilinear sampling algorithm to determine the feature vector corresponding to the initial position coordinates, and flattens the feature vector to the initial feature map corresponding to each frame image to obtain an initial feature trajectory; determines the feature vector of the point to be tracked in the initial feature trajectory The inner product result between the initial feature map of the corresponding time step and the multi-scale similarity feature is extracted from the inner product result to construct the correlation matrix corresponding to the current video frame segment; the correlation matrix, the initial feature trajectory and the initial position trajectory are spliced ​​into a fusion feature, the temporal self-attention and spatial self-attention corresponding to the fusion feature are determined, and the sum of the temporal self-attention and the spatial self-attention is used as the spatiotemporal dependency feature; the spatiotemporal dependency feature is used as the input of the feedforward neural network layer to generate the trajectory update amount corresponding to the initial feature trajectory and the initial position trajectory at the current iteration number, and then the updated initial feature trajectory at the current iteration number is obtained based on the trajectory update amount. The target feature trajectory and the initial position trajectory are obtained after iterative updates for a preset number of iterations. The target feature trajectory is sequentially input into the linear layer and the Sigmoid activation layer to determine the position occlusion state of the last frame image corresponding to the point to be tracked under the current preset window size. When the position occlusion state satisfies the preset window update condition, a new preset window size is determined according to the sliding window update formula corresponding to the preset adaptive sliding window strategy, and a new current video frame segment is re-determined according to the new preset window size. A new target feature trajectory and a new target position trajectory corresponding to the new current video frame segment are re-determined until The new target feature trajectory corresponds to a new position occlusion state that satisfies the preset window update condition; the preset window size is used to slide in the video frame sequence to obtain the next set of video frame fragments, and the initial position coordinates of the point to be tracked in the first frame image of the next set of video frame fragments are determined according to the target position information; return to execute the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame fragment, and obtain the target position trajectory of the point to be tracked corresponding to the next set of video frame fragments; determine the target position trajectory of the point to be tracked corresponding to each video frame fragment in the video frame sequence in turn, and splice each target position trajectory into a final position trajectory according to the time dimension.The embodiment of the present invention extracts the temporal self-attention and spatial self-attention corresponding to the point to be tracked in the current video frame segment and fuses them into spatiotemporal dependent features, which allows information exchange and collaboration between the points to be tracked, and captures and utilizes the spatial and temporal correlations across points, thereby effectively improving the point tracking accuracy and efficiency in the case of target occlusion. At the same time, in the iterative update process of the trajectory, a preset adaptive sliding window strategy is adopted, and the sliding window size of the video frame segment can be dynamically adjusted according to the target occlusion situation, so that the technical solution can effectively cope with long-term occlusion scenarios, further improving the robustness and performance of point trajectory tracking.

[0091] Embodiment 3

[0092] Figure 3 This is a flowchart of a point tracking method provided in Example 3 of the present invention. Based on the above embodiments, this embodiment provides an implementation of a point tracking method, which can realize point tracking in complex occlusion scenes by comprehensively considering the relationship between multiple points and multiple frames and using adaptive sliding windows and Transformer architecture. Figure 3 As shown, a point tracking method provided by Embodiment 3 of the present invention specifically includes the following steps:

[0093] S310: Slide a preset window size in a video frame sequence to obtain a current video frame segment, and determine current position information of a point to be tracked in the current set of video frame segments.

[0094] S320 , calling the trajectory initialization module of the point tracking model to initialize and obtain the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment.

[0095] Figure 4 This is a flowchart of a point tracking method provided in Embodiment 3 of the present invention. Figure 4 As shown, the encoder can be called to extract the feature map corresponding to each frame image from the current video frame fragment, that is, the initial feature map, and then the feature vector of the point to be tracked is calculated by bilinear sampling within the given coordinates (that is, the current position information) in the feature map of the first frame, and the above feature vector is flattened to the initial feature map corresponding to each frame image in the time dimension, so as to obtain the initial feature trajectory of the point to be tracked corresponding to the current video frame fragment, wherein this feature trajectory initialization method assumes the consistency of the appearance of the target (that is, the point to be tracked).

[0096] Similarly, for the initialization operation of the position trajectory, the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment can be determined according to the current position information of the point to be tracked, and then the above initial position coordinates are tiled (copied) to each frame image of the current video frame segment, thereby obtaining the initial position trajectory of the point to be tracked corresponding to the current video frame segment, wherein this position trajectory initialization method is based on the zero speed assumption, that is, in the absence of other motion information, it is assumed that the target remains stationary.

[0097] After performing the trajectory initialization operation, the feature track will be updated to track the appearance changes of the points to be tracked, and the position track will be updated to track the movement of the points to be tracked.

[0098] S330. In the process of iteratively updating the initial feature trajectory and the initial position trajectory corresponding to the current video frame segment, the local similarity measurement module of the point tracking model is called to determine the correlation matrix of the point to be tracked corresponding to the current video frame segment.

[0099] In the embodiment of the present invention, Figure 4 and Figure 5 As shown in , after each iterative update of the initial feature trajectory and the initial position trajectory, the matching degree between the position trajectory and its associated features and the initial feature map can be evaluated, that is, the inner product operation of each feature vector in the updated initial feature trajectory and the initial feature map of the corresponding time step is performed to obtain a similarity map. This step returns a series of similarity score maps, in which a large positive value indicates that the target feature is highly similar to the convolution feature of the position. Finally, a spatial pyramid is used to obtain multi-scale similarity measurements in the similarity map corresponding to each point to be tracked, and a set of multi-scale scoring blocks C is formed. k , the scoring block set C k That is, the correlation matrix of the point to be tracked corresponding to the current video frame segment.

[0100] S340, splicing the correlation matrix, the initial feature trajectory and the initial position trajectory into a fusion feature, and calling the multi-point information interaction module of the model to be tracked to determine the target feature trajectory and the target position trajectory of the point to be tracked corresponding to the current video frame segment.

[0101] In the embodiment of the present invention, Figure 4 and Figure 6 As shown, the Transformer decoder can be called to obtain the fusion features F fuse The corresponding temporal self-attention F traj and spatial self-attention F point , and then take the sum of temporal self-attention and spatial self-attention as the spatiotemporal dependency feature F crossThen, the captured spatiotemporal dependency features are input into the feedforward neural network layer of the point tracking model for processing to generate the initial feature trajectory F under the current iteration number K K and the initial position trajectory X K The corresponding trajectory update amounts ΔF and ΔX are then applied to the current state by addition to obtain the updated initial feature trajectory F under the current iteration number K. K =F K-1 +ΔF and initial position trajectory X K =X K-1 +ΔX. Among them, Figure 6 f t 、f t+1 and f t+2 They respectively represent the feature vectors corresponding to time steps t, t+1 and t+2 in the feature trajectory of a point to be tracked.

[0102] At the same time, the Adaptive Sliding Window (ASW) strategy is seamlessly integrated into the Transformer decoder architecture. This integration takes full advantage of the Transformer's advantages in handling long-term dependencies and occlusions, and provides higher performance and adaptability compared to models with fixed window sizes. By adopting the adaptive sliding window strategy, the window size can be dynamically adjusted according to the occlusion prediction. That is, if the point to be tracked is predicted to be occluded in the last frame of the video frame segment, the window size is increased to obtain more contextual information, thereby improving the point tracking performance; conversely, if the point to be tracked is predicted to be unoccluded, a smaller window size can be used for processing, thereby improving the point tracking efficiency.

[0103] After iterative updating for a preset number of iterations according to the adaptive sliding window strategy, the target feature trajectory and the target position trajectory of the point to be tracked corresponding to the current video frame segment can be obtained.

[0104] S350, sliding in the video frame sequence to obtain the next set of video frame segments, taking the position coordinates of the point to be tracked corresponding to the last frame image in the current video frame segment as the initial position coordinates in the first frame image of the next set of video frame segments, and determining the target position trajectory of the point to be tracked corresponding to the next set of video frame segments.

[0105] S360: sequentially determine the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence, and splice the target position trajectories into a final position trajectory according to the time dimension.

[0106] Furthermore, based on the above-mentioned embodiments of the invention, in order to further improve the performance of the point tracking model, especially to maintain the consistency of target features when processing time series data, the embodiments of the invention adopt a time contrast loss mechanism. The time contrast loss strengthens the model's learning of the continuity of the target in the time dimension by minimizing the distance between the feature representations of the same target at different time steps and maximizing the distance between different target or background features.

[0107] In the multi-point information interaction module of the embodiment of the present invention, the time contrast loss is used to optimize the feature trajectory to ensure that the model can accurately track the target even when the target changes dynamically or is occluded. In this way, the model can not only capture the characteristics of the target at a single time step, but also understand the behavior and change pattern of the target in the time series. Specifically, the time contrast loss calculates the feature trajectory F k In the loss function, the similarity of features between adjacent time steps is calculated and compared with the similarity of features of different targets or backgrounds. Through this comparison, the model is trained to distinguish the consistency of the target in consecutive time steps from the difference between the background or other targets. The introduction of this loss function significantly improves the robustness and accuracy of the model in complex scenes, especially when the target moves quickly or is partially occluded.

[0108] In practice, the temporal contrast loss can be implemented by constructing a positive pair and multiple negative pairs. A positive pair consists of the feature representations of the same object at different time steps, while a negative pair consists of the feature representations of the object and the background or other objects. The loss function encourages the model to bring the feature representations of the positive pair closer together, while pushing the feature representations of the negative pair farther apart, thus forming a sharp boundary in the feature space.

[0109] Through this method, the point tracking model of the embodiment of the present invention can effectively process time series data, improve performance in dynamic environments, and maintain a high degree of tracking accuracy in the presence of occlusion and appearance changes.

[0110] The point tracking method proposed in the embodiment of the present invention has demonstrated excellent point tracking performance through quantitative evaluation on multiple benchmark datasets such as FlyingThings++, CroHD, BADJA, and TAP-Vid-DAVIS. Compared with existing point tracking methods, the point tracking method proposed in this embodiment significantly reduces the error rate of occluded points, improves tracking accuracy, and surpasses competitors in multiple indicators. This method also demonstrates excellent robustness in complex actual scenes. By exchanging information between point trajectories, it can still maintain accurate tracking when encountering challenges such as occlusion and large-scale motion changes, providing an effective solution to the point tracking problem in the field of computer vision.

[0111] like Figure 7 As shown, the point tracking method proposed in the embodiment of the present invention can still maintain accurate tracking in scenes with severe occlusion and large motion changes. By exchanging information between point trajectories, the method exhibits consistent tracking performance. When some points are occluded, the shared information is used to "fill in the gaps", effectively improving the point tracking accuracy and performance in the case of target occlusion.

[0112] The technical solution of the embodiment of the present invention is to obtain the current video frame segment by sliding a preset window size in a video frame sequence, and determine the current position information of the point to be tracked in the current group of video frame segments; call the trajectory initialization module of the point tracking model to initialize and obtain the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment; in the process of iteratively updating the initial feature trajectory and initial position trajectory corresponding to the current video frame segment, call the local similarity measurement module of the point tracking model to determine the correlation matrix of the point to be tracked corresponding to the current video frame segment; splice the correlation matrix, the initial feature trajectory and the initial position trajectory into a fusion feature, and call the multi-point information interaction module of the model to be tracked to determine the target feature trajectory and target position trajectory of the point to be tracked corresponding to the current video frame segment; slide in the video frame sequence to obtain the next group of video frame segments, use the position coordinates of the point to be tracked in the current video frame segment corresponding to the last frame image as the initial position coordinates in the first frame image of the next group of video frame segments, and determine the target position trajectory of the point to be tracked corresponding to the next group of video frame segments; determine the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence in turn, and splice each target position trajectory into a final position trajectory according to the time dimension. The embodiments of the present invention have the following beneficial effects: ① The introduction of an adaptive sliding window strategy can effectively cope with the challenge of long-term occlusion. Compared with the traditional fixed window method, the sliding window size of the video frame segment can be dynamically adjusted according to the target occlusion situation, so that the model can be more flexible in handling occlusion, thereby improving the robustness of point trajectory tracking and having stronger adaptability; ② By introducing a multi-point information interaction module, information exchange and collaboration between multiple points can be allowed, thereby capturing and utilizing the spatial and temporal correlations between points. This innovation improves the accuracy of point trajectory tracking, especially in the presence of occlusion; ③ The Transformer decoder architecture is adopted, which can better handle long-term dependencies and occlusion situations, and provides more powerful point trajectory tracking capabilities. By comprehensively utilizing temporal self-attention and spatial self-attention operations, the tracking performance between points is improved, so that the model can better capture the motion trajectory of the points.

[0113] Embodiment 4

[0114] Figure 8 This is a schematic diagram of the structure of a point tracking device provided in Embodiment 4 of the present invention. Figure 8 As shown, the device comprises:

[0115] A trajectory initialization module 41 is used to determine an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked;

[0116] A feature determination module 42, used to determine the time-space dependent features corresponding to the point to be tracked in the current video frame segment;

[0117] The target trajectory determination module 43 is used to iteratively update the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features, and a preset number of iterations to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust the preset window size corresponding to the current video frame segment according to the target feature trajectory;

[0118] The final trajectory determination module 44 is used to determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined.

[0119] The technical solution of the embodiment of the present invention is as follows: a trajectory initialization module determines an initial feature trajectory and an initial position trajectory corresponding to a point to be tracked in a current video frame segment based on current position information of the point to be tracked; a feature determination module determines the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; a target trajectory determination module iteratively updates the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, spatiotemporal dependency features, and a preset number of iterations, to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust a preset window size corresponding to the current video frame segment according to the target feature trajectory; a final trajectory determination module determines the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and uses the target position information as the current position information of the point to be tracked, and returns to the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined. The embodiment of the present invention realizes information exchange between points to be tracked and captures complex dependencies across points by extracting the corresponding spatiotemporal dependency features of the points to be tracked in the current video frame segment, thereby effectively improving the point tracking accuracy and efficiency under target occlusion. At the same time, in the iterative update process of the trajectory, a preset adaptive sliding window strategy is adopted, which can dynamically adjust the sliding window size of the video frame segment, so that the technical solution can effectively cope with long-term occlusion scenarios, further improving the robustness and performance of point trajectory tracking.

[0120] Further, based on the above-mentioned embodiment of the invention, the trajectory initialization module 41 includes:

[0121] An initial feature extraction unit, used to call a convolutional neural network of a pre-trained point tracking model to extract an initial feature map corresponding to each frame image in the current video frame segment;

[0122] A position trajectory initialization unit, used to determine the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment according to the current position information, and to flatten the initial position coordinates to each frame image to obtain an initial position trajectory;

[0123] The feature trajectory initialization unit is used to call a preset bilinear sampling algorithm to determine the feature vector corresponding to the initial position coordinates in the initial feature map corresponding to the first frame image, and to flatten the feature vector to the initial feature map corresponding to each frame image to obtain the initial feature trajectory.

[0124] Further, based on the above-mentioned embodiment of the invention, the feature determination module 42 includes:

[0125] A correlation matrix construction unit is used to determine the inner product result between the feature vector of the point to be tracked in the initial feature trajectory and the initial feature map of the corresponding time step, and extract multi-scale similarity features from the inner product result to construct a correlation matrix corresponding to the current video frame segment;

[0126] The spatiotemporal dependency feature determination unit is used to splice the correlation matrix, the initial feature trajectory and the initial position trajectory into a fusion feature, determine the temporal self-attention and spatial self-attention corresponding to the fusion feature, and use the sum of the temporal self-attention and the spatial self-attention as the spatiotemporal dependency feature.

[0127] Further, based on the above-mentioned embodiment of the invention, the target trajectory determination module 43 includes:

[0128] A trajectory updating unit is used to use the spatiotemporal dependent features as inputs of the feedforward neural network layer, generate trajectory update amounts corresponding to the initial feature trajectory and the initial position trajectory at the current iteration number, and then obtain the updated initial feature trajectory and the initial position trajectory at the current iteration number based on the trajectory update amounts;

[0129] A first target trajectory determination unit, configured to obtain a target feature trajectory and a target position trajectory after performing iterative updates for a preset number of iterations;

[0130] The occlusion prediction unit is used to input the target feature trajectory into the linear layer and the Sigmoid activation layer in sequence to determine the position occlusion state of the last frame image corresponding to the point to be tracked under the current preset window size;

[0131] A window updating unit, configured to determine a new preset window size according to a sliding window updating formula corresponding to a preset adaptive sliding window strategy when the position occlusion state satisfies a preset window updating condition, and to redetermine a new current video frame segment according to the new preset window size;

[0132] The second target trajectory determination unit is used to re-determine a new target feature trajectory and a new target position trajectory corresponding to the new current video frame segment until the new target feature trajectory corresponds to a new position occlusion state that does not meet the preset window update condition.

[0133] Further, based on the above-mentioned embodiment of the invention, the final trajectory determination module 44 includes:

[0134] A position coordinate determining unit, used to slide in the video frame sequence using a preset window size to obtain a next set of video frame segments, and determine the initial position coordinates of the point to be tracked in the first frame image of the next set of video frame segments according to the target position information;

[0135] A target position trajectory determining unit is used to return to the step of determining an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment, and obtain a target position trajectory corresponding to the next set of video frame segments of the point to be tracked;

[0136] The final position trajectory determination unit is used to sequentially determine the target position trajectory of the point to be tracked corresponding to each video frame segment in the video frame sequence, and splice the target position trajectories into the final position trajectory according to the time dimension.

[0137] Furthermore, based on the above-mentioned embodiments of the invention, during the training process of the point tracking model, a time contrast loss mechanism is used to optimize the feature trajectory, wherein the time contrast loss mechanism is constructed based on positive sample pairs and negative sample pairs.

[0138] Further, based on the above-mentioned embodiment of the invention, the sliding window update formula is expressed as: W' = W + 2 n ; Wherein, W represents the preset window size; W' represents the new preset window size; n represents the number of times the position occlusion state meets the preset window update condition during the feature trajectory update process for the same video frame segment.

[0139] The point tracking device provided in the embodiment of the present invention can execute the point tracking method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0140] Embodiment 5

[0141] Fig. 9 A schematic diagram of an electronic device 50 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0142] like Fig. 9As shown, the electronic device 50 includes at least one processor 51, and a memory connected to the at least one processor 51 in communication, such as a read-only memory (ROM) 52, a random access memory (RAM) 53, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 51 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 52 or the computer program loaded from the storage unit 58 to the random access memory (RAM) 53. In the RAM 53, various programs and data required for the operation of the electronic device 50 can also be stored. The processor 51, the ROM 52, and the RAM 53 are connected to each other via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0143] A number of components in the electronic device 50 are connected to the I / O interface 55, including: an input unit 56, such as a keyboard, a mouse, etc.; an output unit 57, such as various types of displays, speakers, etc.; a storage unit 58, such as a disk, an optical disk, etc.; and a communication unit 59, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 59 allows the electronic device 50 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0144] The processor 51 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 51 executes the various methods and processes described above, such as the point tracking method.

[0145] In some embodiments, the point tracking method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 50 via the ROM 52 and / or the communication unit 59. When the computer program is loaded into the RAM 53 and executed by the processor 51, one or more steps of the point tracking method described above may be performed. Alternatively, in other embodiments, the processor 51 may be configured to perform the point tracking method in any other suitable manner (e.g., by means of firmware).

[0146] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0148] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0150] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0151] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0152] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0153] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A point tracking method, It is characterized in that The method comprises: Determine, based on the current position information of the point to be tracked, an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment; Determine the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; Iteratively updating the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, the spatiotemporal dependency feature, and a preset number of iterations to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is used to adjust a preset window size corresponding to the current video frame segment according to the target feature trajectory; Determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and use the next set of video frame segments as the current video frame segments, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined; The iterative updating of the initial feature trajectory and the initial position trajectory based on the preset adaptive sliding window strategy, the spatiotemporal dependency feature and the preset number of iterations to generate the target feature trajectory and the target position trajectory corresponding to the current video frame segment includes: Using the spatiotemporal dependency feature as the input of the feedforward neural network layer, generating trajectory update amounts corresponding to the initial feature trajectory and the initial position trajectory at the current iteration number, and then obtaining the updated initial feature trajectory and the initial position trajectory at the current iteration number based on the trajectory update amounts; After performing the iterative update for the preset number of iterations, the target feature trajectory and the target position trajectory are obtained; The target feature trajectory is sequentially input into the linear layer and the Sigmoid activation layer to determine the position occlusion state of the point to be tracked corresponding to the last frame image under the current preset window size; When the position occlusion state satisfies a preset window update condition, a new preset window size is determined according to a sliding window update formula corresponding to the preset adaptive sliding window strategy, and a new current video frame segment is re-determined according to the new preset window size; A new target feature track and a new target position track corresponding to the new current video frame segment are re-determined until the new position occlusion state corresponding to the new target feature track does not satisfy the preset window update condition.

2. The method according to claim 1, It is characterized in that The determining, based on the current position information of the point to be tracked, an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment comprises: Calling a convolutional neural network of a pre-trained point tracking model to extract an initial feature map corresponding to each frame image in the current video frame segment; Determine the initial position coordinates of the point to be tracked in the first frame image of the current video frame segment according to the current position information, and spread the initial position coordinates to each frame image to obtain the initial position trajectory; In the initial feature map corresponding to the first frame image, a preset bilinear sampling algorithm is called to determine the feature vector corresponding to the initial position coordinates, and the feature vector is flattened to the initial feature map corresponding to each frame image to obtain the initial feature trajectory.

3. The method according to claim 1, It is characterized in that The determining of the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment includes: Determine an inner product result between a feature vector of the point to be tracked in the initial feature trajectory and an initial feature map of a corresponding time step, and extract a multi-scale similarity feature from the inner product result to construct a correlation matrix corresponding to the current video frame segment; The correlation matrix, the initial feature trajectory and the initial position trajectory are spliced ​​into a fusion feature, the temporal self-attention and the spatial self-attention respectively corresponding to the fusion feature are determined, and the sum of the temporal self-attention and the spatial self-attention is used as the spatiotemporal dependency feature.

4. The method according to claim 1, It is characterized in that The step of determining the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and using the target position information as the current position information of the point to be tracked, and using the next group of video frame segments as the current video frame segments, and returning to the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined, includes: Sliding the preset window size in the video frame sequence to obtain a next set of video frame segments, and determining the initial position coordinates of the point to be tracked in the first frame image of the next set of video frame segments according to the target position information; The next set of video frame segments is used as the current video frame segment, and the step of determining the initial feature trajectory and the initial position trajectory corresponding to the point to be tracked in the current video frame segment is returned to obtain the target position trajectory of the point to be tracked corresponding to the next set of video frame segments; The target position trajectory of each video frame segment corresponding to the point to be tracked is determined in sequence in the video frame sequence, and each of the target position trajectories is spliced ​​into the final position trajectory according to the time dimension.

5. The method according to claim 2, It is characterized in that During the training process of the point tracking model, a time contrast loss mechanism is used to optimize the feature trajectory, wherein the time contrast loss mechanism is implemented based on positive sample pairs and negative sample pairs.

6. The method according to claim 1, It is characterized in that The sliding window update formula is expressed as: W' = W + 2 n ; Wherein, W represents the preset window size; W' represents the new preset window size; n represents the number of times the position occlusion state meets the preset window update condition during the feature trajectory update process for the same video frame segment.

7. A point tracking device, It is characterized in that The device comprises: A trajectory initialization module, used to determine an initial feature trajectory and an initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked; A feature determination module, used to determine the spatiotemporal dependency features corresponding to the point to be tracked in the current video frame segment; A target trajectory determination module, configured to iteratively update the initial feature trajectory and the initial position trajectory based on a preset adaptive sliding window strategy, the spatiotemporal dependency feature, and a preset number of iterations, to generate a target feature trajectory and a target position trajectory corresponding to the current video frame segment; wherein the preset adaptive sliding window strategy is configured to adjust a preset window size corresponding to the current video frame segment according to the target feature trajectory; a final trajectory determination module, configured to determine the target position information of the last frame image of the point to be tracked in the current video frame segment based on the target position trajectory, and use the target position information as the current position information of the point to be tracked, and use the next set of video frame segments as the current video frame segments, and return to the step of determining the initial feature trajectory and initial position trajectory corresponding to the point to be tracked in the current video frame segment based on the current position information of the point to be tracked, until the final position trajectory of the video frame sequence corresponding to the point to be tracked is determined; Wherein, the target trajectory determination module includes: A trajectory updating unit, configured to use the spatiotemporal dependent features as inputs of a feedforward neural network layer, generate trajectory update amounts corresponding to the initial feature trajectory and the initial position trajectory at a current number of iterations, and then obtain the updated initial feature trajectory and the initial position trajectory at the current number of iterations based on the trajectory update amounts; A first target trajectory determination unit, configured to obtain the target feature trajectory and the target position trajectory after performing iterative updates for the preset number of iterations; An occlusion prediction unit is used to input the target feature trajectory into a linear layer and a Sigmoid activation layer in sequence, and determine the position occlusion state of the point to be tracked corresponding to the last frame image under the current preset window size; A window updating unit, configured to determine a new preset window size according to a sliding window updating formula corresponding to the preset adaptive sliding window strategy when the position occlusion state satisfies a preset window updating condition, and to redetermine a new current video frame segment according to the new preset window size; The second target trajectory determination unit is used to re-determine a new target feature trajectory and a new target position trajectory corresponding to the new current video frame segment until the new position occlusion state corresponding to the new target feature trajectory does not meet the preset window update condition.

8. An electronic device, It is characterized in that The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the point tracking method according to any one of claims 1 to 6.

9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the point tracking method according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • Multi-target tracking method and system based on spatial-temporal trajectory association

    CN114913200A

  • Method and system for video-based vehicle tracking adaptable to traffic conditions

    US20150146917A1