Video target tracking method, system and device based on time-space feature fusion
By introducing a spatiotemporal feature fusion module into the video target tracking method, the time and space fusion perception of multi-frame images is solved, and the existing method has insufficient tracking accuracy in complex environments is achieved, and efficient and high-precision video target tracking is achieved.
Patent Information
- Application Number
- CN202510266039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing video target tracking methods are difficult to effectively track fast moving or severely deformed targets in complex environments, and lack effective template feature updates and timing feature extraction methods, resulting in insufficient tracking accuracy and robustness in the case of large changes in the target.
The video target tracking method based on temporal and spatial feature fusion is adopted. By building a model including an encoder, a spatiotemporal and spatial feature fusion module and a decoder, the time and space fusion perception of multi-frame images are performed, cross-temporal and spatial perception features are extracted, and position and bounding box prediction are performed to achieve efficient tracking of the target.
It effectively improves the timing perception ability of the model, enhances the ability to track fast moving or severely deformed targets in complex environments, reduces memory consumption and computing complexity, and achieves efficient and high-precision video target tracking.
Smart Images

Figure CN119784798B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a video target tracking method, system and equipment based on time-space feature fusion. Background Art
[0002] Target tracking and identification (such as littering in road monitoring) plays an important role in traffic safety supervision. The existing monitoring methods for littering in traffic safety supervision are mainly carried out manually, which is not capable of large-scale data processing and supervision. Using the target tracking method in deep learning to batch process video images and locate the time and location of bad behaviors such as vehicle littering provides a feasible solution for traffic safety supervision.
[0003] Compared with most tracking scenarios, tracking traffic spills has the following five main challenges:
[0004] (1) Object size is extremely small. In traffic surveillance scenarios, objects thrown from vehicles, such as peels and cigarette butts, are usually very small. This problem is particularly prominent given that these videos are captured by cameras located next to the road.
[0005] (2) Objects move quickly and have severe motion blur. Objects thrown from a moving vehicle inherit some of the vehicle’s kinetic energy, and from the perspective of a stationary surveillance camera, they appear to be moving very fast. In addition, these thrown objects are subject to external forces such as air resistance and gravity, which can change their trajectories. As a result, the motion blur of thrown objects is quite severe, especially considering the limited frame rate of traffic surveillance cameras.
[0006] (3) The target is easily deformed and occluded. In fact, the most common thrown objects are food packaging and paper products, which are prone to significant changes in shape and appearance during the throwing process. In addition, in a complex and changing traffic environment, thrown garbage may be occluded by buildings, pedestrians, roadside infrastructure or other vehicles, causing the tracking algorithm to lose the target.
[0007] (4) Lighting conditions vary widely. Discarded objects may move quickly from well-lit areas to shadowed or dimly lit areas. In addition, changing lighting conditions during the day from the angle of the sun and from streetlights and vehicle headlights at night can significantly affect the visibility of objects, further complicating the tracking process.
[0008] (5) Complex background. Traffic environments usually have highly complex backgrounds, including moving vehicles, various road surfaces, traffic signs, and other dynamic and static elements. This complex background generates visual noise, making it challenging to detect and track small, fast-moving projectiles.
[0009] In general, the targets and target backgrounds in traffic spillage target tracking scenarios vary greatly and the situation is complex.
[0010] Existing target tracking methods mainly focus on estimating the trajectory of a single target in a series of frames given an initial target template image, which can be divided into correlation filter-based and twin network-based tracking methods.
[0011] Correlation filter-based algorithms track targets by using correlation filters to compare the similarity responses of a reference template with the search area in subsequent frames. Early trackers mainly relied on hand-crafted features to represent targets, but these hand-crafted features and their priors have difficulty capturing high-level semantic information that is critical for positioning and tracking, especially in complex scenes. In recent years, many researchers have tried to combine the powerful representation capabilities of convolutional neural networks with the efficient positioning characteristics of correlation filters to improve tracking performance. For example, convolutional neural networks trained in multiple fields combined with correlation filters can adaptively select and update features to achieve more accurate tracking. However, since correlation filter-based algorithms usually rely on simpler feature and template update mechanisms, although such methods are usually fast, their robustness and adaptability in dynamic or complex environments (such as occlusion, fast motion, and scale and rotation changes) are limited.
[0012] On the other hand, the tracking method based on the twin network determines the target position by measuring the similarity between the candidate area in the new frame and the reference object. The early twin network method extracts features from the template image and the image to be tested through a convolutional network, and uses cross-correlation to measure the similarity between the template and the image to be tested to locate the target position in the image to be tested. Many researchers have improved the feature extraction model by introducing the Transformer model, and improved the similarity measurement network by introducing the attention mechanism and self-attention mechanism, which effectively improves the performance of the twin network and significantly improves its tracking accuracy. However, most methods based on the twin network only extract features from a single frame, resulting in the failure to fully utilize temporal information. This deficiency makes it difficult for such methods to track fast-moving or severely deformed objects in complex environments.
[0013] In general, the existing video target tracking methods still have the following shortcomings:
[0014] (1) Lack of effective template feature update method: Most existing methods focus on global feature extraction, ignoring the changes in target detail information in adjacent frames and failing to update template features in a timely manner according to target changes. This results in insufficient tracker accuracy and robustness when the target changes greatly.
[0015] (2) Lack of effective temporal feature extraction methods: Existing methods ignore the long-term continuity of feature changes during target tracking and make it difficult to predict future changes of the target based on the global temporal features in the video sequence, which affects the accuracy of recognition.
[0016] In summary, although the existing video object tracking has made significant progress in accuracy and robustness, there are still many shortcomings. In particular, in terms of template feature update method and temporal feature extraction method, the existing methods have not yet provided a satisfactory solution. Summary of the invention
[0017] In view of the above technical problems, the present invention provides a video target tracking method, system and device based on time-space feature fusion.
[0018] The technical solution adopted by the present invention to solve the technical problem is:
[0019] A video target tracking method based on time-space feature fusion, the method comprising the following steps:
[0020] S100: Obtaining video data to be processed, the video data to be processed including a plurality of image frames and the position of the target to be tracked in the first frame, and preprocessing the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set;
[0021] S200: Building a video target tracking model, the model includes an encoder module, a spatiotemporal feature fusion module and a decoder module connected in sequence;
[0022] S300: Inputting multiple frames of template image data and the image data to be detected into the encoder module for feature extraction, and outputting image features;
[0023] S400: Inputting the template image target relative position data and image features into the spatiotemporal feature fusion module for spatiotemporal feature fusion processing, and outputting cross-spatiotemporal perception features;
[0024] S500: Output the cross-temporal and spatial perception features to the decoder module for position and bounding box prediction, and output the position result of the target to be detected;
[0025] S600: Train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained, and the real-time video image data of the target to be detected is acquired, pre-processed into template image data, template image target relative position data and image data to be detected, and input into the trained video target tracking model, and the entire video data is processed in a loop to obtain the position of the target to be tracked in the entire video.
[0026] Preferably, the encoder module includes several layers of encoding blocks connected in sequence, the decoder module includes several layers of decoding blocks connected in sequence, the spatiotemporal feature fusion module includes several spatiotemporal feature fusion blocks, and the several layers of encoding blocks and the several layers of decoding blocks other than the top layer are respectively connected through several spatiotemporal feature fusion blocks.
[0027] Preferably, the encoder module is a deep residual network, the deep residual network includes a plurality of residual blocks connected in sequence, and S300 includes:
[0028] The template image data and the image data to be detected are input into the deep residual network for processing, and a number of residual blocks connected in sequence extract a number of image features from multiple frames of image data in sequence.
[0029] Preferably, each spatiotemporal feature fusion block in the spatiotemporal feature fusion module includes a position feature prediction module and a two-dimensional selection scanning module, and S400 includes:
[0030] S410: The position feature prediction module receives the target relative position feature of the template image to obtain the target relative position feature of the image to be detected;
[0031] S420: All target relative position features and corresponding image features are added and concatenated to obtain image features with target position marks, and then processed by a two-dimensional selection scanning module to obtain spatiotemporal information fusion features and output them.
[0032] Preferably, the position feature prediction module includes an input gate, a forget gate and an output gate connected in sequence, and S410 includes:
[0033] S411: The position feature prediction module initializes internal hidden features, state features, normalized state features, and stabilizer state features;
[0034] S412: The input gate concatenates the hidden features and the relative position features of the template image target in the feature dimension, processes them through a linear layer, and obtains the input features of the position feature prediction block;
[0035] S413: The forget gate processes the prediction block input feature, the hidden feature of the previous moment, the state feature, the normalized state feature and the stabilizer state feature to obtain a new state feature, a normalized state feature and a stabilizer state feature;
[0036] S414: The output gate processes the new state feature, the normalized state feature and the stabilizer state feature to obtain a new hidden feature, and uses the hidden feature as an output to obtain an output target relative position feature.
[0037] Preferably, the two-dimensional selection scanning module includes a first normalization layer, a first multi-layer perceptron, a two-dimensional selection scanning layer, a second normalization layer, and a second multi-layer perceptron connected in sequence, and S420 includes:
[0038] S421: adding all target relative position features and corresponding image features and concatenating them to obtain image features with target position marks;
[0039] S422: The first normalization layer receives and processes the image features with the target position mark to obtain the first layer of normalized linear features, the first multilayer perceptron receives and processes the first layer of normalized linear features to obtain the first multilayer perceptron features, the two-dimensional selective scanning layer receives the first multilayer perceptron features to obtain the two-dimensional selective scanning features, the second normalization layer receives and processes the two-dimensional selective scanning features to obtain the second layer of normalized linear features, and the second multilayer perceptron receives and processes the second layer of normalized linear features to obtain the second multilayer perceptron features;
[0040] S423: Crop the second multi-layer perceptron features to obtain spatiotemporal fusion features of the image to be detected.
[0041] Preferably, the decoder module includes a position heat map prediction module and a bounding box regression module, each of which is composed of two multi-layer perceptrons. S500 includes:
[0042] The position heat map prediction module receives the relative position features of the output target and performs position heat map prediction to obtain the position heat map prediction result. The bounding box regression module receives the relative position features of the output target and performs bounding box regression prediction to obtain the bounding box regression prediction result. The position heat map prediction result is flattened to obtain the position of the maximum value, and the bounding box regression result at this position is obtained to obtain the position of the target in the image to be detected.
[0043] Preferably, the preprocessing process in S600 is specifically as follows:
[0044] During the first detection, the first frame image and the annotated tracking target position are preprocessed as template image data and template image target relative position data, the second frame image is preprocessed as the image data to be detected and the target position in the second frame image is detected. In the detection process of subsequent frames, the first frame image and the annotated tracking target position and the image of the previous frame of the current frame and the model predicted tracking target position are preprocessed as template image data and template image target relative position data, the current frame image is preprocessed as the image data to be detected, and the model is used to track and identify the entire video data frame by frame to complete the detection and identification of the video target.
[0045] The video target tracking system based on spatiotemporal feature fusion includes a data set acquisition module, a video target tracking model building module, an encoder module, a spatiotemporal feature fusion module, a decoder module, a model training and target tracking module;
[0046] The data set acquisition module is used to acquire the video data to be processed, which includes multiple image frames and the position of the target to be tracked in the first frame, and preprocess the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set;
[0047] A video target tracking model building module is used to build a video target tracking model. The model includes an encoder module, a spatiotemporal feature fusion module, and a decoder module connected in sequence;
[0048] The encoder module is used to receive multiple frames of template image data and image data to be detected, perform feature extraction, and output image features;
[0049] The spatiotemporal feature fusion module is used to receive the template image target relative position data and image features for processing and output cross-spatiotemporal perception features;
[0050] The decoder module is used to receive and process the cross-temporal and spatial perception features, output the position result of the target to be detected, and cyclically process the entire video data to obtain the position of the target to be tracked in the entire video;
[0051] The model training and target tracking module is used to train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained. The real-time video data to be processed is processed based on the trained video target tracking model to obtain the position of the target to be tracked in the video.
[0052] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a video target tracking method based on time-space feature fusion when executing the computer program.
[0053] The above-mentioned video target tracking method, system and device based on spatiotemporal feature fusion improve the shortcomings of Transformer's insufficient long-term temporal memory capacity and large amount of attention mechanism calculation. It uses the spatiotemporal feature fusion module to perform temporal and spatial fusion perception of multiple frame images, effectively improving the model's temporal perception ability, enhancing the ability to track fast-moving or severely deformed targets in complex environments, and reducing memory consumption and computational complexity, and can achieve video target tracking efficiently and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1Flow chart of a video target tracking method based on time-space feature fusion in one embodiment of the present invention;
[0055] Figure 2 is a schematic diagram of the structure of a target tracking model in one embodiment of the present invention;
[0056] Figure 3 is a schematic diagram of the working process of a position prediction module in one embodiment of the present invention;
[0057] Figure 4 It is a schematic diagram of the working process of the selection scanning module in one embodiment of the present invention. DETAILED DESCRIPTION
[0058] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings.
[0059] In response to the problems in the prior art, the present invention proposes a video target tracking method based on spatiotemporal feature fusion, which improves the shortcomings of the Transformer's insufficient long-term temporal memory capacity and the large amount of attention mechanism computation. The spatiotemporal feature fusion module is used to perform temporal and spatial fusion perception of multiple frame images, thereby improving the ability to track fast-moving or severely deformed objects in complex environments.
[0060] In one embodiment, Figure 1 As shown, a video target tracking method based on time-space feature fusion comprises the following steps:
[0061] S100: Obtaining video data to be processed, the video data to be processed including a plurality of image frames and the position of the target to be tracked in the first frame, and preprocessing the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set;
[0062] S200: Building a video target tracking model, the model includes an encoder module, a spatiotemporal feature fusion module and a decoder module connected in sequence;
[0063] S300: Inputting multiple frames of template image data and the image data to be detected into the encoder module for feature extraction, and outputting image features;
[0064] S400: Inputting the template image target relative position data and image features into the spatiotemporal feature fusion module for spatiotemporal feature fusion processing, and outputting cross-spatiotemporal perception features;
[0065] S500: Output the cross-temporal and spatial perception features to the decoder module for position and bounding box prediction, and output the position result of the target to be detected;
[0066] S600: Train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained, and the real-time video image data of the target to be detected is acquired, pre-processed into template image data, template image target relative position data and image data to be detected, and input into the trained video target tracking model, and the entire video data is processed in a loop to obtain the position of the target to be tracked in the entire video.
[0067] In one embodiment, the encoder module includes several layers of encoding blocks connected in sequence, the decoder module includes several layers of decoding blocks connected in sequence, the spatiotemporal feature fusion module includes several spatiotemporal feature fusion blocks, and the several layers of encoding blocks and the several layers of decoding blocks other than the top layer are respectively connected through several spatiotemporal feature fusion blocks.
[0068] Specifically, Figure 2 As shown in the figure, the encoder module includes 5 layers of sequentially connected encoding blocks, which encode several frames of template images and one frame of image to be detected. The decoder module includes a position heat map prediction module and a bounding box regression module. The encoding block and the decoding block are correspondingly connected through several spatiotemporal feature fusion blocks.
[0069] In one embodiment, the encoder module is a deep residual network, the deep residual network includes a plurality of residual blocks connected in sequence, the residual blocks correspond to encoding blocks, and S300 includes:
[0070] The template image data and the image data to be detected are input into the deep residual network for processing, and a number of residual blocks connected in sequence extract a number of image features from multiple frames of image data in sequence.
[0071] For example, a deep residual network ResNet-50 or ResNet-101 can be used as an encoder module. Taking ResNet-101 as an example, the deep residual network ResNet-101 includes 5 residual blocks connected in sequence. The 5 residual blocks are used as corresponding encoding blocks. The template image data and the image data to be detected are input into the deep residual network ResNet-101 for processing. The 5 residual blocks use pre-trained weights to extract 5 image features from the template image data and the image data to be detected in sequence, including: the first image feature F_1, the second image feature F_2, the third image feature F_3, the fourth image feature F_4 and the fifth image feature F_5, and the fifth image feature F_5 is output as the final acquired image feature.
[0072] In one embodiment, each spatiotemporal feature fusion block in the spatiotemporal feature fusion module includes a position feature prediction module and a two-dimensional selection scanning module, and S400 includes:
[0073] S410: The position feature prediction module receives the target relative position feature of the template image to obtain the target relative position feature of the image to be detected;
[0074] S420: All target relative position features and corresponding image features are added and concatenated to obtain image features with target position marks, and then processed by a two-dimensional selection scanning module to obtain spatiotemporal information fusion features and output them.
[0075] Specifically, the spatiotemporal feature fusion module consists of two modules in series, namely the position feature prediction module (e.g. Figure 3 ) and the 2D selection scanning module (as shown Figure 4 As shown in the figure, the former is used to learn the continuous position change information in the template features, and the latter is used to efficiently extract the global information between the template image and the image to be detected, and perform temporal and spatial fusion perception of multiple frame images, which effectively improves the model's temporal perception ability and enhances the target tracking model's ability to track fast-moving or severely deformed targets in complex environments.
[0076] In one embodiment, Figure 3 As shown, the position feature prediction module includes an input gate, a forget gate and an output gate connected in sequence, and S410 includes:
[0077] S411: The position feature prediction module initializes internal hidden features, state features, normalized state features, and stabilizer state features;
[0078] S412: The input gate concatenates the hidden features and the relative position features of the template image target in the feature dimension, processes them through a linear layer, and obtains the input features of the position feature prediction block;
[0079] S413: The forget gate processes the prediction block input feature, the hidden feature of the previous moment, the state feature, the normalized state feature and the stabilizer state feature to obtain a new state feature, a normalized state feature and a stabilizer state feature;
[0080] S414: The output gate processes the new state feature, the normalized state feature and the stabilizer state feature to obtain a new hidden feature, and uses the hidden feature as an output to obtain an output target relative position feature.
[0081] Specifically, after initializing the hidden features, state features, normalized state features and stabilizer state features inside the position feature prediction module, the relative position features of the template image target are input into the position feature prediction module one by one, processed through the linear layer and the position feature prediction block input features are obtained, the position feature prediction block input features are processed through the input gate, forgetting gate and output gate in turn, the hidden features, state features, normalized state features and stabilizer state features are updated, and the final hidden features are used as output to obtain the relative position features of the image target to be detected.
[0082] In one embodiment, the two-dimensional selection scanning module includes a first normalization layer, a first multi-layer perceptron, a two-dimensional selection scanning layer, a second normalization layer, and a second multi-layer perceptron connected in sequence, and S420 includes:
[0083] S421: adding all target relative position features and corresponding image features and concatenating them to obtain image features with target position marks;
[0084] S422: The first normalization layer receives and processes the image features with the target position mark to obtain the first layer of normalized linear features, the first multilayer perceptron receives and processes the first layer of normalized linear features to obtain the first multilayer perceptron features, the two-dimensional selective scanning layer receives the first multilayer perceptron features to obtain the two-dimensional selective scanning features, the second normalization layer receives and processes the two-dimensional selective scanning features to obtain the second layer of normalized linear features, and the second multilayer perceptron receives and processes the second layer of normalized linear features to obtain the second multilayer perceptron features;
[0085] S423: Crop the second multi-layer perceptron features to obtain spatiotemporal fusion features of the image to be detected.
[0086] Specifically, all target relative position features and corresponding image features are added and concatenated to obtain image features with target position marks. The image features with target position marks are processed in sequence by the first normalization layer, the first multi-layer perceptron, the two-dimensional selective scanning layer, the second normalization layer and the second multi-layer perceptron, and finally cropped to obtain the relative position features of the image target to be detected.
[0087] In one embodiment, the decoder module includes a position heat map prediction module and a bounding box regression module, each of which is composed of two multi-layer perceptrons. S500 includes:
[0088] The position heat map prediction module receives the relative position features of the output target and performs position heat map prediction to obtain the position heat map prediction result. The bounding box regression module receives the relative position features of the output target and performs bounding box regression prediction to obtain the bounding box regression prediction result. The position heat map prediction result is flattened to obtain the position of the maximum value, and the bounding box regression result at this position is obtained to obtain the position of the target in the image to be detected.
[0089] In one embodiment, the preprocessing process in S600 is specifically as follows:
[0090] During the first detection, the first frame image and the annotated tracking target position are preprocessed as template image data and template image target relative position data, the second frame image is preprocessed as the image data to be detected and the target position in the second frame image is detected. In the detection process of subsequent frames, the first frame image and the annotated tracking target position and the image of the previous frame of the current frame and the model predicted tracking target position are preprocessed as template image data and template image target relative position data, the current frame image is preprocessed as the image data to be detected, and the model is used to track and identify the entire video data frame by frame to complete the detection and identification of the video target.
[0091] The present invention applies for a video target tracking method based on spatiotemporal feature fusion. Specifically, a position feature prediction module is designed to extract temporal features from the position information of multi-frame template features, predict the target position of the image to be tested, and provide help information for subsequent target tracking. In addition, a two-dimensional selection scanning module is designed to flatten the input image features with position marks into one-dimensional vectors along four different directions, and dynamically interact with the one-dimensional vectors, reducing the quadratic complexity of the Transformer to linear complexity, thereby achieving faster and more accurate spatiotemporal feature fusion extraction. Sufficient experiments on the road scattered object target tracking dataset have verified the effectiveness of the method proposed in the present invention in video target tracking tasks. Compared with existing target tracking methods, the present invention can more effectively achieve target tracking, and has higher tracking accuracy for targets with large lighting changes, blurred backgrounds, fast movement, and severe deformation.
[0092] The video target tracking system based on spatiotemporal feature fusion includes a data set acquisition module, a video target tracking model building module, an encoder module, a spatiotemporal feature fusion module, a decoder module, a model training and target tracking module;
[0093] The data set acquisition module is used to acquire the video data to be processed, which includes multiple image frames and the position of the target to be tracked in the first frame, and preprocess the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set;
[0094] A video target tracking model building module is used to build a video target tracking model. The model includes an encoder module, a spatiotemporal feature fusion module, and a decoder module connected in sequence;
[0095] The encoder module is used to receive multiple frames of template image data and image data to be detected, perform feature extraction, and output image features;
[0096] The spatiotemporal feature fusion module is used to receive the template image target relative position data and image features for processing and output cross-spatiotemporal perception features;
[0097] The decoder module is used to receive and process the cross-temporal and spatial perception features, output the position result of the target to be detected, and cyclically process the entire video data to obtain the position of the target to be tracked in the entire video;
[0098] The model training and target tracking module is used to train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained. The real-time video data to be processed is processed based on the trained video target tracking model to obtain the position of the target to be tracked in the video.
[0099] For the specific definition of the video target tracking system based on time-space feature fusion, please refer to the definition of the video target tracking method based on time-space feature fusion mentioned above, which will not be repeated here. Each module in the above-mentioned video target tracking system based on time-space feature fusion can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0100] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a video target tracking method based on time-space feature fusion when executing the computer program.
[0101] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0102] The above is a detailed introduction to the video target tracking method, system and device based on time-space feature fusion provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A video target tracking method based on temporal and spatial feature fusion, characterized in that: The method comprises the following steps: S100: Obtaining video data to be processed, the video data to be processed including a plurality of image frames and the position of the target to be tracked in the first frame, and preprocessing the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set; S200: Building a video target tracking model, the model includes an encoder module, a spatiotemporal feature fusion module and a decoder module connected in sequence; S300: Inputting multiple frames of template image data and the image data to be detected into the encoder module for feature extraction, and outputting image features; S400: Inputting the template image target relative position data and image features into the spatiotemporal feature fusion module for spatiotemporal feature fusion processing, and outputting cross-spatiotemporal perception features; each spatiotemporal feature fusion block in the spatiotemporal feature fusion module includes a position feature prediction module and a two-dimensional selection scanning module, S400 includes: S410: The position feature prediction module receives the target relative position feature of the template image to obtain the target relative position feature of the image to be detected; S420: All target relative position features and corresponding image features are added and concatenated to obtain image features with target position marks, and then processed by a two-dimensional selection scanning module to obtain spatiotemporal information fusion features and output; S500: Output the cross-temporal and spatial perception features to the decoder module for position and bounding box prediction, and output the position result of the target to be detected; S600: Train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained, and the real-time video image data of the target to be detected is acquired, pre-processed into template image data, template image target relative position data and image data to be detected, and input into the trained video target tracking model, and the entire video data is processed in a loop to obtain the position of the target to be tracked in the entire video.
2. The method according to claim 1, characterized in that The encoder module includes several layers of encoding blocks connected in sequence, the decoder module includes several layers of decoding blocks connected in sequence, the spatiotemporal feature fusion module includes several spatiotemporal feature fusion blocks, and the several layers of encoding blocks and the several layers of decoding blocks, except the top layer, are respectively connected through several spatiotemporal feature fusion blocks.
3. The method according to claim 2, characterized in that The encoder module is a deep residual network, which includes a plurality of residual blocks connected in sequence, and the residual blocks correspond to encoding blocks. S300 includes: The template image data and the image data to be detected are input into the deep residual network for processing, and a number of residual blocks connected in sequence extract a number of image features from multiple frames of image data in sequence.
4. The method according to claim 3, characterized in that The position feature prediction module includes an input gate, a forget gate, and an output gate connected in sequence, and S410 includes: S411: The position feature prediction module initializes internal hidden features, state features, normalized state features, and stabilizer state features; S412: The input gate concatenates the hidden features and the relative position features of the template image target in the feature dimension, processes them through a linear layer, and obtains the input features of the position feature prediction block; S413: The forget gate processes the prediction block input feature, the hidden feature of the previous moment, the state feature, the normalized state feature and the stabilizer state feature to obtain a new state feature, a normalized state feature and a stabilizer state feature; S414: The output gate processes the new state feature, the normalized state feature and the stabilizer state feature to obtain a new hidden feature, and uses the hidden feature as an output to obtain an output target relative position feature.
5. The method according to claim 4, characterized in that The two-dimensional selection scanning module includes a first normalization layer, a first multi-layer perceptron, a two-dimensional selection scanning layer, a second normalization layer, and a second multi-layer perceptron connected in sequence, and S420 includes: S421: adding all target relative position features and corresponding image features and concatenating them to obtain image features with target position marks; S422: The first normalization layer receives and processes the image features with the target position mark to obtain the first layer of normalized linear features, the first multilayer perceptron receives and processes the first layer of normalized linear features to obtain the first multilayer perceptron features, the two-dimensional selective scanning layer receives the first multilayer perceptron features to obtain the two-dimensional selective scanning features, the second normalization layer receives and processes the two-dimensional selective scanning features to obtain the second layer of normalized linear features, and the second multilayer perceptron receives and processes the second layer of normalized linear features to obtain the second multilayer perceptron features; S423: Crop the second multi-layer perceptron features to obtain spatiotemporal fusion features of the image to be detected.
6. The method according to claim 5, characterized in that The decoder module includes a position heat map prediction module and a bounding box regression module, each consisting of two multi-layer perceptrons. S500 includes: The position heat map prediction module receives the relative position features of the output target and performs position heat map prediction to obtain the position heat map prediction result. The bounding box regression module receives the relative position features of the output target and performs bounding box regression prediction to obtain the bounding box regression prediction result. The position heat map prediction result is flattened to obtain the position of the maximum value, and the bounding box regression result at this position is obtained to obtain the position of the target in the image to be detected.
7. The method according to claim 6, characterized in that The preprocessing process in S600 is as follows: During the first detection, the first frame image and the annotated tracking target position are preprocessed as template image data and template image target relative position data, the second frame image is preprocessed as the image data to be detected and the target position in the second frame image is detected. In the detection process of subsequent frames, the first frame image and the annotated tracking target position and the image of the previous frame of the current frame and the model predicted tracking target position are preprocessed as template image data and template image target relative position data, the current frame image is preprocessed as the image data to be detected, and the model is used to track and identify the entire video data frame by frame to complete the detection and identification of the video target.
8. The video target tracking system based on time-space feature fusion is characterized by: It includes data set acquisition module, video target tracking model building module, encoder module, spatiotemporal feature fusion module, decoder module, model training and target tracking module; The data set acquisition module is used to acquire the video data to be processed, which includes multiple image frames and the position of the target to be tracked in the first frame, and preprocess the video data to be processed into template image data, template image target relative position data and image data to be detected as a data set; A video target tracking model building module is used to build a video target tracking model. The model includes an encoder module, a spatiotemporal feature fusion module, and a decoder module connected in sequence; The encoder module is used to receive multiple frames of template image data and image data to be detected, perform feature extraction, and output image features; The spatiotemporal feature fusion module is used to receive the template image target relative position data and image features for processing and output cross-spatiotemporal perception features; Each spatiotemporal feature fusion block in the spatiotemporal feature fusion module includes a position feature prediction module and a two-dimensional selection scanning module. The spatiotemporal feature fusion module includes: the position feature prediction module receives the target relative position feature of the template image to obtain the target relative position feature of the image to be detected; all the target relative position features are added and spliced with the corresponding image features to obtain the image feature with the target position mark, and then processed by the two-dimensional selection scanning module to obtain the spatiotemporal information fusion feature and output it; The decoder module is used to receive and process the cross-temporal and spatial perception features, output the position result of the target to be detected, and cyclically process the entire video data to obtain the position of the target to be tracked in the entire video; The model training and target tracking module is used to train the video target tracking model based on the training set and the preset loss function. When the preset training end condition is reached, the trained video target tracking model is obtained. The real-time video data to be processed is processed based on the trained video target tracking model to obtain the position of the target to be tracked in the video.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.