Target tracking method and device based on self-mask transformation neural network
The self-mask transformation neural network eliminates the chaotic correlation between the search area and the template, generates target-oriented features, solving the problems of high error detection rate and unstable performance in visual target tracking, and achieving more efficient target tracking effects.
Patent Information
- Application Number
- CN202510410023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Existing visual target tracking techniques have problems with high false detection rates, failed tracking and unstable performance when dealing with drastic deformation, partial occlusion, complex background and scale changes, especially due to the inadequate target-non-target distinction capability due to the chaotic correlation between the search area and the template.
The self-mask transformation neural network is used to eliminate the chaotic correlation between the search area and the template through non-binding position coding, and build a self-learning adaptive target-oriented representation, introduce mask attention, design coded attention to suppress interference from non-target information, and generate target-oriented feature expressions.
It effectively reduces the false detection rate of target tracking, improves the success rate, accuracy and stability of multi-scenario and multi-target tracking, and achieves a more stable target tracking effect.
Smart Images

Figure CN120339330A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of object tracking. Specifically, it relates to an object tracking method and device based on a self-masking transformation neural network. Background Art
[0002] Visual object tracking is a basic task in the field of computer vision, and its goal is to estimate the future state of any target of interest based on initial manual annotations. Therefore, visual object tracking has been widely applied in autonomous driving, human-computer interaction systems, and intelligent monitoring. With the popularization of visual transformation neural networks, the tracking performance of visual object tracking has been further improved, but there are still many challenges that have not been overcome, such as drastic deformation, partial occlusion, complex background, and scale variation.
[0003] Although the self-attention module performs self-modeling and cross-modeling on each token of the template and the search region separately, which is beneficial to connecting the module-search image pair through bidirectional information flow to generate target-oriented features; however, since the search region of visual tracking is usually a cropped area that is 4-5 times larger than the target, this means that most of the tokens in the search region are unnecessary (that is, the background or distractors with appearances similar to the target), and the existing single-stream tracker relationship modeling is the relationship modeling between all tokens in the search region and the template. This will result in the template token naturally performing an undesired relationship modeling with non-target tokens when the discriminative ability in the early stage of feature extraction is not strong enough. On the one hand, as the network deepens and iterates, the template tokens may aggregate to the features of the background or distractors, reducing the quality of the template tokens. On the other hand, non-target tokens will aggregate to the template containing target information, confusing the target in the search tokens and unable to accurately find the correct target. These will all inhibit the ability of the single-stream tracker to distinguish between target and non-target, thereby inhibiting the performance of the tracker, resulting in a high false detection rate, tracking failure, unstable performance, and insufficient adaptability. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an object tracking method and device based on a self-masking transformation neural network to solve the above problems existing in the prior art and achieve stable and accurate object tracking.
[0005] In a first aspect, an object tracking method based on a self-masking transformation neural network is provided, and the method may include:
[0006] Obtain a target video and a tracking template corresponding to the target video; wherein, both the target video and the tracking template contain a tracking target;
[0007] For any video frame in the target video except the first frame, obtain the historical localization information of the tracking target in the previous video frame of the video frame;
[0008] Generate a target search image corresponding to the video frame according to the historical localization information and the video frame;
[0009] Input the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain the guiding feature information of the tracking target;
[0010] Input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the localization information of the tracking target in the video frame.
[0011] In an alternative implementation, the method for obtaining the tracking template includes:
[0012] Obtain the boundary coordinates of the labeled area in the first video frame of the target video;
[0013] Enlarge the boundary coordinates by a first multiple of the configured magnification to obtain enlarged boundary coordinates;
[0014] Crop the tracking template from the first video frame based on the enlarged boundary coordinates.
[0015] In an alternative implementation, the historical localization information includes historical localization boundary coordinates;
[0016] The generating of the target search image corresponding to the video frame according to the historical localization information and the video frame includes:
[0017] Enlarge the historical localization boundary coordinates by a second multiple of the configured magnification to obtain enlarged historical localization boundary coordinates;
[0018] Crop the target search image from the video frame based on the enlarged historical localization boundary coordinates.
[0019] In an alternative implementation, the self-masking transformation neural network includes: a fusion layer, a self-masking transformation neural network encoder, and a self-masking transformation neural network decoder;
[0020] The fusion layer is used to respectively extract the features of the tracking template and the target search image to obtain a first tracking template feature and a first target search image feature; fuse the tracking template feature and the target search image feature to obtain a fused feature;
[0021] The self-masked transformation neural network encoder is used to perform position encoding and feature encoding on the tracking template, the target search image, and the fused features to obtain attention-encoded features;
[0022] The self-masked transformation neural network decoder is used to perform cross-attention and self-attention decoding on the attention-encoded features to obtain the guiding feature information of the tracking target.
[0023] In an optional implementation, the self-masked transformation neural network encoder is specifically used for:
[0024] Perform absolute position encoding on the tracking template and the target search image to obtain a first encoded vector and a second encoded vector; concatenate the first encoded vector and the second encoded vector to obtain a third encoded vector;
[0025] Perform relative position encoding on the fused features to obtain a fourth encoded vector;
[0026] Add the third encoded vector and the fourth encoded vector to obtain a target encoded vector;
[0027] Perform attention encoding on the target encoded vector and the fused features to obtain attention-encoded features.
[0028] In an optional implementation, the attention-encoded features include: a second tracking template feature and a second target search image feature;
[0029] The self-masked transformation neural network decoder is specifically used for:
[0030] Perform matrix multiplication on the second target search image feature and the attention-encoded features to obtain an initial mask;
[0031] After transforming the initial mask through the Sigmoid function, perform thresholding processing to obtain a binarized mask;
[0032] Perform attention decoding on the second target search image feature, the attention-encoded features, and the binarized mask to obtain cross-attention decoding features;
[0033] Perform self-attention decoding on the cross-attention decoding features to obtain the guiding feature information of the tracking target.
[0034] In an optional implementation, input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame, including:
[0035] Generate a guiding feature map of the tracking target based on the guiding feature information;
[0036] Input the guiding feature map into a pre-trained tracking target center point prediction model to obtain the center point score, detection box size, and offset;
[0037] Determine the center point coordinates based on the center point score;
[0038] Generate the positioning information of the tracking target in the video frame according to the center point coordinates, the detection box size, and the offset.
[0039] In a second aspect, a target tracking device based on a self-masking transformation neural network is provided. The device may include:
[0040] An acquisition unit for acquiring a target video and a tracking template corresponding to the target video; wherein both the target video and the tracking template contain a tracking target; for any video frame in the target video except the first frame, acquire the historical positioning information of the tracking target in the previous video frame of the video frame;
[0041] A generation unit for generating a target search image corresponding to the video frame according to the historical positioning information and the video frame;
[0042] An extraction unit for inputting the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain the guiding feature information of the tracking target;
[0043] A prediction unit for inputting the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame.
[0044] In a third aspect, an electronic device is provided. The electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0045] The memory is used to store a computer program;
[0046] The processor is used to implement any of the method steps in the first aspect when executing the program stored on the memory.
[0047] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program realizes any of the method steps in the first aspect when executed by a processor.
[0048] This application uses unbound position encoding to eliminate the chaotic correlation between the search region and the template, constructs a self-masking decoder to provide a self-learning and adaptive target-oriented representation for visual object tracking, introduces masked attention, effectively avoids the interference of non-target information in the search region, and realizes tracking based on generating target-oriented feature expressions with masked attention.
[0049] The present invention encodes the discriminative target information in the tracking template and the target search image by using a self-masking transformation neural network, designs the attention as encoded attention, thereby adaptively suppressing the distraction of irrelevant non-target search tokens. This application designs a new position encoding to effectively eliminate the chaotic correlation between the search region and the template.
[0050] This application effectively reduces the false detection rate of object tracking, can realize object tracking in multiple scenarios and multiple targets, and improves the tracking success rate, accuracy and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required to be used in the embodiments of this application. It should be understood that the following drawings only show some embodiments of this application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0052] Figure 1 It is a flowchart of a target tracking method based on a self-masking transformation neural network provided by an embodiment of this application;
[0053] Figure 2 It is a schematic structural diagram of a self-masking transformation neural network provided by an embodiment of this application;
[0054] Figure 3 It is a schematic structural diagram of a target tracking device based on a self-masking transformation neural network provided by an embodiment of this application;
[0055] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0057] The object tracking method based on the self-masking transformation neural network provided by the embodiments of the present application can be applied to a server or a terminal with strong computing power. The server can be a physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be a user equipment (UE) such as a mobile phone, a smart phone, a laptop computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (PAD), a handheld device, a vehicle-mounted device, a wearable device, a computing device or other processing devices connected to a wireless modem, a mobile station (MS), a mobile terminal, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application.
[0058] The preferred embodiments of the present application will be described below with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. And without conflict, the embodiments and features in the embodiments of the present application can be combined with each other.
[0059] Figure 1 It is a schematic flowchart of an object tracking method based on the self-masking transformation neural network provided by the embodiments of the present application. As Figure 1 shown, the method may include:
[0060] Step S110, obtain a target video and a tracking template corresponding to the target video; for any video frame in the target video except the first frame, obtain the historical positioning information of the tracking target in the previous video frame of the video frame.
[0061] In the embodiments of the present application, the target video is composed of multiple video frames and the time points corresponding to each video frame; both the target video and the tracking template contain the tracking target.
[0062] In the embodiments of the present application, the tracking target is pre-annotated in the first video frame of the target video, and based on the annotation information, a tracking template corresponding to the target video is generated. The specific steps are as follows:
[0063] Obtain the boundary coordinates of the labeled area in the first video frame of the target video; magnify the boundary coordinates by a configured first multiple to obtain the magnified boundary coordinates; based on the magnified boundary coordinates, crop the tracking template from the first video frame.
[0064] In the embodiment of the present application, the boundary coordinates of the labeled area are the boundary coordinates of the detection box, including the upper left corner coordinates and the lower right corner coordinates of the labeled area; the first multiple can be set according to the actual situation.
[0065] For example, the boundary coordinates of the detection box labeled in the first video frame of the target video are (1, 0) and (0, 1), and the first multiple is 2 times. Then the magnified boundary coordinates are (2, 0) and (0, 2). The rectangular area in the first video frame of the target video corresponding to the magnified boundary coordinates (2, 0) and (0, 2) is used as the tracking template.
[0066] In the embodiment of the present application, the historical positioning information includes historical positioning boundary coordinates; the historical positioning information is actually historical tracking box information, that is, the tracking box information obtained by tracking the tracking target in the previous video frame, including the upper left corner coordinates and the lower right corner coordinates of the tracking box.
[0067] Step S120: Generate a target search image corresponding to the video frame according to the historical positioning information and the video frame; input the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain the guiding feature information of the tracking target.
[0068] In the embodiment of the present application, generating a target search image corresponding to the video frame according to the historical positioning information and the video frame includes:
[0069] Magnify the historical positioning boundary coordinates by a configured second multiple to obtain the magnified historical positioning boundary coordinates; based on the magnified historical positioning boundary coordinates, crop the target search image from the video frame.
[0070] For example, the tracking result of the previous video frame of the current video frame, that is, the positioning information Box of the tracking target in the previous video frame of the current video frame m-1 , for the current video frame I m , m ∈ [2, n], according to the tracking box information of the tracking result of the previous frame, magnify it by 5 times and then crop out the target search image S m .
[0071] In the embodiment of the present application, as Figure 2 shown, the self-masking transformation neural network includes:
[0072] A fusion layer, a self-masking transformation neural network encoder, and a self-masking transformation neural network decoder;
[0073] A fusion layer, configured to respectively extract features of a tracking template and a target search image to obtain a first tracking template feature and a first target search image feature; fuse the tracking template feature and the target search image feature to obtain a fused feature;
[0074] A self-masked transformation neural network encoder, configured to perform positional encoding and feature encoding on the tracking template, the target search image, and the fused feature to obtain an attention encoded feature;
[0075] A self-masked transformation neural network decoder, configured to perform cross-attention and self-attention decoding on the attention encoded feature to obtain guiding feature information of a tracking target.
[0076] In an embodiment of this application, the fusion layer is specifically configured to:
[0077] Respectively extract features of the tracking template and the target search image to obtain a first tracking template feature and a first target search image feature Fuse the tracking template feature and the target search image feature to obtain a fused feature Wherein, Generally represents the size of a feature, C represents the number of channels, H represents the height, W represents the width, and B represents the input batch size.
[0078] In an embodiment of this application, the tracking template and the target search image are each downsampled through a shared fully convolutional embedding layer to obtain a first tracking template feature and a first target search image feature
[0079] In an embodiment of this application, the self-masked transformation neural network encoder adopts a transformation neural network encoder architecture to perform several attention encodings on the input features; specifically including:
[0080] Perform absolute positional encoding on the tracking template and the target search image to obtain a first encoded vector and a second encoded vector
[0081] Concatenate and the second encoded vector to obtain a third encoded vector Perform encoding on the fused feature Perform relative position encoding to obtain the fourth encoding vector Add the third encoding vector and the fourth encoding vector to obtain the target encoding vector Perform attention encoding on the target encoding vector and the fused feature to obtain the attention-encoded feature
[0082] In the embodiment of the present application, performing attention encoding on the target encoding vector and the fused feature includes:
[0083]
[0084] wherein, is generated by after passing through a linear layer, and Q i , K i , V i represent the linear mapping parameter matrices of the i-th head; h represents the number of heads in the calculation process, d = C / h, and T represents the transpose; in practical applications, the self-masking transformation neural network encoder performs feature encoding on the tracking template, the target search image, and the fused feature 12 times, that is, the feature encoding process needs to be repeated 12 times.
[0085] In the embodiment of the present application, the attention-encoded feature includes: the second tracking template feature and the second target search image feature; cropping the attention-encoded feature according to the original length (i.e., the sizes of the tracking template and the target search image), the second tracking template feature and the second target search image feature
[0086] In the embodiment of the present application, the self-masking transformation neural network decoder uses cross-attention and self-attention for adaptive decoding to determine the guiding feature information of the tracking target; specifically including:
[0087] Perform matrix multiplication on the second target search image feature and the attention-encoded feature to obtain the initial mask After transforming the initial mask through the Sigmoid function, perform thresholding processing to obtain the binarized mask Perform matrix multiplication on the second target search image feature the attention-encoded feature and the binarized mask Perform attention decoding to obtain cross-attention decoding features For the cross-attention decoding features Perform self-attention decoding to obtain the guiding feature information of the tracking target
[0088] In the embodiment of the present application, the thresholding process includes: mapping all elements below the threshold of 0.5 to negative infinity, and mapping elements greater than or equal to the threshold of 0.5 to 1.
[0089] In the embodiment of the present application, performing attention decoding on the second target search image features, attention encoding features, and the binarized mask includes:[[]]
[0090] For the second target search image features And the feature sequence And the mask Perform attention decoding, and the specific method is as follows:
[0091]
[0092] Wherein, Is generated after passing through the linear layer by ; Is generated after passing through the linear layer by ; Is generated from the binarized mask M0; h represents the number of heads in the calculation process, d = C / h, and T represents transpose.
[0093] In the embodiment of the present application, perform self-attention decoding on the cross-attention decoding features :
[0094]
[0095] Wherein, Is generated after passing through the linear layer by X SA ; h represents the number of heads in the calculation process, d = C / h, and T represents transpose.
[0096] In the embodiment of the present application, the self-mask transformation neural network needs to be trained in two stages, and the training method includes:
[0097] In the first stage, preprocess the training dataset, select two frames with an interval of T in the video sequence, and according to the annotation information, crop the template image and the search image to sizes of 128×128 and 256×256; input the preprocessed training dataset into the deep learning model for training, calculate the joint loss during training, perform backpropagation, and update the model parameters to complete the training;
[0098] In the second stage, the parameters of the quality assessment network are fine-tuned. The training dataset is preprocessed. Two frames with an interval of T in the video sequence are selected. According to the annotation information, the cropped template image and the search image are resized to 128×128 and 256×256 sizes; the preprocessed training dataset is input into the deep learning model for training. During training, the cross-entropy loss is calculated, backpropagation is performed, and the model parameters are updated to complete the training.
[0099] In the embodiment of the present application, the combined loss is expressed by the following formula:
[0100]
[0101] where L iοu represents the intersection over union loss, which is used to measure the distance between the ground truth and the predicted value, L1 represents the mean absolute error loss, and L focal represents the cross-entropy loss, which is used to alleviate the problem of imbalance between positive and negative samples. λ iοu , represents the weight of the corresponding loss function, for example, 5 and 2 respectively. b i and represent the ground truth and the predicted bounding box;
[0102] During the training process of the first stage, the batch size is 80, the learning rate decreases from 0.0001 to 0.00001, the AdamW algorithm is used for iterative training 300 times, and the results of each iteration are saved. The last 100 iterations start training with one-tenth of the overall network learning rate. It should be noted that only the parameters of the transformation neural network are fine-tuned in this stage.
[0103] In the embodiment of the present application, the cross-entropy loss is expressed by the following formula:
[0104]
[0105] where y i represents the ground truth, where the presence of the tracking target is 1 and the absence is 0. p i represents the reliability score of the final prediction.
[0106] During the training process of the second stage, the batch size is 256, the learning rate decreases from 0.0001 to 0.00001, the AdamW algorithm is used for iterative training 40 times, and the results of each iteration are saved. The last 10 iterations start training with one-tenth of the overall network learning rate. It should be noted that only the parameters of the quality assessment network are fine-tuned in this stage, and the parameters of the transformation neural network are frozen throughout the process.
[0107] Step S130: Input the guiding feature information into the pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame.
[0108] In the embodiment of the present application, the tracking target center point prediction model is composed of 3 different convolutional layers.
[0109] In the embodiment of the present application, the guiding feature information is input into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame, including:
[0110] Based on the guiding feature information, a guiding feature map of the tracking target is generated; the guiding feature map is input into a pre-trained tracking target center point prediction model to obtain the center point score, the size of the detection box (or tracking box or bounding box), and the offset; based on the center point score, the center point coordinates are determined; according to the center point coordinates, the size of the detection box, and the offset, the positioning information of the tracking target in the video frame is generated.
[0111] In the embodiment of the present application, the guiding feature information is deformed into a new guiding feature map Guiding feature map The center point score, the size of the detection box (or tracking box or bounding box), and the offset of the prediction result are obtained through 3 different convolutional layers; the center point coordinates of the prediction result are obtained according to the center point score, and the positioning boundary coordinates are obtained according to the obtained width and height and the offset, and finally the positioning information Box is obtained. m 。
[0112] Corresponding to the above method, the embodiment of the present application further provides a target tracking device based on a self-masking transformation neural network, as Figure 3 shown. The target tracking device based on the self-masking transformation neural network includes:
[0113] An acquisition unit 310, configured to acquire a target video and a tracking template corresponding to the target video; wherein, both the target video and the tracking template include a tracking target; for any video frame except the first frame in the target video, acquire the historical positioning information of the tracking target in the previous video frame of the video frame;
[0114] A generation unit 320, configured to generate a target search image corresponding to the video frame according to the historical positioning information and the video frame;
[0115] An extraction unit 330, configured to input the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain the guiding feature information of the tracking target;
[0116] A prediction unit 340, configured to input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame.
[0117] The functions of the functional units of the object tracking device based on the self-masking transformation neural network provided in the above embodiments of the present application can be implemented by the above method steps. Therefore, the specific working processes and beneficial effects of each unit in the object tracking device based on the self-masking transformation neural network provided in the embodiments of the present application will not be repeated here.
[0118] The embodiments of the present application also provide an electronic device, as Figure 4 shown, including a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440.
[0119] The memory 430 is used to store a computer program;
[0120] When the processor 410 is used to execute the program stored on the memory 430, the following steps are implemented:
[0121] Obtain a target video and a tracking template corresponding to the target video; wherein, both the target video and the tracking template contain a tracking target;
[0122] For any video frame except the first frame in the target video, obtain the historical positioning information of the tracking target in the previous video frame of the video frame;
[0123] Generate a target search image corresponding to the video frame according to the historical positioning information and the video frame;
[0124] Input the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain the guiding feature information of the tracking target;
[0125] Input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame.
[0126] The above-mentioned communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0127] The communication interface is used for communication between the above electronic device and other devices.
[0128] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0129] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0130] Since the implementation manners and beneficial effects of the various components of the electronic device in the above embodiments can be seen Figure 1 from the steps in the embodiments shown, therefore, the specific working process and beneficial effects of the electronic device provided in the embodiments of the present application will not be repeated here.
[0131] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute any one of the target tracking methods based on the self-masking transformation neural network in the above embodiments.
[0132] In another embodiment provided by the present application, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute any one of the target tracking methods based on the self-masking transformation neural network in the above embodiments.
[0133] Those skilled in the art should understand that the embodiments in the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments in the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments in the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0134] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a means for implementing the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the specified functions in multiple blocks.
[0135] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the specified functions in multiple blocks.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the specified functions in multiple blocks.
[0137] Although the preferred embodiments in the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0138] Obviously, those skilled in the art can make various changes and variations to the embodiments in the embodiments of the present application without departing from the spirit and scope of the embodiments in the embodiments of the present application. Thus, if these modifications and variations of the embodiments in the embodiments of the present application fall within the scope of the claims of the embodiments of the present application and their equivalent technologies, the embodiments in the embodiments of the present application are also intended to include these changes and variations.
Claims
1. A target tracking method based on a self-masking transformation neural network, characterized in that The method includes: Obtaining a target video and a tracking template corresponding to the target video; wherein, both the target video and the tracking template contain a tracking target; For any video frame in the target video except the first frame, obtaining historical positioning information of the tracking target in the previous video frame of the video frame; Generating a target search image corresponding to the video frame according to the historical positioning information and the video frame; Inputting the tracking template and the target search image into a pre-trained self-masking transformation neural network to obtain guiding feature information of the tracking target; Inputting the guiding feature information into a pre-trained tracking target center point prediction model to obtain positioning information of the tracking target in the video frame.
2. The method according to claim 1, wherein The method for obtaining the tracking template includes: Obtaining the boundary coordinates of the labeled area in the first video frame of the target video; Magnifying the boundary coordinates by a first multiple of the configured magnification to obtain magnified boundary coordinates; Cropping the tracking template from the first video frame based on the magnified boundary coordinates.
3. The method according to claim 1, wherein The historical positioning information includes historical positioning boundary coordinates; The generating of the target search image corresponding to the video frame according to the historical positioning information and the video frame includes: Magnifying the historical positioning boundary coordinates by a second multiple of the configured magnification to obtain magnified historical positioning boundary coordinates; Cropping the target search image from the video frame based on the magnified historical positioning boundary coordinates.
4. The method according to claim 1, wherein The self-masking transformation neural network includes: a fusion layer, a self-masking transformation neural network encoder, and a self-masking transformation neural network decoder; The fusion layer is used to respectively extract the features of the tracking template and the target search image to obtain a first tracking template feature and a first target search image feature; and fuse the tracking template feature and the target search image feature to obtain a fused feature; The self-masking transformation neural network encoder is used to perform position encoding and feature encoding on the tracking template, the target search image, and the fused feature to obtain attention encoding features; The self-masking transformation neural network decoder is used to perform cross-attention and self-attention decoding on the attention encoding features to obtain the guiding feature information of the tracking target.
5. The method according to claim 4, characterized in that The self-masking transformation neural network encoder is specifically used for: Performing absolute position encoding on the tracking template and the target search image to obtain a first encoding vector and a second encoding vector; splicing the first encoding vector and the second encoding vector to obtain a third encoding vector; Performing relative position encoding on the fused feature to obtain a fourth encoding vector; Adding the third encoding vector and the fourth encoding vector to obtain a target encoding vector; Performing attention encoding on the target encoding vector and the fused feature to obtain attention encoding features.
6. The method according to claim 5, characterized in that The attention encoding features include: a second tracking template feature and a second target search image feature; The self-masking transformation neural network decoder is specifically used for: Perform matrix multiplication on the second target search image features and the attention encoding features to obtain an initial mask; After transforming the initial mask through the Sigmoid function, perform thresholding processing to obtain a binarized mask; Perform attention decoding on the second target search image features, the attention encoding features, and the binarized mask to obtain cross-attention decoding features; Perform self-attention decoding on the cross-attention decoding features to obtain the guiding feature information of the tracking target.
7. The method according to claim 1, wherein Input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame, including: Generate a guiding feature map of the tracking target based on the guiding feature information; Input the guiding feature map into a pre-trained tracking target center point prediction model to obtain a center point score, a detection box size, and an offset; Determine the center point coordinates based on the obtained center point score; Generate the positioning information of the tracking target in the video frame according to the center point coordinates, the detection box size, and the offset.
8. An object tracking device based on a self-masking transformation neural network, characterized in that, The device includes: An acquisition unit, configured to acquire a target video and a tracking template corresponding to the target video; wherein both the target video and the tracking template include a tracking target; for any video frame other than the first frame in the target video, acquire the historical positioning information of the tracking target in the previous video frame of the video frame; A generation unit, configured to generate a target search image corresponding to the video frame according to the historical positioning information and the video frame; An extraction unit, configured to input the tracking template and the target search image into a pre-trained self-mask transformation neural network to obtain the guiding feature information of the tracking target; A prediction unit, configured to input the guiding feature information into a pre-trained tracking target center point prediction model to obtain the positioning information of the tracking target in the video frame.
9. An electronic device, characterized in that, The electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is configured to implement the method according to any one of claims 1-7 when executing the program stored on the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-7.