Method and System for Coping with Target Rotation Based on Rotation-Invariant Equivariant Networks
By introducing rotating and other variable twin neural networks and rotation object reference data sets, the tracking problem of twin neural networks in rotation scenarios is solved, and stable target tracking in in-plane and out-of-plane rotation scenarios is achieved, and the accuracy and efficiency of the visual target tracking algorithm are improved.
Patent Information
- Application Number
- CN202111623248.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The existing twin neural network trackers are prone to failure when rotating in and out of the plane of the target, making it difficult to effectively solve the rotation problem, resulting in insufficient accuracy and efficiency of visual target tracking algorithms in complex scenarios.
The rotation isovariant twin neural network is adopted, and the target tracking of in-plane and out-of-plane rotation is achieved by introducing controllable filters and rotation isovariant convolutional layers, combined with the concept of variance of rotation, and the rotation object reference data set is used for training and testing.
It realizes stable target tracking in different rotation scenarios, broadens the application scope of visual target tracking, and improves the accuracy and efficiency of the tracker in complex scenarios.
Smart Images

Figure CN114913077B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision artificial intelligence, and more specifically, to an object tracking method and system that uses a rotation-invariant attention mechanism and an equivariant network to address the problem of object rotation. Background Art
[0002] Visual tracking refers to detecting, extracting, identifying, and tracking moving objects in an image sequence to obtain the motion parameters of the moving objects, such as position, speed, acceleration, and motion trajectory, etc., so as to perform further processing and analysis, realize the understanding of the behavior of the moving objects, and complete a higher-level detection task. In recent years, due to the rapid development of industries such as intelligent monitoring, human-computer interaction, and security monitoring, visual object tracking has become one of the popular research topics in the field of computer vision.
[0003] In recent years, although researchers have proposed a large number of excellent visual object tracking algorithms, due to the complexity and variability of the tracking scenarios, such as similar distractors, partial occlusion, etc., many algorithms still have many defects. The main difficulty lies in the balance between accuracy and efficiency.
[0004] Currently, in the field of visual object tracking, the vast majority of state-of-the-art trackers (such as Siamese trackers) adopt the Siamese networks model. Existing Siamese network trackers are very popular in the field of visual object tracking because they obtain strong discriminative ability from similarity matching and have become the basic framework of most state-of-the-art tracking algorithms. However, although Siamese network trackers generally perform well, they are still prone to failure when one of the two inputs rotates. Because such trackers are mainly translation-equivariant in nature and are not designed to solve the rotation problem. Therefore, in the case of in-plane rotation in the two inputs, conventional Siamese networks will still fail. This is mainly because, as is well known, Siamese networks are a type of neural network architecture that contains two or more identical sub-networks. The convolutional neural networks used by them do not have equivariance to the rotation of the object.
[0005] A more straightforward solution is to force the learning of rotation variables by using a training set with rotation attributes or data augmentation, but this is undoubtedly costly, and data augmentation also has certain limitations.
[0006] Therefore, for a long time, rotation has been one of the long-standing but still unsolved difficult challenges in visual object tracking technology. Existing deep learning-based tracking algorithms use conventional convolutional neural networks, although these convolutional neural networks have certain translation-equivariant properties, they are not designed to solve the rotation problem.
[0007] Therefore, there is an urgent need in the art for a visual object tracking method and system that can solve the in-plane and out-of-plane rotation problems. Summary of the Invention
[0008] A brief overview of one or more aspects is given below to provide a basic understanding of these aspects. This overview is not an exhaustive survey of all contemplated aspects, and is neither intended to identify key or decisive elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0009] Therefore, to solve the in-plane and out-of-plane rotation problems of an object relative to a reference frame, the present application proposes a method using a rotation-equivariant Siamese neural network to solve the object rotation problem. By using the method according to the present application, the problems of in-plane and out-of-plane rotation of the object can be overcome, thereby greatly broadening the application scope and scenarios of visual object tracking.
[0010] Specifically, the present application constructs by using a set of equivariant convolutional layers containing controllable filters, aiming to add rotation-equivariant characteristics to an existing Siamese neural network tracker. This characteristic will allow the tracker to capture rotation changes from the beginning without additional data augmentation. We introduce rotation equivariance for the object localization task in video frames, utilize the concept of group-equivariant CNN, and use controllable filters to make the Siamese tracker rotation-equivariant. This way of combining rotation equivariance introduces built-in weight sharing between different rotation groups and adds a concept of internal rotation and external rotation to the model.
[0011] In addition, in order to more accurately compare the performance advantages and disadvantages of the existing tracker and the tracker proposed in the present application, the present application also proposes the concept of a rotated object benchmark, which is a set of video data sets focusing on in-plane rotation (i.e., 2D rotation, such as a 20-degree plane rotation, etc.) and out-of-plane rotation (i.e., 3D rotation, such as rotating from the front to the side). The annotations include the bounding box of the target object, its orientation in each frame, and the rotation type (in-plane), while the out-of-plane rotation annotation only includes the bounding box of the target object and the rotation type (out-of-plane).
[0012] According to a first aspect of the present application, a method for dealing with object rotation based on a rotation-invariant equivariant network is disclosed, the method comprising:
[0013] Inputting a template image and a video frame image of the object;
[0014] Convolving the template image and the video frame image respectively to obtain their feature maps;
[0015] Perform cross - correlation on the feature maps of the template image and the video frame image;
[0016] Pool the results of the cross - correlation respectively to output heatmaps;
[0017] Compare the output heatmaps with the database annotation results;
[0018] If the results are consistent, determine the position of the target according to the output heatmaps,
[0019] Otherwise, use the output heatmaps for back - training the convolved feature maps.
[0020] According to a preferred embodiment of the present application, obtaining their feature maps by convolving the template image and the video frame image respectively further includes:
[0021] Perform convolution on the template image and the video frame image through rotation - equivariant convolution to obtain in - plane rotation - equivariant feature maps; and
[0022] Perform convolution on the template image and the video frame image through rotation - invariant attention mechanism convolution to obtain out - of - plane rotation - invariant feature maps.
[0023] According to a preferred embodiment of the present application, performing cross - correlation on the feature maps of the template image and the video frame image further includes:
[0024] Perform cross - correlation on the rotation - equivariant feature maps and rotation - invariant feature maps of the template image and the video frame image respectively.
[0025] According to a preferred embodiment of the present application, pooling the results of the cross - correlation respectively to output heatmaps further includes:
[0026] Perform max - pooling operation on the rotation - equivariant feature maps to obtain rotation - equivariant heatmaps; and
[0027] Perform average - pooling operation on the rotation - invariant feature maps to obtain rotation - invariant heatmaps.
[0028] According to a preferred embodiment of the present application, the output heatmaps are used to indicate the predicted target positions, and the database annotation results are the trained data sets in the database used to indicate the correct positions of the targets, and
[0029] If the comparison results are inconsistent, use the output heatmaps together with the database annotation results for back - training the convolved feature maps.
[0030] According to the second aspect of the present application, a system for a method of dealing with target rotation based on a rotation - invariant equivariant network is disclosed. The system includes:
[0031] An input component for inputting a template image and a video frame image of a target;
[0032] A convolutional layer component that convolves the template image and the video frame image respectively to obtain their feature maps;
[0033] A cross - correlation component that performs cross - correlation on the feature maps of the template image and the video frame image;
[0034] A pooling layer component that pools the results of the cross - correlation respectively to output heatmaps;
[0035] A training component that compares the output heatmaps with the database annotation results;
[0036] An output component for determining the position of the target according to the output heatmap when the results are consistent, and
[0037] A back - training component for using the output heatmaps to back - train the convolved feature maps when the results are inconsistent.
[0038] To achieve the foregoing and related purposes, one or more aspects include the features that are fully described hereinafter and particularly pointed out in the appended claims. The following description and the drawings detail certain illustrative features of one or more aspects. However, these features are merely indicative of several of the various ways in which the principles of the various aspects may be employed, and this description is intended to cover all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] For a more specific description of the manner in which the features set forth above are used, reference may be made to the aspects, some of which are illustrated in the drawings. It should be noted, however, that the drawings illustrate only certain typical aspects of the present application and should not be considered as limiting its scope, since the description may admit of other equally effective aspects.
[0040] In the drawings:
[0041] Figure 1 is a framework diagram of a system 100 for coping with target rotation based on a rotation - invariant equivariant network according to an embodiment of the present application; and
[0042] Figure 2 is a flowchart of a method 200 for coping with target rotation based on a rotation - invariant equivariant network according to an embodiment of the present application; and
[0043] Figure 3 is a schematic block diagram of the main aspects of a system and method for coping with target rotation based on a rotation - invariant equivariant network according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following detailed description presented in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details to provide a thorough understanding of the various concepts. However, it will be apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, well-known components are shown in block diagram form to avoid obscuring such concepts.
[0045] It should be understood that, based on this disclosure, other embodiments will be apparent and that system, structural, process, or mechanical changes may be made without departing from the scope of this disclosure.
[0046] Referring Figure 1 and 2 , aspects are depicted with reference to one or more components and one or more methods that can perform the actions or functions described herein. In one aspect, the term "component" as used herein can be one of the parts that make up a system, can be hardware or software or some combination thereof, and can be divided into other components. Although the operations described below in Figure 2 are presented in a particular order and / or as performed by example components, it should be understood that the order of these actions and the components performing the actions can vary depending on the implementation. Additionally, it should be understood that the following actions or functions can be performed by a specially programmed processor, a processor executing specially programmed software or a computer-readable medium, or any other combination of hardware components and / or software components capable of performing the described actions or functions.
[0047] As described above, in the field of visual object tracking, most of its algorithms rely on siamese neural networks. The cross-correlation response result heatmap of the existing siamese neural network is calculated as follows:
[0048] h(z,x) = f(z)*f(x) (1)
[0049] where z and x represent the template frame image and the candidate frame image respectively, f(·) represents the encoding function of the siamese neural network, and * represents the convolution operation.
[0050] Figure 1 illustrates a framework diagram of a system 100 for coping with object rotation based on a rotation-invariant equivariant network according to an embodiment of the present application.
[0051] As Figure 1As shown in [reference], the rotation-equivariant Siamese neural network is based on the existing Siamese neural network structure and is optimized and improved. The existing Siamese neural network mainly includes: an input part, a convolutional layer, and a correlation response output part. The rotation-equivariant Siamese neural network proposed in this application replaces the basic convolutional layer with a rotation-equivariant model and introduces a set of max-pooling models to select the most appropriate cross-correlation encoding, so as to generate the most appropriate direction among multiple heatmaps.
[0052] As Figure 1 shown in [reference], the system mainly includes the following components:
[0053] Input component 101, convolutional layer component 102, cross-correlation component 103, pooling layer component 104, training component 105, output component 106, and backpropagation training component 107. The main functions of each component will be described in detail below with reference to the attached Figure 1 drawings.
[0054] Input component 101
[0055] The input component 101 is mainly used to input the template image and the target video frame image.
[0056] The candidate branches of the network structure (i.e., the input target video frame images) select a set of normal search images as the input and modify the input of the template branch. Among them, n rotation variable sets Z = {z1, z2,..., z n n} are defined, which includes an out-of-plane rotation variable.
[0057] Convolutional layer component 102
[0058] The convolutional layer component 102 is mainly used to perform convolution on the input target image information. It is mainly divided into two parts: rotation-equivariant convolution and rotation-invariant attention mechanism convolution for in-plane rotation and out-of-plane rotation respectively.
[0059] For in-plane rotation, first rotate the first n - 1 variable sets and the entire frame in the plane centered on the target, and then perform cropping to obtain the feature map. Among them, each input image I contains C channels, and each channel is represented by I C , where c ∈ {1, 2,..., C}. The input is then convolved with rotation filters , where see the following expression (2).
[0060] Then, the nth variable is convolved with a set of out-of-plane rotation attention mechanism filters , where see the following expression (3).
[0061] The result features obtained before applying the non-linear activation are shown in the following expression (4). The filter rotates the variant in the equidistant direction θ, which is represented by the set as shown.
[0062]
[0063]
[0064]
[0065] Learn the invariance of the out-of-plane rotation of the target plane through expression (3). During the model training process, a supervised method is adopted to enable this branch to obtain the target invariance features during out-of-plane rotation.
[0066] Generate feature maps through the above expressions (2) and (3) and further process them using group convolution, generalizing the spatial convolution to a wider group of transformations. Similar to the first layer, the steerable filter bank is defined as the following expression (5). The additional exponent introduced for the weight tensor in equation (5) facilitates the group convolution operation along the rotation dimension.
[0067]
[0068] Cross-correlation component 103
[0069] The cross-correlation component is mainly used to perform cross-correlation operations on the feature maps of the convolved template image and the video frame image to obtain the cross-correlation result.
[0070] Pooling layer component 104
[0071] The pooling layer component is mainly used to further process the output of a group of convolutional layers through max pooling in the rotation dimension, obtain the out-of-plane rotation invariance features through average pooling, and the out-of-plane rotation invariance attention mechanism branch still needs to perform isovariance to maintain rotation along the spatial dimension.
[0072] As described above, through the convolutional layer component, two sets of feature maps {φ(z)} and φ(x) are obtained from the two sub-network branches (in-plane rotation and out-of-plane rotation) of the rotation-equivariant siamese network, where {φ(z)} is a feature set containing Λ directions. Then {φ(z)} and φ(x) are convolved to obtain Λ heatmaps The formula is expressed as h i (z,x) = φ(z i ) * φ(x).
[0073] The above heatmap The final output heatmap h(Z, x) is obtained by processing through a global max pooling layer and an average pooling layer respectively. The global max pooling layer is defined as the feature map composed of selecting the maximum values of The global average pooling layer is defined as the feature map composed of selecting the average values of
[0074] Training component 105
[0075] The training component 105 is mainly used to compare the output heatmap (from which the target position can be obtained) with the annotation result in the database (i.e., the correct target position in the training database) to determine whether they are consistent.
[0076] Output component 106
[0077] The output component 106 is used to, if the comparison result is consistent, use the output heatmap to determine the position of the target.
[0078] Backward training component 107
[0079] The backward training component 107 is used, in the case where the comparison result is inconsistent, to use the output heatmap to backward train the neural network model of the convolutional layer component.
[0080] Figure 2 The flowchart of method 200 for target tracking using a rotation-equivariant siamese neural network as shown in Figure 1 is illustrated in
[0081] As shown in Figure 2 the method 200 mainly includes the following steps.
[0082] First, input the template image of the target and the video frame image containing the target through the input component 101 (step 201).
[0083] Second, using the convolutional layer component 102, the template image and the video frame image are respectively convolved through rotation-equivariant convolution and rotation-invariant attention mechanism to obtain an in-plane rotation-equivariant feature map and an out-of-plane rotation-invariant feature map (step 202).
[0084] Specifically, for in-plane rotation, first rotate the first n - 1 variable sets and the entire frame in the plane centered on the target, and then perform cropping to obtain the feature map.
[0085] Each input image I contains C channels, and each channel is represented by I C where c ∈ {1, 2,..., C}. The input is then convolved with rotation filters where Refer to the following expression (2).
[0086] For out-of-plane rotation, for the nth variable, a set of out-of-plane rotation attention mechanism filters is used for convolution operation, where refer to the following expression (3).
[0087] The resulting features obtained before applying the non-linear activation refer to the following expression (4). The filter rotates the variant in the equidistant direction θ and is represented by the set .
[0088]
[0089]
[0090]
[0091] Next, through the cross-correlation component 103, and by performing cross-correlation operations on the rotation-equivariant features of the template image and the video frame image, and the rotation-invariant features of the template image and the video frame image respectively, the cross-correlation operation result is obtained (step 203).
[0092] Then, through the pooling layer component 104, a max-pooling operation is performed on the result of the cross-correlation operation of the rotation-equivariant features to obtain the rotation equivariance along the rotation dimension, that is, the rotation-equivariant heatmap; while an average-pooling operation is performed on the result of the cross-correlation operation of the rotation-invariant features to obtain the rotation equivariance along the spatial dimension, that is, the rotation-invariant heatmap (step 204).
[0093] In the training component 105, the heatmap output by the pooling layer component 104 is compared with the annotation result (i.e., the correct target position) in the database to determine whether they are consistent (step 205).
[0094] If it is consistent with the annotation result in the database, then in step 206, the position of the target is determined according to the output heatmap.
[0095] Otherwise, in step 207, the output heatmap is used for backpropagation training of the neural network model of the convolutional layer component. That is, the output prediction result together with the correct data result is used for backpropagation to correct the network model.
[0096] Figure 3 The schematic block diagram showing the main aspects of the system and method for coping with target rotation based on the rotation-invariant equivariant network according to the embodiments of the present application is illustrated. This block diagram corresponds to the Figure 1 and Figure 2 steps or components shown.
[0097] As described above, the present application is applicable to visual object tracking scenarios with in-plane and out-of-plane rotations. For example, when the video is recorded using a drone camera, other videos recorded from a top view, a camera mounted on a rotating object, and an egocentric video. The present application is also applicable to 3D scenarios such as out-of-plane rotations.
[0098] Different from the prior art methods for object tracking using Siamese networks, the present application has prominent substantial advantages and significant progress compared to the prior art.
[0099] First, the present application expands the in-plane rotation equivariant convolutional network and the out-of-plane rotation attention mechanism, so as to obtain a rotation-equivariant Siamese architecture with in-plane and out-of-plane rotation equivariance, thus being applicable to scenarios where the object rotates in-plane and out-of-plane simultaneously and achieving stable object tracking.
[0100] Second, the maximum pooling layer can achieve maintaining rotation equivariance along the rotation dimension, and the average pooling layer can achieve maintaining rotation equivariance along the spatial dimension.
[0101] Third, for the benchmark test, the present application proposes a rotating object benchmark. This is a new dataset that includes sequences with in-plane rotations of the object while also including sequences with out-of-plane rotations of the object, thus being able to handle the rotation problems of the object both in-plane and out-of-plane, and thereby enabling optimal tracking of the object.
[0102] Aspects, elements, or any part of an element, or any combination of elements according to the present disclosure can be implemented using a "processing system" that includes one or more processors. Examples of processors include: microprocessors, microcontrollers, digital signal processors (DSPs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout the present disclosure. One or more processors in the processing system can execute software. Software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, procedures, functions, etc., regardless of whether it is referred to in terms of software, firmware, middleware, microcode, hardware description language, or other terms. The software can reside on a computer-readable medium. The computer-readable medium can be a non-transitory computer-readable medium. As an example, non-transitory computer-readable media include: magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs)), smart cards, flash memory devices (e.g., memory cards, memory sticks, key drives), random access memory (RAM), read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, removable disks, and any other suitable medium for storing software and / or instructions that can be accessed and read by a computer. As an example, computer-readable media can also include carrier waves, transmission lines, and any other suitable medium for transmitting software and / or instructions that can be accessed and read by a computer. The computer-readable medium can reside within the processing system, outside the processing system, or be distributed across multiple entities including the processing system. The computer-readable medium can be implemented in a computer program product. As an example, the computer program product can include a computer-readable medium in a packaging material. Those skilled in the art will recognize how to best implement the described functionality presented throughout the present disclosure depending on the particular application and overall design constraints imposed on the overall system.
[0103] It should be understood that the specific order or hierarchy of the steps in the disclosed methods is illustrative of exemplary processes. Based on design preferences, it should be understood that the specific order or hierarchy of the steps in the methods or methodologies described herein can be rearranged. The appended method claims present the elements of the various steps in a sample order and are not meant to be limited to the specific order or hierarchy presented, unless specifically recited herein.
[0104] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein but are to be accorded the full scope consistent with the language of the claims, where the singular forms of the elements are not intended to mean "one and only one" (unless specifically stated otherwise) but rather "one or more." The term "some," unless specifically stated otherwise, means one or more. A phrase that recites "at least one" of a list of items refers to any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: at least one a; at least one b; at least one c; at least one a and at least one b; at least one a and at least one c; at least one b and at least one c; and at least one a, at least one b, and at least one c. Elements of the various aspects described throughout this disclosure are expressly incorporated herein by reference to all structural and functional equivalents known to those of ordinary skill in the art currently or hereafter, and are intended to be covered by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is explicitly recited in the claims.
Claims
1. A method for dealing with target rotation based on a rotation-invariant equivariant network, characterized in that The method includes: Inputting a template image and a video frame image of the target; Convolving the template image and the video frame image respectively to obtain their feature maps; Performing cross - correlation on the feature maps of the template image and the video frame image; Pooling the results of the cross - correlation respectively to output heatmaps; Comparing the output heatmaps with the database annotation results; If the results are consistent, determining the position of the target according to the output heatmaps; Otherwise, using the output heatmaps for back - training the convolved feature maps; Convolving the template image and the video frame image respectively to obtain their feature maps further includes: Convolving the template image and the video frame image through rotation - equivariant convolution to obtain in - plane rotation - equivariant feature maps; and Convolving the template image and the video frame image through rotation - invariant attention mechanism convolution to obtain out - of - plane rotation - invariant feature maps; Pooling the results of the cross - correlation respectively to output heatmaps further includes: Performing max - pooling operation on the rotation - equivariant feature maps to obtain rotation - equivariant heatmaps; and Performing average - pooling operation on the rotation - invariant feature maps to obtain rotation - invariant heatmaps; The output heatmaps are used to indicate the predicted target positions, and the database annotation results are the trained data sets in the database used to indicate the correct positions of the targets, and If the comparison results are inconsistent, using the output heatmaps together with the database annotation results for back - training the convolved feature maps.
2. The method according to claim 1, wherein Performing cross - correlation on the feature maps of the template image and the video frame image further includes: Performing cross - correlation on the rotation - equivariant feature maps and rotation - invariant feature maps of the template image and the video frame image respectively.
3. A system for a method of dealing with target rotation based on a rotation-invariant equivariant network, characterized in that, The system includes: An input component for inputting a template image and a video frame image of the target; A convolutional layer component that convolves the template image and the video frame image respectively to obtain their feature maps; A cross - correlation component that performs cross - correlation on the feature maps of the template image and the video frame image; A pooling layer component that pools the results of the cross - correlation respectively to output heatmaps; A training component that compares the output heatmaps with the database annotation results; An output component for determining the position of the target according to the output heatmaps when the results are consistent, and A back - training component for using the output heatmaps for back - training the convolved feature maps when the results are inconsistent; Convolving the template image and the video frame image respectively to obtain their feature maps further includes: Convolving the template image and the video frame image through rotation - equivariant convolution to obtain in - plane rotation - equivariant feature maps; and Convolving the template image and the video frame image through rotation - invariant attention mechanism convolution to obtain out - of - plane rotation - invariant feature maps; Pooling the results of the cross - correlation respectively to output heatmaps further includes: Performing max - pooling operation on the rotation - equivariant feature maps to obtain rotation - equivariant heatmaps; and Performing average - pooling operation on the rotation - invariant feature maps to obtain rotation - invariant heatmaps; The output heatmap is used to indicate the predicted target location, and the database annotation result is the trained data set in the database for indicating the correct location of the target, and If the comparison results are inconsistent, the output heatmap together with the database annotation result is used to reversely train the convolutional feature map.
4. The system according to claim 3, wherein Performing cross-correlation on the feature maps of the template image and the video frame image further includes: Performing cross-correlation on the rotation-equivariant feature map and the rotation-invariant feature map of the template image and the video frame image respectively.
Citation Information
Patent Citations
Target specific response attention target tracking method based on twin network
CN111291679A