Twin visual tracking method based on multi-parallel interactive transformation network
By introducing multiple parallel interactive attention modules into the twin vision tracker, the deep information interaction between the template and the search area is enhanced, and the problem of insufficient distinction capabilities of trackers in the prior art is solved, achieving higher tracking accuracy and robustness.
Patent Information
- Application Number
- CN202510118386.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing twin vision trackers lack distinction capabilities when dealing with target appearance changes, occlusion, background interference, and violent target movement, resulting in a degradation of tracking performance.
A twin visual tracking method based on a multi-parallel interactive transformation network is adopted. By building a multi-parallel interactive transformation network model and training on a preconfigured data training set, a multi-parallel interactive attention module is introduced to enhance the deep information interaction between the template and the search area.
It significantly improves the accuracy and robustness of tracking, enhances the deep information interaction between the template and the search area, and effectively improves the distinction ability of the tracker.
Smart Images

Figure CN120088325A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and specifically refers to a Siamese visual tracking method based on a multi-parallel interaction transformation network. Background Art
[0002] Visual object tracking is one of the most critical components in computer vision. With the continuous development of artificial intelligence and electronic devices, visual tracking plays a crucial role in practical scenarios such as precision guidance, search and rescue missions, and intelligent transportation monitoring. However, factors such as target appearance changes, occlusion, background interference, and the violent movement of the target. These factors will lead to a decline in the performance of the tracker, affecting its robustness and reliability in practical applications.
[0003] In response to the above problems, Siamese trackers have been developed, which have shown excellent performance through powerful convolutional networks and effectively solved the above defects. However, in the existing technology, Siamese trackers use two independent streams with shared weights to describe the feature representations of the template and the search region, but there is no interaction between these two streams. Therefore, these trackers can only achieve the final information association through shallow cross-correlation or correlation filters, which ignores the deep interaction between the template and the search region and may reduce the discrimination ability of the tracker.
[0004] Currently, there is a lack of a good visual tracking method that can effectively improve the discrimination ability of the tracker.
[0005] The information disclosed in this background art section is only intended to increase the understanding of the overall background of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the above defects and provide a Siamese visual tracking method based on a multi-parallel interaction transformation network.
[0007] To solve the above technical problems, the technical solution provided by the present invention is as follows:
[0008] A Siamese visual tracking method based on a multi-parallel interaction transformation network, including: building a multi-parallel interaction transformation network model, the multi-parallel interaction transformation network model being capable of tracking visual objects; training the multi-parallel interaction transformation network model based on a pre-configured data training set; validating the trained multi-parallel interaction transformation network model based on a preset validation method; and deploying the validated multi-parallel interaction transformation network model in a pre-set multi-parallel interaction attention module to enhance the deep information interaction between the template and the search region, thereby effectively coping with changes in the target appearance.
[0009] Optionally, the multi-parallel interaction transformation network model at least includes a framework part and a core part. The framework part at least includes a template area, a transmission branch, and a search area. The core part at least includes a plurality of serially connected multi-parallel interaction attention modules. Each multi-parallel interaction attention module at least includes an intra-region self-attention module, a template interaction block, a search interaction block, and a transmission block.
[0010] Optionally, the template area and the search area share the same network weights.
[0011] Optionally, the calculation process of the intra-region self-attention module is expressed by the following formula:
[0012]
[0013] where and respectively represent the template area and search area features after multi-head self-attention. and respectively represent the template area and search area features after layer normalization and feed-forward network. is the template token. is the search token. f intra is the intra-region self-attention block. LN and FFN respectively represent layer normalization and feed-forward neural network transmission. MHSA is the standard multi-head self-attention function.
[0014] Optionally, the calculation process of MHSA is expressed by the following formula:
[0015]
[0016] where [q intra , k intra , v intra respectively represent the query vector, key vector, and value vector. dk represents the dimension of the key vector. [q intra , k intra , v intra = xU qkv and is the corresponding linear transformation network. U qkv represents the weight matrix. indicates that the matrix belongs to the real number field and the dimension is C × 3C.
[0017] Optionally, training the multi-parallel interaction transformation network model based on the pre-configured data training set includes: initializing all filters, parameters, and weights with random numbers; randomly selecting several template samples and test samples from the configured training data set as the current input; using the template samples as the input to the multi-parallel interaction transformation network and calculating the loss of the network output; updating the network parameters according to the calculated loss of the network output; repeating the above steps within a preset period.
[0018] Optionally, the loss of the network output is calculated by the following formula:
[0019]
[0020] where and λ iou are the weight coefficients of the L 1 loss and the GIoU loss respectively, and B i and are the ground truth bounding box and the predicted bounding box respectively.
[0021] Optionally, validating the trained multi-parallel interaction transformation network model based on a preset validation method includes: extracting the template region feature encoding based on the pre-configured template region using the trained multi-parallel interaction transformation network model; extracting the search region feature encoding based on the pre-selected search region using the trained multi-parallel interaction transformation network model to obtain a response map; obtaining the location corresponding to the maximum value in the response map to obtain the estimated target location in the current frame, thereby performing online target tracking using the trained multi-parallel interaction transformation network model.
[0022] Optionally, the pre-configured template region includes: the first frame in the given sequence of images and the initial rectangular box of the target, and an image patch is cropped with the target as the center and the image size is adjusted as the template region.
[0023] Optionally, the pre-selected search region includes: in the second frame and each subsequent frame, a search region is selected with the position of the target in the previous frame as the center.
[0024] By introducing a multi-parallel interaction attention module, the present invention significantly improves the accuracy and robustness of tracking. While maintaining efficient computation, this method enhances the deep information interaction between the template and the search region, and can effectively improve the discrimination ability of the tracker. Description of the Drawings
[0025] Figure 1 is a flowchart of a siamese visual tracking method based on a multi-parallel interaction transformation network provided by an embodiment of the present invention;
[0026] Figure 2 is the overall network structure diagram provided by the embodiments of the present invention;
[0027] Figure 3 is the structure diagram of the self-attention module within the region provided by the embodiments of the present invention;
[0028] Figure 4 is the test result diagram provided by the embodiments of the present invention;
[0029] Figure 5 is an example flowchart provided by the embodiments of the present invention. Detailed implementation manners
[0030] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.
[0031] As described above, in the existing technology, the Siamese tracker uses two independent streams with shared weights to describe the feature representations of the template and the search region, but there is no interaction between these two streams. Therefore, these trackers can only achieve the final information association through shallow cross-correlation or correlation filters, which ignores the deep interaction between the template and the search region and may reduce the discrimination ability of the tracker.
[0032] In view of this, the present invention proposes a Siamese visual tracking method based on a multi-parallel interaction transformation network to solve the above problems. The present invention is implemented in the following manner:
[0033] Please refer to the attached Figure 1 , Figure 1 is a flowchart of a Siamese visual tracking method based on a multi-parallel interaction transformation network provided by the embodiments of the present invention. As Figure 1 shown, in an embodiment of the present invention, the execution subject may be a controller, and it may include the following steps:
[0034] S100. Build a multi-parallel interaction transformation network model, and the multi-parallel interaction transformation network model can track visual targets.
[0035] S200. Train the multi-parallel interaction transformation network model based on a pre-configured data training set.
[0036] S300. Verify the trained multi-parallel interaction transformation network model based on a preset verification method.
[0037] S400. Deploy the verified multi - parallel interaction transformation network model in a pre - set multi - parallel interaction attention module to enhance the depth information interaction between the template and the search area, so as to effectively cope with the changes in the target appearance.
[0038] Through the above introduction of the multi - parallel interaction attention module, the present invention significantly improves the accuracy and robustness of tracking. While maintaining efficient computing, this method enhances the depth information interaction between the template and the search area, and can effectively improve the discrimination ability of the tracker.
[0039] To further clearly and completely explain the present invention, on the basis of the above - mentioned embodiments, the present invention also provides another preferred embodiment. In another preferred embodiment of the present invention, it includes the following steps:
[0040] Step 1. Build a multi - parallel interaction transformation network model, and the specific process is as follows:
[0041] As Figure 2 shown, this framework consists of three parallel branches: the template area, the transmission branch, and the search area. Each branch is specially designed to enhance the tracking performance by promoting robust interaction and dynamic information integration. Among them, the template branch and the search branch share the same network weights. The core of the multi - parallel interaction transformation network framework is composed of multiple (for example: N m = 8) parallel interaction attention modules connected in series. Each multi - parallel interaction attention module in the network consists of multiple (for example: 5) key components. For example, two intra - region self - attention blocks f intra , one template interaction block f TIB , one search interaction block f SIB , and one transmission block f trans . The input of the i - th multi - parallel interaction attention module is the output of the (i - 1) - th module, that is, the template token search token and transmission token where H t and W t respectively represent the height and width of the template image, where H s and W s respectively represent the height and width of the search image, C represents the number of channels of the feature, N TT and D TT respectively represent the number of transmission tokens and the dimension of each transmission token.
[0042] To describe the basic representations of the template and search tokens, two intra - region self - attention blocks are initially adopted, as Figure 2As shown, these blocks have shared weights, and the standard ViT is applied to capture the global context information within each region. The calculation of the intra-region self-attention block can be expressed as:
[0043]
[0044] where and represent the template region and search region features after multi-head self-attention respectively, and represent the template region and search region features after layer normalization and feed-forward network respectively, is the template token, is the search token, f intra is the intra-region self-attention block, LN and FFN represent layer normalization and feed-forward neural network transfer respectively, and MHSA is the standard multi-head self-attention function. It can be expressed by the following formula:
[0045]
[0046] where, [q intra , k intra , v intra represent the query vector, key vector and value vector respectively, dk represents the dimension of the key vector, [q intra , k intra , v intra = xU qkv and is the corresponding linear transformation network, U qkv represents the weight matrix, means the matrix belongs to the real number field and the dimension is C×3C.
[0047] Step 2: Use the labeled training set to perform end-to-end training on the multi-parallel interaction transformation network model. The specific process is as follows:
[0048] (1) Construct the multi-parallel interaction transformation network model according to Figure 2 and initialize all filters, parameters and weights with random numbers.
[0049] (2) Randomly select N train template samples and test samples from all training sets as the input of the current mini-batch. The test samples are selected from one or more of the test sets in datasets such as GOT-10k, LaSOT, TNL2K, TrackingNet, etc. Among them, the target in the template sample is at the center position of the image patch, while the target in the test sample may be located at any position. The above N trainUse a sample as the input to the multi-parallel interactive transformation network, calculate the loss of the network output according to the following loss function, and then update the network parameters according to the loss:
[0050]
[0051] where and λ iou are the weight coefficients of the L 1 loss and the GIoU loss respectively, and B i and are the ground truth bounding box and the predicted bounding box respectively.
[0052] Repeat step (2) after each mini-batch until the preset number of training times N batch .
[0053] Step 3: Use the network model trained in step 2 for online object tracking. The specific process is as follows:
[0054] Given the first frame I i (i = 1, 2,..., n) in the sequence of images I 1 , and the initial rectangular box B 1 of the target, centered on the target, crop out an image patch that is N t times the size of the target, and resize the image to 128×128×3 as the template region, and use the multi-parallel interactive transformation network to extract the template feature encoding.
[0055] In the second frame and each subsequent frame, centered on the position of the target in the previous frame, select a search region, use the multi-parallel interactive transformation network to extract the search region feature encoding, and pass it through the post-processing module with the template feature encoding to obtain a response map. By finding the position corresponding to the maximum value in the response map, the estimated position of the target in the current frame can be obtained.
[0056] Accordingly, the visual object tracking method based on the multi-parallel interactive transformation network of the present invention makes full use of the global information in the target region and the search region through steps (1) and (2) of the invention for offline training, improves the resolution of the model for the target and the background, and significantly improves the accuracy of visual object tracking through the multi-parallel interactive attention module and the sparse update mechanism.
[0057] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0058] In addition, the terms "system" and "network" in this document are often used interchangeably herein. The term "and / or" in this document is merely a description of the associated relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.
[0059] It should be understood that in the embodiments of the present invention, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0060] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0061] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0062] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.
[0063] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.
[0064] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0065] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by hardware, or by firmware, or by a combination thereof. When implemented in software, the above functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a computer. By way of example but not limitation: the computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection can suitably be a computer-readable medium. For example, if the software is transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technologies such as infrared, radio and microwave from a website, server or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, wireless and microwave are included in the definition of the medium. As used in the present invention, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disk generally magnetically replicates data, while disc optically replicates data with a laser. The above combinations should also be included within the scope of protection of the computer-readable medium.
[0066] In summary, the above description is only a preferred embodiment of the technical solution of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A twin vision tracking method based on multiple parallel interactive transformation networks, characterized in that: include: Building a multi-parallel interactive transformation network model, wherein the multi-parallel interactive transformation network model can track a visual target; Based on a preconfigured data training set, training the multiple parallel interactive transformation network model; Based on a preset verification method, verifying the trained multi-parallel interactive transformation network model; The verified multi-parallel interactive transformation network model is deployed in a pre-set multi-parallel interactive attention module to enhance the deep information interaction between the template and the search area, thereby effectively responding to changes in the target appearance.
2. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 1 is characterized in that: The multi-parallel interactive transformation network model includes at least a framework part and a core part. The framework part includes at least a template area, a transmission branch and a search area. The core part includes at least a number of multi-parallel interactive attention modules connected in series. Each of the multi-parallel interactive attention modules includes at least an intra-region self-attention module, a template interaction block, a search interaction block and a transmission block.
3. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 2 is characterized in that: The template region and the search region share the same network weights.
4. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 2 is characterized in that: The calculation process of the self-attention module in the region is expressed by the following formula: in, and They represent the features of the template area and search area after multi-head self-attention, and Respectively represent the features of the template area and search area after layer normalization and feed-forward network, is the template token, is the search token, f intra is the intra-region self-attention block, LN and FFN represent layer normalization and feed-forward neural network transfer respectively, and MHSA is the standard multi-head self-attention function.
5. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 4 is characterized in that: The calculation process of MHSA is expressed by the following formula: Among them, [q intra , k intra , v intra ] represent the query vector, key vector and value vector respectively, dk represents the dimension of the key vector, [q intra , k intra , v intra ]=xU qkv and is the corresponding linear transformation network, U qkv represents the weight matrix, Denotes that the matrix belongs to the real number domain and has dimension C×3C.
6. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 1 is characterized in that: The training of the multi-parallel interactive transformation network model based on a preconfigured data training set includes: Initialize all filters, parameters and weights with random numbers; Randomly selecting a number of template samples and test samples from the configured training data set as current input; Using the template sample as input to a multi-parallel interactive transformation network, and calculating the loss of the network output; Update network parameters according to the calculated loss of the network output; Repeat the above steps within a preset period.
7. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 6 is characterized in that: The loss of calculating the network output is expressed by the following formula: in, and λ iou are the weight coefficients of L1 loss and GIoU loss, respectively. i and They are the true bounding box and the predicted bounding box respectively.
8. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 1 is characterized in that: Based on a preset verification method, the trained multi-parallel interactive transformation network model is verified, including: Based on the pre-configured template area, the trained multi-parallel interactive transformation network model is used to extract the feature code of the template area; based on the pre-selected search area, the trained multi-parallel interactive transformation network model is used to extract the feature code of the search area to obtain a response map; The unknown corresponding to the maximum value in the corresponding image is obtained to obtain the target position estimated in the current frame, so as to use the trained multi-parallel interactive transformation network model for online target tracking.
9. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 8 is characterized in that: The preconfigured template region includes: the first frame in a given sequence of images, and an initial rectangular frame of the target, with the target as the center, cutting out an image block, and adjusting the image size to serve as the template region.
10. The twin vision tracking method based on multiple parallel interactive transformation networks according to claim 9 is characterized in that: The pre-selected search area includes: in the second frame and each subsequent frame, a search area is selected with the position of the target in the previous frame as the center.
Citation Information
Patent Citations
Visual target tracking method and device
CN114841310A
Layered sparse fusion and Transform combined target tracking method suitable for unmanned aerial vehicle
CN117612034A