Target tracking method based on graph attention fusion network
By adopting a feature fusion method based on graph attention fusion network in target tracking, spatial and channel feature fusion is solved, and the problem of feature fusion in the prior art is only carried out from the spatial dimension, improving the robustness and accuracy of target tracking.
Patent Information
- Application Number
- CN202510715338.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing twin network target tracking algorithm only fuses from the spatial dimension when feature fusion, resulting in fewer channels for effective activation of features after template image features and search image features, which affects the robustness of the tracking target's background interference.
The target tracking method based on the graph attention fusion network is adopted, and the fusion feature is carried out through the feature fusion network based on the graph attention network in the target twin network, and the template image features and search image features are integrated to obtain the fusion features.
The twin network's resolution ability for targets to be tracked is improved, and the tracker's robust tracking is achieved under background interference, which improves the accuracy of tracking target recognition and tracking.
Smart Images

Figure CN120235915A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an object tracking method based on a graph attention fusion network. Background Art
[0002] With the development of science and technology, video single-object tracking, as a basic research topic in the field of computer vision, is the basis for tasks such as navigation and guidance, video surveillance, and scene understanding, and can be widely applied in various fields. Among them, for example, the Siamese network object tracking algorithm can be used for object search. However, when performing feature fusion, only spatial dimension fusion is performed, resulting in fewer effectively activated channels in the features after fusing the template image features and the search image features, and weaker resolution ability for object boundaries. This affects the robustness of the tracker when there is interference in the background of the tracked object, and the object tracking effect is poor. Summary of the Invention
[0003] The present disclosure provides an object tracking method based on a graph attention fusion network, which can improve the resolution ability of the Siamese network for the object to be tracked, can achieve robust tracking of the tracker under background interference, improve the accuracy of tracking object recognition, and improve the accuracy of object tracking. The technical solution of the present disclosure is as follows: According to the first aspect of the embodiments of the present disclosure, there is provided an object tracking method based on a graph attention fusion network, including: Using the object backbone network in the object Siamese network to obtain the first template image feature corresponding to the first template image and the search image features of each search image in the single-object tracking data set; Using the feature fusion network based on the graph attention network in the object Siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image, and obtaining a fusion feature; Inputting the fusion feature into the object tracking head network in the object Siamese network to obtain the first tracking result corresponding to the first object to be tracked in the first template image, where the first tracking result includes the first classification result corresponding to the first object to be tracked and the first regression result corresponding to the first object to be tracked.
[0004] According to some embodiments, the step of using the feature fusion network based on the graph attention network in the object Siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image, and obtaining a fusion feature includes: Using the object spatial feature fusion network based on the graph attention network in the object Siamese network to perform fusion processing on the first template image feature and the search image features of each search image, and obtaining a first spatial fusion feature; The target channel feature fusion network based on the graph attention network in the target twin network is used to perform enhancement processing and fusion processing on the first spatial fusion feature, obtain a first channel fusion feature, and use the first channel fusion feature as the fusion feature.
[0005] According to some embodiments, the fusing the first template image feature and the search image features of the respective search images to obtain a first spatial fusion feature includes: Obtain a first node set corresponding to the first template image feature, where each pixel of the first template image feature is a node; Obtain second node sets corresponding to the search image features of the respective search images, where each pixel of the search image features of the respective search images is a node; Construct a bipartite graph according to the first node set and the second node sets, where the bipartite graph is used to indicate the correspondence between the first template image feature and the search image features of the respective search images; According to the bipartite graph, the first feature vector corresponding to the first node in the first node set, and the second feature vector corresponding to the second node in the second node set, obtain the score between the first node and the second node, and use the Softmax function to regularize the score to obtain weight information, where the first node is any node in the first node set, and the second node is any node in the second node set; According to at least one weight information passed from all nodes in the first subgraph of the bipartite graph to the first node, the second feature vectors corresponding to the nodes in the second node set, and the linear transformation matrix, obtain the aggregated feature corresponding to the first node; Fuse the aggregated feature and the first node feature of the first node to obtain a concatenated feature; Perform spatial feature fusion on at least one concatenated feature corresponding to the first node set to obtain a first spatial fusion feature.
[0006] According to some embodiments, the performing enhancement processing and fusion processing on the first spatial fusion feature to obtain a first channel fusion feature includes: Generate a random graph using the small-world model; Encode the channel features corresponding to the first spatial fusion feature to obtain second node features; Decompose the second node features according to the channel dimension information to obtain decomposed node features; Sort the decomposed node features in a preset order to obtain third node features of the random graph; Perform graph attention convolution operation on the third node feature to obtain a fourth node feature; Perform feature extraction processing and assignment processing on the fourth node feature, and perform splicing processing in the order corresponding to the channel dimension to obtain a feature after channel fusion; Perform decoding processing on the feature after channel fusion to obtain a first channel fusion feature.
[0007] According to some embodiments, wherein the initial Siamese network includes an initial backbone network, an initial spatial feature fusion network based on a graph attention network, an initial channel feature fusion network based on a graph attention network, and an initial tracking head network, and the method further includes: Obtain a training data set, wherein the training data set includes at least one training sample image; Adopt the initial backbone network to initially obtain a second search image feature corresponding to the second search image and training sample image features corresponding to each training sample image in the training data set; Adopt the initial spatial feature fusion network based on the graph attention network to perform fusion processing on the second search image feature and the training sample image to obtain a first spatial fusion feature; Adopt the initial channel feature fusion network based on the graph attention network to enhance and fuse the first spatial fusion feature to obtain a second channel fusion feature; Input the second channel fusion feature into the initial tracking head network to obtain a second tracking result corresponding to a second target to be tracked in the second search image, wherein the second tracking result includes a second classification result and a second regression result corresponding to the second target to be tracked; Obtain a first error between the second classification result and a third classification result and a second error between the second regression result and a third regression result, wherein the third classification result is a classification result obtained by annotating the second target to be tracked, and the third regression result is a regression result obtained by annotating the second target to be tracked; Adjust the model parameters of the initial Siamese network according to the first error and the second error until the initial Siamese network meets the network training requirements to obtain the target Siamese network.
[0008] According to some embodiments, the obtaining the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result includes: Adopt a binary cross-entropy loss function to obtain the first error between the second classification result and the third classification result; Adopt an intersection over union loss function to obtain the second error between the second regression result and the third regression result.
[0009] According to some embodiments, adjusting the model parameters of the initial siamese network according to the first error and the second error until the initial siamese network meets the network training requirements to obtain the target siamese network includes: Obtaining a total error according to a first weight corresponding to the first error, the first error, a second weight corresponding to the second error, and the second error; Using the stochastic gradient descent method to backpropagate the total error and adjusting the model parameters of the initial siamese network until the initial siamese network meets the network training requirements to obtain the target siamese network.
[0010] According to a second aspect of the embodiments of the present disclosure, there is provided an object tracking device based on a graph attention fusion network, including: An image acquisition unit, configured to use a target backbone network in the target siamese network to obtain a first template image feature corresponding to a first template image and search image features of each search image in a single-object tracking data set; A feature fusion unit, configured to use a feature fusion network based on a graph attention network in the target siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature; An object tracking unit, configured to input the fusion feature into a target tracking head network in the target siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked.
[0011] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the object tracking method based on a graph attention fusion network according to any one of the foregoing aspects.
[0012] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the object tracking method based on a graph attention fusion network according to any one of the foregoing aspects.
[0013] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method described in any one of the foregoing aspects.
[0014] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: In some or related embodiments, by using the target backbone network in the target Siamese network, the first template image feature corresponding to the first template image and the search image features of each search image in the single-target tracking data set are obtained; the feature fusion network based on the graph attention network in the target Siamese network is used to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature; the fusion feature is input into the target tracking head network in the target Siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked. Therefore, the topological relationship of the channel dimension can be increased, channel feature fusion can be added on the basis of spatial feature fusion, and the template image feature and the search image feature can be effectively enhanced and fused in the channel dimension, which can reduce the situation that only spatial dimension feature fusion results in fewer effectively activated channels of the fused features of the template image feature and the search image feature and weaker resolution ability for object boundaries, improve the resolution ability of the Siamese network for the target to be tracked, achieve robust tracking of the tracker under background interference, improve the accuracy of tracking target recognition, and improve the accuracy of target tracking.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0017] Figure 1 is a flowchart of a first target tracking method based on a graph attention fusion network provided by an embodiment of the present disclosure; Figure 2 is a flowchart of a second target tracking method based on a graph attention fusion network provided by an embodiment of the present disclosure; Figure 3 is a structural diagram of a Siamese network based on a graph attention fusion network provided by an embodiment of the present disclosure; Figure 4It is a flowchart of a spatial feature fusion network based on graph attention provided by an embodiment of the present disclosure; Figure 5 It is a flowchart of a channel feature fusion network based on graph attention provided by an embodiment of the present disclosure; Figure 6 It is a comparison example schematic diagram of a second object tracking method based on a graph attention fusion network provided by an embodiment of the present disclosure; Figure 7 It is an example schematic diagram of an electronic device shown according to an exemplary embodiment; Figure 8 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0018] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0019] The embodiments of the present disclosure propose an object tracking method, device, electronic device, and storage medium based on a graph attention fusion network. In some embodiments, the object tracking method based on a graph attention fusion network can be interchanged with terms such as an information processing method and a communication method, the object tracking device based on a graph attention fusion network can be interchanged with terms such as an information processing device and a communication device, and terms such as an information processing system and a communication system can be interchanged.
[0020] The embodiments of the present disclosure are not exhaustive, but only for illustration of some embodiments, and do not constitute a specific limitation on the protection scope of the present disclosure. Without contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be arbitrarily exchanged. In addition, the optional implementation manners in an embodiment can be combined arbitrarily; furthermore, the embodiments can be combined arbitrarily. For example, some or all of the steps of different embodiments can be combined arbitrarily, and an embodiment can be arbitrarily combined with the optional implementation manners of other embodiments.
[0021] In each embodiment of the present disclosure, if there is no special explanation and logical conflict, the terms and / or descriptions between the embodiments are consistent and can be cited from each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0022] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended as a limitation on the present disclosure.
[0023] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular form, such as "a", "an", "the", "above-mentioned", "said", "aforementioned", "this", etc., may mean "one and only one", or may also mean "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English translation, the noun following the article can be understood as a singular form of expression or a plural form of expression.
[0024] In the embodiments of the present disclosure, "a plurality of" means two or more.
[0025] In some embodiments, terms such as "at least one of", "one or more", "a plurality of", "multiple", etc. may be used interchangeably.
[0026] Prefix words such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different described objects, and do not impose limitations on the position, order, priority, quantity, content, etc. of the described objects. The statements of the described objects refer to the descriptions in the context of the claims or embodiments, and no redundant limitations should be formed due to the use of prefix words. For example, if the described object is "field", the ordinal numbers before "field" in "first field" and "second field" do not limit the position or order between the "fields", and "first" and "second" do not limit whether the "fields" they modify are in the same message, nor the order of "first field" and "second field". Again, if the described object is "level", the ordinal numbers before "level" in "first level" and "second level" do not limit the priority between the "levels". Again, the quantity of the described object is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the quantity of "device" therein can be one or more. In addition, the objects modified by different prefix words can be the same or different. For example, if the described object is "device", "first device" and "second device" can be the same device or different devices, and their types can be the same or different; again, if the described object is "information", "first information" and "second information" can be the same information or different information, and their content can be the same or different.
[0027] In some embodiments, a "terminal" or "terminal device" may be referred to as a "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc.
[0028] In some embodiments, data, information, etc. may be obtained after obtaining the consent of the user.
[0029] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0030] According to some embodiments, the task of single-object video tracking is as follows: for a video sequence, first initialize the tracker according to the position of the target in the initial frame, then extract the target features and establish a target model, adopt a certain tracking strategy in subsequent frames, estimate the position of the target in the current frame based on the target model, and finally update the target model using the current position and continue tracking the next frame.
[0031] Among them, the Siamese network object tracking algorithm uses a neural network to learn more expressive features through training on a dataset, and applies these features to target localization in the search area. The input of the Siamese network for single-object tracking is two images, one is the template image of the target to be tracked, and the other is the image of the search area with a slightly larger size. The output is the classification result of the target category in the predicted search area and the regression result of the bounding box. Therefore, the network needs to learn from the data how to distinguish whether there is a target in the search image. If there is, the target needs to be identified with a bounding box.
[0032] The structure of the Siamese network is divided into three parts: the backbone network, which is generally modified from the network structure that has achieved excellent performance in the image classification task, and is used to extract the template image features and search image features from the template image and the search image; the spatial feature fusion network, which is used to further enhance and fuse the template image features and search image features in the spatial dimension to obtain the fused features; the tracking head network, which can adopt some convolutional layers or multi-layer perceptrons, for example, to obtain the classification result and the regression result.
[0033] Under the condition of using the same backbone network and tracking head network, the tracker based on the graph neural network constructs a bipartite graph between the template image feature pixels and the search image feature pixels, and designs a spatial feature fusion network based on the graph neural network on the basis of this bipartite graph to enhance and fuse the template image features and search image features in the spatial dimension. However, only fusing features in the spatial dimension ignores the correlation of features in the channel dimension, resulting in fewer effectively activated channels in the features after fusing the template image features and search image features, and weaker ability to distinguish object boundaries, thus affecting the robustness of the tracker when there is interference in the background of the tracked target.
[0034] Figure 1 It is the flowchart of the first object tracking method based on the graph attention fusion network provided by the embodiments of the present disclosure. As Figure 1 shown, this object tracking method based on the graph attention fusion network can be used in the scenario of object tracking for background interference in the process of video single-object tracking, including the following steps: In step S11, the target backbone network in the target Siamese network is used to obtain the first template image features corresponding to the first template image and the search image features of each search image in the single-object tracking data set; According to some embodiments, the execution subject of the embodiments of the present disclosure may be, for example, an electronic device. This electronic device does not specifically refer to a certain fixed device. For example, when the device identifier of the electronic device changes, the electronic device may also change accordingly. For example, when the device structure of the electronic device changes, the electronic device may also change accordingly. Among them, the execution subject of the embodiments of the present disclosure may also be a server, which may be, for example, a single server or a server cluster, and the embodiments of the present disclosure do not limit this.
[0035] In some embodiments, a siamese network is a neural network composed of two or more sub-networks with the same structure and shared weights, and is used to measure the similarity between input samples. The target siamese network may be, for example, a network that has been trained and can be used for image tracking. This target siamese network does not specifically refer to a certain fixed network. For example, when the sub-networks in the target siamese network change, the target siamese network may also change accordingly. Among them, the target siamese network may include, for example, a target backbone network, a feature fusion network based on a graph attention network, and a target tracking head network.
[0036] According to some embodiments, the target backbone network may be used to extract features from an input image. Among them, this target backbone network does not specifically refer to a certain fixed network. For example, the target backbone network may correspond to a template image. For example, template images of different sizes or different pixels may correspond to different backbone networks.
[0037] In some embodiments, the first template image may be, for example, an image including a target to be tracked. This first template image does not specifically refer to a certain fixed image. For example, when the image content corresponding to the first template image changes, the first template image may also change accordingly. For example, when the acquisition time point or reception time point corresponding to the first template image changes, the first template image may also change accordingly. Among them, the "first" in the first template image is used to distinguish it from other template images.
[0038] According to some embodiments, the first template image feature refers to the feature corresponding to the first template image and may be obtained by performing image feature extraction on the first template image. The "first" in the first template image feature is used to distinguish it from other template image features and does not specifically refer to a certain fixed feature. For example, when the feature extraction method corresponding to the first template image feature changes, the first template image feature may also change accordingly.
[0039] According to some embodiments, a single-object tracking data set can be, for example, a collective formed by converging at least one single-object tracking data. The single-object tracking data set does not specifically refer to a certain fixed set. For example, when a certain data in the single-object tracking data set changes, the single-object tracking data set can also change accordingly. For example, when the data volume corresponding to the single-object tracking data set changes, the single-object tracking data set can also change accordingly. Herein, the single-object tracking data set can also be referred to as a single-object search image set, and the embodiments of the present disclosure do not limit this.
[0040] In some embodiments, the search image can be, for example, an image to be determined whether it includes a target to be tracked.
[0041] In some embodiments, use the target backbone network in the target siamese network to obtain the first template image feature corresponding to the first template image and the search image features of each search image in the single-object tracking data set. The order of the first template image feature and the search image features is not limited herein.
[0042] In step S12, use the feature fusion network based on the graph attention network in the target siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image, and obtain a fusion feature; According to some embodiments, the feature fusion network based on the graph attention network can be, for example, a feature fusion network that has been trained and can perform feature fusion. The feature fusion network based on the graph attention network can include, for example, a spatial feature fusion network based on the graph attention network and a channel feature fusion network based on the graph attention network. The feature fusion network based on the graph attention network does not specifically refer to a certain fixed network. For example, when the parameters in the feature fusion network based on the graph attention network change, the feature fusion network based on the graph attention network can also change accordingly.
[0043] In some embodiments, the spatial feature fusion processing can be, for example, the enhancement and fusion processing of the first template image feature and the search image features of each search image in the spatial dimension.
[0044] According to some embodiments, the channel feature fusion processing can be, for example, the enhancement and fusion processing of the first template image feature and the search image features of each search image in the channel dimension.
[0045] In some embodiments, the fusion feature can be, for example, the feature obtained after fusing the first template image feature and the search image features of each search image. The fusion feature does not specifically refer to a certain fixed feature. For example, when the feature fusion network based on the graph attention network changes, the fusion feature can also change accordingly.
[0046] In some embodiments, a feature fusion network based on a graph attention network in the target Siamese network can be used to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image, so as to obtain fused features.
[0047] In step S13, the fused features are input into the target tracking head network in the target Siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked.
[0048] According to some embodiments, the target tracking head network in the target Siamese network is used to obtain a classification result and a regression result. Among them, the first tracking result may include, for example, a first classification result and a first regression result. The "first" in the first classification result is used to distinguish it from other classification results and does not specifically refer to a certain fixed result. For example, when the target Siamese network changes, the first classification result may also change accordingly.
[0049] In some embodiments, the first target to be tracked may be, for example, the target to be tracked in the first template image. The first target to be tracked does not specifically refer to a certain fixed target. The first target to be tracked may be determined according to a target setting instruction. The first target to be tracked does not specifically refer to a certain fixed target.
[0050] In some embodiments, the fused features are input into the target tracking head network in the target Siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked.
[0051] In some or related embodiments, by using the target backbone network in the target Siamese network, the first template image feature corresponding to the first template image and the search image features of each search image in the single-object tracking data set are obtained; the feature fusion network based on the graph attention network in the target Siamese network is used to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain fused features; the fused features are input into the target tracking head network in the target Siamese network to obtain the first tracking result corresponding to the first target to be tracked in the first template image, where the first tracking result includes the first classification result corresponding to the first target to be tracked and the first regression result corresponding to the first target to be tracked. Therefore, the topological relationship in the channel dimension can be increased, channel feature fusion can be added on the basis of spatial feature fusion, the template image feature and the search image feature can be effectively enhanced and fused in the channel dimension, and the situation that only spatial dimension feature fusion results in fewer effectively activated channels of the fused features of the template image feature and the search image feature and weaker resolution ability for object boundaries can be reduced, the resolution ability of the Siamese network for the target to be tracked can be improved, robust tracking of the tracker under background interference can be achieved, the accuracy of tracking target recognition can be improved, and the accuracy of target tracking can be improved.
[0052] Figure 2 is a flowchart of the second object tracking method based on the graph attention fusion network provided by the embodiments of the present disclosure. As Figure 2 shown, the object tracking method based on the graph attention fusion network can be used in an object tracking scenario, and includes the following steps: In step S21, the target backbone network in the target Siamese network is used to obtain the first template image feature corresponding to the first template image and the search image features of each search image in the single-object tracking data set; Among them, the related description is as above, and will not be repeated here.
[0053] According to some embodiments, where the initial Siamese network includes an initial backbone network, an initial spatial feature fusion network based on the graph attention network, an initial channel feature fusion network based on the graph attention network, and an initial tracking head network, the method further includes: Obtain a training data set, where the training data set includes at least one training sample image; Use the initial backbone network to initially obtain the second search image feature corresponding to the second search image and the training sample image features corresponding to each training sample image in the training data set; Use the initial spatial feature fusion network based on the graph attention network to perform fusion processing on the second search image feature and the training sample image to obtain the first spatial fusion feature; The initial channel feature fusion network based on the graph attention network is used to enhance and fuse the first spatial fusion feature to obtain the second channel fusion feature; The second channel fusion feature is input into the initial tracking head network to obtain the second tracking result corresponding to the second target to be tracked in the second search image, where the second tracking result includes the second classification result and the second regression result corresponding to the second target to be tracked; Obtain the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result, where the third classification result is the classification result obtained by annotating the second target to be tracked, and the third regression result is the regression result obtained by annotating the second target to be tracked; Adjust the model parameters of the initial siamese network according to the first error and the second error until the initial siamese network meets the network training requirements, and obtain the target siamese network. Therefore, the parameters can be adjusted according to multiple errors, improving the accuracy of obtaining the target siamese network and the accuracy of target retrieval.
[0054] According to some embodiments, obtaining the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result includes: Use the binary cross-entropy loss function to obtain the first error between the second classification result and the third classification result; Use the intersection over union loss function to obtain the second error between the second regression result and the third regression result. Therefore, different errors can be obtained according to different loss functions, improving the accuracy of error acquisition.
[0055] According to some embodiments, adjusting the model parameters of the initial siamese network according to the first error and the second error until the initial siamese network meets the network training requirements, and obtaining the target siamese network includes: Obtain the total error according to the first weight corresponding to the first error, the first error, the second weight corresponding to the second error, and the second error; Use the stochastic gradient descent method to backpropagate the total error and adjust the model parameters of the initial siamese network until the initial siamese network meets the network training requirements, and obtain the target siamese network.
[0056] According to some embodiments, Figure 3 is a structural diagram of a siamese network based on a graph attention fusion network provided by an embodiment of the present disclosure, as Figure 3As shown in the figure, the Siamese network based on the graph attention fusion network includes a backbone network, a feature fusion network based on the graph attention network, and a tracking head network. Among them, the backbone network is a network structure that can achieve excellent performance in image classification tasks and is used to extract template image features and search image features from the template image and the search image. The feature fusion network based on the graph attention network is used to enhance and fuse the template image features and the search image features to obtain fused features. The tracking head network can use some convolutional layers or multi-layer perceptrons to obtain classification results and regression results. The feature fusion network based on the graph attention network consists of a spatial feature fusion network based on the graph attention network and a channel feature fusion network based on the graph attention network. The spatial feature fusion network based on the graph attention network performs enhancement and fusion of the template image features and the search image features in the spatial dimension, and the channel feature fusion network based on the graph attention network performs enhancement and fusion of the template image features and the search image features in the channel dimension.
[0057] According to some embodiments, such as Figure 5 the Siamese network based on the graph attention fusion network shown in the figure, when performing object tracking, may include: Step 1: Select the MobileNetV2 network that has achieved good performance in image classification tasks as the backbone network of the initial Siamese network in the embodiments of the present disclosure, denoted as , and remove the last fully connected layer of the MobileNetV2 network to meet the requirements of the single-object tracking task. At the same time, in order to make the stride of the backbone network be 8, remove the operations after the convolutional layer of the last downsampling in this network.
[0058] Input the template image and the search image extracted from the training set of the single-object tracking data set into two MobileNetV2 networks with the same structure and shared parameters for feature extraction, and output the template image features and the search image features extracted using MobileNetV2. Its formula can be shown as formula (1) for example: (1) where is a tensor of size , and are the height and width, and is the number of feature channels; is a tensor of size , and are the height and width.
[0059] Step 2: Input the template image features and the search image features into the fusion network based on the graph attention mechanism, and output the relevant operation features after enhancing and fusing the two , and the relevant operation network is denoted as , that is, its formula can be shown as formula (2) for example: (2) where is a tensor with a size of , and are the height and width, is the number of feature channels.
[0060] Step 3: Input the relevant operation features into the tracking head network, denoted as , and output the prediction results of the entire Siamese network for the target to be tracked in the search image, including the classification result and the regression result .
[0061]
[0062] where is a tensor with a size of , is a tensor with a size of , and are the height and width.
[0063] Step 4: Calculate the errors between the classification result and the regression result predicted by the tracking head network and the true classification result and the regression result of the target to be tracked in the search image.
[0064] For the error calculation of the classification result, the binary cross-entropy loss function is used, and for the error calculation of the regression result, the intersection over union loss function is used. The former is denoted as , and the latter is denoted as , then the calculation of the total error can be shown as formula (3) for example: (3) where and are two weight factors, both set to 1 in this disclosure.
[0065] Using the Stochastic Gradient Descent method to backpropagate the calculated error can optimize the parameters in the entire Siamese network.
[0066] Step 5: Continuously extract template images and search images from the training set of the single-object tracking dataset and input them into the Siamese network based on the shunt attention mechanism. Repeat Steps 1 to 4 until the error is less than the threshold of 0.01 set in this disclosure. It is considered that the Siamese network has converged, that is, the target Siamese network is obtained.
[0067] According to some embodiments, the technical solution of the embodiments of this disclosure can run, for example, on a single 3090 graphics card, which can improve the robustness of the single-object tracking method, especially the ability to adapt to interference in the background of the target.
[0068] In step S22, the target space feature fusion network based on the graph attention network in the target Siamese network is used to fuse the first template image feature and the search image features of each search image to obtain the first spatial fusion feature; Among them, the relevant descriptions are as above and will not be elaborated here.
[0069] According to some embodiments, fusing the first template image feature and the search image features of each search image to obtain the first spatial fusion feature includes: Obtain a first node set corresponding to the first template image feature, where each pixel of the first template image feature is a node; Obtain second node sets corresponding to the search image features of each search image, where each pixel of the search image features of each search image is a node; Construct a bipartite graph according to the first node set and the second node sets, where the bipartite graph is used to indicate the correspondence between the first template image feature and the search image features of each search image; According to the bipartite graph, the first feature vector corresponding to the first node in the first node set, and the second feature vector corresponding to the second node in the second node set, obtain the score between the first node and the second node, and use the Softmax function to regularize the score to obtain the weight information, where the first node is any node in the first node set and the second node is any node in the second node set; According to at least one weight information passed from all nodes in the first subgraph of the bipartite graph to the first node, the second feature vectors corresponding to the nodes in the second node set, and the linear transformation matrix, obtain the aggregated feature corresponding to the first node; Fuse the aggregated feature and the first node feature of the first node to obtain the concatenated feature; Perform spatial feature fusion on at least one stitching feature corresponding to the first node set to obtain a first spatial fusion feature.
[0070] According to some embodiments, Figure 4 is a flowchart of a spatial feature fusion network based on graph attention provided by an embodiment of the present disclosure. As Figure 4 shown, input the template image feature and the search image feature into the spatial fusion network based on graph attention (Spatial Fusion Network), denoted as to obtain the enhanced and fused image feature, denoted as . The formula is expressed as formula (4): (4) wherein, specifically, it may include: A. Construct a bipartite graph: Regard each pixel of the template image feature and the search image feature as a node. Let be the set of all nodes on the feature , and be the set of all nodes on the feature .
[0071] Use a complete bipartite graph to construct the correspondence between the template image feature and the search image feature, where: .
[0072] Among them, for example, it can be defined that has two subgraphs , where refers to the empty set; B. Apply the graph attention mechanism: Adopt the graph attention mechanism to enhance and fuse the template image feature and the search image feature.
[0073] For each , let represent the correlation score between node and node . Its formula can be, for example, as shown in formula (5): (5) wherein, are the feature vectors of node and node respectively, represents the number of feature channels, and the score Proportional to the similarity between the features of two nodes.
[0074] Using the inner product between features as a similarity measure, the node features are first linearly transformed, and then the inner product between the transformed feature vectors is taken to calculate the correlation score. The formula can be shown in formula (6), for example: (6) in and is a linear transformation matrix and a learnable parameter.
[0075] According to some embodiments, f is an intermediate variable that can represent h x and h z The degree of matching, The normalized weight can be further calculated based on this matching degree.
[0076] Use the Softmax function to Regularization, its formula can be shown as formula (7): (7) in, Slave Node The perspective of the tracker to the node How much attention to devote.
[0077] Using from All nodes in the The The attention of nodes, computing nodes The aggregate representation of , its formula can be shown as formula (8) for example: (8) in is the linear transformation matrix, is a learnable parameter, is the attention weight.
[0078] Finally, the aggregated features and node features are combined Fusion, to obtain a more powerful feature based on target information, for example, can be shown as formula (9): (9) in Represents a vector concatenation operation.
[0079] For any node calculate , and finally complete the calculation of the graph attention mechanism, and obtain the features after spatial feature fusion using the graph attention mechanism , recorded as the spatial fusion feature.
[0080] In step S23, the target channel feature fusion network based on the graph attention network in the target Siamese network is used to perform enhancement processing and fusion processing on the first spatial fusion feature, obtain the first channel fusion feature, and use the first channel fusion feature as the fusion feature; Among them, the relevant descriptions are as above and will not be elaborated here.
[0081] According to some embodiments, performing enhancement processing and fusion processing on the first spatial fusion feature to obtain the first channel fusion feature includes: Using the small-world model to generate a random graph; Encoding the channel features corresponding to the first spatial fusion feature to obtain the second node features; Decomposing the second node features according to the channel dimension information to obtain the decomposed node features; Sorting the decomposed node features in a preset order to obtain the third node features of the random graph; Performing graph attention convolution operation on the third node features to obtain the fourth node features; Performing feature extraction processing and assignment processing on the fourth node features, and performing splicing processing in the order corresponding to the channel dimension to obtain the features after channel fusion; Performing decoding processing on the features after channel fusion to obtain the first channel fusion feature. The Watts-Strogatz small-world model can be used to generate a random graph, thereby establishing the topological connection relationship between the channels of the features, and using the graph neural network to learn the fusion of the features in the channel dimension. Therefore, by applying the Watts-Strogatz random graph and the graph attention network in the graph neural network, a channel feature fusion module based on the graph attention network can be designed, which together with the spatial feature fusion module based on the graph attention network constitutes a fusion network based on the graph attention network, and performs enhancement and fusion of the template image features and the search image features, so that the single-object tracker can better distinguish the target area and the background area in the search image by learning the topological relationship between the template image pixels and the search image pixels, and adapt to the background interference in the tracking process, such as the low-light background or the background with similar objects, thereby improving the robustness of the tracker.
[0082] According to some embodiments, Figure 5 is a flowchart of a channel feature fusion network based on graph attention provided by an embodiment of the present disclosure. As Figure 5 shown, the spatial fusion feature is input into the channel feature fusion network (Channel Fusion Network) based on graph attention, denoted as , and the further enhanced and fused image features are obtained, denoted as . The formula expression can be as shown in, for example, formula (10): (10) According to some embodiments, the specific implementation of the channel feature fusion network based on graph attention may include: A. Generate a random graph using the Watts-Strogatz small-world model: Set the number of nodes of the random graph to , that is, the number of feature channels. Among them, when generating this random graph, set the average degree of the graph to 2, for example, and set the parameter in the Watts-Strogatz model to 1.
[0083] B. Feature encoding: Encode the channel features to generate Figure 5 node features on First, perform a deformation operation on to transform it into , and then use two learnable linear transformation matrices and for feature encoding. Its formula can be as shown in, for example, formula (11): (11) Among them, is the broadcast product, and is the broadcast sum.
[0084] The parameter matrix , so ; the parameter matrix , so , and it is a four-dimensional tensor.
[0085] Among them, is the channel dimension, and is the dimension of the channel features. Among them, the pre-generated Watts-Strogatz random graph in step A has nodes, and the dimension of the node features is .
[0086] C. Feature decomposition: Decompose along the channel dimension. Its formula can be as shown in, for example, formula (12): (12) The decomposed feature is used as the node feature , the process of assignment is as follows Figure 5 shown, and its formula can be, for example, as shown in formula (13): (13) where is the assignment operation.
[0087] D. Feature transformation: Perform a graph attention convolution operation using the graph attention mechanism to fuse the channel features. The process is as Figure 5 shown, and finally obtain the output of feature transformation. Its formula can be, for example, as shown in formula (14): (14) E. Feature concatenation: Extract the node features on the graph and assign them to the new channel features , and this process is as Figure 5 shown. Its formula can be, for example, as shown in formula (15): (15) where .
[0088] Concatenate in the original order of the channel dimension to obtain the result of fused channel features. Its formula can be, for example, as shown in formula (16): (16) This is a four-dimensional tensor.
[0089] F. Feature decoding: Decode the channel features of to obtain the channel fusion features .
[0090] Use two learnable linear transformation matrices and to perform feature decoding. Its formula can be, for example, as shown in formula (17): (17) where, is the broadcast product operation, is the broadcast summation mechanism.
[0091] The parameter matrix , so ; the parameter matrix , so , and it is a three-dimensional tensor.
[0092] According to some embodiments, the graph attention-based spatial feature fusion process can be performed first, and then the graph attention-based channel feature fusion process can be carried out. The features obtained from the graph attention-based channel feature fusion process are used as the fusion features. Output is performed, and its formula can be, for example, as shown in formula (18): (18) Where, is an assignment operation.
[0093] In step S24, the fusion features are input into the target tracking head network in the target Siamese network to obtain the first tracking result corresponding to the first target to be tracked in the first template image. The first tracking result includes the first classification result corresponding to the first target to be tracked and the first regression result corresponding to the first target to be tracked.
[0094] Where, the related descriptions are as above and will not be elaborated here.
[0095] According to some embodiments, for example, when identifying a bus, Figure 6 shows a comparison schematic diagram of target tracking (before deformation) without using the graph attention fusion network of the present disclosure embodiment and target tracking (after deformation) using the graph attention fusion network of the present disclosure embodiment.
[0096] In some or related embodiments, the target spatial feature fusion network based on the graph attention network in the target Siamese network is used to perform fusion processing on the first template image features and the search image features of each search image to obtain the first spatial fusion features; the target channel feature fusion network based on the graph attention network in the target Siamese network is used to perform enhancement processing and fusion processing on the first spatial fusion features to obtain the first channel fusion features, and the first channel fusion features are used as the fusion features. Therefore, feature fusion in the spatial dimension can be performed on the template image features and the search image features, and then feature fusion in the channel dimension can be carried out, making the network have a stronger ability to distinguish object boundaries and realizing robust tracking of the tracker under background interference.
[0097] The block diagram of a target tracking device based on a graph attention fusion network shown according to an exemplary embodiment. Refer to Figure 7 , the device 700 includes: An image acquisition unit 701, configured to use the target backbone network in the target Siamese network to obtain the first template image features corresponding to the first template image and the search image features of each search image in the single-target tracking data set; The feature fusion unit 702 is configured to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image by using the feature fusion network based on the graph attention network in the target Siamese network, so as to obtain fused features; The target tracking unit 703 is configured to input the fused features into the target tracking head network in the target Siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked.
[0098] According to some embodiments, when the feature fusion unit 702 is configured to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image by using the feature fusion network based on the graph attention network in the target Siamese network to obtain fused features, it is specifically configured to: Perform fusion processing on the first template image feature and the search image features of each search image by using the target spatial feature fusion network based on the graph attention network in the target Siamese network to obtain a first spatial fusion feature; Perform enhancement processing and fusion processing on the first spatial fusion feature by using the target channel feature fusion network based on the graph attention network in the target Siamese network to obtain a first channel fusion feature, and use the first channel fusion feature as the fused feature.
[0099] According to some embodiments, when the feature fusion unit 702 is configured to perform fusion processing on the first template image feature and the search image features of each search image to obtain a first spatial fusion feature, it is specifically configured to: Obtain a first node set corresponding to the first template image feature, where each pixel of the first template image feature is a node; Obtain a second node set corresponding to the search image features of each search image, where each pixel of the search image features of each search image is a node; Construct a bipartite graph according to the first node set and the second node set, where the bipartite graph is used to indicate the correspondence between the first template image feature and the search image features of each search image; Obtain the score between the first node and the second node according to the bipartite graph, the first feature vector corresponding to the first node in the first node set, and the second feature vector corresponding to the second node in the second node set, and regularize the score by using the Softmax function to obtain weight information, where the first node is any node in the first node set, and the second node is any node in the second node set; Obtain the aggregated feature corresponding to the first node according to at least one weight information passed from all nodes in the first subgraph of the bipartite graph, the second feature vectors corresponding to the nodes in the second node set, and the linear transformation matrix; Fuse the aggregated feature and the first node feature of the first node to obtain the concatenated feature; Perform spatial feature fusion on at least one concatenated feature corresponding to the first node set to obtain the first spatial fusion feature.
[0100] According to some embodiments, the feature fusion unit 702, when used for enhancing and fusing the first spatial fusion feature to obtain the first channel fusion feature, specifically is used for: Generate a random graph using the small-world model; Encode the channel features corresponding to the first spatial fusion feature to obtain the second node feature; Decompose the second node feature according to the channel dimension information to obtain the decomposed node feature; Sort the decomposed node features in a preset order to obtain the third node feature of the random graph; Perform graph attention convolution operation on the third node feature to obtain the fourth node feature; Perform feature extraction processing and assignment processing on the fourth node feature, and perform concatenation processing in the order corresponding to the channel dimension to obtain the feature after channel fusion; Perform decoding processing on the feature after channel fusion to obtain the first channel fusion feature.
[0101] According to some embodiments, wherein the initial siamese network includes an initial backbone network, an initial spatial feature fusion network based on the graph attention network, an initial channel feature fusion network based on the graph attention network, and an initial tracking head network, the image acquisition unit 701 is further used for: Obtain a training data set, where the training data set includes at least one training sample image; Use the initial backbone network to initially obtain the second search image feature corresponding to the second search image and the training sample image features corresponding to the training sample images in the training data set; Use the initial spatial feature fusion network based on the graph attention network to fuse the second search image feature and the training sample images to obtain the first spatial fusion feature; Use the initial channel feature fusion network based on the graph attention network to enhance and fuse the first spatial fusion feature to obtain the second channel fusion feature; Input the fused features of the second channel into the initial tracking head network to obtain the second tracking result corresponding to the second target to be tracked in the second search image, where the second tracking result includes the second classification result and the second regression result corresponding to the second target to be tracked; Obtain the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result, where the third classification result is the classification result obtained by annotating the second target to be tracked, and the third regression result is the regression result obtained by annotating the second target to be tracked; Adjust the model parameters of the initial siamese network according to the first error and the second error until the initial siamese network meets the network training requirements, and obtain the target siamese network.
[0102] According to some embodiments, when the image acquisition unit 701 is used to obtain the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result, it is specifically used for: Use the binary cross-entropy loss function to obtain the first error between the second classification result and the third classification result; Use the intersection over union loss function to obtain the second error between the second regression result and the third regression result.
[0103] According to some embodiments, when the image acquisition unit 701 is used to adjust the model parameters of the initial siamese network according to the first error and the second error until the initial siamese network meets the network training requirements and obtain the target siamese network, it is specifically used for: Obtain the total error according to the first weight corresponding to the first error, the first error, the second weight corresponding to the second error, and the second error; Use the stochastic gradient descent method to backpropagate the total error and adjust the model parameters of the initial siamese network until the initial siamese network meets the network training requirements and obtain the target siamese network.
[0104] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0105] In some or related embodiments, an image acquisition unit is configured to use a target backbone network in a target Siamese network to acquire a first template image feature corresponding to a first template image and search image features of each search image in a single-object tracking data set; a feature fusion unit is configured to use a feature fusion network based on a graph attention network in the target Siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature; a target tracking unit is configured to input the fusion feature into a target tracking head network in the target Siamese network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked. Therefore, the topological relationship in the channel dimension can be increased, channel feature fusion can be added on the basis of spatial feature fusion, the template image feature and the search image feature can be effectively enhanced and fused in the channel dimension, and the situation that only spatial dimension feature fusion results in fewer effectively activated channels of the fused features of the template image feature and the search image feature and weaker resolution ability for object boundaries can be reduced, the resolution ability of the Siamese network for the target to be tracked can be improved, robust tracking of the tracker under background interference can be achieved, the accuracy of tracking target recognition can be improved, and the accuracy of target tracking can be improved.
[0106] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device 800 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0107] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0108] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0109] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the above method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the above method in any other suitable manner (e.g., by means of firmware).
[0110] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0112] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0114] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.
[0115] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact operatively through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.
[0116] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0117] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A target tracking method based on a graph attention fusion network, characterized in that Including: Using the target backbone network in the target Siamese network to obtain the first template image feature corresponding to the first template image and the search image features of each search image in the single-object tracking data set; Using the feature fusion network based on the graph attention network in the target Siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature; Inputting the fusion feature into the target tracking head network in the target Siamese network to obtain the first tracking result corresponding to the first target to be tracked in the first template image, where the first tracking result includes the first classification result corresponding to the first target to be tracked and the first regression result corresponding to the first target to be tracked.
2. The method according to claim 1, wherein The step of using the feature fusion network based on the graph attention network in the target Siamese network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature includes: Using the target spatial feature fusion network based on the graph attention network in the target Siamese network to perform fusion processing on the first template image feature and the search image features of each search image to obtain a first spatial fusion feature; Using the target channel feature fusion network based on the graph attention network in the target Siamese network to perform enhancement processing and fusion processing on the first spatial fusion feature to obtain a first channel fusion feature, and using the first channel fusion feature as the fusion feature.
3. The method according to claim 2, wherein The step of performing fusion processing on the first template image feature and the search image features of each search image to obtain a first spatial fusion feature includes: Obtaining a first node set corresponding to the first template image feature, where each pixel of the first template image feature is a node; Obtaining a second node set corresponding to the search image features of each search image, where each pixel of the search image features of each search image is a node; Constructing a bipartite graph according to the first node set and the second node set, where the bipartite graph is used to indicate the corresponding relationship between the first template image feature and the search image features of each search image; According to the bipartite graph, the first feature vector corresponding to the first node in the first node set, and the second feature vector corresponding to the second node in the second node set, obtaining the score between the first node and the second node, and using the Softmax function to regularize the score to obtain weight information, where the first node is any node in the first node set, and the second node is any node in the second node set; According to at least one weight information passed from all nodes in the first subgraph of the bipartite graph to the first node, the second feature vectors corresponding to the nodes in the second node set, and the linear transformation matrix, obtaining the aggregated feature corresponding to the first node; Fusing the aggregated feature and the first node feature of the first node to obtain a concatenated feature; Perform spatial feature fusion on at least one stitching feature corresponding to the first node set to obtain a first spatial fusion feature.
4. The method according to claim 2, characterized in that The performing enhancement processing and fusion processing on the first spatial fusion feature to obtain a first channel fusion feature includes: Generate a random graph using the small-world model; Encode the channel features corresponding to the first spatial fusion feature to obtain second node features; Decompose the second node features according to the channel dimension information to obtain the decomposed node features; Sort the decomposed node features in a preset order to obtain third node features of the random graph; Perform graph attention convolution operation on the third node features to obtain fourth node features; Perform feature extraction processing and assignment processing on the fourth node features, and perform stitching processing in the order corresponding to the channel dimension to obtain the features after channel fusion; Perform decoding processing on the features after channel fusion to obtain a first channel fusion feature.
5. The method according to claim 1, characterized in that, Wherein, The initial Siamese network includes an initial backbone network, an initial spatial feature fusion network based on a graph attention network, an initial channel feature fusion network based on a graph attention network, and an initial tracking head network. The method further includes: Obtain a training data set, wherein the training data set includes at least one training sample image; Use the initial backbone network to initially obtain second search image features corresponding to the second search image and training sample image features corresponding to each training sample image in the training data set; Use the initial spatial feature fusion network based on the graph attention network to perform fusion processing on the second search image features and the training sample images to obtain a first spatial fusion feature; Use the initial channel feature fusion network based on the graph attention network to enhance and fuse the first spatial fusion feature to obtain a second channel fusion feature; Input the second channel fusion feature into the initial tracking head network to obtain a second tracking result corresponding to a second target to be tracked in the second search image, where the second tracking result includes a second classification result and a second regression result corresponding to the second target to be tracked; Obtain a first error between the second classification result and a third classification result and a second error between the second regression result and a third regression result, where the third classification result is the classification result obtained by annotating the second target to be tracked, and the third regression result is the regression result obtained by annotating the second target to be tracked; Adjust the model parameters of the initial Siamese network according to the first error and the second error until the initial Siamese network meets the network training requirements to obtain the target Siamese network.
6. The method according to claim 5, characterized in that The obtaining the first error between the second classification result and the third classification result and the second error between the second regression result and the third regression result includes: Use the binary cross-entropy loss function to obtain the first error between the second classification result and the third classification result; Use the intersection over union loss function to obtain the second error between the second regression result and the third regression result.
7. The method according to claim 5 or 6, characterized in that Adjusting the model parameters of the initial twin network according to the first error and the second error until the initial twin network meets the network training requirements to obtain the target twin network, including: Obtaining a total error according to the first weight corresponding to the first error, the first error, the second weight corresponding to the second error, and the second error; Using the stochastic gradient descent method to backpropagate the total error and adjusting the model parameters of the initial twin network until the initial twin network meets the network training requirements to obtain the target twin network.
8. An object tracking device based on a graph attention fusion network, characterized in that Including: An image acquisition unit, configured to use the target backbone network in the target twin network to obtain the first template image feature corresponding to the first template image and the search image features of each search image in the single target tracking data set; A feature fusion unit, configured to use the feature fusion network based on the graph attention network in the target twin network to perform spatial feature fusion processing and channel feature fusion processing on the first template image feature and the search image features of each search image to obtain a fusion feature; A target tracking unit, configured to input the fusion feature into the target tracking head network in the target twin network to obtain a first tracking result corresponding to a first target to be tracked in the first template image, where the first tracking result includes a first classification result corresponding to the first target to be tracked and a first regression result corresponding to the first target to be tracked.
9. An electronic device, characterized in that, Including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the target tracking method based on the graph attention fusion network according to any one of claims 1 to 7.
10. A storage medium, the storage medium stores instructions, characterized in that, When the instructions are running on an electronic device, the electronic device is caused to execute the target tracking method based on the graph attention fusion network according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cascade multi-mode fusion video target tracking method based on attention model
CN108171141A
Target detection method
CN112232232A
Pedestrian multi-target tracking method based on multivariate difference fusion
CN113221787A
Image attention visual target tracking method
CN113256677A
Twin infrared target tracking method fused with Kalman filtering
CN113344971A