Embedded real-time single target tracking method based on feature clustering twin network
Through the feature extraction, enhancement and fusion method based on feature clustering twin networks, the problem of insufficient accuracy and real-time performance in single-target tracking of videos is solved, and high-precision real-time single-target tracking on embedded platforms is realized.
Patent Information
- Application Number
- CN202510741808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing twin network cannot meet the needs of high precision and real-time in single-target video tracking.
A feature clustering twin network is adopted to extract templates and search image features through lightweight feature extraction networks, feature enhancement and cross-fusion are used for feature clustering, and prediction is performed through tracking head networks, and converging networks are trained to deploy to an embedded platform for real-time single-target tracking.
Real-time high-precision tracking of embedded single targets is realized, improving the accuracy and real-timeness of single target tracking of video.
Smart Images

Figure CN120259367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of electronic information and artificial intelligence, and particularly relates to an embedded real - time single - object tracking method based on a feature - clustering Siamese network. Background Art
[0002] Video single - object tracking, as a basic research topic in the field of computer vision, is the basis for tasks such as navigation and guidance, video surveillance, and scene understanding, and can be widely applied in various fields. In related technologies, video single - object tracking is generally implemented based on a Siamese network. However, the Siamese network can neither meet the requirements of high - precision single - object tracking nor the requirements of single - object real - time tracking. Therefore, there is an urgent need for a more reliable embedded real - time single - object tracking method based on a feature - clustering Siamese network. Summary of the Invention
[0003] The present invention aims to solve at least one of the technical problems in the related technologies to some extent.
[0004] To this end, the first object of the present invention is to propose an embedded real - time single - object tracking method based on a feature - clustering Siamese network, which effectively enhances and fuses the template image features and the search image features through the feature - clustering Siamese network to achieve real - time high - precision tracking of an embedded single object.
[0005] The second object of the present invention is to propose an embedded real - time single - object tracking device based on a feature - clustering Siamese network.
[0006] The third object of the present invention is to propose an electronic device.
[0007] The fourth object of the present invention is to propose a non - transitory computer - readable storage medium storing computer instructions.
[0008] To achieve the above object, the first - aspect embodiment of the present invention proposes an embedded real - time single - object tracking method based on a feature - clustering Siamese network, and the method includes: Select a lightweight feature extraction network to extract features from the template image and the search image respectively to obtain the template image features and the search image features, where the template image and the search image are obtained by cropping each frame of the video sequence of the single object to be tracked; Through a preset correlation operation network for feature clustering, perform feature enhancement on the template image features and the search image features respectively and then perform feature cross - fusion to obtain correlation operation features, where the correlation operation network includes a dual - branch feature self - enhancement module based on feature clustering constructed by a neural network for image feature clustering and enhancement, and a dual - branch feature cross - fusion module based on feature clustering constructed by a neural network for image feature clustering and cross - fusion; Tracking the relevant operation features through a tracking head network based on network tracking to obtain the prediction results of the relevant operation network for the single target to be tracked in the search image, where the prediction results include the classification result of the category of the single target to be tracked and the regression result of the bounding box; Calculating the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to the set threshold, backpropagate the calculated error value to adjust the network parameters in the relevant operation network and the feature extraction network; Until the error value corresponding to the adjusted network parameters is less than the set threshold, obtain the converged relevant operation network and feature extraction network; Construct a feature clustering siamese network by the relevant operation network and the feature extraction network and deploy it to the embedded platform for embedded real-time single target tracking.
[0009] To achieve the above object, an embodiment of the second aspect of the present invention proposes an embedded real-time single target tracking device based on a feature clustering siamese network, the device includes: A feature extraction module, configured to select a lightweight feature extraction network to extract features from the template image and the search image respectively, to obtain the template image features and the search image features, where the template image and the search image are obtained by cropping each frame of the video sequence of the single target to be tracked; A feature enhancement and cross-fusion module, configured to respectively perform feature enhancement on the template image features and the search image features through a preset relevant operation network for feature clustering, and then perform feature cross-fusion to obtain relevant operation features, where the relevant operation network includes a feature clustering-based double-branch feature self-enhancement module constructed by a neural network for image feature clustering and enhancement, and a feature clustering-based double-branch feature cross-fusion module constructed by a neural network for image feature clustering and cross-fusion; A prediction module, configured to track the relevant operation features through a tracking head network based on network tracking to obtain the prediction results of the relevant operation network for the single target to be tracked in the search image, where the prediction results include the classification result of the category of the single target to be tracked and the regression result of the bounding box; A calculation module, configured to calculate the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to the set threshold, backpropagate the calculated error value to adjust the network parameters in the relevant operation network and the feature extraction network; An adjustment module, configured to obtain the converged relevant operation network and feature extraction network until the error value corresponding to the adjusted network parameters is less than the set threshold; A tracking module, which is used to deploy a feature clustering Siamese network composed of a related operation network and a feature extraction network to an embedded platform for embedded real-time single-object tracking.
[0010] To achieve the above object, an embodiment of the third aspect of the present invention provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect.
[0011] To achieve the above object, an embodiment of the fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method described in the first aspect.
[0012] The embedded real-time single-object tracking method, device, electronic device, and storage medium based on a feature clustering Siamese network provided by the embodiments of the present invention extract the template image features and search image features of a video sequence through a lightweight feature extraction network; through a related operation network for feature clustering, the template image features and search image features are respectively enhanced and then cross-fused to obtain related operation features; the classification result of the single object category to be tracked and the regression result of the bounding box obtained by tracking the related operation features through a tracking head network, and a converged related operation network and feature extraction network are trained to form a feature clustering Siamese network deployed to an embedded platform for embedded real-time single-object tracking. Thus, the template image features and search image features are effectively enhanced and fused through the feature clustering Siamese network, realizing real-time high-precision tracking of an embedded single object.
[0013] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0014] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where: Figure 1 It is a schematic flowchart of an embedded real-time single-object tracking method based on a feature clustering Siamese network provided by an embodiment of the present invention; Figure 2 It is a structural diagram of a dual-branch feature self-enhancement module based on feature clustering provided by an embodiment of the present invention; Figure 3 It is a structural diagram of a dual-branch feature cross-fusion module based on feature clustering provided by an embodiment of the present invention; Figure 4 A structure diagram of a Siamese network based on feature clustering provided by an embodiment of the present invention; Figure 5 A schematic structural diagram of an embedded real-time single-object tracking device based on a feature clustering Siamese network provided by an embodiment of the present invention. Detailed implementation manners
[0015] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0016] It should be noted that, in the technical solution of the present invention, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of relevant laws and regulations.
[0017] The following refers to the accompanying drawings to describe an embedded real-time single-object tracking method, device, electronic device, and storage medium based on a feature clustering Siamese network according to an embodiment of the present invention.
[0018] Figure 1 A flowchart of an embedded real-time single-object tracking method based on a feature clustering Siamese network provided by an embodiment of the present invention.
[0019] As Figure 1 shown, the method includes the following steps: Step 101, select a lightweight feature extraction network, and perform feature extraction on the template image and the search image respectively to obtain the template image feature and the search image feature, where the template image and the search image are obtained by cropping each frame of the video sequence of the single object to be tracked.
[0020] In some possible implementation manners, select the convolutional neural network AlexNet, which has achieved good performance in the image classification task, as the lightweight feature extraction network in the present invention, and modify the AlexNet network according to the requirements of the single-object tracking task. Input the template image and the search image extracted from the training set of the single-object tracking dataset into two AlexNet networks with the same structure and shared parameters respectively, and output the template image feature and the search image feature obtained by feature extraction using the AlexNet network.
[0021] Specifically, in the case of selecting the AlexNet network as the feature extraction network in the present invention, it is denoted as , the last fully connected layer of the AlexNet network is removed to meet the requirements of the single-object tracking task. At the same time, in order to make the stride of the feature extraction network be 8, the convolutional layer of the last downsampling in the network and the subsequent operations are removed.
[0022] The template images and search images extracted from the training set of the single-object tracking dataset are respectively input into two AlexNet networks with the same structure and shared parameters for feature extraction, and the template image features and search image features extracted by using the AlexNet network are output, that is:
[0023] Among them, is a tensor with a size of , and are the height and width, and c is the number of feature channels; is a tensor with a size of , and are the height and width.
[0024] Step 102, through a preset correlation operation network for feature clustering, the template image features and the search image features are respectively enhanced and then cross-fused to obtain correlation operation features. Among them, the correlation operation network includes a dual-branch feature self-enhancement module based on feature clustering constructed by a multi-layer neural network for image feature clustering and enhancement, and a dual-branch feature cross-fusion module based on feature clustering constructed by a neural network for image feature clustering and cross-fusion.
[0025] Specifically, the template image features and the search image features are input into the correlation operation network based on feature clustering, and the correlation operation features after enhancing and fusing the two are output. The correlation operation network is denoted as , that is:
[0026] Among them, is a tensor with a size of , and are the height and width, is the number of feature channels.
[0027] In some possible embodiments, through a preset correlation operation network for feature clustering, the template image features and the search image features are respectively enhanced and then cross-fused to obtain correlation operation features, including: inputting the template image features into a dual-branch feature self-enhancement module based on feature clustering to obtain self-enhanced first template image features, and inputting the search image features into the same dual-branch feature self-enhancement module based on feature clustering to obtain self-enhanced first search image features; simultaneously inputting the first template image features and the first search image features into a dual-branch feature cross-fusion module based on feature clustering to obtain second template image features with feature cross-fusion and second search image features with feature cross-fusion; using the second template image features and the second search image features as correlation operation features can effectively enhance and fuse the template image features and the search image features.
[0028] Step 103, track the correlation operation features through a tracking head network based on network tracking to obtain the prediction result of the single target to be tracked in the search image by the correlation operation network, where the prediction result includes the classification result of the category of the single target to be tracked and the regression result of the bounding box.
[0029] In some possible embodiments, the correlation operation features are input into the tracking head network, denoted as , and the prediction result of the single target to be tracked in the entire pair of search images is output, where it includes the classification result and the regression result , .
[0030] Among them, is a tensor with a size of , is a tensor with a size of, and are the height and width.
[0031] Figure 2 This is the structure diagram of a dual-branch feature self-enhancement module based on feature clustering provided by the embodiments of the present invention, as shown in Figure 2As shown, feature extraction is performed on the template image (e.g., a single-person or single-bicycle video frame) and the search image (e.g., a multi-person or multi-vehicle video frame) to obtain the template image features and the search image features. Then, based on the set clustering centers of the template image features, aggregation and divergence operations are performed on the first set of template image feature points (obtained by two-dimensional position encoding of the template image features), and after being processed by a multi-layer perceptron, the self-enhanced first template image features are obtained. Based on the clustering centers of the search image features, aggregation and divergence operations are performed on the first set of search image feature points (obtained by two-dimensional position encoding of the template image features), and after being processed by a multi-layer perceptron, the self-enhanced first search image features are obtained.
[0032] Optionally, input the template image features into a dual-branch feature self-enhancing module based on feature clustering to obtain the self-enhanced first template image features, and input the search image features into the same dual-branch feature self-enhancing module based on feature clustering to obtain the self-enhanced first search image features, including: The dual-branch feature self-enhancing module based on feature clustering splits the template image features into multiple template image feature glyphs (two-dimensional position encoding) by pixel, assigns a two-dimensional coordinate to each template image feature glyph, and regards each template image feature glyph as a point on the two-dimensional plane to form the first set of template image feature points; The dual-branch feature self-enhancing module based on feature clustering splits the search image features into multiple first search image feature glyphs, assigns a two-dimensional coordinate to each first search image feature glyph in the same way as the template image feature glyphs, and regards each first search image feature glyph as a point on the same two-dimensional plane as the template image feature glyphs to form the first set of search image feature points; Select a point in the first set of template image feature points as the template image feature clustering center, perform aggregation and divergence operations on the first set of template image feature points, and after being processed by a multi-layer perceptron, obtain the self-enhanced first template image features; Select a point in the first set of search image feature points as the search image feature clustering center, perform aggregation and divergence operations on the first set of search image feature points, and after being processed by a multi-layer perceptron, obtain the self-enhanced first search image features.
[0033] Specifically, input the template image features into the dual-branch feature self-enhancing module (Self-Enhancing Module) based on feature clustering, denoted as , to obtain the first template image features, denoted as ; input the search image features into the same dual-branch feature self-enhancing module based on feature clustering , to obtain the first search image features, denoted as , the superscript k indicates the k-th dual-branch feature self-enhancement module. The formula is expressed as: .
[0034] Among them, the specific implementation of the dual-branch feature self-enhancement module based on feature clustering is as follows: 1. If the dual-branch feature self-enhancement module based on feature clustering is the first module, that is, k = 1, then for the input template image features and search image features perform two-dimensional position encoding, that is, divide the template image features into the first template image feature glyphs by pixel, and divide the search image features into the first search image feature glyphs:
[0035] Among them, is a tensor with a size of , representing the m-th first template image feature glyph. There are a total of such first template image feature glyphs, and are the height and width of the template image features; is a tensor with a size of , representing the n-th first search image feature glyph. There are a total of such first search image feature glyphs, and are the height and width of the search image features.
[0036] According to the position of each first template image feature glyph in the original template image features, encode to obtain a two-dimensional vector. Suppose the first template image feature glyph is taken from the pixel in the i-th row and j-th column, then the two-dimensional vector assigned to it is:
[0037] Similarly, according to the position of each first search image feature glyph in the original search image features, encode to obtain a two-dimensional vector. Suppose the first search image feature glyph is taken from the pixel in the i-th row and j-th column, then the two-dimensional vector assigned to it is:
[0038] Incorporate the two-dimensional vector as a feature into each first template image feature glyph or first search image feature glyph to obtain new features:
[0039] Among them, is the tensor concatenation operation, is a tensor of size , is also a tensor of size .
[0040] Then, use a fully connected layer to restore the number of new features c + 2 to the number of features c of the original template image:
[0041] Among them, is a fully connected layer with an input dimension of c + 2 and an output dimension of c, is an assignment operation.
[0042] The above process completes the two-dimensional position encoding of the first template image feature glyph and the first search image feature glyph. Regarding the glyphs with two-dimensional position encoding as points on a two-dimensional plane, the first template image feature point set and the first search image feature point set can be obtained and output. These two point sets are and .
[0043] If the dual-branch feature self-enhancement module based on feature clustering is not the first module, that is, , then the input template image feature and the search image feature are divided into the first template image feature glyph and the second search image feature glyph by pixels, and without two-dimensional position encoding, they are directly output as the first template image feature point set and the first search image feature point set.
[0044] 2. Select a point from the first template image feature point set as the template image feature clustering center, and first perform an aggregation operation on the first template image feature point set:
[0045] Among them, is a function, and are learnable scalars, is the m-th point in the template feature point set, and are the height and width of the template image feature. Measures and the template image feature clustering center point The similarity between them is expressed in the present invention through the inner product operation plus the learnable fully connected layer as:
[0046] Then, according to the similarity, the aggregated feature Diverge out:
[0047] Among them, and are learnable scalars.
[0048] Finally, it is processed by a multi-layer perceptron Process:
[0049] Among them, is an assignment operation.
[0050] The newly processed template image feature point set is restored to the template image feature by pixels , and it is used as the first template image feature Output.
[0051] 3. Select a point from the first search image feature point set as the clustering center of the search image feature, and first perform an aggregation operation on the search image feature point set:
[0052] Among them, is function, and are learnable scalars, is the nth point in the first search feature point set, and are the height and width of the search image feature. Measure and the similarity between the search image feature clustering center point In the present invention, it is represented by the inner product operation plus a learnable fully connected layer It is expressed as:
[0053] Then, the aggregated feature According to the similarity Diverge out:
[0054] Among them, and are learnable scalars.
[0055] Finally, it is processed by a multi-layer perceptron Process:
[0056] Among them, is an assignment operation.
[0057] The newly processed search image feature point set is restored to the search image features by pixels again , and use it as the first search image feature output.
[0058] Figure 3 This is a structural diagram of a dual-branch feature cross-fusion module based on feature clustering provided by an embodiment of the present invention. As Figure 3 shown, feature extraction is performed on the template image (such as a single-person or single-bicycle video frame) and the search image (such as a multi-person or multi-bicycle video frame) to obtain the template image features and the search image features. Then, clustering operations and divergence operations are performed on the second template image feature point set (obtained by performing two-dimensional position encoding on the first template image features) according to the set template image feature clustering center. After being processed by a multi-layer perceptron, the cross-fused second template image features are obtained. According to the search image feature clustering center, clustering operations and divergence operations are performed on the second search image feature point set (obtained by performing two-dimensional position encoding on the first template image features), and after being processed by a multi-layer perceptron, the cross-fused second search image features are obtained.
[0059] Optionally, input the first template image features and the first search image features into the dual-branch feature cross-fusion module based on feature clustering at the same time to obtain the cross-fused second template image features and the cross-fused second search image features, including: the dual-branch feature cross-fusion module based on feature clustering splits the first template image features into multiple second template image feature glyphs by pixels, forms a set of two-dimensional coordinate points, and constitutes the second template image feature point set; the dual-branch feature cross-fusion module based on feature clustering splits the first search image features into multiple search image feature glyphs by pixels, forms another set of two-dimensional coordinate points, and constitutes the second search image feature point set; select a point in the second search image feature point set as the template image feature clustering center, perform clustering operations and divergence operations on the second template image feature point set, and after being processed by a multi-layer perceptron, obtain the cross-fused second template image features; select a point in the second template image feature point set as the search image feature clustering center, perform clustering operations and divergence operations on the second search image feature point set, and after being processed by a multi-layer perceptron, obtain the cross-fused second search image features.
[0060] Specifically, input the first template image features into the dual-branch feature cross-fusion module (Cross-Fusing Module) based on feature clustering, denoted as , to obtain the second template image features, denoted as ; Input the first search image feature into the same dual-branch feature cross-fusion module based on feature clustering , and obtain the second search image feature, denoted as , where the superscript k indicates the k-th dual-branch feature cross-fusion module. The formula is expressed as: .
[0061] Among them, the specific implementation of the dual-branch feature cross-fusion module based on feature clustering is as follows: 1. Divide the input first template image feature into the second template image feature glyphs by pixel, and the division is the same as that of the first search image feature into the second search image feature glyphs:
[0062] Among them, is a tensor with a size of , representing the m-th second template image feature glyph. There are a total of such second template image feature glyphs, and are the height and width of the first template image feature; is a tensor with a size of , representing the n-th second search image feature glyph. There are a total of such second search image feature glyphs, and are the height and width of the first search image feature.
[0063] Regarding the second template image feature glyphs and the second search image feature glyphs as point sets in a two-dimensional plane, the second template image feature point set and the second search image feature point set can be obtained and output. These two point sets are and .
[0064] 2. Select a point in the second template image feature point set as the template image feature clustering center, and first perform an aggregation operation on the second template image feature point set:
[0065] Among them, is function, and are learnable scalars, is the m-th point in the second template feature point set, and are the height and width of the first template image feature. Measure The similarity with the clustering center point of the first template image features In the present invention, through the inner product operation plus a learnable fully connected layer is expressed as:
[0066] Then, the aggregated features According to the similarity are spread out:
[0067] Wherein, and are learnable scalars.
[0068] Finally, it is processed by a multi-layer perceptron :
[0069] Wherein, is an assignment operation.
[0070] The newly obtained second template image feature point set after processing is restored to the first template image features pixel by pixel and used as the second template image features for output.
[0071] 3. Select a point from the second search image feature point set as the search image feature clustering center, and first perform an aggregation operation on the second search image feature point set:
[0072] Wherein, is function, and are learnable scalars, is the nth point in the second search feature point set, and are the height and width of the first search image feature. Measure the similarity between and the search image feature clustering center In the present invention, through the inner product operation plus a learnable fully connected layer
[0073] Then, the aggregated features According to the similarity are spread out:
[0074] Among them, and are learnable scalars.
[0075] Finally, it passes through a multi-layer perceptron for processing:
[0076] Among them, is an assignment operation.
[0077] The newly processed second search image feature point set is reverted to the search image feature by pixels and used as the second search image feature for output.
[0078] In addition, the present invention can output the second search image feature of cross-fusion as the relevant operation feature for output, , is an assignment operation.
[0079] Step 104: Calculate the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to the set threshold, backpropagate the calculated error value to adjust the network parameters in the relevant operation network and the feature extraction network.
[0080] In some possible implementation manners, the classification result and the regression result predicted by the tracking head network and the regression result are used to calculate the error with the true classification result
[0081] The error calculation of the classification result uses the binary cross entropy loss function (Binary Cross Entropy Loss), and the error calculation of the regression result uses the intersection over union loss function (Intersection Over Union Loss, IOULoss). The former is denoted as , and the latter is denoted as , then the calculation of the error value loss is as follows:
[0082] Among them, and are two weight factors, which are both set to 1 in the present invention.
[0083] Use the Stochastic Gradient Descent (SGD) method to perform backpropagation on the calculated error value to optimize the network parameters in the relevant operation network and feature extraction network.
[0084] Step 105. Until the error value corresponding to the adjusted network parameters is less than the set threshold, obtain the converged relevant operation network and feature extraction network.
[0085] Optionally, the set threshold can be 0.01, but not limited to this. When the error value is less than 0.01, it is considered that the relevant operation network and feature extraction network have converged.
[0086] Step 106. Construct a feature clustering Siamese network from the relevant operation network and feature extraction network and deploy it to an embedded platform for embedded real-time single-object tracking.
[0087] Optionally, the feature clustering Siamese network can be deployed to the Horizon development hardware platform to test the real-time performance of the feature clustering Siamese network for single-object tracking on the embedded platform, and high-precision real-time single-object tracking at 30 frames per second can be performed.
[0088] The embedded real-time single-object tracking method based on the feature clustering Siamese network according to the embodiments of the present invention extracts the template image features and search image features of the video sequence through a lightweight feature extraction network; through the relevant operation network of feature clustering, the template image features and search image features are respectively feature-enhanced and then feature-cross-fused to obtain relevant operation features; through the tracking head network, the classification result of the single object category to be tracked and the regression result of the bounding box obtained by tracking the relevant operation features are used to train a converged relevant operation network and feature extraction network, so as to constitute a feature clustering Siamese network deployed to an embedded platform for embedded real-time single-object tracking. Thus, the template image features and search image features are effectively enhanced and fused through the feature clustering Siamese network, realizing real-time high-precision tracking of embedded single objects.
[0089] Figure 4 This is a structure diagram of a Siamese network based on feature clustering provided by an embodiment of the present invention. The structure diagram of the Siamese network based on feature clustering is used to execute the embedded real-time single-object tracking method based on the feature clustering Siamese network. Specifically, through the feature extraction network, the template image and the search image are respectively feature-extracted to obtain the template image features and search image features; through a preset relevant operation network of feature clustering (a multi-layer dual-branch feature self-enhancement module based on feature clustering and a dual-branch feature cross-fusion module based on feature clustering), the template image features and search image features are respectively feature-enhanced ( ), ( ) After feature cross - fusion ( ), ( ) to obtain relevant operation features; then track the relevant operation features through a tracking head network to obtain the prediction results of the relevant operation network for the single target to be tracked in the search image. Among them, the prediction results include the classification results of the category of the single target to be tracked and the regression results of the bounding box. Furthermore, a converged relevant operation network and a feature extraction network are trained and deployed to an embedded platform as a feature - clustering Siamese network for real - time high - precision tracking.
[0090] To implement the above - mentioned embodiments, the present invention also proposes an embedded real - time single - target tracking device based on a feature - clustering Siamese network.
[0091] Figure 5 It is a schematic structural diagram of an embedded real - time single - target tracking device based on a feature - clustering Siamese network provided by an embodiment of the present invention.
[0092] As Figure 5 shown, the embedded real - time single - target tracking device 50 based on a feature - clustering Siamese network includes: a feature extraction module 51, a feature enhancement and cross - fusion module 52, a prediction module 53, a calculation module 54, an adjustment module 55, and a tracking module 56.
[0093] The feature extraction module 51 is used to select a lightweight feature extraction network to extract features from the template image and the search image respectively, to obtain the template image features and the search image features, where the template image and the search image are obtained by cropping each frame of the video sequence of the single target to be tracked; The feature enhancement and cross - fusion module 52 is used to perform feature enhancement on the template image features and the search image features respectively through a preset relevant operation network for feature clustering, and then perform feature cross - fusion to obtain relevant operation features, where the relevant operation network includes a feature - clustering - based double - branch feature self - enhancement module constructed by a neural network for image feature clustering and enhancement, and a feature - clustering - based double - branch feature cross - fusion module constructed by a neural network for image feature clustering and cross - fusion; The prediction module 53 is used to track the relevant operation features through a tracking head network based on network tracking to obtain the prediction results of the relevant operation network for the single target to be tracked in the search image, where the prediction results include the classification results of the category of the single target to be tracked and the regression results of the bounding box; A calculation module 54 is configured to calculate the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to a set threshold, the calculated error value is backpropagated to adjust the network parameters in the correlation operation network and the feature extraction network; An adjustment module 55 is configured to obtain a converged correlation operation network and a feature extraction network until the error value corresponding to the adjusted network parameters is less than the set threshold; A tracking module 56 is configured to deploy the correlation operation network and the feature extraction network to form a feature clustering Siamese network on an embedded platform for embedded real-time single target tracking.
[0094] Further, in a possible implementation manner of the embodiment of the present invention, the feature enhancement and cross-fusion module 52 includes: A feature enhancement unit is configured to input the template image feature into a dual-branch feature self-enhancement module based on feature clustering to obtain a self-enhanced first template image feature, and input the search image feature into the same dual-branch feature self-enhancement module based on feature clustering to obtain a self-enhanced first search image feature; A feature cross-fusion unit is configured to input the first template image feature and the first search image feature into a dual-branch feature cross-fusion module based on feature clustering to obtain a second template image feature with feature cross-fusion and a second search image feature with feature cross-fusion; An output unit is configured to use the second template image feature and the second search image feature as correlation operation features.
[0095] Further, in a possible implementation manner of the embodiment of the present invention, the feature enhancement unit is specifically configured to: The dual-branch feature self-enhancement module based on feature clustering splits the template image feature into multiple first template image feature tokens by pixels, assigns a two-dimensional coordinate to each first template image feature token, and regards each first template image feature token as a point on a two-dimensional plane to form a first template image feature point set; The dual-branch feature self-enhancement module based on feature clustering splits the search image feature into multiple first search image feature tokens by pixels, assigns a two-dimensional coordinate to each first search image feature token in the same way as the first template image feature token, and regards each first search image feature token as a point on the same two-dimensional plane as the first template image feature token to form a first search image feature point set; Select a point in the first template image feature point set as the template image feature clustering center, perform aggregation and divergence operations on the first template image feature point set, and then obtain the self-enhanced first template image feature after being processed by a multi-layer perceptron; Select a point in the first search image feature point set as the search image feature clustering center, perform aggregation and divergence operations on the first search image feature point set, and then obtain the self-enhanced first search image feature after being processed by the multi-layer perceptron.
[0096] Further, in a possible implementation manner of the embodiment of the present invention, the feature cross-fusion unit is specifically used for: The dual-branch feature cross-fusion module based on feature clustering splits the first template image feature into multiple second template image feature glyphs by pixels, forms a set of two-dimensional coordinate points, and composes a second template image feature point set; The dual-branch feature cross-fusion module based on feature clustering splits the first search image feature into multiple search image feature glyphs by pixels, forms another set of two-dimensional coordinate points, and composes a second search image feature point set; Select a point in the second search image feature point set as the template image feature clustering center, perform aggregation and divergence operations on the second template image feature point set, and then obtain the cross-fused second template image feature after being processed by a multi-layer perceptron; Select a point in the second template image feature point set as the search image feature clustering center, perform aggregation and divergence operations on the second search image feature point set, and then obtain the cross-fused second search image feature after being processed by a multi-layer perceptron.
[0097] It should be noted that the foregoing explanation of the method embodiment also applies to the device of this embodiment, and will not be elaborated here.
[0098] The embedded real-time single-object tracking device based on the feature clustering Siamese network in the embodiment of the present invention extracts the template image feature and the search image feature of the video sequence through a lightweight feature extraction network; through the related operation network of feature clustering, the template image feature and the search image feature are respectively feature-enhanced and then feature-cross-fused to obtain related operation features; through the tracking head network, the classification result of the single-object category to be tracked and the regression result of the bounding box obtained by tracking the related operation features are used to train a convergent related operation network and a feature extraction network, so as to form a feature clustering Siamese network deployed to an embedded platform for embedded real-time single-object tracking. Thus, the template image feature and the search image feature are effectively enhanced and fused through the feature clustering Siamese network, realizing real-time high-precision tracking of the embedded single object.
[0099] To implement the above embodiments, the present invention also provides an electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the foregoing method.
[0100] To implement the above embodiments, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the foregoing method.
[0101] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0102] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0103] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0104] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0105] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0106] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0107] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0108] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. An embedded real-time single-object tracking method based on a feature clustering Siamese network, characterized in that The method includes: Select a lightweight feature extraction network to extract features from the template image and the search image respectively, obtaining the template image features and the search image features, where the template image and the search image are obtained by cropping each frame of the video sequence of the single target to be tracked; Through a preset correlation operation network for feature clustering, perform feature enhancement on the template image features and the search image features respectively and then perform feature cross-fusion to obtain correlation operation features, where the correlation operation network includes a dual-branch feature self-enhancement module based on feature clustering constructed by a neural network for performing image feature clustering and enhancement in multiple layers, and a dual-branch feature cross-fusion module based on feature clustering constructed by a neural network for performing image feature clustering and cross-fusion; Select a tracking head network based on network tracking to track the correlation operation features, so as to obtain the prediction result of the single target to be tracked in the search image by the correlation operation network, where the prediction result includes the classification result of the category of the single target to be tracked and the regression result of the bounding box; Calculate the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to the set threshold, backpropagate the calculated error value to adjust the network parameters in the correlation operation network and the feature extraction network; Until the error value corresponding to the adjusted network parameters is less than the set threshold, obtain the converged correlation operation network and feature extraction network; Construct a feature clustering-based Siamese network with the correlation operation network and the feature extraction network and deploy it to an embedded platform for embedded real-time single target tracking.
2. The method according to claim 1, characterized in that, The step of, through a preset correlation operation network for feature clustering, perform feature enhancement on the template image features and the search image features respectively and then perform feature cross-fusion to obtain correlation operation features, where the correlation operation network includes a dual-branch feature self-enhancement module based on feature clustering constructed by a neural network for performing image feature clustering and enhancement in multiple layers, and a dual-branch feature cross-fusion module based on feature clustering constructed by a neural network for performing image feature clustering and cross-fusion, includes: Input the template image features into the dual-branch feature self-enhancement module based on feature clustering to obtain the self-enhanced first template image features, and input the search image features into the same dual-branch feature self-enhancement module based on feature clustering to obtain the self-enhanced first search image features; Input the first template image features and the first search image features into the dual-branch feature cross-fusion module based on feature clustering simultaneously to obtain the second template image features with feature cross-fusion and the second search image features with feature cross-fusion; Use the second template image features and the second search image features as the correlation operation features.
3. The method according to claim 2, wherein The step of inputting the template image features into the dual-branch feature self-enhancement module based on feature clustering to obtain the self-enhanced first template image features, and inputting the search image features into the same dual-branch feature self-enhancement module based on feature clustering to obtain the self-enhanced first search image features, includes: The dual-branch feature self-enhancement module based on feature clustering splits the template image features into multiple first template image feature glyphs by pixel, assigns a two-dimensional coordinate to each first template image feature glyph, and regards each first template image feature glyph as a point on a two-dimensional plane to form a first template image feature point set; The dual-branch feature self-enhancement module based on feature clustering splits the search image features into multiple first search image feature glyphs by pixel, assigns a two-dimensional coordinate to each first search image feature glyph in the same way as the first template image feature glyph, and regards each first search image feature glyph as a point on the same two-dimensional plane as the first template image feature glyph to form a first search image feature point set; Select a point in the first template image feature point set as the template image feature clustering center, perform an aggregation operation and a divergence operation on the first template image feature point set, and then obtain the self-enhanced first template image feature after being processed by a multi-layer perceptron; Select a point in the first search image feature point set as the search image feature clustering center, perform an aggregation operation and a divergence operation on the first search image feature point set, and then obtain the self-enhanced first search image feature after being processed by the multi-layer perceptron.
4. The method according to claim 2, wherein The step of simultaneously inputting the first template image feature and the first search image feature into the dual-branch feature cross-fusion module based on feature clustering to obtain the second template image feature with feature cross-fusion and the second search image feature with feature cross-fusion includes: The dual-branch feature cross-fusion module based on feature clustering splits the first template image feature into multiple second template image feature glyphs by pixel, constitutes a set of two-dimensional coordinate points, and forms a second template image feature point set; The dual-branch feature cross-fusion module based on feature clustering splits the first search image feature into multiple search image feature glyphs by pixel, constitutes another set of two-dimensional coordinate points, and forms a second search image feature point set; Select a point in the second search image feature point set as the template image feature clustering center, perform an aggregation operation and a divergence operation on the second template image feature point set, and then obtain the second template image feature with cross-fusion after being processed by a multi-layer perceptron; Select a point in the second template image feature point set as the search image feature clustering center, perform an aggregation operation and a divergence operation on the second search image feature point set, and then obtain the second search image feature with cross-fusion after being processed by a multi-layer perceptron.
5. An embedded real-time single-object tracking device based on a feature clustering Siamese network, characterized in that, The device includes: A feature extraction module, which is used to select a lightweight feature extraction network to respectively extract features from the template image and the search image to obtain template image features and search image features, where the template image and the search image are obtained by cropping each frame of the video sequence of the single target to be tracked; A feature enhancement and cross - fusion module, which is used to perform feature enhancement on the template image feature and the search image feature respectively through a preset operation network related to feature clustering, and then perform feature cross - fusion to obtain related operation features. Among them, the related operation network includes a dual - branch feature self - enhancement module based on feature clustering constructed by a multi - layer neural network for image feature clustering and enhancement, and a dual - branch feature cross - fusion module based on feature clustering constructed by a neural network for image feature clustering and cross - fusion; A prediction module, which is used to track the related operation features by selecting a tracking head network based on network tracking to obtain the prediction result of the related operation network for the single target to be tracked in the search image. Among them, the prediction result includes the classification result of the single target category to be tracked and the regression result of the bounding box; A calculation module, which is used to calculate the error between the classification result and the regression result and the true classification result and the true regression result of the single target to be tracked in the search image. When the error value is greater than or equal to the set threshold, the calculated error value is back - propagated to adjust the network parameters in the related operation network and the feature extraction network; An adjustment module, which is used to obtain a converged related operation network and feature extraction network until the error value corresponding to the adjusted network parameters is less than the set threshold; A tracking module, which is used to deploy the related operation network and the feature extraction network as a feature - clustering Siamese network to an embedded platform for embedded real - time single - target tracking.
6. The device according to claim 5, characterized in that The feature enhancement and cross - fusion module includes: A feature enhancement unit, which is used to input the template image feature into the dual - branch feature self - enhancement module based on feature clustering to obtain the self - enhanced first template image feature, and input the search image feature into the same dual - branch feature self - enhancement module based on feature clustering to obtain the self - enhanced first search image feature; A feature cross - fusion unit, which is used to input the first template image feature and the first search image feature into the dual - branch feature cross - fusion module based on feature clustering simultaneously to obtain the second template image feature with feature cross - fusion and the second search image feature with feature cross - fusion; An output unit, which is used to use the second template image feature and the second search image feature as related operation features.
7. The device according to claim 6, characterized in that, The feature enhancement unit is specifically used for: The dual - branch feature self - enhancement module based on feature clustering splits the template image feature into multiple first template image feature tokens by pixels, assigns a two - dimensional coordinate to each first template image feature token, and regards each first template image feature token as a point on a two - dimensional plane to form a first template image feature point set; The dual - branch feature self - enhancement module based on feature clustering splits the search image feature into multiple first search image feature tokens by pixels, assigns a two - dimensional coordinate to each first search image feature token in the same way as the first template image feature token, and regards each first search image feature token as a point on the same two - dimensional plane as the first template image feature token to form a first search image feature point set; Select a point in the first template image feature point set as the template image feature clustering center, perform aggregation and divergence operations on the first template image feature point set, and then obtain the self-enhanced first template image feature after being processed by a multi-layer perceptron; Select a point in the first search image feature point set as the search image feature clustering center, perform aggregation and divergence operations on the first search image feature point set, and then obtain the self-enhanced first search image feature after being processed by the multi-layer perceptron.
8. The device according to claim 6, characterized in that, The feature cross-fusion unit is specifically configured to: The dual-branch feature cross-fusion module based on feature clustering splits the first template image feature into multiple second template image feature glyphs by pixels, forms a set of two-dimensional coordinate points, and constitutes a second template image feature point set; The dual-branch feature cross-fusion module based on feature clustering splits the first search image feature into multiple search image feature glyphs by pixels, forms another set of two-dimensional coordinate points, and constitutes a second search image feature point set; Select a point in the second search image feature point set as the template image feature clustering center, perform aggregation and divergence operations on the second template image feature point set, and then obtain the cross-fused second template image feature after being processed by a multi-layer perceptron; Select a point in the second template image feature point set as the search image feature clustering center, perform aggregation and divergence operations on the second search image feature point set, and then obtain the cross-fused second search image feature after being processed by a multi-layer perceptron.
9. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-4.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Embedded twin network real-time tracking method applied to mobile platform
CN113850189A
Three-dimensional point cloud single target tracking method based on regional self-attention mechanism
CN115909010A
Target tracking method based on double attention mechanism
CN116563337A
Twin network single target tracking method and device
CN116934807A
Lightweight unmanned aerial vehicle real-time target tracking method based on twin network and attention mechanism
CN118570254A