Dynamic target tracking method based on LBP features and semantic features
Through a dynamic target tracking method based on LBP features and semantic features, the ResNet-50 network and attention mechanism are used to solve the accuracy and robustness of target tracking in complex scenarios, and the stable tracking and computing efficiency improvement under complex conditions is achieved.
Patent Information
- Application Number
- CN202510913600.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-03
AI Technical Summary
The existing target tracking methods are prone to feature matching failures in complex scenarios and difficult to track accurately when target motion state changes dramatically. The perceived reliability in bad weather and complex lighting environments is reduced, and large search areas lead to increased computational volume and serious background interference.
The dynamic target tracking method based on LBP features and semantic features is adopted. The target features are extracted through the ResNet-50 backbone network combined with bilinear interpolation, the inter-frame offset is calculated and the motion amplification coefficient and uncertainty estimation are introduced, the search area is reduced, the feature quality is judged using attention weights and entropy values, and the target classification and position regression prediction head are used for precise positioning.
The stable and continuous tracking of the target under complex conditions is achieved, reducing the amount of calculation and reducing background interference, and improving the accuracy and robustness of the tracking.
Smart Images

Figure CN120411176B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target tracking, and in particular to a dynamic target tracking method based on LBP features and semantic features. Background Art
[0002] With the increasing development of deep learning and target detection technologies, more and more researchers are focusing on the field of target tracking. Target tracking is a crucial component of autonomous driving technology and has been widely used in intelligent transportation, road video surveillance, low-altitude perception, and other fields. Specifically, target tracking aims to identify and predict targets in different time frames and establish associations between these frames to achieve continuous positioning and tracking of the target. During the target tracking process, because the target is in motion, the size, shape, and posture of the target in the image will change dynamically. In addition, real-world scenes are relatively complex and often suffer from problems such as excessive lighting, poor distinction between the target and the background, complex target motion, and target occlusion. Existing target tracking methods focus on solving these problems.
[0003] However, existing target tracking methods still face many challenges when deployed in practical applications. For example, in complex scenes, targets may occlude each other or undergo drastic deformation, which can easily lead to feature matching failure. At the same time, when the target's motion state changes drastically, it is difficult for the tracking algorithm to achieve accurate and stable tracking. In addition, image quality degradation caused by severe weather such as rain, snow, fog and haze, and feature weakening in complex lighting environments, will significantly reduce the algorithm's perception reliability, resulting in target positioning deviations and even tracking interruption. In order to include all possible positions of the target in the next frame, the scale is set to a large value, making it difficult to effectively and accurately determine and narrow the search area. The larger search area increases the amount of calculation and also introduces too much cluttered background, making it more difficult to identify the target. Summary of the Invention
[0004] Based on this, it is necessary to propose a dynamic target tracking method based on LBP features and semantic features to address the above problems.
[0005] A dynamic target tracking method based on LBP features and semantic features, wherein the target appears in multiple frames of images at different times, wherein each time corresponds to a frame of image, the method comprising:
[0006] Determine a template area and a first search area according to a frame image where the target is located;
[0007] determining a template region texture feature of the template region and a first search region texture feature of the first search region;
[0008] The template region is operated by a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensional semantic feature of the template region; the first search region is operated by a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensional semantic feature of the first search region;
[0009] Determine a template region feature based on the template region texture feature and the template region dimensionality reduction semantic feature; determine a first search region feature based on the first search region texture feature and the first search region dimensionality reduction semantic feature;
[0010] Determining an attention weight according to the template region feature and the first search region feature; concatenating multiple attention weights along the channel dimension to obtain an attention result;
[0011] determining an attention weight peak concentration according to the attention weight, and determining an entropy value according to the attention result;
[0012] If the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, the first search area is re-determined and the first search area features are extracted; otherwise, the interaction features are determined based on the first search area features and the attention results;
[0013] The target classification prediction head operates on the interaction features to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; the position regression prediction head operates on the interaction features to obtain the target bounding box position of the target; a Hanning window is used to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally, the bounding box position with the highest confidence score is selected as the tracking result.
[0014] In one embodiment, determining the template area and the first search area according to a frame image where the target is located includes:
[0015] Determining an initial frame image corresponding to the target, marking the target in the initial frame image to obtain a marked area, and cropping the marked area to obtain a template area;
[0016] Determining the coordinates of a first center point of the second frame image and the coordinates of a second center point of the third frame image, and determining the offset coordinates of the target based on the first center point coordinates and the second center point coordinates; the second frame image and the third frame image are adjacent and both are located before the image to be tracked, the third frame target image is adjacent to the image to be tracked, and the image to be tracked is an image of the target at a certain moment;
[0017] determining first offset coordinates of the target in the third frame image according to the third center point coordinates of the initial frame image and the first center point coordinates of the second frame image, and determining second offset coordinates of the target in the image to be tracked according to the second center point coordinates of the third frame image and the first center point coordinates of the second frame image;
[0018] Determining the predicted center point coordinates of the first search area in the image to be tracked according to the second offset coordinates and the second center point coordinates;
[0019] Determine the center point coordinates of the second search area in the third frame image according to the first offset coordinates and the first center point coordinates;
[0020] Determining a prediction error based on the coordinates of the second center point and the coordinates of the center point of the second search area; and adjusting the side length of the first search area to update the side length of the first search area;
[0021] The image to be tracked is clipped according to the predicted center point coordinates and side lengths to obtain a first search area.
[0022] In one embodiment, determining the template region texture feature of the template region and the first search region texture feature of the first search region comprises:
[0023] Determine multiple LBP feature maps of the template area, and splice the multiple LBP feature maps of the template area to obtain a first spliced feature map; perform a dimensionality reduction operation on the first spliced feature map to obtain a first reduced dimensionality feature map; the first reduced dimensionality feature map corresponds to a template area texture feature representing the template area;
[0024] Determine multiple LBP feature maps of the first search area, splice the multiple LBP feature maps of the first search area to obtain a second spliced feature map; perform a dimensionality reduction operation on the second spliced feature map to obtain a second reduced dimensionality feature map; the second reduced dimensionality feature map corresponds to the first search area texture feature representing the first search area.
[0025] In one embodiment,
[0026] The operation of the template region by using the ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the dimensionality reduction semantic features of the template region includes:
[0027] A first semantic feature map, a second semantic feature map, and a third semantic feature map are obtained by performing multi-layer convolution operations on the template area through a ResNet-50 backbone network; the first semantic feature map, the second semantic feature map, and the third semantic feature map are optimized respectively to obtain a first optimized semantic feature, a second optimized semantic feature, and a third optimized semantic feature; the first optimized semantic feature, the second optimized semantic feature, and the third optimized semantic feature are unified in resolution through bilinear interpolation upsampling, and then spliced to obtain a template area spliced semantic feature, and the template area spliced semantic feature is reduced in dimension to obtain a template area reduced-dimensional semantic feature;
[0028] The first search region is operated by using a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the dimensionality reduction semantic features of the first search region, including:
[0029] Multi-layer convolution operations are performed on the first search area through the ResNet-50 backbone network to obtain the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map; the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map are optimized respectively to obtain the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature; the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature are unified in resolution through bilinear interpolation upsampling and then spliced to obtain the spliced semantic feature of the first search area, and the spliced semantic feature of the first search area is reduced in dimension to obtain the reduced-dimensional semantic feature of the first search area.
[0030] In one embodiment, determining the template region feature according to the template region texture feature and the template region dimensionality reduction semantic feature includes:
[0031] Normalizing the template region texture features and the template region dimensionality reduction semantic features to obtain template region normalized texture features and template region normalized semantic features, respectively; determining a template region gating weight map according to the template region normalized texture features and template region normalized semantic features;
[0032] The template area fusion feature is determined based on the normalized texture feature of the template area at the feature point, the normalized semantic feature of the template area at the feature point, the first weight at the feature point in the template area gated weight map, and the second weight at the feature point in the template area gated weight map. All template area fusion features are combined to obtain the template area feature; the feature point is the feature of the corresponding point of the pixel point in the template area in the normalized texture feature of the template area and the normalized semantic feature of the template area.
[0033] In one embodiment, determining the first search area feature according to the first search area texture feature and the first search area dimensionality reduction semantic feature includes:
[0034] Normalizing the texture features of the first search area and the dimensionality-reduced semantic features of the first search area to obtain normalized texture features of the first search area and normalized semantic features of the first search area, respectively; determining a gating weight map of the first search area based on the normalized texture features of the first search area and the normalized semantic features of the first search area;
[0035] The first search area fusion feature is determined based on the first search area normalized texture feature at the feature point, the first search area normalized semantic feature at the feature point, the third weight at the feature point in the first search area gated weight map, and the fourth weight at the feature point in the first search area gated weight map, and all the first search area fusion features are combined to obtain the first search area feature; the feature point is the feature of the corresponding point of the pixel point in the first search area in the normalized texture feature of the first search area and the normalized semantic feature of the first search area.
[0036] In one embodiment, determining the attention weight according to the template region feature and the first search region feature includes:
[0037] Projecting the template region feature to obtain a query vector, and dividing the query vector into a first number of query vector units;
[0038] Projecting the first search area feature to obtain a key vector and a value vector; dividing the key vector into a first number of key vector units, and dividing the value vector into a first number of value vector units; the first number is the number of channels;
[0039] Performing random orthogonal operations on the query vector unit and the key vector unit to obtain a query vector random orthogonal matrix and a key vector random orthogonal matrix respectively;
[0040] An attention weight is obtained according to the value vector unit, the query vector random orthogonal matrix and the key vector random orthogonal matrix.
[0041] In one embodiment, determining the attention weight peak concentration according to the attention weight and determining the entropy value according to the attention result includes:
[0042] The largest of the plurality of attention weights is the maximum response value, and the average of the plurality of attention weights is the response mean; determining the attention weight peak concentration according to the maximum response value and the response mean;
[0043] The attention result is normalized to obtain a corresponding probability distribution; and an entropy value is determined based on the probability distribution.
[0044] In one embodiment, determining the interaction feature based on the first search area feature and the attention result includes:
[0045] Performing a linear transformation on the first search area feature and performing a residual connection with the attention result to obtain a connection feature;
[0046] Normalizing the connection features to obtain normalized features;
[0047] Performing a nonlinear transformation on the normalized features to obtain nonlinear enhanced features;
[0048] The normalized features and the nonlinear enhanced features are residually connected to obtain residual features, and the residual features are normalized to obtain interactive features.
[0049] A dynamic target tracking system based on LBP features and semantic features, the system comprising:
[0050] A first determining module, configured to determine a template area and a first search area according to a frame image where the target is located;
[0051] a second determining module, configured to determine a template region texture feature of the template region and a first search region texture feature of the first search region;
[0052] An operation module is used to operate the template area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the template area; and to operate the first search area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the first search area;
[0053] A third determining module is configured to determine a template region feature based on the template region texture feature and the template region dimensionality reduction semantic feature; and determine a first search region feature based on the first search region texture feature and the first search region dimensionality reduction semantic feature;
[0054] a splicing module, configured to determine an attention weight based on the template region feature and the first search region feature; and splice multiple attention weights along a channel dimension to obtain an attention result;
[0055] a fourth determining module, configured to determine an attention weight peak concentration according to the attention weight, and determine an entropy value according to the attention result;
[0056] a judgment module, configured to, when the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, re-determine the first search area and extract the first search area features; otherwise, determine the interaction features based on the first search area features and the attention result;
[0057] The fifth determination module is used to operate the interactive features through the target classification prediction head to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; operate the interactive features through the position regression prediction head to obtain the target bounding box position of the target; use the Hanning window to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally select the bounding box position with the highest confidence score as the tracking result.
[0058] The present invention determines a template area and a first search area according to a frame image where the target is located; determines the template area texture features of the template area and the first search area texture features of the first search area; operates the template area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain template area dimensionality reduction semantic features; operates the first search area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the first search area dimensionality reduction semantic features; determines the template area features according to the template area texture features and the template area dimensionality reduction semantic features; determines the first search area features according to the first search area texture features and the first search area dimensionality reduction semantic features; determines the attention weight according to the template area features and the first search area features; and splices multiple attention weights along the channel dimension to obtain the attention structure. result; determine the attention weight peak concentration according to the attention weight, and determine the entropy value according to the attention result; if the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, redetermine the first search area and extract the first search area feature; otherwise, determine the interaction feature according to the first search area feature and the attention result; operate the interaction feature through the target classification prediction head to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; operate the interaction feature through the position regression prediction head to obtain the target bounding box position of the target; use the Hanning window to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; finally select the bounding box position with the highest confidence score as the tracking result.
[0059] The present invention calculates the offset of the target in the inter-frame image and introduces technologies such as motion magnification coefficient and uncertainty estimation to accurately predict and narrow the first search area, thereby reducing the amount of calculation while reducing the interference of background information on the target. The present invention uses ResNet-50 to extract target features, and performs dimensionality reduction processing on them while ensuring the discriminability and robustness of target features, effectively avoiding the problems of excessive calculation and high complexity. By accurately locating and narrowing the first search area where the target is located, as well as robust feature extraction and collaborative interaction, the ability to continuously and stably track the target is enhanced under complex conditions such as difficult target motion fitting, complex lighting conditions, and target occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0061] in:
[0062] Figure 1 FIG. 1 is a diagram illustrating an application environment of a dynamic target tracking method based on LBP features and semantic features in one embodiment;
[0063] Figure 2 Flowchart of a dynamic target tracking method based on LBP features and semantic features in one embodiment;
[0064] Figure 3 1 is a structural block diagram of a dynamic target tracking system based on LBP features and semantic features in one embodiment;
[0065] Figure 4 FIG. 1 is a network structure diagram of a dynamic target tracking method based on LBP features and semantic features in one embodiment;
[0066] Figure 5 FIG. 1 is a structural block diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0068] With the increasing development of deep learning and object detection technologies, a growing number of researchers are focusing on target tracking. Object tracking is a crucial component of autonomous driving technology and has been widely applied in fields such as intelligent transportation, road video surveillance, and low-altitude perception. Specifically, target tracking aims to identify and predict targets in different time frames and establish correlations between these frames to achieve continuous positioning and tracking of the target. During the target tracking process, the target's motion causes dynamic changes in its size, shape, and pose within the image. Furthermore, real-world scenes are relatively complex, often experiencing issues such as excessive lighting, poor target-background distinction, complex target motion, and target occlusion. Existing target tracking methods focus on addressing these issues. However, existing target tracking methods still face numerous challenges in practical deployment. For example, in complex scenes, targets may occlude each other and undergo drastic deformation, which can easily lead to feature matching failures. Furthermore, when the target's motion changes dramatically, tracking algorithms struggle to achieve accurate and stable tracking. Furthermore, image quality degradation caused by inclement weather like rain, snow, fog, and haze, as well as feature weakening in complex lighting environments, significantly reduce the algorithm's perceptual reliability, leading to target positioning errors and even tracking interruption. To encompass all possible locations of the target in the next frame, a large search area is required, making it difficult to accurately determine and narrow the search area. This large search area increases the computational effort and introduces excessive background clutter, making target identification more difficult.
[0069] In order to solve the above technical problems, the present application provides a dynamic target tracking method based on LBP features and semantic features.
[0070] Figure 1 FIG1 is an application environment diagram of a dynamic target tracking method based on LBP features and semantic features in one embodiment. Figure 1, the dynamic target tracking method based on LBP features and semantic features is applied to a dynamic target tracking system based on LBP features and semantic features. The dynamic target tracking system based on LBP features and semantic features includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal. The mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server or a server cluster composed of multiple servers. The terminal 110 is used to determine the template area and the first search area according to a certain frame image where the target is located, and the server 120 is used to determine the template area texture features of the template area and the first search area texture features of the first search area; the template area is operated by the ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the template area dimensionality reduction semantic features; the first search area dimensionality reduction semantic features are obtained by operating the first search area by the ResNet-50 backbone network combined with bilinear interpolation upsampling; the template area features are determined according to the template area texture features and the template area dimensionality reduction semantic features; the template area features are determined according to the first search area texture features and the first search area dimensionality reduction semantic features. The first search area feature is determined by the feature; the attention weight is determined according to the template area feature and the first search area feature; the multiple attention weights are spliced along the channel dimension to obtain the attention result; the attention weight peak concentration is determined according to the attention weight, and the entropy value is determined according to the attention result; if the attention weight peak concentration is less than the first threshold or the entropy value is greater than the second threshold, the first search area is re-determined and the first search area feature is extracted; otherwise, the interaction feature is determined according to the first search area feature and the attention result; the interaction feature is operated on by the target classification prediction head to obtain the category of the target, and the interaction feature is operated on by the position regression prediction head to obtain the position of the target. The present invention calculates the offset of the target in the inter-frame image and introduces technologies such as motion magnification coefficient and uncertainty estimation to accurately predict and reduce the first search area, thereby reducing the amount of calculation and reducing the interference of background information on the target. The present invention extracts target features through ResNet-50, performs dimensionality reduction processing on the target while ensuring the discriminability and robustness of the target features, effectively avoiding the problems of excessive calculation and high complexity. By accurately locating and narrowing the first search area where the target is located, as well as robust feature extraction and collaborative interaction, the ability to continuously and stably track the target is enhanced under complex conditions such as difficult target motion fitting, complex lighting conditions, and target occlusion.
[0071] like Figure 2As shown, in one embodiment, a dynamic target tracking method based on LBP features and semantic features is provided, wherein the target appears in multiple frames of images at different times, where each time corresponds to a frame of image. The method can be applied to both a terminal and a server. This embodiment uses application to a terminal as an example. The dynamic target tracking method based on LBP features and semantic features specifically includes the following steps:
[0072] S10: determining a template area and a first search area according to a frame image where the target is located;
[0073] S20: Determine a template region texture feature of the template region and a first search region texture feature of the first search region;
[0074] S30: operating the template region through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensionality semantic feature of the template region; operating the first search region through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensionality semantic feature of the first search region;
[0075] S40: determining a template region feature according to the template region texture feature and the template region dimensionality reduction semantic feature; determining a first search region feature according to the first search region texture feature and the first search region dimensionality reduction semantic feature;
[0076] S50: determining an attention weight according to the template region feature and the first search region feature; and concatenating multiple attention weights along the channel dimension to obtain an attention result;
[0077] S60: determining an attention weight peak concentration according to the attention weight, and determining an entropy value according to the attention result;
[0078] S70: If the attention weight peak concentration is less than the first threshold or the entropy value is greater than the second threshold, re-determine the first search area and extract the first search area features; otherwise, determine the interaction features based on the first search area features and the attention result;
[0079] S80: operating the interactive features through the target classification prediction head to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; operating the interactive features through the position regression prediction head to obtain the target bounding box position of the target; using the Hanning window to reduce the confidence scores of multiple bounding box positions in the current frame image, the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally selecting the bounding box position with the highest confidence score as the tracking result.
[0080] The present invention calculates the offset of the target in the inter-frame image and introduces technologies such as motion magnification coefficient and uncertainty estimation to accurately predict and narrow the first search area, thereby reducing the amount of calculation while reducing the interference of background information on the target. The present invention uses ResNet-50 to extract target features, and performs dimensionality reduction processing on them while ensuring the discriminability and robustness of target features, effectively avoiding the problems of excessive calculation and high complexity. By accurately locating and narrowing the first search area where the target is located, as well as robust feature extraction and collaborative interaction, the ability to continuously and stably track the target is enhanced under complex conditions such as difficult target motion fitting, complex lighting conditions, and target occlusion.
[0081] In one embodiment, determining the template area and the first search area according to a frame image where the target is located in step S10 includes:
[0082] S101: Determine an initial frame image corresponding to the target, mark the target in the initial frame image to obtain a marked area, and crop the marked area to obtain a template area;
[0083] S102: Determine the coordinates of a first center point of a second frame image and a second center point of a third frame image, and determine the offset coordinates of the target based on the first center point coordinates and the second center point coordinates; the second frame image and the third frame image are adjacent and both are located before the image to be tracked, the third frame target image is adjacent to the image to be tracked, and the image to be tracked is an image of the target at a certain moment;
[0084] S103: determining first offset coordinates of the target in the third frame image according to the third center point coordinates of the initial frame image and the first center point coordinates of the second frame image, and determining second offset coordinates of the target in the image to be tracked according to the second center point coordinates of the third frame image and the first center point coordinates of the second frame image;
[0085] S104: Determine the predicted center point coordinates of the first search area in the image to be tracked according to the second offset coordinates and the second center point coordinates;
[0086] S105: Determine the center point coordinates of the second search area in the third frame image according to the first offset coordinates and the first center point coordinates;
[0087] S106: determining a prediction error according to the coordinates of the second center point and the coordinates of the center point of the second search area; and adjusting the side length of the first search area to update the side length of the first search area;
[0088] S107: Clipping the image to be tracked according to the predicted center point coordinates and side lengths to obtain a first search area.
[0089] Specifically, steps S101-S107 are implemented by the following formula:
[0090] (1)
[0091] (2)
[0092] in, is the coordinate of the third center point of the initial frame image, is the coordinate of the first center point of the second frame image, is the coordinate of the second center point of the third frame image, ( , ) is the first offset coordinate of the target in the third frame image, ( , ) is the second offset coordinate of the target in the image to be tracked.
[0093] Then the motion amplification factor is introduced to compensate for the conservatism of the linear model and make the prediction more accurate. This process is expressed as:
[0094] (3)
[0095] (4)
[0096] in, is the motion amplification factor, is the predicted center point coordinate of the first search area in the image to be tracked, ( , ) are the center coordinates of the second search area in the third frame image.
[0097] Because the complexity of real-world scenarios may lead to certain errors in the offset, uncertainty estimation is used to quantify the prediction error of the offset:
[0098] (5)
[0099] in, is the prediction error, take and The larger one in the direction. , ) are the coordinates of the center point of the second search area in the third frame target image.
[0100] Then, the error sensitivity coefficient is introduced to regulate the prediction error to dynamically adjust the side length of the first search area. This process is expressed as follows:
[0101] (6)
[0102] in, is the side length of the adjusted first search area, is the preset side length of the first search area, Represents the prediction error.
[0103] By predicting the center point coordinates The position of the first search area can be determined, and the size of the first search area can be determined by adjusting the side length of the first search area. The image to be tracked is clipped according to the predicted center point coordinates and side length to obtain the first search area. Finally, in order to prevent the error accumulation from being too large and causing the prediction to deviate seriously from the actual situation, each time The frame will perform the first search area doubling operation, expanding the first search area times to accurately reposition the target and eliminate cumulative errors.
[0104] In one embodiment, determining the template region texture feature of the template region and the first search region texture feature of the first search region in step S20 includes:
[0105] S201: Determine multiple LBP feature maps of the template area, and splice the multiple LBP feature maps of the template area to obtain a first spliced feature map; perform a dimensionality reduction operation on the first spliced feature map to obtain a first reduced dimensionality feature map; the first reduced dimensionality feature map corresponds to a template area texture feature representing the template area.
[0106] Specifically, we first define three sets of radius sets of different sizes: And the number of sampling points The template area contains The first area with radius The second area with radius The third region with a radius of 1000 is obtained; P sampling points are obtained in the first, second, and third regions respectively; then bilinear interpolation is used to obtain the grayscale values of non-integer coordinates of the first, second, and third regions to avoid sampling bias when calculating the sampling point positions. Finally, the sampling point difference of each region is calculated based on the grayscale value and the P sampling points of each region. The sampling point difference is binary-encoded and decimal-converted in turn to generate LBP feature maps of three different scales of the template region: ;right Perform channel splicing to obtain the first splicing feature map of the template area ; Then, in order to make the model lightweight, a 1×1 convolution is used on the first concatenated feature map Perform dimensionality reduction to reduce computational complexity and obtain the first dimensionality reduction feature map ; The first dimensionality reduction feature map corresponds to the template area texture feature representing the template area;
[0107] S202: Determine multiple LBP feature maps of the first search area, and splice the multiple LBP feature maps of the first search area to obtain a second spliced feature map; perform a dimensionality reduction operation on the second spliced feature map to obtain a second reduced dimensionality feature map; the second reduced dimensionality feature map corresponds to the first search area texture feature representing the first search area.
[0108] Specifically, similarly, first define three sets of radius sets of different sizes And the number of sampling points The first search area contains The fourth area with a radius of The fifth area with radius The sixth region with a radius of 0.001 is obtained; P sampling points are obtained in the fourth, fifth, and sixth regions respectively; then bilinear interpolation is used to obtain the grayscale values of non-integer coordinates of the fourth, fifth, and sixth regions to avoid sampling bias when calculating the sampling point positions. Finally, the sampling point difference of each region is calculated based on the grayscale value and the P sampling points of each region. The sampling point difference is binary-encoded and decimal-converted in turn to generate three LBP feature maps of different scales in the first search region: ;right Perform channel splicing to obtain the second splicing feature map of the first search area , Then, in order to make the model lightweight, 1×1 convolution is used on the second concatenated feature map Perform dimensionality reduction to reduce computational complexity and obtain the second dimensionality reduction feature map ; The second dimensionality reduction feature map corresponds to the first search area texture feature representing the first search area.
[0109] Regions with different radii can capture texture features of different granularities. The texture features of the template region and the texture features of the first search region retain global and local texture information.
[0110] In one embodiment,
[0111] The operation of the template region by using the ResNet-50 backbone network combined with bilinear interpolation upsampling in step S30 to obtain the dimensionality reduction semantic features of the template region includes:
[0112] A first semantic feature map, a second semantic feature map, and a third semantic feature map are obtained by performing multi-layer convolution operations on the template area through a ResNet-50 backbone network; the first semantic feature map, the second semantic feature map, and the third semantic feature map are optimized respectively to obtain a first optimized semantic feature, a second optimized semantic feature, and a third optimized semantic feature; the first optimized semantic feature, the second optimized semantic feature, and the third optimized semantic feature are unified in resolution through bilinear interpolation upsampling, and then spliced to obtain a template area spliced semantic feature, and the template area spliced semantic feature is reduced in dimension to obtain a template area reduced-dimensional semantic feature;
[0113] The operation of the first search area by using the ResNet-50 backbone network combined with bilinear interpolation upsampling in step S30 to obtain the dimensionality reduction semantic features of the first search area includes:
[0114] Multi-layer convolution operations are performed on the first search area through the ResNet-50 backbone network to obtain the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map; the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map are optimized respectively to obtain the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature; the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature are unified in resolution through bilinear interpolation upsampling and then spliced to obtain the spliced semantic feature of the first search area, and the spliced semantic feature of the first search area is reduced in dimension to obtain the reduced-dimensional semantic feature of the first search area.
[0115] Specifically, we first use the ResNet-50 backbone network to perform multi-layer convolution operations on the template area to obtain the semantic feature maps of the corresponding layers. Here we select the semantic feature maps of the third, fourth, and fifth layers: ; Perform multi-layer convolution operations on the first search area through the ResNet-50 backbone network to obtain the semantic feature maps of the corresponding layers. Here we select the semantic feature maps of the third, fourth, and fifth layers: Then, the channel weight of each layer feature is adjusted by the following formula to highlight the channel features that are important for the target tracking task and improve the expression ability of semantic information.
[0116] , ∈[3,4,5](7)
[0117] , ∈[3,4,5](8)
[0118] in, Optimized semantic features for the template region, It is the optimized semantic feature of the first search area, and GAP stands for global average pooling, which is used to compress the features into a vector of channel dimension, remove redundant spatial dimensions and retain global semantic information. This represents a one-dimensional convolution with a kernel size of k=3, which is used to capture long-range dependencies between channels and dynamically adjust channel weights. Using one-dimensional convolution instead of a fully connected layer to capture dependencies between channels can effectively reduce the number of parameters and computational complexity. represents the activation function, generates channel attention weights, and constrains the values of channel attention weights between [0, 1] to achieve adaptive weighting of feature channels; It represents element-wise multiplication, that is, the attention weight is multiplied by the original feature element-by-element to enhance important features and suppress unimportant features.
[0119] Optimize the third-layer semantic feature map of the template area through formula (7) , the fourth layer semantic feature map , the fifth layer semantic feature map Get the first optimized semantic feature of the template area , the second optimized semantic feature and the third optimized semantic feature ; Optimize the third-layer semantic feature map of the first search area through formula (8) , the fourth layer semantic feature map , the fifth layer semantic feature map , get the first optimized semantic feature of the first search area , the second optimized semantic feature and the third optimized semantic feature .
[0120] In order to balance the details and semantic information and improve the global expression ability of the features, the first optimized semantic features of the template area are upsampled by bilinear interpolation. , the second optimized semantic feature and the third optimized semantic feature After unifying the resolution, multi-scale splicing is performed according to the channel dimension to obtain the template area splicing semantic features The first optimized semantic features of the first search area are obtained by bilinear interpolation upsampling. , the second optimized semantic feature and the third optimized semantic feature After unifying the resolution, multi-scale splicing is performed according to the channel dimension to obtain the splicing semantic features of the first search area However, the template region splicing semantic features Concatenate semantic features with the first search area The feature dimension is high and there is repeated or inefficient information. Therefore, convolution is used to splice semantic features of the template area. Perform feature dimensionality reduction to obtain the template area dimensionality reduction semantic features , use convolution to concatenate semantic features of the first search area Perform feature dimensionality reduction to obtain the first search area dimensionality reduction semantic features The convolution is a convolution with a kernel of 3×3, a stride of 1, and a padding of 1. This can significantly reduce the amount of subsequent computation while retaining key semantic information through the local receptive field and removing redundant information. Finally, the processed template area is used to reduce the semantic features. and the first search area dimension reduction semantic features Output.
[0121] In one embodiment, determining the template region feature according to the template region texture feature and the template region dimensionality reduction semantic feature in step S40 includes:
[0122] S401: normalizing the template region texture features and the template region dimensionality reduction semantic features to obtain normalized template region texture features and normalized template region semantic features respectively; determining a template region gating weight map according to the normalized template region texture features and the normalized template region semantic features;
[0123] S402: Determine the template region fusion feature based on the normalized texture feature of the template region at the feature point, the normalized semantic feature of the template region at the feature point, the first weight at the feature point in the template region gated weight map, and the second weight at the feature point in the template region gated weight map, and combine all the template region fusion features to obtain the template region feature; the feature point is the feature of the corresponding point of the pixel point in the template region in the normalized texture feature of the template region and the normalized semantic feature of the template region.
[0124] Determining the first search area feature according to the first search area texture feature and the first search area dimensionality reduction semantic feature in step S40 includes:
[0125] S403: Normalizing the texture features of the first search area and the dimensionality-reduced semantic features of the first search area to obtain normalized texture features of the first search area and normalized semantic features of the first search area, respectively; determining a gating weight map of the first search area according to the normalized texture features of the first search area and the normalized semantic features of the first search area;
[0126] S404: Determine the first search area fusion feature according to the first search area normalized texture feature at the feature point, the first search area normalized semantic feature at the feature point, the third weight at the feature point in the first search area gated weight map, and the fourth weight at the feature point in the first search area gated weight map, and combine all the first search area fusion features to obtain the first search area feature; the feature point is the feature of the corresponding point of the pixel point in the first search area in the normalized texture feature of the first search area and the normalized semantic feature of the first search area.
[0127] Specifically, the standardized texture features and semantic features are first concatenated along the channel dimension to generate a joint feature map. Then, a 1×1 convolution is used to compress the channel dimension and extract cross-channel interaction information. The nonlinear activation function ReLU is used on the reduced-dimensional features to enhance feature separability. Then, a 3×3 convolution is used to extract spatial relationships and fuse local spatial context information to generate a preliminary weight distribution. Finally, the Sigmoid activation function is used to constrain the weight values to the [0,1] interval to obtain the final weight map. The generation process of the gated weight map can be expressed as follows:
[0128] (9)
[0129] (10)
[0130] in, is the template region gating weight map, is the gating weight map of the first search area, is a 1×1 convolution, is a nonlinear activation function, is a 3×3 convolution, Normalized texture features of the template area, Normalized semantic features of the template area, Normalize the texture features of the first search area, Normalize the semantic features of the first search area, Indicates channel splicing, is the Sigmoid activation function.
[0131] (11)
[0132] in, is the normalized texture feature of the template area At the feature point ( ), is the normalized semantic feature of the template area At the feature point ( ), is the normalized texture feature of the template area at the feature point ( ), is the normalized semantic feature of the template area At the feature point ( ), It is the template region fusion feature. All template region fusion features are finally combined to obtain the template region feature. .
[0133] (12)
[0134] in, is the normalized texture feature of the first search area At the feature point ( ), is the normalized semantic feature of the first search area At the feature point ( ), is the normalized texture feature of the first search area at the feature point ( ), is the normalized semantic feature of the first search area At the feature point ( ), It is the fusion feature of the first search area. All the fusion features of the first search area are finally combined to obtain the first search area feature. .
[0135] In one embodiment, determining the attention weight according to the template region feature and the first search region feature in step S50 includes:
[0136] S501: Projecting the template region feature to obtain a query vector, and dividing the query vector into a first number of query vector units;
[0137] S502: Project the first search area feature to obtain a key vector and a value vector; divide the key vector into a first number of key vector units, and divide the value vector into a first number of value vector units; the first number is the number of channels;
[0138] S503: Perform random orthogonal operations on the query vector unit and the key vector unit respectively to obtain a query vector random orthogonal matrix and a key vector random orthogonal matrix;
[0139] S504: Obtain an attention weight according to the value vector unit, the query vector random orthogonal matrix and the key vector random orthogonal matrix.
[0140] Specifically, the grouped linear attention mechanism is used to establish a global long-range association between the template region and the first search region. The query vector obtained by projection is: ,in, is the template region feature, the first search region feature Projection to get key vector and value vector: ,in , , The design of projecting template region features into a query vector fully considers computational efficiency, task requirements, and target positioning. The relatively small query region not only reduces computational complexity but also guides attention to key areas in the first search region, rather than blindly matching the template.
[0141] To further reduce the complexity, the query vector , key vector Sum value vector Divide into G groups, the number of channels is G, and obtain query vector units respectively , key vector unit , value vector unit , (1, 2, 3, ...., G). Each group calculates attention independently to avoid cross-group interference and enhance local feature interaction. At the same time, different groups focus on the feature patterns of different channel subspaces, which can improve matching robustness. When calculating attention, a linear kernel function is used instead of the Softmax kernel function to avoid exponential operations and significantly improve computational efficiency. Specifically, a random orthogonal matrix is used To approximate the Softmax kernel function:
[0142] (13)
[0143] (14)
[0144] in, is a random orthogonal matrix of the query vector, is a random orthogonal matrix of key vectors, is an orthogonal weight matrix, is the bias term.
[0145] Attention weight The calculation process is shown as follows:
[0146] (15)
[0147] in, is the attention weight, is the value vector unit, is a hyperparameter that controls the numerical range of the dot product result, ensuring gradient stability and thus optimizing the model training effect. After calculating the attention of each group, the final attention result A is obtained by concatenating them along the channel dimension.
[0148] In one embodiment, for determining the attention weight peak concentration according to the attention weight in step S60, determining the entropy value according to the attention result includes:
[0149] S601: The largest of the plurality of attention weights is a maximum response value, and the average of the plurality of attention weights is a response mean; and the peak concentration of the attention weights is determined according to the maximum response value and the response mean;
[0150] S602: Normalize the attention result to obtain a corresponding probability distribution; and determine an entropy value based on the probability distribution.
[0151] Specifically, the largest of the multiple attention weights is the maximum response value, and the average of the multiple attention weights is the response mean; the calculation process of the attention weight peak concentration is as follows:
[0152] (16)
[0153] in, is the peak concentration of attention weight, is the number of channels, is the maximum response value, is the response mean; It is a very small constant, and its main function is to prevent the denominator from being zero and ensure numerical stability.
[0154] The calculation of entropy value requires normalizing the attention result A to obtain the corresponding probability distribution P, and then calculating the entropy value by the following formula:
[0155] (17)
[0156] Among them, AE is the entropy value and P is the probability distribution.
[0157] If the attention weight peak concentration If the entropy value AE is less than the first threshold or greater than the second threshold, the first search area is judged to be of low confidence, triggering the abnormal recovery mechanism and expanding the first search area to the original times, re-determine the first search area and extract the first search area features; otherwise, the first search area is determined to be high confidence, and according to the first search area features And the attention result A determines the interaction feature F.
[0158] In the above-mentioned step S70, the interaction feature is determined based on the first search area feature and the attention result, such as Figure 4 Shown, including:
[0159] S701: Performing a linear transformation on the first search area feature and performing a residual connection with the attention result to obtain a connection feature;
[0160] S702: performing normalization processing on the connection features to obtain normalized features;
[0161] S703: Performing nonlinear transformation on the normalized features to obtain nonlinear enhanced features;
[0162] S704: Performing a residual connection on the normalized feature and the nonlinear enhancement feature to obtain a residual feature, and performing normalization on the residual feature to obtain an interactive feature.
[0163] The target classification prediction head operates on the interactive features to obtain the category of the target, and the position regression prediction head operates on the interactive features to obtain the position of the target (wherein the target classification prediction head and the position regression prediction head are both existing technologies and are not protected points of this application).
[0164] This application also provides a dynamic target tracking system based on LBP features and semantic features, such as Figure 3 As shown, the system includes:
[0165] A first determining module 10 is configured to determine a template area and a first search area according to a frame image where the target is located;
[0166] A second determining module 20 is configured to determine a template region texture feature of the template region and a first search region texture feature of the first search region;
[0167] An operation module 30 is configured to operate the template region through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the template region; and operate the first search region through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the first search region;
[0168] A third determining module 40 is configured to determine a template region feature based on the template region texture feature and the template region dimensionality reduction semantic feature; and determine a first search region feature based on the first search region texture feature and the first search region dimensionality reduction semantic feature;
[0169] a splicing module 50 for determining an attention weight based on the template region feature and the first search region feature; and splicing multiple attention weights along the channel dimension to obtain an attention result;
[0170] a fourth determining module 60, configured to determine an attention weight peak concentration according to the attention weight, and determine an entropy value according to the attention result;
[0171] a judgment module 70 configured to, when the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, re-determine the first search area and extract the first search area features; otherwise, determine the interaction features based on the first search area features and the attention result;
[0172] The fifth determination module 80 is used to operate the interactive features through the target classification prediction head to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; operate the interactive features through the position regression prediction head to obtain the target bounding box position of the target; use the Hanning window to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally select the bounding box position with the highest confidence score as the tracking result.
[0173] The present invention calculates the offset of the target in the inter-frame image and introduces technologies such as motion magnification coefficient and uncertainty estimation to accurately predict and narrow the first search area, thereby reducing the amount of calculation while reducing the interference of background information on the target. The present invention uses ResNet-50 to extract target features, and performs dimensionality reduction processing on them while ensuring the discriminability and robustness of target features, effectively avoiding the problems of excessive calculation and high complexity. By accurately locating and narrowing the first search area where the target is located, as well as robust feature extraction and collaborative interaction, the ability to continuously and stably track the target is enhanced under complex conditions such as difficult target motion fitting, complex lighting conditions, and target occlusion.
[0174] Figure 5 FIG1 shows an internal structure diagram of a computer device in an embodiment. The computer device can be a terminal or a server. Figure 5As shown, the computer device includes a processor, a memory and a network interface connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement a dynamic target tracking method based on LBP features and semantic features. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may implement a dynamic target tracking method based on LBP features and semantic features. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0175] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0176] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0177] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A dynamic target tracking method based on LBP features and semantic features, wherein the target appears in multiple frames of images at different times, wherein each time corresponds to a frame of image, characterized in that: The method comprises: Determining an initial frame image corresponding to the target, marking the target in the initial frame image to obtain a marked area, and cropping the marked area to obtain a template area; Determining the coordinates of a first center point of the second frame image and the coordinates of a second center point of the third frame image, and determining the offset coordinates of the target based on the first center point coordinates and the second center point coordinates; the second frame image and the third frame image are adjacent and both are located before the image to be tracked, the third frame target image is adjacent to the image to be tracked, and the image to be tracked is an image of the target at a certain moment; determining first offset coordinates of the target in the third frame image according to the third center point coordinates of the initial frame image and the first center point coordinates of the second frame image, and determining second offset coordinates of the target in the image to be tracked according to the second center point coordinates of the third frame image and the first center point coordinates of the second frame image; Determining the predicted center point coordinates of the first search area in the image to be tracked according to the second offset coordinates and the second center point coordinates; Determine the center point coordinates of the second search area in the third frame image according to the first offset coordinates and the first center point coordinates; Determining a prediction error based on the coordinates of the second center point and the coordinates of the center point of the second search area; and adjusting the side length of the first search area to update the side length of the first search area; Clipping the image to be tracked according to the predicted center point coordinates and side length to obtain a first search area; determining a template region texture feature of the template region and a first search region texture feature of the first search region; The template region is operated by a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensional semantic feature of the template region; the first search region is operated by a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a reduced-dimensional semantic feature of the first search region; Determine a template region feature based on the template region texture feature and the template region dimensionality reduction semantic feature; determine a first search region feature based on the first search region texture feature and the first search region dimensionality reduction semantic feature; Determining an attention weight according to the template region feature and the first search region feature; concatenating multiple attention weights along the channel dimension to obtain an attention result; determining an attention weight peak concentration according to the attention weight, and determining an entropy value according to the attention result; If the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, the first search area is re-determined and the first search area features are extracted; otherwise, the interaction features are determined based on the first search area features and the attention results; The target classification prediction head operates on the interaction features to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; the position regression prediction head operates on the interaction features to obtain the target bounding box position of the target; a Hanning window is used to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally, the bounding box position with the highest confidence score is selected as the tracking result.
2. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: The determining of the template region texture feature of the template region and the first search region texture feature of the first search region comprises: Determine multiple LBP feature maps of the template area, and splice the multiple LBP feature maps of the template area to obtain a first spliced feature map; perform a dimensionality reduction operation on the first spliced feature map to obtain a first reduced dimensionality feature map; the first reduced dimensionality feature map corresponds to a template area texture feature representing the template area; Determine multiple LBP feature maps of the first search area, splice the multiple LBP feature maps of the first search area to obtain a second spliced feature map; perform a dimensionality reduction operation on the second spliced feature map to obtain a second reduced dimensionality feature map; the second reduced dimensionality feature map corresponds to the first search area texture feature representing the first search area.
3. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: The operation of the template region by using the ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the dimensionality reduction semantic features of the template region includes: A first semantic feature map, a second semantic feature map, and a third semantic feature map are obtained by performing multi-layer convolution operations on the template area through a ResNet-50 backbone network; the first semantic feature map, the second semantic feature map, and the third semantic feature map are optimized respectively to obtain a first optimized semantic feature, a second optimized semantic feature, and a third optimized semantic feature; the first optimized semantic feature, the second optimized semantic feature, and the third optimized semantic feature are unified in resolution through bilinear interpolation upsampling, and then spliced to obtain a template area spliced semantic feature, and the template area spliced semantic feature is reduced in dimension to obtain a template area reduced-dimensional semantic feature; The first search region is operated by using a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain the dimensionality reduction semantic features of the first search region, including: Multi-layer convolution operations are performed on the first search area through the ResNet-50 backbone network to obtain the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map; the fourth semantic feature map, the fifth semantic feature map and the sixth semantic feature map are optimized respectively to obtain the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature; the fourth optimized semantic feature, the fifth optimized semantic feature and the sixth optimized semantic feature are unified in resolution through bilinear interpolation upsampling and then spliced to obtain the spliced semantic feature of the first search area, and the spliced semantic feature of the first search area is reduced in dimension to obtain the reduced-dimensional semantic feature of the first search area.
4. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: Determining the template region feature according to the template region texture feature and the template region dimensionality reduction semantic feature comprises: Normalizing the template region texture features and the template region dimensionality reduction semantic features to obtain template region normalized texture features and template region normalized semantic features, respectively; determining a template region gating weight map according to the template region normalized texture features and template region normalized semantic features; The template area fusion feature is determined based on the normalized texture feature of the template area at the feature point, the normalized semantic feature of the template area at the feature point, the first weight at the feature point in the template area gated weight map, and the second weight at the feature point in the template area gated weight map. All template area fusion features are combined to obtain the template area feature; the feature point is the feature of the corresponding point of the pixel point in the template area in the normalized texture feature of the template area and the normalized semantic feature of the template area.
5. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: The determining of the first search area feature according to the first search area texture feature and the first search area dimensionality reduction semantic feature comprises: Normalizing the texture features of the first search area and the dimensionality-reduced semantic features of the first search area to obtain normalized texture features of the first search area and normalized semantic features of the first search area, respectively; determining a gating weight map of the first search area based on the normalized texture features of the first search area and the normalized semantic features of the first search area; The first search area fusion feature is determined based on the first search area normalized texture feature at the feature point, the first search area normalized semantic feature at the feature point, the third weight at the feature point in the first search area gated weight map, and the fourth weight at the feature point in the first search area gated weight map, and all the first search area fusion features are combined to obtain the first search area feature; the feature point is the feature of the corresponding point of the pixel point in the first search area in the normalized texture feature of the first search area and the normalized semantic feature of the first search area.
6. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: Determining the attention weight according to the template region feature and the first search region feature includes: Projecting the template region feature to obtain a query vector, and dividing the query vector into a first number of query vector units; Projecting the first search area feature to obtain a key vector and a value vector; dividing the key vector into a first number of key vector units, and dividing the value vector into a first number of value vector units; the first number is the number of channels; Performing random orthogonal operations on the query vector unit and the key vector unit to obtain a query vector random orthogonal matrix and a key vector random orthogonal matrix respectively; An attention weight is obtained according to the value vector unit, the query vector random orthogonal matrix and the key vector random orthogonal matrix.
7. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: Determining the attention weight peak concentration according to the attention weight, and determining the entropy value according to the attention result includes: The largest of the plurality of attention weights is the maximum response value, and the average of the plurality of attention weights is the response mean; determining the attention weight peak concentration according to the maximum response value and the response mean; The attention result is normalized to obtain a corresponding probability distribution; and an entropy value is determined based on the probability distribution.
8. The dynamic target tracking method based on LBP features and semantic features according to claim 1, characterized in that: Determining the interaction feature according to the first search area feature and the attention result includes: Performing a linear transformation on the first search area feature and performing a residual connection with the attention result to obtain a connection feature; Normalizing the connection features to obtain normalized features; Performing a nonlinear transformation on the normalized features to obtain nonlinear enhanced features; The normalized features and the nonlinear enhanced features are residually connected to obtain residual features, and the residual features are normalized to obtain interactive features.
9. A dynamic target tracking system based on LBP features and semantic features, characterized in that: The system comprises: The first determination module is used to determine the initial frame image corresponding to the target, mark the target in the initial frame image to obtain a marked area, and crop the marked area to obtain a template area; determine the first center point coordinates of the second frame image and the second center point coordinates of the third frame image, and determine the offset coordinates of the target according to the first center point coordinates and the second center point coordinates; the second frame image and the third frame image are adjacent and both are located before the image to be tracked, the third frame target image is adjacent to the image to be tracked, and the image to be tracked is a frame image of the target at a certain moment; determine the target in the third frame image according to the third center point coordinates of the initial frame image and the first center point coordinates of the second frame image determining the second offset coordinates of the target in the image to be tracked according to the second center point coordinates of the third frame image and the first center point coordinates of the second frame image; determining the predicted center point coordinates of a first search area in the image to be tracked according to the second offset coordinates and the second center point coordinates; determining the center point coordinates of a second search area in the third frame image according to the first offset coordinates and the first center point coordinates; determining a prediction error according to the second center point coordinates and the center point coordinates of the second search area; and adjusting the side length of the first search area to update the side length of the first search area; and cropping the image to be tracked according to the predicted center point coordinates and the side length to obtain a first search area; a second determining module, configured to determine a template region texture feature of the template region and a first search region texture feature of the first search region; An operation module is used to operate the template area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the template area; and to operate the first search area through a ResNet-50 backbone network combined with bilinear interpolation upsampling to obtain a dimensionality reduction semantic feature of the first search area; A third determining module is configured to determine a template region feature based on the template region texture feature and the template region dimensionality reduction semantic feature; and determine a first search region feature based on the first search region texture feature and the first search region dimensionality reduction semantic feature; a splicing module, configured to determine an attention weight based on the template region feature and the first search region feature; and splice multiple attention weights along a channel dimension to obtain an attention result; a fourth determining module, configured to determine an attention weight peak concentration according to the attention weight, and determine an entropy value according to the attention result; a judgment module, configured to, when the attention weight peak concentration is less than a first threshold or the entropy value is greater than a second threshold, re-determine the first search area and extract the first search area features; otherwise, determine the interaction features based on the first search area features and the attention result; The fifth determination module is used to operate the interactive features through the target classification prediction head to obtain the category of the target and the confidence score corresponding to the target bounding box position of the target; operate the interactive features through the position regression prediction head to obtain the target bounding box position of the target; use the Hanning window to reduce the confidence scores of multiple bounding box positions in the current frame image, and the multiple bounding box positions are other bounding box positions except the bounding box position closest to the target bounding box position in the previous frame image; and finally select the bounding box position with the highest confidence score as the tracking result.
Citation Information
Patent Citations
Apparent enhancement depth target tracking method
CN113592911A
Multi-feature target tracking method fusing LBP and attention based on twin network
CN118628765A