A video target tracking method, device, medium and equipment
By employing a multimodal fusion and occlusion-aware tracking strategy, the problems of insufficient fusion and robustness of infrared and visible light images in video target tracking are solved, achieving high-precision and stable target tracking in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN GANXIN TECH CO LTD
- Filing Date
- 2025-06-11
- Publication Date
- 2026-04-10
AI Technical Summary
Existing infrared and visible light image fusion and video target tracking technologies suffer from problems such as coarse fusion strategies, difficulty in fully aligning differences between modes, and insufficient tracking robustness, resulting in poor target tracking performance in complex environments.
By employing a multimodal fusion and occlusion-aware tracking strategy, infrared and visible light images are acquired simultaneously. After preprocessing, a target tracking model is constructed. By combining a multimodal fusion module, a video tracking module, a feature extraction and alignment unit, an adaptive feature fusion unit, and an occlusion-aware update unit, accurate alignment of thermal information from infrared images with texture details from visible light images and highly robust tracking are achieved.
It significantly improves the identifiability and tracking stability of targets in complex environments, and can maintain accurate and continuous target positioning in scenarios such as low light, occlusion, and high-speed movement. It has strong anti-interference ability and occlusion recovery ability, and achieves target tracking effect with higher robustness and accuracy.
Smart Images

Figure CN120598992B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of computer vision and artificial intelligence, and particularly relates to a video target tracking method, device, medium and equipment. BACKGROUND
[0002] The existing infrared image and visible light image fusion and video target tracking technology still faces many challenges in actual application, mainly in aspects of rough fusion strategy, difficulty in fully aligning inter-modal differences and insufficient tracking robustness. Traditional methods mostly adopt simple image-level or feature-level splicing, which is difficult to take into account the thermal characteristics of infrared images and the detailed texture of visible light images, resulting in problems such as structural blur, thermal weakening or edge distortion in the fused image; in addition, in video tracking, the existing methods have limited modeling capabilities for occlusion, adaptive update and multiple frames, and are prone to target drift, mismatch and even tracking failure in complex dynamic scenes, which seriously restricts their practicality in high-reliability scenarios such as night vision monitoring, intelligent transportation and border security.
[0003] Therefore, an improved method capable of realizing accurate alignment of multi-modal features, adaptive fusion and high-robustness tracking is urgently needed to improve image quality and tracking continuity. SUMMARY
[0004] In view of the deficiencies in the prior art, the purpose of the present application is to provide a video target tracking method, device, medium and equipment, which can realize high precision, high robustness and strong continuity of target tracking in complex environments through the joint design of multi-modal fusion and occlusion perception tracking strategy.
[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0006] A video target tracking method, the method comprising: synchronously acquiring an infrared image and a visible light image containing a target to be measured; pre-processing the infrared image and the visible light image; constructing a target to be measured tracking model and training the model; inputting the pre-processed infrared image and visible light image into the trained target to be measured tracking model to identify and track the position of the target to be measured.
[0007] Optionally, the pre-processing of the infrared image comprises: performing size unification and thermal value normalization processing on the infrared image; performing local contrast enhancement and nonlinear thermal value mapping on the normalized infrared image; constructing a spatial Gaussian attention map based on thermal intensity and guiding weighted smoothing processing.
[0008] Optionally, the visible light image is preprocessed, including: performing size and brightness normalization processing on the visible light image; performing illumination correction and local contrast enhancement on the normalized visible light image; performing edge extraction and detail enhancement on the illumination corrected visible light image; converting the detail enhanced visible light image to Lab space and enhancing a / b channel color contrast.
[0009] Optionally, the to-be-tested target tracking model includes: the multi-modal fusion module, configured to fuse the preprocessed infrared image and visible light image to obtain a fused image; and the video tracking module, configured to identify and track the position of the to-be-tested target based on the fused image.
[0010] Optionally, the multi-modal fusion module includes: a feature extraction and alignment unit, configured to extract multi-scale features of the preprocessed infrared image and visible light image and perform spatial alignment; an adaptive feature fusion unit, configured to adaptively fuse the aligned features to obtain a fused image; and an optimization unit, configured to perform structural optimization and reconstruction on the fused image.
[0011] Optionally, the video tracking module includes: an initialization tracking unit, configured to construct a multi-modal template to provide tracking starting reference information; a lightweight temporal correlation unit, configured to combine previous frame residual information and current frame features to predict target displacement; and an occlusion awareness update unit, configured to detect an occlusion state based on a similarity heat map change and guide a template update strategy to maintain tracking stability.
[0012] The application also provides a video target tracking method and device, the device including: an acquisition unit configured to synchronously acquire an infrared image and a visible light image containing a to-be-tested target; a preprocessing unit configured to preprocess the infrared image and the visible light image; a model construction and training unit configured to construct a to-be-tested target tracking model and train the model; and an identification unit configured to input the preprocessed infrared image and visible light image into the trained to-be-tested target tracking model to identify and track the position of the to-be-tested target.
[0013] The application also provides a storage medium including instructions that, when executed on a computer, cause the computer to perform the method of any one of the preceding.
[0014] The application also provides an electronic device including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method of any one of the preceding when executing the program.
[0015] Compared with the prior art, the application has the following beneficial effects:
[0016] The application fuses the thermal information of an infrared image and the texture details of a visible light image, constructs a multi-modal feature fusion mechanism guided by structural consistency, and combines front frame residual modeling, saliency enhancement, and occlusion perception strategies to significantly improve the recognizability and tracking stability of a target in a complex environment. Compared with traditional methods, the application can maintain accurate and continuous target positioning in low-light, occlusion, high-speed motion, and other scenes, has strong anti-interference and occlusion recovery capabilities, achieves higher robustness and precision target tracking effects, and has good practical value and promotion prospects. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a flowchart of a video target tracking method provided by an embodiment of the application;
[0018] Figure 2 is a schematic diagram of an original infrared image provided by another embodiment of the application;
[0019] Figure 3 is a schematic diagram of an original visible light image provided by another embodiment of the application;
[0020] Figure 4 is a schematic diagram of fusion by a traditional method;
[0021] Figure 5 is a schematic diagram of fusion based on the method described in the application;
[0022] Figure 6 is a schematic diagram of tracking effects of a traditional method;
[0023] Figure 7 is a schematic diagram of tracking effects based on the method described in the application;
[0024] Figure 8 is a schematic diagram of tracking accuracy of a target under high-speed motion by a traditional method;
[0025] Figure 9 is a schematic diagram of tracking accuracy of a target under high-speed motion based on the method described in the application;
[0026] Figure 10 is a schematic diagram of occlusion recovery by a traditional method;
[0027] Figure 11 is a schematic diagram of occlusion recovery based on the method described in the application;
[0028] Figure 12 is a schematic diagram of a video target tracking device provided by another embodiment of the application. DETAILED DESCRIPTION
[0029] The specific embodiments of the present application will be described in detail below with reference to the drawings. Although specific embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided so that the present application can be more thoroughly and completely understood, and so that the scope of the present application can be conveyed to those skilled in the art.
[0030] It should be noted that some terms are used in the specification and claims to refer to certain components. Those skilled in the art will understand that the same component can be referred to by different names. The specification and claims of the present application do not distinguish components based on the difference in names, but rather on the difference in function. As mentioned throughout the specification and claims, "comprising" or "including" is an open term, and should be interpreted as "including but not limited to". The subsequent description is a preferred embodiment of the present application, and is intended to illustrate the general principles of the specification, and not to limit the scope of the present application. The scope of protection of the present application is defined by the appended claims.
[0031] For the sake of understanding the embodiments of the present application, the following will be further explained with specific embodiments as examples in conjunction with the drawings, and each drawing does not constitute a limitation on the embodiments of the present application.
[0032] Figure 1 is a flowchart of a video target tracking method provided by an exemplary embodiment of the present application, as shown in Figure 1 The method comprises:
[0033] S100: synchronously acquiring an infrared image and a visible light image containing a target to be tracked;
[0034] S200: pre-processing the infrared image and the visible light image;
[0035] S300: constructing a target tracking model to be tested, and training the model;
[0036] S400: inputting the pre-processed infrared image and visible light image into the trained target tracking model to be tested, to identify and track the position of the target to be tracked.
[0037] In another exemplary embodiment, in step S200, the pre-processing of the infrared image comprises:
[0038] S201: performing size unification and normalization processing on the infrared image;
[0039] In this step, due to the inconsistency of inter-frame resolution caused by the infrared images from different sensor devices, the input image is first subjected to scale normalization processing, i.e., the infrared images are uniformly scaled to a preset standard size (such as 224x224 or 512x512), so as to guarantee the consistency of the subsequent model input dimension.
[0040] After the size unification is completed, in order to eliminate the absolute difference of the thermal value distribution of different infrared images, the infrared images are further subjected to normalization processing, i.e., the pixel intensity is compressed to a standard interval [0, 1]. The normalization not only can improve the numerical stability of image processing, but also can provide a unified thermal value reference basis for the next step of local contrast enhancement processing.
[0041] S202: performing local contrast enhancement and non-linear thermal value mapping on the infrared image subjected to the normalization processing;
[0042] In this step, the infrared image subjected to the size unification and thermal value normalization processing has the standardized thermal signal expression capability. However, in the scene with medium and low temperature background or non-obvious thermal gradient (such as night distance monitoring, complex urban background), the target thermal source region may still be indistinguishable from the background, which limits the subsequent target detection and tracking effect. In order to enhance the distinguishability of the target region, the CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm is introduced in this step, which dynamically adjusts the gray scale distribution in the local image window, enhances the local contrast of the image, highlights the edge structure and local thermal change of the thermal source region, and suppresses the influence of the large-area low-contrast region on the overall visual perception. At the same time, the CLAHE algorithm sets a contrast limit threshold, which effectively avoids the problem of noise amplification caused by over-enhancement in the high dynamic range infrared image.
[0043] On the basis of the CLAHE enhancement, the non-linear thermal value mapping method is further adopted in this embodiment, such as the logarithmic transformation (applying the logarithmic function mapping to the normalized thermal value graph, improving the response amplitude of the low thermal value area (cold area, medium and low temperature area), and enhancing the visual saliency of the weak thermal signal (such as a long-distance pedestrian, a low-heat equipment) in the image) or gamma correction (by adjusting the gamma parameter (γ∈ [0.4, 1.2]), the brightness enhancement degree of the medium and low thermal value area is controlled, the overall visual effect and detail performance of the image are considered, and the edge definition and contrast of the thermal source are further improved), so as to lift the response value of the medium and low thermal area, thereby amplifying the weak thermal signal (such as a long-distance pedestrian, an edge thermal source vehicle).
[0044] The above processing mode can simulate the sensitive response of the human eye to temperature difference, actively amplify the weak thermal signal in the infrared image, and at the same time maintain the detail fidelity of the high thermal area (such as a vehicle engine, a human core thermal area), so as to effectively improve the saliency and discrimination ability of the target region.
[0045] S203: Construct a spatial Gaussian attention map based on the thermal intensity and guide the weighted smoothing process.
[0046] In this step, in order to further highlight the target heat source region in the infrared image and suppress the background redundant signal, a spatial thermal sensation weight map is constructed based on the enhanced thermal value map. The specific method is as follows: first, according to the thermal value map after local contrast enhancement and nonlinear thermal value mapping, the high temperature region in the image is identified, and these heat source regions are taken as thermal sensation center points. On this basis, a two-dimensional Gaussian distribution function is constructed around each heat source center to generate a corresponding spatial thermal sensation attention map (Attention Map). The attention map gives different response weights to different positions in the image, and the weight value of the high temperature region is close to 1, and the weight of the cold background region far from the heat source gradually tends to 0, thereby forming a continuous and smooth spatial weight map.
[0047] Subsequently, the attention map is used as a mask to perform weighted smoothing processing on the original enhanced image: in the heat source aggregation region, the edge structure is kept clear to retain the thermal sensation features of the key target; and in the background region where the thermal gradient changes gently or there is no significant heat source distribution, the blur suppression is performed to reduce the response strength of the redundant noise.
[0048] In another exemplary embodiment, in step S200, the pre-processing of the visible light image includes the following steps:
[0049] S2001: Perform size and brightness standardization processing on the visible light image;
[0050] In this step, in order to ensure consistent input dimensions and distribution characteristics with the infrared image, the visible light image is first subjected to basic standardization processing. Specifically, it includes:
[0051] Size standardization: uniformly scale the image to a set size (such as 224x224 or 512x512) to facilitate subsequent model alignment and fusion;
[0052] Brightness standardization: normalize the image pixel value to [0, 1] or maintain it in the standard 8-bit RGB range;
[0053] Dynamic range compression: in view of the brightness difference under different shooting scenes, the brightness of the whole image is processed through linear adjustment or gamma correction (γ∈[1.5, 2.2]) to eliminate the unevenness of bright and dark areas.
[0054] This step serves as the starting point of the pre-processing and provides a unified calculation basis for subsequent illumination enhancement and texture extraction.
[0055] S2002: Perform illumination correction and local contrast enhancement on the standardized visible light image;
[0056] In this step, the normalized visible light image may still have problems such as local underexposure and structural blur in night or complex lighting scenes, which affects the accuracy of subsequent feature extraction and target detection. To improve the lighting robustness of the visible light image, this embodiment introduces a multi-scale Retinex algorithm (such as MSRCR) to perform local dynamic compression and global contrast enhancement processing on the normalized visible light image. This algorithm decomposes the reflection and illumination components of the image through multi-scale logarithmic transformation, dynamically compresses the local strong light area, and enhances the detail performance of the weak light area, which can simulate the visual compensation mechanism of the human eye under different brightness conditions. The algorithm specifically includes: constructing a Gaussian filter pyramid under different scales to realize multi-scale modeling of the local lighting distribution of the image; applying logarithmic dynamic range compression to suppress local overexposure or underexposure effects and enhance the detail performance of dark areas; introducing a color restoration mechanism to maintain the naturalness of image colors while ensuring brightness enhancement effects and preventing color distortion and artifacts.
[0057] After the above processing, the visible light image as a whole can present the effect of enhanced edge contrast, clear object outline, rich details, and balanced lighting, which helps to improve the image quality under complex lighting conditions.
[0058] S2003: Edge extraction and detail enhancement are performed on the visible light image after illumination correction.
[0059] In this step, on the basis of an image with good lighting balance, in order to further tap its structural detail information, this step uses the Sobel operator or Canny edge detection algorithm to extract the significant edge regions of the image (such as the window frame, license plate outline, and portrait edge), and generates an edge mask image.
[0060] Then, a de-sharpening mask is constructed, and the edge map and the original image are superimposed to enhance the local texture details. This structure enhancement process not only helps the model to accurately recognize boundary information, but also significantly improves the expression ability of the visible light image for high-frequency features, providing edge positioning reference for multi-modal feature alignment.
[0061] S2004: Convert the visible light image after detail enhancement to Lab space and enhance the color contrast of a / b channels.
[0062] In this step, the color and structure in the visible light image are often coupled, which affects the accuracy of boundary modeling. To effectively decouple color features and structural features, this step converts the image from RGB space to Lab color space, where L channel represents brightness and a / b channel represents color component.
[0063] By performing nonlinear stretching and normalization on the a / b channel, the color contrast between different object regions in the visible light image can be effectively improved, and the ability to distinguish the target and the background in visual perception can be enhanced. This operation is particularly suitable for targets with color distinguishing features (such as vehicle / background, clothes / road, etc.), further promoting the complementary fusion of color and texture between modalities.
[0064] In another example embodiment, as shown in Figure 2 The target tracking model to be tested includes a multi-modal fusion module for performing multi-modal fusion on the preprocessed infrared image and visible light image to obtain a multi-modal fusion image; and a video tracking module for identifying and tracking the position of the target to be tested based on the multi-modal fusion image.
[0065] In this embodiment, the multi-modal fusion module includes a feature extraction and alignment unit, an adaptive feature fusion unit, and an optimization unit. The feature extraction and alignment unit is configured to extract multi-scale features of the preprocessed infrared image and visible light image and perform spatial alignment. The feature extraction and alignment unit includes an infrared branch and a visible light branch, which are respectively configured to extract features from the preprocessed infrared image and visible light image.
[0066] The infrared branch is configured to extract thermal feature information from the preprocessed infrared image and obtain an infrared feature map. The infrared branch includes a thermal guidance attention module, a multi-scale pyramid residual extraction module, and a deformation convolution enhancement layer. First, the thermal guidance attention module generates a thermal weight map w(x, y) (high temperature region close to 1, background region close to 0) by sequentially performing thermal value normalization (mapping the pixel thermal value in the infrared image to the standard interval [0, 1] to eliminate the thermal value scale difference between different images), Gaussian smoothing (applying Gaussian filtering to the normalized image to reduce the influence of local noise and highlight the overall trend of thermal distribution), and nonlinear stretching (using nonlinear mapping methods such as logarithmic transformation or gamma correction to dynamically stretch the smoothed thermal value map, further improving the response value in the low-thermal region). This thermal weight map not only serves as a spatial attention mask and multiplies the original infrared image to form a front attention map to enhance the feature expression of the high-thermal target region in space, but also dynamically adjusts the weight distribution of each channel of the infrared image feature map through the embedded channel attention mechanism (CBAM), thereby guiding the model to focus on the high-thermal target region (such as human body, vehicle) and suppress background interference.
[0067] On the basis of heat-guided, the multi-scale pyramid residual extraction module adopts a four-branch structure, respectively using 1x1, 3x3, 5x5 standard convolution and 3x3 dilated convolution to extract different spatial scale features of heat source. Among them, the 1x1 convolution branch (1x1 convolution can reorganize the features of each pixel point in the channel dimension, has strong local response ability, and is helpful to retain the key area with high heat intensity but small area) is used to reorganise the channel features, and is matched with batch normalization and ReLU activation function, responsible for capturing the local high temperature pixel area (heat spot center, such as heat source concentration point of vehicle engine) in the preprocessed infrared image, so as to avoid the "heat feature averaging" caused by subsequent convolution or fusion operation, thereby enhancing the model's attention to the tiny target in the input feature map. The 3x3 convolution branch is responsible for capturing the edge profile information of the medium scale heat target (such as car window, tire) in the preprocessed infrared image. Compared with 1x1 convolution, 3x3 convolution has stronger local spatial modeling ability and can identify the heat-cold boundary jump features existing in the image. The 5x5 convolution branch is responsible for extracting the area with wide temperature distribution (such as vehicle body) in the preprocessed infrared image, so as to strengthen the model's perception ability to the low-frequency heat structure such as background heat block and heat group profile. When the heat source area distribution is large, small scale convolution (such as 1x1 convolution and 3x3 convolution) is difficult to perceive the overall temperature change trend, while 5x5 convolution can provide wider structure capturing ability. The 3x3 dilated convolution branch is responsible for capturing the area with fuzzy edge and uneven heat diffusion (such as heat tail produced by vehicle in motion) in the preprocessed infrared image. 3x3 dilated convolution allows skip sampling, which can effectively extract long-distance temperature change information to enhance the response to heat gradient change band, so as to make up for the fuzzy area which cannot be effectively modeled by the previous three branches, thereby enhancing the model's perception robustness to continuous heat distribution and fuzzy profile.
[0068] The output features of the above four branches are spliced in the channel dimension, then compressed and fused in the channel dimension by 1x1 convolution to form a unified scale heat feature map. Then, residual connection is performed with the preprocessed original infrared image, and finally the enhanced heat feature map under unified scale is obtained for the subsequent infrared-visible light feature alignment and fusion process.
[0069] In order to further improve the model's adaptability to non-rigid structure and fuzzy edge area, the infrared branch also introduces a deformation convolution enhancement layer, which includes an offset generation submodule and a deformation convolution submodule. The offset generation submodule takes the enhanced heat feature map under unified scale as input, extracts the local structure (such as heat boundary, texture gradient) in the enhanced heat feature map by using 3x3 standard convolution, and predicts a two-dimensional offset vector in (x, y) direction for each convolution sampling point . The deformation convolution submodule takes the two-dimensional offset vector and the enhanced heat feature map under the unified scale as input, at each convolution sampling point, according to the two-dimensional offset vector The sampling point position is dynamically adjusted, and then weighted convolution is performed to finally output the heat feature enhancement map
[0070]
[0071] wherein, represents the heat feature enhancement map The response value at the current output feature point position reflects the heat feature expression after deformation sampling and weighted convolution; represents the input infrared branch feature map; represents the total number of sampling points covered by the convolution kernel; represents the th sampling position under the standard convolution window; represents the learnable offset of the th point; represents the weight of the convolution kernel at the th sampling point.
[0072] In summary, the infrared branch is driven by heat attention, and the joint action of multi-scale residual modeling and deformation perception enhancement modules constructs a complete, high-precision, and sensitive feature extraction path for the target heat area, providing a high-quality structural foundation and heat source alignment capability for subsequent infrared-visible light feature fusion.
[0073] The visible light branch is used for deep feature extraction of the preprocessed visible light image to obtain a visible light feature map with clear structure, rich texture, adaptive illumination, and modal alignment capability. The visible light branch includes a guided edge enhancement module, a multi-channel frequency perception module, and a color-structure alignment module. First, the guided edge enhancement module performs gradient calculation on the preprocessed visible light image through the Sobel operator to extract a preliminary edge image. This edge image can effectively highlight the significant structural outlines in the visible light image, such as window frames, license plate edges, and other detailed information. Then, the edge image is converted into a spatial attention mask and embedded into the main convolutional neural network (CNN) as weight guide information to multiply with the feature map elements, thereby guiding the main network to focus more on the structural significant area during feature extraction and generating an edge attention mask graph to improve the model's perception of high-frequency edge details such as boundaries and textures in the visible light image, while suppressing invalid features in non-target areas, thereby helping to improve the geometric expression accuracy and discrimination ability of the visible light feature map in multi-modal fusion.
[0074] On the basis of structural enhancement, the multi-channel frequency perception module introduces a frequency domain analysis mechanism to mine the discriminative features carried by different frequency components in the image. The multi-channel frequency perception module includes a frequency domain transformation submodule, a channel attention regulation submodule, and a frequency selective enhancement submodule, aiming to realize the separate modeling and adaptive enhancement of different frequency components (low-frequency structure and high-frequency detail) in the visible light image, to improve the texture resolution and boundary clarity. Among them, the frequency domain transformation submodule first performs a two-dimensional fast Fourier transform (2D FFT) on the input multi-channel visible light feature map, converts each channel from the spatial domain to the frequency domain, and obtains a frequency spectrum map containing amplitude spectrum and phase spectrum information, wherein the amplitude spectrum reflects the local change intensity of the image (such as texture information), and the phase spectrum maintains the structure position information. Further, to improve the frequency representation effect, the frequency spectrum map is further centralized, and a high-pass and low-pass mask is used to extract high-frequency regions (such as license plate characters, window frame lines) and low-frequency regions (such as road background, vehicle contour) respectively.
[0075] The channel attention regulation submodule respectively performs global average pooling and maximum pooling on the frequency spectrum map in the channel dimension to extract channel response statistical features, and generates a channel attention vector through a fully connected layer with two shared parameters, which is used to weight the high-frequency and low-frequency channel feature maps, thereby realizing adaptive enhancement of the "high-response frequency channel". Through this frequency and channel collaborative modeling method, the network can dynamically focus on the frequency band and channel with the largest texture expression contribution, improving the discriminative ability of key texture regions.
[0076] The frequency selective enhancer module is responsible for precisely mining and enhancing the frequency features that are discriminative in the image in the frequency domain. The module first receives the frequency spectrum processed by fast Fourier transform (FFT), and introduces a set of learnable two-dimensional complex frequency domain filters (such as high-pass, band-pass, band-stop, etc.) to cover different frequency regions from high-frequency details (such as license plate characters, window frame contours) to low-frequency structures (such as vehicle body contours, background blocks). Unlike traditional static masks, these filters can be dynamically adjusted according to the target texture distribution in the specific scene during the training process, so as to enhance the key frequency band features of the target region. To avoid missing important information, the module retains the frequency response characteristics after each filtering output, and integrates the full frequency domain information through a weighted fusion method to improve the collaborative expression ability of global and local information. Then, all the filtered and enhanced frequency spectrum is converted back to the spatial domain by inverse Fourier transform to reconstruct the image feature map containing enhanced texture. Finally, the image feature map is fused with the original spatial feature map through residual connection to ensure that the enhanced features improve the discriminability while maintaining the original structure without distortion. The frequency selective enhancer module can achieve frequency selective enhancement of image texture, and effectively improve the sensitivity and discrimination ability of the model to detailed targets in complex scenes.
[0077] The color-structure alignment module is used to realize decoupled expression of color information and structure information and cross-modal alignment, and the color-structure alignment module includes a feature decoupling sub-module, an illumination adaptive reorganization sub-module, and a structure alignment correction sub-module. The feature decoupling sub-module is used to convert the input visible light image into a YUV color space or a Lab space, so as to realize preliminary separation of image color (U / V or a / b channel) and structure (Y or L channel) information. Then, the color component and the brightness component are respectively subjected to convolution coding to extract color feature maps and structure feature maps. To further strengthen the independence of the two types of features, the module introduces a channel separation convolution and an orthogonal constraint loss mechanism (in a deep network, for example, two feature vectors 、 represent different semantic information, but they are highly correlated, so they may not be effectively decoupled, resulting in blurred or redundant effects after fusion. Through orthogonal constraint, the 、 between them are forced to be as orthogonal as possible, that is, the included angle is 90° in geometry, and the inner product is 0 in mathematics. For this application, the color features and the structure features can be made to overlap as little as possible in the feature space, so as to achieve better semantic separation), so that the color branch is mainly responsible for capturing semantic features such as hue and color temperature, and the structure branch focuses on edges, shapes and textures.
[0078] On the basis of decoupling features, the illumination adaptive recombination sub-module generates an illumination mask by calculating the local brightness gradient of the structural feature map, to regulate the recombination weight of color and structure. On this basis, the illumination adaptive recombination sub-module introduces a lightweight Gated Linear Unit (GLU), which performs pixel-level weighted recombination of color and structure features according to pixel position, dynamically suppresses color artifacts in unstable illumination areas, and strengthens the structural expression of the target boundary area, thereby generating a fusion feature map that is robust to illumination changes.
[0079] The structure alignment correction sub-module acts on the structural feature map of the visible light image, and the core goal is to correct the areas with offset or inconsistency in the visible light structural feature by using the thermal structural feature extracted from the infrared image as a guide. First, the thermal structural feature extracted from the infrared image is input into a lightweight attention guide module, and is spatially aligned with the structural feature in the visible light image to form an edge alignment guide signal. Then, through a residual connection mechanism, the part of the visible light structure map that fails to accurately align with the infrared heat source area is dynamically corrected. Finally, an aligned structure map with consistent edge position and clear structure expression is output, which is used for subsequent modal fusion.
[0080] The adaptive feature fusion unit is used to adaptively fuse the aligned features to obtain a fused image. The adaptive feature fusion unit includes a channel semantic enhancement sub-module, a spatial semantic gating sub-module, and a structure consistency guide sub-module. The channel semantic enhancement sub-module first receives the feature maps output by the infrared branch and the visible light branch respectively, and performs independent global average pooling operations to obtain the global response and peak response information of each channel. Then, a shared two-layer fully connected network (MLP) is used to non-linearly map the two sets of pooling results to generate an infrared branch channel weight vector and a visible light branch channel weight vector , which respectively reflect the importance of the channels in the respective modalities.
[0081] To further introduce cross-modal dependency modeling, the two sets of channel weights are normalized to each other to improve complementarity:
[0082]
[0083]
[0084] wherein, represents the difference-aware infrared channel attention weight, and represents the importance distribution of infrared in fusion; represents the importance distribution of visible light in fusion after difference perception.
[0085] The cross-modal dependency modeling is a complementary modeling strategy, which can improve the discrimination ability after fusion by comparing the channel weights with each other, for example, if is obviously greater than , tends to be biased, on the contrary, if is greater than , will pay more attention to the modality, in addition, by modeling the difference through softmax, the model can be more inclined to allocate more weight in the modality with significant semantic difference, thereby improving the effect of cross-modal feature fusion.
[0086] The spatial semantic gating sub-module adopts a shared lightweight convolutional block to generate spatial response maps at position (x, y) for infrared and visible light feature maps respectively , In order to preserve key target regions (such as human heat sources, car window edges), the two spatial maps are further fused at the pixel level by complementary weighting:
[0087]
[0088] wherein, represents the response value of the fused spatial attention mask at position (x, y); represents a Sigmoid activation function, is a trainable weight used to dynamically balance the contributions of infrared and visible light in the spatial domain.
[0089] The final spatial mask is used as a gating mechanism to regulate the spatial response of the fused feature map, suppressing background noise (such as trees, shadows) and highlighting target regions.
[0090] Considering the natural differences in perception between infrared and visible light, direct fusion often leads to structural inconsistencies. To this end, the present application designs a structural consistency guiding sub-module, which first extracts gradient maps (using Sobel or Laplace operators) from the two branches respectively, obtaining structure maps , Then calculate the structure difference map :
[0091]
[0092] Further, the structure difference map is input into a lightweight residual network as guiding information, specifically, the structure difference map First, two consecutive 3x3 convolutional layers are used to extract local structural change features and enhance non-linear expression capabilities, followed by a ReLU function.
[0093]
[0094] wherein, is the fusion map after channel attention and spatial gating, and the final output ensures the integrity of the infrared heat source while maintaining the spatial consistency of the structural edges.
[0095] Finally, the output is normalized by the Sigmoid activation function to ensure that the structure correction weight is in the interval [0, 1], which can be used to locally enhance the edges and fine-tune the position of the fused features.
[0096] The optimization unit aims to optimize and reconstruct the image features after preliminary fusion, making up for information loss and edge blur that may occur during modal fusion, thereby improving the overall quality and expression ability of the fused image. The optimization unit includes a global semantic compensation module, an edge refinement enhancement module, and a multi-scale reconstruction module to achieve multi-level optimization from semantic completion to edge restoration. First, the global semantic compensation module introduces a Transformer encoder structure to model the global context of the preliminary fused image, focusing on capturing long-distance cross-regional dependencies. Since there may be information incompleteness or insufficient semantic complementarity between infrared and visible light images in some areas, direct fusion may result in structural disconnection or regional semantic loss. Through the multi-head attention mechanism of Transformer, the global semantic compensation module can dynamically mine the potential semantic relationships between distant regions in the preliminary fused image features, thereby achieving semantic completion of the fused image in terms of content and improving overall consistency and expression integrity.
[0097] Secondly, to solve the problems of edge blur and structure breakage in the modal fusion process, the edge refinement enhancement module uses an image gradient guided mechanism to repair structural details. Specifically, the fusion image and its Sobel edge map are jointly input into a shallow convolutional network, and the edge is refined and enhanced through feature residual learning to enhance the continuity and distinguishability of image structural details and avoid edge blur or loss after modal fusion. This module focuses on strengthening high-frequency features such as object contours and structural boundaries in the image, so that the fusion result has a more clear and continuous edge expression while maintaining the integrity of the heat source information, improving the readability and recognizability of the image.
[0098] Finally, the multi-scale reconstruction module introduces a spatial pyramid structure (such as Atrous Spatial Pyramid Pooling, ASPP) to obtain image features under different receptive fields and realize collaborative modeling of global contours and local textures. This module extracts features from the fusion image through multiple sampling scales (scale 1 (rate = 1): equivalent to a normal 3x3 convolution, focusing on fine-grained local texture features; scale 2 (rate = 6): receptive field expansion, capturing medium-scale structures (such as windows, human contours); scale 3 (rate = 12): can extract wide-area background structures (such as car bodies, roads); scale 4 (rate = 18): used to extract global semantic information and low-frequency contours in the image), and uses a cascaded fusion method (concatenating the feature maps extracted under different sampling scales in the channel dimension, and then using a 1x1 convolution for channel compression and information integration to output a unified feature map) to uniformly reconstruct the feature maps under multiple scales and output an optimized image feature map with consistent scales and complete structures. This module can effectively preserve the semantic consistency of the image at a large scale and the edge details at a small scale, providing a stable and balanced resolution image input for subsequent high-precision tasks.
[0099] In summary, the optimization unit not only improves the clarity and contrast of the fusion image, but also helps to enhance its structural consistency and semantic expression ability in multi-target scenes, providing a higher quality input basis for subsequent target detection and video tracking tasks.
[0100] Further, the video tracking module aims to fuse the multi-modal features of infrared images and visible light images, and to identify and continuously locate the dynamic behavior of the target to be tracked in the video. The video tracking module comprises a dynamic initialization unit, a lightweight temporal correlation unit, and an occlusion-aware updating unit. The dynamic initialization unit is configured to construct multi-modal templates to provide tracking starting reference information, and to ensure that the system has multi-modal understanding ability for the target at the starting stage of the video sequence. The dynamic initialization unit introduces a semantic response template pool, which extracts high semantic features of the target region in the first frame of the video as the main template , which is used as a whole matching reference; meanwhile, an infrared auxiliary template and a visible light auxiliary template are constructed, wherein the infrared auxiliary template mainly focuses on the infrared thermal structure (anti-occlusion) to improve the robustness in the occlusion environment; the visible light auxiliary template focuses on the visible light texture details (anti-mismatch), and strengthens the texture and edge features, which helps to reduce the mismatch caused by texture weakening. In the subsequent frames, the three types of templates are modally weighted and fused to obtain the fusion template of the current frame :
[0101]
[0102]
[0103] wherein, represents the weighting coefficient of the main template in the current frame, reflecting the credibility of the main template at the current time; represents the weighting coefficient of the infrared template, reflecting the contribution of the infrared feature to the current target; represents the weighting coefficient of the visible light template, reflecting the contribution of the visible light feature to the current target.
[0104] The lightweight temporal correlation module aims to establish the short-time dependence of the target state and identify the motion trend between consecutive frames while maintaining the computational efficiency, and to predict the possible position of the target in the next frame. The lightweight temporal correlation module comprises a pre-frame guided residual update mechanism and a dual-channel saliency enhancement mechanism. Specifically, the pre-frame guided residual update mechanism uses the temporal residual information between the current frame feature map and the previous frame feature map to guide the update of the current target position, so that the model can better capture the motion trend and structural changes. Specifically, first, the current frame feature map and the previous frame feature map are obtained, then the temporal residual information between the two frames is calculated to capture the motion direction and change amplitude of the target:
[0105]
[0106] The dual-channel saliency enhancement mechanism respectively performs saliency weighting on the infrared channel and the visible light channel to respectively extract an infrared saliency map and a visible light saliency map :
[0107]
[0108]
[0109] wherein, represents a Sigmoid activation function.
[0110] temporal residual information and the infrared saliency map and the visible light saliency map are cascaded and subjected to 1x1 convolution to obtain a target center displacement vector :
[0111]
[0112] wherein, represents a predicted target center displacement vector.
[0113] The final estimated target search region of the next frame is:
[0114]
[0115] wherein, represents an estimated position of the target region of the next frame; represents a position of the target region of the current frame.
[0116] By estimating the target search region of the next frame, the prediction and update of the target position can be completed. This module realizes cross-frame dependency modeling through the simplest structure, effectively improves the continuity and anti-interference ability of tracking, and avoids the high computational burden of traditional complex temporal models in embedded environments.
[0117] The occlusion awareness updating unit is used to detect the occlusion state based on the similarity heat map change and guide the template updating strategy to maintain the tracking stability, so as to improve the robustness and continuity of video target tracking in complex occlusion scenes. Specifically, this module calculates the chord similarity between the current candidate region feature and the fusion template of the current frame to generate a similarity heat map of the current frame :
[0118]
[0119] wherein, represents the feature vector of the current frame at position (x, y) in the candidate region; represents the fusion template vector of the current frame; represents the L2 norm of a vector, used for normalization operation.
[0120] After obtaining the current frame similarity heat map , a comparison analysis is performed with the previous frame similarity heat map. If a locally continuous low response region (i.e., a "low response hole") is detected in the target region, and its distribution position is significantly offset from the previous frame salient region, the system determines that the target may have partial occlusion or deformation out of the template representation range.
[0121] To enhance the robustness of the tracking system to the occlusion situation, the module introduces an occlusion confidence map as a position weight to adjust the participation degree of each region in the current frame feature map. Specifically, the original feature map is point-by-point weighted with the occlusion confidence map (ranging between [0, 1], 1 indicating trustworthiness and 0 indicating occlusion):
[0122]
[0123] wherein, represents the valid feature value of the current frame at position (x, y).
[0124] By assigning a position weight to the current feature map, only the high-confidence region can be retained as valid features to participate in tracking update, while the influence of the occluded region on position judgment is weakened.
[0125] Further, if occlusion is detected, the system will call the infrared auxiliary template to compensate for redundant features, so as to extract valid information from the unoccluded region and guide the model to continuously lock the target.
[0126] In summary, the video tracking module takes semantic template initialization, temporal residual modeling, and occlusion adaptive update as the core mechanism, and builds a lightweight, efficient, and robust multi-modal target tracking scheme. Its structure takes into account the calculation efficiency and tracking continuity, and can adapt to various complex environments such as motion, occlusion, and light changes, significantly enhancing the practicality and stability of the system in multi-target video analysis tasks.
[0127] In another exemplary embodiment, in step S300, the target tracking model to be tested is trained by the following steps:
[0128] S301: Data preparation and preprocessing: synchronously collect multiple pairs of infrared images and visible light images, preprocess the two types of images respectively, and simultaneously label the real positions of the target in the preprocessed infrared images and visible light images to obtain a training sample set;
[0129] S302: Feature extraction and fusion module training: input the preprocessed infrared images and visible light into the infrared and visible light feature extraction network to extract thermal and texture structure features respectively; the fusion module introduces channel attention, spatial gating mechanism and structure consistency guidance strategy to realize adaptive fusion of multi-modal features; the fusion effect is optimized through supervised training of the structure reconstruction error of the original image.
[0130] S303: Template matching and position regression training: extract candidate region features in each frame, and calculate similarity heat maps with the initialized main template, infrared template and visible light template; construct a supervised heat map based on the labeled position to optimize the matching accuracy, and simultaneously train the position regression branch to predict the target displacement vector and realize continuous tracking of the target trajectory.
[0131] S304: Occlusion perception module training: compare the similarity heat map changes of the current frame and the previous frame to construct an occlusion confidence map for suppressing low confidence regions in the feature map; the trained model maintains tracking stability under occlusion conditions.
[0132] S305: End-to-end joint optimization: integrate the fusion module, tracking module and occlusion perception module into a unified model, jointly train in an end-to-end manner, and optimize the model performance through multiple loss functions (including reconstruction loss, similarity loss, displacement prediction loss and occlusion suppression loss); when the loss functions tend to be stable on the validation set and meet the preset convergence conditions, the training process is considered complete, and finally the optimized model parameter for multi-modal image fusion and robust tracking is obtained.
[0133] In this embodiment, the reconstruction loss function is expressed as follows:
[0134]
[0135] wherein, represents the weight of the L1 loss term; represents the fusion image output by the model; represents the reference image; represents the weight of the SSIM loss term; represents the L1 norm, which is used to measure the absolute difference between the pixels of two images; represents the Structural Similarity Index Measure (SSIM), which is used to measure the similarity of images in structure, brightness and contrast.
[0136] Similarity loss function It is expressed as follows:
[0137]
[0138] in, Indicates the current frame's position Similarity heatmap; This represents a supervised heatmap constructed based on ground truth (GT) location information. This represents the square of the L2 norm, used to measure heatmap differences.
[0139] Displacement prediction loss function It is expressed as follows:
[0140]
[0141] in, This represents the target displacement vector predicted by the model (the change in the center position from the current frame to the next frame). Represents the actual target displacement vector; This represents smoothed L1 loss.
[0142] Occlusion suppression loss function It is expressed as follows:
[0143]
[0144] in, This represents the occlusion confidence value predicted at position (x, y) in the current frame, with a value range of [0, 1]. This represents the supervision mask for the occlusion area, with a value range of [0,1], where 1 represents the visible area and 0 represents the occlusion area; This represents the binary cross-entropy loss function, used for binary classification probability evaluation.
[0145] Total loss function It is expressed as follows:
[0146]
[0147] in, , , and These represent the weighting coefficients for reconstruction loss, similarity loss, displacement prediction loss, and occlusion suppression loss, respectively.
[0148] Below, this application uses Figure 2 The original infrared image shown and Figure 3Using the original visible light image shown as an example, the method described in this application and the conventional method are compared and explained.
[0149] Figure 4 Images generated using traditional image fusion methods primarily employ image-level weighting or feature stitching strategies. While these methods can demonstrate modal fusion effects in certain areas, the overall image still suffers from significant structural blurring, unclear edges, and insufficient thermal representation. The boundary between the target and background lacks clear segmentation, making it difficult to form a salient fusion representation, especially in scenes with complex backgrounds or weak target heat sources. Figure 5 The image shown is fused using the method described in this application. Through the introduction of multimodal feature alignment, adaptive fusion, and optimized reconstruction mechanisms, this image retains the thermal response characteristics of the infrared image while accurately preserving the structural details and texture information of the visible light image. The fusion result exhibits clear boundaries, strong contrast, and a more prominent target region, demonstrating higher perceptual consistency and discriminative power, providing a more stable feature base for subsequent target detection and tracking.
[0150] Figure 6 The results demonstrate the effectiveness of traditional target tracking methods. While the tracking bounding box (white rectangle) roughly covers the vehicle, it exhibits a noticeable offset, slightly tilted to one side. This indicates that the tracker is prone to positioning errors and exhibits poor stability under conditions of background interference or unclear textures. Furthermore, the overall image grayscale is low, with the grayscale values of the trees in the background and the vehicle being similar, resulting in weak heat-texture contrast and difficulty in distinguishing the target from the background, thus interfering with the algorithm's accuracy. Details on the vehicle body, such as brand logos and window outlines, are blurred, which can lead to gradual degradation of target features during long-term tracking, ultimately causing tracking failure. In contrast, Figure 7 The tracking performance of the proposed method is demonstrated. The fused image features are clear, the tracking bounding box accurately covers the vehicle body, exhibits good symmetry, and high edge fit, fully illustrating that the proposed method enhances target saliency and boundary clarity through multimodal fusion, making it easier for the model to extract stable ROIs (Regions of Interest). The fused image combines the thermal response of the infrared image with the detailed texture of the visible light image, resulting in clear vehicle edges, soft lighting, and good segmentation between the background and the target, effectively suppressing interference from irrelevant background information. Furthermore, this method can preserve high-frequency texture features such as vehicle body reflections and corner contours, providing continuous and identifiable structural information for multi-frame target recognition and relocalization, significantly improving the robustness and accuracy of video target tracking.
[0151] Figure 8The tracking effect of the traditional method in the target high-speed motion scene is shown. Although the tracking frame still covers the vehicle body, there is obvious delay and position drift, especially when the vehicle speed is fast or the target turns quickly, the position of the frame cannot be updated in real time, showing the "tail" effect, which is difficult to closely fit the vehicle contour. The root cause of this phenomenon is that the traditional method is difficult to resist the image blur caused by high-speed motion, and the target edge becomes blurred, which makes the model unable to extract discriminative features stably in the case of rapid appearance change, and is prone to misjudgment of background as target, showing insufficient fault tolerance. In addition, the overall image generated by the traditional method is blurred, the contours of the vehicle head and tail are not clear, and the frequency components of the background and the target are highly overlapped, which further aggravates the interference in the tracking process and reduces the positioning accuracy. In contrast, Figure 9 The performance of the present method under the same high-speed motion condition is shown. The tracking frame always closely follows the vehicle edge and stably remains in the center position of the motion path, without obvious deviation or drift phenomenon, showing good spatiotemporal coherence and dynamic adaptability. This advantage is due to the fact that in the image fusion stage of the present method, the infrared image enhances the heat source boundary and the visible light image enhances the high-frequency texture details, thereby forming clear and high-recognizability structural features, which improves the robustness of the model to motion blur scenes. At the same time, the present method introduces a reinforcement learning strategy and a lightweight time series prediction mechanism, so that the model has a certain motion trend perception ability and can adaptively adjust the tracking strategy. Even if the target appears complex changes such as acceleration and deflection, the target position can still be locked stably, and there is almost no tracking failure.
[0152] Figure 10 With Figure 11 Both of them show a typical occlusion scene: a vehicle is partially occluded by a large tree in front during driving, and only part of the target area is visible, which is a severe test of the tracking system's target integrity judgment and trajectory coherence prediction ability. In Figure 10 , when the traditional method is used for tracking, although the tracking frame roughly covers the vehicle body, after the occlusion occurs, especially in the vehicle head area, the frame body obviously deviates and drifts, the target positioning at the occlusion boundary is not accurate, and even shows frame jumping or unstable fluctuation. The main reason for this problem is that the traditional method highly depends on the pixel information of the visible area in the current frame, lacks the perception ability of historical trajectory or occlusion completion, and is difficult to infer from the context whether the occluded area belongs to the target as a whole. In addition, the occluded part and the background area have small difference in image gray level, and the boundary is blurred, which makes the algorithm easily misidentify them as background area, resulting in tracking failure or loss of target.
[0153] In contrast, Figure 11The patent method demonstrates the tracking performance in the same occlusion context, and the tracking frame still stably covers the entire vehicle body. Even if the vehicle head is blocked by trees, the frame body does not shrink or shift significantly, showing good occlusion tolerance and target consistency judgment ability. This robustness is due to the fusion of infrared thermal images and visible light images in the method, which constructs a redundant perception image. The introduction of the occlusion confidence map and the context temporal modeling mechanism enables the system to predict the spatial location of the occlusion area through the semantic clues of the historical frames and surrounding frames even when some information is missing. In addition, the occlusion area still retains identifiable features such as the vehicle head heat source and tire profile in the fusion image, even if it is partially invisible. The model can still maintain the tracking continuity of the target as a whole through the temporal features modeled by the Transformer structure, thereby significantly improving the tracking stability and recovery ability in the occlusion environment. In summary, by comparing Figures 4 to 11 It can be seen that the video target tracking method proposed in the present application is superior to the traditional method in multiple key performance dimensions. In terms of image fusion quality, the method effectively combines infrared thermal and visible light structural information to generate a fusion image with clearer boundaries and more prominent target regions. In terms of target tracking performance, the fusion features enhance the target discrimination ability, making the tracking frame more stable and more accurate. Under high-speed motion conditions, the method successfully realizes the precise connection of target positions between consecutive frames through temporal residual modeling and motion prediction mechanism. In the occlusion scene, the introduced occlusion perception and confidence weighting strategy significantly improves the adaptability and tracking continuity of the occlusion area. Overall, the method has significant advantages in robustness, accuracy, and adaptability to complex environments, and has stronger practical value and promotion prospects.
[0154] In another exemplary embodiment, as shown in Figure 12 The present application also provides a video target tracking device, which comprises: an acquisition module 100 for synchronously acquiring infrared images and visible light images containing a target to be measured; a preprocessing module 200 for preprocessing the infrared images and the visible light images; a model construction and training module 300 for constructing a target to be measured tracking model and training the model; and an identification module 400 for inputting the preprocessed infrared images and visible light images into the trained target to be measured tracking model to identify and track the position of the target to be measured.
[0155] In another exemplary embodiment, the present application also provides a storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method of any one of the preceding embodiments.
[0156] In another example embodiment, the application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of the preceding method embodiments when executing the program.
[0157] The above merely describes the preferred embodiments of the application, and is not intended to limit the patent scope of the application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, based on the content of the specification and drawings of the application, are also included in the patent protection scope of the application.
Claims
1. A method of video object tracking, characterized by, The method comprises: synchronously acquiring an infrared image and a visible light image containing a target to be detected; preprocessing the infrared image and the visible light image, comprising: performing size unification and thermal value normalization processing on the infrared image; performing local contrast enhancement and nonlinear thermal value mapping on the normalized infrared image; constructing a spatial Gaussian attention map based on thermal intensity and guiding weighted smoothing processing, comprising: identifying a thermal center point in the image according to the thermal value map after local contrast enhancement and nonlinear thermal value mapping; constructing a two-dimensional Gaussian distribution function around each thermal source center to generate a corresponding spatial thermal attention map, wherein the attention map gives different response weights to different positions in the image, thereby forming a continuous and smooth spatial weighted map; subsequently, using the attention map as a mask, performing weighted smoothing processing on the original enhanced image: in the thermal source aggregation area, the edge structure is kept clear to retain the thermal features of the key target; and in the background area where the thermal gradient changes gently or there is no significant thermal source distribution, fuzzy suppression is performed to reduce the response strength of redundant noise; preprocessing the visible light image, comprising: performing size and brightness standardization processing on the visible light image; performing illumination correction and local contrast enhancement on the normalized visible light image; performing edge extraction and detail enhancement on the illumination corrected visible light image; converting the detail enhanced visible light image to Lab space and enhancing the color contrast of a / b channels; constructing a target to be detected tracking model and training the model; the target to be detected tracking model comprises: a multi-modal fusion module for fusing the preprocessed infrared image and visible light image to obtain a fused image; the multi-modal fusion module comprises: a feature extraction and alignment unit for extracting multi-scale features of the preprocessed infrared image and visible light image and performing spatial alignment; The feature extraction and alignment unit comprises an infrared branch and a visible light branch, the infrared branch is used for extracting thermal feature information in the preprocessed infrared image and obtaining an infrared feature map, and specifically comprises a thermal sensation guided attention module, a multi-scale pyramid residual extraction module and a deformation convolution enhancement layer, wherein the thermal sensation guided attention module generates a thermal sensation weight map by sequentially performing thermal value normalization, Gaussian smoothing and nonlinear stretching; the multi-scale pyramid residual extraction module adopts a four-branch structure and extracts different spatial scale features of heat sources by using 1*1, 3*3, 5*5 standard convolution and 3*3 dilated convolution respectively; the deformation convolution enhancement layer comprises an offset generation submodule and a deformation convolution submodule, wherein the offset generation submodule takes an enhanced thermal feature map under a unified scale as input, extracts local structures in the enhanced thermal feature map by using 3*3 standard convolution, and predicts a two-dimensional offset vector in the (x, y) direction for each convolution sampling point ; the deformation convolution submodule takes the two-dimensional offset vector and the enhanced thermal feature map under the unified scale as input, dynamically adjusts the sampling point position according to the two-dimensional offset vector at each convolution sampling point, then performs weighted convolution, and finally outputs a thermal feature enhancement map ; the visible light branch is used for deep feature extraction on the preprocessed visible light image to obtain a visible light feature map, specifically comprising a guided edge enhancement module, a multi-channel frequency perception module and a color-structure alignment module, wherein the guided edge enhancement module performs gradient calculation on the preprocessed visible light image through a Sobel operator to extract a preliminary edge image; the multi-channel frequency perception module introduces a frequency domain analysis mechanism to mine the discriminative features carried by different frequency components in the image; and the color-structure alignment module is used for decoupled expression of color information and structural information and cross-modal alignment; inputting the preprocessed infrared image and visible light image into the trained target to be detected tracking model to identify and track the target to be detected.
2. The video object tracking method according to claim 1, characterized by, the target to be detected tracking model further comprises: a video tracking module for identifying and tracking the target to be detected based on the fused image.
3. The video object tracking method according to claim 2, characterized by, the multi-modal fusion module further comprises: an adaptive feature fusion unit for adaptively fusing the aligned features to obtain a fused image; an optimization unit for structure optimization and reconstruction of the fused image.
4. The video object tracking method according to claim 2, characterized by, the video tracking module comprises: an initialization tracking unit for constructing a multi-modal template to provide tracking starting reference information; A lightweight temporal correlation unit is configured to combine the previous frame residual information and the current frame features to predict the target displacement. An occlusion-aware updating unit is configured to detect the occlusion state based on the similarity heat map change and guide the template updating strategy to maintain the tracking stability.
5. A video object tracking apparatus characterized by comprising: The device comprises: An acquisition module is configured to synchronously acquire an infrared image and a visible light image containing a target to be detected; A preprocessing module is configured to preprocess the infrared image and the visible light image, including: performing size unification and thermal value normalization processing on the infrared image; performing local contrast enhancement and nonlinear thermal value mapping on the normalized infrared image; constructing a spatial Gaussian attention map based on thermal intensity and guiding weighted smoothing processing, including: identifying a thermal center point in the image according to the thermal value map after local contrast enhancement and nonlinear thermal value mapping; constructing a two-dimensional Gaussian distribution function around each thermal source center to generate a corresponding spatial thermal attention map, wherein the attention map gives different response weights to different positions in the image, thereby forming a continuous and smooth spatial weighted map; performing preprocessing on the visible light image, including: performing size and brightness standardization processing on the visible light image; performing illumination correction and local contrast enhancement on the normalized visible light image; performing edge extraction and detail enhancement on the illumination-corrected visible light image; converting the detail-enhanced visible light image to Lab space and enhancing the color contrast of a / b channels; A model construction and training module is configured to construct a target to be detected tracking model and train the model; The target to be detected tracking model comprises: A multi-modal fusion module is configured to fuse the preprocessed infrared image and the visible light image to obtain a fused image; The multi-modal fusion module comprises: A feature extraction and alignment unit is configured to extract multi-scale features of the preprocessed infrared image and the visible light image and perform spatial alignment; The feature extraction and alignment unit comprises an infrared branch and a visible light branch, the infrared branch is used for extracting thermal feature information in the preprocessed infrared image and obtaining an infrared feature map, and specifically comprises a thermal sensation guided attention module, a multi-scale pyramid residual extraction module and a deformation convolution enhancement layer, wherein the thermal sensation guided attention module generates a thermal sensation weight map by sequentially performing thermal value normalization, Gaussian smoothing and nonlinear stretching; the multi-scale pyramid residual extraction module adopts a four-branch structure and extracts different spatial scale features of heat sources by using 1*1, 3*3, 5*5 standard convolution and 3*3 dilated convolution respectively; the deformation convolution enhancement layer comprises an offset generation submodule and a deformation convolution submodule, wherein the offset generation submodule takes an enhanced thermal feature map under a unified scale as input, extracts local structures in the enhanced thermal feature map by using 3*3 standard convolution, and predicts a two-dimensional offset vector in the (x, y) direction for each convolution sampling point ; the deformation convolution submodule takes the two-dimensional offset vector and the enhanced thermal feature map under the unified scale as input, dynamically adjusts the sampling point position according to the two-dimensional offset vector at each convolution sampling point, then performs weighted convolution, and finally outputs a thermal feature enhancement map ; The visible light branch is configured to extract deep features of the preprocessed visible light image to obtain a visible light feature map, specifically including a guided edge enhancement module, a multi-channel frequency perception module and a color-structure alignment module, wherein the guided edge enhancement module performs gradient calculation on the preprocessed visible light image through a Sobel operator to extract a preliminary edge image; the multi-channel frequency perception module introduces a frequency domain analysis mechanism to mine discriminative features carried by different frequency components in the image; and the color-structure alignment module is configured to realize decoupled expression of color information and structural information and cross-modal alignment; An identification module is configured to input the preprocessed infrared image and the visible light image into the trained target to be detected tracking model to identify and track the position of the target to be detected.
6. A storage medium, characterized by It comprises instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1-4.
7. An electronic device, comprising: The electronic device comprises: a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the method of any one of claims 1-4 when executing the program. The electronic device comprises: a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the method of any one of claims 1-4 when executing the program.
Citation Information
Patent Citations
Bimodal target tracking method and device based on infrared and visible light images
CN117078719A