Multi-domain adaptive target tracking method and system based on domain customization adaptation

By using a text diffusion model and a candidate response optimal transmission mechanism, a unified multi-domain adaptive tracking framework is constructed, which solves the problem of unstable target tracking performance under adverse weather conditions, achieves efficient feature collaboration and layout consistency among multiple domains, and reduces deployment and maintenance costs.

CN121074091AActive Publication Date: 2025-12-05BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511246778.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-05
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing technologies have unstable target tracking performance under adverse weather conditions, require separate model training for each weather condition, and have high deployment and maintenance costs, and lack efficient interaction mechanisms between multiple domains.

Method used

A text diffusion model is used to synthesize data on various severe weather scenarios. A unified multi-domain adaptive tracking framework is constructed through a candidate response optimal transmission mechanism and a scenario adaptive mapping module to achieve cross-domain feature collaboration and layout consistency. The framework is combined with a teacher-student network to perform cross-domain knowledge transfer and specialized domain customization.

Benefits of technology

It maintains high-quality target status prediction under various severe weather conditions, reduces dependence on labeled data, improves tracking performance stability and efficiency, reduces computational complexity, and adapts to various complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074091A_ABST
    Figure CN121074091A_ABST
Patent Text Reader

Abstract

The invention provides a multi-domain adaptive target tracking method and system based on domain customization adaptation. The method comprises the steps of a data synthesis stage and a multi-weather universal adaptation stage. In the data synthesis stage, a preset text diffusion model is adopted to diffuse and output a plurality of target domain search frames; the multi-weather universal adaptation stage comprises the following steps of: constructing the plurality of target domain search frames into a first set, constructing the plurality of target domain search frames and the source domain search frames into a second set, and respectively extracting two frames corresponding to the same source domain search frame from the first set and the second set; inputting the frames extracted from the first set into a preset teacher network, and outputting a first positioning score graph by the teacher network; inputting the frames extracted from the second set into a preset student network, wherein the student network outputs a second positioning score graph; and calculating a first loss function based on the first positioning score map and the second positioning score map, and training a student model and a teacher model based on the first loss function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of target tracking, and in particular to a multi-domain adaptive target tracking method and system based on domain customization adaptation. BACKGROUND

[0002] In recent years, with the growing demand for practical applications such as autonomous driving and intelligent monitoring, visual object tracking (VOT) in adverse weather conditions has attracted widespread attention. However, the image quality under adverse weather conditions often decreases significantly, leading to a sharp decline in tracking performance based on visible light cameras. Traditional methods usually rely on multi-modal sensors such as visible light + depth (RGB-D) or visible light + thermal imaging (RGB-T) to learn cross-modal target representation through a large number of labeled samples, but the high cost of data acquisition limits its application range. To reduce the dependence on labeled data, some works attempt to use only RGB images for domain adaptation. The common approach is to unify the target representation through image enhancement or synthesis. For example, RGB and depth images are fused to generate a foggy scene, and feature alignment is performed on a Siamese tracker. HighlightNet excavates potential target features to improve UAV tracking performance in low-light scenes. UDAT migrates daytime semantic knowledge to the night domain through a Transformer bridge layer. Although these methods have achieved good results in a single adverse weather condition, the tracking performance still fluctuates greatly in complex scenes where multiple weather conditions coexist.

[0003] At the same time, controllable text-to-image generation technology (T2I) has made breakthrough progress. Early scene translation based on GAN requires training the model from scratch in a specific domain, with limited versatility. Methods such as GLIDE, DALL-E, and ControlNet based on diffusion models can generate high-quality images and edits with only text descriptions, without the need for large-scale labeling, providing a new paradigm for multi-weather data synthesis.

[0004] On the other hand, various techniques have emerged in the field of multi-target domain adaptation (MTDA): independent domain models and buffer merging to achieve knowledge fusion; graph matching and self-training strategies to enhance cross-domain generalization; unsupervised alignment methods based on optimal transport (OT), such as SOOD, to maintain cross-domain layout consistency through transport planning. However, how to build an integrated tracking adaptation framework for fog, night, rain, and other adverse weather conditions remains a difficult problem to be solved. Existing methods either require large-scale parameter training or updating for each weather condition, or lack an efficient interaction mechanism between multiple domains, resulting in high deployment and maintenance costs. SUMMARY

[0005] In view of this, the embodiment of the present application provides a multi-domain adaptive target tracking method and system based on domain customization adaptation to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present application provides a multi-domain adaptive target tracking method based on domain customization adaptation, the steps of the method include a data synthesis stage and a multi-weather universal adaptation stage; The steps of the data synthesis stage include: Input the source domain search frame in the training data into a preset text diffusion model, and the text diffusion model outputs a plurality of target domain search frames, each target domain search frame corresponding to a weather scene; The steps of the multi-weather universal adaptation stage include: The plurality of target domain search frames are constructed into a first set, and the plurality of target domain search frames and the source domain search frame are constructed into a second set, and two frames corresponding to the same source domain search frame are extracted from the first set and the second set respectively; Input the frame extracted from the first set into a preset teacher network, and the teacher network outputs a first positioning score map; input the frame extracted from the second set into a preset student network, and the student network outputs a second positioning score map; Construct a mask map based on the first positioning score map, and determine a first response position map and a second response position map based on the mask map corresponding to the first positioning score map and the second positioning score map respectively; Determine a cost matrix based on the first response position map and the second response position map, calculate a first loss function based on the cost matrix, and train the student model and the teacher model based on the first loss function.

[0007] With the above scheme, a unified multi-domain adaptive tracking framework is proposed, which integrates target tracking problems in different domains, such as sunny, rainy, foggy, night and other complex weather and scene conditions, into the same model architecture for processing, avoiding the overhead of training a model for each scene in the traditional method. Through the three-stage process of pre-training, cross-domain adaptation and scene customization, a full-link design from general feature learning to specialized scene optimization is realized, and high-quality target state prediction is maintained under various severe weather conditions.

[0008] In some embodiments of the present application, in the step of constructing a mask map based on the first positioning score map, each position in the first positioning score map is determined based on a preset confidence threshold, the each position is determined to be a first mask value or a second mask value, and a mask map is obtained.

[0009] In some embodiments of the present application, in the step of determining the first response position map and the second response position map based on the mask map corresponding to the first positioning score map and the second positioning score map respectively, normalizing the first positioning score map and the second positioning score map respectively to obtain the first probability distribution map and the second probability distribution map; processing the first probability distribution map and the second probability distribution map based on the mask map respectively to obtain the first response position map and the second response position map.

[0010] In some embodiments of the present application, in the step of determining the cost matrix based on the first response position map and the second response position map, the positions in the first positioning score map, the second positioning score map, the first response position map and the second response position map are sequentially numbered respectively, and the value of each position in the cost matrix is calculated using the following formula: wherein C i,j represents the value of the position with horizontal coordinate i and vertical coordinate j in the cost matrix, represents the value of the position numbered i in the second response position map, represents the value of the position numbered j in the first response position map, represents the value of the position numbered i in the second positioning score map, represents the value of the position numbered j in the first positioning score map, and Z represents a normalization factor and a is a preset weight parameter.

[0011] In some embodiments of the present application, in the step of calculating the first loss function based on the cost matrix, the first loss function is calculated using the following formula: L OT = min<C, P> - eH(P); wherein L OT represents the value of the first loss function, C represents the cost matrix, P represents the optimal transport plan matrix solved by the Sinkhorn algorithm, e represents a preset calculation parameter, and min<C, P> represents the product of the value of one position in the two matrices C and P respectively, and the minimum value in all products.

[0012] By using the above scheme, the response distribution of the source domain and the target domain is constructed, and the optimal transport matching is performed in the high-confidence candidate region, so that the distribution of the features of the source domain and the target domain in the target region remains consistent, and the Sinkhorn algorithm is used for efficient solution through entropy regularization.

[0013] In some embodiments of the present application, the steps of the method further include a specialized domain customization stage, and the steps of the specialized domain customization stage include: inputting the target domain search frame corresponding to the source domain search frame in the above step into the student network, and outputting a process vector by a self-attention mechanism module of the student network; inputting the target domain search frame corresponding to the source domain search frame in the above step into the student network, and outputting a process vector by a self-attention mechanism module of the student network; inputting the score vector and the process vector into a cross-attention module of the student network, and outputting a first output vector corresponding to the prompt score vector and a second output vector corresponding to the process vector by the cross-attention module; determining a corresponding pooling conversion vector based on the first output vector, and determining a feature vector based on the pooling conversion vector and the second output vector, wherein each dimension of the feature vector corresponds to a picture area in the target domain search frame; determining a target area in the target domain search frame based on the source domain template frame; determining a value of a picture area corresponding to a center position of the target area in the feature vector as a positive sample; screening values of picture areas outside the target area in the feature vector corresponding dimensions to obtain a preset number of values in the feature vector as a negative sample set; calculating a second loss function based on the negative sample set and the positive sample, and training the student network based on the second loss function.

[0014] In some embodiments of the present application, each dimension in the second output vector corresponds to a picture area in the target domain search frame, and a corresponding area vector is arranged in the dimension, and in the step of determining the feature vector based on the pooling conversion vector and the second output vector, a cosine similarity of the pooling conversion vector and each area vector in the second output vector is calculated to obtain the feature vector.

[0015] In some embodiments of the present application, in the step of screening values of picture areas outside the target area in the feature vector corresponding dimensions to obtain a preset number of values in the feature vector as a negative sample set, values of the picture areas outside the target area in the feature vector corresponding dimensions are screened, and a preset number of maximum values are obtained as the negative sample set.

[0016] In some embodiments of the present application, in the step of calculating the second loss function based on the negative sample set and the positive sample, the second loss function is calculated by the following formula: wherein, L cont represents a value of the second loss function, S + represents the positive sample, represents a value of an i-th Δ sample in the negative sample set, and q Δdenotes the feature value of the th sample in the negative sample set, and β is a preset calculation parameter.

[0017] The second aspect of the present application also provides a multi-domain adaptive target tracking system based on domain customization adaptation, which comprises a computer device, the computer device comprising a processor and a memory, the memory storing computer instructions, and the processor being configured to execute the computer instructions stored in the memory, so as to realize the steps of the method as described above.

[0018] The third aspect of the present application also provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to realize the steps of the multi-domain adaptive target tracking method based on domain customization adaptation as described above.

[0019] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and will become apparent to those skilled in the art upon examination of the following detailed description and drawings in which

[0020] Those skilled in the art will appreciate that the objects and advantages of the application can not be limited to the specifically described above, and the above and other objects that can be achieved by the application will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0021] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments of the application and, together with the description, serve to explain the principles of the application.

[0022] Figure 1 A schematic diagram of an embodiment of the multi-domain adaptive target tracking method based on domain customization adaptation of the present application; Figure 2 A schematic diagram of the steps of the specialized domain customization stage of the multi-domain adaptive target tracking method based on domain customization adaptation of the present application; Figure 3 A comparison chart of tracking effects in different weathers of the present application; Figure 4 A schematic diagram of the execution architecture of an embodiment of the present application; Figure 5 A schematic diagram of the execution architecture of the specialized domain customization stage of the present application. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed explanations will be given below in conjunction with the embodiments and drawings. Herein, the illustrative embodiments of the present application and their explanations are used to explain the present application but not as limitations to the present application.

[0024] It is also needed to be explained herein that, in order not to obscure the present application with unnecessary details, only the structures and / or processing steps closely related to the solutions according to the present application are shown in the drawings, and other details not closely related to the present application are omitted.

[0025] As shown in Figure 1 and 3 The present application proposes a multi-domain adaptive target tracking method based on domain customization adaptation, the steps of which include a data synthesis stage and a multi-weather universal adaptation stage. The steps of the data synthesis stage include: Step S110, input the source domain search frame in the training data to a preset text diffusion model, the text diffusion model outputs a plurality of target domain search frames, each target domain search frame corresponds to a weather scene; In the specific implementation process, the text diffusion model adopts a Stable Diffusion Turbo model; In the data synthesis stage, the existing source domain data (such as video data set under normal daytime weather) is used for initial training of the target tracking model. At the same time, in order to make up for the problem of insufficient training samples under severe weather conditions, part of the pictures of the source domain are input into the text diffusion model based on Stable Diffusion Turbo, and the specified source domain clean pictures are translated into images by using text prompts, to generate target severe weather samples highly consistent with the visual distribution of real scenes.

[0026] Traditional visual tracking methods often rely on multi-modal sensors such as RGB-D and RGB-T, which have high hardware purchase and labeling costs, are not conducive to large-scale deployment, and multi-modal data acquisition is limited by the environment and equipment, with insufficient available labeled samples and difficulty in covering multiple severe weather scenes. Although some methods try to synthesize or enhance unlabeled data, a large number of sunny or single-domain labeled samples are still needed for pre-training or calibration, which reduces the original intention of reducing the dependence on labeling of adaptive methods, and there is a gap between the synthesized data and the real data in distribution, which still needs manual fine-tuning after direct migration; Domain adaptation methods based on image enhancement or synthesis are mostly designed for a single severe weather (such as only foggy, night or rainy), and lack overall adaptation strategies when multiple weather conditions occur at the same time. Deploying a single domain adaptive model trained in a complex environment with multiple weather conditions will cause the tracking performance to fluctuate sharply and have poor stability. Although existing unsupervised alignment (such as optimal transport) or self-training strategies can enhance inter-domain consistency, they are often computationally complex and have large parameter sizes, which are not suitable for real-time tracking scenarios, while independent domain models or buffer fusion are lightweight but difficult to guarantee feature coordination and layout consistency between different weather domains. The scheme diffuses the original data through a text diffusion model, so that the feature coordination and layout consistency between different weather domains are consistent, the data stability is ensured, and the data annotation amount is reduced. In view of the problems of high cost, difficulty in popularization of the above-mentioned multi-modal dependence, and decline of tracking performance caused by multi-domain migration, a unified multi-domain adaptive target tracking method based on a candidate response optimal transport mechanism and a scene adaptive mapping module is proposed by using a diffusion model. First, the diffusion model of text-to-image translation is used to automatically synthesize videos or images of multiple types of severe weather scenes, significantly reducing the dependence on real labeled data and enriching the diversity of training samples. Then, the candidate response optimal transport mechanism is introduced to realize the integrated optimization of the source domain and the target domain tracking results, achieving universal alignment effect in all weather conditions. Finally, according to the actual needs of different scenes, the scene adaptive mapping module is introduced to realize fast domain adaptation. The method achieves the best level of tracking in multiple weather conditions.

[0027] The method further includes a pre-training phase, in which the student model is initially trained using source domain search frames.

[0028] As shown in Figure 4 The steps of the multi-weather universal adaptation stage include: Step S210, a plurality of target domain search frames are constructed into a first set, and a plurality of target domain search frames and source domain search frames are constructed into a second set, and two frames corresponding to the same source domain search frame are extracted from the first set and the second set, respectively. Step S220, the frames extracted from the first set are input into a preset teacher network, and the teacher network outputs a first positioning score map; the frames extracted from the second set are input into a preset student network, and the student network outputs a second positioning score map; step S230, a mask map is constructed based on the first positioning score map, and a first response position map and a second response position map are determined based on the mask map corresponding to the first positioning score map and the second positioning score map, respectively. In the specific implementation process, the high confidence area is screened, and the mask map guided by the teacher is defined. For any position of the first positioning score map, if it is greater than or equal to the threshold τ, it is assigned a value of 1, otherwise a value of 0. In the step of determining the first response position map and the second response position map based on the mask map corresponding to the first positioning score map and the second positioning score map, respectively, the high response position is further screened by using the following formula: wherein, represents element-wise multiplication, M represents a mask map, represents a second positioning score map, represents a first positioning score map, represents a second response position map, represents a first response position map.

[0029] In step S240, a cost matrix is determined based on the first response position map and the second response position map, a first loss function is calculated based on the cost matrix, and the student model and the teacher model are trained based on the first loss function.

[0030] In the multi-weather universal adaptation stage, the teacher-student network structure is used for cross-domain knowledge transfer, so that the model can maintain strong generalization ability under different weather conditions. The synthesized target domain data and the source domain data are input into the teacher-student structure at the same time, the teacher network parameters are updated through the exponential moving average (EMA), and the progressive transfer from the source domain to multiple target domains is realized; then the candidate response optimal transport mechanism is introduced, and the optimal transport method is used to align the response confidence distribution of the teacher and the student in different domains, especially the consistency constraint is performed on the high positioning score area, so that the cross-domain positioning accuracy is improved, and the high-precision tracking of global samples is realized, and a unified multi-domain adaptive tracking model is obtained.

[0031] The scheme can dynamically adjust the interpretation of the network to the input features, and efficiently focus on the target area through the contrast learning method, so that the model maintains stable performance when facing various weather and scene changes. A large number of experiments show that the algorithm used in the present scheme performs excellently in tracking positioning consistency, and its performance is significantly better than that of the existing advanced method.

[0032] By using the above scheme, the present scheme proposes a unified multi-domain adaptive tracking framework, which integrates target tracking problems in different domains, such as sunny, rainy, foggy, night and other complex weather and scene conditions, into the same model architecture for processing, avoiding the overhead of training a model for each scene in the traditional method. Through the pre-training, cross-domain adaptation and scene customization three-stage process, the full-link design from general feature learning to specialized scene optimization is realized, and high-quality target state prediction is maintained under various severe weather conditions.

[0033] In some embodiments of the present application, in the step of constructing a mask map based on the first positioning score map, each position in the first positioning score map is judged based on a preset confidence threshold, the each position is determined to be a first mask value or a second mask value, and the mask map is obtained.

[0034] In some embodiments of the present application, in the step of determining the first response position map and the second response position map based on the mask map corresponding to the first positioning score map and the second positioning score map respectively, normalizing the first positioning score map and the second positioning score map to obtain the first probability distribution map and the second probability distribution map respectively; Specifically, for the positioning output of different tracking heads of the student network and the teacher network, the second probability distribution map S of the student network is obtained S ∈R H×W , and the first probability distribution map S of the teacher network is obtained T ∈R H×W The first positioning score map and the second positioning score map are normalized by a softmax function to obtain a probability distribution:

[0035] The first probability distribution map and the second probability distribution map are processed based on the mask map to obtain the first response position map and the second response position map.

[0036] In some embodiments of the present application, in the step of determining the cost matrix based on the first response position map and the second response position map, the positions in the first positioning score map, the second positioning score map, the first response position map and the second response position map are sequentially numbered respectively, and the value of each position in the cost matrix is calculated using the following formula: Wherein, C i,j represents the value of the position with horizontal coordinate i and vertical coordinate j in the cost matrix, represents the value of the number i in the second response position map, represents the value of the number j in the first response position map, represents the value of the number i in the second positioning score map, represents the value of the number j in the first positioning score map, and Z represents a normalization factor and a is a preset weight parameter.

[0037] In some embodiments of the present application, in the step of calculating the first loss function based on the cost matrix, the first loss function is calculated using the following formula: L OT = min<C, P> - εH(P); Wherein, L OT represents the value of the first loss function, C represents the cost matrix, P represents the optimal transport plan matrix solved by the Sinkhorn algorithm, ε represents a preset calculation parameter, and min<C, P> represents the product of the value of one position in the two matrices C and P respectively, and the minimum value in all products.

[0038] In practical implementation, the Sinkhorn algorithm can efficiently solve the optimal transmission plan, thereby minimizing the transmission cost and achieving consistent alignment between the source and target domains, thus significantly improving the consistency and robustness of cross-domain positioning.

[0039] For the optimal transmission plan P, the following closed-form expression holds: Where vector u∈R H ,v∈R W The marginal constraints are satisfied by iteratively solving using the Sinkhorn algorithm. P1 n =a,P T 1 n =b This minimizes transmission costs, ensures consistent alignment between the source and target domains, and significantly improves the consistency and robustness of cross-domain positioning.

[0040] like Figure 2 and 5 As shown, in some embodiments of the present invention, the method further includes a specialized domain customization stage, the steps of which include: Step S310: Input the target domain search frame corresponding to the special domain into the preprocessing module to obtain the visual cue vector; Step S320: Input the source domain search frame corresponding to the target domain search frame in step S310 into the student network, and the self-attention mechanism module of the student network outputs the process vector. Specifically, the preprocessing module employs a lightweight residual network to perform linear transformations on the features of the target search region and reshape them using multiple downsampling layers. It then obtains the corresponding prompt score vector score∈R through global average pooling and the softmax function. K×L′ Meanwhile, a library of learnable tokens, B∈R, initialized from Gaussian random variables... L′×C Starting from this point, it is mapped to the value V∈R through a fully connected layer. L′×C Thus, we obtain the feature preparation for query scores and key-value branches.

[0041] The cue score vector is then multiplied by V to obtain the structured visual cue vector T∈R. K×C This vector hint incorporates unique feature information from the target scene, can encode the potential spatial structure of the target area under different weather conditions, and shares contextual similarity with the source domain template features in the embedding space.

[0042] Step S330, input the score vector and the process vector into a cross attention module of the student network, the cross attention module outputs a first output vector corresponding to the prompt score vector and a second output vector corresponding to the process vector; Step S340, determine a corresponding pooling conversion vector based on the first output vector, determine a feature vector based on the pooling conversion vector and the second output vector, each dimension of the feature vector corresponds to a picture area in the target domain search frame; Step S350, determine a target area in the target domain search frame based on the source domain template frame; Step S360, determine the value of the picture area of the center position of the target area in the corresponding dimension of the feature vector as a positive sample; Step S370, screen the values of the picture areas outside the target area in the corresponding dimensions of the feature vector, and screen a preset number of values of the dimensions in the feature vector as a negative sample set; Step S380, calculate a second loss function based on the negative sample set and the positive sample, and train the student network based on the second loss function.

[0043] In order to make the obtained visual prompt contain not only global visual features, an instance-level feature representation is used, so that based on contrast learning, the feature information of the foreground and the background is extracted according to the position of the annotation box, and contrast learning is performed.

[0044] The teacher model and the student model have the same structure, and the model parameters of the student model are synchronized to the teacher model every preset time interval.

[0045] By using the above scheme, in the special domain customization stage, the scheme further optimizes the model for specific severe weather scenes. We freeze the backbone feature extraction network, insert a scene adaptive mapping module based on contrast learning, and integrate visual information of specific scene weather into the network through contrast learning. Through contrast learning and a learnable token mechanism, the invariant domain features in specific severe weather scenes are fully learned, and the model can further improve the precision and robustness under specific target domain conditions, realizing fast migration and fine optimization.

[0046] In the special domain customization stage, according to the feature differences of different scenes (such as rainy days, foggy days, night, etc.), the general features extracted by the backbone network are mapped and adjusted, the visual vector prompt words are inserted to integrate additional information, and contrast learning is used, so that the integrated visual prompt can distinguish the foreground and the background, efficiently search the target area, and thus the model can maintain high target tracking precision in various complex environments, such as Figure 3 as shown.

[0047] In some embodiments of the present application, each dimension in the second output vector corresponds to a picture region in the target domain search frame, and a corresponding region vector is arranged in the dimension. In the step of determining the feature vector based on the pooling conversion vector and the second output vector, the cosine similarity between the pooling conversion vector and each region vector in the second output vector is calculated to obtain the feature vector.

[0048] Given the output prompt word of the cross-attention layer, it is subjected to a global average pooling to obtain an aggregated pooling conversion vector, and the cosine similarity is calculated with the feature patches of the search graph.

[0049] In some embodiments of the present application, in the step of screening the values of the picture regions outside the target region in the corresponding dimensions of the feature vector, a preset number of values in the feature vector are screened as the negative sample set.

[0050] In some embodiments of the present application, in the step of calculating the second loss function based on the negative sample set and the positive sample, the second loss function is calculated using the following formula: wherein, L cont represents the value of the second loss function, S + represents the positive sample, represents the value of the th Δ sample in the negative sample set, q Δ represents the feature value of the th Δ sample in the negative sample set, and β is a preset calculation parameter.

[0051] wherein, l represents a calculation parameter, u represents any negative sample in the negative sample set, S represents the negative sample set, and S u the value of the th u sample in the negative sample set.

[0052] By using the above scheme, the positive and negative sample scores are compared and learned, the foreground features and the background features are forced to be significantly distinguished in the feature space, and finally the similarity score of the target region is always higher than that of the background region, thereby effectively improving the accuracy of target detection and tracking. At the same time, the entropy regularization constraint is introduced at the distribution level to avoid the output distribution of the model being excessively sharp during the training process, thereby reducing the risk of model collapse caused by abnormal samples or pseudo-label noise, so that the performance of the model in a complex environment is more stable.

[0053] In the specific implementation process, the model trained through the multi-weather universal adaptation stage or the special domain customization stage is used for target tracking.

[0054] The beneficial effects of the present solution include: 1. The present solution distinguishes from the existing methods of single training or single domain adaptation, and designs a three-stage tracking process from data preparation to real-time tracking, which solves the three core problems of sample scarcity, cross-domain generalization and rapid migration. The present application uses a pre-trained diffusion model to synthesize a small amount of samples, inputs mixed data into the teacher-student network, uses the probability response map output by the network to constrain and align the high-probability matching area, and realizes efficient tracking in all-weather. And additional optimization for specific weather scenarios, introducing a scene mapping mechanism based on contrast learning, separates foreground and background, allowing the model to learn invariant domain features in the same scene, enabling rapid and refined fine-tuning to meet different scene requirements. A large number of experiments on multiple severe weather tracking benchmarks show that the present framework outperforms existing advanced methods in accuracy and robustness, significantly improving cross-domain tracking performance.

[0055] 2. The present solution aims to establish a positioning information bridge between the source domain and the target domain, align the response patterns of the source domain and the target domain on the positioning score map, and realize the positioning score Figure 1 consistency between the source domain and the target domain, align the response distribution of the teacher network and the student network in the high-confidence area through the optimal transport method, thereby reducing the noise interference of pseudo-labels while allowing the model to focus on reliable prediction areas during training, improving the cross-domain robustness and positioning accuracy in the target tracking process.

[0056] 3. Recent research has attempted to expand the camouflage target samples through GAN or simple affine transformation, but the generation of diversified and quality-controlled camouflage scenes is still insufficient. The method of the present application further integrates a data synthesis strategy based on a diffusion model, which uses text descriptions to synthesize samples from the source domain to the target domain, allowing the model to be trained on a more diverse and challenging distribution, thereby significantly improving the robustness and generalization ability of tracking in multiple severe weather scenarios.

[0057] 4. The multi-domain adaptive tracking framework proposed in the present application can handle target tracking tasks in different domains (including different weather, lighting and scene conditions) under a unified model architecture, avoiding the cumbersome process of designing or training models for different domains in traditional methods, improving the consistency and scalability of the system. At the same time, it supports special domain customization, which can adapt and adapt flexibly to the specific domain conditions required in actual work, reducing the training and deployment cost, and has higher computational efficiency compared to the method of retraining the entire network, and achieves the best effect in a specific domain.

[0058] 5. Traditional pseudo-label tracking often ignores the positioning errors caused by high-confidence false detections and spatial drift, leading to model mis-convergence. The method proposed in this application introduces a candidate response optimal transport mechanism, constructs a confidence and position dual cost matrix based on optimal transport, and fine-aligns the source domain and target domain output distributions, effectively filtering pseudo-label noise and significantly improving cross-domain positioning consistency, significantly enhancing tracking robustness in complex environments.

[0059] The embodiment of the application also provides a multi-domain adaptive target tracking system based on domain customization adaptation, which comprises a computer device, the computer device comprising a processor and a memory, the memory storing computer instructions, and the processor being configured to execute the computer instructions stored in the memory, so that the system implements the steps implemented by the method described above.

[0060] The embodiment of the application also provides a computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps implemented by the multi-domain adaptive target tracking method based on domain customization adaptation described above. The computer readable storage medium can be a tangible storage medium, such as random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.

[0061] Those skilled in the art should understand that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of both. Whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link.

[0062] It is to be expressly understood that the invention is not limited to the specific configurations and process described above and illustrated in the accompanying drawings. For the sake of clarity, detailed descriptions of known methods are omitted. In the above-described embodiments, several specific steps are described and illustrated as examples. However, the method processes of the present invention are not limited to the specific steps described and illustrated, and various changes, modifications and additions can be made thereto by one of ordinary skill in the art without departing from the spirit of the present invention, and the order of the steps can be changed.

[0063] In the present invention, features described and / or illustrated with respect to one embodiment can be used in the same or a similar way in one or more other embodiments, and / or in combination with or instead of features of other embodiments.

[0064] The above description is merely illustrative of the application, and is not intended to limit the scope of the application. Various modifications and changes can be made by one of ordinary skill in the art without departing from the spirit and scope of the application. Any modification, equivalent replacement, improvement, and the like made within the spirit and principle of the application should be included in the scope of the application.

Claims

1. A multi-domain adaptive target tracking method based on domain customized adaptation, characterized in that, The steps of the method include a data synthesis stage and a multi-weather universal adaptation stage. The steps of the data synthesis stage include: inputting a source domain search frame in training data into a preset text diffusion model, the text diffusion model outputting a plurality of target domain search frames, each target domain search frame corresponding to a weather scene; The steps of the multi-weather universal adaptation stage include: constructing a plurality of target domain search frames into a first set, and constructing a plurality of target domain search frames and source domain search frames into a second set, and extracting two frames corresponding to the same source domain search frame from the first set and the second set respectively; inputting the frame extracted from the first set into a preset teacher network, the teacher network outputting a first positioning score map; inputting the frame extracted from the second set into a preset student network, the student network outputting a second positioning score map; based on the first positioning score map, a mask map is constructed, and based on the mask map, a first response position map and a second response position map are determined respectively corresponding to the first positioning score map and the second positioning score map; based on the first response position map and the second response position map, a cost matrix is determined, based on the cost matrix, a first loss function is calculated, and based on the first loss function, a student model and a teacher model are trained.

2. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 1, characterized in that, In the step of constructing a mask map based on the first positioning score map, each position in the first positioning score map is judged based on a preset confidence threshold, the each position is judged to be a first mask value or a second mask value, and the mask map is obtained. 3.The multi-domain adaptive target tracking method based on domain customization adaptation of claim 1, wherein, In the step of determining the first response position map and the second response position map based on the mask map corresponding to the first positioning score map and the second positioning score map respectively, the first positioning score map and the second positioning score map are normalized respectively to obtain a first probability distribution map and a second probability distribution map; based on the mask map, the first probability distribution map and the second probability distribution map are processed respectively to obtain the first response position map and the second response position map.

4. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 3, characterized in that, In the step of determining the cost matrix based on the first response position map and the second response position map, the positions in the first positioning score map, the second positioning score map, the first response position map and the second response position map are sequentially numbered respectively, and the value of each position in the cost matrix is calculated using the following formula: where C i,j denotes the value at position i in the horizontal coordinate and j in the vertical coordinate of the cost matrix, denotes the value at index i in the second response position map, denotes the value at index j in the first response position map, denotes the value at index i in the second positioning score map, denotes the value at index j in the first positioning score map, and Z denotes a normalization factor, and a denotes a predetermined weight parameter.

5. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 1, characterized in that, In the step of calculating the first loss function based on the cost matrix, the first loss function is calculated using the following formula: L OT = min < C, P > - εH(P); wherein L OT represents the value of the first loss function, C represents the cost matrix, P represents the optimal transport plan matrix solved by the Sinkhorn algorithm, ε represents a preset calculation parameter, min < C, P > represents the value of one position in the two matrices C and P respectively, and the minimum value in all products is selected.

6. The multi-domain adaptive target tracking method based on domain customized adaptation according to any one of claims 1-5, characterized in that, The steps of the method also include a specialized domain customization stage, and the steps of the specialized domain customization stage include: inputting a target domain search frame corresponding to a specialized domain into a pre-processing module to obtain a visual prompt vector; inputting the target domain search frame corresponding to the source domain search frame in the above step into the student network, and the self-attention mechanism module of the student network outputs a process vector; inputting the score vector and the process vector into the cross-attention module of the student network, the cross-attention module outputs a first output vector corresponding to the prompt score vector and a second output vector corresponding to the process vector; determining a corresponding pooling conversion vector based on the first output vector, determining a feature vector based on the pooling conversion vector and a second output vector, each dimension of the feature vector corresponding to a picture region in a target domain search frame; determining a target region in the target domain search frame based on a source domain template frame; determining a value of a picture region corresponding to a center position of the target region in the feature vector as a positive sample; screening values of picture regions outside the target region in the feature vector to obtain a preset number of values in the feature vector as a negative sample set; calculating a second loss function based on the negative sample set and the positive sample, and training the student network based on the second loss function.

7. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 6, characterized in that, Each dimension in the second output vector corresponds to a picture region in the target domain search frame, and a corresponding region vector is set in the dimension. In the step of determining the feature vector based on the pooling conversion vector and the second output vector, a cosine similarity of the pooling conversion vector and each region vector in the second output vector is calculated to obtain the feature vector.

8. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 6, characterized in that, In the step of screening values of picture regions outside the target region in the feature vector to obtain a preset number of values in the feature vector as a negative sample set, values of picture regions outside the target region in the feature vector are screened to obtain a preset number of maximum values as the negative sample set.

9. The multi-domain adaptive target tracking method based on domain customized adaptation according to claim 6, characterized in that, In the step of calculating the second loss function based on the negative sample set and the positive sample, the second loss function is calculated using the following formula: wherein, L cont represents a value of the second loss function, S + represents a positive sample, represents a value of the th Δ sample in the negative sample set, q Δ represents a feature value of the th Δ sample in the negative sample set, and β is a preset calculation parameter.

10. A multi-domain adaptive target tracking system based on domain customized adaptation, characterized in that, The system comprises a computer device, the computer device comprising a processor and a memory, the memory storing computer instructions, and the processor being configured to execute the computer instructions stored in the memory, so that the system implements the steps of the method according to any one of claims 1 to 9. determining a corresponding pooling conversion vector based on the first output vector, determining a feature vector based on the pooling conversion vector and a second output vector, each dimension of the feature vector corresponding to a picture region in a target domain search frame; determining a target region in the target domain search frame based on a source domain template frame; determining a value of a picture region corresponding to a center position of the target region in the feature vector as a positive sample; screening values of picture regions outside the target region in the feature vector to obtain a preset number of values in the feature vector as a negative sample set; calculating a second loss function based on the negative sample set and the positive sample, and training the student network based on the second loss function. Each dimension in the second output vector corresponds to a picture region in the target domain search frame, and a corresponding region vector is set in the dimension. In the step of determining the feature vector based on the pooling conversion vector and the second output vector, a cosine similarity of the pooling conversion vector and each region vector in the second output vector is calculated to obtain the feature vector. In the step of screening values of picture regions outside the target region in the feature vector to obtain a preset number of values in the feature vector as a negative sample set, values of picture regions outside the target region in the feature vector are screened to obtain a preset number of maximum values as the negative sample set. In the step of calculating the second loss function based on the negative sample set and the positive sample, the second loss function is calculated using the following formula: The system comprises a computer device, the computer device comprising a processor and a memory, the memory storing computer instructions, and the processor being configured to execute the computer instructions stored in the memory, so that the system implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Analysis device, method, and program

    WO2024247192A1