Visible light infrared target tracking training method and device based on unpaired data
By using unpaired data training methods in visible infrared target tracking, modal specific network, modal sharing module, modal adaptive attention module and modal adaptation module, the problem of insufficient data volume is solved and robust tracking performance is achieved.
Patent Information
- Application Number
- CN202210095429.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing visible infrared target tracking algorithms rely on large-scale registration data, resulting in insufficient data volume.
A visible light infrared target tracking training method based on unpaired data is proposed. By setting a modal specific network, a modal sharing module, a modal adaptive attention module and a modal adaptation module, learning and enhancement between unpaired visible light data and thermal infrared data modes is achieved.
It effectively avoids the problem of insufficient data amount required for training, fully mines and utilizes dual-modal information on a limited data set, realizes mutual enhancement between unpaired multimodal data, and trains a robust visible light infrared tracker.
Smart Images

Figure CN114445461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a visible light infrared target tracking training method and device based on unpaired data. Background Art
[0002] Object tracking has always been an important topic in the field of computer vision. In recent years, object tracking technology has made great breakthroughs and is widely used in the fields of intelligent transportation, unmanned driving, and robotics. The task of object tracking is to predict the size and position of the object in the subsequent frames given the size and position of the object in the initial frame of a video sequence.
[0003] Most target tracking algorithms are based on a single visible light modality and have superior tracking performance under this modality. However, in complex environments or extreme conditions, such as haze weather and low light, the robustness of tracking still needs to be improved. In recent years, more and more sensor technologies have been applied to the field of target tracking, such as thermal infrared sensors and depth sensors. Thermal infrared sensors form images by capturing the temperature information of the target and have low sensitivity to lighting conditions. At the same time, visible light data can make up for the shortcomings of thermal infrared images, such as blurred edges and less detailed information. The complementarity of visible light data and thermal infrared data can help the algorithm achieve robust tracking.
[0004] At present, the research focus of most visible light infrared tracking algorithms is to complement the advantages of multiple modalities and fuse the information of each modality to achieve more robust tracking results. However, visible light infrared tracking algorithms often use matching visible light data and thermal infrared data. For example, the invention patent application with application number 201510521038.8 discloses a real-time all-weather target tracking method based on visible light and infrared images, which requires the registration of visible light images and infrared images before target tracking and detection. Thermal infrared data can be obtained through thermal infrared sensors, but the matching visible light thermal infrared data requires a lot of manual selection and manual annotation, which brings certain challenges to the production of data sets. The number of public standard visible light thermal infrared data sets is small, and the advantages of thermal infrared modalities cannot be fully utilized. Summary of the invention
[0005] The technical problem to be solved by the present invention is how to solve the problem of visible light infrared target tracking performance training relying on large-scale registration data.
[0006] The present invention solves the above technical problems through the following technical means:
[0007] On the one hand, an embodiment of the present invention provides a visible light infrared target tracking training method based on unpaired data, the method comprising the following steps:
[0008] Acquire an unpaired visible light image and a thermal infrared image, and generate candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples;
[0009] Using the candidate samples to train a visible light infrared tracker;
[0010] Among them, the visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image serves as the input of the first modality-specific network and the input of the modality sharing module, and the thermal infrared image serves as the input of the second modality-specific network and the input of the modality sharing module. The visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module.
[0011] The present invention sets a first modality specific network and a second modality specific network to respectively extract features of visible light images and thermal infrared images, sets a modality sharing module to extract similar features of visible light data and thermal infrared data, further strengthens the connection between modalities, sets a modality adaptive attention module to achieve learning and enhancement between unpaired visible light data and thermal infrared data modalities, liberates the power of unpaired visible light infrared data, effectively avoids the problem of insufficient data required for training, fully mines and utilizes dual-modal information on a limited data set, achieves mutual enhancement between unpaired multimodal data, and trains a robust visible light infrared tracker.
[0012] Furthermore, the first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as inputs of the last two convolutional layers of the modality sharing network;
[0013] The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
[0014] Further, the modality-adaptive attention module includes a first fully-connected layer and a second fully-connected layer with shared weights, a third fully-connected layer and a fourth fully-connected layer specific to the modality, and a fifth fully-connected layer shared by the modality;
[0015] The visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimensionality reduction processing is performed to obtain the visible light modal features and the thermal infrared modal features after dimensionality reduction;
[0016] The third fully connected layer and the fourth fully connected layer respectively process the reduced-dimensional visible light modal features and the reduced-dimensional thermal infrared modal features using a QKV mechanism to form attention matrices corresponding to the two modalities;
[0017] The attention matrices corresponding to the two modalities are passed through the fifth fully connected layer to form a modality shared query set;
[0018] After multiplying the modality shared query set with the attention matrices corresponding to the two modalities respectively, enhanced feature maps corresponding to the two modalities are obtained.
[0019] Further, the modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities;
[0020] A random activation function of neurons is added after the first two fully connected layers connected in sequence;
[0021] The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
[0022] Furthermore, the method further comprises:
[0023] The visible light infrared tracker is trained using the cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is:
[0024] Loss=-(y*log(y^)+(1-y)*log(1-y^));
[0025] The overall network of the visible light infrared tracker is optimized by stochastic gradient descent method.
[0026] On the other hand, the present invention provides a visible light infrared target tracking training device based on unpaired data, the device comprising:
[0027] An acquisition module, used to acquire an unpaired visible light image and a thermal infrared image, and generate candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples;
[0028] A training module, used for training a visible light infrared tracker using the candidate samples;
[0029] Among them, the visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image serves as the input of the first modality-specific network and the input of the modality sharing module, and the thermal infrared image serves as the input of the second modality-specific network and the input of the modality sharing module. The visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module.
[0030] Furthermore, the first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as inputs of the last two convolutional layers of the modality sharing network;
[0031] The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
[0032] Further, the modality-adaptive attention module includes a first fully-connected layer and a second fully-connected layer with shared weights, a third fully-connected layer and a fourth fully-connected layer specific to the modality, and a fifth fully-connected layer shared by the modality;
[0033] The visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimensionality reduction processing is performed to obtain the visible light modal features and the thermal infrared modal features after dimensionality reduction;
[0034] The third fully connected layer and the fourth fully connected layer respectively process the reduced-dimensional visible light modal features and the reduced-dimensional thermal infrared modal features using a QKV mechanism to form attention matrices corresponding to the two modalities;
[0035] The attention matrices corresponding to the two modalities are passed through the fifth fully connected layer to form a modality shared query set;
[0036] After multiplying the modality shared query set with the attention matrices corresponding to the two modalities respectively, enhanced feature maps corresponding to the two modalities are obtained.
[0037] Further, the modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities;
[0038] A random activation function of neurons is added after the first two fully connected layers connected in sequence;
[0039] The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
[0040] Furthermore, the training module includes:
[0041] A training unit is used to train the visible light infrared tracker using a cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is:
[0042] Loss=-(y*log(y^)+(1-y)*log(1-y^));
[0043] The optimization unit is used to optimize the overall network of the visible light infrared tracker by using a stochastic gradient descent method.
[0044] The advantages of the present invention are:
[0045] (1) The present invention sets a first modality-specific network and a second modality-specific network to respectively extract features of visible light images and thermal infrared images, sets a modality sharing module to extract similar features of visible light data and thermal infrared data, further strengthens the connection between modalities, sets a modality-adaptive attention module to achieve learning and enhancement between unpaired visible light data and thermal infrared data modalities, liberates the power of unpaired visible light infrared data, effectively avoids the problem of insufficient data required for training, fully mines and utilizes dual-modal information on a limited data set, achieves mutual enhancement between unpaired multimodal data, and trains a robust visible light infrared tracker.
[0046] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of a visible light infrared target tracking training method based on unpaired data in a first embodiment of the present invention;
[0048] Figure 2 is an overall flow chart of the visible light infrared target tracking training method based on unpaired data in the first embodiment of the present invention;
[0049] Figure 3 is a structural diagram of the visible light infrared tracker in the present invention;
[0050] Figure 4 It is a structural diagram of a visible light infrared target tracking training device based on unpaired data in the second embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] like Figures 1 to 3 As shown, the embodiment of the present invention proposes a visible light infrared target tracking training method based on unpaired data, comprising the following steps:
[0053] S10, acquiring an unpaired visible light image and a thermal infrared image, and generating candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples;
[0054] It should be noted that by sampling 8 consecutive frames in a pair of unpaired multimodal videos, there are manually marked detection boxes in the images to indicate the current tracking target, and the sampling data of the two modalities are randomly distributed in a Gaussian manner to generate 256 positive samples and 768 negative samples respectively.
[0055] S20, training a visible light infrared tracker using the candidate samples;
[0056] Among them, the visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image serves as the input of the first modality-specific network and the input of the modality sharing module, and the thermal infrared image serves as the input of the second modality-specific network and the input of the modality sharing module. The visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module.
[0057] Among them, the first modality specific network and the second modality specific network are used to extract feature maps of visible light images and thermal infrared images respectively, the modality sharing module is used to extract similar features of visible light feature maps and thermal infrared feature maps, the modality adaptive attention module is used to realize inter-modal learning and enhancement of unpaired data, and the modality adaptation module is used to divide the modalities and realize target tracking.
[0058] This embodiment sets a modality sharing module to extract similar features of visible light data and thermal infrared data, further strengthens the connection between modalities, and sets a modality adaptive attention module to achieve learning and enhancement between unpaired visible light data and thermal infrared data modalities, thereby liberating the power of unpaired visible light infrared data, effectively avoiding the problem of insufficient data required for training, and fully mining and utilizing bimodal information on a limited data set to achieve mutual enhancement between unpaired multimodal data.
[0059] In one embodiment, the first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as inputs of the last two convolutional layers of the modality sharing network;
[0060] The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
[0061] It should be noted that the samples of each modality are input into their own modality-specific networks and transmitted in parallel, and the samples of the two modalities are input into the modality sharing module together. The first modality-specific network, the second modality-specific network and the modality sharing module all draw on the first three convolutional layers of the VGG network, and the convolution kernel sizes are 7×7, 5×5, and 3×3 respectively. The results of the first two convolutional layers of the first modality-specific network and the second modality-specific network are used as the input of the next convolutional layer in the modality sharing module. The results of the modality sharing module and the first modality-specific network, and the results of the modality sharing module and the second modality-specific network are fused to obtain the final modality features.
[0062] It should be noted that the modality sharing module is used to extract similar features of visible light data and thermal infrared data, further strengthen the connection between the modalities, help fuse the data of the two modalities, and further achieve feature complementarity between the modalities. Although unpaired visible light data and thermal infrared data cannot be simply fused using a general fusion network, there are great similarities between visible light and thermal infrared data, and there are shared distribution features. This embodiment extracts the shared features of the two modalities through the modality sharing module to balance the fusion ratio of the two modalities.
[0063] In one embodiment, the modality-adaptive attention module includes a first fully-connected layer and a second fully-connected layer with shared weights, a third fully-connected layer and a fourth fully-connected layer that are modality-specific, and a fifth fully-connected layer that is modality-shared;
[0064] The visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimensionality reduction processing is performed to obtain the visible light modal features and the thermal infrared modal features after dimensionality reduction;
[0065] The third fully connected layer and the fourth fully connected layer respectively process the reduced-dimensional visible light modal features and the reduced-dimensional thermal infrared modal features using a QKV mechanism to form attention matrices corresponding to the two modalities;
[0066] The attention matrices corresponding to the two modalities are passed through the fifth fully connected layer to form a modality shared query set;
[0067] After multiplying the modality shared query set with the attention matrices corresponding to the two modalities respectively, enhanced feature maps corresponding to the two modalities are obtained.
[0068] Specifically, the modality-adaptive attention module consists of two first and second fully connected layers of size 512 with shared weights, two third and fourth fully connected layers of size 64 that are modality-specific, and a fifth fully connected layer of size 64 that is modality-shared. Because the first fully connected layer is based on modality sharing, each modality will go through the fifth fully connected layer to obtain the features of the modality. First, the feature map output by the modality sharing module is flattened to a size of 512, and then reduced to 64 through the first and second fully connected layers, and then the visible light and thermal infrared specific keys K and values V are formed through the modality-specific third and fourth fully connected layers, and two modality-specific sub-attention matrices V*K are formed through vector points. The fifth fully connected modality is used to form a modality-shared query set Q, and the resulting modality-specific attention matrix is multiplied to obtain the final mutually enhanced features of the two modalities.
[0069] The modality-adaptive attention module set in this embodiment is an attention module that can realize information interaction, has a powerful attention mechanism, automatically learns the specific gradient information preferences of the two modalities, and then in the process of overall network optimization, coordinates and optimizes the entire network, fully utilizes the advantages of each modality, and realizes robust visible light infrared target tracking. It liberates the power of unpaired visible light infrared data and fully taps the potential of single modality data.
[0070] In one embodiment, the modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities;
[0071] A random activation function of neurons is added after the first two fully connected layers connected in sequence;
[0072] The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
[0073] Specifically, the modality adaptation module includes four fully connected layers of size 1024, 512, 2, 2. Two fully connected layers of size 2, 2 are parallel to form a modality connection layer. The two fully connected layers of size 1024 and 512 are followed by a dropout (random activation of neurons) regularization method to reduce the risk of overfitting. Finally, the two fully connected layers of size 2 divided by modality contain softmax layers to calculate the positive and negative scores f for each candidate sample feature in parallel. + (x i ) and f - (x i ), calculate the target probability of the candidate sample, the detection box with the highest target probability is the predicted target tracking result, and finally predict the target position by the following formula:
[0074]
[0075] Among them, xi represents the i-th sample, f + (x i ) represents the target probability of obtaining samples, f - (x i ) represents the background probability of obtaining the sample, and x* is the predicted target position.
[0076] In some embodiments, the method further comprises:
[0077] The visible light infrared tracker is trained using the cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is:
[0078] Loss=-(y*log(y^)+(1-y)*log(1-y^));
[0079] The overall network of the visible light infrared tracker is optimized by stochastic gradient descent method.
[0080] Specifically, the training process of the visible light infrared tracker is as follows:
[0081] (1) The parameters of the modality-specific module and the modality-shared module are initialized using the VGG pre-trained model. The modality-adaptive attention module and the modality-adaptive module are randomly initialized. The modality-specific module and the modality-shared module consist of three convolutional layers and ReLU (non-linear layer). The first two convolutional layers also add LRN (local response function) and MaxPool (maximum pooling function). The convolution kernel sizes are 7*7*96, 5*5*256, and 3*3*512.
[0082] (2) Use manually labeled visible light infrared data that does not require pairing to train the entire network, and use Gaussian distribution to randomly select 256 candidate samples near the true value box.
[0083] (3) In the first stage, thermal infrared data or visible light data are randomly input to train the modality sharing module, and the stochastic gradient algorithm SGD is used to update the network parameters. The parameters of the modality sharing module are updated by each training data. In the second stage, thermal infrared data and visible light data are simultaneously input to train other modules of each modality, and the stochastic gradient algorithm SGD is used to update the network parameters. Each branch corresponding to each modality is iteratively updated using its corresponding video sequence. The final model is saved for the online tracking stage.
[0084] In the model training process, this embodiment trains the modality-sharing module and the modality-specific module separately, and uses unpaired multimodal data for training. Using unpaired multimodal data for training solves the problem of reliance on large-scale aligned training data in RGBT tracking, and can make full use of existing thermal infrared data sets and visible light data sets, saving a lot of manpower and time costs. The trained tracker reveals the power of unpaired RGBT data and effectively plays the advantages of each modality.
[0085] It should be noted that it is difficult to perfectly pair the existing visible light and thermal infrared data sets, and all existing visible light infrared target tracking algorithms are designed for the characteristics of paired two modal data. This will result in the module design not being able to fully exert its performance. This embodiment designs a modality-adaptive attention module for the enhancement of unpaired data, which can achieve target tracking using unpaired modality data.
[0086] In one embodiment, the process of tracking a target using a trained visible light infrared tracker is as follows:
[0087] (1) According to the visible light thermal infrared paired video sequence, extract the first frame truth box of the video sequence, and use the pre-trained parameters to initialize the network model to obtain a new layer. At this time, the learning rate of the first two fully connected layers of the modality adaptation module is set to 0.001, and the learning rate of the last fully connected layer is set to 0.0005. After initialization, use Gaussian distribution sampling to generate 256 candidate samples.
[0088] (2) The candidate samples are sent to the corresponding modality-specific modules respectively, and then to the modality sharing module in parallel. The results of each modality-specific module are sent to the modality sharing module of the next layer, and the candidate samples are sent to the modality sharing module together. The results of the modality sharing module and the modality-specific module are fused according to the modality to obtain the modality features. The modality features are sent to the modality adaptive attention module, and the modality sharing is fully connected to form the modality sharing query set Q. The modality-specific attention matrix generated is multiplied to obtain the final mutually enhanced modality features. In the last convolutional layer, the enhanced feature maps of different modalities are spliced in the channel dimension to obtain an overall feature map, which is then sent to the last modality adaptation module and sent to the softmax function in the last convolutional layer to obtain the binary classification score and predict the target position.
[0089] (3) When the target probability of the predicted sample is greater than 0.5, the tracking is considered successful. When the target probability of the predicted sample is less than 0.5, the tracking fails and a short-term update is performed. If the number of frames in the positive and negative sample data sets exceeds 20, the negative sample areas of the earliest frames are discarded. 32 positive samples and 96 negative samples are extracted from the positive and negative sample sets to fine-tune the parameters of the fully connected layer. The iteration is repeated 10 times, and the learning rate is set to 0.00003.
[0090] (4) During the online tracking process, long-term updates are performed every 8 frames. If the number of frames in the positive and negative sample data sets exceeds 100, the positive sample areas of the earliest frames are discarded. 32 positive samples and 96 negative samples are extracted from the positive and negative sample sets to fine-tune the parameters of the fully connected layer. The training is repeated 10 times, and the learning rate is set to 0.00003. If the conditions for short-term updates and long-term updates are not met, the next frame is tracked directly and the model is not updated.
[0091] It should be noted that both long-term and short-term updates are used to adapt to changes in the appearance of the tracked target, and the model parameters are updated using sample data. Short-term updates are updated immediately when tracking fails (immediate adjustments), and long-term updates are used to adapt to changes in the target during tracking (adjustments are made at specified frame intervals).
[0092] In this embodiment, the method of the present invention and some existing methods are tested on the public data sets GTOT, RGBT234 and LaSHeR, and the test results are evaluated with other trackers in terms of SR (success rate) and PR (precision rate). The results are shown in Table 1:
[0093] Table 1
[0094]
[0095] Among them, _ indicates that the experiment was not conducted on the data set, and UMT indicates the target tracking method used in the present invention (other methods are comparative methods). From the data in Table 1, it can be observed that the success rate and accuracy of the method of the present invention are relatively high on the existing data set, and the tracking performance is uniformly improved to a certain extent.
[0096] In addition, if Figure 4 As shown, the embodiment of the present invention also proposes a visible light infrared target tracking training device based on unpaired data, the device comprising:
[0097] An acquisition module 10, configured to acquire an unpaired visible light image and a thermal infrared image, and generate candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples;
[0098] A training module 20, configured to train a visible light infrared tracker using the candidate samples;
[0099] Among them, the visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image serves as the input of the first modality-specific network and the input of the modality sharing module, and the thermal infrared image serves as the input of the second modality-specific network and the input of the modality sharing module. The visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module serves as the input of the modality adaptive attention module.
[0100] In one embodiment, the first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as inputs of the last two convolutional layers of the modality sharing network;
[0101] The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
[0102] In one embodiment, the modality-adaptive attention module includes a first fully-connected layer and a second fully-connected layer with shared weights, a third fully-connected layer and a fourth fully-connected layer that are modality-specific, and a fifth fully-connected layer that is modality-shared;
[0103] The visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimensionality reduction processing is performed to obtain the visible light modal features and the thermal infrared modal features after dimensionality reduction;
[0104] The third fully connected layer and the fourth fully connected layer respectively process the reduced-dimensional visible light modal features and the reduced-dimensional thermal infrared modal features using a QKV mechanism to form attention matrices corresponding to the two modalities;
[0105] The attention matrices corresponding to the two modalities are passed through the fifth fully connected layer to form a modality shared query set;
[0106] After multiplying the modality shared query set with the attention matrices corresponding to the two modalities respectively, enhanced feature maps corresponding to the two modalities are obtained.
[0107] In one embodiment, the modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities;
[0108] A random activation function of neurons is added after the first two fully connected layers connected in sequence;
[0109] The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
[0110] In one embodiment, the training module 20 includes:
[0111] A training unit is used to train the visible light infrared tracker using a cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is:
[0112] Loss=-(y*log(y^)+(1-y)*log(1-y^));
[0113] The optimization unit is used to optimize the overall network of the visible light infrared tracker by using a stochastic gradient descent method.
[0114] It should be noted that other embodiments or implementation methods of the visible light infrared target tracking training device based on unpaired data of the present invention can refer to the above-mentioned method embodiments, which will not be repeated here.
[0115] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0116] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0117] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0118] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0119] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A visible light infrared target tracking training method based on unpaired data, characterized in that: The method comprises: Acquire an unpaired visible light image and a thermal infrared image, and generate candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples; Using the candidate samples to train a visible light infrared tracker; The visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image is used as the input of the first modality-specific network and the input of the modality sharing module, the thermal infrared image is used as the input of the second modality-specific network and the input of the modality sharing module, the visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module is used as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module is used as the input of the modality adaptive attention module; The modality-adaptive attention module includes a first fully connected layer and a second fully connected layer with shared weights, a modality-specific third fully connected layer and a fourth fully connected layer, and a modality-shared fifth fully connected layer; the visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimension reduction processing is performed to obtain the reduced visible light modal features and the reduced thermal infrared modal features; the reduced visible light modal features and the reduced thermal infrared modal features are respectively processed by the third fully connected layer and the fourth fully connected layer using the QKV mechanism to form attention matrices corresponding to the two modalities; the attention matrices corresponding to the two modalities are formed into a modality-shared query set through the fifth fully connected layer; the modality-shared query set is multiplied by the attention matrices corresponding to the two modalities to obtain enhanced feature maps corresponding to the two modalities.
2. The visible light infrared target tracking training method based on unpaired data as claimed in claim 1, characterized in that: The first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as the inputs of the last two convolutional layers of the modality sharing module; The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
3. The visible light infrared target tracking training method based on unpaired data as claimed in claim 1, characterized in that: The modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities; A random activation function of neurons is added after the first two fully connected layers connected in sequence; The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
4. The visible light infrared target tracking training method based on unpaired data as claimed in claim 3, characterized in that: The method further comprises: The visible light infrared tracker is trained using the cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is: Loss=-(y*log(y^)+(1-y)*log(1-y^)); The overall network of the visible light infrared tracker is optimized by stochastic gradient descent method.
5. A visible light infrared target tracking training device based on unpaired data, characterized in that: The device comprises: An acquisition module, used to acquire an unpaired visible light image and a thermal infrared image, and generate candidate samples based on the visible light image and the thermal infrared image, wherein the candidate samples include positive samples and negative samples; A training module, used for training a visible light infrared tracker using the candidate samples; The visible light infrared tracker includes a modality-specific module, a modality sharing module, a modality adaptive attention module and a modality adaptation module connected in sequence, and the modality-specific module includes a first modality-specific network and a second modality-specific network; the visible light image is used as the input of the first modality-specific network and the input of the modality sharing module, the thermal infrared image is used as the input of the second modality-specific network and the input of the modality sharing module, the visible light modality feature obtained by adding the output of the first modality-specific network to the output of the modality sharing module is used as the input of the modality adaptive attention module, and the thermal infrared modality feature obtained by adding the output of the second modality-specific network to the output of the modality sharing module is used as the input of the modality adaptive attention module; The modality-adaptive attention module includes a first fully connected layer and a second fully connected layer with shared weights, a modality-specific third fully connected layer and a fourth fully connected layer, and a modality-shared fifth fully connected layer; the visible light modal features and the thermal infrared modal features are respectively used as inputs of the first fully connected layer and the second fully connected layer, and dimension reduction processing is performed to obtain the reduced visible light modal features and the reduced thermal infrared modal features; the reduced visible light modal features and the reduced thermal infrared modal features are respectively processed by the third fully connected layer and the fourth fully connected layer using the QKV mechanism to form attention matrices corresponding to the two modalities; the attention matrices corresponding to the two modalities are formed into a modality-shared query set through the fifth fully connected layer; the modality-shared query set is multiplied by the attention matrices corresponding to the two modalities to obtain enhanced feature maps corresponding to the two modalities.
6. The visible light infrared target tracking training device based on unpaired data as claimed in claim 5, characterized in that: The first modality specific network, the second modality specific network and the modality sharing module each include three convolutional layers connected in sequence, and the outputs of the first two convolutional layers of the first modality specific network and the second modality specific network serve as the inputs of the last two convolutional layers of the modality sharing module; The output of the last convolutional layer of the first modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module, and the output of the last convolutional layer of the second modality specific network and the output of the last convolutional layer of the modality sharing module are added as the input of the modality adaptive attention module.
7. The visible light infrared target tracking training device based on unpaired data as claimed in claim 6, characterized in that: The modality adaptation module includes two fully connected layers and a modality connection layer connected in sequence, and the modality connection layer includes a visible light modality fully connected layer and a thermal infrared modality fully connected layer corresponding to the two modalities; A random activation function of neurons is added after the first two fully connected layers connected in sequence; The visible light modality fully connected layer and the thermal infrared modality fully connected layer contain a softmax layer, and the softmax layer is used to calculate the score values of the positive and negative samples in the candidate samples and predict the target position.
8. The visible light infrared target tracking training device based on unpaired data according to claim 7, characterized in that: The training module includes: A training unit is used to train the visible light infrared tracker using a cross entropy loss function generated by the negative sample score value y and the positive score value y^, wherein the cross entropy loss function is: Loss=-(y*log(y^)+(1-y)*log(1-y^)); The optimization unit is used to optimize the overall network of the visible light infrared tracker by using a stochastic gradient descent method.
Citation Information
Patent Citations
All-weather target real-time tracking method based on visible light and infrared images
CN106485245A
Multi-vision modal data acquisition system and acquisition method
CN107343180A
Target tracking method and device and related equipment
CN112150508A