Track platform target tracking method and device based on visible light-infrared image
Through the combination of zero-sample image recovery and cross-modal regulation attention model, the interference problem of visible light images in complex environments is solved, and efficient and accurate tracking of orbital platform targets is achieved.
Patent Information
- Application Number
- CN202510767950.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the prior art, visible light images are susceptible to interference in complex environments, target tracking accuracy is low, and traditional deep learning methods have limited generalization capabilities in multimodal image fusion.
The zero-sample image recovery strategy is used to preprocess the visible light image, combined with the cross-modal regulatory attention model, feature fusion and target prediction are performed through Transformer's visible light-infrared tracker, the noise injection mechanism is used to reduce the number of iterations, and tracking accuracy is improved through collaborative marking elimination strategies.
Improve image recovery quality with fewer iterations, enhance the robustness and accuracy of target tracking, and perform well in complex scenarios.
Smart Images

Figure CN120279065A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of track perception, and in particular to a method and device for tracking objects on a track platform based on visible light-infrared images. Background Art
[0002] With the rapid development of rail transit, the safety monitoring and object tracking on track platforms have become important links in ensuring passenger safety and operation efficiency. Traditional track platform monitoring systems mainly rely on visible light cameras. However, visible light images are easily disturbed in complex environments (such as insufficient light, haze, night, etc.), resulting in a decline in image quality and thus affecting the accuracy of object tracking. To address this problem, infrared imaging technology has gradually been introduced into track platform monitoring systems. Infrared images can provide clear object information under low light and adverse weather conditions, but their resolution is low and they lack color information.
[0003] In terms of image restoration, traditional image restoration methods mainly rely on signal processing and optimization techniques, such as model-based methods (such as Wiener filtering, inverse filtering, etc.), variational methods, and sparse representation. These methods model image degradation through mathematical models and restore images through optimization techniques. With the development of deep learning, deep learning-based image restoration methods have gradually become the mainstream. These methods are mainly divided into task-specific deep neural networks and zero-shot image restoration. Task-specific deep neural networks train specialized networks for each image restoration task, but these networks usually perform well on specific tasks and are very sensitive to changes in the observation model during testing, with limited generalization ability. Summary of the Invention
[0004] The present invention provides a method and device for tracking objects on a track platform based on visible light-infrared images, which solves the problems in the prior art that visible light images are easily disturbed in complex environments and the object tracking accuracy is low, and improves the robustness and accuracy of object tracking.
[0005] The present invention provides a method for tracking objects on a track platform based on visible light-infrared images, including the following steps: Preprocess the first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and a noise injection mechanism; Input the second visible light image and the infrared image into a cross-modal modulation attention model to obtain the predicted object position; the cross-modal modulation attention model includes a visible light-infrared tracker based on Transformer, Transformer blocks, a cross-modal modulation attention unit, and a tracking prediction head.
[0006] A method for tracking objects on an orbital platform based on visible-light and infrared images according to the present invention preprocesses a first visible-light image based on a zero-shot image restoration strategy to obtain a second visible-light image, which specifically includes: obtaining the pseudo-inverse of the first visible-light image, initializing based on the pseudo-inverse of the first visible-light image to generate a starting point of the target image; performing image restoration based on the starting point of the target image to obtain a second visible-light image; accelerating the image restoration process through back-projection guidance; and reducing the number of iterations in the image restoration process through a noise injection mechanism.
[0007] A method for tracking objects on an orbital platform based on visible-light and infrared images according to the present invention inputs the second visible-light image and the infrared image into a cross-modal modulation attention model to obtain a predicted target position, which specifically includes: inputting the second visible-light image and the infrared image into a visible-light-infrared tracker based on Transformer to obtain a landmark sequence; the second visible-light image corresponding to the infrared image; inputting the landmark sequence into a Transformer block for processing to obtain a feature sequence; inputting the feature sequence into a cross-modal modulation attention unit to obtain an updated feature sequence; merging the updated feature sequence along the channel dimension and inputting it into a tracking prediction head to output the position of the predicted target.
[0008] A method for tracking objects on an orbital platform based on visible-light and infrared images according to the present invention inputs the second visible-light image and the infrared image into a visible-light-infrared tracker based on Transformer to obtain a landmark sequence, which specifically includes: inputting the second visible-light image and the infrared image into a visible-light-infrared tracker based on Transformer and dividing them into image patches of the same size; flattening the divided image patches to obtain an image patch sequence; determining a template feature and a search region feature according to the image patch sequence; and connecting the template feature and the search region feature to obtain a landmark sequence.
[0009] A method for tracking target on rail transit platform based on visible light-infrared images according to the present invention. When inputting the feature sequence into the cross-modal modulation attention unit to obtain the updated feature sequence, it specifically includes: inputting the feature sequence into the cross-modal modulation attention unit to obtain the query matrix and keyword matrix of the visible light modality, and the query matrix and keyword matrix of the infrared modality; determining the original visible light correlation map according to the query matrix and keyword matrix of the visible light modality; determining the original infrared correlation map according to the query matrix and keyword matrix of the infrared modality; determining the aggregated information according to the original visible light correlation map and the original infrared correlation map; performing an attention operation on the aggregated information to obtain the initial visible light modulation correlation map and the initial infrared modulation correlation map; determining the final visible light image attention map according to the original visible light correlation map and the initial visible light modulation correlation map; determining the final infrared image attention map according to the original infrared correlation map and the initial infrared modulation correlation map; determining the corresponding updated feature sequence according to the final visible light image attention map and the final infrared image attention map.
[0010] A method for tracking target on rail transit platform based on visible light-infrared images according to the present invention. During the process of inputting the feature sequence into the cross-modal modulation attention unit to obtain the updated feature sequence, the irrelevant markers to the target are eliminated through the collaborative marker elimination strategy.
[0011] A method for tracking target on rail transit platform based on visible light-infrared images according to the present invention. The eliminating the irrelevant markers to the target through the collaborative marker elimination strategy specifically includes: determining the search area defined in the current frame image; extracting the marker features from the search area, where the marker features include the markers from the visible light and infrared modalities; calculating the attention weight between each marker in the search area and the template marker; obtaining the fused weight according to the attention weights of the visible light and infrared modalities; sorting all the markers in the search area according to the fused weight, and selecting the top k markers with the highest weights as the target candidates, and deleting the markers other than the top k markers.
[0012] The present invention also provides a device for tracking target on rail transit platform based on visible light-infrared images, including the following modules: A preprocessing module, configured to preprocess the first visible light image based on the zero-shot image restoration strategy to obtain the second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and noise injection mechanism. A target tracking module for inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain a predicted target position; the cross-modal regulation attention model includes a visible light-infrared tracker based on Transformer, a Transformer block, a cross-modal regulation attention unit, and a tracking prediction head.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for tracking an orbital platform target based on visible light-infrared images as described in any one of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for tracking an orbital platform target based on visible light-infrared images as described in any one of the above.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for tracking an orbital platform target based on visible light-infrared images as described in any one of the above.
[0016] The method and device for tracking an orbital platform target based on visible light-infrared images provided by the present invention have the following beneficial effects: By preprocessing the first visible light image through a zero-shot image restoration strategy, a high-quality second visible light image is obtained, effectively solving the problems that visible light images are vulnerable to interference and quality degradation in complex environments; at the same time, through a visible light-infrared tracker based on Transformer and a cross-modal regulation attention unit, multi-modal feature fusion of the second visible light image and the infrared image is performed, enhancing the robustness and accuracy of target tracking. Therefore, the present invention realizes the efficient completion of the image restoration task with fewer iteration times, effectively fuses the multi-modal information of visible light and infrared images, and significantly improves the accuracy and efficiency of orbital platform target tracking. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is one of the flow diagrams of the method for tracking an orbital platform target based on visible light-infrared images provided by the present invention.
[0019] Figure 2It is the second schematic flow chart of the method for tracking targets at railway stations based on visible light-infrared images provided by the present invention.
[0020] Figure 3 It is a schematic diagram of the cross-modal regulated attention model architecture provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of the device for tracking targets at railway stations based on visible light-infrared images provided by the present invention.
[0022] Figure 5 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] In terms of image restoration, before the emergence of deep learning, image restoration mainly relied on signal processing and optimization techniques. These methods include model-based methods, variational methods, and sparse representation. Model-based methods such as Wiener filtering and inverse filtering model image degradation through mathematical models and restore images through optimization techniques. Variational methods optimize the image restoration problem by defining an energy function (such as total variation regularization). Sparse representation uses the sparsity of images (such as wavelet transform) to restore images, assuming that images are sparse in a certain transform domain. With the development of deep learning, many new technologies have emerged in the field of image restoration. Deep learning-based methods are mainly divided into task-specific deep neural networks and zero-shot image restoration. Task-specific deep neural networks train a dedicated deep neural network for each image restoration task (such as denoising, super-resolution, deblurring, etc.), but these networks usually perform well on specific tasks and are very sensitive to changes in the observation model during testing, with limited generalization ability. Zero-shot image restoration methods do not rely on networks trained for specific tasks, but use pre-trained generative models (such as diffusion models, generative adversarial networks, etc.) as signal priors and adapt to the observation model during testing, avoiding the limitations of task-specific networks.
[0025] In the aspect of object tracking, existing visible-light-infrared tracking methods based on Transformer extract unimodal features through self-attention and enhance multimodal feature interaction through cross-modal attention. However, this method calculates correlations independently in self-attention and is vulnerable to low-quality data, resulting in inaccurate correlation weights, thus limiting the tracking performance.
[0026] It can be seen that there is an urgent need to provide a robust object tracking method that can effectively handle various image degradation tasks and fuse multimodal information.
[0027] A visible-light-infrared orbital station object tracking method adopting a zero-shot image restoration strategy proposed by the present invention can be applied to the processing of various image degradation tasks and effectively utilize multimodal information.
[0028] The present invention provides a visible-light-infrared orbital station object tracking method adopting a zero-shot image restoration strategy, which can be applied to the processing of various image degradation tasks and effectively utilize multimodal information. As Figure 1 shown in the general flowchart of a visible-light-infrared orbital station object tracking method adopting a zero-shot image restoration strategy provided by the present invention, it includes the following steps: The input visible-light image is first processed by the zero-shot image restoration strategy to obtain a visible-light image without degradation effects, and then together with the infrared image, it is input into the cross-modal modulation attention model. After being processed by the Transformer block, a feature sequence is obtained, and then processed by the cross-modal modulation attention unit. At the same time, a collaborative label elimination strategy is applied in a specific layer to obtain an updated feature sequence, which is then input into the tracking prediction head to output the position of the predicted target.
[0029] The advantages of the present invention compared with the previous technologies can be summarized as the following points: (1) The present invention adopts a zero-shot image restoration strategy, which can efficiently complete tasks such as image super-resolution, deblurring, and restoration with fewer iteration times, significantly reducing the computational complexity while improving the restoration quality.
[0030] (2) Through the cross-modal modulation attention mechanism, the present invention fuses visible-light and infrared multimodal information, enhances feature interaction and correlation calculation, and improves the robustness and accuracy of object tracking, especially performing better in complex scenarios.
[0031] The following combines Figures 2 - 5 to specifically describe the embodiments of the present invention.
[0032] Figure 2 is a schematic flowchart of an orbital station object tracking method based on visible-light-infrared images provided by the present invention. As Figure 2 shown, the method includes the following steps: S210. Preprocess the first visible light image based on the zero-shot image restoration strategy to obtain the second visible light image. The zero-shot image restoration strategy includes input initialization, back-projection guidance, and noise injection mechanism.
[0033] According to a method for tracking an orbital platform target based on visible light-infrared images provided by the present invention, preprocess the first visible light image based on the zero-shot image restoration strategy to obtain the second visible light image, specifically including: obtaining the pseudo-inverse of the first visible light image, initializing based on the pseudo-inverse of the first visible light image to generate the starting point of the target image; performing image restoration based on the starting point of the target image to obtain the second visible light image; accelerating the image restoration process through back-projection guidance; reducing the number of iterations in the image restoration process through the noise injection mechanism.
[0034] Specifically, the present invention designs a zero-shot image restoration strategy based on a consistency model, which can effectively complete tasks such as image super-resolution, deblurring, and restoration with fewer iterations. This strategy includes several parts: input initialization, back-projection guidance, and noise injection mechanism. Input initialization is used to preprocess the first visible light image and provide a suitable starting point for subsequent image restoration; back-projection guidance is used to adjust the intermediate results of the image restoration process; the noise injection mechanism can effectively utilize noise and reduce the number of iterations.
[0035] Specifically, the zero-shot image restoration strategy based on the consistency model includes: input initialization, back-projection guidance, and noise injection mechanism.
[0036] The goal of the image restoration task is to restore a high-quality image from the degraded first visible light image where and and The relationship between them can usually be represented by a linear model:
[0037] where is the measurement matrix, is the additive noise.
[0038] Initialization refers to the process of preprocessing the first visible light image before the start of the restoration algorithm. Its purpose is to provide a suitable starting point for subsequent restoration operations so as to improve the efficiency and restoration quality of the algorithm.
[0039] Traditional image restoration methods based on diffusion models usually use pure noise for initialization. Specifically, the input of the initialization is usually a Gaussian noise vector, that is where represents a Gaussian distribution, represents the time step, represents the identity matrix. This initialization method ignores the target image contained in the first visible light image and only relies on the generation ability of the model. To make the generated image consistent with the first visible light image , the diffusion model introduces data fidelity guidance during the iteration process. However, this guidance mechanism usually requires multiple iterations (i.e., a large number of neural function evaluations) to achieve good results.
[0040] The consistency model attempts to more directly utilize the information of the first visible light image during initialization. Specifically, the initialization input is initialized through the pseudo-inverse of the first visible light image , that is, , where is the pseudo-inverse of the measurement matrix , defined as , , is the initial noise level, and the measurement matrix represents the process of image degradation (such as blurring, downsampling, etc.). This method utilizes the structural information in the first visible light image and provides a starting point closer to the target image for the restoration process. A certain level of noise is injected during this process, but the noise injection not only retains the generation ability of the model but also provides the necessary "degrees of freedom" for the subsequent restoration process, enabling the model to better adapt to different restoration tasks. This initialization method reduces the burden of subsequent iterations, enabling the consistency model to achieve efficient image restoration with fewer neural function evaluations.
[0041] Back-projection guidance is used to accelerate the image restoration process. The core is to utilize the first visible light image and the pseudo-inverse of the measurement matrix to adjust the intermediate results during the restoration process through back-projection operations. Specifically, the update formula for back-projection guidance is
[0042] where, represents the restored image at the time step , represents the initial estimated image at the time step τn, is the guidance scaling factor used to control the weight of the data fidelity term, represents the noise level; is the data fidelity term guided by backprojection, (AA T ) -1 / 2 represents the square root of the pseudo-inverse of the measurement matrix A, and ||·||2 represents the L2 norm, which is used to calculate the Euclidean distance of a vector; is the gradient of this data fidelity term, represents the pseudo-inverse of the measurement matrix A, Ax represents the degradation operation of the restored image x through the measurement matrix A, and y represents the observed data, that is, the first visible light image.
[0043] The noise injection mechanism includes noise level decoupling and noise injection segmentation, which allows the image to make more effective use of noise during the restoration process and reduces the number of iterations.
[0044] Noise level decoupling means separating the noise level used in the denoising operation from the noise level of the noise injection step, so that these two processes can be optimized and adjusted independently.
[0045] Noise injection segmentation divides the injected noise into random noise and estimated noise. Random noise refers to the unpredictable noise added during the image restoration or generation process, which helps the model explore different parts of the data distribution, increases the diversity of samples, and prevents the model from converging to the local optimal solution prematurely. It follows a Gaussian distribution . Estimated noise is calculated based on the difference between the currently estimated signal and the true signal, which can help the model converge to the true signal faster, expressed as
[0046] Then the noise injection can be expressed as
[0047] where is a hyperparameter used to balance between random noise and estimated noise.
[0048] After the visible light image is processed by the zero-shot image restoration strategy, an image without degradation effect is obtained, which is convenient for jointly predicting the position of the target with the infrared image.
[0049] S220. Input the second visible light image and the infrared image into the cross-modal modulation attention model to obtain the predicted target position. The cross-modal modulation attention model includes a Transformer-based visible light-infrared tracker, Transformer blocks, cross-modal modulation attention units, and a tracking prediction head.
[0050] Specifically, the present invention designs a visible light-infrared orbital platform target tracking method, which is achieved through a cross-modal regulation attention model. As Figure 3 shown, the model mainly includes a visible light-infrared tracker based on Transformer, a cross-modal regulation attention unit, and proposes a cross-modal regulation attention mechanism and a collaborative marker elimination strategy. By means of a unified attention model, single-modal self-correlation, cross-modal feature interaction, and search-template correlation calculations are simultaneously performed to improve the tracking performance.
[0051] According to a visible light-infrared image-based orbital platform target tracking method provided by the present invention, inputting a second visible light image and an infrared image into the cross-modal regulation attention model to obtain a predicted target position, specifically including: inputting the second visible light image and the infrared image into a visible light-infrared tracker based on Transformer to obtain a flag sequence; the second visible light image corresponds to the infrared image; inputting the flag sequence into a Transformer block for processing to obtain a feature sequence; inputting the feature sequence into the cross-modal regulation attention unit to obtain an updated feature sequence; merging the updated feature sequence along the channel dimension and inputting it into a tracking prediction head to output the position of the predicted target.
[0052] According to a visible light-infrared image-based orbital platform target tracking method provided by the present invention, inputting a second visible light image and an infrared image into a visible light-infrared tracker based on Transformer to obtain a flag sequence, specifically including: inputting the second visible light image and the infrared image into a visible light-infrared tracker based on Transformer, and respectively dividing them into image blocks of the same size; flattening the divided image blocks to obtain an image block sequence; determining a template feature and a search area feature according to the image block sequence; connecting the template feature and the search area feature to obtain a flag sequence.
[0053] Specifically, the model mainly includes a visible light-infrared tracker based on Transformer, and the tracker mainly includes a Transformer block, a cross-modal regulation attention unit, and a tracking prediction head.
[0054] The tracker consists of two branches of visible light and infrared, and these two branches share parameters and independently process different modalities.
[0055] Given an input pair of visible light and infrared template images and a pair of search area images , where , represents the image height, , represents the image height, and 3 represents 3 channels.
[0056] First, these images are divided into image patches of size , and then they are flattened to obtain a sequence of image patches and , where , respectively representing the number of patches of the template and the search box. Use a patch embedding layer with parameters and learnable positional encoding and to obtain the template features , and the search area features , , as shown in the following formula:
[0057] Then these features are concatenated to obtain the flag sequence , .
[0058] According to a method for tracking an orbital platform target based on visible light-infrared images provided by the present invention, the feature sequence is input into a cross-modal regulation attention unit to obtain an updated feature sequence, which specifically includes: inputting the feature sequence into the cross-modal regulation attention unit to obtain a query matrix and a key matrix in the visible light modality, and a query matrix and a key matrix in the infrared modality; determining an original visible light correlation map according to the query matrix and the key matrix in the visible light modality; determining an original infrared correlation map according to the query matrix and the key matrix in the infrared modality; determining aggregated information according to the original visible light correlation map and the original infrared correlation map; performing an attention operation on the aggregated information to obtain an initial visible light regulation correlation map and an initial infrared regulation correlation map; determining a final visible light image attention map according to the original visible light correlation map and the initial visible light regulation correlation map; determining a final infrared image attention map according to the original infrared correlation map and the initial infrared regulation correlation map; and determining a corresponding updated feature sequence according to the final visible light image attention map and the final infrared image attention map.
[0059] Specifically, the flag sequence is sent to the Transformer block at the layer for processing, and the formula for the forward propagation process is as follows: , is the feature sequence output by the last Transformer block. Send , into the cross-modal regulation attention unit.
[0060] The query and keyword matrix in the visible light modality can be expressed as:
[0061] where and represent the linear projection weights of the query and the key respectively. For the visible light branch, the process of generating its feature correlation map can be expressed as:
[0062] For the infrared branch, the same processing is adopted to obtain . and are both divided into four parts and , which have different roles in tracking prediction. Each part is simply named TT, TS, ST, SS according to the query-key pair used to calculate the correlation. Among them, ST is a special part that controls the information flow from the template to the search box and has a significant impact on the tracking result. Due to the spatio-temporal aligned multi-modal image pairs, STs within different branches have significant correlations.
[0063] To achieve adaptive correlation regulation, a cross-modal regulation attention mechanism is designed to enhance the interaction in the cross-modal regulation attention unit using the correlation maps of the two modalities. The purpose of the cross-modal regulation attention mechanism is to regulate ST and also consider SS to facilitate the regulation of the final attention map. Specifically, the aggregated information of the two branches is as follows:
[0064] where LN represents the normalization layer, is a learnable linear projection weight for embedding the correlation of the two branches. Then we perform an attention operation on to obtain the regulated correlation map .
[0065]
[0066] where is the number of template tokens, and represent the linear projection weights of the query and the key in the cross-modal regulation attention unit. Next, and are separated from the initial regulated correlation maps and .
[0067]
[0068] where are learnable linear projection weights.
[0069] Finally, the obtained is added to the original correlation map to obtain the final regulatory correlation map. Generate the final visible light image attention map A rgb The process can be described as follows: ,
[0070] where C represents the dimensionality size of the flag.
[0071] The cross-modal regulatory attention unit is a symmetric structure, and the parameters at the corresponding positions of the two branches are shared. After being processed by the cross-modal regulatory attention unit, the updated feature sequence is obtained , .
[0072] Merge these features along the channel dimension and input them into the tracking prediction head to obtain the predicted bounding box , and the position of the bounding box is the position of the target.
[0073] According to a method for tracking an orbital platform target based on visible light-infrared images provided by the present invention, in the process of inputting the feature sequence into the cross-modal regulatory attention unit to obtain the updated feature sequence, the markers irrelevant to the target are eliminated through a collaborative marker elimination strategy.
[0074] According to a method for tracking an orbital platform target based on visible light-infrared images provided by the present invention, the markers irrelevant to the target are eliminated through a collaborative marker elimination strategy, which specifically includes: determining the search area defined in the current frame image; extracting marker features from the search area, and the marker features include markers from the visible light and infrared modalities; for each marker in the search area, calculating the attention weight between it and the template marker; obtaining the fused weight according to the attention weights of the visible light and infrared modalities; sorting all the markers in the search area according to the fused weight, and selecting the top k markers with the highest weights as target candidates, and deleting the markers other than the top k markers.
[0075] Specifically, a collaborative marker elimination strategy is also proposed, which is applied to a specific layer of the cross-modal regulatory attention unit to improve the tracking efficiency and accuracy. The core idea of this strategy is to make a judgment by combining the attention weights of the visible light and infrared modalities to eliminate non-target markers.
[0076] First, extract features from the search region, which include markers from visible light and infrared modalities. For each search region marker, calculate the correlation or attention weight between it and the template marker. These weights reflect the similarity between each marker and the target. Since the visible light and infrared modalities provide complementary information, the collaborative marker elimination strategy combines the attention weights of the two modalities to obtain a more accurate distinction between the target and non-target. Sort the search region markers using the fused weights and select the top k markers with the highest weights as target candidates. Exclude those markers with lower weights, which are considered less relevant to the target.
[0077] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: (1) Improvement in image restoration efficiency: The present invention adopts a zero-shot image restoration strategy, significantly reducing the number of iterations for tasks such as image super-resolution, deblurring, and inpainting, improving computational efficiency, and at the same time enhancing the quality of image restoration.
[0078] (2) Enhancement of target tracking robustness: By means of a cross-modal modulation attention mechanism, the present invention enhances the interaction and correlation calculation of visible light and infrared image features, improving the accuracy and robustness of target tracking.
[0079] Next, a description is provided for the rail transit platform target tracking device based on visible light-infrared images provided by the present invention. The rail transit platform target tracking device based on visible light-infrared images described below can be correspondingly referred to the rail transit platform target tracking method described above.
[0080] As Figure 4 shown, a rail transit platform target tracking device based on visible light-infrared images provided by the present invention includes: A preprocessing module 410, configured to preprocess the first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and noise injection mechanisms; A target tracking module 420, configured to input the second visible light image and the infrared image into a cross-modal modulation attention model to obtain a predicted target position; the cross-modal modulation attention model includes a visible light-infrared tracker based on Transformer, Transformer blocks, a cross-modal modulation attention unit, and a tracking prediction head.
[0081] Figure 5 Illustrates a schematic physical structure diagram of an electronic device, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete their mutual communication through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute an orbital platform target tracking method based on visible light-infrared images. The method includes: preprocessing a first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and a noise injection mechanism; inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain a predicted target position; the cross-modal regulation attention model includes a Transformer-based visible light-infrared tracker, Transformer blocks, a cross-modal regulation attention unit, and a tracking prediction head.
[0082] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods provided in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0083] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the orbital platform target tracking method based on visible light-infrared images provided by the above-mentioned various methods. The method includes: preprocessing a first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and a noise injection mechanism; inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain a predicted target position; the cross-modal regulation attention model includes a Transformer-based visible light-infrared tracker, Transformer blocks, a cross-modal regulation attention unit, and a tracking prediction head.
[0084] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for tracking an orbital platform target based on visible light-infrared images provided by the above-mentioned various methods. The method includes: preprocessing a first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; the zero-shot image restoration strategy includes input initialization, back-projection guidance, and a noise injection mechanism; inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain a predicted target position; the cross-modal regulation attention model includes a Transformer-based visible light-infrared tracker, Transformer blocks, a cross-modal regulation attention unit, and a tracking prediction head.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. An orbital platform target tracking method based on visible light-infrared images, characterized in that Including: Preprocessing the first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image; The zero-shot image restoration strategy includes input initialization, back-projection guidance, and noise injection mechanism; Inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain the predicted target position; the cross-modal regulation attention model includes a Transformer-based visible light-infrared tracker, Transformer blocks, a cross-modal regulation attention unit, and a tracking prediction head.
2. The method for tracking an orbital platform target based on visible light-infrared images according to claim 1, wherein Preprocessing the first visible light image based on a zero-shot image restoration strategy to obtain a second visible light image, specifically including: Obtaining the pseudo-inverse of the first visible light image, initializing according to the pseudo-inverse of the first visible light image, and generating the starting point of the target image; Performing image restoration based on the starting point of the target image to obtain a second visible light image; Accelerating the image restoration process through back-projection guidance; Reducing the number of iterations in the image restoration process through a noise injection mechanism.
3. The method for tracking an orbital platform target based on visible-infrared images according to claim 1, wherein, The step of inputting the second visible light image and the infrared image into a cross-modal regulation attention model to obtain the predicted target position specifically includes: Inputting the second visible light image and the infrared image into a Transformer-based visible light-infrared tracker to obtain a landmark sequence; the second visible light image corresponds to the infrared image; Inputting the landmark sequence into Transformer blocks for processing to obtain a feature sequence; Inputting the feature sequence into a cross-modal regulation attention unit to obtain an updated feature sequence; Merging the updated feature sequence along the channel dimension and inputting it into a tracking prediction head to output the position of the predicted target.
4. The method for tracking an orbital platform target based on visible-infrared images according to claim 3, wherein The step of inputting the second visible light image and the infrared image into a Transformer-based visible light-infrared tracker to obtain a landmark sequence specifically includes: Inputting the second visible light image and the infrared image into a Transformer-based visible light-infrared tracker and dividing them into image patches of the same size; Flattening the divided image patches to obtain an image patch sequence; Determining a template feature and a search region feature according to the image patch sequence; Connecting the template feature and the search region feature to obtain a landmark sequence.
5. The method for tracking the target of an orbital platform based on visible light-infrared images according to claim 3, wherein The step of inputting the feature sequence into a cross-modal regulation attention unit to obtain an updated feature sequence specifically includes: Inputting the feature sequence into a cross-modal regulation attention unit to obtain a query matrix and a key matrix for the visible light modality, and a query matrix and a key matrix for the infrared modality; Determining an original visible light correlation map according to the query matrix and the key matrix for the visible light modality; Determining an original infrared correlation map according to the query matrix and the key matrix for the infrared modality; Determining aggregated information according to the original visible light correlation map and the original infrared correlation map; Performing an attention operation on the aggregated information to obtain an initial visible light regulation correlation map and an initial infrared regulation correlation map; Determining a final visible light image attention map according to the original visible light correlation map and the initial visible light regulation correlation map; Determine the final infrared image attention map according to the original infrared correlation map and the initial infrared regulation correlation map; Determine the corresponding updated feature sequence according to the final visible light image attention map and the final infrared image attention map.
6. The method for tracking an orbital platform target based on visible light-infrared images according to claim 3, characterized in that, In the process of inputting the feature sequence into the cross-modal regulation attention unit to obtain the updated feature sequence, eliminate the tags irrelevant to the target through the collaborative tag elimination strategy.
7. The method for tracking an orbital platform target based on visible-infrared images according to claim 6, wherein, The elimination of the tags irrelevant to the target through the collaborative tag elimination strategy specifically includes: Determine the search area delimited in the current frame image; Extract tag features from the search area, and the tag features include tags from visible light and infrared modalities; For the tags in each search area, calculate the attention weight between it and the template tag; Obtain the fused weight according to the attention weights of the visible light and infrared modalities; Sort all the tags in the search area according to the fused weight, select the top k tags with the highest weight as target candidates, and delete the tags other than the top k tags.
8. An orbital platform target tracking device based on visible light-infrared images, characterized in that, Include: A preprocessing module for preprocessing the first visible light image based on the zero-shot image restoration strategy to obtain a second visible light image; The zero-shot image restoration strategy includes input initialization, back-projection guidance, and noise injection mechanism; A target tracking module for inputting the second visible light image and the infrared image into the cross-modal regulation attention model to obtain the predicted target position; the cross-modal regulation attention model includes a visible light-infrared tracker based on Transformer, a Transformer block, a cross-modal regulation attention unit, and a tracking prediction head.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for tracking the target of the rail transit platform based on visible light-infrared images according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for tracking the target of the rail transit platform based on visible light-infrared images according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target tracking method and system under visible light and infrared images
CN116758117A
Multi-attention RGBT target tracking method based on visible light guidance
CN118365675A
RGBT target tracking method based on convolution attention fusion
CN120088292A
Coarse-to-fine heterologous image matching method based on edge guidance
WO2024148969A1
Object detection method and apparatus, device, and storage medium
WO2024183181A1