Remote sensing target tracking method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610345212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-20
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-03-20
AI Technical Summary
现有目标跟踪技术主要基于单模态图像或传统多模态特征融合,对于跨模态的目标跟踪的研究仍具有较大的挑战性,主要存在跨模态一致性差、在轨成像扰动复杂导致精度低、跨星融合出现对象漂移或者丢失以及算力受限等问题,难以在跨模态、多帧和边缘设备条件下实现高精度、稳健的目标跟踪
[0020]在本申请实施例中,该方法在遥感数据序列依次确定目标帧并输入多模态大模型,以分别通过第一编码器和第二编码器对其中的第一图像和第二图像进行编码得到第一图像特征和第二图像特征,融合后得到目标图像特征。根据目标图像特征确定目标帧中至少一个检测目标对应的候选区域,再根据前一帧对应的目标识别结果和至少一个候选区域,确定目标帧的目标识别结果以及多模态大模型的输出结果。本申请通过两种类型的图像分别编码针对性的提取特征得到准确的图像特征,提高输出结果的准确性。同时,通过前一帧的识别结果辅助当前帧的识别保证连续跟踪稳健,避免跟踪目标漂移。
Smart Images

Figure CN122289952B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and includes, but is not limited to, a remote sensing target tracking method, apparatus, electronic device, and storage medium. Background Technology
[0002] Remote sensing image target tracking is of great significance in military monitoring, marine surveillance, airport security, disaster tracking, and other fields. Remote sensing images include multiple modalities such as optical and SAR (Synthetic Aperture Radar) images. Existing target tracking technologies are mainly based on single-modal images or traditional multimodal feature fusion. Research on cross-modal target tracking remains challenging, mainly due to poor cross-modal consistency, low accuracy caused by complex on-orbit imaging perturbations, object drift or loss during cross-satellite fusion, and computational limitations. It is difficult to achieve high-precision and robust target tracking under cross-modal, multi-frame, and edge device conditions. Summary of the Invention
[0003] In view of this, the remote sensing target tracking method, apparatus, electronic device and storage medium provided in the embodiments of this application can perform high-precision and robust target tracking based on remote sensing images.
[0004] The remote sensing target tracking method, apparatus, electronic device, and storage medium provided in this application are implemented as follows: One aspect of this application provides a remote sensing target tracking method, the method comprising: In the remote sensing data sequence, the i-th frame is determined as the target frame. The target frame includes a first image and a second image. The first image is an optical image, and the second image is a synthetic aperture radar image. The target frame is input into a multimodal large model, and the first image and the second image are encoded by the first encoder and the second encoder respectively to obtain the first image features and the second image features. The target image features are obtained by fusing the first image features and the second image features; Based on the features of the target image, determine at least one candidate region corresponding to the detected target in the target frame; Based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region, determine the target recognition result of the target frame and the output result of the multimodal large model.
[0005] In one possible implementation, the target image features are obtained by fusing the first image features and the second image features, including: The first image features and the second image features are pruned according to the pruning strategy; The first and second image features, after feature pruning, are weighted and fused to obtain the target image features.
[0006] In one possible implementation, the target recognition result includes at least one target location, target features, and target velocity corresponding to the detected target; Based on the target recognition result corresponding to the (i-1)th frame in multi-frame remote sensing data and at least one candidate region, determine the target recognition result of the target frame and the output result of the multimodal large model, including: Extract features from at least one candidate region to obtain candidate features for the corresponding detection target; Based on the candidate features of each detected target and the target features in the target recognition result corresponding to the detected target in the (i-1)th frame, the corresponding semantic consistency is calculated; Based on the candidate region location of each detected target, and the target position and target velocity in the target recognition result of the detected target in the (i-1)th frame, the corresponding motion consistency is calculated; The optimal solution is calculated based on at least one semantic consistency and motion consistency corresponding to each detected target, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution.
[0007] In one possible implementation, an optimal solution is calculated based on at least one semantic consistency and motion consistency corresponding to each detected target, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution, including: Calculate the matching confidence score based on at least one semantic consistency and motion consistency for each detected target; The optimal solution is determined based on the matching confidence of each detected target according to the Hungarian algorithm; Based on the candidate regions corresponding to each optimal solution and the target recognition results corresponding to the (i-1)th frame, determine the target position, target features and target velocity of at least one detected target in the target frame, and obtain the target recognition results of the target frame. The output of the multimodal large model is determined based on the candidate region corresponding to the optimal solution of each detection target and the matching confidence.
[0008] In one possible implementation, the method also includes: Determine a training set that includes multiple sets of sample remote sensing data sequences and the corresponding annotation results for each frame therein; Each frame of the remote sensing data sequence of each sample in the training set is used as the input of the multimodal large model. The prediction result corresponding to each frame is determined, and the model loss of the multimodal large model is determined based on the annotation result and prediction result corresponding to each frame in the remote sensing data sequence of each sample. Adjust the multimodal large model based on the model loss until the convergence condition is met.
[0009] In one possible implementation, each frame of the remote sensing data sequence from each sample in the training set is used as input to the multimodal large model. The prediction result corresponding to each frame is determined, and the model loss of the multimodal large model is determined based on the annotation result and prediction result corresponding to each frame in each sample remote sensing data sequence, including: For the i-th frame of the remote sensing data sequence of target samples in the training set, determine the first sample image and the second sample image therein; The prediction results, cross-modal contrast loss, and perturbation consistency loss are determined based on the first and second sample images of the multimodal large model. The matching loss is determined based on the annotation and prediction results of the i-th frame; The model loss of the multimodal large model is obtained by calculating the weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss.
[0010] In one possible implementation, the prediction result, cross-modal contrast loss, and perturbation consistency loss are determined based on the first and second sample images according to the multimodal large model, including: For the first sample image and the second sample image of the i-th frame of the target sample remote sensing data sequence, calculate the first similarity with each positive sample pair in the training set, and the second similarity with each negative sample pair. Positive samples are frames in the first sample image that match the detected target and frames in the second sample image that match the detected target. Calculate the cross-modal contrast loss based on the first and second similarities of each frame in the training set; And / or, based on the multimodal large model, the prediction results, cross-modal contrast loss, and perturbation consistency loss are determined based on the first and second sample images, including: Perturbations are added to the first sample image and the second sample image to obtain the first perturbation image and the second perturbation image; The first encoder extracts features from the first sample image and the first perturbation image respectively to obtain the first sample features and the first perturbation features; The second encoder is used to extract features from the second sample image and the second perturbation image respectively, so as to obtain the second sample features and the second perturbation features. The perturbation consistency loss is obtained by calculating the squared difference between the first sample feature and the first perturbation feature, and the sum of the squared differences between the second sample feature and the second perturbation feature. And / or, the method also includes: The importance of each channel in the first and second sample images is calculated based on the model loss, the features of the first sample, and the features of the second sample. The pruning strategy is determined based on the importance of each channel.
[0011] Another aspect of the embodiments of this application provides a remote sensing target tracking device, the device comprising: The target frame determination module is used to determine the i-th frame as the target frame in the remote sensing data sequence. The target frame includes a first image and a second image. The first image is an optical image and the second image is a synthetic aperture radar image. The image encoding module is used to input the target frame into the multimodal large model so that the first image and the second image are encoded by the first encoder and the second encoder respectively to obtain the first image features and the second image features. The feature fusion module is used to fuse the first image features and the second image features to obtain the target image features; The region recognition module is used to determine at least one candidate region corresponding to a detected target in the target frame based on the features of the target image. The result determination module is used to determine the target recognition result of the target frame and the output result of the multimodal large model based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region.
[0012] In one possible implementation, the feature fusion module is further used for: The first image features and the second image features are pruned according to the pruning strategy; The first and second image features, after feature pruning, are weighted and fused to obtain the target image features.
[0013] In one possible implementation, the target recognition result includes at least one target location, target features, and target velocity corresponding to the detected target; The result determination module is further used for: Extract features from at least one candidate region to obtain candidate features for the corresponding detection target; Based on the candidate features of each detected target and the target features in the target recognition result corresponding to the detected target in the (i-1)th frame, the corresponding semantic consistency is calculated; Based on the candidate region location of each detected target, and the target position and target velocity in the target recognition result of the detected target in the (i-1)th frame, the corresponding motion consistency is calculated; The optimal solution is calculated based on at least one semantic consistency and motion consistency corresponding to each detected target, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution.
[0014] In one possible implementation, the result determination module is further used for: Calculate the matching confidence score based on at least one semantic consistency and motion consistency for each detected target; The optimal solution is determined based on the matching confidence of each detected target according to the Hungarian algorithm; Based on the candidate regions corresponding to each optimal solution and the target recognition results corresponding to the (i-1)th frame, determine the target position, target features and target velocity of at least one detected target in the target frame, and obtain the target recognition results of the target frame. The output of the multimodal large model is determined based on the candidate region corresponding to the optimal solution of each detection target and the matching confidence.
[0015] In one possible implementation, the device further includes: The training set determination module is used to determine the training set, which includes multiple sets of sample remote sensing data sequences and the corresponding annotation results of each frame therein; The loss determination module is used to take each frame of the remote sensing data sequence of each sample in the training set as input to the multimodal large model, determine the prediction result corresponding to each frame, and determine the model loss of the multimodal large model based on the annotation result and prediction result corresponding to each frame in the remote sensing data sequence of each sample. The model tuning module is used to adjust the multimodal large model according to the model loss until the convergence condition is met.
[0016] In one possible implementation, the loss determination module is further used for: For the i-th frame of the remote sensing data sequence of target samples in the training set, determine the first sample image and the second sample image therein; The prediction results, cross-modal contrast loss, and perturbation consistency loss are determined based on the first and second sample images of the multimodal large model. The matching loss is determined based on the annotation and prediction results of the i-th frame; The model loss of the multimodal large model is obtained by calculating the weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss.
[0017] In one possible implementation, the loss determination module is further used for: For the first sample image and the second sample image of the i-th frame of the target sample remote sensing data sequence, calculate the first similarity with each positive sample pair in the training set, and the second similarity with each negative sample pair. Positive samples are frames in the first sample image that match the detected target and frames in the second sample image that match the detected target. Calculate the cross-modal contrast loss based on the first and second similarities of each frame in the training set; And / or, the loss determination module is further used for: Perturbations are added to the first sample image and the second sample image to obtain the first perturbation image and the second perturbation image; The first encoder extracts features from the first sample image and the first perturbation image respectively to obtain the first sample features and the first perturbation features; The second encoder is used to extract features from the second sample image and the second perturbation image respectively, so as to obtain the second sample features and the second perturbation features. The perturbation consistency loss is obtained by calculating the squared difference between the first sample feature and the first perturbation feature, and the sum of the squared differences between the second sample feature and the second perturbation feature. And / or, the device further includes: The importance determination module is used to calculate the importance of each channel in the first sample image and the second sample image based on the model loss, the features of the first sample, and the features of the second sample. The strategy determination module is used to determine the pruning strategy based on the importance of each channel.
[0018] The electronic device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.
[0019] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method provided in this application embodiment.
[0020] In this embodiment, the method sequentially determines target frames in a remote sensing data sequence and inputs them into a multimodal large model. First and second images are encoded using a first encoder and a second encoder respectively to obtain first image features and second image features, which are then fused to obtain target image features. Based on the target image features, candidate regions corresponding to at least one detected target in the target frame are determined. Then, based on the target recognition result of the previous frame and at least one candidate region, the target recognition result of the target frame and the output result of the multimodal large model are determined. This application obtains accurate image features by encoding two types of images separately and extracting features accordingly, improving the accuracy of the output results. Simultaneously, the recognition result of the previous frame assists the recognition of the current frame, ensuring robust continuous tracking and avoiding target drift. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a remote sensing target tracking method according to an embodiment of this application is shown; Figure 2 A schematic diagram illustrating the training process of a multimodal large model according to an embodiment of this application is shown. Figure 3 A schematic diagram illustrating a method for determining perturbation consistency loss according to an embodiment of this application is shown; Figure 4 A schematic diagram illustrating one method of determining matching loss according to an embodiment of this application is shown; Figure 5 A schematic diagram showing a target tracking result according to an embodiment of this application; Figure 6 A schematic diagram of a remote sensing target tracking device according to an embodiment of this application is shown; Figure 7 A schematic diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0025] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0026] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0027] The remote sensing target tracking method of this application embodiment can be executed by any electronic device, including but not limited to mobile phones, wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), tablet computers, laptops, vehicle terminals, PCs (Personal Computers), etc. Alternatively, it may include devices such as satellites. The functions implemented by this method can be achieved by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. Therefore, the electronic device includes at least a processor and a storage medium.
[0028] The remote sensing target tracking method of this application can be used in any application scenario that tracks targets based on remote sensing data. For example, it can be applied to the application scenario of identifying and tracking ships in the field of military and security monitoring. Alternatively, it can also be applied to the application scenario of identifying and tracking rescue teams in the field of emergency management and disaster prevention and mitigation.
[0029] The related technologies have the following problems when performing target tracking based on remote sensing data: Poor cross-modal consistency: Optical images are affected by illumination and cloud cover; SAR images are affected by speckle noise, resulting in inherent inconsistencies in the multimodal feature space. Direct fusion can easily lead to feature misalignment and tracking drift.
[0030] In-orbit imaging is subject to complex perturbations: satellite attitude jitter, orbital changes, and nadir oscillations can cause significant spatiotemporal shifts in the target under different modes. Traditional target tracking methods struggle to guarantee accuracy under low frame rate conditions.
[0031] Multi-satellite coordination is difficult: different orbits, perspectives and sensors cause inconsistencies in target positions, and cross-satellite fusion is prone to ID drift or target loss.
[0032] Spaceborne computing power is limited: direct deployment of large models results in high power consumption and high latency. Traditional pruning methods are difficult to take into account multimodal characteristics, which affects on-orbit real-time performance.
[0033] Therefore, the technical problem solved by the embodiments of this application is how to achieve high-precision and robust target tracking under cross-modal, multi-frame and edge device conditions.
[0034] The remote sensing target tracking scheme of this application embodiment will be described in detail below with reference to the accompanying drawings.
[0035] Figure 1 A flowchart illustrating a remote sensing target tracking method according to an embodiment of this application is shown. Figure 1 As shown, the remote sensing target tracking method of this application embodiment may include the following steps S10-S50.
[0036] For ease of description, the remote sensing target tracking method of this application embodiment is described using an electronic device as the execution subject. It should be understood that the execution subject of this application embodiment can also be a processor or chip in an electronic device, and this application embodiment does not impose any limitations.
[0037] Step S10: Determine the i-th frame as the target frame in the remote sensing data sequence.
[0038] In one possible implementation, the electronic device can acquire a sequence of remote sensing data comprising multiple remote sensing data in a temporal acquisition order, wherein each frame of remote sensing data includes different types of image data, such as optical images, synthetic aperture radar images, lidar images, and thermal infrared images. The different types of images may also include multiple image channels; for example, optical images may also include visible light images corresponding to different imaging bands.
[0039] Optionally, after acquiring the current remote sensing data sequence, the electronic device sequentially determines the i-th frame as the target frame according to the order of the remote sensing data, and extracts the first image and the second image included in the target frame. The first image is an optical image, and the second image is a synthetic aperture radar (SAR) image. Optical images and SAR images are two different modalities. Optical images are characterized by strong structural features and color, while SAR images are grayscale images with more noise and obvious texture features.
[0040] Step S20: Input the target frame into a multimodal large model to encode the first image and the second image respectively through the first encoder and the second encoder to obtain the first image features and the second image features.
[0041] In one possible implementation, after acquiring the first and second images of the target frame, the electronic device can input the target frame into a multimodal large model to identify whether a detection target exists in the target frame, and determine the location of the detection target if it exists. The detection target in this application embodiment varies depending on the application scenario. For example, in a target tracking scenario in the military and security monitoring field, the detection target could be a specific ship. In a target tracking scenario in the emergency management and disaster prevention and mitigation field, the detection target could be a vehicle of a rescue team.
[0042] In some embodiments, the multimodal large model may include different encoders for the first image and the second image, used to encode images of different modalities in a targeted manner to achieve feature extraction. That is, the first image and the second image of the target frame can be input into the first encoder and the second encoder in the multimodal large model, respectively, and the first image and the second image can be encoded by the first encoder and the second encoder, respectively, to obtain the first image features and the second image features.
[0043] Optionally, the first encoder and the second encoder described above can employ a lightweight Transformer structure, comprising several multi-head attention layers and feedforward network layers, for mapping optical images or SAR images to high-dimensional semantic feature vectors. The first encoder and the second encoder have the same structure but independent parameters to preserve modality-specific information. For example, in the first image and the second image respectively... , The first encoder is The second encoder is In this case, the first image feature can be determined as The second image feature is .
[0044] Step S30: Fuse the first image features and the second image features to obtain the target image features.
[0045] In one possible implementation, the multimodal large model may further include a feature fusion layer. After obtaining the first image features and the second image features through feature encoding, the first image features and the second image features can be input into the feature fusion layer for feature fusion to obtain the target image features. Optionally, to reduce redundant computation and lower hardware resource consumption and latency, embodiments of this application may first prune the first image features and the second image features according to a pruning strategy at the feature fusion layer. Then, the first image features and the second image features after feature pruning are weighted and fused to obtain the target image features. The pruning strategy may include at least one channel feature from the first image features to be removed, and at least one channel feature from the second image features.
[0046] In some embodiments, after pruning the first and second image features, the multimodal large model performs weighted fusion of the pruned channel features to obtain the target image features. This weighted fusion method can be expressed by the formula... To achieve, among which, These are channel features from the first or second image, features extracted from different modalities. These are the dynamic weights of each channel's features. After feature fusion of each channel is completed, the target image features are obtained and output by the feature fusion layer.
[0047] Step S40: Determine at least one candidate region corresponding to a detection target in the target frame based on the target image features.
[0048] In one possible implementation, the multimodal large model of this application embodiment may further include a target detection layer for identifying at least one detected target based on the target image features output by the feature fusion layer. This identification method can convert the target image features into location information corresponding to at least one detected target using a target tracking head, and output the location information corresponding to the detected target as a candidate region. The target detection layer can be a lightweight RPN or similar detection method, and each output candidate region has the corresponding identification information of the detected target.
[0049] Step S50: Based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region, determine the target recognition result of the target frame and the output result of the multimodal large model.
[0050] In one possible implementation, the candidate region of at least one detected target in the target frame determined by the multimodal large model is the target location independently identified based on the features of this frame. After determining the candidate region, the multimodal large model can also match the candidate region with the target detection result of the previous frame to associate the same target in consecutive frames, ensuring the stability of the detected target in the tracking result and avoiding target drift. The target recognition result includes the target location, target features, and target velocity corresponding to at least one detected target. The target location corresponding to each detected target is the position of the detected target in the corresponding frame, which can be represented by the center coordinates of the detection box containing the detected target. The target features corresponding to each detected target can be the image semantic features obtained by feature extraction of the region of the detected target within the detection box in the corresponding frame. Each detected target can be calculated using the target location of the detected target in at least two consecutive frames and the frame rate.
[0051] Optionally, in this embodiment of the application, after determining at least one candidate region of a detected target in the target frame and the target recognition result corresponding to the previous frame (i.e., the (i-1)th frame), the multimodal large model can first extract features from at least one candidate region to obtain candidate features of the corresponding detected target. Then, based on the candidate features of each detected target and the target features in the target recognition result corresponding to the detected target in the (i-1)th frame, the corresponding semantic consistency is calculated. Based on the position of the candidate region of each detected target and the target position and target velocity in the target recognition result corresponding to the detected target in the (i-1)th frame, the corresponding motion consistency is calculated. Based on at least one semantic consistency and motion consistency corresponding to each detected target, the optimal solution is calculated, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution.
[0052] In some embodiments, semantic consistency refers to the degree of matching between the semantic features of the detected target in the previous frame and the semantic features within the candidate region of the current frame. The semantic consistency between each detected target and its corresponding candidate region can be calculated using the formula... Calculate the candidate features of each detected target in candidate region i. Target features in the target recognition result of the i-th frame The cosine similarity.
[0053] in, It represents the cross-modal semantic feature vector of the candidate region corresponding to the i-th detected target in the target frame. It contains the semantic information of the detected target in the target frame, including the shape, texture and category of the detected target. It is the feature vector of the detected target in the previous frame, representing the semantic features of the target in historical frames. It represents the static features of the target at previous moments, including the target's shape, texture, and category information. ,when When =1, it means and They are completely similar; the two feature vectors point in the same direction, indicating that the target is the same in both frames. When =-1, it means and They are completely opposite, representing different goals; when =0 indicates and They are spatially unrelated.
[0054] In other embodiments, motion consistency refers to the dynamic characteristics of the detected target, i.e., the trajectory of the detected target over time. Since the position of the detected target changes over time during target tracking, a motion model is needed to predict the future position of the detected target, and historical motion information is then used for matching. Optionally, the motion consistency between each detected target and its corresponding candidate region can be calculated by first determining the target's position in the previous frame. and target speed To predict its position in the target frame. ,in, Indicates the inter-frame time interval. This indicates the position of the detected target in the target frame, obtained based on motion prediction.
[0055] Since the motion characteristics (such as velocity and direction) of each detected target are usually relatively stable over a short period of time during target tracking, it is possible to predict the position. A Gaussian kernel function is used to measure the matching degree between the predicted location and the candidate region. The closer the predicted location is to the center of the candidate region, the higher the motion consistency. This motion consistency is measured using a Gaussian kernel function. It means that, among them, This indicates the predicted position of the detected target in the target frame. The coordinates of the center position of the candidate region. This represents the motion uncertainty control parameter, which controls the uncertainty of the predicted position; a smaller value indicates better control. This indicates that the predictions of motion are more accurate.
[0056] In some embodiments, after determining the semantic and motion consistency between each detected target and each candidate region, the matching degree between the detected target and each candidate region is determined based on the semantic and motion consistency, and the detected target and candidate region are matched based on the matching degree. That is, the embodiments of this application can calculate the corresponding matching confidence based on at least one semantic and motion consistency corresponding to each detected target. The optimal solution is determined based on the matching confidence of each detected target according to the Hungarian algorithm. Based on the candidate regions corresponding to each optimal solution and the target recognition result corresponding to the (i-1)th frame, the target position, target features, and target velocity corresponding to at least one detected target in the target frame are determined, and the target recognition result of the target frame is obtained. The output result of the multimodal large model is determined based on the candidate regions corresponding to the optimal solution of each detected target and the matching confidence.
[0057] The matching confidence score corresponding to the detected target can be expressed by the formula. The weighted sum is obtained by calculating the sum, where, Semantic consistency weights are used to control semantic consistency. and motion consistency The proportion of influence in the final match can be dynamically adjusted. ,when When the size is large, multimodal large models pay more attention to semantic consistency. When the value is small or close to 0, the multimodal large model focuses more on motion consistency and is suitable for situations where the appearance of the target changes significantly or becomes blurred during motion. It is the confidence level of the match between the target and the i-th candidate region; It is the semantic consistency score, which represents the similarity between the semantic features of the detected target in the current frame and the target features in the previous frame; It is the motion consistency score, which represents the degree of matching between the predicted position of the detected target in the current frame and the position of the candidate region.
[0058] After determining the matching confidence of each detection target and each candidate region, the Hungarian algorithm can be used. Solve for the optimal match to ensure that the target in the target frame is optimally associated with the detected target in the previous frame. This is the index of the finally selected target candidate region, that is, the candidate region in the target frame with the highest matching degree with the target detected in the previous frame. The target candidate region index refers to the position representation or number of the candidate region in the target frame that best matches the target in the previous frame; that is, among all candidate regions, the region with the highest matching confidence is selected for target association. The Hungarian algorithm solves for the optimal match by minimizing or maximizing the cost in the cost matrix. Here, we want to maximize the matching confidence of the candidate region. In order to find the best match.
[0059] Based on the aforementioned technical features, in the model inference scenario, this embodiment of the application obtains accurate image features by encoding two types of images separately for targeted feature extraction, thereby improving the accuracy of the output results. Simultaneously, semantic feature matching and motion feature matching are performed between the recognition result of the previous frame and the current frame. The position of the detected target in the current frame is jointly determined based on the matching results, ensuring robust continuous tracking and preventing positional drift of the tracked target in different frames.
[0060] In one possible implementation, the electronic device in this embodiment may further perform model training before inference based on the multimodal large model, or periodically optimize and train the multimodal large model after inference is completed. Optionally, the model training process may include determining a training set comprising multiple sets of sample remote sensing data sequences and corresponding annotation results for each frame therein. Each frame of each sample remote sensing data sequence in the training set is used as input to the multimodal large model, the prediction result corresponding to each frame is determined, and the model loss of the multimodal large model is determined based on the annotation results and prediction results corresponding to each frame in each sample remote sensing data sequence. The multimodal large model is adjusted according to the model loss until the convergence condition is met.
[0061] Figure 2 This diagram illustrates the training process of a multimodal large model according to an embodiment of this application. Figure 2 As shown, the multimodal large model training framework of this application embodiment can be composed of three modules: a perturbation-invariant cross-modal semantic space module, a semantic-motion consistency fusion module, and an adaptive pruning module. During model training, the electronic device can use each frame of the remote sensing data sequence of each sample in the training set as input to the multimodal large model. The perturbation-invariant cross-modal semantic space module and the semantic-motion consistency fusion module calculate the model loss of the multimodal large model. Then, the adaptive pruning module calculates the importance of each channel in the first and second sample images based on the model loss, the features of the first sample, and the features of the second sample, and determines the pruning strategy based on the importance of each channel.
[0062] Optionally, during the training of the multimodal large model, the electronic device can determine the first sample image and the second sample image for the i-th frame of the remote sensing data sequence of the target samples in the training set. Then, the perturbation-invariant cross-modal semantic space module and the semantic-motion consistency fusion module determine the prediction result, cross-modal contrast loss, and perturbation consistency loss based on the first and second sample images according to the multimodal large model. The matching loss is then determined based on the annotation and prediction results of the i-th frame. The weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss is calculated to obtain the model loss of the multimodal large model. Further, the adaptive pruning module calculates the importance of each channel in the first and second sample images based on the model loss, the features of the first sample, and the features of the second sample, and then determines the pruning strategy based on the importance of each channel.
[0063] In some embodiments, the perturbation-invariant cross-modal semantic space module can apply controllable perturbations to optical and SAR images during training and introduce perturbation consistency constraints to learn shared semantic representations with cross-modal robustness. Specifically, this module can stabilize the image features of the first and second sample images by adding perturbations during the modeling process, and then input them into the first and second encoders respectively for encoding to extract image features. Finally, cross-modal contrastive loss and perturbation consistency loss are determined by contrastive learning on the extracted image features. The determination of the cross-modal loss ensures that the features of the same detected target obtained during training are as close as possible in the shared semantic space, while the features of different detected targets are far apart, ensuring optical-SAR alignment in the same semantic space. The determination of the consistency loss provides a foundation for subsequent semantic-motion consistent cross-frame association mechanisms and modal adaptive pruning, providing reliable information support for accurate target tracking.
[0064] Figure 3 This diagram illustrates a method for determining perturbation consistency loss according to an embodiment of this application. Figure 3 As shown, the method for determining the perturbation consistency loss in this embodiment of the application can be as follows: after determining the first sample image and the second sample image in any frame, perturbation is first added to the first sample image and the second sample image to obtain the first perturbation image and the second perturbation image. Then, feature extraction is performed on the first sample image and the first perturbation image respectively by the first encoder to obtain the first sample feature and the first perturbation feature. Feature extraction is then performed on the second sample image and the second perturbation image respectively by the second encoder to obtain the second sample feature and the second perturbation feature. The squared difference between the first sample feature and the first perturbation feature, and the sum of the squared differences between the second sample feature and the second perturbation feature are calculated to obtain the perturbation consistency loss.
[0065] Optionally, to ensure the stability of the multimodal large model against small noise in the input multimodal target, a cross-modal interpretable multi-level perturbation is designed. A perturbation generation function is introduced to simulate the input, improving the robustness of each modality image and ensuring the stability of features in each frame under perturbation. This includes perturbing the first sample image with random brightness variations, local occlusion, and noise superposition, and the second sample image with speckle noise, local speckle interference, and interpolation anomalies, to generate corresponding first and second perturbation images. The perturbation addition formula includes... and ,in, , These are perturbation samples from optical images and SAR images, respectively. This is the perturbation generation function.
[0066] In some embodiments, the perturbation method may be to randomly add noise to a first sample image of type optical image to simulate random errors in image data; to translate the pixel position of the image to simulate slight changes in target position; and to adjust the brightness and contrast of the image to simulate changes in illumination and image quality. In a second sample image of type SAR image, the image is scaled to simulate scale changes of different targets, and speckle noise is added to simulate noise in a real SAR image.
[0067] For example, the optical image perturbation generation process may include the addition of Gaussian noise. ,in The mean is The variance is The Gaussian distribution. This is a sample with noise added. Simulation of image translation. ,in The original image is in position pixels, It is the horizontal translation amount. This is the vertical translation amount. Brightness and contrast adjustment. , It is the contrast factor. It is a scaling factor. The SAR image perturbation generation process may include adding speckle noise. Scaling operation s is the scaling factor.
[0068] Optionally, the perturbation samples generated in this embodiment are random, with a random factor set to randomly generate different types of changes in the optical and SAR images. At this point, it still retains the image itself, and then, on top of this, an unperturbed SAR image and an optical image are mixed. Perturbation samples are generated for each mode. and The embodiments of this application can calculate the corresponding perturbation consistency loss DILoss, enabling the encoder to maintain a stable representation of the perturbation and ensuring consistent output features. The formula for calculating the perturbation consistency loss can be... The perturbation consistency loss is calculated by extracting features from the images before and after the perturbation is added into the encoder, transforming them into a set of feature vectors. and The similarity between optical images and SAR images is calculated using the motion consistency loss formula. If they are similar, the consistency is strong, which helps the model learn consistency information.
[0069] Considering the varying difficulty of the objectives, to improve the robustness of the multimodal large model to challenging scenarios, this embodiment of the application can also employ dynamic weights for the consistency of perturbations, adaptively adjusting according to the objective difficulty. The dynamic weights used in the perturbation consistency loss can ultimately result in the following loss calculation: ,in, It is a dynamic weight, based on the current input sample. and disturbance The difficulty is used to adjust the intensity of the loss.
[0070] In target recognition and tracking tasks, some targets are easy to identify and track correctly, while others are more difficult. For example, in images with complex backgrounds, it is difficult to distinguish the target from the background, and targets that are relatively small in the image are easily overlooked. Therefore, a "coefficient" is needed to adjust the state. This coefficient adjusts the difficulty of the current sample and the perturbed samples, assigning a larger loss weight to the difficult samples.
[0071] In other embodiments, after feature extraction of the optical image and SAR image by the dual-branch encoder, the two branches output... Two feature vectors are used to represent the semantic information of the optical image and the SAR image. This semantic information exists in both modalities and needs to be consistent for output. To ensure that the features of the same detected target in the optical image and the SAR image are similar, while the features of different detected targets are dissipated, this application introduces a cross-modal contrast loss. This cross-modal contrast loss is determined by calculating, for the first sample image and the second sample image of the i-th frame of the target sample remote sensing data sequence, a first similarity with each positive sample pair in the training set, and a second similarity with each negative sample pair. Positive samples are frames where the detected target in the first sample image matches the detected target in the second sample image, and negative samples are frames where the detected target in the first sample image matches the detected target in the second sample image. The cross-modal contrast loss is calculated based on the first and second similarities corresponding to each frame in the training set.
[0072] For example, embodiments of this application design a cross-modal alignment loss function to calculate cross-modal contrast loss, the alignment loss function being: , where N is the size of the frames included in the training set, i represents the i-th positive sample pair, and j represents all samples in the batch; As a negative sample candidate, represents the SAR feature of the j-th sample within the batch; τ is a temperature coefficient used to control the sharpness of the exponential function. The loss is to maximize the similarity of positive sample pairs relative to all negative sample pairs. For example, if a frame contains a ship target in both the optical and SAR images, that frame is identified as a positive sample pair. Conversely, if a frame contains a vehicle target in both the optical and SAR images, that frame is identified as a negative sample pair.
[0073] Figure 4 This diagram illustrates a method for determining matching loss according to an embodiment of this application. Figure 4 As shown, the electronic device can input any frame from any sample remote sensing data sequence into the multimodal large model during the training process to determine the prediction result. Specifically, when the electronic device determines the prediction result of the (i-1)th frame in the arbitrary sample remote sensing data sequence, the corresponding target recognition result can be obtained based on encoder feature extraction and subsequent target recognition. Further, after the electronic device inputs the i-th frame from the arbitrary sample remote sensing data sequence into the multimodal large model, feature extraction, fusion, and candidate region identification are performed sequentially. Then, based on the candidate region identification result, the prediction result of the (i-1)th frame, and the candidate region identification result of the current frame, the prediction result corresponding to the current frame is determined. After determining the prediction results for each frame in the training set, the cross-entropy loss is calculated based on the annotation results and prediction results of each frame to obtain the corresponding matching loss.
[0074] After obtaining the matching loss, perturbation consistency loss, and cross-modal contrast loss, a weighted sum of each loss can be calculated to obtain the corresponding model loss. After determining the model loss, the electronic device can also perform adaptive redundancy pruning based on the model loss to determine the pruning strategy. This adaptive redundancy pruning optimizes the model by removing unnecessary branches or channels, making it more efficient. The modal adaptive redundancy pruning method not only optimizes the model structure but also ensures cross-modal semantic consistency across different modalities, such as optical images and SAR images, while reducing redundant computation, lowering hardware resource consumption, and reducing latency.
[0075] Optionally, the core of this adaptive redundancy pruning is to score the importance of each channel in the image. To save computational resources and achieve higher efficiency in real-time and resource-constrained environments, embodiments of this application can perform pruning operations during both the training and inference phases. During the training phase, the goal of pruning is to optimize the model structure by dynamically evaluating the importance of each channel branch for each modality, thereby reducing redundant computation and improving training efficiency. The importance of each channel branch can be determined using a formula... The calculation yielded that, and These are the outputs of channels or feature layers in the optical and SAR branches, respectively. and It is modal adaptive weights. This represents the importance score of a specific channel or feature layer. and It is the gradient of the corresponding channel or feature layer with respect to the loss function L.
[0076] For example, in the feature extraction stage, branches of the optical image and the SAR image can be extracted by the encoder, for example... The optical branch represents the output feature of a specific channel or feature layer in an optical image, indicating the feature information extracted from the optical image. SAR branches represent the output features of a specific channel or feature layer in a SAR image, indicating the feature information extracted from the SAR image. After calculating the importance of each channel, a preset threshold is applied. and importance The ratio is used to determine the contribution of each channel to the task. If the corresponding importance is less than a threshold, it is determined that the channel's contribution to the task is insufficient, and the branch needs to be pruned. This reduces redundant computation in the model, lowers computing power and latency, while maintaining cross-modal semantic representation capabilities.
[0077] Figure 5 A schematic diagram illustrating a target tracking result according to an embodiment of this application is shown. Figure 5 As shown, the target tracking method based on the embodiments of this application performs target identification and tracking, and the obtained target tracking results are continuous and robust, greatly optimizing the target drift situation.
[0078] Based on the aforementioned technical features, this application constructs a perturbation-invariant cross-modal semantic space, effectively overcoming modal differences and imaging interference between optical and SAR images, and significantly improving the consistency of target features in complex environments. Simultaneously, this application combines a semantic-motion consistency cross-frame association mechanism for target tracking, significantly reducing the target drift rate in tracking results between different frames, achieving continuous and stable target tracking. Furthermore, through a modal adaptive redundancy pruning strategy for spaceborne GPUs, the computational load and memory consumption of the model are significantly reduced while maintaining high accuracy, enabling efficient real-time deployment on spaceborne edge devices. This solution combines high robustness, high accuracy, and lightweight advantages, and can be widely applied in fields such as military reconnaissance, marine monitoring, and disaster emergency response, powerfully promoting the technological leap from ground analysis to on-orbit real-time remote sensing intelligent processing.
[0079] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0080] Based on the foregoing embodiments, this application provides a remote sensing target tracking device, which includes a control module. The control module can implement the remote sensing target tracking method of this application embodiment through a processor; of course, the remote sensing target tracking method can also be implemented through specific logic circuits. In the implementation process, the processor of the control device can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.
[0081] Figure 6 A schematic diagram of a remote sensing target tracking device according to an embodiment of this application is shown. Figure 6 As shown, the remote sensing target tracking device in this application embodiment includes: The target frame determination module 60 is used to determine the i-th frame as the target frame in the remote sensing data sequence. The target frame includes a first image and a second image. The first image is an optical image and the second image is a synthetic aperture radar image. Image encoding module 61 is used to input the target frame into a multimodal large model so that the first image and the second image are encoded by the first encoder and the second encoder respectively to obtain the first image features and the second image features. Feature fusion module 62 is used to fuse first image features and second image features to obtain target image features; Region recognition module 63 is used to determine candidate regions corresponding to at least one detected target in the target frame based on the features of the target image; The result determination module 64 is used to determine the target recognition result of the target frame and the output result of the multimodal large model based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region.
[0082] In one possible implementation, the feature fusion module 62 is further used for: The first image features and the second image features are pruned according to the pruning strategy; The first and second image features, after feature pruning, are weighted and fused to obtain the target image features.
[0083] In one possible implementation, the target recognition result includes at least one target location, target features, and target velocity corresponding to the detected target; Result determination module 64 is further used for: Extract features from at least one candidate region to obtain candidate features for the corresponding detection target; Based on the candidate features of each detected target and the target features in the target recognition result corresponding to the detected target in the (i-1)th frame, the corresponding semantic consistency is calculated; Based on the candidate region location of each detected target, and the target position and target velocity in the target recognition result of the detected target in the (i-1)th frame, the corresponding motion consistency is calculated; The optimal solution is calculated based on at least one semantic consistency and motion consistency corresponding to each detected target, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution.
[0084] In one possible implementation, the result determination module 64 is further used for: Calculate the matching confidence score based on at least one semantic consistency and motion consistency for each detection target; The optimal solution is determined based on the matching confidence of each detected target according to the Hungarian algorithm; Based on the candidate regions corresponding to each optimal solution and the target recognition results corresponding to the (i-1)th frame, determine the target position, target features and target velocity of at least one detected target in the target frame, and obtain the target recognition results of the target frame. The output of the multimodal large model is determined based on the candidate region corresponding to the optimal solution of each detection target and the matching confidence.
[0085] In one possible implementation, the device further includes: The training set determination module is used to determine the training set, which includes multiple sets of sample remote sensing data sequences and the corresponding annotation results of each frame therein; The loss determination module is used to take each frame of the remote sensing data sequence of each sample in the training set as input to the multimodal large model, determine the prediction result corresponding to each frame, and determine the model loss of the multimodal large model based on the annotation result and prediction result corresponding to each frame in the remote sensing data sequence of each sample. The model tuning module is used to adjust the multimodal large model according to the model loss until the convergence condition is met.
[0086] In one possible implementation, the loss determination module is further used for: For the i-th frame of the remote sensing data sequence of target samples in the training set, determine the first sample image and the second sample image therein; The prediction results, cross-modal contrast loss, and perturbation consistency loss are determined based on the first and second sample images of the multimodal large model. The matching loss is determined based on the annotation and prediction results of the i-th frame; The model loss of the multimodal large model is obtained by calculating the weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss.
[0087] In one possible implementation, the loss determination module is further used for: For the first sample image and the second sample image of the i-th frame of the target sample remote sensing data sequence, calculate the first similarity with each positive sample pair in the training set, and the second similarity with each negative sample pair. Positive samples are frames in the first sample image that match the detected target and frames in the second sample image that match the detected target. Calculate the cross-modal contrast loss based on the first and second similarities of each frame in the training set; And / or, the loss determination module is further used for: Perturbations are added to the first sample image and the second sample image to obtain the first perturbation image and the second perturbation image; The first encoder extracts features from the first sample image and the first perturbation image respectively to obtain the first sample features and the first perturbation features; The second encoder is used to extract features from the second sample image and the second perturbation image respectively, so as to obtain the second sample features and the second perturbation features. The perturbation consistency loss is obtained by calculating the squared difference between the first sample feature and the first perturbation feature, and the sum of the squared differences between the second sample feature and the second perturbation feature. And / or, the device further includes: The importance determination module is used to calculate the importance of each channel in the first sample image and the second sample image based on the model loss, the features of the first sample, and the features of the second sample. The strategy determination module is used to determine the pruning strategy based on the importance of each channel.
[0088] It should be noted that, in the embodiments of this application... Figure 6 The module division of the remote sensing target tracking device shown is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, exist as separate physical units, or be integrated into one unit by two or more units. The integrated units can be implemented in hardware, as software functional units, or a combination of both.
[0089] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0090] Figure 7 A schematic diagram of an electronic device according to an embodiment of this application is shown. For example... Figure 7 As shown in the figure, this application provides an electronic device, which can be a server, and its internal structure diagram can be as follows. Figure 7 As shown, the electronic device includes a processor 720, a memory, and a transceiver 740 connected via a system bus 710. The processor 720 provides computing and control capabilities. The memory includes a non-volatile storage medium 731 and internal memory 732. The non-volatile storage medium 731 stores an operating system, computer programs, and a database. The internal memory 732 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium 731. The database stores data. The transceiver 740 communicates with an external terminal via a network connection. The computer program is executed by the processor 720 to implement the aforementioned methods.
[0091] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor 720, implements the steps of the method provided in the above embodiments.
[0092] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0093] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0094] In one possible implementation, the shooting prompting device provided in this application can be implemented as a computer program, which can be configured as follows: Figure 7 The device operates on the electronic device shown. The memory of the electronic device can store the various program modules that make up the above-described apparatus. The computer program composed of the various program modules causes the processor 720 to execute the steps of the methods in the various embodiments of this application described in this specification.
[0095] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium, storage medium, and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0096] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, phrases such as "in one possible implementation," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0097] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0098] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0100] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0101] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0102] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0103] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0104] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0105] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0106] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0107] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A remote sensing target tracking method, characterized in that, The method includes: In the remote sensing data sequence, the i-th frame is determined as the target frame. The target frame includes a first image and a second image. The first image is an optical image, and the second image is a synthetic aperture radar image. The target frame is input into a multimodal large model to encode the first image and the second image respectively through a first encoder and a second encoder to obtain the first image features and the second image features; The process of fusing the first image features and the second image features to obtain target image features includes: pruning the first image features and the second image features according to a pruning strategy; and performing weighted fusion of the pruned first image features and the second image features to obtain target image features. Based on the target image features, determine at least one candidate region corresponding to a detected target in the target frame; Based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region, determine the target recognition result of the target frame and the output result of the multimodal large model; The importance of each channel in the first sample image and the second sample image is calculated based on the model loss, the first sample feature, and the second sample feature. The first sample feature is the feature extracted from the first sample image by the first encoder, and the second sample feature is the feature extracted from the second sample image by the second encoder. The pruning strategy is determined based on the importance of each channel; The calculation process of the model loss includes: For the i-th frame of the remote sensing data sequence of target samples in the training set, determine the first sample image and the second sample image therein; The prediction result, cross-modal contrast loss, and perturbation consistency loss are determined based on the first sample image and the second sample image according to the multimodal large model. The matching loss is determined based on the annotation and prediction results of the i-th frame; The model loss of the multimodal large model is obtained by calculating the weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss.
2. The method according to claim 1, characterized in that, The target recognition result includes at least one target location, target features, and target velocity corresponding to the detected target; The step of determining the target recognition result of the target frame and the output result of the multimodal large model based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region includes: Extract features from at least one of the candidate regions to obtain candidate features corresponding to the detection target; Based on the candidate features of each detected target and the target features in the target recognition result of the detected target in the (i-1)th frame, the corresponding semantic consistency is calculated; Based on the candidate region location of each detected target, and the target position and target velocity in the target recognition result of the detected target in the (i-1)th frame, the corresponding motion consistency is calculated; The optimal solution is calculated based on at least one semantic consistency and motion consistency corresponding to each of the detected targets, and the target recognition result of the target frame and the output result of the multimodal large model are determined based on the candidate region corresponding to the optimal solution.
3. The method according to claim 2, characterized in that, The step of calculating the optimal solution based on at least one semantic consistency and motion consistency corresponding to each detected target, and determining the target recognition result of the target frame and the output result of the multimodal large model based on the candidate region corresponding to the optimal solution, includes: Calculate the matching confidence score based on at least one semantic consistency and motion consistency for each of the detected targets; The optimal solution is determined based on the matching confidence of each of the aforementioned detection targets according to the Hungarian algorithm; Based on the candidate regions corresponding to each optimal solution and the target recognition result corresponding to the (i-1)th frame, determine the target position, target features and target velocity of at least one detected target in the target frame, and obtain the target recognition result of the target frame; The output of the multimodal large model is determined based on the candidate region corresponding to the optimal solution of each detection target and the matching confidence.
4. The method according to claim 1, characterized in that, The method further includes: Determine a training set that includes multiple sets of sample remote sensing data sequences and the corresponding annotation results for each frame therein; Each frame of the remote sensing data sequence of each sample in the training set is used as the input of the multimodal large model. The prediction result corresponding to each frame is determined, and the model loss of the multimodal large model is determined based on the annotation result and prediction result corresponding to each frame in the remote sensing data sequence of each sample. The multimodal large model is adjusted according to the model loss until the convergence condition is met.
5. The method according to claim 4, characterized in that, The step of determining the prediction result, cross-modal contrast loss, and perturbation consistency loss based on the multimodal large model using the first sample image and the second sample image includes: For the first sample image and the second sample image of the i-th frame of the target sample remote sensing data sequence, calculate the first similarity with each positive sample pair in the training set, and the second similarity with each negative sample pair. The positive sample is the frame in the first sample image where the detected target matches the frame in the second sample image, and the negative sample is the frame in the first sample image where the detected target does not match the frame in the second sample image. Calculate the cross-modal contrast loss based on the first similarity and second similarity corresponding to each frame in the training set; The step of determining the prediction result, cross-modal contrast loss, and perturbation consistency loss based on the multimodal large model using the first sample image and the second sample image includes: Perturbations are added to the first sample image and the second sample image to obtain a first perturbation image and a second perturbation image; The first encoder is used to extract features from the first sample image and the first perturbation image respectively to obtain the first sample feature and the first perturbation feature; The second encoder is used to extract features from the second sample image and the second perturbation image respectively, to obtain the second sample features and the second perturbation features; The perturbation consistency loss is obtained by calculating the sum of the squared difference between the first sample feature and the first perturbation feature, and the squared difference between the second sample feature and the second perturbation feature.
6. A remote sensing target tracking device, characterized in that, The device includes: The target frame determination module is used to determine the i-th frame as the target frame in the remote sensing data sequence. The target frame includes a first image and a second image, wherein the first image is an optical image and the second image is a synthetic aperture radar image. The image encoding module is used to input the target frame into a multimodal large model, so as to encode the first image and the second image respectively through the first encoder and the second encoder to obtain the first image features and the second image features; The feature fusion module is used to fuse the first image features and the second image features to obtain target image features, including: pruning the first image features and the second image features according to a pruning strategy; and performing weighted fusion on the first image features and the second image features after feature pruning to obtain target image features. The region identification module is used to determine, based on the features of the target image, at least one candidate region corresponding to a detected target in the target frame; The result determination module is used to determine the target recognition result of the target frame and the output result of the multimodal large model based on the target recognition result corresponding to the (i-1)th frame in the multi-frame remote sensing data and at least one candidate region. The importance of each channel in the first sample image and the second sample image is calculated based on the model loss, the first sample feature, and the second sample feature. The first sample feature is the feature extracted from the first sample image by the first encoder, and the second sample feature is the feature extracted from the second sample image by the second encoder. The pruning strategy is determined based on the importance of each channel; The calculation process of the model loss includes: For the i-th frame of the remote sensing data sequence of target samples in the training set, determine the first sample image and the second sample image therein; The prediction result, cross-modal contrast loss, and perturbation consistency loss are determined based on the first sample image and the second sample image according to the multimodal large model. The matching loss is determined based on the annotation and prediction results of the i-th frame; The model loss of the multimodal large model is obtained by calculating the weighted sum of the cross-modal contrast loss, perturbation consistency loss, and matching loss.
7. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
FPGA-based neural network acceleration method and system
CN117521752A
An apparatus, a method and a computer program for running a neural network
EP3633990A1