A Remote Sensing Target Open Set Detection and Recognition Method Based on Text Retrieval
Through the remote sensing target open-set detection and recognition method based on text retrieval, the knowledge distillation method of multimodal fusion features and diffusion model is used to solve the problems of insufficient generalization capabilities of the model and errors in multi-objective recognition in the prior art, and more accurate and efficient remote sensing target detection and recognition are achieved.
Patent Information
- Application Number
- CN202410846208.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-06-27
AI Technical Summary
The prior art performs poorly when applying identification models on closed-set data to untrained sample data, and is prone to misclassification or position labeling errors when multiple objects of the same type in the image.
The remote sensing target open-set detection and recognition method based on text retrieval is adopted. Through the alignment and enhancement of multimodal fusion features, combined with the knowledge distillation method of the diffusion model, the alignment and implicit fusion of the target features and pre-trained visual-text features are achieved, and the generalization ability of the detection model is improved.
Deeply mining the cross-domain and cross-modal feature correlation of remote sensing data, improves the detection and identification ability of open set targets, enhances the generalization ability of the model, and reduces the possibility of misclassification and position labeling.
Smart Images

Figure CN118864809B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision computing technology, and in particular to an open set detection and recognition method for remote sensing targets based on text retrieval. Background Art
[0002] Target detection and recognition technology is one of the core issues in the field of computer vision. Its task is to find all the targets (objects) of interest in the image data and determine their categories and locations. In recent years, with the development of artificial intelligence technology, target detection and recognition technology based on deep learning has gradually been used in complex application scenarios such as industrial production, national defense and military. Among them, the use of target detection to achieve intelligent recognition of remote sensing data is a hot research field with great development prospects. By matching the user's text retrieval input, and then finding feature matching points from the feature map, and setting the prior frame to filter and adjust it, to obtain the final prediction frame, the target can be detected and recognized. This technology can achieve accurate visual perception of remote sensing visual images, providing users with a deeper panoramic understanding.
[0003] In recent years, the performance of deep learning target detection and recognition technology has been improved. However, researchers have found that models with good recognition performance on closed-set data perform poorly when applied to sample data that has never been trained and has no accompanying information. This is because the training images and application images have two completely unrelated data distributions. At the same time, when there are multiple objects of the same type in the image, the model may misclassify them as a whole or mislabel their locations. In order to solve the above problems,
[0004] Therefore, it is an urgent problem for those skilled in the art to propose an open set detection and recognition method for remote sensing targets based on text retrieval to solve the difficulties existing in the prior art. Summary of the invention
[0005] In view of this, the present invention provides an open-set detection and recognition method for remote sensing targets based on text retrieval. Based on the multimodal fusion features of remote sensing data, the present invention first uses a query embedding set to perform semantically independent detection to capture multi-dimensional features of scene targets, and then uses a knowledge distillation method based on a diffusion model to complete the alignment (semantic feature similarity) and implicit fusion between target features and pre-trained visual-text features, thereby completing the efficient transfer of knowledge and improving the generalization ability of the detection model from visible data sets to unknown data sets.
[0006] In order to achieve the above object, the present invention adopts the following technical solution:
[0007] A remote sensing target open set detection and recognition method based on text retrieval comprises the following steps:
[0008] S1, feature extraction step: obtain remote sensing raw data, process the remote sensing raw data based on the perception model, obtain the remote sensing visual features and text features of the image and map them to the visual-text joint embedding space;
[0009] S2, feature alignment step: the visual features and text features obtained in S1 are aligned in the visual-text joint embedding space through the alignment module to obtain aligned features;
[0010] S3, feature enhancement step: enhancing the features after alignment processing obtained in S2 to obtain target visual-text multimodal fusion features;
[0011] S4, feature conversion step: based on the target feature set and the diffusion model, align and fuse to obtain the open set category features;
[0012] S5, optimization model step: detect and identify the image based on the open set category features obtained in S4, calculate the loss value based on the prediction result, and optimize the target detection network model according to the loss value; based on the optimized target detection network model, perform target detection and identification on the remote sensing open set data.
[0013] In the above method, optionally, the perception model content in S1 includes:
[0014] The visual encoding network and text encoding network are used to map remote sensing vision and text features into the visual-text joint embedding space respectively, and the mapping process is optimized using the bidirectional sorting loss function.
[0015] In the above method, optionally, the content of S2 includes:
[0016] S201, multimodal feature completion, using the generator model to output the completed features of visual and textual modalities, and using adversarial learning strategies to evaluate the completed features;
[0017] S202, alignment algorithm, in the visual-text joint embedding space, takes the center of the multimodal features of the same concept as the anchor point, shortens the distance between the multimodal features of the same concept object and evenly distributes them around the center anchor point, and makes the multimodal features of different semantic objects stay away from mutual exclusion;
[0018] S203: Using visual features and text features as positive and negative sample pairs in a comparative learning process, updating parameters as the model learning process progresses, and obtaining aligned features.
[0019] In the above method, optionally, the feature enhancement step in S3 includes:
[0020] S301, feature enhancement: for the features after alignment processing obtained in S203, feature enhancement is performed through complementary masks in a multimodal feature enhancement module to obtain an incomplete single-modal feature vector, and then remote sensing visual enhancement features and text enhancement features are obtained through mask feature reconstruction technology;
[0021] S302, multimodal feature fusion: input remote sensing visual enhancement features and text enhancement features into the cross-modal self-attention Transformer module to obtain multimodal complementary features after cross-domain cross-modal attention weighting;
[0022] The calculation formula is as follows:
[0023]
[0024] Among them, CrossAtt(·) calculates cross-modal attention, m is a single modality, n is the target sample, Q m is the feature vector of the visible part after masking, K n and V n are the other modal eigenvectors, d k Indicates Q m and K n The characteristic dimension of
[0025] S303: Fusing the multimodal complementary features obtained in S302 with the unimodal features after mask processing, providing guidance for the mask reconstruction process based on the cross-domain multimodal complementary information, enhancing the in-domain features of modal complementarity, and outputting the cross-domain multimodal fusion feature F fuse .
[0026] In the above method, optionally, the content of S4 includes:
[0027] S401, taking the cross-domain multimodal fusion features of the remote sensing data obtained in S3 as input, performing block processing to obtain a plurality of block features of the same size, and adding position coding to different block features using a sinusoidal coding rule to obtain block features with position information;
[0028] S402, randomly initialize the target feature set to convert the target into a learnable query embedding set, and adopt a self-attention transformation model to continuously extract key semantic structure information from the input cross-domain multimodal fusion features, and then use the cross-attention mechanism to drive the initialized query embedding set to capture different target structure information, thereby forming a captured target feature set;
[0029] S403, using a decoder that only focuses on the position information in the feature set to regress the prediction frame, and during the training process, comparing the prediction set y with the target true value Matching associations, specifically deconstructing it into a bipartite graph matching task, so that one query embedding in the set corresponds to one target, where the number of query embeddings should be much larger than the number of targets;
[0030] S404, using the diffusion model to transfer the high confidence prior knowledge obtained by pre-training, and integrating the semantic structure information in S402 into the capture target features;
[0031] S405. Classify the open set targets using a non-parametric metric method, map the features of the query feature set and the target semantic features extracted by the CLIP text encoder to the same feature space by minimizing the cosine distance to achieve alignment, and then use an adaptive fusion network to achieve selective fusion, thereby obtaining the open set category features.
[0032] The above method optionally uses the Hungarian matching function to calculate the similarity loss between the regressed predicted box and the true box, as shown in the following formula:
[0033]
[0034] Where N represents the total number of targets in the prediction set or the true set. Represents the category similarity loss between the predicted box and the true box, y is the predicted set, is the target true value, is the predicted probability, c i For the prediction sample, Represents the similarity loss between the predicted box and the true box of the regression.
[0035] It can be seen from the above technical solution that, compared with the prior art, the remote sensing target open set detection and recognition method based on text retrieval of the present invention has the following beneficial effects:
[0036] 1. Deeply explore the correlation between cross-domain and cross-modal (visual-text pair) features of remote sensing data, so as to synchronously align and fuse cross-domain multimodal features and give full play to the advantages of multi-domain complementarity;
[0037] 2. Aiming at the problem of open set detection with uncertain detection targets, this paper starts from the perspective of prior knowledge transfer and utilizes the knowledge distillation method based on the diffusion model to transfer the high-dimensional semantic information obtained in the large-scale pre-trained visual-text model to realize open set target detection and recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0039] Figure 1 A flow chart of a remote sensing target open set detection and recognition method based on text retrieval provided by the present invention;
[0040] Figure 2 A schematic diagram of feature extraction provided by the present invention;
[0041] Figure 3 A schematic diagram of joint feature space alignment provided by the present invention;
[0042] Figure 4 A schematic diagram of cross-module feature fusion provided by the present invention;
[0043] Figure 5 This is a schematic diagram of target detection provided by the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0046] The present invention can be used in many general or special computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or devices, etc.
[0047] Reference Figure 1 As shown, the present invention discloses a remote sensing target open set detection and recognition method based on text retrieval, comprising the following steps:
[0048] S1, feature extraction step: obtain remote sensing raw data, process the remote sensing raw data based on the perception model, obtain the remote sensing visual features and text features of the image and map them to the visual-text joint embedding space;
[0049] S2, feature alignment step: the visual features and text features obtained in S1 are aligned in the visual-text joint embedding space through the alignment module to obtain aligned features;
[0050] S3, feature enhancement step: enhancing the features after alignment processing obtained in S2 to obtain target visual-text multimodal fusion features;
[0051] S4, feature conversion step: based on the target feature set and the diffusion model, align and fuse to obtain the open set category features;
[0052] S5, optimization model step: detect and identify the image based on the open set category features obtained in S4, calculate the loss value based on the prediction result, and optimize the target detection network model according to the loss value; based on the optimized target detection network model, perform target detection and identification on the remote sensing open set data.
[0053] Furthermore, the perception model content in S1 includes:
[0054] The visual encoding network and text encoding network are used to map remote sensing vision and text features into the visual-text joint embedding space respectively, and the mapping process is optimized using the bidirectional sorting loss function.
[0055] Specifically, Figure 2 As shown, the perception model in step S1 mainly uses the visual encoding network E v And the text encoding network E t The visual feature f v (v) and text features f t (t) are mapped to the visual-text joint embedding space. Where D represents the input layer and M represents the feature intermediate layer. Assuming the original remote sensing input data X, the modal feature can be expressed as:
[0056]
[0057] f v (v=E v (X v )
[0058] f t (t) = E t (X t )
[0059] Using bidirectional sorting loss The mapping process is optimized, and the specific expression is as follows:
[0060]
[0061] in, Represents the size of each batch, m is a hyperparameter, cos(·,·) represents cosine similarity, [·] + represents the hinge function max(0,·).
[0062] Furthermore, the content of S2 includes:
[0063] S201, multimodal feature completion, using the generator model to output the completed features of visual and textual modalities, and using adversarial learning strategies to evaluate the completed features;
[0064] S202, alignment algorithm, such as Figure 3 As shown in the figure, in the visual-text joint embedding space, the center of the multimodal features of the same concept is used as the anchor point, the multimodal features of the same concept objects are shortened and evenly distributed around the central anchor point, and the multimodal features of different semantic objects are kept away from mutual exclusion;
[0065] S203: Using visual features and text features as positive and negative sample pairs in a comparative learning process, updating parameters as the model learning process progresses, and obtaining aligned features.
[0066] Specifically, step S2 needs to perform multimodal feature completion before feature alignment, taking the cross-domain fusion features as knowledge priors, and using a generator model based on the adversarial learning strategy to output the completed features of the visual and textual modalities. The same adversarial learning strategy is used to evaluate the completed features to enhance the rationality of the missing modality feature completion results.
[0067] The loss function is calculated during training as follows:
[0068]
[0069] Where N represents the total number of training samples, M represents the total number of modalities, and Δ nm is 0 or 1, which is used to indicate whether mode m is missing. n To integrate features across domains, is the original feature of the missing mode, and They represent contrastive learning loss and missing modality reconstruction loss respectively, and the overall loss function is given by and The weighted sum of the two, λ 1 and 2 Represents the weight coefficient.
[0070] The feature alignment method can be specifically explained as taking the center of the multimodal features of the same concept as the anchor point in the joint visual-text embedding space, shortening the distance of the multimodal features of the same concept object as much as possible and evenly distributing them around the center anchor point, and making the multimodal features of different semantic objects far away from mutual exclusion. Finally, the multimodal features of the above two types of properties are used as positive and negative sample pairs in the comparative learning process, and the parameters are continuously updated as the model learning process progresses. Finally, the alignment loss function is as follows:
[0071]
[0072] in is the similarity metric score between the current input modality feature and the multimodal feature center (positive sample pair), is the similarity measure score between the current input object feature and other object features (negative sample pairs), τ is used to adjust the model's attention to positive and negative sample pairs, N is the set of all object samples, and M is the set of all modalities.
[0073] Further, such as Figure 4 As shown, the feature enhancement steps in S3 include:
[0074] S301, feature enhancement: for the features after alignment processing obtained in S203, feature enhancement is performed through complementary masks in a multimodal feature enhancement module to obtain an incomplete single-modal feature vector, and then remote sensing visual enhancement features and text enhancement features are obtained through mask feature reconstruction technology;
[0075] S302, multimodal feature fusion: input remote sensing visual enhancement features and text enhancement features into the cross-modal self-attention Transformer module to obtain multimodal complementary features after cross-domain cross-modal attention weighting;
[0076] The calculation formula is as follows:
[0077]
[0078] Among them, CrossAtt(·) calculates cross-modal attention, m is a single modality, n is the target sample, Q m is the feature vector of the visible part after masking, K n and V n are the other modal eigenvectors, d k Indicates Qm and K n The characteristic dimension of
[0079] S303: Fusing the multimodal complementary features obtained in S302 with the unimodal features after mask processing, providing guidance for the mask reconstruction process based on the cross-domain multimodal complementary information, enhancing the in-domain features of modal complementarity, and outputting the cross-domain multimodal fusion feature F fuse .
[0080] Specifically, the calculation result obtained by the above formula is the multimodal complementary feature after cross-domain and cross-modal attention weighting, which is further fused with the unimodal feature after masking, and the fused feature vector is used to predict the original content of the masked part. Since the original feature vector before masking is known, there is no need to provide manual annotation information during the training process. The prediction results of the masked autoencoder can be supervised directly by reading the original feature data, and the reconstruction loss between the prediction results and the original features can be calculated.
[0081] Different from the traditional mask autoencoder that only considers a single modality, this scheme uses the complementary information of cross-domain multimodal to guide the mask reconstruction process, thereby achieving the intra-domain feature enhancement of modal complementarity, and finally outputs the cross-domain multimodal fusion feature F fuse .
[0082] Further, such as Figure 5 As shown, the contents of S4 include:
[0083] S401, taking the cross-domain multimodal fusion features of the remote sensing data obtained in S3 as input, performing block processing to obtain a plurality of block features of the same size, and adding position coding to different block features using a sinusoidal coding rule to obtain block features with position information;
[0084] S402, randomly initialize the target feature set to convert the target into a learnable query embedding set, and adopt a self-attention transformation model to continuously extract key semantic structure information from the input cross-domain multimodal fusion features, and then use the cross-attention mechanism to drive the initialized query embedding set to capture different target structure information, thereby forming a captured target feature set;
[0085] S403, using a decoder that only focuses on the position information in the feature set to regress the prediction frame, and during the training process, comparing the prediction set y with the target true value Matching associations, specifically deconstructing it into a bipartite graph matching task, so that one query embedding in the set corresponds to one target, where the number of query embeddings should be much larger than the number of targets;
[0086] S404, using the diffusion model to transfer the high confidence prior knowledge obtained by pre-training, and integrating the semantic structure information in S402 into the capture target features;
[0087] S405. Classify the open set targets using a non-parametric metric method, map the features of the query feature set and the target semantic features extracted by the CLIP text encoder to the same feature space by minimizing the cosine distance to achieve alignment, and then use an adaptive fusion network to achieve selective fusion, thereby obtaining the open set category features.
[0088] Specifically, after using the inverse process of the diffusion model to accurately remove standard Gaussian noise, the semantic structure information extracted by the CLIP pre-trained image model is effectively integrated into the captured target feature set.
[0089] When training for the open set recognition task, a non-parametric metric will be used for classification because the number of categories is not fixed. Given that CLIP image features are embedded in the open set multimodal fusion feature set, and the image features extracted by the pre-trained CLIP are naturally aligned with the text features, open set target recognition can be achieved by fine-tuning the CLIP text encoder model. By minimizing the cosine distance, the model can effectively map the features of the query feature set and the target semantic features extracted by the CLIP text encoder to the same feature space to achieve alignment, and then use the adaptive fusion network to achieve selective fusion, that is, to obtain the open set category features by alignment and fusion.
[0090] To achieve open set recognition, the diffusion model is used to transfer pre-trained prior knowledge. In the process of adding noise, the original features are kept unchanged while the open set knowledge is integrated. The KL divergence is used to constrain the open set feature x to force its feature distribution to tend to the Gaussian distribution N(μ,I), so as to naturally integrate it into the capture target feature set. In the inverse denoising process, the network is trained to learn the conditional distribution probability to remove the Gaussian noise with the distribution function N(0,I). The training loss is as follows:
[0091]
[0092] Among them, t is the current diffusion layer number, x t is the input image of the current diffusion layer, ε is the standard normal distribution, ε θ is the noise distribution removed in the back propagation of the diffusion model.
[0093] Furthermore, the Hungarian matching function is used to calculate the similarity loss between the regressed predicted box and the true box, as shown in the following formula:
[0094]
[0095] Where N represents the total number of targets in the prediction set or the true set. Represents the category similarity loss between the predicted box and the true box, y is the predicted set, is the target true value, is the predicted probability, c i For the prediction sample, Represents the similarity loss between the predicted box and the true box of the regression.
[0096] Specifically, in order to eliminate semantic interference to achieve positioning and ensure that each target has a unique embedded feature to match it, a decoder that only focuses on the position information in the feature set is proposed to be used for regression prediction box.
[0097] In the prediction stage, the present invention first needs to perform template prompt preprocessing on the category text retrieval. The template uses "pictures about {}", and the content in brackets can be any open category that needs to be identified. Then the category text is sent to the pre-trained text encoder to extract text features, and a non-parametric similarity calculation is performed between it and the query feature set to find the text feature that is most similar to each feature in the feature set. The corresponding word is the category to be identified. In this way, the knowledge transfer from the visible class to the invisible class is realized, thereby completing the target positioning anchor box R for open set prediction. box and identification type R cls .
[0098] In addition, the loss formula corresponding to target detection is calculated as follows:
[0099]
[0100] Among them, IoU(·,·) represents the value used to calculate the real box b i,j and prediction box As a function of the IoU score between them, i and j jointly determine the location of the target pixel.
[0101] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0102] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, computer software, or a combination of both.
[0103] In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0104] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A remote sensing target open set detection and recognition method based on text retrieval, characterized in that: The following steps are involved: S1, feature extraction step: obtain remote sensing raw data, process the remote sensing raw data based on the perception model, obtain the remote sensing visual features and text features of the image and map them to the visual-text joint embedding space; S2, feature alignment step: the visual features and text features obtained in S1 are aligned in the visual-text joint embedding space through the alignment module to obtain aligned features; S3, feature enhancement step: enhancing the features after alignment processing obtained in S2 to obtain target visual-text multimodal fusion features; S4, feature conversion step: based on the target feature set and the diffusion model, align and fuse to obtain the open set category features; S5, model optimization step: detect and identify the image based on the open set category features obtained in S4, calculate the loss value based on the prediction result, and optimize the target detection network model according to the loss value; Based on the optimized target detection network model, target detection and recognition are performed on remote sensing open set data; S4 includes: S401, taking the cross-domain multimodal fusion features of the remote sensing data obtained in S3 as input, performing block processing to obtain a plurality of block features of the same size, and adding position coding to different block features using a sinusoidal coding rule to obtain block features with position information; S402, randomly initialize the target feature set to convert the target into a learnable query embedding set, and adopt a self-attention transformation model to continuously extract key semantic structure information from the input cross-domain multimodal fusion features, and then use the cross-attention mechanism to drive the initialized query embedding set to capture different target structure information, thereby forming a captured target feature set; S403, using a decoder that only focuses on the position information in the feature set to regress the prediction frame, and during the training process, comparing the prediction set y with the target true value Matching associations, specifically deconstructing it into a bipartite graph matching task, so that one query embedding in the set corresponds to one target, where the number of query embeddings should be much larger than the number of targets; S404, using the diffusion model to transfer the high confidence prior knowledge obtained by pre-training, and integrating the semantic structure information in S402 into the capture target features; S405. Classify the open set targets using a non-parametric metric method, map the features of the query feature set and the target semantic features extracted by the CLIP text encoder to the same feature space by minimizing the cosine distance to achieve alignment, and then use an adaptive fusion network to achieve selective fusion, thereby obtaining the open set category features.
2. The method for open set detection and recognition of remote sensing targets based on text retrieval according to claim 1, characterized in that: The perception model content in S1 includes: The visual encoding network and text encoding network are used to map remote sensing vision and text features into the visual-text joint embedding space respectively, and the mapping process is optimized using the bidirectional sorting loss function.
3. The method for open set detection and recognition of remote sensing targets based on text retrieval according to claim 1, characterized in that: The contents of S2 include: S201, multimodal feature completion, using the generator model to output the completed features of visual and textual modalities, and using adversarial learning strategies to evaluate the completed features; S202, alignment algorithm, in the visual-text joint embedding space, takes the center of the multimodal features of the same concept as the anchor point, shortens the distance between the multimodal features of the same concept object and evenly distributes them around the center anchor point, and makes the multimodal features of different semantic objects stay away from mutual exclusion; S203: Using visual features and text features as positive and negative sample pairs in a comparative learning process, updating parameters as the model learning process progresses, and obtaining aligned features.
4. The method for open set detection and recognition of remote sensing targets based on text retrieval according to claim 1, characterized in that: The feature enhancement steps in S3 include: S301, feature enhancement: for the features after alignment processing obtained in S203, feature enhancement is performed through complementary masks in a multimodal feature enhancement module to obtain an incomplete single-modal feature vector, and then remote sensing visual enhancement features and text enhancement features are obtained through mask feature reconstruction technology; S302, multimodal feature fusion: input remote sensing visual enhancement features and text enhancement features into the cross-modal self-attention Transformer module to obtain multimodal complementary features after cross-domain cross-modal attention weighting; The calculation formula is as follows: Among them, CrossAtt(·) calculates cross-modal attention, m is a single modality, n is the target sample, Q m is the feature vector of the visible part after masking, K n and V n are the other modal eigenvectors, d k Indicates Q m and K n The characteristic dimension of S303: Fusing the multimodal complementary features obtained in S302 with the unimodal features after mask processing, providing guidance for the mask reconstruction process based on the cross-domain multimodal complementary information, enhancing the in-domain features of modal complementarity, and outputting the cross-domain multimodal fusion feature F fuse .
5. The method for open set detection and recognition of remote sensing targets based on text retrieval according to claim 1, characterized in that: The Hungarian matching function used is used to calculate the similarity loss between the regressed prediction box and the true box, as shown in the following formula: Where N represents the total number of targets in the prediction set or the true set. Represents the category similarity loss between the predicted box and the true box, y is the predicted set, is the target true value, is the predicted probability, c i For the prediction sample, Represents the similarity loss between the predicted box and the true box of the regression.
Citation Information
Patent Citations
Remote sensing image-text retrieval method based on guiding visual semantic alignment
CN117009569A
Multi-modal visual target tracking method based on self-distillation symmetric adapter
CN117710414A
Adaptive text summarization method with multi-modal anchor points
CN118035433A