Cross-modal remote sensing target detection method and system based on space consistency constraint and deep feature alignment

By adopting the method of aligning with deep features in cross-modal remote sensing object detection, the problem of poor cross-modal capability caused by feature hierarchical differences in the prior art is solved, and high-precision object detection of remote sensing data with domain differences is achieved.

CN120182852APending Publication Date: 2025-06-20SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510277887.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing cross-modal methods do not fully consider feature hierarchical differences, resulting in poor cross-modal capabilities and cannot achieve high-precision object detection for remote sensing data with domain differentials.

Method used

A cross-modal remote sensing object detection method based on spatial consistency constraints and deep features is adopted. Through the improved teacher-student network model, combining the spatial consistency alignment module and the Gaussian distribution consistency alignment module are used to achieve multi-level alignment of shallow and deep features.

Benefits of technology

It significantly improves the detection accuracy and model generalization performance in domain differential scenarios, reduces labeling costs, and realizes high-precision cross-modal remote sensing object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182852A_ABST
    Figure CN120182852A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal remote sensing target detection method and system based on spatial consistency constraint and deep feature alignment, belongs to the field of computer vision and remote sensing science and technology and the technical field of machine learning and deep learning, and solves the problem that a conventional cross-modal method does not fully consider feature hierarchy difference. According to the invention, an improved teacher-student network model is constructed and comprises a student branch network, a teacher branch network and an optimization module for performing pseudo-label optimization, non-monitoring learning and supervised learning on the student branch network and the teacher branch network; training the improved teacher-student network model by adopting the target domain data set and the source domain data set to obtain a trained improved teacher-student network model; and carrying out cross-modal remote sensing target detection on a to-be-detected target domain image by adopting the trained improved teacher-student network model. The method is used for cross-modal remote sensing target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A cross-modal remote sensing object detection method and system based on spatial consistency constraint and deep feature alignment, which is used for cross-modal remote sensing object detection and belongs to the fields of computer vision, remote sensing science and technology, as well as machine learning and deep learning technologies. Background Art

[0002] Remote sensing object detection is the process of automatically identifying and locating ground objects using remote sensing technology and computer vision algorithms. As an important task in computer vision, remote sensing object detection has a wide range of applications. In urban planning and management, it can extract information such as buildings, roads, and green spaces to support land use management and infrastructure construction. Through remote sensing object detection, information such as farmland distribution, vegetation cover, and land use can also be obtained to assist in agriculture and resource management. In addition, it also plays an important role in monitoring changes in ecosystems such as forests, wetlands, and grasslands and geological disaster monitoring. This technology also has great potential in the fields of water resource management, traffic planning, disaster prevention and mitigation, etc.

[0003] Traditional remote sensing object detection usually relies on a large amount of labeled data for training. Especially when detecting objects in different regions or using different sensors, the performance of the model will significantly decline. If object detection needs to be carried out in a new area or using a new type of remote sensing data, usually large-scale data annotation and model training need to be carried out again. Moreover, traditional methods often only perform well under the conditions of the training dataset. Once facing changes in data distribution, the performance of the model will significantly decline.

[0004] Most of the existing methods achieve cross-modal object detection through adversarial learning or simple feature alignment methods. The influence of different-level features on cross-modal objects is not considered. Therefore, the following technical problems exist in the existing technology:

[0005] 1. The existing cross-modal methods do not fully consider the differences in feature hierarchy, resulting in poor cross-modal ability and inability to achieve high-precision object detection for remote sensing data with domain differences;

[0006] 2. Traditional object detection relies on a large amount of labeled data, and the annotation cost is high;

[0007] 3. The problems of unstable training and insufficient alignment ability caused by traditional methods based on gradient reversal layers and adversarial learning in deep feature alignment. Summary of the Invention

[0008] The purpose of the present invention is to provide a cross-modal remote sensing object detection method and system based on spatial consistency constraint and deep feature alignment, so as to solve the problem that the existing cross-modal methods do not fully consider the differences in feature hierarchy, resulting in poor cross-modal ability and inability to achieve high-precision object detection for remote sensing data with domain differences.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A cross-modal remote sensing target detection method based on spatial consistency constraint and deep feature alignment, comprising the following steps:

[0011] Step 1, construct an improved teacher-student network model, including a student branch network, a teacher branch network, and an optimization module for pseudo-label optimization, unsupervised learning, and supervised learning of the student branch network and the teacher branch network. The student branch network includes a student branch backbone network, a spatial consistency alignment module for spatially aligning the shallow features extracted by the student branch backbone network, a Gaussian distribution consistency alignment module for aligning the deep features extracted by the student branch backbone network in terms of Gaussian distribution, and a student branch detector for predicting the results after Gaussian distribution consistency alignment. The teacher branch network includes a teacher branch backbone network and a teacher branch detector connected in sequence;

[0012] Step 2, use the target domain dataset and the source domain dataset to train the improved teacher-student network model to obtain a trained improved teacher-student network model;

[0013] Step 3, use the trained improved teacher-student network model to perform cross-modal remote sensing target detection on the target domain image to be detected.

[0014] Further, the specific implementation steps of the spatial consistency alignment module in Step 1 are as follows:

[0015] Step 1.11, perform feature partitioning and local geometric perception on each target domain shallow feature and each source domain shallow feature of the target domain dataset and the source domain dataset extracted by the student branch backbone network to obtain weighted weights. The specific steps are as follows:

[0016] Divide each target domain shallow feature and each source domain shallow feature of the target domain dataset and the source domain dataset extracted by the student branch backbone network into blocks of the same size for feature partitioning;

[0017] After partitioning, calculate the geometric richness of each partition using the L2 norm and use it as the weighted weight of the partition geometric consistency loss. The formula is:

[0018] w i =||p i ||2

[0019] where w i represents the weight value of the i-th partition after calculation, and ||p i ||2 represents calculating the L2 norm for the i-th partition;

[0020] Step 1.12. Based on the results obtained in Step 1.11, construct an overall function with multiple loss joint constraints to perform hierarchical spatial alignment on the shallow features of each target domain and the shallow features of each source domain. The overall function is the constructed shallow feature space consistency function, and the formula is:

[0021]

[0022] Among them, M represents the number of partitions of the shallow features of the target domain or the source domain, i represents the i-th partition, and cosine_similarity i represents the cosine similarity loss of the i-th partition, MSE i represents the MSE loss of the i-th partition, L1 i represents the L1 loss of the i-th partition, N represents the dimension of the shallow features of the target domain and the source domain, X j represents the j-th shallow feature of the source domain, Y j represents the j-th shallow feature of the target domain, "·" represents multiplication, X represents all the shallow features of the source domain in the i-th partition, and Y represents all the shallow features of the target domain in the i-th partition.

[0023] Furthermore, the specific implementation steps of the Gaussian distribution consistency alignment module in Step 1 are as follows:

[0024] Step 1.21. Use the Box-Cox nonlinear transformation to perform a class Gaussian transformation on each deep feature of the source domain and each deep feature of the target domain in the source domain dataset and the target domain data extracted by the student branch backbone network respectively. The formula is:

[0025]

[0026] Among them, x represents the input data, that is, the deep feature of the source domain or the deep feature of the target domain, λ represents the hyperparameter of the transformation, and y represents the result of the class Gaussian transformation of the deep feature of the source domain or the deep feature of the target domain;

[0027] Step 1.22. After the class Gaussian transformation, use the KL divergence loss to perform Gaussian distribution consistency alignment on the class Gaussian distribution results of each deep feature of the source domain and each deep feature of the target domain. The formula of the KL divergence loss is:

[0028]

[0029] Among them, μ s and σ s represent the mean and variance of each deep feature of the source domain after the transformation respectively, μ t and σ t represent the mean and variance of each deep feature of the target domain after the change respectively, D KL represents the KL divergence loss, Symbol representing the normal distribution.

[0030] Furthermore, the specific implementation steps of the optimization module in step 1 are as follows:

[0031] For the current iteration round, for the teacher branch pseudo-labels obtained by processing each image in the target domain dataset based on the teacher branch, a progressive dynamic threshold method is used to screen the dynamic threshold of the teacher branch pseudo-labels, that is, the current dynamic parameter is calculated through the current iteration number of the teacher branch network and the iteration number at the end of the dynamic selection phase, and then the weight of the current maximum threshold is calculated using the obtained dynamic parameter. Finally, the screening threshold of the current teacher branch pseudo-labels is obtained using the weight and the maximum threshold. The calculation formula is:

[0032] γ = C / L

[0033]

[0034] Among them, C represents the current iteration number of the teacher branch network, L represents the iteration number at the end of the dynamic selection phase. The overall formula of ω is designed to imitate the shape of the Sigmoid function. Among them, α represents the gain, β represents the offset, Max_thershold is the set maximum threshold, which is also the maximum value that thershold can reach. ω and γ represent dynamic parameters, and thershold represents the screening threshold;

[0035] Based on the screening threshold, dynamic threshold screening is performed on the teacher branch pseudo-labels in the current iteration round, and the final teacher branch pseudo-labels are obtained. The formula is:

[0036]

[0037] where, box score represents the class confidence in each teacher branch pseudo-label predicted by the teacher branch network, Pseudo useful represents the retained teacher branch pseudo-labels, Pseudo dropout represents the discarded teacher branch pseudo-labels, Pseudo instance represents the final teacher branch pseudo-labels;

[0038] Based on the final teacher branch pseudo-labels obtained after threshold screening and the prediction results predicted by the student branch network for the target domain, the classification loss is calculated using the loss calculation function during the unsupervised process. The formula is:

[0039]

[0040] Among them, O and V respectively represent the number and the number of categories of all the final teacher branch pseudo-labels obtained in the current iteration round. For each final teacher branch pseudo-label o, use y o,v to represent the indication value that the category in the o-th final teacher branch pseudo-label is v, and p o,v to represent the probability that the category is v in the o-th prediction result obtained by the student branch network for the target domain prediction, represents the category loss of the unsupervised branch;

[0041] When calculating the unsupervised detection box regression loss, there is a case of category mismatch between the final teacher branch pseudo-label and the detection box in the prediction result obtained by the student branch network for the target domain prediction. Soft weighting is used for adjustment. The specific steps are as follows:

[0042] First, for the current iteration round, check whether the category in the prediction result obtained by the student branch network for the target domain prediction is similar to the category of the final teacher branch pseudo-label;

[0043] For the current iteration round, when multiple categories are predicted, judge whether the similarity between the prediction result obtained by the student branch network for the target domain prediction and each category of the final teacher branch pseudo-label is consistent. The judgment formula is:

[0044]

[0045] Among them, similarity represents the category similarity result, S class represents the category predicted in the prediction result obtained by the student branch network for the target domain prediction, and T class represents the category predicted in the teacher branch pseudo-label. True means similar, that is, the categories are consistent, and False means dissimilar, that is, the categories are inconsistent;

[0046] For the current iteration round, when a single category is predicted, since there is only one category, the difference in category confidence is used to judge whether the prediction result obtained by the student branch network for the target domain prediction is consistent with the category of the final teacher branch pseudo-label. The specific formula is as follows:

[0047]

[0048] Among them, and respectively represent the category confidence in the final teacher branch pseudo-label and the prediction result obtained by the student branch network for the target domain prediction, and thershold class is the threshold for judging whether the categories in the final teacher branch pseudo-label and the prediction result obtained by the student branch network for the target domain prediction are similar;

[0049] Subsequently, weights are assigned to each final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction based on the similarity results. That is, when the similarity is True, its weight proportion is large, denoted by w1, and when the similarity is False, its weight proportion is small, denoted by w2. The specific soft weighting formula is as follows:

[0050]

[0051] Among them, represents the regression box loss of the unsupervised branch, L True represents the regression box loss when the similarity is True. The specific loss is calculated by Smooth-L1, L False represents the regression box loss when the similarity is False. The calculation method and process are the same as those of L True ; p h and y h respectively represent the consistent final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction. p k and y k respectively represent the inconsistent final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction. H and K respectively represent the number of correct categories and the number of incorrect categories after comparing the final teacher branch pseudo-label with the prediction results obtained by the student branch network for the target domain prediction;

[0052] After obtaining the class loss of the unsupervised branch and the regression box loss of the unsupervised branch, the overall loss of the unsupervised branch is defined as:

[0053]

[0054] Then, the supervised loss is calculated based on the prediction results of the student branch network for the source domain and the ground truth labels of the images corresponding to the prediction results;

[0055] Based on the unsupervised loss and the supervised loss, during training, the improved teacher-student network model is optimized and trained.

[0056] Furthermore, both the student branch backbone network and the teacher branch backbone network are ResNet101.

[0057] Furthermore, the specific steps for training the improved teacher-student network model in step 2 are as follows:

[0058] Step 2.1: Extract shallow features and deep features from each image input to the target domain dataset and the source domain dataset based on the student branch backbone network. The shallow features include target domain shallow features and source domain shallow features, and the deep features include target domain deep features and source domain deep features;

[0059] Construct the spatial consistency alignment of the shallow features of the target domain and the shallow features of the source domain by using feature partitioning and local geometry perception;

[0060] Perform deep feature alignment on the deep features of the target domain and the deep features of the source domain using Gaussian distribution;

[0061] Input the aligned deep features of the target domain and the deep features of the source domain into the student branch detector, and respectively obtain the prediction results of the target domain and the source domain. The prediction results include detection boxes, categories, and category confidence levels;

[0062] Step 2.2: First, use the teacher branch backbone network to extract deep features from each image input in the target domain dataset, and obtain the teacher branch pseudo-labels through the teacher branch detector. The teacher branch pseudo-labels include detection boxes, categories, and category confidence levels;

[0063] Step 2.3: Use dynamic threshold screening and pseudo-label soft weighting to optimize the teacher branch pseudo-labels and perform unsupervised learning with the prediction results obtained by the student branch network for the target domain. At the same time, perform supervised learning on the prediction results and true labels obtained by the student branch network for the source domain;

[0064] Step 2.4: If the iteration end condition is met, that is, the trained improved teacher-student network model is obtained. Otherwise, adjust the parameters of the improved teacher-student network model and execute Step 2.1 again.

[0065] A cross-modal remote sensing target detection system based on spatial consistency constraint and deep feature alignment, comprising:

[0066] Model construction module: Construct an improved teacher-student network model, including a student branch network, a teacher branch network, and an optimization module for pseudo-label optimization, unsupervised learning, and supervised learning of the student branch network and the teacher branch network. The student branch network includes a student branch backbone network, a spatial consistency alignment module for spatially aligning the shallow features extracted by the student branch backbone network, a Gaussian distribution consistency alignment module for Gaussian distribution alignment of the deep features extracted by the student branch backbone network, and a student branch detector for predicting the results of the Gaussian distribution alignment. The teacher branch network includes a teacher branch backbone network and a teacher branch detector connected in sequence;

[0067] Model training module: Use the target domain dataset and the source domain dataset to train the improved teacher-student network model to obtain the trained improved teacher-student network model;

[0068] Target detection module: Use the trained improved teacher-student network model to perform cross-modal remote sensing target detection on the target domain image to be detected.

[0069] Furthermore, the specific implementation steps of the spatial consistency alignment module in the model construction module are as follows:

[0070] Step 1.11: For the target domain shallow features and source domain shallow features of the target domain dataset and source domain dataset extracted by the student branch backbone network, perform feature partitioning and local geometric perception respectively to obtain weighted weights. The specific steps are as follows:

[0071] Partition the target domain shallow features and source domain shallow features of the target domain dataset and source domain dataset extracted by the student branch backbone network into blocks of the same size for feature partitioning;

[0072] After partitioning, use the L2 norm to calculate the geometric richness of each partition, and use it as the weighted weight of the partition geometric consistency loss. The formula is:

[0073] w i = ||p i ||2

[0074] where w i represents the weight value of the i-th partition after calculation, and ||p i ||2 means calculating the L2 norm for the i-th partition;

[0075] Step 1.12: Based on the results obtained in Step 1.11, construct an overall function with multi-loss joint constraints to perform hierarchical spatial alignment on the target domain shallow features and source domain shallow features. The overall function is the constructed shallow feature spatial consistency function. The formula is:

[0076]

[0077] where M represents the number of partitions of the target domain shallow features or source domain shallow features, i represents the i-th partition, cosine_similarity i represents the cosine similarity loss of the i-th partition, MSE i represents the MSE loss of the i-th partition, L1 i represents the L1 loss of the i-th partition, N represents the dimension of the target domain shallow features and source domain shallow features, X j represents the j -th source domain shallow feature, Y j represents the j -th target domain shallow feature, "·" represents multiplication, X represents all the source domain shallow features in the i-th partition, and Y represents all the target domain shallow features in the i-th partition.

[0078] Furthermore, the specific implementation steps of the Gaussian distribution consistency alignment module in the model construction module are as follows:

[0079] Step 1.21: Use the Box-Cox nonlinear transformation to perform a Gaussian-like transformation on each source domain deep feature and each target domain deep feature of the source domain dataset and the target domain data extracted by the student branch backbone network. The formula is as follows:

[0080]

[0081] where \(x\) represents the input data, that is, the source domain deep feature or the target domain deep feature, \(\lambda\) represents the hyperparameter of the transformation, and \(y\) represents the result of the Gaussian-like transformation of the source domain deep feature or the target domain deep feature.

[0082] Step 1.22: After the Gaussian-like transformation, use the KL divergence loss to align the Gaussian distribution results of each source domain deep feature and each target domain deep feature obtained by the transformation. The formula for the KL divergence loss is as follows:

[0083]

[0084] where \(\mu\) s and \(\sigma\) s represent the mean and variance of each source domain deep feature after the transformation respectively, \(\mu\) t and \(\sigma\) t represent the mean and variance of each target domain deep feature after the transformation respectively, \(D\) KL represents the KL divergence loss, represents the symbol of the normal distribution.

[0085] Furthermore, the specific implementation steps of the optimization module in the model construction module are as follows:

[0086] For the current iteration round, use the progressive dynamic threshold method to perform dynamic threshold screening on the teacher branch pseudo-labels obtained by processing each image in the target domain dataset based on the teacher branch, that is, calculate the current dynamic parameter through the current iteration number of the teacher branch network and the iteration number at the end of the dynamic selection stage, then calculate the weight of the current maximum threshold using the obtained dynamic parameter, and finally obtain the screening threshold of the current teacher branch pseudo-label using the weight and the maximum threshold. The calculation formula is as follows:

[0087] \(\gamma = C / L\)

[0088]

[0089] Among them, C represents the current iteration number of the teacher branch network, L represents the iteration number at the end of the dynamic selection stage. The exponent of the overall formula of ω is designed to imitate the shape of the Sigmoid function. Among them, α represents the gain, β represents the offset, Max_thershold is the set maximum threshold and also the maximum value that the thershold can reach. ω and γ represent dynamic parameters, and thershold represents the screening threshold;

[0090] Based on the screening threshold, dynamic threshold screening is performed on the teacher branch pseudo-labels in the current iteration round, and the final teacher branch pseudo-labels are obtained. The formula is:

[0091]

[0092] Among them, box score represents the class confidence in each teacher branch pseudo-label predicted by the teacher branch network, Pseudo useful represents the retained teacher branch pseudo-labels, Pseudo dropout represents the discarded teacher branch pseudo-labels, Pseudo ins tan ce represents the final teacher branch pseudo-labels;

[0093] Based on the final teacher branch pseudo-labels obtained after threshold screening and the prediction results obtained by the student branch network for the target domain, the classification loss is calculated using the loss calculation function during the unsupervised process. The formula is:

[0094]

[0095] Among them, O and V respectively represent the number and the number of classes of all the final teacher branch pseudo-labels obtained in the current iteration round. For each final teacher branch pseudo-label o, y o,v represents the indicator value that the class in the o-th final teacher branch pseudo-label is v, and p o,v represents the probability that the class is v in the o-th prediction result obtained by the student branch network for the target domain, represents the class loss of the unsupervised branch;

[0096] When calculating the unsupervised detection box regression loss, there is a situation of class mismatch between the final teacher branch pseudo-labels and the detection boxes in the prediction results obtained by the student branch network for the target domain. Soft weighting is used for adjustment. The specific steps are as follows:

[0097] First, for the current iteration round, check whether the class in the prediction result obtained by the student branch network for the target domain is similar to the class of the final teacher branch pseudo-labels;

[0098] For the current iteration round, when multiple categories are predicted, it is necessary to determine whether the similarity of the prediction results obtained by the student branch network for the target domain is consistent with the categories of the final teacher branch pseudo-labels. The judgment formula is as follows:

[0099]

[0100] Among them, similarity represents the category similarity result, S class represents the category predicted in the prediction results obtained by the student branch network for the target domain, T class represents the category predicted in the teacher branch pseudo-labels. True indicates similarity, that is, the categories are consistent, and False indicates dissimilarity, that is, the categories are inconsistent;

[0101] For the current iteration round, when a single category is predicted, since there is only one category, the difference in category confidence is used to determine whether the prediction results obtained by the student branch network for the target domain are consistent with the category of the final teacher branch pseudo-labels. The specific formula is as follows:

[0102]

[0103] Among them, and respectively represent the category confidences in the prediction results obtained by the final teacher branch pseudo-labels and the student branch network for the target domain. The thershold class is the threshold used to determine whether the categories in the prediction results obtained by the final teacher branch pseudo-labels and the student branch network for the target domain are similar;

[0104] Subsequently, based on the similarity results, weights are assigned to each final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain. That is, when the similarity is True, its weight ratio is large, denoted by w1, and when the similarity is False, its weight ratio is small, denoted by w2. The specific soft weighting formula is as follows:

[0105]

[0106] Among them, represents the regression box loss of the unsupervised branch, L True represents the regression box loss when the similarity is True. The specific loss is calculated by Smooth-L1. L False represents the regression box loss when the similarity is False. The calculation method and process are the same as those of L True ; p h and y hrespectively represent the consistent final teacher branch pseudo - label and the prediction result obtained by the student branch network for the target domain prediction, p k and y k respectively represent the inconsistent final teacher branch pseudo - label and the prediction result obtained by the student branch network for the target domain prediction. H and K respectively represent the number of correct classes and the number of wrong classes after comparing the final teacher branch pseudo - label with the prediction result obtained by the student branch network for the target domain prediction;

[0107] After obtaining the class loss of the unsupervised branch and the regression box loss of the unsupervised branch, the overall loss of the unsupervised branch is defined as:

[0108]

[0109] Then, based on the prediction result of the student branch network for the source domain and the ground - truth label of the image corresponding to the prediction result, the supervised loss is calculated;

[0110] Based on the unsupervised loss and the supervised loss, during training, the improved teacher - student network model is optimized and trained.

[0111] Compared with the prior art, the advantages of the present invention are as follows:

[0112] First, through the cross - modal remote - sensing target detection method, the present invention overcomes the problems of traditional target detection that rely on a large amount of labeled data and have weak cross - modal capabilities, and realizes high - precision target detection for remote - sensing data with domain differences. That is, through the hierarchical feature alignment strategy and the pseudo - label dynamic optimization mechanism, high - robustness target detection of multi - modal remote - sensing data is realized, significantly improving the detection accuracy and model generalization performance in the domain - difference scenario, and at the same time reducing the labeling cost;

[0113] Second, the present invention first analyzes the important influence of shallow features and deep features on cross - modal remote - sensing target detection. On this basis, through multi - loss joint constraints, the spatial consistency alignment of shallow features is realized, and combined with the characteristics of the Gaussian distribution, the deep - feature alignment is completed, laying a foundation for realizing high - precision cross - modal target detection. Finally, through pseudo - label optimization, the quality of pseudo - labels is improved, further improving the overall training quality of the network, ensuring the stability and alignment ability of training, enhancing the accuracy of cross - modal target detection. Experiments on the DOTA to DIOR dataset prove that the detection accuracy of the present invention is about 3% higher than that of the existing Adaptive Teacher method in terms of the mAP50 index. Brief Description of the Drawings

[0114] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0115] Figure 1 Schematic diagram of the overall research idea of the present invention;

[0116] Figure 2 Schematic diagram of the spatial consistency constraint in the present invention;

[0117] Figure 3 Schematic diagram of the visualization of the consistency of shallow features in the present invention, where (a) represents no spatial consistency constraint and (b) represents having a spatial consistency constraint;

[0118] Figure 4 Schematic diagram of the structure of the Gaussian-like distribution transformation in the present invention;

[0119] Figure 5 Schematic diagram of feature alignment based on Gaussian distribution in the present invention;

[0120] Figure 6 Schematic diagram of the t-SNE 3D visualization of deep features before and after transformation in the present invention, where (a) represents the t-SNE visualization result of feature points before adding consistency alignment and Gaussian distribution alignment, and (b) represents the visualization result after adding consistency alignment and Gaussian distribution alignment;

[0121] Figure 7 Schematic diagram of the structure of the pseudo-label optimization mechanism in the present invention;

[0122] Figure 8 Schematic diagram of the improvement effect of pseudo-labels in the present invention, where (a) represents the curve of the change in pseudo-label quality without adding pseudo-label optimization, and (b) represents the curve of the change in pseudo-label quality after adding pseudo-label optimization; Detailed implementation manners

[0123] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0124] The present invention studies a cross-modal remote sensing object detection method that takes into account spatial consistency and deep feature alignment.

[0125] First, clarify the important role of the shallow features of remote sensing images in cross-modal remote sensing object detection. Aiming at the characteristics of shallow features, a multi-loss constraint function is combined to achieve the spatial consistency alignment of shallow features, improve the sensitivity of the model to shallow features, and then improve the detection ability of the remote sensing cross-modal object detection model.

[0126] Secondly, aiming at the problems of training instability and insufficient alignment ability caused by traditional methods based on gradient reversal layers and adversarial learning in deep feature alignment, the present invention combines the mathematical characteristics of the Gaussian distribution, and directly realizes the high-precision feature alignment of the network by adjusting the mean and variance attributes of the distribution, avoiding problems such as training instability and insufficient alignment ability brought by using methods such as adversarial learning.

[0127] A cross-modal remote sensing object detection method based on spatial consistency constraint and deep feature alignment, comprising the following steps:

[0128] Step 1, construct an improved teacher-student network model, including a student branch network, a teacher branch network, and an optimization module for pseudo-label optimization, unsupervised learning, and supervised learning of the student branch network and the teacher branch network. The student branch network includes a student branch backbone network, a spatial consistency alignment module for spatially consistent alignment of the shallow features extracted by the student branch backbone network, a Gaussian distribution consistency alignment module for Gaussian distribution consistency alignment of the deep features extracted by the student branch backbone network, and a student branch detector for predicting the results after Gaussian distribution consistency alignment. The teacher branch network includes a teacher branch backbone network and a teacher branch detector connected in sequence; the prediction structures obtained by the student branch network and the teacher branch network both include detection boxes, categories, and category confidence levels.

[0129] Both the student branch backbone network and the teacher branch backbone network are ResNet101.

[0130] There are problems of feature scale difference and scale inconsistency in multi-source data, and different geometric deformations exist in the shallow features of different data. The cosine similarity loss, MSE loss, and L1 loss are combined to effectively constrain the spatial consistency of the model. Therefore, the specific implementation steps of the spatial consistency alignment module are as follows:

[0131] Step 1.11, perform feature partitioning and local geometric perception on each target domain shallow feature and each source domain shallow feature of the target domain dataset and the source domain dataset extracted by the student branch backbone network respectively to obtain weighted weights. The specific steps are as follows:

[0132] The target domain shallow features and source domain shallow features of the target domain dataset and source domain dataset extracted by the student branch backbone network are divided into blocks of the same size for feature partitioning; the advantage of the partitioning operation is that it can achieve geometric consistency alignment within a local area, effectively avoiding the differences caused by large-scale geometric differences in global feature alignment.

[0133] After partitioning, the geometric richness of each partition is calculated using the L2 norm, and it is used as the weighted weight of the partition geometric consistency loss. The formula is:

[0134] w i = ||p i ||2

[0135] where w i represents the weight value of the i-th partition after calculation, ||p i ||2 represents calculating the L2 norm for the i-th partition. Geometric richness can naturally capture and amplify feature intensity;

[0136] Step 1.12: To better achieve the alignment of shallow features between the source domain and the target domain, thereby enhancing the ability of cross-modal object detection, a method of using joint multi-loss is proposed for model training constraints. That is, based on the results obtained in Step 1.11, an overall function of multi-loss joint constraints is constructed to perform hierarchical spatial alignment on the shallow features of each target domain and the shallow features of each source domain. The overall function is the constructed shallow feature space consistency function, and the formula is:

[0137]

[0138] where M represents the number of partitions of the target domain shallow features or source domain shallow features, i represents the i-th partition, cosine_similarity i represents the cosine similarity loss of the i-th partition. The cosine similarity loss uses the cosine value of the angle between the source domain data and the target domain data to measure the consistency of the data. MSE i represents the MSE loss of the i-th partition. The MSE loss measures the similarity of two features by calculating the squared difference between them, and is more suitable for measuring the global differences between features. L1 i represents the L1 loss of the i-th partition, which is less sensitive to outliers and helps to maintain the sparsity of features. Therefore, when dealing with noisy data, the L1 loss can be more robust. N represents the dimension of the target domain shallow features and source domain shallow features, X j represents the j th source domain shallow feature, Y j represents the jOne shallow feature in the target domain, "·" represents multiplication, X represents all source domain shallow features in the i-th partition, and Y represents all target domain shallow features in the i-th partition.

[0139] The spatial consistency alignment module can improve the model's ability to capture diverse information and enhance the model's robustness by combining multiple losses. Moreover, the multi-loss function constraint can also alleviate the deficiencies of a single optimization objective. As a result, it can effectively improve the accuracy of remote sensing cross-modal object detection.

[0140] To address the complex problem of feature distribution between different domains, the present invention uses the method of class Gaussian distribution transformation to solve this situation. The class Gaussian distribution transformation is of great significance in domain adaptation and feature alignment. Its core role is to map the complex distribution to a space close to the Gaussian distribution, thereby reducing the difficulty of distribution alignment and improving the accuracy and robustness of alignment. Therefore, the specific implementation steps of the Gaussian distribution consistency alignment module are as follows:

[0141] Step 1.21: Perform class Gaussian transformation on each source domain deep feature and each target domain deep feature of the source domain dataset and the target domain data extracted by the student branch backbone network using the Box-Cox nonlinear transformation. The formula is:

[0142]

[0143] Among them, x represents the input data, that is, the source domain deep feature or the target domain deep feature, λ represents the transformation hyperparameter, and y represents the result of the class Gaussian transformation of the source domain deep feature or the target domain deep feature;

[0144] Perform relevant class Gaussian transformations on the deep features from the source domain and the target domain. This is because the deep features of the source domain and the target domain usually have different distributions and may differ in terms of mean, standard deviation, skewness, and other statistical characteristics. By applying the class Gaussian transformation, the features of the source domain and the target domain can be transformed into a form close to the Gaussian distribution, thereby simplifying the alignment task.

[0145] Since the feature distribution of the network may change with each iteration, it is necessary to re-estimate λ each time using the maximum likelihood estimation method to ensure the effectiveness of the class Gaussian transformation.

[0146] Step 1.22: After the class Gaussian transformation, perform Gaussian distribution consistency alignment on the class Gaussian distribution results of each source domain deep feature and each target domain deep feature obtained by the transformation using the KL divergence loss. The formula for the KL divergence loss is:

[0147]

[0148] Among them, μ s and σ sRepresent the mean and variance of the deep features in each transformed source domain respectively, μ t and σ t Represent the mean and variance of the deep features in each transformed target domain respectively, D KL Denote the KL divergence loss, Denote the symbol of the normal distribution.

[0149] After applying the Gaussian-like transformation, it is necessary to enforce the Gaussian distribution consistency alignment through the KL divergence loss. KL divergence is a commonly used method to measure the difference between two probability distributions. In the present invention, the deep features of the source domain and the target domain are modeled as Gaussian distributions, and KL divergence is used to quantify the difference between them. If the deep feature distributions of the source domain and the target domain are inconsistent.

[0150] Aiming at the local difference problem in the source domain and the target domain, the KL divergence constraint is adopted to guide the distribution alignment process in a quantitative way by measuring the difference between the deep feature distributions of the source domain and the target domain. It realizes the compactness of intra-class features and the separation of inter-class features by forcing the KL divergence between distributions to be minimized, effectively solving the distribution shift problem. At the same time, KL divergence has the ability to handle high-dimensional features and asymmetric distribution differences, can retain the unique characteristics of the target domain, and avoid information loss. Compared with other alignment strategies, it has a clear optimization goal, high interpretability and robustness, and improves the efficiency of feature alignment and the generalization performance of the model.

[0151] The specific steps of the optimization module are as follows:

[0152] Traditional cross-modal object detection methods usually screen pseudo-labels through a fixed threshold. The advantage of this method is that it is simple and easy to implement, but its disadvantages are also very obvious. In the initial stage of model training, the performance of the teacher model is poor and it cannot generate a large number of high-quality labels. Therefore, if a large pseudo-label screening threshold is adopted, a large number of useful labels may be filtered out, thus prolonging the model training time. On the other hand, if a small pseudo-label screening threshold is set, a large number of low-quality pseudo-label noises may be introduced into the training process, thereby affecting the training effect of the model.

[0153] To solve this problem, this paper proposes to use a progressive dynamic threshold method, so that the threshold for pseudo-label screening can gradually increase from a lower level to a higher level. For the current iteration round, for the teacher branch pseudo-labels obtained by processing the image data in the target domain dataset based on the teacher branch, the progressive dynamic threshold method is used for dynamic threshold screening of the teacher branch pseudo-labels, that is, the current dynamic parameter is calculated through the current iteration number of the teacher branch network and the iteration number at the end of the dynamic selection stage, and then the weight of the current maximum threshold is calculated using the obtained dynamic parameter. Finally, the screening threshold of the current teacher branch pseudo-labels is obtained by using the weight and the maximum threshold, and the calculation formula is:

[0154]

[0155] Among them, C represents the current iteration number of the teacher branch network, L represents the iteration number at the end of the dynamic selection stage. The whole formula of ω is designed to imitate the shape of the Sigmoid function. Among them, α represents the gain, β represents the offset, Max_thershold is the set maximum threshold and also the maximum value that the thershold can reach. ω and γ represent dynamic parameters, and thershold represents the screening threshold;

[0156] Based on the screening threshold, dynamic threshold screening is performed on the teacher branch pseudo-labels in the current iteration round to obtain the final teacher branch pseudo-labels. The formula is:

[0157]

[0158] Among them, box score represents the class confidence in each teacher branch pseudo-label predicted by the teacher branch network, Pseudo useful represents the retained teacher branch pseudo-labels, Pseudo dropout represents the discarded teacher branch pseudo-labels, Pseudo ins tan ce represents the final teacher branch pseudo-labels;

[0159] After threshold screening, the teacher branch pseudo-labels have reduced some noise, but there is still the influence of noise. This is because the threshold only removes the bounding boxes with low confidence, and there may still be misclassification problems in the bounding box classes. Therefore, based on the final teacher branch pseudo-labels obtained after threshold screening and the prediction results obtained by the student branch network for the target domain, the classification loss is calculated using the loss calculation function in the unsupervised process. The formula is:

[0160]

[0161] Among them, O and V respectively represent the number and the number of classes of all the final teacher branch pseudo-labels obtained in the current iteration round. For each final teacher branch pseudo-label o, y o,v represents the indication value that the class of the o-th final teacher branch pseudo-label is v, and p o,v represents the probability that the class is v in the o-th prediction result obtained by the student branch network for the target domain. represents the class loss of the unsupervised branch;

[0162] When calculating the unsupervised detection box regression loss, there is a class mismatch between the finally obtained pseudo-labels of the teacher branch and the prediction results obtained by the student branch network for the target domain (that is, for the current iteration round, the prediction results obtained by the student branch network for the images in the target domain dataset). If the loss between the two detection boxes is simply calculated, it may cause the model to learn incorrect spatial relationships, thereby affecting its detection ability. The present invention uses soft weighting for adjustment. This weighting method allows the model to comprehensively use the results of the student branch network and the teacher branch network as feedback for model training, rather than relying solely on the confidence of the teacher model as feedback. The specific steps are as follows:

[0163] First, for the current iteration round, check whether the classes in the prediction results obtained by the student branch network for the target domain are similar to the classes in the finally obtained pseudo-labels of the teacher branch;

[0164] For the current iteration round, when multiple classes are predicted, determine whether the similarities between the prediction results obtained by the student branch network for the target domain and the classes of the finally obtained pseudo-labels of the teacher branch are consistent. The judgment formula is:

[0165]

[0166] where similarity represents the class similarity result, S class represents the class predicted in the prediction results obtained by the student branch network for the target domain, and T class represents the class predicted in the pseudo-labels of the teacher branch. True indicates similarity, that is, the classes are consistent, and False indicates dissimilarity, that is, the classes are inconsistent;

[0167] For the current iteration round, when a single class is predicted, since there is only one class, the difference in class confidence is used to determine whether the prediction results obtained by the student branch network for the target domain are consistent with the classes of the finally obtained pseudo-labels of the teacher branch. The specific formula is as follows:

[0168]

[0169] where and represent the class confidences in the finally obtained pseudo-labels of the teacher branch and the prediction results obtained by the student branch network for the target domain respectively, and thershold class is the threshold for determining whether the classes in the finally obtained pseudo-labels of the teacher branch and the prediction results obtained by the student branch network for the target domain are similar;

[0170] Subsequently, based on the similarity results, weights are assigned to each final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction. That is, when the similarity is True, its weight ratio is large, denoted by w1, and when the similarity is False, its weight ratio is small, denoted by w2. To prevent the model from relying too much on certain labels during training, the present invention defines the sum of w1 and w2 as 2. For example, w1 is set to 1.2 and w2 is set to 0.8. The purpose of this setting is to make the more correct labels more effective during training while suppressing the noise impact brought by possible incorrect labels. The specific execution process can be defined as follows: The specific soft weighting formula is:

[0171]

[0172] Wherein, represents the regression box loss of the unsupervised branch, L True represents the regression box loss when the similarity is True. The specific loss is calculated by Smooth-L1, L False represents the regression box loss when the similarity is False. The calculation method and process are the same as those of L True ; p h and y h respectively represent the consistent final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction. p k and y k respectively represent the inconsistent final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain prediction. H and K respectively represent the number of correct categories and the number of incorrect categories after comparing the final teacher branch pseudo-label with the prediction results obtained by the student branch network for the target domain prediction;

[0173] After obtaining the class loss of the unsupervised branch and the regression box loss of the unsupervised branch, the overall loss of the unsupervised branch is defined as:

[0174]

[0175] Then, based on the prediction results of the student branch network for the source domain and the true labels of the images corresponding to the corresponding prediction results, the supervised loss is calculated;

[0176] Based on the unsupervised loss and the supervised loss, during training, the improved teacher-student network model is optimized and trained.

[0177] Step 2: Use the target domain dataset and the source domain dataset to train the improved teacher-student network model to obtain the trained improved teacher-student network model;

[0178] The specific steps for training the improved teacher-student network model are as follows:

[0179] Step 2.1: Based on the student branch backbone network, extract shallow features and deep features from each image input in the target domain dataset and the source domain dataset. The shallow features include target domain shallow features and source domain shallow features, and the deep features include target domain deep features and source domain deep features;

[0180] Use feature partitioning and local geometric perception to construct spatial consistency alignment for the target domain shallow features and the source domain shallow features;

[0181] Based on the target domain deep features and the source domain deep features, use Gaussian distribution for deep feature alignment;

[0182] Input the aligned target domain deep features and source domain deep features into the student branch detector, and respectively obtain the prediction results of the target domain and the source domain. The prediction results include detection boxes, categories, and category confidence levels;

[0183] Step 2.2: First, use the teacher branch backbone network to extract deep features from each image input in the target domain dataset, and obtain the teacher branch pseudo-labels through the teacher branch detector. The teacher branch pseudo-labels include detection boxes, categories, and category confidence levels;

[0184] Step 2.3: Use dynamic threshold screening and pseudo-label soft weighting to optimize the teacher branch pseudo-labels and perform unsupervised learning with the prediction results obtained by the student branch network for the target domain, and at the same time perform supervised learning on the prediction results and the true labels obtained by the student branch network for the source domain;

[0185] Step 2.4: If the iteration end condition is met, that is, the trained improved teacher-student network model is obtained. Otherwise, adjust the parameters of the improved teacher-student network model and execute Step 2.1 again.

[0186] Step 3: Use the trained improved teacher-student network model to perform cross-modal remote sensing target detection on the target domain image to be detected. The target domain image to be detected can be buildings, cars, stadiums, ships, etc. in the remote sensing field.

Claims

1. A cross-modal remote sensing target detection method based on spatial consistency constraints and deep feature alignment, characterized in that: The steps include: Step 1, construct an improved teacher-student network model, including a student branch network, a teacher branch network, and an optimization module for pseudo-label optimization, unsupervised learning and supervised learning of the student branch network and the teacher branch network. The student branch network includes a student branch backbone network, a spatial consistency alignment module for spatially aligning shallow features extracted from the student branch backbone network, a Gaussian distribution consistency alignment module for Gaussian distribution consistency alignment of deep features extracted from the student branch backbone network, and a student branch detector for predicting the results after Gaussian distribution consistency alignment. The teacher branch network includes a teacher branch backbone network and a teacher branch detector connected in sequence. Step 2: Use the target domain data set and the source domain data set to train the improved teacher-student network model to obtain a trained improved teacher-student network model; Step 3: Use the trained improved teacher-student network model to perform cross-modal remote sensing target detection on the target domain image to be detected.

2. According to claim 1, a cross-modal remote sensing target detection method based on spatial consistency constraint and deep feature alignment is characterized in that: The specific implementation steps of the spatial consistency alignment module in step 1 are: Step 1.11: For each shallow feature of the target domain and each shallow feature of the source domain extracted by the student branch backbone network, feature partitioning and local geometric perception are performed to obtain weighted weights. The specific steps are as follows: Divide the shallow features of each target domain and each source domain of the target domain dataset and the source domain dataset extracted by the student branch backbone network into blocks of the same size for feature partitioning; After partitioning, the L2 norm is used to calculate the geometric richness of each partition and used as the weighted weight of the partition geometric consistency loss. The formula is: w i =||p i ||2 Among them, w i Represents the calculated weight value of the ith partition, ||p i ||2 means calculating the L2 norm for the i-th partition; Step 1.12: Based on the results obtained in step 1.11, an overall function of multi-loss joint constraints is constructed to hierarchically align the shallow features of each target domain with the shallow features of each source domain. The overall function is the constructed shallow feature space consistency function, and the formula is: Among them, M represents the number of shallow features in the target domain or shallow features in the source domain, i represents the i-th partition, and cosine_similarity i Represents the cosine similarity loss and MSE of the i-th partition i represents the MSE loss of the i-th partition, L1 i represents the L1 loss of the i-th partition, N represents the dimension of the shallow features of the target domain and the shallow features of the source domain, X j represents the jth source domain shallow feature, Y j represents the j-th target domain shallow feature, "·" represents multiplication, X represents all the source domain shallow features in the ith partition, and Y represents all the target domain shallow features in the ith partition.

3. According to claim 2, a cross-modal remote sensing target detection method based on spatial consistency constraints and deep feature alignment is characterized in that: The specific implementation steps of the Gaussian distribution consistency alignment module in step 1 are: Step 1.21: Use Box-Cox nonlinear transformation to perform Gaussian-like transformation on each source domain deep feature and each target domain deep feature of the source domain data set and target domain data extracted by the student branch backbone network. The formula is: Among them, x represents the input data, that is, the deep features of the source domain or the deep features of the target domain, λ represents the hyperparameter of the conversion, and y represents the result of the Gaussian-like transformation of the deep features of the source domain or the deep features of the target domain; Step 1.22: After the Gaussian-like transformation, the Gaussian-like distribution results of each source domain deep feature and each target domain deep feature obtained by the transformation are aligned with the Gaussian distribution using the KL divergence loss. The formula of the KL divergence loss is: Among them, μ s and σ s Represent the mean and variance of each deep feature of the source domain after transformation, μ t and σ t Represent the mean and variance of the deep features of each target domain after the change, D KL represents the KL divergence loss and N represents the sign of the normal distribution.

4. According to claim 3, a cross-modal remote sensing target detection method based on spatial consistency constraint and deep feature alignment is characterized in that: The specific implementation steps of the optimization module in step 1 are: For the current iteration round, based on the teacher branch pseudo-labels obtained by processing each image in the target domain dataset by the teacher branch, the progressive dynamic threshold method is used to perform dynamic threshold screening of the teacher branch pseudo-labels, that is, the current dynamic parameters are calculated by the current number of iterations of the teacher branch network and the number of iterations at the end of the dynamic selection phase, and then the obtained dynamic parameters are used to calculate the weight of the current maximum threshold. Finally, the weight and maximum threshold are used to obtain the screening threshold of the current teacher branch pseudo-label. The calculation formula is: γ=C / L Among them, C represents the current iteration number of the teacher branch network, L represents the iteration number at the end of the dynamic selection phase, and the overall formula of ω is to imitate the shape of the Sigmoid function, where α represents the gain and β represents the offset. Max_thershold is the maximum threshold set, and thershold The maximum value that can be achieved, ω and γ represent dynamic parameters, and thershold represents the screening threshold; Based on the screening threshold, the teacher branch pseudo labels in the current iteration round are dynamically screened to obtain the final teacher branch pseudo labels. The formula is: Among them, box score Pseudo represents the category confidence in each teacher branch pseudo label predicted by the teacher branch network. useful Pseudo labels of teacher branches that are retained, dropout Pseudo labels of teacher branches that are discarded instance represents the final teacher branch pseudo-label; Based on the final teacher branch pseudo-label obtained after threshold screening and the prediction results of the target domain predicted by the student branch network, the classification loss is calculated using the loss calculation function in the unsupervised process. The formula is: Among them, O and V represent the number and number of categories of all final teacher branch pseudo labels obtained in the current iteration round, respectively. For each final teacher branch pseudo label o, y o,v represents the indicator value of the category v in the oth final teacher branch pseudo label, p o,v represents the probability of category v in the oth prediction result obtained by the student branch network for the target domain, represents the category loss of the unsupervised branch; When calculating the unsupervised detection box regression loss, there is a category mismatch between the final teacher branch pseudo label and the detection box in the prediction result obtained by the student branch network for the target domain. Soft weighting is used for adjustment. The specific steps are as follows: First, for the current iteration, check whether the categories in the prediction results of the student branch network for the target domain are similar to the categories of the final teacher branch pseudo labels; For the current iteration round, when multiple categories are predicted, it is judged whether the prediction results of the student branch network for the target domain are consistent with the similarities of the final teacher branch pseudo labels of each category. The judgment formula is: Among them, similarity represents the category similarity result, S class represents the predicted category in the prediction results of the student branch network for the target domain, T class Represents the category predicted in the teacher branch pseudo-label. True means similarity, that is, the category is consistent, and False means dissimilarity, that is, the category is inconsistent. For the current iteration round, when a single category is predicted, since there is only one category, the difference in category confidence is used to determine whether the prediction result of the student branch network for the target domain is consistent with the final teacher branch pseudo-label category. The specific formula is as follows: in, and They represent the category confidence in the prediction results obtained by the final teacher branch pseudo label and the student branch network for the target domain prediction, respectively. class It is the threshold used to determine whether the categories in the final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain are similar; Subsequently, based on the similarity results, weights are assigned to each final teacher branch pseudo label and the prediction results obtained by the student branch network for the target domain prediction. That is, when the similarity is True, its weight accounts for a large proportion, represented by w1, and when the similarity is False, its weight accounts for a small proportion, represented by w2. The specific soft weighting formula is: in, represents the regression box loss of the unsupervised branch, L True Represents the regression box loss when the similarity is True. The specific loss is calculated by Smooth-L1. False Represents the regression box loss when the similarity is False. The calculation method and process are the same as L True The calculation process is the same as p h and h They represent the consistent final teacher branch pseudo-label and the prediction results of the student branch network on the target domain, respectively. k and k They represent the inconsistent final teacher branch pseudo labels and the prediction results obtained by the student branch network for the target domain prediction, respectively. H and K represent the number of correct categories and the number of incorrect categories after comparing the final teacher branch pseudo labels with the prediction results obtained by the student branch network for the target domain prediction, respectively. After obtaining the category loss of the unsupervised branch and the regression box loss of the unsupervised branch, the overall loss of the unsupervised branch is defined as: Then, based on the student branch network, the supervision loss is calculated for the prediction results of the source domain and the true labels of the images corresponding to the corresponding prediction results; Based on unsupervised loss and supervised loss, the improved teacher-student network model is optimized and trained during training.

5. The cross-modal remote sensing target detection method based on spatial consistency constraint and deep feature alignment according to claim 1, characterized in that: The student branch backbone network and the teacher branch backbone network are both ResNet101.

6. The cross-modal remote sensing target detection method based on spatial consistency constraint and deep feature alignment according to claim 5, characterized in that: The specific steps of training the improved teacher-student network model in step 2 are: Step 2.1, based on the student branch backbone network, shallow features and deep features are extracted for each image input into the target domain dataset and the source domain dataset, where the shallow features include shallow features of the target domain and shallow features of the source domain, and the deep features include deep features of the target domain and deep features of the source domain; The shallow features of the target domain and the shallow features of the source domain are aligned and constructed with spatial consistency by using feature partitioning and local geometric perception. Gaussian distribution is used to align deep features based on the target domain deep features and the source domain deep features; The aligned deep features of the target domain and the deep features of the source domain are input into the student branch detector to predict the prediction results of the target domain and the source domain respectively. The prediction results include the detection box, category, and category confidence. Step 2.2: First, use the teacher branch backbone network to perform deep feature extraction on each image input into the target domain dataset, and obtain the teacher branch pseudo label through the teacher branch detector. The teacher branch pseudo label includes the detection box, category, and category confidence. Step 2.3, use dynamic threshold screening and pseudo-label soft weighting to optimize the teacher branch pseudo-label and perform unsupervised learning with the prediction results of the target domain predicted by the student branch network, and supervised learning of the prediction results and true labels of the source domain predicted by the student branch network; Step 2.4: If the iteration end condition is met, the trained improved teacher-student network model is obtained; otherwise, adjust the parameters of the improved teacher-student network model and execute step 2.1 again.

7. A cross-modal remote sensing target detection system based on spatial consistency constraints and deep feature alignment, characterized in that: include: Model construction module: construct an improved teacher-student network model, including a student branch network, a teacher branch network, and an optimization module for pseudo-label optimization, unsupervised learning, and supervised learning of the student branch network and the teacher branch network. The student branch network includes a student branch backbone network, a spatial consistency alignment module for spatial consistency alignment of shallow features extracted from the student branch backbone network, a Gaussian distribution consistency alignment module for Gaussian distribution consistency alignment of deep features extracted from the student branch backbone network, and a student branch detector for predicting the results after Gaussian distribution consistency alignment. The teacher branch network includes a teacher branch backbone network and a teacher branch detector connected in sequence. Model training module: Use the target domain data set and the source domain data set to train the improved teacher-student network model to obtain a trained improved teacher-student network model; Target detection module: The trained improved teacher-student network model is used to perform cross-modal remote sensing target detection on the target domain image to be detected.

8. The cross-modal remote sensing target detection system based on spatial consistency constraint and deep feature alignment according to claim 7, characterized in that: The specific implementation steps of the spatial consistency alignment module in the model building module are: Step 1.11: For each shallow feature of the target domain and each shallow feature of the source domain extracted by the student branch backbone network, feature partitioning and local geometric perception are performed to obtain weighted weights. The specific steps are as follows: Divide the shallow features of each target domain and each source domain of the target domain dataset and the source domain dataset extracted by the student branch backbone network into blocks of the same size for feature partitioning; After partitioning, the L2 norm is used to calculate the geometric richness of each partition and used as the weighted weight of the partition geometric consistency loss. The formula is: w i =||p i ||2 Among them, w i Represents the calculated weight value of the ith partition, ||p i ||2 means calculating the L2 norm for the i-th partition; Step 1.12: Based on the results obtained in step 1.11, an overall function of multi-loss joint constraints is constructed to hierarchically align the shallow features of each target domain with the shallow features of each source domain. The overall function is the constructed shallow feature space consistency function, and the formula is: Among them, M represents the number of shallow features in the target domain or shallow features in the source domain, i represents the i-th partition, and cosine_similarity i Represents the cosine similarity loss and MSE of the i-th partition i represents the MSE loss of the i-th partition, L1 i represents the L1 loss of the i-th partition, N represents the dimension of the shallow features of the target domain and the shallow features of the source domain, X j represents the jth source domain shallow feature, Y j represents the j-th target domain shallow feature, "·" represents multiplication, X represents all the source domain shallow features in the ith partition, and Y represents all the target domain shallow features in the ith partition.

9. The cross-modal remote sensing target detection system based on spatial consistency constraint and deep feature alignment according to claim 7, characterized in that: The specific implementation steps of the Gaussian distribution consistency alignment module in the model building module are: Step 1.21: Use Box-Cox nonlinear transformation to perform Gaussian-like transformation on each source domain deep feature and each target domain deep feature of the source domain data set and target domain data extracted by the student branch backbone network. The formula is: Among them, x represents the input data, that is, the deep features of the source domain or the deep features of the target domain, λ represents the hyperparameter of the conversion, and y represents the result of the Gaussian-like transformation of the deep features of the source domain or the deep features of the target domain; Step 1.22: After the Gaussian-like transformation, the Gaussian-like distribution results of each source domain deep feature and each target domain deep feature obtained by the transformation are aligned with the Gaussian distribution using the KL divergence loss. The formula of the KL divergence loss is: Among them, μ s and σ s Represent the mean and variance of each deep feature of the source domain after transformation, μ t and σ t Represent the mean and variance of the deep features of each target domain after the change, D KL represents the KL divergence loss, The symbol for the normal distribution.

10. The cross-modal remote sensing target detection system based on spatial consistency constraint and deep feature alignment according to claim 7, characterized in that: The specific implementation steps of the optimization module in the model building module are: For the current iteration round, based on the teacher branch pseudo-labels obtained by processing each image in the target domain dataset by the teacher branch, the progressive dynamic threshold method is used to perform dynamic threshold screening of the teacher branch pseudo-labels, that is, the current dynamic parameters are calculated by the current number of iterations of the teacher branch network and the number of iterations at the end of the dynamic selection phase, and then the obtained dynamic parameters are used to calculate the weight of the current maximum threshold. Finally, the weight and maximum threshold are used to obtain the screening threshold of the current teacher branch pseudo-label. The calculation formula is: γ=C / L Among them, C represents the current iteration number of the teacher branch network, L represents the iteration number at the end of the dynamic selection phase, and the overall formula of ω is to imitate the shape of the Sigmoid function, where α represents the gain and β represents the offset. Max_thershold is the maximum threshold set, and thershold The maximum value that can be achieved, ω and γ represent dynamic parameters, and thershold represents the screening threshold; Based on the screening threshold, the teacher branch pseudo labels in the current iteration round are dynamically screened to obtain the final teacher branch pseudo labels. The formula is: Among them, box score Pseudo represents the category confidence in each teacher branch pseudo label predicted by the teacher branch network. useful Pseudo labels of teacher branches that are retained, dropout Pseudo labels of teacher branches that are discarded instance represents the final teacher branch pseudo-label; Based on the final teacher branch pseudo-label obtained after threshold screening and the prediction results of the target domain predicted by the student branch network, the classification loss is calculated using the loss calculation function in the unsupervised process. The formula is: Among them, O and V represent the number and number of categories of all final teacher branch pseudo labels obtained in the current iteration round, respectively. For each final teacher branch pseudo label o, y o,v represents the indicator value of the category v in the oth final teacher branch pseudo label, p o,v represents the probability of category v in the oth prediction result obtained by the student branch network for the target domain, represents the category loss of the unsupervised branch; When calculating the unsupervised detection box regression loss, there is a category mismatch between the final teacher branch pseudo label and the detection box in the prediction result obtained by the student branch network for the target domain. Soft weighting is used for adjustment. The specific steps are as follows: First, for the current iteration, check whether the categories in the prediction results of the student branch network for the target domain are similar to the categories of the final teacher branch pseudo labels; For the current iteration round, when multiple categories are predicted, it is judged whether the prediction results of the student branch network for the target domain are consistent with the similarities of the final teacher branch pseudo labels of each category. The judgment formula is: Among them, similarity represents the category similarity result, S class represents the predicted category in the prediction results of the student branch network for the target domain, T class Represents the category predicted in the teacher branch pseudo-label. True means similarity, that is, the category is consistent, and False means dissimilarity, that is, the category is inconsistent. For the current iteration round, when a single category is predicted, since there is only one category, the difference in category confidence is used to determine whether the prediction result of the student branch network for the target domain is consistent with the final teacher branch pseudo-label category. The specific formula is as follows: in, and They represent the category confidence in the prediction results obtained by the final teacher branch pseudo label and the student branch network for the target domain prediction, respectively. class It is the threshold used to determine whether the categories in the final teacher branch pseudo-label and the prediction results obtained by the student branch network for the target domain are similar; Subsequently, based on the similarity results, weights are assigned to each final teacher branch pseudo label and the prediction results obtained by the student branch network for the target domain prediction. That is, when the similarity is True, its weight accounts for a large proportion, represented by w1, and when the similarity is False, its weight accounts for a small proportion, represented by w2. The specific soft weighting formula is: in, represents the regression box loss of the unsupervised branch, L True Represents the regression box loss when the similarity is True. The specific loss is calculated by Smooth-L1. False Represents the regression box loss when the similarity is Fal se. The calculation method and process are the same as L True The calculation process is the same as p h and h They represent the consistent final teacher branch pseudo-label and the prediction results of the student branch network on the target domain, respectively. k and k They represent the inconsistent final teacher branch pseudo labels and the prediction results obtained by the student branch network for the target domain prediction, respectively. H and K represent the number of correct categories and the number of incorrect categories after comparing the final teacher branch pseudo labels with the prediction results obtained by the student branch network for the target domain prediction, respectively. After obtaining the category loss of the unsupervised branch and the regression box loss of the unsupervised branch, the overall loss of the unsupervised branch is defined as: Then, based on the student branch network, the supervision loss is calculated for the prediction results of the source domain and the true labels of the images corresponding to the corresponding prediction results; Based on unsupervised loss and supervised loss, the improved teacher-student network model is optimized and trained during training.

Citation Information

Cited By

  • Passenger abnormal behavior recognition method and device in elevator monitoring night vision mode

    CN120452068A

  • Method and device for training polypropylene film defect detection model and polypropylene film defect detection method

    CN120783149A

  • Weak supervision target detection method guided by cross-modal pseudo tag

    CN120953596A

  • A weakly supervised object detection method guided by cross-modal pseudo labels

    CN120953596B

  • Bridge area water area ship target detection method, equipment and medium

    CN121214301A