Building roof extraction method based on optical and SAR remote sensing images
By designing adaptive feature alignment and cross-modal multi-scale feature fusion modules in building roof extraction method, the problem of ignoring the low-level feature and semantic gap when fusion of optical and SAR remote sensing images is solved, and a higher precision building roof extraction is achieved.
Patent Information
- Application Number
- CN202411849538.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-17
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-16
AI Technical Summary
When the existing multimodal fusion semantic segmentation algorithms fuse optical and SAR remote sensing images, low-level features and supplementary information are usually ignored. Due to different imaging mechanisms, there is a semantic gap between the modals, and direct feature fusion cannot fully utilize the advantages of information complementarity.
Adaptive feature alignment module and cross-modal multi-scale feature fusion module are designed. By adaptive alignment and multi-scale feature fusion of high-level features of optical and SAR images, the channel self-attention mechanism is used to fuse discriminative features and obtain the fused features.
Effectively integrate the features of optical and SAR images, make full use of complementary information, improve the accuracy of building roof extraction, and solve the problem of semantic gap between modals.
Smart Images

Figure CN120014257A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of remote sensing image semantic segmentation, and in particular relates to a building roof extraction method based on optical and SAR remote sensing images. Background Art
[0002] The construction of distributed rooftop photovoltaic systems plays an important role in the transition to renewable energy power generation. The extraction of building roof area is the first problem to be solved in the evaluation of distributed rooftop photovoltaic potential, which directly affects the accuracy of rooftop photovoltaic resource assessment.
[0003] High-resolution remote sensing images mainly include optical images and synthetic aperture radar (SAR) images. Optical images provide high spatial resolution, rich spectral and texture information, but are easily affected by weather. SAR can work in all weather conditions and penetrate certain ground objects. Therefore, optical images and SAR images are complementary to ground information. Compared with single-modal image segmentation, multi-modal segmentation that fuses the two can improve the accuracy of segmentation.
[0004] In recent years, the research on multimodal fusion methods has become a hot topic for many scholars. Some multimodal fusion semantic segmentation algorithms have been gradually proposed and applied to the extraction of building roofs. However, they still face challenges when fusing different modal data. First, existing methods usually focus on high-level features while ignoring low-level features and other supplementary information. Secondly, due to different imaging mechanisms, there is a semantic gap between different modal data. Direct feature fusion will not be able to fully utilize the complementary information advantages between different modal data. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a building roof extraction method based on optical and SAR remote sensing images, designs an adaptive feature alignment module and a cross-modal multi-scale feature fusion module, which are used for the alignment of heterogeneous modal data and the feature fusion of multimodal data, respectively, and makes full use of the complementary information provided by optical and SAR images.
[0006] The technical scheme for achieving the purpose of the present invention is as follows:
[0007] A building roof extraction method based on optical and SAR remote sensing images, comprising:
[0008] Obtain low-level features of optical and SAR images and high-level features
[0009] According to the high-level features Identify relevant information between two modalities from the spatial dimension and align modal features;
[0010] High-level features of the aligned optical and SAR images obtained Obtain multi-scale feature representations through dilated spatial convolutional pooling pyramids
[0011] and Each of them is a group, and the channel self-attention mechanism is used to adaptively fuse the discriminant features by applying weights to each modality and selectively discarding irrelevant components to obtain the fused features.
[0012] Will After 4 times upsampling, After joint splicing, they pass through a convolution layer and then undergo 4x upsampling to form a building roof extraction network.
[0013] According to the described method, the adaptive feature alignment process is as follows:
[0014] The high-level features of the acquired optical and SAR images are converted into the RKHS space;
[0015] Feature alignment is performed by minimizing the distance between the transformed feature distributions of the two. The calculation formula is as follows:
[0016]
[0017] Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is a unified representation for feature conversion to RKHS;
[0018] The semantic distribution alignment loss is added to the loss function, and its calculation formula is as follows:
[0019] L=L dice +L sda
[0020] L sda =βMMD 2 (F opt ,F sar )
[0021] Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, as the loss function for class imbalanced semantic segmentation, L sda It is the semantic distribution alignment loss, which is obtained by minimizing the square of the MMD distance between the feature distributions of optical and SAR images multiplied by the coefficient β.
[0022] According to the described method, the cross-modal multi-scale feature fusion process is as follows:
[0023] Superposition F opt and F sar becomes
[0024] The entire spatial feature on a Channel is encoded into a global feature through feature compression operation;
[0025] Perform adaptive calibration on global features, learn the relationship between channels, and obtain the weights of different channels;
[0026] Perform weight calibration operation and multiply the weight E of different channels by the original feature map;
[0027] Will Perform convolution operation to finally obtain the fused feature F opt-sar .
[0028] According to the described method, the calculation formula of the feature compression operation is as follows:
[0029]
[0030] Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position;
[0031] The calculation formula for the adaptive calibration operation of global features is as follows:
[0032] E=F ex (S, W)=σ(g(S, W))=σ(W2δ(W1S))
[0033] Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function;
[0034] The calculation formula for weight calibration operation is as follows:
[0035]
[0036] Among them: F sacle represents the weight calibration operation, is the feature after superposition, and E is the weight after learning;
[0037] A building roof extraction device based on optical and SAR remote sensing images, comprising:
[0038] Feature acquisition module, used to obtain low-level features of optical and SAR images and high-level features
[0039] Feature alignment module, used for the high-level features Identify relevant information between two modalities from the spatial dimension and align modal features;
[0040] Multi-scale feature acquisition module for obtaining high-level features of optical and SAR images Obtain multi-scale feature representations through dilated spatial convolutional pooling pyramids
[0041] The feature fusion module is used to obtain high and low fusion features respectively. and Each of them is a group, and the channel self-attention mechanism is used to adaptively fuse the discriminant features by applying weights to each modality and selectively discarding irrelevant components to obtain the fused features.
[0042] Upsampling module, used to After 4 times upsampling, After joint splicing, they pass through a convolution layer and then undergo 4x upsampling to form a building roof extraction network.
[0043] According to the device, the adaptive feature alignment process is as follows:
[0044] The high-level features of the acquired optical and SAR images are converted into the RKHS space;
[0045] Feature alignment is performed by minimizing the distance between the transformed feature distributions of the two. The calculation formula is as follows:
[0046]
[0047] Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is a unified representation for feature conversion to RKHS;
[0048] The semantic distribution alignment loss is added to the loss function, and its calculation formula is as follows:
[0049] L=L dice +L sda
[0050] L sda =βMMD2 (F opt ,F sar )
[0051] Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, as the loss function for class imbalanced semantic segmentation, L sda It is the semantic distribution alignment loss, which is obtained by minimizing the square of the MMD distance between the feature distributions of optical and SAR images multiplied by the coefficient β.
[0052] According to the device, the cross-modal multi-scale feature fusion process is as follows:
[0053] Superposition F opt and F sar becomes
[0054] The entire spatial feature on a Channel is encoded into a global feature through feature compression operation;
[0055] Perform adaptive calibration on global features, learn the relationship between channels, and obtain the weights of different channels;
[0056] Perform weight calibration operation and multiply the weight E of different channels by the original feature map;
[0057] Will Perform convolution operation to finally obtain the fused feature F opt-sar .
[0058] According to the device, the calculation formula of the feature compression operation is as follows:
[0059]
[0060] Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position;
[0061] The calculation formula for the adaptive calibration operation of global features is as follows:
[0062] E=F ex (S, W)=σ(g(S, W))=σ(W2δ(W1S))
[0063] Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function;
[0064] The calculation formula for weight calibration operation is as follows:
[0065]
[0066] Among them: F sacle represents the weight calibration operation, is the feature after superposition, and E is the weight after learning.
[0067] Based on the above technical solution, the present invention has the following beneficial effects:
[0068] The present invention takes into account the semantic gap problem between different modal data. Through adaptive feature alignment and cross-modal multi-scale feature fusion, the features of optical and SAR images can be effectively fused, and the complementary information provided by optical and SAR images can be fully utilized to improve the extraction accuracy of building roofs. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 This is a structural diagram of the network model of the present invention.
[0070] Figure 2 This is a structural diagram of the cross-modal multi-scale feature fusion of the present invention. DETAILED DESCRIPTION
[0071] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. The specific implementation methods described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0072] Figure 1 The network model structure diagram of the present invention is shown, which adopts the classic encoder-decoder structure. It includes the following steps: (1) In the encoder part, the optical and SAR images are input into two independent feature extractors respectively to obtain the low-level features and high-level features of the optical and SAR images respectively:
[0073] (1.1) Preprocess the optical image and SAR image, perform spatial registration, and crop the image size to 512X512. Input the preprocessed images into a convolutional network with ResNet-50 as the backbone;
[0074] (1.2) Obtaining low-level features of optical and SAR images at the second layer Its size is 128X128, and high-level features are obtained in the fourth layer Its size is 32X32.
[0075] (2) Due to different imaging mechanisms, different modal images have different representations of the same object. These differences cause the model to ignore the fact that the two input images have the same semantic meaning, making it difficult to capture hidden semantic representations. The purpose of the adaptive feature alignment module is to better learn the features of a single modality before fusion, convert the high-level features of optical and SAR images into the same semantic space, and perform feature alignment by minimizing the distance between the feature distributions of the two. By connecting an adaptive feature alignment module, the relevant information between the two modalities is identified from the spatial dimension and the modal features are aligned:
[0076] (2.1) The high-level features of the acquired optical and SAR images are converted into the RKHS space;
[0077] (2.2) The maximum average difference is used to estimate the distance between the distributions, and feature alignment is performed by minimizing the distance between the two feature distributions. The calculation formula is as follows:
[0078]
[0079] Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is the unified representation after feature transformation;
[0080] (2.3) Add semantic distribution alignment loss to the loss function, and its calculation formula is as follows:
[0081] L=L dice +L sda
[0082] L sda =βMMD 2 (F opt ,F sar )
[0083] Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, which is often used as the loss function for category imbalanced semantic segmentation. sda is the semantic distribution alignment loss, which is obtained by multiplying the square of the previously calculated MMD by the coefficient β.
[0084] (3) The high-level features of the optical and SAR images of size 32X32 obtained in (1.2) at the fourth layer Multi-scale feature representations of the two images are obtained through dilated spatial convolutional pooling pyramids. The atrous spatial convolution pooling pyramid samples the given input in parallel with atrous convolutions of different sampling rates, which is equivalent to capturing the context of the feature image at multiple scales.
[0085] (4) The size of 128X128 obtained in (1.2) on the second layer and (3) obtained Each is a group, such as Figure 2 As shown in the figure, the cross-modal multi-scale feature fusion module is input respectively. This module is based on the channel attention mechanism and adaptively assigns fusion weights to each modal feature according to the contribution of the multi-modal feature. The weight coefficient is learned based on the statistical information extracted from the global feature and is used to adjust the feature response of the channel and reduce interference, thereby enhancing the representation ability of the fused feature, adaptively fusing the discriminant features, and obtaining the fused features. Here, the optical characteristics are uniformly represented as F opt , SAR feature is expressed as F sar , the processing process is as follows:
[0086] (4.1) Superposition F opt and F sar becomes
[0087] (4.2) The entire spatial feature on a Channel is encoded into a global feature through feature compression operation. The calculation formula is as follows:
[0088]
[0089] Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position;
[0090] (4.3) Perform adaptive calibration on the global features, as shown in Formula 2, learn the relationship between each channel, and obtain the weights of different channels. The calculation formula is as follows:
[0091] E=(F ex (S, W)=σ(g(S, W))=σ(W2δ(W1S))
[0092] Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function;
[0093] (4.4) Perform weight calibration operation and multiply the weight E of different channels by the original feature map. The calculation formula is as follows:
[0094]
[0095] Among them: F sacle represents the weight calibration operation, is the feature after superposition in step (1), E is the weight after learning in step (3); (4.5) Perform convolution operation to finally obtain the fused feature F opt-sar , where the low-level features are The high-level features are
[0096] (5) The high-level fusion features of size 32X32 obtained in (4.5) After 4 times upsampling, the size becomes 128X128, which is fused with the low-level features of size 128X128 obtained in (4.5). After joint splicing, a convolution layer is performed, the output size is 128X128, and then a semantic segmentation result of size 512 is obtained after 4 times upsampling.
[0097] A building roof extraction device based on optical and SAR remote sensing images, comprising:
[0098] Feature acquisition module, used to obtain low-level features of optical and SAR images and high-level features
[0099] Feature alignment module, used for the high-level features Identify relevant information between two modalities from the spatial dimension and align modal features;
[0100] Multi-scale feature acquisition module for obtaining high-level features of optical and SAR images Obtain multi-scale feature representations through dilated spatial convolutional pooling pyramids
[0101] The feature fusion module is used to obtain high and low fusion features respectively. and Each of them is a group, and the channel self-attention mechanism is used to adaptively fuse the discriminant features by applying weights to each modality and selectively discarding irrelevant components to obtain the fused features.
[0102] Upsampling module, used to After 4 times upsampling, After joint splicing, they pass through a convolution layer and then undergo 4x upsampling to form a building roof extraction network.
[0103] The device and the adaptive feature alignment process are as follows:
[0104] The high-level features of the acquired optical and SAR images are converted into the RKHS space;
[0105] Feature alignment is performed by minimizing the distance between the transformed feature distributions of the two. The calculation formula is as follows:
[0106]
[0107] Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is a unified representation for feature conversion to RKHS;
[0108] The semantic distribution alignment loss is added to the loss function, and its calculation formula is as follows:
[0109] L=L dice +L sda
[0110] L sda =βMMD 2 (F opt ,F sar )
[0111] Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, as the loss function for class imbalanced semantic segmentation, L sda It is the semantic distribution alignment loss, which is obtained by minimizing the square of the MMD distance between the feature distributions of optical and SAR images multiplied by the coefficient β.
[0112] According to the device, the cross-modal multi-scale feature fusion process is as follows:
[0113] Superposition F opt and F sar becomes
[0114] The entire spatial feature on a Channel is encoded into a global feature through feature compression operation;
[0115] Perform adaptive calibration on global features, learn the relationship between channels, and obtain the weights of different channels;
[0116] Perform weight calibration operation and multiply the weight E of different channels by the original feature map;
[0117] Will Perform convolution operation to finally obtain the fused feature F opt-sar .
[0118] For the device, the calculation formula of the feature compression operation is as follows:
[0119]
[0120] Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position;
[0121] The calculation formula for the adaptive calibration operation of global features is as follows:
[0122] E=F ex (S, W)=σ(g(S, W))=σ(W2δ(W1S))
[0123] Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function;
[0124] The calculation formula for weight calibration operation is as follows:
[0125]
[0126] Among them: F sacle represents the weight calibration operation, is the feature after superposition, and E is the weight after learning.
[0127] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0128] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0129] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0131] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0132] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A building roof extraction method based on optical and SAR remote sensing images, characterized in that: include: Obtain low-level features of optical and SAR images and high-level features According to the high-level features Identify relevant information between two modalities from the spatial dimension and align modal features; High-level features of the aligned optical and SAR images obtained Obtain multi-scale feature representations through dilated spatial convolutional pooling pyramids and Each of them is a group, and the channel self-attention mechanism is used to adaptively fuse the discriminant features by applying weights to each modality and selectively discarding irrelevant components to obtain the fused features. Will After 4 times upsampling, After joint splicing, they pass through a convolution layer and then undergo 4x upsampling to form a building roof extraction network.
2. The method according to claim 1, characterized in that The adaptive feature alignment process is as follows: The high-level features of the acquired optical and SAR images are converted into the RKHS space; Feature alignment is performed by minimizing the distance between the transformed feature distributions of the two. The calculation formula is as follows: Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is a unified representation for feature conversion to RKHS; The semantic distribution alignment loss is added to the loss function, and its calculation formula is as follows: L=L dice +L sda L sda =βMMD 2 (F opt ,F sar ) Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, as the loss function for class imbalanced semantic segmentation, L sda It is the semantic distribution alignment loss, which is obtained by minimizing the square of the MMD distance between the feature distributions of optical and SAR images multiplied by the coefficient β.
3. The method according to claim 1, characterized in that The cross-modal multi-scale feature fusion process is as follows: Superposition F opt and F sar becomes The entire spatial feature on a Channel is encoded into a global feature through feature compression operation; Perform adaptive calibration on global features, learn the relationship between channels, and obtain the weights of different channels; Perform weight calibration operation and multiply the weight E of different channels by the original feature map; Will Perform convolution operation to finally obtain the fused feature F opt-sar .
4. The method according to claim 3, characterized in that The calculation formula for the feature compression operation is as follows: Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position; The calculation formula for the adaptive calibration operation of global features is as follows: E=F ex (S, W)=σ(g(S, W))=σ(W2δ(W1S)) Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function; The calculation formula for weight calibration operation is as follows: Among them: F sacle represents the weight calibration operation, is the feature after superposition, and E is the weight after learning.
5. A building roof extraction device based on optical and SAR remote sensing images, characterized in that: include: Feature acquisition module, used to obtain low-level features of optical and SAR images and high-level features Feature alignment module, used for the high-level features Identify relevant information between two modalities from the spatial dimension and align modal features; Multi-scale feature acquisition module for obtaining high-level features of optical and SAR images Obtain multi-scale feature representations through dilated spatial convolutional pooling pyramids The feature fusion module is used to obtain high and low fusion features respectively. and Each of them is a group, and the channel self-attention mechanism is used to adaptively fuse the discriminant features by applying weights to each modality and selectively discarding irrelevant components to obtain the fused features. Upsampling module, used to After 4 times upsampling, After joint splicing, they pass through a convolution layer and then undergo 4x upsampling to form a building roof extraction network.
6. The device according to claim 5, characterized in that The adaptive feature alignment process is as follows: The high-level features of the acquired optical and SAR images are converted into the RKHS space; Feature alignment is performed by minimizing the distance between the transformed feature distributions of the two. The calculation formula is as follows: Among them: F opt is the optical advanced feature, F sar It is the advanced feature of SAR. It is a unified representation for feature conversion to RKHS; The semantic distribution alignment loss is added to the loss function, and its calculation formula is as follows: L=L dice +L sda L sda =βMMD 2 (F opt ,F sar ) Among them: L is the loss function which consists of two parts, namely L dice and L sda , L dice is Dice Loss, as the loss function for class imbalanced semantic segmentation, L sda It is the semantic distribution alignment loss, which is obtained by minimizing the square of the MMD distance between the feature distributions of optical and SAR images multiplied by the coefficient β.
7. The device according to claim 5, characterized in that The cross-modal multi-scale feature fusion process is as follows: Superposition F opt and F sar becomes The entire spatial feature on a Channel is encoded into a global feature through feature compression operation; Perform adaptive calibration on global features, learn the relationship between channels, and obtain the weights of different channels; Perform weight calibration operation and multiply the weight E of different channels by the original feature map; Will Perform convolution operation to finally obtain the fused feature F opt-sar .
8. The device according to claim 7, characterized in that The calculation formula for the feature compression operation is as follows: Among them: F sq represents the feature compression operation, H represents the height of the input image, W represents the width of the input image, and i, j represents the current pixel position; The calculation formula for the adaptive calibration operation of global features is as follows: E=F ex (S,W)=σ(g(S,W))=σ(W2δ(W1S)) Among them: F ex represents the adaptive calibration operation, They represent the weight matrices of the two fully connected layers, δ is the ReLU function, and σ is the Sigmoid function; The calculation formula for weight calibration operation is as follows: Among them: F sacle represents the weight calibration operation, is the feature after superposition, and E is the weight after learning.
Citation Information
Patent Citations
Building extraction method in remote sensing image, electronic equipment and storage medium
CN115345866A
Optical and SAR image registration method, device and equipment based on position awareness
CN116883466A
SAR and optical image matching method
CN117036754A
Land coverage classification method, system and equipment fusing visible light and SAR (Synthetic Aperture Radar) image
CN117115649A
Multi-modal image segmentation method based on cross attention
CN117911426A
Cited By
Cross-modal sea surface target identification method and system for high-dynamic unmanned platform
CN121544965A
Method and system for automatically extracting height of ancillary facility of building
CN121810758A