A cross-modal visual positioning method and system based on directional structure enhancement and modal perception affine remapping
Patent Information
- Application Number
- CN202610974191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]为了克服现有跨模态视觉地理定位方法在模态差异大、弱纹理场景缺少强语义地标时,存在定位准确性差的问题,本发明提供一种基于方向结构增强与模态感知仿射重映射的跨模态视觉定位方法及系统
[0025]本发明提供的跨模态视觉定位方法,针对荒漠、农田及低密度道路等弱纹理区域中显著语义地标匮乏的固有困境,首先对初始特征图执行方向信息与空间结构信息的显式增强处理,有效捕捉并放大沙丘脊线、地貌条带分界及稀疏道路延伸等各向异性结构线索,使增强后特征图在弱纹理区域具备连续且判别力强的方向响应,为后续匹配提供可靠的空间几何约束,从而从根源上克服了现有方法因依赖颜色或局部纹理而导致的无显著目标区域匹配失效问题。
Smart Images

Figure CN122820833A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of cross-modal visual retrieval, UAV geolocation, and remote sensing image matching, and particularly to a cross-modal visual positioning method and system based on orientation structure enhancement and modal perception affine remapping. Background Technology
[0002] UAV geolocation is a crucial foundational technology for UAV autonomous navigation, disaster inspection, border patrol, emergency rescue, and low-altitude remote sensing. In scenarios involving satellite signal rejection, interference, or insufficient accuracy, achieving geolocation by matching airborne images with a satellite base map library becomes a key technological approach.
[0003] Thermal infrared imaging exhibits greater environmental adaptability under conditions such as nighttime, low light, smoke, and complex weather, making UAV thermal infrared imagery of significant value in practical positioning applications. However, large-scale baseline maps are typically visible light satellite images, and the imaging mechanisms of thermal infrared and visible light differ greatly: thermal infrared reflects temperature and thermal radiation distribution, while visible light reflects reflection, color, and texture, resulting in significant differences in brightness, texture, and contrast for the same ground feature in the two modes.
[0004] Existing visual geolocation methods typically rely on salient semantic information such as texture consistency, buildings, road intersections, and street layouts between visible light images. When applied to cross-modal retrieval from thermal infrared UAV images to visible light satellite images, these methods are susceptible to modal appearance differences, leading to large cross-modal feature distances at the same location, while areas in different locations but with similar appearances may be incorrectly matched. Furthermore, in weakly textured areas such as deserts, sandy areas, bare land, sparse shrubs, farmland, and low-density roads, scenes often lack strong semantic landmarks such as dense buildings and clear street blocks. Summary of the Invention
[0005] To overcome the problem of poor localization accuracy in existing cross-modal visual geolocation methods when dealing with scenes with large modal differences and weak textures lacking strong semantic landmarks, this invention provides a cross-modal visual geolocation method and system based on directional structure enhancement and modality-aware affine remapping. This invention prioritizes extracting stable cross-modal structures and then performs modal correction in the structure space to improve localization accuracy in scenes with weak textures and large modal differences.
[0006] In a first aspect, the present invention provides a cross-modal visual localization method based on orientation structure enhancement and modality-aware affine remapping, comprising:
[0007] The system acquires thermal infrared images of drones as query images and collects visible light satellite images labeled with geographic location information to form a search database.
[0008] The process involves generating global descriptors for the query image and each visible light satellite image in the retrieval database, specifically including: extracting an initial feature map from the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure-enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image; sensing the modality type of the input image and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; and performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor for the input image.
[0009] Calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the search database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
[0010] Furthermore, a directional structure enhancement module is used to enhance the directional and spatial structure information in the initial feature map to generate a directional structure enhanced feature map of the input image; the directional structure enhancement module includes a semantic preservation branch, a directional response branch, a low-frequency structure branch, and a fusion layer;
[0011] The semantic preservation branch is used to preserve the original context and discrimination information in the initial feature map; the directional response branch is used to extract responses from the initial feature map in different directions and perform adaptive weighted fusion on the extracted responses; the low-frequency structure branch is used to extract low-frequency spatial organization information in the initial feature map; the fusion layer is used to perform adaptive weighted fusion on the features output by the three branches; and the output of the fusion layer is residually connected with the initial feature map to generate a directional structure enhanced feature map.
[0012] Furthermore, the directional response branch uses directional-sensitive depthwise convolution to extract responses from the initial feature map in the horizontal, vertical, and diagonal directions.
[0013] Furthermore, the low-frequency structural branch is used to extract low-frequency spatial organization information from the initial feature map, including: first smoothing and denoising the initial feature map, and then using a convolutional layer with a kernel size of not less than 5×5 to extract low-frequency spatial organization information.
[0014] Furthermore, a shared feature extraction network is used to extract the initial feature map of the input image; the shared feature extraction network is used to map input images of different modalities to the same feature space.
[0015] Furthermore, the modality type of the perceived input image is used to perform channel-level remapping of the fused feature map using affine parameters corresponding to the modality, including:
[0016] The fused feature map is normalized. The modality type of the input image is perceived based on the modality identifier carried by the input image. Learnable scaling and translation parameters are selected according to the modality type to scale and translate the features of each channel of the normalized feature map.
[0017] Secondly, the present invention provides a cross-modal visual positioning system based on orientation structure enhancement and modality-aware affine remapping, comprising:
[0018] The data processing unit is used to acquire UAV thermal infrared images as query images and collect visible light satellite images labeled with geographical location information to form a retrieval database.
[0019] The feature extraction unit is used to generate global descriptors for the query image and each visible light satellite image in the retrieval database, respectively. Specifically, it includes: extracting an initial feature map of the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure-enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image; sensing the modality type of the input image and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; and performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor of the input image.
[0020] The positioning unit is used to calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the retrieval database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
[0021] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in the first aspect.
[0022] Fourthly, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the first aspect.
[0023] Fifthly, the present invention provides a computer program product, comprising a computer program, characterized in that the computer program, when executed by a processor, implements the method described in the first aspect.
[0024] The beneficial effects of this invention are as follows:
[0025] The cross-modal visual localization method provided by this invention addresses the inherent challenge of a lack of significant semantic landmarks in weakly textured regions such as deserts, farmland, and low-density roads. First, it performs explicit enhancement processing on the initial feature map to obtain directional and spatial structural information. This effectively captures and amplifies anisotropic structural cues such as dune ridges, landform strip boundaries, and sparse road extensions. As a result, the enhanced feature map has a continuous and discriminative directional response in weakly textured regions, providing reliable spatial geometric constraints for subsequent matching. This fundamentally overcomes the problem of existing methods failing to match regions without significant targets due to reliance on color or local texture.
[0026] Based on this, the present invention does not simply aggregate the features of the backbone network, but adaptively fuses the feature maps before and after the enhancement of the directional structure, so that the fused features simultaneously contain the original context information and the enhanced directional structure response. Furthermore, it introduces channel-level affine remapping based on modality type awareness to accurately identify the thermal infrared or visible light modality to which the current input belongs. Thus, it only corrects the brightness, texture and contrast shifts caused by differences in imaging mechanisms, effectively avoiding the drawback of traditional methods that confuse modality-specific appearance information with cross-modality common structure. While preserving stable structural response, it directionally suppresses modality-specific noise.
[0027] Crucially, this invention strictly follows a progressive processing order of "first directional structural enhancement, then modality-aware affine remapping"—first ensuring that the cross-modal stable structure dominates in the feature space, and then inputting the directional structural enhancement features into the modality-aware affine remapping module. This allows the optimization of affine parameters to be naturally guided to correct statistical distribution differences rather than erasing structural information. This eliminates the risk of weakening the structural contour due to direct alignment and overcomes the legacy problem of inconsistent cross-modal feature distributions after simple structural enhancement. Ultimately, this ensures that the thermal infrared and visible light modal features fall into a unified discrimination space after affine remapping.
[0028] In summary, this invention significantly improves the robustness and retrieval accuracy of cross-modal visual localization in weakly textured scenes. Attached Figure Description
[0029] Figure 1 A flowchart illustrating a cross-modal visual localization method based on orientation structure enhancement and modality-aware affine remapping provided in an embodiment of the present invention;
[0030] Figure 2 A framework diagram of a cross-modal visual localization method based on orientation structure enhancement and modality-aware affine remapping provided in an embodiment of the present invention;
[0031] Figure 3 This is a schematic diagram of the directional structure enhancement module provided in an embodiment of the present invention;
[0032] Figure 4 This is a schematic diagram of the modality-aware affine remapping module provided in an embodiment of the present invention;
[0033] Figure 5 This is a schematic diagram of the structure of a cross-modal visual positioning system based on orientation structure enhancement and modal perception affine remapping, provided in an embodiment of the present invention.
[0034] Figure 6 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0036] like Figure 1 As shown, this embodiment of the invention provides a cross-modal visual localization method based on orientation structure enhancement and modality-aware affine remapping, including:
[0037] S101: Acquire thermal infrared images of UAVs as query images, and collect visible light satellite images labeled with geographical location information to form a search database;
[0038] S102: Generate global descriptors for the query image and each visible light satellite image in the retrieval database, respectively; specifically including: extracting an initial feature map of the input image; enhancing the directional and spatial structure information in the initial feature map to generate a directional structure enhanced feature map of the input image; fusing the feature maps of the input image before and after directional structure enhancement, perceiving the modal type of the input image, and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modally corrected structural feature map of the input image to reduce the modal statistical differences between thermal infrared and visible light images; performing channel transformation and global aggregation on the modally corrected structural feature map of the input image to generate a global descriptor of the input image;
[0039] Specifically, for weakly textured areas such as deserts, farmland, and low-density roads, the truly stable cross-modal correspondence cues are not color or fine-grained texture, but rather spatial organization and directional structures such as dune flow direction, landform stripes, surface undulation trends, regional outlines, road extension directions, and plot boundaries. Therefore, this step specifically enhances the directional structure of the initial feature map, effectively capturing and amplifying anisotropic directional structural cues such as dune ridges, landform strip boundaries, and sparse road extensions. This results in the enhanced feature map having a clear and continuous directional structural response in weakly textured areas, providing reliable spatial geometric constraints for subsequent feature matching.
[0040] S103: Calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the retrieval database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
[0041] The cross-modal visual localization method provided in this invention specifically enhances the directional and spatial structural information of the initial feature map for weak texture regions such as deserts, farmland, and low-density roads. It effectively captures and amplifies anisotropic directional structural cues such as dune ridges, landform strip boundaries, and sparse road extensions, so that the enhanced feature map has obvious and continuous directional structural responses in weak texture regions. This provides reliable spatial geometric constraints for subsequent feature matching, thereby solving the matching failure problem of existing methods when there are no significant semantic landmarks in weak texture regions.
[0042] Based on this, compared with the feature fusion method that usually simply aggregates the features of the backbone network, the embodiments of the present invention fuse the feature maps before and after the directional structure enhancement and introduce channel-level affine remapping based on modality type awareness. This can accurately identify whether the current input belongs to the thermal infrared or visible light modality. Then, only the brightness, texture and contrast shifts caused by the difference in imaging principle are corrected. This effectively avoids the drawback of "mixing modality-specific information with structural information" in traditional methods. Thus, while retaining the common structural response across modalities, modality-specific noise is suppressed in a targeted manner, which greatly improves the generalization ability of features across modalities.
[0043] Furthermore, this embodiment of the invention first performs directional structure enhancement to ensure that the cross-modal stable structure is dominant in the feature space; then, the features before and after directional structure enhancement are fused, so that the fused features simultaneously contain the original context and the enhanced directional structure response; finally, affine remapping is applied based on the perceived modality type. The above processing sequence ensures that the input to affine correction is no longer the original modality features, but rather the enhanced features that have been fused with the prior directional structure, thereby ensuring that the learning direction of the affine parameters is naturally guided to "correct statistical differences" rather than "erasure structural information". This avoids the weakening of the structural contour caused by direct modality alignment and overcomes the problem that it is difficult to eliminate cross-modal differences after simple structure enhancement, so that the final modality-corrected structural feature map has both discriminative power and robustness in cross-modal visual localization tasks.
[0044] Based on the above embodiments, such as Figure 2 As shown in the framework diagram, this embodiment of the invention also provides a cross-modal visual localization method based on orientation structure enhancement and modality-aware affine remapping, including the following steps:
[0045] S201: Construct cross-modal geolocation samples.
[0046] Specifically, thermal infrared UAV images are acquired as query images, and visible light satellite images of the same area are acquired as search database images. Based on the geographic coordinates attached to the query image, coordinate comparison is performed with the location labels of each satellite image in the search database. Query image-satellite image pairs with the same coordinates or a spatial distance less than a set tolerance value are considered positive samples, while query image-satellite image pairs with different coordinates but a spatial distance greater than or equal to the set tolerance value are considered negative samples. The image data is encapsulated into H5 format files, including query image files for training, validation, and testing phases, as well as search database image files.
[0047] S202: Shared feature extraction.
[0048] Specifically, thermal infrared drone images visible light satellite images Input the shared feature extraction network separately Each of them obtains its corresponding initial feature map. and :
[0049]
[0050]
[0051] In this step, the shared feature extraction network is used. To map thermal infrared / visible modal images to the same feature space, a visual backbone network such as Swing Transformer, convolutional neural network, ConvNeXt, or other networks capable of outputting two-dimensional feature maps can be used.
[0052] S203: Directional structural reinforcement.
[0053] Specifically, the initial feature map F (including Input direction structure enhancement module. For example... Figure 3 As shown, the directional structure enhancement module includes a semantic preservation branch, a directional response branch, a low-frequency structure branch, and a fusion layer. The semantic preservation branch preserves the original context and discriminative information; the directional response branch extracts responses from the initial feature map in different directions and performs adaptive weighted fusion on the extracted responses; the low-frequency structure branch extracts low-frequency spatial organization information from the initial feature map; the fusion layer performs adaptive weighted fusion on the features output from the three branches; and the output of the fusion layer is residually concatenated with the initial feature map to generate a directional structure enhanced feature map.
[0054] In this embodiment, the semantic preservation branch adopts an identity mapping method, without performing additional transformations or downsampling on the initial feature map. It directly uses the high-level feature map output by the shared feature network as the semantic feature output, thereby fully preserving its existing global context, scene category, and local semantic discrimination information, avoiding the loss of semantic information caused by subsequent structural processing. The directional response branch uses multiple direction-sensitive deep convolution operators to extract responses from the initial feature map in different directions. Then, it calculates the weights of each directional response using softmax or other normalization functions, and performs adaptive weighted fusion of different directional responses according to the weights. In this embodiment, the number of directions is 8, used to cover discrete directions such as horizontal, vertical, and diagonal. The low-frequency structure branch first smooths and denoises the initial feature map, and then extracts low-frequency spatial organization information (such as landform stripes, dune flow direction, sparse road trends, and large-scale structures such as plot boundaries) through convolutional layers with a kernel size of not less than 5×5, reducing modality-specific noise and fine texture interference.
[0055] Let the output of the semantically preserved branch be The output of the directional response branch is The output of the low-frequency structure branch is The outputs of the three branches are adaptively weighted and fused to obtain the fused features. It outputs directional structural enhancement features in the form of residuals. ,Right now By using residual output, the directional structure is strengthened while preserving the original semantic discrimination information, which can prevent the loss of the original discriminative semantics due to excessively strong low-frequency structural branches.
[0056] S204: Modality-aware affine remapping.
[0057] Specifically, enhance the directional structure features Input modality-aware affine remapping module. For example... Figure 4 As shown, firstly... Perform GroupNorm or other normalization processing to obtain Then, based on the modal identifier carried by the input image, its corresponding mode is determined, and the affine parameters of the corresponding mode are selected. When the input image is a thermal infrared UAV image, the thermal infrared mode parameters are selected. and When the input image is a visible light satellite image, select the visible light mode parameters. and Based on the selected affine parameters and Modal correction is performed on the normalized feature map to output the modally corrected structural feature map. :
[0058]
[0059] in, This indicates the modal identifier carried by the input image. When the input image is a thermal infrared UAV image... When the input image is a visible light satellite image .
[0060] S205: Global descriptor generation.
[0061] Specifically, the modality-corrected structural feature maps undergo channel transformation and global aggregation to generate global descriptors. One possible implementation is to first map local features to fixed-dimensional local descriptors using 1×1 convolution and batch normalization, and then perform global aggregation using NetVLAD (or generalized average pooling, or attention aggregation) to obtain global descriptors for the thermal infrared query images. Global descriptors of visible light satellite images .
[0062] S206: Perturbation Enhancement and Robust Training, and Triplet Training.
[0063] Specifically, the shared feature extraction network, the orientation structure enhancement module, and the modality-aware affine remapping module are trained.
[0064] Perturbation Enhancement and Robust Training: To improve the robustness of the modality-aware affine remapping module, perturbation enhancement and robust training can be performed on... and Incorporating Gaussian random perturbations allows the model to maintain a stable cross-modal structure under different statistical shifts, preventing overfitting to a fixed statistical distribution. As one possible implementation, the standard deviation of the scaling parameter perturbation can be set to 0.02, and the standard deviation of the translation parameter perturbation can be set to 0.01.
[0065] Triple training: Construct triple samples consisting of a thermal infrared query image, a corresponding positive visible light satellite image from a geographical location, and a corresponding negative visible light satellite image from a geographical location. Through triple-based retrieval constraints, ensure that the global descriptor distance between the query image and the positive sample image is less than the global descriptor distance between the query image and the negative sample image. Optionally, a hard negative sample mining mechanism can be combined to improve retrieval discriminability.
[0066] In this embodiment, perturbation enhancement and robust training are embedded regularization strategies within the triplet training framework. Specifically, in each iteration of forward computation, the affine parameters of the modality-aware affine remapping module are first processed. and Gaussian random noise is superimposed, and the perturbed affine parameters are then applied to feature remapping. Subsequently, a triplet loss function consisting of thermal infrared queries, positive samples corresponding to geographical locations, and negative samples not corresponding to geographical locations is calculated. That is, the Gaussian perturbation forces the modality-aware affine remapping module to still satisfy the triplet distance constraint under the condition that the statistical parameters undergo random shifts. This enables the finally trained global descriptor to maintain inter-class discriminative power while possessing inherent robustness to cross-modal statistical distribution fluctuations.
[0067] S207: Search and locate.
[0068] Specifically, in the application phase, the thermal infrared UAV image to be located is input into the trained model to obtain the query descriptor; the visible light satellite image in the retrieval library is input into the same model or pre-calculated offline to obtain the retrieval library descriptor; the distance or similarity between the query descriptor and each retrieval library descriptor is calculated, and the corresponding visible light satellite image and its geographic coordinates are output as the positioning result according to the similarity from high to low.
[0069] This invention adopts a technical approach of "directional structure enhancement + modality-aware affine remapping + global aggregation retrieval" to address the significant modal differences between thermal infrared UAV images and visible light satellite images, as well as the lack of strong semantic landmarks in weak texture areas. It first enhances cross-modal stable structural features, and then performs modal statistical correction on the enhanced structural features. This significantly reduces the differences between thermal infrared and visible light imaging, while strengthening the terrain structure clues in weak texture scenes, ultimately improving the accuracy and adaptability of cross-modal geolocation.
[0070] Based on the same inventive concept, such as Figure 5As shown, this embodiment of the invention provides a cross-modal visual positioning system based on orientation structure enhancement and modality-aware affine remapping, including a data processing unit, a feature extraction unit, and a positioning unit.
[0071] The data processing unit is used to acquire UAV thermal infrared images as query images and collect visible light satellite images labeled with geographic location information to form a search library;
[0072] The feature extraction unit is used to generate global descriptors for the query image and each visible light satellite image in the retrieval database, respectively. Specifically, it includes: extracting an initial feature map of the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image; sensing the modality type of the input image and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; and performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor of the input image.
[0073] The positioning unit is used to calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the retrieval database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
[0074] It should be noted that the cross-modal visual positioning system based on orientation structure enhancement and modal perception affine remapping provided in this embodiment of the invention is for implementing the above method. Its specific functions can be referred to in the above method embodiments, and will not be repeated here.
[0075] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 can call logic instructions in the memory 603 to execute a cross-modal visual positioning method based on orientation structure enhancement and modality-aware affine remapping. This method includes: acquiring a UAV thermal infrared image as a query image and collecting visible light satellite images labeled with geographic location information to form a retrieval library; generating global descriptors for the query image and each visible light satellite image in the retrieval library, respectively; specifically including: extracting an initial feature map of the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure-enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image, perceiving the modality type of the input image, and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor of the input image; calculating the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the retrieval library, and outputting visible light satellite images with high similarity and their corresponding geographic location information.
[0076] Furthermore, when the logical instructions in the aforementioned memory 603 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0077] This invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute a cross-modal visual localization method based on orientation structure enhancement and modal perception affine remapping provided in the above-described method embodiments.
[0078] This invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements a cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping provided in the above-described method embodiments.
[0079] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping, characterized in that, include: The system acquires thermal infrared images of drones as query images and collects visible light satellite images labeled with geographic location information to form a search database. The process involves generating global descriptors for the query image and each visible light satellite image in the retrieval database, specifically including: extracting an initial feature map from the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure-enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image; sensing the modality type of the input image and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; and performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor for the input image. Calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the search database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
2. The cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping as described in claim 1, characterized in that, A directional structure enhancement module is used to enhance the directional and spatial structure information in the initial feature map to generate a directional structure enhanced feature map of the input image; the directional structure enhancement module includes a semantic preservation branch, a directional response branch, a low-frequency structure branch, and a fusion layer; The semantic preservation branch is used to preserve the original context and discrimination information in the initial feature map; the directional response branch is used to extract responses from the initial feature map in different directions and perform adaptive weighted fusion on the extracted responses; the low-frequency structure branch is used to extract low-frequency spatial organization information in the initial feature map; the fusion layer is used to perform adaptive weighted fusion on the features output by the three branches; and the output of the fusion layer is residually connected with the initial feature map to generate a directional structure enhanced feature map.
3. The cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping according to claim 2, characterized in that, The directional response branch uses directionally sensitive depthwise convolution to extract responses from the initial feature map in the horizontal, vertical, and diagonal directions.
4. The cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping according to claim 2, characterized in that, The low-frequency structural branch is used to extract low-frequency spatial organization information from the initial feature map, including: first, smoothing and denoising the initial feature map, and then using a convolutional layer with a kernel size of not less than 5×5 to extract low-frequency spatial organization information.
5. The cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping according to claim 1, characterized in that, An initial feature map of the input image is extracted using a shared feature extraction network; the shared feature extraction network is used to map input images of different modalities to the same feature space.
6. The cross-modal visual localization method based on orientation structure enhancement and modal-aware affine remapping according to claim 1, characterized in that, The modality type of the perceived input image is used to perform channel-level remapping of the fused feature map using the affine parameters of the corresponding modality, including: The fused feature map is normalized. The modality type of the input image is perceived based on the modality identifier carried by the input image. Learnable scaling and translation parameters are selected according to the modality type to scale and translate the features of each channel of the normalized feature map.
7. A cross-modal visual positioning system based on orientation structure enhancement and modal-aware affine remapping, characterized in that, include: The data processing unit is used to acquire UAV thermal infrared images as query images and collect visible light satellite images labeled with geographical location information to form a retrieval database. The feature extraction unit is used to generate global descriptors for the query image and each visible light satellite image in the retrieval database, respectively. Specifically, it includes: extracting an initial feature map of the input image; enhancing the orientation and spatial structure information in the initial feature map to generate an orientation structure-enhanced feature map of the input image; fusing the feature maps before and after orientation structure enhancement of the input image; sensing the modality type of the input image and using affine parameters of the corresponding modality to perform channel-level remapping on the fused feature map to generate a modality-corrected structural feature map of the input image; and performing channel transformation and global aggregation on the modality-corrected structural feature map of the input image to generate a global descriptor of the input image. The positioning unit is used to calculate the similarity between the global descriptor of the query image and the global descriptors of all visible light satellite images in the retrieval database, and output the visible light satellite images with high similarity and their corresponding geographical location information.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.