High-precision radio environment map reconstruction method based on multi-modal fusion

By fusing RSRP, synthetic aperture radar images, and building information through CMCTNet, the modal fusion problem in radio environment map reconstruction is solved, and high-precision and robust radio environment map reconstruction is achieved, which is suitable for complex propagation environments under sparse measurement conditions.

CN120669192AActive Publication Date: 2025-09-19PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV

Patent Information

Application Number
CN202510763557.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Traditional radio environment map reconstruction methods lack accuracy in complex propagation environments, especially under sparse measurement conditions, where it is difficult to effectively fuse synthetic aperture radar images and building information, resulting in a decline in model performance.

Method used

A cross-modal complete tensor network (CMCTNet) is adopted to achieve high-precision radio environment map reconstruction by combining RSRP, synthetic aperture radar imagery and building information data through a multimodal fusion attention module and a skip connection decoder.

Benefits of technology

The reconstruction accuracy and stability of the radio environment map are significantly improved under sparse measurement conditions, especially the ability to maintain high-resolution signal recovery in complex environments, which is better than traditional interpolation methods and deep learning baseline methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669192A_ABST
    Figure CN120669192A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of radio communication, and particularly discloses a multi-modal fusion-based high-precision radio environment map reconstruction method, which comprises the following steps of: S01, respectively processing a plurality of pieces of input information by adopting a plurality of parallel encoders to generate a plurality of corresponding potential feature representations; and S02, performing cross-modal dynamic fusion on the plurality of potential features by adopting a fusion attention module to obtain fused potential feature representation, which comprises the following steps: constructing spatial adaptive attention maps, simultaneously modeling space and channel dependency by adopting a dual attention mechanism, grouping the spatial adaptive attention maps, and obtaining the fused potential feature representation. The sub-branches are processed through space or channel attention, and finally shuffling and connection are carried out to promote cross-channel interaction; and S03, adopting a decoder to reconstruct the fused potential feature representation, including reserving high-resolution spatial features through jump connection, and transferring position information from the encoder to the decoder to obtain a high-precision radio environment map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radio communication technology, and in particular to a high-precision radio environment map reconstruction method based on multimodal fusion. Background Art

[0002] As wireless communication systems evolve towards full commercial use of 5G and pre-research for 6G, the high-precision construction of radio environment maps (REMs) has become the core foundation for achieving dynamic management of spectrum resources, intelligent interference suppression, and autonomous network optimization. In practical application scenarios such as drone-assisted networks, vehicular communications, and dynamic cognitive radio systems, the availability of reliable REMs determines the effectiveness of key operations such as spectrum switching, load balancing, and adaptive deployment. However, due to the high cost, time, and infrastructure requirements associated with intensive field measurements, constructing REMs with fine spatial resolution is difficult, especially in large-scale or dynamically changing environments. Traditional REM reconstruction methods mainly rely on interpolation algorithms, which are computationally efficient, but their effectiveness decreases rapidly in heterogeneous environments characterized by complex propagation phenomena such as shadowing, diffraction, and multipath scattering.

[0003] Deep learning techniques have become a powerful tool for REM estimation. These methods often incorporate auxiliary information such as building layout, terrain elevation, or transmitter locations to improve generalization and robustness.

[0004] Synthetic aperture radar (SAR) can capture high-resolution surface and structural information regardless of weather or lighting conditions, making it ideal for mapping consistent environments. However, SAR imagery is rarely incorporated into data-driven REM construction due to the following technical issues: RF measurements and SAR imagery differ significantly in spatial characteristics, scale, and statistical properties, making feature alignment and fusion difficult. Furthermore, noise, speckle, suspension, and shadow effects in SAR imagery can introduce irrelevant or misleading features that, if not handled properly, can degrade model performance. Summary of the Invention

[0005] To address the above issues, the present invention aims to provide a high-precision radio environment map reconstruction method based on multimodal fusion. A Cross-Modal Completion TensorNetwork (CMCTNet) is proposed to reconstruct high-fidelity radio environment maps (REMs) under sparse measurement conditions by fusing multimodal information from RSRP, synthetic aperture radar imagery, and building information data. It integrates three dedicated encoders for each modality, a fusion attention module (FAM) that adaptively captures cross-modal dependencies, and a skip-connection decoding pipeline that enables fine-grained signal recovery in complex environments.

[0006] To achieve the purpose of the present invention, the technical solution of the present invention is: a high-precision radio environment map reconstruction method based on multimodal fusion, comprising:

[0007] Step S01: using multiple parallel encoders to process multiple input information respectively to generate corresponding multiple potential feature representations;

[0008] Step S02: using a fusion attention module to dynamically fuse the multiple latent features across modalities to obtain a fused latent feature representation, including: constructing a spatially adaptive attention map, and using a dual attention mechanism to simultaneously model spatial and channel dependencies, grouping the spatially adaptive attention map, processing it through spatial or channel attention sub-branches, and finally shuffling and concatenating to promote cross-channel interaction;

[0009] Step S03: reconstructing the fused latent feature representation using a decoder, including retaining spatial features with high resolution through skip connections and transferring location information from the encoder to the decoder to obtain a high-precision radio environment map.

[0010] Preferably, in step S01, the multiple parallel encoders include three parallel encoders. Using three parallel encoders enables modality-specific feature extraction and cross-modality synergy enhancement: through independent RSRP, synthetic aperture radar image, and building information data encoding paths, the model can separately preserve the propagation characteristics of wireless signals, ground object scattering characteristics, and building geometric constraints, avoiding feature confusion caused by shared encoders.

[0011] Preferably, in step S01, the plurality of input information includes a sparse RSRP map, a SAR image and building data.

[0012] Preferably, in step S01, the processing of multiple input information separately includes: performing initial feature extraction on the multiple input information using discrete wavelet transform, decomposing the input information into multiple frequency bands, performing multi-level feature extraction processing on the multiple frequency bands through separate convolution layers, and generating potential feature representations.

[0013] Preferably, the multiple frequency bands include LL, LH, HL and HH. These four frequency bands represent different frequency information of the input features: the LL channel captures the low-frequency structural features of the image, mainly including overall shape information such as contours and background; the LH channel responds to high-frequency changes in the horizontal direction and can highlight horizontal edge features; the HL channel reflects the edges or textures in the vertical direction; and the HH channel focuses on retaining high-frequency details in the diagonal direction of the image, such as sharp corners, oblique boundaries, etc. This frequency division mechanism enables the model to distinguish information of different frequencies at a shallow stage, thereby improving its ability to understand structure and edges. In this way, CMCTNet can fully explore the spatial details and structural clues hidden in the input in the early stages of the network, laying a more robust foundation for subsequent multimodal feature extraction and fusion, especially in sparse signal recovery tasks, which can enhance the model's ability to retain key texture and edge information.

[0014] Preferably, in step S02, during the cross-modal dynamic fusion, the latent feature representation extracted from the sparse RSRP map is used as the query Q, and the fused latent feature representation extracted from the SAR image and the building data is used as the key K' and the value V.

[0015] Preferably, in step S02, the spatial adaptive attention map is constructed based on the following attention calculation formula:

[0016]

[0017] Where Q is the RSRP query tensor, K' and V are the key and value of modality fusion, d k is the dimension of the key vector.

[0018] Preferably, in step S03, the decoder outputs a dense prediction map for estimating the complete RSRP distribution in the discrete area;

[0019] The dense prediction graph is: in, The complete RSRP distribution is estimated by a multimodal deep learning framework, N x and N y They represent the number of grid cells along the x and y directions after the target area is discretized.

[0020] Preferably, each encoder comprises several dense blocks. Through a layer-by-layer densely connected structure, each convolutional layer can directly access the feature maps of all previous layers, achieving cascaded reuse of the original signal propagation pattern and high-level abstract features, thereby significantly improving the network's ability to model the spatial variation of RSRP, even in synthetic aperture radar images and building information data.

[0021] Preferably, the forward propagation process within the dense block includes: first performing the first convolution on the input to obtain the feature x1, then splicing x1 with the original input in the channel dimension as the input of the second convolution block to obtain x2; then splicing x1, the original input and x2 in the channel dimension as the input of the third convolution block to obtain the final output x3.

[0022] Beneficial effects of the above technical solution:

[0023] The high-precision radio environment map reconstruction method based on multimodal fusion proposed in this paper introduces synthetic aperture radar imagery as a new modality into data-driven radio environment map construction, providing a robust structural prior to improve propagation modeling in challenging environments. Through a unified multimodal deep learning framework, RSRP samples, SAR imagery, and building data are effectively integrated through a multi-attention mechanism to learn obstacle perception signal patterns. Using a new benchmark dataset, the synthesized RSRP map is matched with co-registered synthetic aperture radar imagery and geographic information system-derived building information maps, enabling repeatable research and comparative evaluation. Simulation experiments verify that the performance of CMCTNet is significantly better than traditional interpolation methods and the latest DL-based baseline methods, especially under extremely sparse sampling and non-uniform measurement conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 A flowchart of a method for reconstructing a high-precision radio environment map based on multimodal fusion provided by one embodiment of the present invention;

[0026] Figure 2 A schematic diagram of an encoder-fusion-decoder architecture according to an embodiment of the present invention;

[0027] Figure 3 A schematic diagram of the internal structure of a dense block provided by one embodiment of the present invention;

[0028] Figure 4 A schematic diagram of the architecture of a mechanism provided by an embodiment of the present invention;

[0029] Figure 5 A visual comparison of REM reconstructed at different sampling levels provided for one embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following further describes the implementation methods of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than an exhaustive list of all embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other unless there is a conflict.

[0031] The terms "first," "second," and the like (if any) in the specification and claims are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated (if any) or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0032] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0033] Example 1

[0034] One embodiment of the present invention provides a high-precision radio environment map reconstruction method based on multimodal fusion, for reconstructing high-fidelity radio environment maps from sparsely sampled RSRP measurements. This method effectively combines structural priors extracted from SAR imagery and building information maps to improve signal recovery under extreme data sparsity conditions. It also learns a spatially consistent cross-modal representation while accounting for urban propagation conditions, material-induced attenuation, and structural occlusion.

[0035] A high-precision radio environment map reconstruction method based on multimodal fusion Figure 1As shown, it includes: step S01: using multiple parallel encoders to process multiple input information separately to generate corresponding multiple potential feature representations; step S02: using a fusion attention module to dynamically fuse the multiple potential features across modalities to obtain a fused potential feature representation, including: constructing a spatial adaptive attention map, and using a dual attention mechanism to simultaneously model spatial and channel dependencies, grouping the spatial adaptive attention map, processing it through spatial or channel attention sub-branches, and finally shuffling and connecting to promote cross-channel interaction; step S03: using a decoder to reconstruct the fused potential feature representation, including retaining spatial features with high resolution through jump connections, and transferring position information from the encoder to the decoder to obtain a high-precision radio environment map.

[0036] CMCTNet is a multimodal deep learning network for high-fidelity Radio Environment Map (REM) reconstruction tasks. Figure 2 As shown in the figure, the overall "encoder-fusion-decoder" framework is employed to recover a complete electromagnetic signal distribution map from sparse RSRP measurements. The network first receives three input modalities: a sparse RSRP map, a SAR remote sensing image, and a building information map. Three independently designed modal encoders are used for feature extraction. Each encoder front-end is embedded with wavelet blocks to achieve multi-scale frequency domain modeling, enhancing the recognition of texture edges, structural contours, and spatial details. Subsequently, deep feature representations are extracted from each modal feature using dense blocks. After encoding, the three modal features are fed into a fusion attention module. This module utilizes a self-attention mechanism with RSRP features as queries and SAR and building information as keys. It dynamically guides the cross-modal feature fusion through channel-spatial joint attention to model occlusion-aware propagation characteristics. The fused unified features are then fed into the decoder for progressive upsampling to the original spatial resolution. Meanwhile, skip connections are used to incorporate shallow features from the encoder to preserve critical signal boundaries and local variations. Finally, the network outputs a fully reconstructed RSRP distribution map. CMCTNet utilizes the collaborative enhancement of multimodal features to demonstrate superior completion and structure preservation capabilities in low sampling rates and complex urban scenes, providing an efficient and robust solution for radio environment perception under sparse sensing conditions.

[0037] In step S01, the multiple parallel encoders include three parallel encoders. The multiple input information includes sparse RSRP maps, SAR images, and building data. Processing the multiple input information separately includes: performing initial feature extraction on the multiple input information using discrete wavelet transforms (DWTs), decomposing the input information into multiple frequency bands, and performing multi-level feature extraction on each of the multiple frequency bands through separate convolutional layers to generate latent feature representations. In this embodiment, to enhance the network's ability to model local signal discontinuities and structural patterns, a wavelet-based depthwise separable convolution module is constructed. This module, as one of the core structures of each modal encoder, is embedded in the network front-end to improve feature extraction. Unlike traditional convolution methods that operate only in the spatial domain, this module integrates multiple layers of discrete wavelet transforms (DWTs) and inverse wavelet transforms (IWTs) to achieve frequency-domain enhanced feature representation. First, the input feature map is decomposed by DWT to generate four subbands: LL, LH, HL, and HH, corresponding to different directional and frequency components, respectively. These four subbands are then processed using a dedicated channel-grouped convolution, and feature amplitudes are further adjusted using a learnable scaling module to enhance the representation of high-frequency components such as building edges, signal obstruction, and multipath reflections. During the multi-layer wavelet decomposition process, the low-level components extracted at each layer are retained and passed to the next layer for further decomposition. The processed high-frequency components (LH, HL, and HH) are reconstructed along with the current low-level components during the decoding phase and restored to the previous resolution scale using the IWT. This encoding-decoding wavelet processing framework enables the module to aggregate multi-scale feature information while maintaining consistent spatial resolution. The module also features a parallel spatial domain convolution branch to extract shallow local features and perform an element-by-element weighted fusion with the reconstructed wavelet features, thereby combining the advantages of both the frequency and spatial domains. Furthermore, the module supports optional spatial downsampling and maintains channel independence through grouped convolution to avoid information aliasing.

[0038] In order to enhance the expressive power and information flow efficiency of the feature extraction module, a dense block is designed and introduced in the encoder. This module draws on the dense connection idea of ​​DenseNet and uses a cascade feature fusion mechanism to effectively alleviate the problems of gradient disappearance and feature redundancy in deep networks. The internal structure of the dense block is as follows Figure 3As shown, in the forward propagation, the first convolution is performed on the input to obtain the feature x1, and then x1 is spliced ​​with the original input in the channel dimension and used as the input of the second convolution block to obtain x2; then x1, the original input and x2 are spliced ​​in the channel dimension and used as the input of the third convolution block to obtain the final output x3. In this embodiment, the dense block contains three convolution units connected in series. The input of each convolution operation includes not only the output of the previous layer but also the original input features, thereby realizing cross-layer feature fusion and reuse. First, the input tensor is processed by the first convolution unit to generate the first-stage feature; this feature is spliced ​​with the input tensor and sent to the second convolution unit to output further enhanced features; then, the second-stage output is spliced ​​again with the first two features and sent to the third convolution unit to finally obtain the fused output feature. Through this multiple splicing and convolution operation, the dense block can effectively capture local spatial details and cross-scale structural features, and is particularly suitable for modeling non-stationary electromagnetic propagation patterns in complex urban environments. To construct the basic convolutional unit used in this module, a lightweight convolutional encapsulation module, Conv2D, was designed. It consists of a 3×3 standard convolutional layer, a batch normalization layer, and a ReLU activation function in series, ensuring stable training and enhanced nonlinear expression capabilities. This structure maintains spatial consistency of features while also achieving good parameter efficiency and computational stability.

[0039] After encoding, the extracted modality-specific features are aligned and fused through the Fusion Attention Module (FAM), which is the core of CMCTNet's cross-modal reasoning capability. This module uses a scaled dot-product attention mechanism to dynamically model the interaction between RSRP signal features (as queries) and environmental priors extracted from SAR and building data (as keys and values). In order to inject building geometry information into the attention process, building features are element-wise added to the SAR key tensor to form K'=K SAR +X B The calculation formula for attention output is:

[0040]

[0041] Where Q is the RSRP query tensor, K' and V are the key and value of modality fusion, d k is the dimension of the key vector. This formulation enables the model to construct spatially adaptive attention maps that account for geometric obstacles, material-dependent scattering, and other propagation-related interactions.

[0042] In order to further improve the representation ability of the model in the process of multimodal information fusion, a dual attention mechanism module is introduced such as Figure 4As shown in the figure, the channel attention (Channel Attention) and spatial attention (Spatial Attention) are combined, supplemented by the channel shuffle operation, which effectively enhances the diversity and discriminability of feature expression. In this embodiment, the dual attention mechanism module first reorganizes the input features according to the specified number of groups G, so that each channel subset can calculate the channel attention and spatial attention separately. For the channel attention path, the module first performs global average pooling on the input features to extract channel-level statistical features, then adjusts the channel response strength through learnable scaling and bias parameters, and finally obtains the channel weight through the Sigmoid activation function to enhance the key signal dimension. The spatial attention path uses GroupNorm to normalize the spatial dimension, retaining position-sensitive features, and then highlights the regional response with spatial distinction through the same weight bias adjustment and activation operation. The outputs of the two paths are spliced ​​and fused in the channel dimension, and finally the original order is broken up by the channel shuffle operation, promoting information interaction and feature mixing between different channel groups, thereby effectively alleviating the redundancy problem between attention channels. The module has a lightweight overall structure and learnable parameters, which can significantly enhance the network's modeling capabilities for complex structures, edge transitions, and inter-modal coupling features while maintaining low computational overhead.

[0043] The decoder component of CMCTNet reconstructs the full RSRP distribution from the fused latent representation. Its architecture is similar to the RSRP encoder, but with the addition of skip connections to preserve high-resolution spatial features throughout the network. These connections directly transfer positional information from the encoder to the decoder, which is particularly important when reconstructing fine-grained signal variations that may be lost in deeper layers. The decoder output is a dense prediction map: Used to estimate the complete RSRP distribution within a discrete area. The complete RSRP distribution is estimated by a multimodal deep learning framework, N x and N y They represent the number of grid cells along the x and y directions after the target area is discretized.

[0044] The REM reconstruction problem can be formulated as a tensor completion task, where the goal is to reconstruct the complete discretized RSRP distribution from sparse measurements and associated SAR images. Consider a two-dimensional region of interest A with length X and width Y. The region is discretized into a uniform rectangular grid G ​​with size G∈ The length of each grid cell is Δx and the width is Δy, satisfying the following relationship:

[0045]

[0046] Where (N x ) and (N y ) represent the number of grid cells along the x-direction and y-direction, respectively. Figure 1 The spatial coordinates associated with each grid cell ((i,j)) are defined as:

[0047] s(i,j)=(i·Δ x ,j·Δ y )

[0048] where i∈{1,2,…,N x},j∈{1,2,…,N y In this spatially discretized model, the signal power at each grid location is represented by a function (P(s,f)) that maps a coordinate-frequency pair ((s,f)) to a corresponding RSRP measurement:

[0049] P(s(i,j),f)=RSRP at s(i,j)for frequency f

[0050] Define the fully observed but unknown RSRP distribution as a tensor:

[0051] Y∈R Nx×Ny ,

[0052] Among them, each element (Y (i,j) ) represents the RSRP value at grid cell ((i, j)). Due to practical limitations, only a limited number of RSRP measurements are available, resulting in an incomplete representation of the signal distribution.

[0053] To improve the accuracy of RSRP reconstruction and mitigate the impact of missing data, additional environmental information from remote sensing data, especially SAR imagery and architectural maps, is incorporated. SAR imagery provides information about structure and surface roughness that affect radio signal propagation. The SAR-derived feature tensor is defined as: SAR ∈R Nx×Ny .

[0054] Similarly, the building information map that captures the impact of urban structure on RF propagation is represented as: B ∈R Nx×Ny .

[0055] Given sparse RSRP observations Y OBS and auxiliary data X SAR and X B , estimating the complete RSRP distribution via a multimodal deep learning framework

[0056]

[0057] where f(·) represents the proposed multimodal learning framework. This problem setting enables the multimodal learning framework to exploit the spatial correlation between RF signal measurements and environmental priors. In particular, SAR imagery and building information maps provide complementary cues about propagation obstacles, surface discontinuities, and structural layout, which significantly affect signal attenuation and coverage. To train and evaluate the proposed multimodal REM reconstruction framework, a dataset is constructed that integrates simulated RSRP measurements, co-registered SAR imagery, and building information maps.

[0058] RSRP data is derived from the BARTLab radiation pattern dataset, where urban radio propagation is simulated using the Altair Feko electromagnetic modeling tool. Each simulation assumes a building height of 10 meters and includes three 720842A2 antennas mounted at a height of 30 meters, each transmitting at 46.00 dBm. Although the dataset supports five frequency bands from 1750 MHz to 5750 MHz, this example restricts the study to the 1750 MHz band to ensure spectral consistency across all samples.

[0059] The original dataset contains approximately 2,000 simulated RSRP maps with spatial resolutions ranging from (140 × 268) to (561 × 1,313) pixels. These maps are constructed on a structured grid aligned with a terrain layout extracted from OpenStreetMap. To incorporate environmental context, Sentinel-1 SAR imagery was retrieved from the Google Earth Engine (GEE) platform. Only urban areas with reliable SAR coverage were retained, resulting in a carefully curated multimodal subset of 1,600 samples.

[0060] Prior to training, all input modalities undergo a series of preprocessing steps to ensure spatial alignment and resolution consistency. First, RSRP maps are center-cropped and uniformly resized to (256×256) pixels using bilinear interpolation. This normalization facilitates batch and architecture consistency. Second, SAR images are temporally averaged across multiple acquisitions to mitigate speckle noise while preserving stable surface scattering patterns. Third, seasonal alignment is enforced by selecting SAR acquisitions from the same time period as the RSRP simulation, thereby reducing differences due to environmental changes. Finally, the building information map is rasterized and resampled to match the spatial resolution of the RSRP and SAR data.

[0061] After preprocessing, each data sample is represented as a spatially aligned triplet (Y OBS ,X SAR ,X B ), where Y OBSrepresents sparse RSRP measurement, X SAR is the SAR backscatter intensity map, X B This dataset design enables the proposed CMCTNet to learn the joint spatial relationship between RF propagation and environmental structure, thereby achieving robust reconstruction under conditions of high spatial heterogeneity and extreme measurement sparsity.

[0062] In summary, CMCTNet integrates wavelet-based feature extraction, densely connected spatial modeling, and adaptive multimodal attention to achieve robust and high-resolution REM reconstruction. Its architectural components are specifically tailored to the challenges of cross-modal heterogeneity, signal sparsity, and environmental complexity inherent in real-world wireless sensing scenarios.

[0063] Simulation results evaluation

[0064] In the CMCTNet model, input data consists of three modalities: RSRP data, SAR imagery, and building height maps, each with an input size of 256×256. First, preliminary spatial feature extraction is performed on each of these three input types using independent wavelet blocks, generating 32-channel low-level feature maps. Next, the RSRP, SAR, and building data enter their respective encoder branches, which consist of multiple dense blocks and are downsampled layer by layer using max pooling after each layer to extract more abstract features. In the RSRP encoder, the features are upscaled from 32 channels to 64, 64, and 128 channels, ultimately extracting 256-channel deep semantic features. The SAR and building branches follow the same structural process, extracting their own 256-channel high-level features. At this point, three deep semantic feature maps with half the spatial size and 256 channels are generated for each of the three modalities. The high-level features of the three modalities are then fed into a fusion attention module. This module adaptively fuses the complementary information between the three modalities using a multi-channel attention mechanism to form a comprehensive representation. Finally, a double attention mechanism is added to enhance inter-channel feature selection. The fused feature map then passes through two convolutional modules to further extract a 512-channel deep fused representation. A dropout layer is then used for regularization to prevent overfitting. The decoding stage employs a stepwise upsampling architecture: the fused features are first upsampled to 1 / 8 of the original spatial size (i.e., the same as before the last downsampling layer in the encoder), and then sequentially concatenated with the feature maps of the RSRP and building branches at the same level. The concatenated features at each upsampled step are further fused using a dense block to recover more spatial detail. Specifically, the 512-channel upsampled features are concatenated with two 256-channel features to form a 1024-channel representation. These features are then compressed and fused to 256 channels using a dense block. The same process is repeated at each subsequent step, ultimately restoring the 64-channel representation. In the final stage, the 64-channel decoded features are first reduced to 32 channels through a 7×7 convolution and activated by a ReLU. A 3×3 convolution then outputs the final prediction, the recovered RSRP distribution map, with an output channel of 1, consistent with the input. This entire process enables high-fidelity reconstruction of a complete electromagnetic environment map from sparse RSRP and auxiliary modalities (SAR and buildings).

[0065] In order to rigorously evaluate the reconstruction performance of the proposed CMCTNet under sparse sampling conditions, four widely used quantitative indicators are adopted: mean absolute error (MAE), root mean square error (RMSE), structural similarity index measure (SSIM) and peak signal-to-noise ratio (PSNR). MAE and RMSE are standard indicators of numerical accuracy, while SSIM and PSNR are used to evaluate the structural similarity and perceptual quality between the reconstructed radio environment map (REM) and the ground truth. These indicators are defined as follows:

[0066]

[0067]

[0068] Together, these four metrics provide a robust framework for evaluating the accuracy and perceptual consistency of reconstructed REM.

[0069] To improve the model's generalization and robustness with limited training data, data augmentation techniques are applied during training. Each training example is randomly flipped horizontally, vertically, or rotated 90 degrees with a probability of 50%. These geometric augmentations force the model to learn propagation patterns that are translationally and rotationally invariant, which is particularly important in heterogeneous urban environments, where signal behavior can vary significantly depending on layout orientation and structural symmetry.

[0070] To benchmark CMCTNet, it is compared with two representative deep learning based REM reconstruction methods:

[0071] 1. RobUNet: A UNet-based radio map construction algorithm that uses an encoder-decoder architecture to capture multi-scale spatial features. By leveraging hierarchical features, RobUNet achieves better generalization under limited measurements in complex urban environments.

[0072] 2. DeepREM: A two-branch model that combines U-Net and conditional generative adversarial network (CGAN) structures to estimate REM from sparse observations. Notably, DeepREM does not require auxiliary information, making it suitable for scenarios with minimal prior information about the environment.

[0073] All comparative experiments were conducted under consistent data partitioning and sampling protocols to ensure fairness.

[0074] Experiments were conducted using uniform random sampling with sampling rates of 5%, 3%, and 1%. For comparison, two baseline methods, DeepREM and RobUNet, were also included. The evaluation focused on four quantitative metrics—MAE, RMSE, SSIM, and PSNR—that capture the numerical accuracy and perceptual quality of the reconstructed wireless environment map.

[0075] Table 1 shows the MAE and RMSE results. CMCTNet consistently achieves the lowest reconstruction error at all sampling rates. As the sampling rate decreases, the performance gap between CMCTNet and the baseline becomes more pronounced, highlighting the proposed model's robustness under extremely sparse conditions. Notably, due to the lack of a structural prior, DeepREM's performance degrades significantly at lower sampling rates. RobUNet maintains relatively stable performance but still lags behind CMCTNet, particularly in low-sampling scenarios.

[0076] Table 1 Comparison of MAE and RMSE under uniform sampling

[0077]

[0078] To further evaluate structural consistency and perceptual fidelity, Table 2 reports the SSIM and PSNR scores at the same sampling rate. CMCTNet once again demonstrates superior performance, producing reconstructions with sharper spatial boundaries and higher visual quality. Even at an extreme 1% sampling rate, CMCTNet maintains an SSIM of 0.876 and a PSNR of 27.39 dB, significantly outperforming RobUNet and DeepREM.

[0079] Table 2 Comparison of SSIM and PSNR under uniform sampling

[0080]

[0081] At the most challenging 1% sampling rate, DeepREM's performance drops significantly, achieving only 0.498 SSIM and 20.76dB PSNR. In contrast, CMCTNet continues to produce high-fidelity reconstruction results, demonstrating its strong generalization ability and superior modeling of spatial and structural dependencies.

[0082] Figure 5 A visual comparison of REM reconstructed at different sampling levels is shown. CMCTNet shows sharper signal transitions and more accurate propagation boundaries, especially near building edges.

[0083] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation methods of the present invention. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation methods here. Any obvious changes or modifications derived from the technical solution of the present invention are still within the scope of protection of the present invention.

Claims

1. A high-precision radio environment map reconstruction method based on multimodal fusion, characterized in that: include: Step S01: using multiple parallel encoders to process multiple input information respectively to generate corresponding multiple potential feature representations; Step S 02: Using a fusion attention module to dynamically fuse the multiple latent features across modalities to obtain a fused latent feature representation, including: constructing a spatially adaptive attention map and using a dual attention mechanism to simultaneously model spatial and channel dependencies, grouping the spatially adaptive attention map, processing it through spatial or channel attention sub-branches, and finally shuffling and concatenating to promote cross-channel interaction; Step S03: reconstructing the fused latent feature representation using a decoder, including retaining spatial features with high resolution through skip connections and transferring location information from the encoder to the decoder to obtain a high-precision radio environment map.

2. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 1, characterized in that: In step S01 , the plurality of parallel encoders include three parallel encoders.

3. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 1, characterized in that: In step S01 , the plurality of input information includes a sparse RSRP map, a SAR image, and building data.

4. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 1, characterized in that: In step S01, the processing of the multiple input information separately includes: performing initial feature extraction on the multiple input information using discrete wavelet transform, decomposing the input information into multiple frequency bands, performing multi-level feature extraction processing on the multiple frequency bands through separate convolution layers, and generating potential feature representations.

5. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 4, characterized in that: The plurality of frequency bands include LL, LH, HL, and HH.

6. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 3, characterized in that: In step S02 , during the cross-modal dynamic fusion, the latent feature representation extracted from the sparse RSRP map is used as the query Q, and the fused latent feature representation extracted from the SAR image and the building data is used as the key K′ and the value V.

7. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 6, characterized in that: In step S02, the spatial adaptive attention map is constructed based on the following attention calculation formula: Where Q is the RSRP query tensor, K' and V are the key and value of modality fusion, d k is the dimension of the key vector.

8. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 1, characterized in that: In step S03, the decoder outputs a dense prediction map for estimating the complete RSRP distribution in the discrete area; The dense prediction graph is: in, The complete RSRP distribution is estimated by a multimodal deep learning framework, N x and N y They represent the number of grid cells along the x and y directions after the target area is discretized.

9. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 2, characterized in that: Each encoder includes several dense blocks.

10. The high-precision radio environment map reconstruction method based on multimodal fusion according to claim 9, wherein the forward propagation process within the dense block comprises: First, the first convolution is performed on the input to obtain the feature x1, and then x1 is spliced ​​with the original input in the channel dimension as the input of the second convolution block to obtain x2; then x1, the original input and x2 are spliced ​​in the channel dimension as the input of the third convolution block to obtain the final output x3.

Citation Information

Patent Citations

  • Multi-modal image segmentation method based on cross attention

    CN117911426A

  • Infrared and visible light fusion method based on multi-scale feature interaction enhancement

    CN119091269A

  • Semantic segmentation model and segmentation method for high-resolution remote sensing image

    CN119206229A

  • Pet imaging method and apparatus based on optical flow registration, and device and storage medium

    WO2025035380A1

Cited By

  • Single base station non-line-of-sight positioning method and device based on electromagnetic map

    CN122534594A