Electromagnetic propagation loss prediction method and system based on multi-modal fusion

By multimodal fusion of tabular data with satellite image data, and using cross-attention mechanism and adversarial training technology to extract features, the problem of low prediction accuracy of electromagnetic propagation loss in complex environments is solved, and higher prediction accuracy and robustness are achieved.

CN120185741APending Publication Date: 2025-06-20SHANGHAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510256420.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the prediction accuracy of electromagnetic propagation loss in complex environments is low, and the prediction accuracy and stability of the model in different environments are insufficient.

Method used

The electromagnetic propagation loss prediction method based on multimodal fusion is adopted, and the tabular data is combined with satellite image data, and the features are extracted through the cross attention mechanism, and the fusion features are used to train the fusion features to generate more information-rich fusion features.

Benefits of technology

It significantly improves the prediction accuracy and robustness of the model in different environments, and can more comprehensively capture various influencing factors in the propagation environment, improving the accuracy and reliability of the prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185741A_ABST
    Figure CN120185741A_ABST
Patent Text Reader

Abstract

The invention provides an electromagnetic propagation loss prediction method based on multi-modal fusion, and belongs to the field of electromagnetic propagation loss prediction, and the method comprises the steps: obtaining table data and a satellite image corresponding to the table data geographically, carrying out the graying processing of the table data, and carrying out the preprocessing of the satellite image, performing feature extraction on the grayscale image and the preprocessed satellite image to obtain table features and satellite image features; respectively calculating the attention weight of the table feature to the satellite image feature and the attention weight of the satellite image feature to the table feature based on a cross attention mechanism, generating the table feature and the satellite image feature after interaction based on the two attention weights, and fusing the table feature and the satellite image feature after interaction by adopting adversarial training. Training a preset propagation loss prediction model based on the fused features, inputting the to-be-detected signal into the trained propagation loss prediction model, and outputting propagation loss; the invention further provides a prediction system. The problem of low propagation loss prediction precision in a complex environment is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electromagnetic propagation loss prediction, and particularly to an electromagnetic propagation loss prediction method and system based on multimodal fusion. Background Art

[0002] The planning and optimization of urban communication networks is a key challenge in the field of wireless communication. In urban environments, due to complex terrain characteristics such as high-rise buildings and narrow roads, signal propagation is often significantly affected by occlusion and multipath effects, resulting in increased propagation loss and insufficient communication coverage. In addition, the dynamically changing population density and traffic flow in cities further increase the complexity of propagation loss prediction. To improve the prediction accuracy of electromagnetic propagation loss, in the prior art, for example, the Chinese patent application for invention "A Method for Predicting Wireless Channel Propagation Loss Driven by Measured Data" with the publication number CN117220803A adopts different strategies according to the sufficiency of measured data: when the data is sufficient, an artificial neural network is used for training; when the data is insufficient, a dual-weight neuron network is adopted, first pre-trained with simulation data and then optimized with measured data; when the data is missing, through parameter migration between scenarios, combined with the measured and simulation data of the reference scenario and the simulation data of the scenario to be measured for training. This prediction method proposes a solution of simulation data and parameter migration, but the training effect still depends on the authenticity and diversity of the data. In the case of complete absence of measured data, the similarity between the reference scenario and the scenario to be measured determines the migration effect, and the reliability of the prediction result may not be guaranteed; moreover, due to the complexity of the network structure and parameters, the dual-weight neuron structure and the multi-stage training method may lead to an increase in model complexity and a large computational overhead.

[0003] The Chinese patent application for invention "A Method for Predicting Over-the-Horizon Propagation Loss Based on SL-TrellisNets Network" with the publication number CN114611415A proposes a short-term and long-term parallel network time convolution (SL-TrellisNets) network to improve the prediction accuracy of over-the-horizon propagation loss. The model complexity is too large. Using the short-term and long-term parallel network time convolution (SL-TrellisNets) model may increase the computational complexity and require high computing resources. Especially when dealing with large-scale data or real-time applications, there may be problems of insufficient efficiency; this method requires a large amount of environmental data as input, and the quality and accuracy of these data directly affect the prediction effect and are dependent on the input features.

[0004] The Chinese patent application for invention "Method and Device for Determining Model Parameters, Method and Device for Loss Prediction" with the publication number CN118199767A obtains the training samples and loss values of the first model in the source scenario, migrates the samples to the sample pool of the second model in the target scenario, constructs training samples, trains the second model with the training samples, records the training parameters and loss values, calculates the loss values of the first model and the second model through weighted calculation, selects the loss value corresponding to the weighted minimum value as the target loss value, and determines the training parameters corresponding to it as the final parameters of the second model. However, it does not clearly solve the problem of insufficient or low-quality samples in the target scenario, which may lead to limited model training effects. Moreover, repeatedly training the second deep learning model to optimize the parameters incurs high computational costs and low efficiency.

[0005] The data sources of the propagation loss prediction model include tabular data and image data. The tabular data contains environmental parameters such as temperature, humidity, and building materials, which have a direct impact on signal propagation. However, single tabular data can only provide quantitative information and usually lacks geospatial information. Satellite images, on the other hand, can provide both qualitative and quantitative information and contain rich spatial features such as terrain undulations and building distributions, which can make up for the lack of spatial distribution information in tabular data. Under different terrain and environmental conditions, a single data source may have data missing or noise problems in some areas, and a single data source cannot accurately reflect the changes in signal propagation. Therefore, there is an urgent need for a precise prediction method that combines multi-modal data to optimize the coverage and overall performance of urban communication networks. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to solve the problem of low prediction accuracy of propagation loss in complex environments in the prior art, so as to improve the accuracy and stability of model prediction in different environments.

[0007] The present invention solves the above technical problems through the following technical solutions: An electromagnetic propagation loss prediction method based on multi-modal fusion, the method includes:

[0008] Obtain tabular data and satellite images geographically corresponding to the tabular data, perform grayscale processing on the tabular data to obtain a grayscale image, and perform preprocessing on the satellite images to obtain preprocessed satellite images;

[0009] Extract features from the grayscale image and the preprocessed satellite images respectively to obtain tabular features and satellite image features;

[0010] Based on the cross-attention mechanism, calculate the attention weights of the tabular features to the satellite image features and the attention weights of the satellite image features to the tabular features respectively, generate the interacted tabular features and satellite image features based on the two attention weights, and adopt adversarial training to fuse the interacted tabular features and satellite image features to obtain the fused features;

[0011] Train a preset propagation loss prediction model based on the fused features to obtain a trained propagation loss prediction model;

[0012] Input the signal to be measured into the trained propagation loss prediction model and output the propagation loss.

[0013] The present invention combines tabular data with satellite image data to combine the diversity of satellite images with the numerical stability of tabular data. The tabular data includes various key environmental parameters, geographical parameters, and signal parameters, providing an accurate numerical description. The satellite image makes up for the lack of spatial distribution information in tabular data, such as the specific distribution of buildings, road grids, and open areas, by providing high-resolution geographical images. Based on the cross-attention mechanism, deep interaction between tabular data and satellite image features is achieved, allowing two-way interactive learning between tabular data and satellite image features before fusion, effectively capturing the potential associations and complementary information between the two modalities, not only enhancing the mutual influence between the two modality features, but also promoting the generation of more informative fused features. Through the deep fusion of multi-modal features, the model can more comprehensively capture various influencing factors in the propagation environment, and can significantly improve the prediction accuracy and robustness of the model in different environments.

[0014] Preferably, the process of grayscale processing the tabular data includes:

[0015] Perform missing value processing, outlier detection, and data format unification on the tabular data in sequence to obtain preprocessed tabular data;

[0016] Normalize the preprocessed tabular data to obtain normalized tabular data;

[0017] Map the normalized tabular data into a two-dimensional pixel grid to obtain a grayscale image.

[0018] Preferably, the process of preprocessing the satellite image includes:

[0019] Crop the corresponding area of the satellite image in the satellite image according to the geographical coordinates corresponding to the tabular data;

[0020] Perform resolution unification, color space conversion, denoising, image enhancement, and normalization on the cropped satellite image in sequence to obtain preprocessed satellite image.

[0021] Preferably, a CNN network is used to extract features from the grayscale image. The grayscale image sequentially passes through a first convolutional pooling layer, a second convolutional pooling layer, a third convolutional pooling layer, and a fully connected layer. In the first convolutional pooling layer, the convolutional layer uses 32 convolutional kernels of size 3×3, with a stride of 1, and the ReLU activation function. In the second convolutional pooling layer, the convolutional layer uses 64 convolutional kernels of size 3×3, with a stride of 1, and the ReLU activation function. In the third convolutional pooling layer, the convolutional layer uses 128 convolutional kernels of size 3×3, with a stride of 1, and the ReLU activation function. The multi-dimensional features extracted from the grayscale image through the three convolutional pooling layers are flattened and then connected to a 128-dimensional fully connected layer. The fully connected layer uses the ReLU activation function. Batch normalization and Dropout operations are introduced after each convolutional layer and fully connected layer.

[0022] Preferably, a DNN network is used to extract features from the preprocessed satellite image. The preprocessed satellite image sequentially passes through a global average pooling layer and a fully connected layer. The global average pooling layer is used to extract deep features from the preprocessed satellite image and perform dimensionality reduction. The feature vector after dimensionality reduction is input into a 256-dimensional fully connected layer to generate 256-dimensional satellite image features. The fully connected layer uses the ReLU activation function.

[0023] Preferably, the attention weight A of the tabular features to the satellite image features tab2sat is:

[0024]

[0025] The attention weight A of the satellite image features to the tabular features sat2tab is:

[0026]

[0027] where Q tab , K tab are respectively the query and key of the tabular feature X tab , Q tab = W q1 X tab , K tab = W k1 X tab , Q sat , K sat are respectively the query and key of the satellite image feature X sat , Q sat = W q2 X sat , K sat = W k2 X sat , W q1 , W k1 , W q2 , Wk2 is the weight matrix, and d k is the dimension of the key vector.

[0028] Preferably, the process of fusing the tabular features and satellite image features after interaction by adversarial training includes:

[0029] Input the tabular features and satellite image features after interaction into the trained generator to obtain the features to be evaluated after fusion;

[0030] Input the features to be evaluated after fusion into the trained discriminator to obtain a scalar value. If the scalar value is 1, it means that the features to be evaluated after fusion are real features. If the scalar value is 0, it means that the features to be evaluated after fusion are generated pseudo-features, and use the real features as the fused features.

[0031] Preferably, the generator and the discriminator are alternately trained, and the loss function L G for the generator training is:

[0032] L G = -logD(Fusion Feature)

[0033] The loss function L D for the discriminator training is:

[0034] L D = -[logD(Real Feature) + log(1 - D(Fake Feature))].

[0035] Preferably, the fused features are reduced in dimension and non-linearly transformed through a fully connected layer to obtain the 256-dimensional fused feature z1 as:

[0036] z1 = ReLU(W1h fusion + b1)

[0037] where W1 ∈ R d×384 , b1 ∈ R d , and d is the set feature dimension.

[0038] The present invention also provides an electromagnetic propagation loss prediction system based on multimodal fusion. The system includes:

[0039] A data preprocessing module for obtaining tabular data and satellite images geographically corresponding to the tabular data, performing grayscale processing on the tabular data to obtain a grayscale image, and preprocessing the satellite images to obtain preprocessed satellite images;

[0040] A feature extraction module for respectively extracting features from the grayscale image and the preprocessed satellite images to obtain tabular features and satellite image features;

[0041] A multi-modal feature fusion module, which is used to calculate the attention weights of tabular features to satellite image features and the attention weights of satellite image features to tabular features respectively based on the cross-attention mechanism, generate the interacted tabular features and satellite image features based on the two attention weights, and adopt adversarial training to fuse the interacted tabular features and satellite image features to obtain the fused features;

[0042] A model training module, which is used to train a preset propagation loss prediction model based on the fused features to obtain a trained propagation loss prediction model;

[0043] A prediction module, which is used to input a signal to be measured into the trained propagation loss prediction model and output the propagation loss.

[0044] The advantages provided by the present invention are as follows:

[0045] (1) The present invention combines tabular data with satellite image data to realize the combination of the diversity of satellite images and the numerical stability of tabular data. The tabular data includes various key environmental parameters, geographical parameters, and signal parameters, providing an accurate numerical description. The satellite image makes up for the lack of spatial distribution information in tabular data, such as the specific distribution of buildings, road grids, and open areas, by providing high-resolution geographical images. Based on the cross-attention mechanism, deep interaction between tabular data and satellite image features is realized, allowing two-way interactive learning between tabular data and satellite image features before fusion, effectively capturing the potential associations and complementary information between the two modalities, not only enhancing the mutual influence between the two modal features, but also promoting the generation of more informative fused features. Through the deep fusion of multi-modal features, the model can more comprehensively capture various influencing factors in the propagation environment, and can significantly improve the prediction accuracy and robustness of the model in different environments.

[0046] (2) The grayscale processing of the tabular data in the present invention preprocesses the original numerical data (such as temperature, humidity, altitude, etc.), including missing value filling, outlier detection, and normalization operations, standardizes the tabular data, and maps it into a two-dimensional grayscale image, thereby endowing the data with spatial structure characteristics. This processing method not only meets the input requirements of the deep learning model, but also enhances the expression ability of the spatial distribution characteristics of the data.

[0047] (3) The satellite image preprocessing in the present invention obtains satellite remote sensing data covering the target area. Through operations such as resolution unification, denoising, color space conversion, and data augmentation, this module effectively eliminates the image quality differences and optimizes the image feature expression ability. The cropping operation ensures the precise correspondence between the image and the tabular data in the geographical range, providing a reliable basis for subsequent feature extraction.

[0048] (4) The present invention extracts hierarchical features from grayscale images through a lightweight convolutional neural network (CNN), capturing both local information of the data and integrating global features. The model generates a 128-dimensional feature vector through a combination of convolutional, pooling, and fully connected layers to represent the deep features of tabular data. By introducing a global average pooling layer and a fully connected layer, high-dimensional feature representation of the input satellite images is achieved, and finally a 256-dimensional feature vector is generated as a compact representation of the image.

[0049] (5) Through adversarial training in the present invention, the mutual game between the generator and the discriminator makes the model more robust when processing data of different modalities, improving the accuracy and reliability of electromagnetic propagation loss prediction. Especially in the face of noise, missing data, or data imbalance, the prediction accuracy can be effectively improved. Through the adversarial game in the feature space, multi-modal information of tabular data and satellite images can be more effectively fused, enhancing the feature representation ability. At the same time, the model has strong adaptability and can achieve high-precision propagation loss prediction under different propagation environments and various input data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flowchart of the electromagnetic propagation loss prediction method based on multi-modal fusion provided by an embodiment of the present invention;

[0051] Figure 2 is a schematic diagram of the electromagnetic propagation loss prediction method based on multi-modal fusion provided by an embodiment of the present invention;

[0052] Figure 3 is a flow block diagram of the electromagnetic propagation loss prediction method based on multi-modal fusion provided by an embodiment of the present invention;

[0053] Figure 4 is a schematic diagram of the electromagnetic propagation loss prediction system based on multi-modal fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely with reference to specific embodiments and the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0055] As Figure 1 shown, this embodiment provides an electromagnetic propagation loss prediction method based on multi-modal fusion, including the following steps:

[0056] Step 1: Obtain tabular data and satellite images geographically corresponding to the tabular data. Perform grayscale processing on the tabular data to obtain a grayscale image, and preprocess the satellite images to obtain preprocessed satellite images.

[0057] Tabular data usually includes distance-related parameters, frequency-related parameters, antenna parameters, etc. In special environments or when physical fields need to be considered, environmental-related parameters also need to be introduced for predicting propagation loss. In the present invention, the tabular data is obtained through simulation using the IIU-R traditional model.

[0058] The process of performing grayscale processing on the tabular data includes:

[0059] S111: Successively perform missing value processing, outlier detection, and data format unification on the tabular data to obtain preprocessed tabular data. Among them, the method for missing value processing is to check and process the missing values in the tabular data, and use mean filling, interpolation method, or delete the records containing missing values. Outlier detection uses methods such as box plots and the 3σ principle to identify and process outliers. Data format unification can ensure that all numerical data types are consistent and unify the data units, facilitating subsequent calculations. Through missing value processing, outlier detection, and data format unification, the data quality is ensured, providing reliable input for the model.

[0060] S112: Normalize the preprocessed tabular data to obtain normalized tabular data. The normalization operation uniformly scales data with different dimensions to the grayscale range [0, 255] to eliminate the dimensional differences between features and avoid some features having too much impact on the model. In the present invention, the Min-Max normalization method is adopted to scale the preprocessed tabular data to the interval [0, 1]. The normalized tabular data x′ is:

[0061]

[0062] Then map the normalized data to the grayscale range of [0, 255]:

[0063] Gray = x ′ × 255

[0064] S113: Map the normalized tabular data into a two-dimensional pixel grid to obtain a grayscale image.

[0065] Map the normalized tabular data into a two-dimensional pixel grid, and the resulting grayscale image endows the numerical data with a spatial structure, making it suitable for processing by a convolutional neural network. Unify the image size to meet the input requirements of the deep learning model and ensure the consistency of subsequent model processing. Map the normalized tabular data into the pixel grid of a two-dimensional grayscale image according to the row and column structure. For high-dimensional data, different dimensions are used as different channels of the image to generate a multi-channel grayscale image. The present invention uses bilinear or bicubic interpolation methods to adjust the generated grayscale image to a fixed resolution (such as 224×224 pixels) to meet the input requirements of the subsequent deep learning model.

[0066] The grayscale processing of tabular data preprocesses the original numerical data (such as temperature, humidity, altitude, etc.), including missing value filling, outlier detection, and normalization operations, standardizes the tabular data, and maps it into a two-dimensional grayscale image, thereby endowing the data with spatial structure characteristics. This processing method not only meets the input requirements of the deep learning model but also enhances the expression ability of the spatial distribution characteristics of the data.

[0067] The process of preprocessing satellite images includes:

[0068] S121. According to the geographical coordinates corresponding to the tabular data, crop the satellite image of the corresponding area in the satellite image; obtain the geographical and environmental information of the target area, provide additional spatial features for the model, enrich the data source, and improve the comprehensiveness of prediction. Obtain the image covering the target area from the satellite data source to ensure the geographical correspondence with the tabular data. Obtain the satellite remote sensing image covering the target area, and crop the satellite image of the corresponding area according to the geographical coordinates corresponding to the tabular data.

[0069] S122. Successively perform resolution unification, color space conversion, denoising, image enhancement, and normalization on the cropped satellite image to obtain the preprocessed satellite image.

[0070] Resolution unification adjusts all images to the same resolution and size. Color space conversion converts the image from RGB to grayscale or extracts specific band information according to requirements. Denoising processing uses methods such as Gaussian filters and median filters to remove noise and improve the image quality. Image enhancement adopts techniques such as histogram equalization and contrast stretching to enhance the details and contrast of the image. Normalization processing normalizes the pixel values to the interval [0, 1]. By resolution unification, color space conversion, denoising, and image enhancement, the quality differences between different images are eliminated, and the model's ability to extract image features is improved. In addition, the present invention increases data diversity through operations such as random rotation, flipping, translation, and scaling, and can also adjust brightness, contrast, and saturation to simulate different lighting conditions. Through geometric transformation and color perturbation, the data diversity is increased, overfitting of the model is prevented, and the generalization ability is improved.

[0071] Satellite image preprocessing obtains satellite remote sensing data covering the target area. Through operations such as resolution unification, denoising, color space conversion, and data augmentation, this module effectively eliminates image quality differences and optimizes the image feature expression ability. In addition, the cropping operation ensures the precise correspondence between the image and the tabular data in terms of geographical scope, providing a reliable basis for subsequent feature extraction.

[0072] Step 2: Feature extraction is performed on the grayscale image and the preprocessed satellite image respectively to obtain tabular features and satellite image features.

[0073] Feature representation learning aims to extract meaningful features from data of different modalities, providing a solid foundation for subsequent feature fusion and prediction. In the present invention, hierarchical features are extracted from the grayscale image through a lightweight convolutional neural network (CNN), which captures both local information of the data and integrates global features. The model generates a 128-dimensional feature vector through a combination of convolutional, pooling, and fully connected layers to represent the deep features of the tabular data.

[0074] In the present invention, a CNN network is used to extract features from the grayscale image. The grayscale image sequentially passes through the first convolutional pooling layer, the second convolutional pooling layer, the third convolutional pooling layer, and the fully connected layer. In the first convolutional pooling layer, the convolutional layer uses 32 convolutional kernels of size 3×3 with a stride of 1 and the activation function ReLU. In the second convolutional pooling layer, the convolutional layer uses 64 convolutional kernels of size 3×3 with a stride of 1 and the activation function ReLU. In the third convolutional pooling layer, the convolutional layer uses 128 convolutional kernels of size 3×3 with a stride of 1 and the activation function ReLU. The multi-dimensional features extracted by the grayscale image passing through the three convolutional pooling layers are flattened and then connected to a 128-dimensional fully connected layer. The fully connected layer uses the activation function ReLU, and batch normalization and Dropout operations are introduced after each convolutional layer and fully connected layer.

[0075] The first convolutional pooling layer can effectively extract the low-level features of the input data, such as edges and textures. The feature map is downsampled by a max pooling layer with a window size of 2×2, which not only reduces the spatial dimension of the feature map to compress the data volume but also retains the key features. The second convolutional pooling layer further extracts the intermediate-level features of the image, such as local patterns and shape information. Subsequently, the feature map is dimensionally reduced using the same 2×2 max pooling layer as the first layer to further compress the data and strengthen the important features. The third convolutional pooling layer focuses on extracting high-level features and global structure information and can capture more complex patterns. Subsequently, the final feature compression is completed through a 2×2 max pooling layer to prepare a high-quality feature representation for the subsequent fully connected layer. After three layers of convolution and pooling operations, the network flattens the extracted multi-dimensional feature map and connects it to a 128-dimensional fully connected layer. The fully connected layer uses the ReLU activation function to enhance the expression ability of non-linear features, thereby generating a compact feature vector suitable for subsequent tasks.

[0076] To improve the generalization ability of the model and reduce the risk of overfitting, batch normalization and Dropout operations are introduced after each convolutional layer and fully connected layer. Batch normalization helps to accelerate the training process and balance the distribution of inputs in each layer of the network, while Dropout effectively reduces the network's dependence on individual neurons by randomly discarding some neurons, thereby improving the robustness and generalization ability of the model.

[0077] The present invention uses a DNN network to extract features from the preprocessed satellite images. The preprocessed satellite images sequentially pass through a global average pooling layer and a fully connected layer. The global average pooling layer is used to extract deep features from the preprocessed satellite images and perform dimensional reduction. The dimensionally reduced feature vector is input into a 256-dimensional fully connected layer to generate 256-dimensional satellite image features, and the fully connected layer uses the ReLU activation function.

[0078] The present invention utilizes a pre-trained deep model to fully exploit its feature extraction ability learned from large-scale image data. By introducing a global average pooling layer (Global Average Pooling, GAP) and a fully connected layer, a high-dimensional feature representation of the input satellite image is achieved, and finally a 256-dimensional feature vector is generated as a compact representation of the image. To further alleviate the overfitting problem, the first few layers of the deep convolutional neural network are also frozen, and only the weights of the custom layer are updated. By adding a global average pooling layer, it is used to reduce the dimension of the deep features, replacing the traditional fully connected layer to avoid the problem of excessive number of parameters. GAP maps the feature map from the spatial dimension to the channel dimension by calculating the average value of each channel of the feature map. A 256-dimensional fully connected layer is added with the activation function ReLU. This layer further learns high-level feature representations and compresses the feature vector to 256 dimensions to maintain computational efficiency. After the above custom layer processing, a 256-dimensional feature vector is finally generated as the high-dimensional feature representation of the satellite image for subsequent tasks.

[0079] Step 3: Calculate the attention weights of the table features to the satellite image features and the attention weights of the satellite image features to the table features respectively based on the cross-attention mechanism, generate the interacted table features and satellite image features based on the two attention weights, and adopt adversarial training to fuse the interacted table features and satellite image features to obtain the fused features.

[0080] For the table features (128-dimensional) and the satellite image features (256-dimensional), corresponding query (Query), key (Key), and value (Value) vectors are generated respectively.

[0081] Q tab =W q1 X tab ,K tab =W k1 X tab, V tab =W v1 X tab

[0082] Q sat =W q2 X sat ,K sat =W k2 X sat, V sat =W v2 X sat

[0083] where W q1 ,W q2 、W k1 、W k2 、W v1, W v2 is the weight matrix.

[0084] The query vector of the table features interacts with the key vector of the satellite image to calculate the attention weights. At the same time, the query vector of the satellite image also interacts with the key vector of the table features to calculate another set of attention weights. Both sets of weights are calculated through the Dot-Product Attention mechanism, ensuring the accuracy and effectiveness of the interaction. Subsequently, the calculated attention weights are normalized by the Softmax function to ensure that the sum of all weights is 1, thus guaranteeing the stability of the feature interaction. Finally, the normalized attention weights are used to perform weighted summation on their respective value vectors to generate the feature representation after interaction.

[0085] The attention weight A of the table features on the satellite image features tab2sat is:

[0086]

[0087] The attention weight A of the satellite image features on the table features sat2tab is:

[0088]

[0089] where d k is the dimension of the key vector.

[0090] In the prior art, the concatenation of table features and image features only retains their respective independent information. In contrast, the present invention enables the features of the two modalities to mutually correct and enhance each other through a bidirectional attention mechanism. The traditional concatenation method may lead to a sharp increase in the feature dimension (such as directly concatenating to obtain 384 dimensions), while the present invention effectively filters out key information through attention-weighted fusion, reducing redundant calculations. The Cross-Attention mechanism realizes the deep interaction between table data and satellite image features, allowing the table data and satellite image features to perform bidirectional interactive learning before fusion. The core of the Cross-Attention mechanism lies in using the attention mechanism to enable the features of one modality to dynamically "focus on" the features of the other modality, thereby effectively capturing the potential associations and complementary information between the two modalities. This not only enhances the mutual influence between the features of the two modalities but also promotes the generation of more informative fusion features.

[0091] By combining the grayscale processing of tabular data and the feature extraction of satellite images, and using the cross-attention mechanism of the cross-modal interaction module to achieve the efficient fusion of cross-modal data, the present invention solves the problems of low prediction accuracy in complex environments and difficult feature fusion in the prior art, and improves the generalization ability and real-time prediction performance of the model. The fusion method can utilize the feature extraction ability of image data and reduce the dependence on the processing of complex tabular data.

[0092] The process of using adversarial training to fuse the tabular features and satellite image features after interaction includes:

[0093] Input the tabular features and satellite image features after interaction into the trained generator to obtain the fused features to be evaluated;

[0094] Input the fused features to be evaluated into the trained discriminator to obtain a scalar value. If the scalar value is 1, it means that the fused features to be evaluated are real features. If the scalar value is 0, it means that the fused features to be evaluated are generated pseudo-features, and use the real features as the fused features.

[0095] The task of the generator is to generate fused features based on the interaction features of tabular data and satellite images. It receives the tabular data features and satellite image features processed by the cross-attention mechanism. The tabular data includes environmental parameters (such as temperature, humidity, etc.), and the satellite image provides geospatial information (such as terrain, buildings, etc.). Through the cross-attention mechanism, the tabular data and satellite image features interact to form a joint feature vector.

[0096] The generator further processes these interaction features using a convolutional neural network (CNN). Low-level spatial information is extracted through multiple convolutional layers, and the modeling ability for environmental influencing factors is enhanced through fully connected layers. The input interaction features undergo a non-linear transformation to generate a fused feature vector, which contains spatial and environmental information from tabular data and satellite images. The fused feature vector output by the generator will be passed as input to the discriminator for evaluation.

[0097] The task of the discriminator is to evaluate the authenticity of the generated fused features and compare them with the fused features in the actual data. The discriminator receives the fused feature vector generated by the generator or the real fused features from the actual data. The features are further extracted through a convolutional neural network (CNN), focusing on high-level spatial structure information, and comprehensively processed through fully connected layers to determine whether the feature conforms to the distribution of the real data.

[0098] The key to adversarial training lies in the game process between the generator and the discriminator. The generator optimizes the fused features it outputs by maximizing the misjudgment probability of the discriminator for the features it generates. Specifically, the goal of the generator is to optimize the feature generation process so that the discriminator has difficulty distinguishing between "real" and "generated" features. The discriminator is optimized by accurately distinguishing between generated features and real features, and its goal is to minimize the misjudgment rate. For the generator, the generator loss function L G is used to make it generate "real" fused features as much as possible. For the discriminator, the discriminator loss function L D is used to optimize its discrimination ability.

[0099] The generator and the discriminator are alternately trained, and the Adam optimizer is used to update the network parameters to minimize the loss functions of the generator and the discriminator. The loss function L G for generator training is as follows:

[0100] L G = -logD(Fusion Feature)

[0101] The loss function L D for discriminator training is as follows:

[0102] L D = -[logD(Real Feature) + log(1 - D(Fake Feature))].

[0103] Since the dimension of the fused features is relatively high (384 dimensions), the fused features need to be reduced in dimension and non-linearly transformed through a fully connected layer to obtain the 256-dimensional fused feature z1 as follows:

[0104] z1 = ReLU(W1h fusion + b1)

[0105] where W1 ∈ R d×384 , b1 ∈ R d , and d is the set feature dimension.

[0106] Through adversarial training, the mutual game between the generator and the discriminator makes the model more robust when processing data of different modalities, improving the accuracy and reliability of electromagnetic propagation loss prediction. Especially in the face of noise, missing data, or data imbalance, it can effectively improve the prediction accuracy. Through the adversarial game in the feature space, the multi-modal information of tabular data and satellite images can be more effectively fused, enhancing the feature representation ability. At the same time, the model is made to have strong adaptability, and high-precision propagation loss prediction can be achieved under different propagation environments and various input data.

[0107] Step 4: Train a preset propagation loss prediction model based on the fused features to obtain a trained propagation loss prediction model; the fused features are first input into the feature extraction layer in the model. Deep feature extraction and prediction further process the fused features using a deep learning model to extract high-level features and achieve accurate prediction of electromagnetic propagation loss.

[0108] The feature extraction layer uses a five-layer convolutional neural network to perform in-depth feature extraction on the fused features, capturing complex environmental patterns and spatial relationships. It processes the basic features of the signal (such as frequency, power), extracts their important information, and prevents the signal features from being weakened during the fusion process. The environmental features and signal features are concatenated to form a comprehensive feature vector. Through the fully connected layer, the predicted value of the electromagnetic propagation loss is output, completing the mapping from features to prediction.

[0109] Independently train the feature extraction networks of each modality to ensure that they can effectively capture key features. On this basis, freeze the parameters of the feature extraction networks and focus on training the multi-modal fusion module and the prediction output layer to optimize the fusion effect and enhance the overall performance of the model. Subsequently, unfreeze all network layers and perform end-to-end fine-tuning to further improve the coordination and overall performance of each module of the model.

[0110] When training the preset propagation loss prediction model, the mean squared error (MSE) is selected as the loss function, which is suitable for regression tasks and can effectively measure the difference between the predicted value and the true value. At the same time, the Adam optimizer is used, and its excellent adaptive learning rate mechanism is used to accelerate the convergence of the model, improving the training efficiency and effect. Among them, the mean squared error L:

[0111]

[0112] Among them, is the predicted value, y i is the true value, and N is the number of samples.

[0113] After each training stage, use the validation set to evaluate the model to guide the adjustment and optimization strategy of the model, and measure the accuracy using the mean absolute percentage error (MAPE). At the same time, comprehensively evaluate the generalization ability of the model through K-fold cross-validation, effectively avoiding the evaluation bias caused by data partitioning. Through residual analysis and case analysis, deeply explore the performance deficiencies of the model in specific scenarios, so as to provide a clear direction for the further improvement of the model.

[0114] Step 5: Input the signal to be measured into the trained propagation loss prediction model and output the propagation loss.

[0115] In the optimal usage state of urban communication, the present invention combines tabular data with satellite image data. The combination of the diversity of satellite images and the numerical stability of tabular data significantly enhances the adaptability of the model in different terrains and environments, effectively solving the problem of low prediction accuracy of traditional methods in special environments. The propagation loss is accurately predicted using a deep learning model. The tabular data includes various key environmental parameters, geographical parameters, and signal parameters, providing an accurate numerical description. The satellite images, by providing high-resolution geographical images, make up for the lack of spatial distribution information in the tabular data, such as the specific distribution of buildings, road grids, and open areas. Through the deep fusion of multi-modal features, the model can more comprehensively capture various influencing factors in the propagation environment, thus significantly improving the prediction accuracy and robustness, providing a reliable solution for propagation loss prediction in complex environments.

[0116] The present invention first performs grayscale processing on the tabular data, converting it into a grayscale image with spatial structure characteristics. At the same time, it extracts geographical features from the satellite images to generate cross-modal fusion features. These fusion features can comprehensively reflect various influencing factors in the propagation environment, including environmental conditions, terrain characteristics, and spatial distribution. Finally, the model accurately predicts the distribution of propagation loss within the urban area based on the fusion features, providing a scientific basis for optimizing the coverage and performance of communication networks.

[0117] The present invention also provides an electromagnetic propagation loss prediction system based on multi-modal fusion, which can be used to execute the method of the present invention. For details not disclosed in the system of the present invention, please refer to the method of the present invention and will not be elaborated here. The system includes:

[0118] A data preprocessing module, used to obtain tabular data and satellite images geographically corresponding to the tabular data, perform grayscale processing on the tabular data to obtain a grayscale image, and preprocess the satellite images to obtain preprocessed satellite images;

[0119] A feature extraction module, used to extract features from the grayscale image and the preprocessed satellite images respectively to obtain tabular features and satellite image features;

[0120] A multi-modal feature fusion module, used to calculate the attention weights of tabular features to satellite image features and the attention weights of satellite image features to tabular features respectively based on the cross-attention mechanism, generate the interacted tabular features and satellite image features based on the two attention weights, and adopt adversarial training to fuse the interacted tabular features and satellite image features to obtain the fused features;

[0121] A model training module, used to train a preset propagation loss prediction model based on the fused features to obtain a trained propagation loss prediction model;

[0122] A prediction module, configured to input a signal to be measured into a trained propagation loss prediction model and output the propagation loss.

[0123] The design structures of the modules of the present invention are clear, the functions are independent and clear, which is not only convenient for optimization, but also has good scalability. The grayscale processing of the tabular data endows it with spatial structure characteristics, and the preprocessing of the satellite image ensures its matching with the tabular data in terms of geographical scope. These basic operations lay a solid foundation for subsequent multi-modal feature fusion and propagation loss prediction, ensuring the efficiency and reliability of the overall system.

[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. The electromagnetic propagation loss prediction method based on multi-modal fusion is characterized by: Methods include: Acquire the table data and the satellite image geographically corresponding to the table data, grayscale the table data to obtain the grayscale image, and preprocess the satellite image to obtain the preprocessed satellite image; Feature extraction is performed on the grayscale image and the preprocessed satellite image respectively to obtain table features and satellite image features; Based on the cross attention mechanism, the attention weights of the table features on the satellite image features and the attention weights of the satellite image features on the table features are calculated respectively. The interactive table features and satellite image features are generated based on the two attention weights. The interactive table features and satellite image features are fused by adversarial training to obtain the fused features. A preset propagation loss prediction model is trained based on the fused features to obtain a trained propagation loss prediction model; The signal to be tested is input into the trained propagation loss prediction model, and the propagation loss is output.

2. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1 is characterized in that: The process of graying the table data includes: The table data is processed with missing values, detected with outliers, and the data format is unified in turn to obtain the preprocessed table data; Normalizing the preprocessed tabular data to obtain normalized tabular data; The normalized tabular data is mapped to a two-dimensional pixel grid to obtain a grayscale image.

3. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1 is characterized in that: The process of preprocessing satellite images includes: According to the geographic coordinates corresponding to the table data, the satellite image of the corresponding area is cut out from the satellite image; The cropped satellite image is subjected to resolution unification, color space conversion, denoising, image enhancement, and normalization in sequence to obtain the preprocessed satellite image.

4. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1 is characterized in that: The CNN network is used to extract features of grayscale images. The grayscale images pass through the first convolution pooling layer, the second convolution pooling layer, the third convolution pooling layer and the fully connected layer in sequence. The convolution layer in the first convolution pooling layer uses 32 convolution kernels of size 3×3, with a step size of 1, and uses the activation function ReLU. The convolution layer in the second convolution pooling layer uses 64 convolution kernels of size 3×3, with a step size of 1, and uses the activation function ReLU. The convolution layer in the third convolution pooling layer uses 128 convolution kernels of size 3×3, with a step size of 1, and uses the activation function ReLU. The multi-dimensional features extracted by the grayscale image through the three convolution pooling layers are flattened and connected to the 128-dimensional fully connected layer. The fully connected layer uses the activation function ReLU. Batch normalization and Dropout operations are introduced after each convolution layer and fully connected layer.

5. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1 is characterized in that: The DNN network is used to extract features from the preprocessed satellite images. The preprocessed satellite images pass through the global average pooling layer and the fully connected layer in sequence. The global average pooling layer is used to extract deep features from the preprocessed satellite images and reduce their dimensionality. The feature vector after dimensionality reduction is input into the 256-dimensional fully connected layer to generate 256-dimensional satellite image features. The fully connected layer uses the activation function ReLU.

6. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1, characterized in that: The attention weight of the table feature to the satellite image feature A tab2sat for: Attention weight A of satellite image features to table features sat2tab for: Among them, Q tab , K tab They are the table features X tab The query,key,Q tab =W q1 X tab , K tab =W k1 X tab , Q sat , K sat They are satellite image features X sat The query,key,Q sat =W q2 X sat , K sat =W k2 X sat , W q1 , W k1 , W q2 , W k2 is the weight matrix, d k is the key vector dimension.

7. The electromagnetic propagation loss prediction method based on multimodal fusion according to claim 1 is characterized by: The process of using adversarial training to fuse the interactive table features and satellite image features includes: Input the interacted table features and satellite image features into the trained generator to obtain the fused features to be evaluated; The fused features to be evaluated are input into the trained discriminator to obtain a scalar value. If the scalar value is 1, it means that the fused features to be evaluated are real features. If the scalar value is 0, it means that the fused features to be evaluated are generated pseudo features, and the real features are used as the fused features.

8. The electromagnetic propagation loss prediction method based on multi-modal fusion according to claim 7 is characterized by: The generator and discriminator are trained alternately, and the loss function of the generator training is L G for: L G =-logD(Fusion Feature) The loss function L for discriminator training D for: L D =-[logD(Real Feature)+log(1-D(Fake Feature))]。 9. The electromagnetic propagation loss prediction method based on multi-modal fusion according to claim 7, characterized in that: The fused features are reduced in dimension and transformed nonlinearly through the fully connected layer, and the 256-dimensional fused feature z1 is obtained as follows: z1=ReLU(W1h fusion +b1) Where W1∈R d×384 , b1∈R d , d is the set feature dimension.

10. The electromagnetic propagation loss prediction system based on multi-modal fusion is characterized by: The system includes: A data preprocessing module is used to obtain the table data and the satellite image corresponding to the geographical location of the table data, grayscale the table data to obtain the grayscale image, and preprocess the satellite image to obtain the preprocessed satellite image; A feature extraction module is used to extract features from the grayscale image and the preprocessed satellite image to obtain table features and satellite image features; The multimodal feature fusion module is used to calculate the attention weights of the table features on the satellite image features and the attention weights of the satellite image features on the table features based on the cross attention mechanism, generate the interactive table features and satellite image features based on the two attention weights, and use adversarial training to fuse the interactive table features and satellite image features to obtain the fused features; A model training module, used for training a preset propagation loss prediction model based on the fused features to obtain a trained propagation loss prediction model; The prediction module is used to input the signal to be tested into the trained propagation loss prediction model and output the propagation loss.

Citation Information

Patent Citations

  • Over-the-horizon propagation loss prediction method based on SL-Trellis Nets network

    CN114611415A

  • Wireless channel propagation loss prediction method driven by measured data

    CN117220803A

  • Model parameter determination method and device and loss prediction method and device

    CN118199767A