Transform network-based troposphere beyond-visual-range communication channel path loss prediction method

By constructing a troposphere over-horizontal communication channel path loss prediction model based on Transformer network, combining WRF and NPS models to obtain environmental parameters, and using the two-dimensional parabolic equation method to predict, the problems of low path loss prediction accuracy and difficulty in fusion of environmental information in the prior art are solved, and more efficient path loss prediction is achieved.

CN119945596APending Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510107955.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult for the prior art to quickly and accurately predict the path loss of the offshore troposphere over-horizontal communication channel, especially in complex sea transmission environments, and existing deep learning models are difficult to effectively integrate environmental information and retain long-term dependencies in the information.

Method used

Using the troposphere over-horizontal communication channel path loss prediction method based on Transformer network, a prediction model including input layer, position encoding, encoder, decoder and output layer is constructed, and feature extraction and prediction are performed through an encoder and decoder composed of an embedded layer, normalization, self-attention layer and fully connected layer. At the same time, a subset of environmental meteorological parameters and evaporation waveguide height characteristics are obtained by using the WRF mesoscale numerical meteorological mode and NPS model, and path loss prediction is performed by combining the two-dimensional parabolic equation method.

Benefits of technology

By retaining the continuity of environmental information and time steps, the model can effectively capture the long-term dependence between evaporative waveguide height and path loss, improving the path loss prediction accuracy of the troposphere over-the-range visual communication link, and is suitable for complex and dynamic changes at sea.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945596A_ABST
    Figure CN119945596A_ABST
Patent Text Reader

Abstract

The invention relates to a troposphere over-the-horizon communication channel path loss prediction method based on a Transform network, and belongs to the field of communication technologies and deep learning. The method comprises the following steps: firstly, constructing a Transform network path loss prediction model; secondly, the method constructs data sets of a path loss prediction model, including a training data set and a test data set. The construction process of the data set comprises the following steps: firstly, based on a WRF mesoscale numerical meteorological mode, inverting marine troposphere beyond-visual-range environmental meteorological parameters, and performing large-area high-resolution long-time-efficiency evaporation waveguide height forecast by using an NPS model to obtain an evaporation waveguide height forecast value; secondly, predicting channel path loss by using a two-dimensional parabolic equation method, combining meteorological parameters, evaporation waveguide height and working parameters of communication equipment to form an X input feature, and taking a path loss predicted value as a Y input feature; and finally, model training optimization and performance evaluation are carried out by the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of communication technology and deep learning, and relates to a method for predicting path loss of tropospheric beyond-horizon communication channels based on a Transformer network. Background Art

[0002] As a crucial component of integrated air-space-ground-sea networks, maritime wireless communications face numerous challenges, one of which is path loss in the complex maritime transmission environment. Cross-sea beyond-line-of-sight communications primarily utilize tropospheric scatter and atmospheric waveguide communication. The time-varying nature of the ocean channel poses significant challenges in obtaining the channel response. Path loss calculation provides an effective way to evaluate signals in maritime environments, making accurate path loss prediction a key research topic.

[0003] Problems with existing technologies:

[0004] Deterministic methods such as ray tracing and parabolic equation method are difficult to achieve rapid predictions in applications due to their high complexity and the fact that their calculation time increases as the propagation range changes.

[0005] Most existing time series predictions based on deep learning models directly use measured data to create datasets, which are expensive to produce, difficult to obtain, and time-consuming. The measured data lacks the retention of environmental information, which affects the accuracy and effectiveness of the path loss prediction model. Existing deep learning models find it difficult to effectively integrate environmental information and retain long-term dependencies in the information.

[0006] Therefore, the present invention aims to provide a method for predicting path loss of tropospheric beyond-horizon communication channels that can effectively integrate environmental information and retain long-term dependencies in the information, so as to solve the problems existing in the prior art. Summary of the Invention

[0007] In view of this, the present invention aims to provide a Transformer network-based method for predicting path loss in tropospheric beyond-horizon communication channels. First, a Transformer network-based prediction model for tropospheric beyond-horizon communication channels is constructed, comprising an input layer, position encoding, an encoder, a decoder, and an output layer. The encoder and decoder consist of an embedding layer, a normalization layer, a self-attention layer, and a fully connected layer. Secondly, considering the environmental characteristics of microwave and scattering communication links at sea, an optimal combination of vertical and horizontal stratification and microphysics is set for a large region. Based on the WRF mesoscale numerical meteorological model, the environmental meteorological parameters of the communication link are inverted to construct a feature subset of meteorological parameters. The NPS model is used to predict the evaporation duct height over a large region with high resolution and long-term duration, obtaining a feature subset of evaporation duct height forecast values. Path loss is deterministically calculated using a two-dimensional parabolic equation, and a dataset is constructed by continuously arranging meteorological parameters, communication system operating parameters, and evaporation duct height in the time dimension. Finally, the datasets are grouped and embedded into the Transformer initial model for training and hyperparameter optimization. The prediction model performance is evaluated considering the characteristics of microwave and scattering communication applications at sea, verifying the accuracy and effectiveness of the method. This method enhances the model's applicability to natural scenarios by retaining environmental information in the dataset. Furthermore, by strictly ensuring the continuity of the time step, it captures the temporal information of the evaporation duct height and effectively preserves the long-term dependencies that may exist in time series data such as evaporation duct height and path loss. For multivariate time series such as path loss, this method can effectively predict the path loss variation with distance on a tropospheric beyond-horizon communication link, given atmospheric pressure, air temperature, relative humidity, sea surface temperature, wind speed, evaporation duct height, communication frequency, transmitting antenna height, and receiving antenna height. This method offers superior prediction accuracy and can be used for intelligent path loss prediction in maritime tropospheric beyond-horizon communications, providing a basis for performance and efficiency evaluation of equipment such as microwave and scattering communications.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] The method comprises the following steps:

[0010] Step 1: Considering the time-varying, environmental dynamics, and long-term dependence of the maritime tropospheric beyond-horizon communication channel on environmental meteorological parameters, a path loss prediction model for the tropospheric beyond-horizon communication channel based on a Transformer network is constructed. The model includes an input layer, position encoding, encoder, decoder, and output layer. The encoder and decoder consist of an embedding layer, normalization layer, self-attention layer, and fully connected layer.

[0011] Step 2: Construct a dataset for the path loss prediction model, including a training dataset and a test dataset. First, based on the environmental characteristics of microwave and scattering communication links at sea, an optimal combination of vertical and horizontal stratification and microphysics schemes is set for a large area. Based on the WRF mesoscale numerical meteorological model, the environmental meteorological parameters on the communication link are inverted to construct a characteristic subset of meteorological parameters. Second, the NPS model is used to predict the evaporation duct height over a large area with high resolution and long duration, obtaining a characteristic subset of the evaporation duct height prediction values.

[0012] Step 3: Predict channel path loss using the two-dimensional parabolic equation method to obtain a feature subset of path loss prediction values. Meanwhile, the operating parameters of microwave and scattering communications are combined to form an operating parameter feature subset. Finally, the meteorological parameter feature subset, the evaporation duct height feature subset, and the operating parameter feature subset are combined and grouped in the time domain to form the X input feature. The path loss prediction value feature subset is then constructed as the Y input feature.

[0013] Step 4: Model training optimization and performance evaluation: First, optimize the model hyperparameters based on the Tree-structured Parzen Estimator Approach (TPE) algorithm. Second, compare with the existing LSTM and GRU models to verify the feasibility and effectiveness of the proposed method for path loss prediction under offshore evaporation duct conditions in terms of mean absolute error, root mean square error, mean absolute percentage error, and relative error.

[0014] Optionally, the step 1 is specifically as follows:

[0015] Step 1-1: Prepare the embedding and position encoding of the input features, and find the corresponding vector embedding for each input token. Assume that the embedding dimension is d. In order to allow the model to perceive the sequence of environmental meteorological parameters and evaporation duct height, as well as the order of each position in the corresponding path loss sequence, positional encoding needs to be added. The common position encoding method is to use sine and cosine functions. The formula is as follows

[0016] PE (pos,2i) =sin(pos / 10000^(2i / d)) (1)

[0017] PE (pos,2i+1) =cos(pos / 10000^(2i / d)) (2)

[0018] Where pos is the position number of the input feature sequence (starting from 0), and i is half of the dimension subscript, which is used to distinguish between even and odd dimensions. Assuming the length of the input feature sequence is L, the position encoding and embedding are added to obtain a sequence representation with position information:

[0019]

[0020] Step 1-2: Multi-head self-attention mechanism construction, first perform single-head self-attention calculation, and generate query vector (Q), key vector (K) and value vector (V) from the input sequence X of each encoder through linear transformation

[0021] Q=XW Q (4)

[0022] K=XW K (5)

[0023] V=XW V (6)

[0024] in d k for (h is the number of attention heads), then calculate the attention score and weighted sum

[0025]

[0026] Z i =softmax(scores)×V (7) Assume that the input feature sequence L is a path loss sequence. When calculating the attention score scores of the first row of the path loss sequence, it is necessary to use each element in the input feature to score the path loss sequence. These scores are calculated by dot product of the key vectors of all elements of the input sequence and the query vector of the first row of the path loss sequence. The role of the Softmax function is to normalize the scores of all elements of the feature sequence. The obtained scores are all positive and the sum is 1, and finally the result vector Zi of the self-attention calculation is obtained. Multi-head attention can divide the input into h heads, calculate the above self-attention separately, and then horizontally splice their results together:

[0027] Z=concat(head1, head2,..., head h )W 0 (8)

[0028] Each head i =Z i ,

[0029] Steps 1-3: Residual connection (Add ResNet) and layer normalization (Layer Normalization, LN). After the input feature sequence passes through the multi-head attention mechanism to obtain the matrix Z, it is not directly passed to the fully connected neural network, but passes through the Add & Normalize layer. The expression is

[0030] LN(X+Z) (9)

[0031] Where X represents the input of the multi-head attention or feedforward network, and Z represents the output of the head attention or feedforward network;

[0032] Steps 1-4: Construct a feedforward network (FFN), expressed as

[0033] FFN(X)=max(0,XW1+b1)W2+b2 (10)

[0034] The weight matrix

[0035] Steps 1-5: Special processing in the decoder. In the multi-head self-attention at the decoder, a mask is required to prevent "seeing" future tokens (i.e., to ensure causality): the position t′>t is masked in the attention score, and the weight is set to 0 after the softmax.

[0036] scores[t, t′]=-∞ (11)

[0037] Taking the path loss feature sequence as an example, using a mask can block the path loss sequence results for the next moment when calculating the path loss sequence for the current moment. Subsequently, in the second attention sub-layer, the decoder interacts with the encoder output, helping the decoder "reference" the context of the original input sequence when generating predictions.

[0038] Steps 1-6: Construct the output layer. The final output layer of the decoder usually adds a linear mapping and softmax to output the next step of path loss sequence prediction. That is:

[0039] P(y t |y1,...,y t-1 , X)=softmax(ZW o +b o ) (12)

[0040] Where Z comes from the last layer output of the decoder.

[0041] Optionally, the step 2 is specifically as follows:

[0042] Step 2-1: Use the WRF mesoscale numerical meteorological model, select appropriate horizontal and vertical layers to obtain large-area, high-resolution environmental information, select the optimal microphysics scheme, and set the time span and time resolution to invert the environmental meteorological parameters in the tropospheric beyond-horizon link area;

[0043] Step 2-2: In post-processing, traverse each set moment and use the variable extraction function in wrf-python to extract all the inverted meteorological parameters at all grid points. According to the longitude and latitude coordinates of the two ends of the required link corresponding to the position in the grid, set the starting and ending points of the link, use the fitting function to generate the line pixel coordinates and remove duplicates, calculate the interpolation position of each point on the link, average the meteorological parameter values ​​at each point, and use them as the meteorological parameters for the link at the current time step to construct the tropospheric beyond-horizon link channel environment meteorological parameter data subset;

[0044] Step 2-3: The prediction of the evaporation duct height characteristic parameters needs to be calculated through the atmospheric correction refractive index. The atmospheric refractive index N is determined by the atmospheric pressure p (unit: hPa), the atmospheric temperature T (unit: K) and the water vapor partial pressure e (unit: hPa). The empirical relationship is:

[0045]

[0046] The calculation formula for water vapor partial pressure e is:

[0047]

[0048]

[0049] R h is the relative humidity of the atmosphere. In order to better study the effect of atmospheric refractive index on electromagnetic wave propagation, the earth's surface is approximately treated as a plane, and the atmospheric corrected refractive index M (unit M) is redefined. The relationship between it and the atmospheric refractive index is:

[0050]

[0051] Z is the altitude (unit: m), r e is the mean earth curvature, which is 6371 km. Substituting it into formula (15) yields

[0052] M=N+0.157z (17) In the NPS model, the vertical profile of temperature T and specific humidity q in the near-surface layer is expressed as

[0053]

[0054]

[0055] Where T0 and q0 are the sea surface temperature and specific humidity, T(z) and q(z) are the atmospheric temperature and specific humidity at the height z, respectively. * ,q * are the characteristic scales of potential temperature θ and specific humidity q, ψ h is the temperature universal function, κ is the Karman constant, Γ dis the dry adiabatic lapse rate, which is about 0.00976K / m, z 0t is the roughness height of atmospheric temperature, and L is the similarity length. The water vapor pressure profile, atmospheric temperature, and pressure are calculated using formulas (14), (18), and (19), and the results are then substituted into formulas (13) and (17) to obtain the atmospheric corrected refractive index profile. A subset of meteorological parameter data is input into the NPS model to obtain a high-resolution, long-term evaporation duct height forecast for a large area.

[0056] Optionally, the step three is specifically as follows:

[0057] Step 3-1: The corresponding recursive formula of the two-dimensional parabolic equation method is

[0058]

[0059] in and are Fourier transform and inverse transform respectively, is the refractive index term, which reflects the influence of space medium on electromagnetic waves. is the diffraction term, reflecting the diffraction effect of obstacles on the propagation path on the electromagnetic wave. Here, p = k0sinθ is the angular spectrum domain variable, and θ is the angle between the electromagnetic wave and the horizontal. Given the initial field distribution u(x0,z), the next step of the field distribution u(x0+Δx,z) can be obtained, thereby iteratively solving the field in the entire computational space. For the current tropospheric beyond-horizon link environment, the range of system parameters for microwave scattering communication equipment, including frequency, transmitting antenna height, and receiving antenna height, is specified. Combined with meteorological parameters and the predicted evaporation duct height, the two-dimensional bidirectional step-by-step parabolic equation method is used to predict the path loss on this link.

[0060] Step 3-2: Construct the dataset required for path loss sequence model training. The feature dimension of the dataset is (number of samples, time steps, feature dimension), where the number of samples represents the total number of combinations of meteorological parameters and communication equipment system parameters, and the time step represents the total number of time steps under the current time span and time resolution. The dataset is divided into X set and Y set, where the features of the X dataset are air temperature (AT), sea surface temperature (SST), atmospheric pressure (AP), relative humidity (RH), wind speed (WS), evaporation duct height (EDH), communication frequency (f), and transmitting antenna height (h). t ), receiving antenna height (h r), according to the system parameters of the communication equipment (f, h t 、h r ) is a sample, strictly ensuring that a sample contains all time step information. For example, if the time span is 72 hours and the time resolution is 1 hour, the total number of time steps is 72. The dimension of a sample in the X dataset is (72, 9), and the dimension of the X dataset is (number of samples, 72, 9). The Y dataset is characterized by a path loss (PL) sequence. Each row of PL corresponds to a combination of meteorological parameters and system parameters in a sample. Similarly, the continuity of time steps is guaranteed. After the arrangement is completed, the dataset is saved.

[0061] Optionally, the step 4 is specifically as follows:

[0062] Step 4-1: To compare and evaluate the models, we built a Long Short Term Memory (LSTM) and a Gate Recurrent Unit (GRU) model for comparison.

[0063] Step 4-2: Data loading and preprocessing:

[0064] (1) Read the dataset through the loading function and obtain the training data in groups (divided into X_train and Y_train).

[0065] (2) For X_train and Y_train, flatten the original (number of samples, time steps, feature dimensions) data to (number of samples × time steps, feature dimensions), normalize it, and then reshape it back to its original shape, while saving the normalizer.

[0066] (3) Divide the data into 70% training sets and 30% validation sets, convert them into tensors and encapsulate them into iterable batch data using data loaders;

[0067] Step 4-3: Model definition:

[0068] (1) The LSTM model is defined, which consists of the following parts: the input gate, which determines to what extent the current input is written into the cell state; the forget gate, which determines how much past information the current cell state can retain and what parts need to be forgotten; the output gate, which determines the content of the hidden state output at the current moment; the cell state: an information channel that runs through the time series, can store information over a long time step, and selectively update or retain content under the gating mechanism;

[0069] (2) The GRU model is defined, which includes the following parts: the update gate, which combines the functions of the "input gate" and the "forget gate" to use a single gate to determine whether to retain past information and accept new information at the current moment; the reset gate, which helps the network decide how much historical information to discard, thereby more flexibly capturing short-term dependencies;

[0070] (3) In the feedforward method, the input is transposed (from (number of samples, time steps, feature dimensions) to (sequence length, number of samples, feature dimensions)), then passed through the embedding layer and Transformer calculation, and finally transposed back to the original shape before being sent to the output layer;

[0071] Step 4-4: Training and validation functions:

[0072] (1) Define a function to train a single training round (epoch) and return the training error (mean square error MSE). The mean square error is defined as

[0073]

[0074] Where N represents the total number of samples, y i is the target vector, is the prediction vector, ||·||2 represents the 2nd-order norm;

[0075] (2) Define a function to perform inference on the validation set and return the validation error (mean square error MSE);

[0076] Step 4-5: Hyperparameter Optimization (Bayesian Optimization):

[0077] (1) Define the hyperparameter search space (learning rate lr, embedding dimension d, attention head h, encoder layer number en, decoder layer number dn, feedforward network dimension d ff , dropout rate).

[0078] (2) Define the parameter optimization function: instantiate and train the model based on the hyperparameters of the current experiment and return the validation loss.

[0079] (3) Use the tree-structured Parsons estimator algorithm to perform Bayesian optimization to find the optimal hyperparameters.

[0080] (4) Record the test process to obtain the optimal hyperparameters and save them;

[0081] Steps 4-6: Retrain the path loss prediction model using the optimal hyperparameters:

[0082] (1) Save the optimal hyperparameter values ​​and instantiate the final Transformer model.

[0083] (2) Set the training parameters, loss function and optimizer (Adam).

[0084] (3) Perform training in a defined training and validation cycle, and set up an early stopping mechanism (stop if the validation loss does not improve within a certain number of rounds).

[0085] (4) Whenever a better validation loss is obtained, save the current model;

[0086] Steps 4-7: Load the best model and evaluate it:

[0087] (1) Load the saved best model.

[0088] (2) Use the validation set for predictive reasoning and collect and concatenate the output and target values.

[0089] (3) Perform inverse normalization on the predicted results and target results.

[0090] (4) Calculate and output the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and the prediction probability (P err5 ) and the prediction probability with relative error ≤ 1% (P err1 ) and other evaluation indicators and record them in the log. The evaluation indicator expression is as follows

[0091]

[0092] The beneficial effects of the present invention are that it can use historical data to train the model, does not rely on explicit physical modeling or specific assumptions, but learns the radio wave propagation characteristics from large-scale data, and can adapt to more complex and dynamically changing sea surface environments. Compared with the data set constructed by measured data, this method enhances the applicability of the model to natural scenes by retaining environmental information in the data set, while strictly ensuring the continuity of the time step to capture the time information of the evaporation duct height, and can effectively retain the long-term dependencies that may exist in time series data such as the evaporation duct height and path loss. Transformer can learn a universal feature expression and has greater migration and generalization potential for different environmental parameters (such as different sea conditions, weather, etc.). Since Transformer only needs one feedforward in the inference stage, the prediction speed is faster, and the calculation time is usually in the millisecond level, which can meet real-time or quasi-real-time requirements. When processing long time series data, the model based on LSTM or GRU relies on the forward and backward memory transmission. The longer the sequence length, the more difficult and slower the training. The Transformer-based tropospheric beyond-horizon channel path loss prediction model, with its self-attention mechanism and parallel processing capabilities, is able to simultaneously focus on global long-range dependencies and different feature subspaces. Compared with existing LSTM and GRU models, the MAE is reduced by 61.17% and 71.9%, respectively; the RMSE is reduced by 67.68% and 74.27%, respectively; the MAPE is improved by 0.32% and 0.52%, respectively; and the prediction probability of a relative error ≤ 1% is increased by 11.32% and 16.9%, respectively. By preserving the spatiotemporal correlations of environmental features and evaporation ducts, this method can achieve dynamic, high-resolution, and long-term predictions for tropospheric beyond-horizon communication links. This method offers superior prediction accuracy and efficiency, and can be used for intelligent path loss prediction in maritime tropospheric beyond-horizon communications, providing a basis for performance and effectiveness evaluation of equipment such as microwave and scattering communications.

[0093] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0095] Figure 1 This is a flow chart of the path loss prediction method for tropospheric beyond-horizon communication channels based on Transformer networks;

[0096] Figure 2This is a schematic diagram of the Transformer structure. The number of input arrows for the multi-head attention is only for illustration and does not represent a restriction on the actual input features.

[0097] Figure 3 This is a schematic diagram of environmental meteorological parameter extraction. Taking sea surface temperature as an example, after inverting the current sea surface temperature at the grid point, the transmitting and receiving sites are set, and the line pixel coordinates are generated using a fitting function and duplicates are removed. The interpolated position of each point on the link is calculated.

[0098] Figure 4 The figure shows a comparison of the path loss prediction results of different methods, including six evaluation indicators, where the prediction time is the prediction time of a single sample. DETAILED DESCRIPTION

[0099] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0100] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0101] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0102] like Figure 1 As shown, the method for predicting path loss of tropospheric beyond-horizon communication channels based on a Transformer network of the present invention comprises the following steps:

[0103] Step 1: Considering the time-varying, environmental dynamics, and long-term dependence of the maritime tropospheric beyond-horizon communication channel on environmental meteorological parameters, a path loss prediction model for the tropospheric beyond-horizon communication channel based on a Transformer network is constructed. The model includes an input layer, position encoding, encoder, decoder, and output layer. The encoder and decoder consist of an embedding layer, normalization layer, self-attention layer, and fully connected layer.

[0104] Step 1-1: Prepare the embedding and position encoding of the input features, and find the corresponding vector embedding for each input token. Assume that the embedding dimension is d. In order to allow the model to perceive the sequence of environmental meteorological parameters and evaporation duct height, as well as the order of each position in the corresponding path loss sequence, positional encoding needs to be added. The common position encoding method is to use sine and cosine functions. The formula is as follows

[0105] PE (pos,2i) =sin(pos / 10000^(2i / d)) (1)

[0106] PE (pos,2i + 1) =cos(pos / 10000^(2i / d)) (2)

[0107] Where pos is the position number of the input feature sequence (starting from 0), and i is half of the dimension subscript, which is used to distinguish between even and odd dimensions. Assuming the length of the input feature sequence is L, the position encoding and embedding are added to obtain a sequence representation with position information:

[0108]

[0109] Step 1-2: Multi-head self-attention mechanism construction, first perform single-head self-attention calculation, and generate query vector (Q), key vector (K) and value vector (V) from the input sequence X of each encoder through linear transformation

[0110] Q=XW Q (4)

[0111] K=XW K (5)

[0112] V=XW V (6)

[0113] in d k Usually (h is the number of attention heads), then calculate the attention score and weighted sum

[0114]

[0115] Z i =softmax(scores)×V (7)

[0116] Assume that the input feature sequence L is a path loss sequence. When calculating the attention score scores of the first row of the path loss sequence, each element in the input feature needs to be scored for the path loss sequence. These scores are calculated by taking the dot product of the key vectors of all elements of the input sequence and the query vector of the first row of the path loss sequence. The role of the Softmax function is to normalize the scores of all elements of the feature sequence. The obtained scores are all positive and sum to 1. Finally, the result vector Z of the self-attention calculation is obtained. i Multi-head attention can divide the input into h heads, calculate the above self-attention separately, and then splice their results horizontally:

[0117] Z=concat(head1, head2,..., head h )W 0 (8)

[0118] Each head i =Z i ,

[0119] Steps 1-3: Residual connection (Add ResNet) and layer normalization (Layer Normalization, LN). After the input feature sequence passes through the multi-head attention mechanism to obtain the matrix Z, it is not directly passed to the fully connected neural network, but passes through the Add & Normalize layer. The expression is

[0120] LN(X+Z) (9)

[0121] Where X represents the input of the multi-head attention or feedforward network, and Z represents the output of the head attention or feedforward network;

[0122] Steps 1-4: Construct a feedforward network (FFN), expressed as

[0123] FFN(X)=max(0,XW1+b1)W2+b2 (10)

[0124] The weight matrix

[0125] Steps 1-5: Special processing in the decoder. In the multi-head self-attention at the decoder, a mask is required to prevent "seeing" future tokens (i.e., to ensure causality): the position t′>t is masked in the attention score, and the weight is set to 0 after the softmax.

[0126] scores[t, t′]=-∞ (11)

[0127] Taking the path loss feature sequence as an example, using a mask can block the path loss sequence results for the next moment when calculating the path loss sequence for the current moment. Subsequently, in the second attention sub-layer, the decoder interacts with the encoder output, helping the decoder "reference" the context of the original input sequence when generating predictions.

[0128] Steps 1-6: Construct the output layer. The final output layer of the decoder usually adds a linear mapping and softmax to output the next step of path loss sequence prediction. That is:

[0129] P(y t |y1,...,y t-1 , X)=softmax(ZW o +b o ) (12)

[0130] Where Z comes from the last layer output of the decoder.

[0131] Figure 2 Schematic diagram of the Transformer structure;

[0132] Step 2: Construct a dataset for the path loss prediction model, including a training dataset and a test dataset. First, based on the environmental characteristics of microwave and scattering communication links at sea, an optimal combination of vertical and horizontal stratification and microphysics schemes is set for a large area. Based on the WRF mesoscale numerical meteorological model, the environmental meteorological parameters on the communication link are inverted to construct a characteristic subset of meteorological parameters. Second, the NPS model is used to predict the evaporation duct height over a large area with high resolution and long duration, obtaining a characteristic subset of the evaporation duct height prediction values.

[0133] Figure 3 Schematic diagram for extracting meteorological parameters (taking sea surface temperature as an example).

[0134] Step 2-1: Use the WRF mesoscale numerical meteorological model, select appropriate horizontal and vertical layers to obtain large-area, high-resolution environmental information, select the optimal microphysics scheme, and set the time span and temporal resolution to invert the environmental meteorological parameters of the tropospheric beyond-horizon link area. For example, this method uses the following scheme: horizontal resolution of 27 km, vertical resolution of 60 layers, integration step of 180 seconds, RRTM scheme for longwave radiation, Dudhia scheme for shortwave radiation, YSU scheme for boundary layer, Noah scheme for land surface processes, and Kain-Fritsch scheme for cumulus parameterization;

[0135] Step 2-2: In post-processing, traverse each set moment and use the variable extraction function in wrf-python to extract all the inverted meteorological parameters at all grid points. According to the longitude and latitude coordinates of the two ends of the required link corresponding to the position in the grid, set the starting and ending points of the link, use the fitting function to generate the line pixel coordinates and remove duplicates, calculate the interpolation position of each point on the link, average the meteorological parameter values ​​at each point, and use them as the meteorological parameters for the link at the current time step to construct the tropospheric beyond-horizon link channel environment meteorological parameter data subset;

[0136] Step 2-3: The prediction of the evaporation duct height characteristic parameters needs to be calculated through the atmospheric correction refractive index. The atmospheric refractive index N is determined by the atmospheric pressure p (unit: hPa), the atmospheric temperature T (unit: K) and the water vapor partial pressure e (unit: hPa). The empirical relationship is:

[0137]

[0138] The calculation formula for water vapor partial pressure e is:

[0139]

[0140]

[0141] R h is the relative humidity of the atmosphere. In order to better study the effect of atmospheric refractive index on electromagnetic wave propagation, the earth's surface is approximately treated as a plane, and the atmospheric corrected refractive index M (unit M) is redefined. The relationship between it and the atmospheric refractive index is:

[0142]

[0143] Z is the altitude (unit: m), r e is the mean earth curvature, which is 6371 km. Substituting it into formula (15) yields

[0144] M=N+0.157z (17) In the NPS model, the vertical profile of temperature T and specific humidity q in the near-surface layer is expressed as

[0145]

[0146]

[0147] Where T0 and q0 are the sea surface temperature and specific humidity, T(z) and q(z) are the atmospheric temperature and specific humidity at the height z, respectively. * ,q * are the characteristic scales of potential temperature θ and specific humidity q, ψ h is the temperature universal function, κ is the Karman constant, Γ d is the dry adiabatic lapse rate, which is about 0.00976K / m, z 0t is the roughness height of atmospheric temperature, and L is the similarity length. The water vapor pressure profile, atmospheric temperature, and pressure are calculated using formulas (14), (18), and (19), and the results are then substituted into formulas (13) and (17) to obtain the atmospheric corrected refractive index profile. A subset of meteorological parameter data is input into the NPS model to obtain a high-resolution, long-term evaporation duct height forecast for a large area.

[0148] Step 3: Predict channel path loss using the two-dimensional parabolic equation method to obtain a feature subset of path loss prediction values. Meanwhile, the operating parameters of microwave and scattering communications are combined to form an operating parameter feature subset. Finally, the meteorological parameter feature subset, the evaporation duct height feature subset, and the operating parameter feature subset are combined and grouped in the time domain to form the X input feature. The path loss prediction value feature subset is then constructed as the Y input feature.

[0149] Step 3-1: The corresponding recursive formula of the two-dimensional parabolic equation method is

[0150]

[0151] in and are Fourier transform and inverse transform respectively, is the refractive index term, which reflects the influence of space medium on electromagnetic waves. is the diffraction term, reflecting the diffraction effect of obstacles on the propagation path on the electromagnetic wave. Here, p = k0sinθ is the angular spectrum domain variable, and θ is the angle between the electromagnetic wave and the horizontal. Given the initial field distribution u(x0,z), the next step of the field distribution u(x0+Δx,z) can be obtained, thereby iteratively solving the field in the entire computational space. For the current tropospheric beyond-horizon link environment, the range of system parameters for microwave scattering communication equipment, including frequency, transmitting antenna height, and receiving antenna height, is specified. Combined with meteorological parameters and the predicted evaporation duct height, the two-dimensional bidirectional step-by-step parabolic equation method is used to predict the path loss on this link.

[0152] Step 3-2: Construct the dataset required for path loss sequence model training. The feature dimension of the dataset is (number of samples, time steps, feature dimension), where the number of samples represents the total number of combinations of meteorological parameters and communication equipment system parameters, and the time step represents the total number of time steps under the current time span and time resolution. The dataset is divided into X set and Y set, where the features of the X dataset are air temperature (AT), sea surface temperature (SST), atmospheric pressure (AP), relative humidity (RH), wind speed (WS), evaporation duct height (EDH), communication frequency (f), and transmitting antenna height (h). t ), receiving antenna height (h r ), according to the system parameters of the communication equipment (f, h t 、h r ) is a sample, strictly ensuring that a sample contains all time step information. For example, if the time span is 72 hours and the time resolution is 1 hour, the total number of time steps is 72. The dimension of a sample in the X dataset is (72, 9), and the dimension of the X dataset is (number of samples, 72, 9). The Y dataset is characterized by a path loss (PL) sequence. Each row of PL corresponds to a combination of meteorological parameters and system parameters in a sample. Similarly, the continuity of time steps is guaranteed. After the arrangement is completed, the dataset is saved.

[0153] Step 4: Model training optimization and performance evaluation: First, optimize the model hyperparameters based on the tree-structured Parsons estimator algorithm; second, compare with the existing LSTM and GRU models to verify the feasibility and effectiveness of the proposed method for path loss prediction under offshore evaporation duct conditions in terms of mean absolute error, root mean square error, mean absolute percentage error, and relative error.

[0154] Step 4-1: To compare and evaluate the models, we built a Long Short Term Memory (LSTM) and a Gate Recurrent Unit (GRU) model for comparison.

[0155] Step 4-2: Data loading and preprocessing:

[0156] (1) Read the dataset through the loading function and obtain the training data in groups (divided into X_train and Y_train).

[0157] (2) For X_train and Y_train, flatten the original (number of samples, time steps, feature dimensions) data to (number of samples × time steps, feature dimensions), normalize it, and then reshape it back to its original shape, while saving the normalizer.

[0158] (3) Divide the data into 70% training sets and 30% validation sets, convert them into tensors and encapsulate them into iterable batch data using data loaders;

[0159] Step 4-3: Model definition:

[0160] (1) The LSTM model is defined, which consists of the following parts: the input gate, which determines to what extent the current input is written into the cell state; the forget gate, which determines how much past information the current cell state can retain and what parts need to be forgotten; the output gate, which determines the content of the hidden state output at the current moment; the cell state: an information channel that runs through the time series, can store information over a long time step, and selectively update or retain content under the gating mechanism;

[0161] (2) The GRU model is defined, which includes the following parts: the update gate, which combines the functions of the "input gate" and the "forget gate" to use a single gate to determine whether to retain past information and accept new information at the current moment; the reset gate, which helps the network decide how much historical information to discard, thereby more flexibly capturing short-term dependencies;

[0162] (3) In the feedforward method, the input is transposed (from (number of samples, time steps, feature dimensions) to (sequence length, number of samples, feature dimensions)), then passed through the embedding layer and Transformer calculation, and finally transposed back to the original shape before being sent to the output layer;

[0163] Step 4-4: Training and validation functions:

[0164] (2) Define a function to train a single training round (epoch) and return the training error (mean square error MSE). The mean square error is defined as

[0165]

[0166] Where N represents the total number of samples, y i is the target vector, is the prediction vector, ||·||2 represents the 2nd-order norm;

[0167] (2) Define a function to perform inference on the validation set and return the validation error (mean square error MSE);

[0168] Step 4-5: Hyperparameter Optimization (Bayesian Optimization):

[0169] (1) Define the hyperparameter search space (learning rate lr, embedding dimension d, attention head h, encoder layer number en, decoder layer number dn, feedforward network dimension d ff , dropout rate).

[0170] (2) Define the parameter optimization function: instantiate and train the model based on the hyperparameters of the current experiment and return the validation loss.

[0171] (3) Use the tree-structured Parsons estimator algorithm to perform Bayesian optimization to find the optimal hyperparameters.

[0172] (4) Record the test process to obtain the optimal hyperparameters and save them;

[0173] The hyperparameter combinations obtained in this method are: learning rate lr = 0.0021, embedding dimension d = 352, attention head h = 8, encoder layer number en = 2, decoder layer number dn = 1, feedforward network dimension d ff =192, dropout rate = 0.3971)

[0174] Steps 4-6: Retrain the path loss prediction model using the optimal hyperparameters:

[0175] (1) Save the optimal hyperparameter values ​​and instantiate the final Transformer model.

[0176] (2) Set the training parameters, loss function and optimizer (Adam).

[0177] (3) Perform training in a defined training and validation cycle, and set up an early stopping mechanism (stop if the validation loss does not improve within a certain number of rounds).

[0178] (4) Whenever a better validation loss is obtained, save the current model;

[0179] Steps 4-7: Load the best model and evaluate it:

[0180] (1) Load the saved best model.

[0181] (2) Use the validation set for predictive reasoning and collect and concatenate the output and target values.

[0182] (3) Perform inverse normalization on the predicted results and target results.

[0183] (4) Calculate and output the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and the prediction probability (P err5 ) and the prediction probability with relative error ≤ 1% (P err1) and other evaluation indicators and record them in the log. The evaluation indicator expression is as follows

[0184]

[0185]

[0186] Figure 4 The following is a comparison chart of path loss prediction results using different methods;

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for predicting path loss of tropospheric beyond-horizon communication channels based on Transformer network, characterized in that: The method comprises the following steps: Step 1: Considering the time-varying, environmental dynamic changes and long-term dependence on environmental meteorological parameters of the marine tropospheric beyond-horizon communication channel, a tropospheric beyond-horizon communication channel path loss prediction model based on the Transformer network is constructed, which includes an input layer, position encoding, encoder, decoder and output layer. The encoder and decoder are composed of an embedding layer and normalization, a self-attention layer and a fully connected layer. Step 2: Construct a data set for the path loss prediction model, including a training data set and a test data set: First, based on the environmental characteristics of microwave communication and scattering communication links at sea, set the optimal combination of vertical and horizontal stratification and microphysics in a large area, and invert the environmental meteorological parameters on the communication link based on the WRF mesoscale numerical meteorological model to construct a characteristic subset of meteorological parameters; secondly, use the NPS model to predict the evaporation duct height in a large area with high resolution and long duration, and obtain the characteristic subset of the evaporation duct height prediction value; Step 3: Based on the two-dimensional parabolic equation method, the channel path loss is predicted to obtain the feature subset of the path loss prediction value, and the working parameters of microwave and scattering communication are used to form the working parameter feature subset; finally, the meteorological parameter feature subset, the evaporation waveguide height feature subset and the working parameter feature subset are combined and grouped in the time domain to construct the X input feature, while the path loss prediction value feature subset is constructed as the Y input feature; Step 4: Model training optimization and performance evaluation: First, optimize the model's hyperparameters based on the Tree-structured Parzen Estimator Approach (TPE) algorithm; second, compare with the existing LSTM and GRU models to verify the feasibility and effectiveness of the proposed method for path loss prediction under offshore evaporation duct conditions in terms of mean absolute error, root mean square error, mean absolute percentage error and relative error.

2. The method for predicting path loss of tropospheric beyond-horizon communication channels based on Transformer network according to claim 1 is characterized in that: The step 1 is specifically as follows: Step 1-1: Prepare the embedding and position encoding of the input features, find the corresponding vector embedding for each input token, assuming that the embedding dimension is d; in order for the model to perceive the sequence of environmental meteorological parameters and evaporation duct height, as well as the order of each position in the corresponding path loss sequence, positional encoding needs to be added. The common position encoding method is to use sine and cosine functions. The formula is as follows: ON (pos,2i) =sin(pos / 10000^(2i / d)) (1) ON (pos,2i+1) =cos(pos / 10000^(2i / d)) (2) Where pos is the position number of the input feature sequence, starting from 0, and i is half of the dimension subscript, which is used to distinguish even dimensions from odd dimensions. Assuming the length of the input feature sequence is L, the position encoding and embedding are added to obtain a sequence representation with position information: Step 1-2: Multi-head self-attention mechanism construction, first perform single-head self-attention calculation, and generate query vector Q, key vector K and value vector V from the input sequence X of each encoder through linear transformation: Q=XW Q (4) K=XW K (5) V=XW V (6) in d k for h is the number of attention heads, and then the attention scores are calculated and weighted summed Z i =softmax(scores)×V (7) Assume that the input feature sequence L is a path loss sequence. When calculating the attention score scores of the first row of the path loss sequence, each element in the input feature needs to be scored for the path loss sequence. These scores are calculated by dot product of the key vectors of all elements of the input sequence and the query vector of the first row of the path loss sequence. The role of the Softmax function is to normalize the scores of all elements of the feature sequence. The scores obtained are all positive and the sum is 1. Finally, the result vector Z of the self-attention calculation is obtained. i ; Multi-head attention divides the input into h heads, calculates the above self-attention separately, and then concatenates their results horizontally: Z=concat(head1,head2,…,head h )W 0 (8) Each head i =Z i , Steps 1-3: Residual connection Add ResNet and layer normalization (Layer Normalization, LN). After the input feature sequence is passed through the multi-head attention mechanism to obtain the matrix Z, it is not directly passed into the fully connected neural network, but passes through the Add&Normalize layer. The expression is: LN(X+Z) (9) Where X represents the input of the multi-head attention or feedforward network, and Z represents the output of the head attention or feedforward network; Step 1-4: Construct a feedforward network (FFN), the expression is: FFN(X)=max(0,XW1+b1)W2+b2 (10) The weight matrix Steps 1-5: Special processing in the decoder. In the multi-head self-attention at the decoder end, a mask is needed to prevent "seeing" future tokens, that is, to ensure causality: the position t'>t is masked in the attention score, and the weight is set to 0 after softmax. scores[t,t]=-∞ (11) Taking the path loss feature sequence as an example, a mask is used to mask the result of the path loss sequence at the next moment when calculating the path loss sequence at the current moment; thereafter, in the second attention sublayer, the decoder interacts with the output of the encoder to help the decoder "reference" the context information of the original input sequence when generating predictions; Steps 1-6: Construct the output layer. A linear mapping and softmax will be added to the final output layer of the decoder to output the next step of path loss sequence prediction; that is: P(y t |y1,…,y t-1 ,X)=softmax(ZW o +b o ) (12) Where Z comes from the last layer output of the decoder.

3. The method for predicting path loss of tropospheric beyond-horizon communication channels based on Transformer network according to claim 1, characterized in that: The steps 2 and 3 are specifically as follows: Step 2-1: Use the WRF mesoscale numerical meteorological model, select appropriate horizontal and vertical layers to obtain large-area high-resolution environmental information, select the optimal microphysical scheme, set the time span and time resolution, and invert the environmental meteorological parameters of the tropospheric beyond-horizon link area; Step 2-2: In post-processing, traverse each set moment, and use the variable extraction function in wrf-python to extract all the meteorological parameters inverted at all grid points; set the start and end points of the link according to the longitude and latitude coordinates of the two ends of the required link corresponding to the position in the grid, use the fitting function to generate the line pixel coordinates and remove the duplicates, calculate the interpolation position of each point on the link, average the meteorological parameter values ​​at each point, and use them as the meteorological parameters on the link at the current time step to construct the tropospheric beyond-horizon link channel environment meteorological parameter data subset; Step 2-3: The prediction of the characteristic parameters of the evaporation duct height needs to be calculated through the atmospheric correction refractive index. The atmospheric refractive index N is determined by the atmospheric pressure p, in hPa, the atmospheric temperature T, in K, and the water vapor partial pressure e, in hPa. The empirical relationship is: The calculation formula of water vapor partial pressure e is: R h is the relative humidity of the atmosphere. In order to better study the effect of atmospheric refractive index on electromagnetic wave propagation, the earth's surface is approximately treated as a plane, and the atmospheric corrected refractive index M (unit M) is redefined. The relationship between it and the atmospheric refractive index is: Z is the altitude, in meters, r e is the mean earth curvature, which is 6371 km. Substituting it into formula (15), we get: M=N+0.157z (17) In the NPS model, the vertical profile of temperature T and specific humidity q in the near-surface layer is expressed as: Where T0 and q0 are the sea surface temperature and specific humidity, T(z) and q(z) are the atmospheric temperature and specific humidity at the height z, θ * ,q * are the characteristic scales of potential temperature θ and specific humidity q, ψ h is the temperature universal function, κ is the Karman constant, Γ d is the dry adiabatic lapse rate, about 0.00976K / m, z 0t is the roughness height of atmospheric temperature, L is the similarity length; the water vapor pressure profile, atmospheric temperature and pressure are calculated by formula (14), formula (18) and formula (19), and the results are then substituted into formula (13) and formula (17) to obtain the atmospheric corrected refractive index profile; the meteorological parameter data subset is input into the NPS model to obtain a large-area high-resolution and long-term evaporation duct height forecast; Step 3-1: The corresponding recursive formula of the two-dimensional parabolic equation method is in and are Fourier transform and inverse transform, respectively. is the refractive index term, which reflects the influence of space medium on electromagnetic waves. is the diffraction term, which reflects the diffraction effect of obstacles on the propagation path on the electromagnetic wave, where p = k0sinθ is the variable in the angular spectrum domain, θ is the angle between the electromagnetic wave and the horizontal direction, and when the initial field distribution u(x0,z) is given, the field distribution u(x0+Δx,z) at the next step is obtained, thereby iteratively solving the field in the entire calculation space; for the current tropospheric beyond-horizon link environment, the range of system parameters of microwave scattering communication equipment is specified, including frequency, transmitting antenna height and receiving antenna height, combined with meteorological parameters and the predicted evaporation waveguide height, and the two-dimensional bidirectional step-by-step parabolic equation method is used to predict the path loss on the link; Step 3-2: Construct the data set required for path loss sequence model training. The feature dimension of the data set is (number of samples, time step, feature dimension), where the number of samples represents the total number of combinations of meteorological parameters and communication equipment system parameters, and the time step represents the total number of time steps under the current time span and time resolution; the data set is grouped into X set and Y set, where the features of the X data set are air temperature (AT), sea surface temperature (SST), atmospheric pressure (AP), relative humidity (RH), wind speed (WS), evaporation duct height (EDH), communication frequency (f), transmitting antenna height (h t ), receiving antenna height (h r ), according to the system parameters of the communication equipment (f, h t 、h r ) is a sample, and it is strictly guaranteed that a sample contains all the time step information. For example, if the time span is 72 hours and the time resolution is 1 hour, the total number of time steps is 72. The dimension of a sample in the X dataset is (72,9), and the dimension of the X dataset is (number of samples, 72, 9). The feature of the Y dataset is the path loss (PL) sequence. The PL of each row corresponds to a combination of meteorological parameters and system parameters in a sample. The continuity of the time step is also guaranteed. After the arrangement is completed, the dataset is saved.

4. The method for predicting path loss of tropospheric beyond-horizon communication channels based on Transformer network according to claim 3 is characterized in that: The step 4 is specifically as follows: Step 4-1: In order to compare and evaluate the models, the Long Short Term (LSTM) and Gate Recurrent Unit (GRU) models were constructed for comparison; Step 4-2: Data loading and preprocessing: (1) Read the data set through the loading function and obtain the training data in groups, which are divided into (X_train, Y_train); (2) For X_train and Y_train, flatten the original (number of samples, time steps, feature dimensions) data to (number of samples × time steps, feature dimensions), normalize it, and then reshape it back to the original shape, while saving the normalizer; (3) Divide the data into 70% training set and 30% validation set, convert them into tensors and encapsulate them into iterable batch data using data loaders; Step 4-3: Model definition: (1) The LSTM model is defined, which includes the following parts: the input gate, which determines to what extent the current input is written into the cell state; the forget gate, which determines how much past information the current cell state can retain and which parts need to be forgotten; Output gate, which determines the content of the hidden state output at the current moment; Cell state: an information channel that runs through the time series, capable of storing information over a long time step, and selectively updating or retaining content under a gating mechanism; (2) The GRU model is defined, which includes the following parts: the update gate, which combines the functions of the "input gate" and the "forget gate" to use a gate to determine whether to retain past information and accept new information at the current moment; the reset gate, which helps the network decide how much historical information to discard, thereby more flexibly capturing short-term dependencies; (3) In the feedforward method, the input is transposed from (number of samples, time steps, feature dimension) to (sequence length, number of samples, feature dimension), then passed through the embedding layer and Transformer calculation, and finally transposed back to the original shape and sent to the output layer; Step 4-4: Training and validation functions: (1) Define a function to train a single training round and return the training error, i.e., the mean square error (MSE). The mean square error is defined as Where N represents the total number of samples, y i is the target vector, is the prediction vector, ||·||2 represents the 2nd-order norm; (2) Define a function to perform inference on the validation set and return the validation error, i.e., mean square error (MSE); Step 4-5: Hyperparameter optimization: (1) Define the hyperparameter search space, including learning rate lr, embedding dimension d, attention head h, encoder layer number en, decoder layer number dn, feedforward network dimension d ff , dropout rate; (2) Define the parameter optimization function: instantiate and train the model according to the hyperparameters of the current experiment and return the validation loss; (3) Use the tree-structured Parsons estimator algorithm for Bayesian optimization to find the optimal hyperparameters; (4) Record the test process to obtain the optimal hyperparameters and save them; Steps 4-6: Retrain the path loss prediction model using optimal hyperparameters: (1) Save the optimal hyperparameter values ​​and instantiate the final Transformer model; (2) Set training parameters, loss function, and optimizer; (3) Perform training in a defined training and validation cycle, and set an early stopping mechanism to stop if the validation loss does not improve within a certain number of rounds; (4) Whenever a better validation loss is obtained, save the current model; Steps 4-7: Load the best model and evaluate it: (1) Load the saved best model; (2) Use the validation set for predictive reasoning and collect and concatenate the output and target values; (3) Perform inverse normalization on the predicted results and target results; (4) Calculate and output the mean absolute error (MAE), root mean square error (RMSE), mean absolute percentage error (MAPE), and prediction probability (P) with a relative error ≤ 5% err5 and the prediction probability P with relative error ≤ 1% err1 The evaluation index is recorded in the log. The evaluation index expression is as follows:

Citation Information

Cited By

  • Waveguide / scattering mode discrimination method based on evaporation waveguide height

    CN120342436A

  • Urban land subsidence prediction method based on improved Transform model

    CN120724159A