Channel knowledge map construction method based on improved transformer and improved transformer system

By improving the Transformer to construct a channel knowledge map, removing the encoder and residual connections, and designing a dynamic normalization module, the problems of high computational overhead and low accuracy of existing methods are solved, and efficient and accurate path loss prediction is achieved.

CN121235045BActive Publication Date: 2026-03-17HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511785278.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-17
Estimated Expiration
2045-12-01

AI Technical Summary

Technical Problem

Existing methods for constructing channel knowledge maps employ a joint structure of a CNN backbone network and a Transformer encoder, resulting in heavy computational overhead and impacting the accuracy of path loss prediction in critical areas.

Method used

An improved channel knowledge map construction method for Transformer is adopted. By removing the encoder and residual connection and designing a dynamic normalization module, a pure decoder model is constructed. Path loss prediction is performed by combining a masked multi-head self-attention mechanism and a feedforward network.

Benefits of technology

It reduces computational complexity and the number of model parameters, improves the generation accuracy and generalization ability of the channel knowledge map, enhances the robustness of the model, reduces the risk of overfitting, and improves the accuracy of path loss prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121235045B_ABST
    Figure CN121235045B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of wireless communication, in particular to a channel knowledge map construction method based on an improved Transform and an improved Transform system. The channel knowledge map construction method comprises the following steps: acquiring an overhead snapshot of a target communication area and encoding the overhead snapshot into an n*n binary environment coding matrix E, wherein a matrix element '1' represents that a corresponding subarea has a building, and a matrix element '0' represents that the corresponding subarea has no building; inputting the binary environment coding matrix E into a pre-trained improved Transform neural network model, wherein the improved Transform neural network model is a pure decoder model constructed based on a Transform decoder architecture, and a standard Transform encoder part and a multi-head encoder-decoder attention layer are removed; and outputting a corresponding path loss prediction matrix H through the improved Transform neural network model to complete the channel knowledge map construction of the target communication area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and more specifically to a method for constructing a channel knowledge map based on an improved Transformer and an improved Transformer system. Background Technology

[0002] The sixth-generation (6G) wireless communication system aims to achieve ultra-wideband, low latency, and massive connectivity, supporting key applications such as digital twins, sensory integration, and autonomous driving. Channel characteristics directly determine the performance of core technologies such as beamforming and resource allocation. Especially in higher frequency bands (millimeter wave, terahertz), signal attenuation and environmental sensitivity intensify, making accurate channel modeling a prerequisite for 6G system development. Among numerous channel modeling methods, channel knowledge maps (CKMs) are one feasible approach. By quantifying the mapping relationship between the spatial environment and channel parameters (such as path loss), high-quality channel knowledge maps can effectively reduce the pilot overhead required for traditional channel estimation and improve network efficiency. Therefore, the construction of CKMs has become one of the important research directions in this field.

[0003] However, existing CKM construction methods have inherent bottlenecks: analytical methods, such as the Free Space Path Loss (FSPL) model, rely solely on distance and frequency for calculation, completely ignoring environmental features like buildings and terrain, leading to significant prediction errors in complex scenes. While methods based on Convolutional Neural Networks (CNNs) excel at extracting local spatial features, thus surpassing FSPL models in accuracy, their local receptive fields struggle to capture long-range spatial relationships between scattered obstacles, and their filling mechanisms cause prediction artifacts in edge regions, making it difficult to achieve globally high-precision CKM construction.

[0004] The Transformer architecture possesses the ability to capture long-range dependencies, thus holding promise for improving the construction accuracy of CKM (Path Loss Mechanism). A CKM construction method employing a joint structure of a CNN backbone and a Transformer encoder utilizes CNN to extract local environmental features and Transformer to model global correlations. However, it still requires spatial feature capture, inherently inheriting the inherent conflict between the position-sensitive nature of CNNs and CKM, leading to heavy computational overhead and impacting the path loss prediction accuracy in key regions. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a channel knowledge map construction method and an improved Transformer system based on an improved Transformer. The aim is to solve the problem that existing CKM construction methods use a joint structure of CNN backbone network and Transformer encoder, which inherits the conflict between the position-sensitive characteristics of CNN and CKM, resulting in heavy computational overhead and affecting the path loss prediction accuracy in key areas.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A method for constructing a channel knowledge map based on an improved Transformer includes the following steps:

[0008] Obtain a top-down snapshot of the target communication area and encode the top-down snapshot into an n×n binary environment coding matrix E, where matrix element '1' indicates that there are buildings in the corresponding sub-region and matrix element '0' indicates that there are no buildings in the corresponding sub-region;

[0009] The binary environment encoding matrix E is input into a pre-trained improved Transformer neural network model, wherein the improved Transformer neural network model is a pure decoder model built on the Transformer decoder architecture, and the standard Transformer encoder part and multi-head encoder-decoder attention layer are removed.

[0010] The improved Transformer neural network model is used to process the data and output a corresponding path loss prediction matrix H to complete the construction of the channel knowledge map for the target communication area.

[0011] Furthermore, the improved Transformer decoder architecture consists of multiple stacked decoding layers with identical structures, and the data processing flow within each decoding layer includes:

[0012] The first processing stage – attention and dynamic normalization – includes:

[0013] The input data first enters the multi-head attention mechanism module, which is used to capture long-range spatial dependencies between different locations in the environment;

[0014] The output of the multi-head attention mechanism is residually concatenated with the original input.

[0015] The results of the residual connection are sent to the dynamic normalization module for processing;

[0016] The second processing stage – forward feedback and normalization – includes:

[0017] The dynamically normalized data is passed to the feedforward network module, which performs the operation FFN(x)=ReLU(xW1+b1)W2+b2 to perform nonlinear feature transformation, where x is the input tensor, W1 and W2 are weight matrices, and b1 and b2 are bias vectors.

[0018] The residual connections of the feedforward network module were removed to enhance the model's flexibility in processing features of different dimensions.

[0019] The output of the feedforward network is processed by a normalization module, which retains traditional layer normalization to ensure the stability of the output features.

[0020] Furthermore, the dynamic normalization module is a dynamic hyperbolic tangent layer, whose operation function is: DyT(x) = γ·tanh(αx) + β, where x is the input tensor, α is a trainable scaling factor, and γ and β are learnable channel-wise vector parameters.

[0021] Furthermore, the data processing flow within the improved Transformer decoder architecture also includes:

[0022] The third processing stage – output range limitation – includes:

[0023] The normalized data enters the output range limiting module, which restricts the predicted path loss value to a preset physical range. Values ​​outside the range will be forcibly set as boundary values.

[0024] Furthermore, the preset physical range is [0, 250].

[0025] An improved Transformer model for performing the above-described channel knowledge map construction method includes:

[0026] Masked multi-head self-attention mechanism is used to model long-range spatial dependencies in input data;

[0027] The dynamic normalization module normalizes the output of the mask multi-head self-attention mechanism.

[0028] The feedforward network module is used to perform the operation FFN(x)=ReLU(xW1+ b1)W2+b2 to perform nonlinear feature transformation, where x is the input tensor, W1 and W2 are weight matrices, b1 and b2 are bias vectors, and the feedforward network module removes residual connections.

[0029] The layer normalization module is used to process the output of the feedforward network module.

[0030] A method for calculating the accuracy of channel knowledge map construction, whereby the accuracy is characterized by mean absolute error (MAE):

[0031] ;

[0032] Where N represents the total number of samples, H represents the height of each two-dimensional CKM matrix (i.e., the number of rows), and W represents the width of each two-dimensional CKM matrix (i.e., the number of columns). This represents the actual path loss value at position (h, w) in the i-th sample. Let |·| represent the predicted path loss value at position (h, w) in the i-th sample, where |·| represents the absolute value operator.

[0033] The channel knowledge map construction method and improved Transformer system described in this invention have the following advantages:

[0034] A top-down snapshot of the target communication area is obtained and encoded into an n×n binary environment coding matrix E, where a matrix element '1' indicates the presence of buildings in the corresponding sub-region and a matrix element '0' indicates the absence of buildings in the corresponding sub-region. This effectively captures building information in complex real-world communication environments, while reducing computational complexity and preserving key building features. It also highlights the core correlation between building distribution and path loss, which is the channel knowledge map referred to in this invention.

[0035] Meanwhile, improvements such as removing the encoder, removing residual connections, and designing dynamic normalization eliminate the entire encoder computation branch and the cross-attention computation between the encoder and the encoder. The number of model parameters and computational overhead are greatly reduced. The simplified model structure reduces the risk of overfitting, enhances the model's generalization ability and structural robustness, and has higher channel knowledge map generation accuracy. Attached Figure Description

[0036] Figure 1 This is a flowchart of the channel knowledge map construction method based on the improved Transformer of the present invention;

[0037] Figure 2 This is a top-view snapshot of the target communication area of ​​the present invention;

[0038] Figure 3 This is a flowchart of the neural network training process based on the improved Transformer of the present invention;

[0039] Figure 4 This is a comparison diagram of the effects of the channel knowledge map construction method based on the improved Transformer of the present invention with other traditional channel knowledge map construction methods. Detailed Implementation

[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0041] like Figures 1 to 2 As shown, this invention provides a method for constructing a channel knowledge map based on an improved Transformer, including the following steps:

[0042] Obtain a top-down snapshot of the target communication area and encode the top-down snapshot into an n×n binary environment coding matrix E, where matrix element '1' indicates that there are buildings in the corresponding sub-region and matrix element '0' indicates that there are no buildings in the corresponding sub-region;

[0043] The binary environment encoding matrix E is input into a pre-trained improved Transformer neural network model, which is a pure decoder model built on the Transformer decoder architecture and removes the standard Transformer encoder part and the multi-head encoder-decoder attention layer.

[0044] By improving the Transformer neural network model, a corresponding path loss prediction matrix H is output to complete the construction of the channel knowledge map of the target communication area.

[0045] A top-down snapshot of the target communication area is obtained and encoded into an n×n binary environment coding matrix E, where a matrix element '1' indicates the presence of buildings in the corresponding sub-region and a matrix element '0' indicates the absence of buildings in the corresponding sub-region. This effectively captures building information in complex real-world communication environments, while reducing computational complexity and preserving key building features. It also highlights the core correlation between building distribution and path loss, which is the channel knowledge map referred to in this invention.

[0046] Meanwhile, improvements such as removing the encoder, removing residual connections, and designing dynamic normalization eliminate the entire encoder computation branch and the cross-attention computation between the encoder and the encoder. The number of model parameters and computational overhead are greatly reduced. The simplified model structure reduces the risk of overfitting, enhances the model's generalization ability and structural robustness, and has higher channel knowledge map generation accuracy.

[0047] Considering an outdoor communication environment, the transmission distance is between tens and hundreds of meters. The communication system includes a communication base station, which can be assumed to be located at the center of the reference coordinate system. With the communication base station as the center, the effective communication area is assumed to be a square with an area of... Where L is the side length of the region. Note that the actual communication area is not regular in shape and can be defined by using the area S or L as the circumscribed square of the actual communication area. Within the communication area, cameras can be deployed to obtain quick top-view images, such as... Figure 1 As shown, a 3D example is illustrated: gray cuboids represent buildings, red pentagrams represent transmitters, and blue dots represent receivers (not all receivers are shown to avoid clutter). Thick dashed lines indicate sub-region boundaries, and thin dashed lines indicate the boundaries of smaller regions. Using the environment coding method proposed in this application, the environment can be transformed into the following matrix:

[0048] ;

[0049] In some embodiments, matrix E is considered to generate a path loss dataset for inputting an improved Transformer architecture. The dataset can be generated using tools such as WirelessInsite, where the input environment encoding matrix E generates a building distribution, and ray tracing methods are used to generate a dataset such as... Figure 3 The path loss at each grid point is shown, which is then used to generate the CKM matrix data H.

[0050] The improved Transformer architecture described in this embodiment can also be used for measured channel path loss datasets. By randomly shuffling and splitting the CKM matrix dataset, training, validation, and test sets can be formed. The ratio of the three datasets can be set to 7:1.5:1.5. After obtaining the complete dataset, the training set of the CKM matrix is ​​input into the improved Transformer architecture for neural network training to obtain the implicit neural network parameters and the mapping relationship H=f(E).

[0051] like Figure 3 As shown, the improved Transformer decoder architecture consists of multiple identical decoding layers stacked together. The data processing flow within each decoding layer includes:

[0052] The first processing stage – attention and dynamic normalization – includes:

[0053] The input data first enters the multi-head attention mechanism module, which is used to capture long-range spatial dependencies between different locations in the environment;

[0054] The output of the multi-head attention mechanism is residually concatenated with the original input.

[0055] The results of the residual connection are sent to the dynamic normalization module for processing;

[0056] The second processing stage – forward feedback and normalization – includes:

[0057] The dynamically normalized data is passed to the feedforward network module, which performs the operation FFN(x) = ReLU(xW1+ b1)W2+b2 to perform nonlinear feature transformation, where x is the input tensor, W1 and W2 are weight matrices, and b1 and b2 are bias vectors.

[0058] The residual connections of the feedforward network module were removed to enhance the model's flexibility in handling features of different dimensions.

[0059] The output of the feedforward network is processed by a normalization module, which retains traditional layer normalization to ensure the stability of the output features.

[0060] By preserving residual connections in the multi-head attention mechanism, the effective propagation of long-range spatial dependencies is ensured, feature reuse rate is improved, the gradient vanishing problem in deep networks is alleviated, training stability is enhanced, and accuracy is improved when capturing the correlation between scattered obstacles.

[0061] By removing residual connections in the FFN stage and adding more neurons to enhance the model's ability to capture features of complex environments, the dimensionality matching limitation is broken, and flexible feature dimension transformation is supported. This improves the model's adaptability to inputs at different resolutions, increases the capacity of feature transformation, and enhances the fitting ability in complex environments. This allows the model to adapt more flexibly and effectively to environmental encoding inputs at different resolutions and complex feature transformation requirements, thereby improving the prediction accuracy of specific path loss modeling tasks and enhancing the adaptability of constructing CKMs at different scales.

[0062] Furthermore, the dynamic normalization module is a dynamic hyperbolic tangent layer, whose operation function is: DyT(x)=γ·tanh(αx)+β, where x is the input tensor, α is a trainable scaling factor, and γ and β are learnable channel-wise vector parameters.

[0063] Unlike fixed layer normalization, DyT dynamically adjusts the distribution of different feature channels through a trainable scaling factor α and channel-wise parameters γ and β. This makes it particularly effective at adapting to the sparse and discontinuous building distribution characteristics in channel environments, thereby improving the prediction accuracy of path loss. It reduces prediction artifacts in edge regions, eliminates mean and variance calculations, lowers normalization computational overhead, and utilizes the saturation properties of the tanh function to provide natural gradient clipping. This improves training stability, reduces memory usage, and supports larger batch training. Since DyT layers do not require calculation of the mean and variance of input features, the computation process is simpler than traditional layer normalization, further improving the model's computational efficiency.

[0064] Furthermore, improvements to the data processing flow within the Transformer decoder architecture also include:

[0065] The third processing stage – output range limitation – includes:

[0066] The normalized data enters the output range limiting module, which restricts the predicted path loss value to a preset physical range. Values ​​outside the range will be forcibly set as boundary values.

[0067] Furthermore, the preset physical range is [0, 250].

[0068] Path loss is a value with a defined physical range. This technique constrains the model's output within a reasonable physical range (e.g., [0, 250] dB), effectively avoiding unrealistic outliers in the model output. This greatly improves the practicality and reliability of the constructed channel knowledge map, providing a reliable data foundation for subsequent network planning.

[0069] After the decoder output is processed by linear transformation and the Softmax activation function, the probability distribution of the target label is finally generated.

[0070] By inputting the environment encoding matrix E and the corresponding CKM matrix training dataset into the improved Transformer architecture described above, the neural network can be trained to capture the core features of environmental factors affecting path loss, i.e., to train the implicit expression of H=f(E). By inputting the environment encoding matrix corresponding to the test set into the trained neural network, the corresponding CKM prediction matrix H can be obtained by substituting it into the trained f, thereby realizing the construction of the channel knowledge map.

[0071] This invention also provides an improved Transformer model for performing the above-described channel knowledge map construction method, comprising:

[0072] Masked multi-head self-attention mechanism is used to model long-range spatial dependencies in input data;

[0073] The dynamic normalization module normalizes the output of the mask multi-head self-attention mechanism;

[0074] The feedforward network module is used to perform the operation FFN(x)=ReLU(xW1+ b1)W2+ b2 to perform nonlinear feature transformation, where x is the input tensor, W1 and W2 are weight matrices, and b1 and b2 are bias vectors. The feedforward network module removes residual connections.

[0075] The layer normalization module is used to process the output of the feedforward network module.

[0076] By removing the encoder, removing residual connections, and designing dynamic normalization, the computational branches of the entire encoder and the cross-attention computation between the encoder are eliminated. The number of model parameters and computational overhead are greatly reduced. The simplified model structure reduces the risk of overfitting, enhances the model's generalization ability and structural robustness, and has higher accuracy in generating channel knowledge maps.

[0077] This invention also provides a method for calculating the accuracy of channel knowledge map construction, wherein the construction accuracy is characterized by the mean absolute error (MAE).

[0078] ;

[0079] Where N represents the total number of samples, H represents the height of each two-dimensional CKM matrix (i.e., the number of rows), and W represents the width of each two-dimensional CKM matrix (i.e., the number of columns). This represents the actual path loss value at position (h, w) in the i-th sample. Let |·| represent the predicted path loss value at position (h, w) in the i-th sample, where |·| represents the absolute value operator.

[0080] Channel knowledge maps were constructed on 1500 test samples using the channel knowledge map construction method of this application, CKM generation based on free space path loss, CKM generation based on convolutional neural network (CNN), CKM generation based on RMTransformer[2], and CKM generation scheme based on traditional Transformer. Then, the channel knowledge map construction accuracy was evaluated using the channel knowledge map construction accuracy calculation method. Figure 4 The mean absolute error distribution plot shown is composed of... Figure 4 It is evident that the channel knowledge map construction method of this application has the lowest MAE.

[0081] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. An improved Transformer-based channel knowledge map construction method, characterized in that, The method comprises the following steps: obtaining an overhead snapshot of a target communication area, and encoding the overhead snapshot into an n*n binary environment coding matrix E, wherein a matrix element '1' represents that a corresponding sub-area has a building, and a matrix element '0' represents that a corresponding sub-area has no building; inputting the binary environment coding matrix E into a pre-trained improved Transformer neural network model, wherein the improved Transformer neural network model is a pure decoder model based on a Transformer decoder architecture, and a standard Transformer encoder part and a multi-head encoder-decoder attention layer are removed; processing by the improved Transformer neural network model to output a corresponding path loss prediction matrix H, so as to complete the channel knowledge map construction of the target communication area; the improved Transformer decoder architecture is stacked by a plurality of decoding layers with the same structure, and the data processing flow in the decoding layer includes: the first processing stage includes attention and dynamic normalization, which comprises: the input data first enters a multi-head attention mechanism module for capturing long-range spatial dependence relationships in the environment; the output of the multi-head attention mechanism is connected in residual connection with the original input; the result of the residual connection is sent to a dynamic normalization module for processing; the second processing stage includes forward feedback and normalization, which comprises: the data after dynamic normalization is transmitted to a forward feedback network module, and the forward feedback network module performs FFN(x) = ReLU(xW1+b1)W2+b2 operation for nonlinear feature transformation, wherein x is an input tensor, W1 and W2 are weight matrices, and b1 and b2 are bias vectors; the residual connection of the forward feedback network module is removed to enhance the flexibility of the model in processing different dimensional features; the output of the forward feedback network is processed through a normalization module, and traditional layer normalization is reserved here to ensure the stability of the output features; the dynamic normalization module is a dynamic hyperbolic tangent layer, and its operation function is DyT(x) = γ·tanh(αx)+β, wherein x is an input tensor, α is a trainable scaling factor, and γ and β are learnable channel-wise vector parameters.

2. The method of claim 1, wherein, the data processing flow in the improved Transformer decoder architecture further includes: the third processing stage includes output range limitation, which comprises: the normalized data enters an output range limitation module, and the range limitation module limits the predicted value of the path loss in a preset physical range, and the value exceeding the range is forced to be set as a boundary value. 3.The method of claim 2, wherein, the preset physical range is [0, 250].

4. The method of claim 1 to 3, wherein, the channel knowledge map construction accuracy is represented by mean absolute error MAE, which comprises: ; where N denotes the total number of samples, H represents the height of each two-dimensional CKM matrix, i.e., the number of matrix rows, and W represents the width of each two-dimensional CKM matrix, i.e., the number of matrix columns, represents the real path loss value at position (h, w) in the i-th sample, represents the predicted path loss value at position (h, w) in the i-th sample, and | · | represents the absolute value operator.

5. An improved Transformer system performing the channel knowledge map construction method according to any one of claims 1 to 3, characterized in that, mask multi-head self-attention mechanism for modeling long-range spatial dependence relationships of input data; a dynamic normalization module for normalizing the output of the mask multi-head self-attention mechanism; ​ a forward feedback network module configured to perform an FFN(x) = ReLU(xW1+ b1)W2+ b2 operation for nonlinear feature transformation, where x is an input tensor, W1 and W2 are weight matrices, b1 and b2 are bias vectors, and the forward feedback network module removes a residual connection; a layer normalization module configured to process an output of the forward feedback network module.

Citation Information

Patent Citations

  • Crack image segmentation method based on double encoders in complex environment

    CN117058382A

  • DoH detection method based on self-attention BiLSTM

    CN118174956A