Environment and channel state information joint reconstruction method and system

By employing a joint reconstruction method of environment and channel state information, and utilizing a conditional variational autoencoder network model, a high-resolution channel fingerprint is reconstructed from a low-resolution channel fingerprint. This solves the problems of high pilot overhead and estimation complexity in channel state information acquisition methods, and achieves high-precision joint reconstruction of channel state information and environment information, thereby improving the transmission efficiency of wireless communication.

CN121968187APending Publication Date: 2026-05-01SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for obtaining channel state information in large-scale antenna arrays face challenges such as high pilot overhead and estimation complexity. How to recover a high-resolution channel fingerprint from a low-resolution channel fingerprint while ensuring accuracy has become a significant engineering challenge.

Method used

A joint reconstruction method using environment and channel state information is adopted. Through a conditional variational autoencoder network model, a low-resolution channel fingerprint is used as a conditional input to reconstruct a high-resolution channel fingerprint. The coupling between environment information and channel state information is defined using a binary representation method to construct a conditional variational autoencoder network, thereby achieving end-to-end training and obtaining optimal parameters.

Benefits of technology

It significantly reduces the pilot overhead for channel state information measurement and the measurement overhead for environment awareness, and achieves joint reconstruction of high-precision channel state information and environment information, thereby improving the transmission efficiency of wireless communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121968187A_ABST
    Figure CN121968187A_ABST
Patent Text Reader

Abstract

The invention discloses an environment and channel state information joint reconstruction method and system. Wireless channel measurement data and environment information are collected in a target area through downlink detection and perception, the environment information in the target area is defined through a binary representation method and serves as a mask to be coupled with channel state information, and a novel channel fingerprint coupling the environment information and the channel state information at the same time is obtained; by constructing a conditional variation auto-encoder network model, a high-resolution channel fingerprint is used as a reconstruction target, a low-resolution channel fingerprint is used as condition input, the high-resolution channel fingerprint is mapped to a potential space through an encoder, and the high-resolution channel fingerprint is reconstructed through a decoder under the given low-resolution channel fingerprint condition. According to the invention, on the premise of ensuring high fidelity of channel state information, accurate environment perception is realized, pilot frequency overhead and perception measurement overhead in an actual system are effectively reduced, and the method has good scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for joint reconstruction of environment and channel state information Technical Field

[0001] This invention belongs to the field of communication technology and relates to a method and system for joint reconstruction of environment and channel state information. Background Technology

[0002] With the advent of the information age, sixth-generation (6G) mobile communication systems widely adopt ultra-large-scale multiple-input multiple-output (MIMO) technology to achieve intelligent ubiquitous networks and massive device connectivity. However, the large-scale antenna array significantly increases the channel dimension, posing serious challenges to traditional pilot-based channel state information (CSI) acquisition methods in terms of pilot overhead and estimation complexity.

[0003] To reduce the overhead of acquiring real-time channel state information, academia and industry have proposed the concepts of channel fingerprint (CF) or channel knowledge map (CKM). These methods significantly reduce reliance on real-time channel state information measurements by pre-constructing a location-related channel knowledge database and querying data during communication. Existing channel fingerprint construction methods can be broadly categorized into model-based and data-driven approaches: the former indirectly constructs channel fingerprints based on channel statistical models, while the latter directly interpolates measurement data or uses deep learning methods for image-to-image prediction.

[0004] Channel fingerprint resolution is a key factor affecting system performance. High-resolution channel fingerprints can preserve more granular channel state information, but they also lead to a significant increase in measurement costs, storage overhead, and privacy risks. Therefore, in practical systems, only low-resolution channel fingerprints are usually available. How to recover high-resolution channel fingerprints from low-resolution channel fingerprints while maintaining accuracy has become a highly challenging engineering problem.

[0005] In recent years, image super-resolution (ISR) methods have received extensive research in the field of computer vision. The task of reconstructing high-resolution images from low-resolution images is naturally similar to the mapping of channel fingerprints from low to high resolution. Variational autoencoders (VAEs) demonstrate strong capabilities in modeling complex data distributions, which often makes them superior to traditional autoregressive networks when facing the ill-conditioned inverse problem of image super-resolution. However, directly using variational autoencoders can lead to significant generation uncertainty in the output of image super-resolution. Summary of the Invention

[0006] Purpose of the invention: To address the shortcomings of existing technologies, the present invention aims to provide a method and system for joint reconstruction of environmental and channel state information, which can recover high-resolution information from low-resolution channel measurement and environmental perception information while ensuring reconstruction accuracy, thereby enabling environmental perception while predicting channel state information.

[0007] Technical Solution: To achieve the above objectives, the present invention provides a method for joint reconstruction of environment and channel state information, which defines a channel fingerprint that simultaneously couples environment information and channel state information, and proposes a joint reconstruction scheme for environment and channel state information. The method includes the following steps:

[0008] The base station divides the target area into grids and collects wireless channel measurement data and environmental information within the target area through downlink detection and sensing.

[0009] The environmental information within the target area is defined using a binary representation method and coupled with the channel state information as a mask. All grids within the area are arranged in spatial order to form a new type of two-dimensional grid channel fingerprint that simultaneously contains environmental information and channel state information.

[0010] The target region is sampled using different spatial resolutions, the channel fingerprint is quantified and described, and then converted into a two-dimensional image format to obtain a high-resolution channel fingerprint and its paired low-resolution channel fingerprint.

[0011] A Conditional Variational Autoencoder (CVAE) network model is constructed, which takes high-resolution channel fingerprints as the reconstruction target and low-resolution channel fingerprints as the conditional input. The encoder maps the high-resolution channel fingerprints to the latent space, and the decoder reconstructs the high-resolution channel fingerprints under the condition of low-resolution channel fingerprints, so as to achieve end-to-end training and obtain the optimal parameters of the model.

[0012] The decoder in the training model is used to achieve joint reconstruction of high-resolution channel state information and environment.

[0013] Preferably, the step of collecting wireless channel measurement data and environmental information within the target area through downlink detection and sensing includes:

[0014] The target area is divided into grids according to the preset spatial resolution;

[0015] Downlink probe signals are sent to the spatial locations corresponding to each grid cell, and the downlink probe signals are received by a receiver set in an open area. Based on the measurement results fed back by the receiver, the channel state information at the corresponding spatial location is obtained.

[0016] Based on the downlink detection signals sent to the spatial locations corresponding to each grid cell, environmental perception is performed grid by grid. Based on the perception results, it is determined whether there are obstructions at the spatial locations corresponding to each grid cell, and environmental occupancy status information corresponding to the spatial resolution is obtained.

[0017] Preferably, the step of defining and distinguishing environmental information within the target area using a binary representation method and coupling it with channel state information as a mask includes:

[0018] Based on the environmental occupancy status information obtained from downlink detection and sensing, the area covered by obstructions is defined as an impassable area, and the corresponding grid cell is assigned a value. Open areas without obstructions are defined as passable areas, and the channel state information measured within the passable areas is linearly scaled and mapped to a preset range.

[0019] The binary-encoded environmental information and the scaled channel state information are arranged spatially on the same two-dimensional grid to form a channel fingerprint that characterizes both the environment and the channel characteristics; where the channel fingerprint is... Corresponding to areas that are impassable.

[0020] Preferably, in the sampling of the target region using different spatial resolutions, the target region is divided into uniform grids, and the spatial resolution is determined by the granularity parameter. Control is performed, with the X and Y axes divided into... 1. Equally spaced intervals are obtained A 3D channel fingerprint, satisfying the following between high and low resolution: , These represent the granularity parameters of high-resolution and low-resolution channel fingerprints, respectively.

[0021] Preferably, the objective of the conditional variational autoencoder network model is to find an optimal set of parameters that maximizes the conditional log-likelihood probability of the high-resolution channel fingerprint with respect to the low-resolution channel fingerprint. ;in, Indicates model parameters, and These represent high-resolution channel fingerprints and low-resolution channel fingerprints, respectively.

[0022] The loss function used for model training includes the reconstruction loss calculated from the pixel-level error between the high-resolution true channel fingerprint and the reconstructed channel fingerprint, and the regularization term calculated from the Kullback-Leibler divergence between the approximate posterior distribution and the prior standard normal distribution of the latent variables:

[0023] ;

[0024] in, Represents latent spatial variables; Indicates to Solve for the expected value; This represents the auxiliary data distribution used to help fit the true data distribution; and Let represent the true latent spatial variable distribution and the true high-resolution channel fingerprint prior distribution, respectively. It follows a standard normal distribution.

[0025] Preferably, the encoder adopts a dual-branch parallel structure. The first branch is used to extract the feature information of the high-resolution channel fingerprint, and the second branch is used to extract the feature representation of the low-resolution channel fingerprint. The first branch passes through several downsampling modules to make the feature dimension consistent with the second branch. After passing through multiple residual modules and convolution operations, the features extracted by the two branches are added element-wise and input into the subsequent convolutional layer to model the mean and variance of the latent variable space. The decoder is used to concatenate the sampled latent variable features with the low-resolution channel fingerprint and output the reconstructed high-resolution channel fingerprint through multiple residual modules and upsampling modules.

[0026] Preferably, the decoder sequentially passes the concatenated features through four cascaded residual modules to fully explore the deep nonlinear correlation between the latent space and the conditional input; at the same time, a corresponding skip connection structure is introduced to alleviate the gradient vanishing problem in deep networks; the decoder adopts a residual learning strategy to learn the residual between the predicted high-resolution channel fingerprint and the nearest neighbor interpolation result, and the final reconstruction result is obtained by adding the predicted residual to the nearest neighbor interpolation result.

[0027] Preferably, during the model training phase, a simulated or measured urban environment dataset is used. The dataset includes multiple urban building floor plans, multiple transmitter locations, and corresponding channel state information. Based on this dataset, matching high- and low-resolution channel fingerprint sample pairs are generated for training the conditional variational autoencoder model. During the inference phase, the low-resolution channel fingerprint and latent spatial variable samples randomly sampled from a standard normal distribution are input into the decoder module of the trained conditional variational autoencoder to achieve joint reconstruction of high-precision channel state information and environmental information.

[0028] The present invention also provides a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the joint reconstruction method of environment and channel state information.

[0029] The present invention also provides a wireless communication system, including a base station and multiple user terminals, wherein the base station or its edge server is used to implement the steps of the joint reconstruction method based on the environmental and channel state information.

[0030] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: This invention utilizes the potential of a reciprocal paradigm that simultaneously reconstructs the physical environment from partial channel state prior information to promote environmental awareness. Specifically, on the one hand, the environment determines channel state information; on the other hand, channel state information also implicitly contains rich environmental information. By arranging environmental and channel state information in spatial order within the same two-dimensional grid, a new paradigm of channel fingerprinting that simultaneously represents both channel state and environmental information is constructed. A conditional variational autoencoder is used to fully mine the prior distribution of high-resolution channel fingerprints given a low-resolution channel fingerprint, achieving high-precision joint reconstruction of channel state and environmental information. The proposed joint reconstruction method of environment and channel state information can significantly reduce the pilot overhead for channel state information measurement and the measurement overhead for environmental awareness in practical systems while ensuring reconstruction accuracy. It achieves environmental awareness while predicting channel state information and outperforms traditional interpolation and existing deep learning methods in terms of reconstruction accuracy. The reconstructed channel state and environmental information can serve as important prior information for information transmission in large-scale wireless communication, thereby improving overall transmission efficiency. Attached Figure Description

[0031] Figure 1 is a general flowchart of an embodiment of the present invention.

[0032] Figure 2 is an example of the joint reconstruction of environment and channel state information in an embodiment of the present invention.

[0033] Figure 3 is a framework diagram of the conditional variational autoencoder network model in an embodiment of the present invention.

[0034] Figure 4 is a schematic diagram of the residual module, upsampling module and downsampling module in an embodiment of the present invention.

[0035] Figure 5 is a schematic diagram comparing the reconstruction performance indicators of the reconstruction method in this embodiment of the invention with those of existing methods.

[0036] Figure 6 is a visual reconstruction diagram of the reconstruction method in the embodiments of the present invention and the existing method. Detailed Implementation

[0037] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0038] As shown in Figure 1, this embodiment of the invention discloses a joint reconstruction method for environment and channel state information, applicable to base station sides or edge servers. First, a channel fingerprint representation that simultaneously includes environment and channel state information is defined. Based on this, a joint reconstruction mechanism based on a conditional variational autoencoder is designed. This mechanism uses a low-resolution channel fingerprint as a conditional input, establishes and optimizes its corresponding lower bound of evidence, thereby achieving high-precision synchronous reconstruction of environment and channel state information.

[0039] Specifically, the joint reconstruction method of environment and channel state information described in this embodiment mainly includes the following steps:

[0040] S1. The base station divides the target area into grids and collects wireless channel measurement data and environmental information within the target area through downlink detection and sensing.

[0041] S2. Use binary representation to define the environmental information within the target area and couple it with the channel state information as a mask. Arrange all grids in the area in spatial order to form a new type of two-dimensional grid channel fingerprint that simultaneously contains environmental information and channel state information.

[0042] S3. Sample the target region using different spatial resolutions, quantify and describe the channel fingerprint, and convert it into a two-dimensional image format to obtain a high-resolution channel fingerprint and its paired low-resolution channel fingerprint.

[0043] S4. Construct a conditional variational autoencoder network model, using high-resolution channel fingerprints as the reconstruction target and low-resolution channel fingerprints as the conditional input. The encoder maps the high-resolution channel fingerprints to the latent space, and the decoder reconstructs the high-resolution channel fingerprints under the given low-resolution channel fingerprint conditions, thus achieving end-to-end training and obtaining the optimal parameters of the model.

[0044] S5. Utilize the decoder in the trained model to achieve joint reconstruction of high-resolution channel state information and environment.

[0045] In some possible implementations, the step S1 of collecting wireless channel measurement data and environmental information within the target area through downlink detection and sensing may include:

[0046] S11. Divide the target area into grids according to the preset spatial resolution.

[0047] S12. Send downlink probe signals to the spatial locations corresponding to each grid cell. Receivers set up in open areas such as streets receive the downlink probe signals and feed back the corresponding measurement results to the base station. The base station then obtains the channel state information at the corresponding spatial locations within each passable area.

[0048] S13. Simultaneously, the base station performs environmental perception grid by grid based on downlink probe signals sent to the spatial locations corresponding to each grid cell. The higher the grid resolution, the higher the density of downlink probe signals sent by the base station to the target area. Based on the perception results, the base station determines whether there are buildings or large obstructions at the spatial location corresponding to each grid cell, thereby obtaining environmental occupancy status information corresponding to the spatial resolution.

[0049] In some possible implementations, step S2, which uses a binary representation method to define and distinguish environmental information within the target area and couples it with channel state information as a mask, may include:

[0050] S21. Based on the environmental occupancy status information obtained from downlink detection and sensing, define grid cells with buildings or large obstructions as impassable areas, and assign the corresponding grid cells a value. Open areas without obstructions, such as streets, are defined as passable areas, and the channel state information measured within these areas is linearly scaled to map them to a preset range.

[0051] S22. Arrange the binary-encoded environment information and the scaled channel state information spatially on the same two-dimensional grid to construct a channel fingerprint that simultaneously represents the environment information and the channel state information. Where the channel fingerprint is... Corresponding to areas that are impassable.

[0052] In some possible implementations, step S3 involves sampling the region using different spatial resolutions, quantizing the obtained high-resolution channel fingerprint 8-bit into a 2D image, and using it as the reconstruction target of the conditional variational autoencoder. The paired low-resolution channel fingerprint 8-bit is quantized into a 2D image and used as conditional information input to the conditional variational autoencoder, thereby achieving high-precision joint reconstruction of environmental and channel state information.

[0053] In some possible implementations, step S3 involves uniformly meshing the target region using granularity parameters. Controlling spatial resolution, dividing the space into sections along the X and Y axes respectively. 1. Equally spaced intervals, thus obtaining a size of A high-resolution channel fingerprint is obtained, and a preset resolution correspondence is maintained between high- and low-resolution channel fingerprints: .

[0054] In some possible implementations, step S4 involves finding an optimal set of parameters that maximizes the conditional log-likelihood of the high-resolution channel fingerprint given a low-resolution channel fingerprint. .in, Indicates model parameters, and These represent high-resolution and low-resolution channel fingerprints, respectively. By deriving the Evidence Lower Bound (ELBO) of the objective function, the original conditional log-likelihood maximization problem is transformed into optimizing the ELBO, thus obtaining the loss function used in model training. The loss function consists of two parts: one part is the reconstruction error term corresponding to the pixel-level deviation between the high-resolution real channel fingerprint and the reconstructed channel fingerprint output by the model; the other part is the regularization term corresponding to the Kullback-Leibler divergence of the approximate posterior distribution relative to the prior standard normal distribution of the latent variables. .in, express Latent spatial variables, Indicates to Solve for the expected value; This represents the auxiliary data distribution used to help fit the true data distribution; and Let represent the true latent spatial variable distribution and the true high-resolution channel fingerprint prior distribution, respectively. It follows a standard normal distribution.

[0055] In some possible implementations, in step S4, to maximize the conditional log-likelihood probability, a joint reconstruction network based on environment and channel state information using a conditional variational autoencoder is constructed. This network mainly consists of two modules: an encoder and a decoder. The encoder extracts and compresses features from the high-resolution channel fingerprint with the assistance of low-resolution conditional information through several residual modules and downsampling modules, mapping it to a lower-dimensional latent space to obtain the distribution representation of the latent variables. The decoder, through several residual modules and upsampling modules, progressively reconstructs the high-resolution channel fingerprint under the condition of simultaneously providing latent variable samples and the low-resolution channel fingerprint, generating a high-fidelity reconstruction result with the same size as the original high-resolution channel fingerprint.

[0056] The residual module consists of several convolutional layers, normalization layers, and nonlinear activation layers connected in series, with skip connections between the input and output. By adding the original features and the features after convolutional transformation element-wise, it effectively alleviates the gradient vanishing and performance degradation problems that occur in deep networks as the number of layers increases, thereby enhancing the network's ability to express complex channel and environmental features and improving the overall reconstruction performance. The downsampling module uses convolutional layers with preset strides in conjunction with corresponding activation units to compress and scale the input channel fingerprint feature map in the spatial dimension, while retaining the low-dimensional representation that is most critical for subsequent modeling, thus playing a role in feature extraction and dimensionality compression. The upsampling module consists of convolutional layers and pixel shaving operations. First, the features are expanded in the channel dimension through convolutional layers, and then the pixel shaving mechanism is used to reorganize the information in the channel direction into the spatial dimension, thereby achieving a step-by-step increase in the resolution of the feature map. This allows for more refined reconstruction of its high-frequency details and boundary structures when recovering high-resolution channel fingerprints, ensuring the consistency of the reconstruction results in terms of visual structure and numerical accuracy.

[0057] In practical applications, during the model training phase, a simulated or measured urban environment dataset is used. This dataset includes multiple urban building floor plans, multiple transmitter locations, and corresponding channel state information. Based on this dataset, matching high- and low-resolution channel fingerprint sample pairs are generated to train the conditional variational autoencoder model. During the inference phase, the low-resolution channel fingerprints and latent spatial variable samples randomly sampled from a standard normal distribution are input into the decoder module of the trained conditional variational autoencoder to achieve joint reconstruction of high-precision channel state information and environmental information.

[0058] For example, taking channel gain as an example of channel state information, as shown in Figure 2, the specific steps of the joint reconstruction method of environment and channel state information in the implementation process may include: (1) The base station configures environmental sensing equipment and channel measurement module, divides the target area into grids, and obtains the physical environment map of the target area and the channel state information of each passable area. (2) The binary masking strategy is used to assign the building and other large-scale obstruction areas to the value. The channel state information measured within the passable area is linearly scaled to... (3) Sample regions with different spatial resolutions and convert them into 2D image format by 8-bit quantization to obtain high and low resolution channel fingerprint sample pairs. (4) Convert the conditional log-likelihood probability to the confidence lower bound and set it as the loss function to train the encoder and decoder modules. (5) The encoder module uses the low-resolution channel fingerprint as a condition and maps the high-resolution channel fingerprint to the latent space through the residual module and the downsampling module. (6) The decoder module uses the low-resolution channel fingerprint as a condition and maps the latent space back to the high-resolution channel fingerprint through the residual module and the upsampling module. (7) Divide the dataset into a training set (6464) and a validation set (1616) in a 4:1 ratio. (8) Use the Adam optimization algorithm to train the above model. The batch size is set to 16 and the learning rate is set to 0.0005 and decayed by half every 40 rounds. This training strategy is used to traverse the training set for 200 rounds. (9) During the inference phase, the low-resolution channel fingerprint and potential spatial samples are input into the pre-trained decoder module to achieve joint reconstruction of high-precision channel state information and environmental information.

[0059] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that the method of the present invention is not only applicable to the specific system model given in the examples below, but also applicable to system models with other configurations.

[0060] I. System Configuration

[0061] Consider the target area The internal wireless communication system includes a base station (BS) and User Equipment (UE) location In the transmission environment Below, position The baseband signal received at the location can be modeled as follows:

[0062] (1)

[0063] in, Indicates spatial response, Indicates Hermitian conjugation. Indicates the transmission power is The launch symbol, For additive noise, its one-sided power spectral density is In this embodiment, the channel state information considered is the channel gain, which consists of two parts: large-scale fading (LSF) and small-scale fading (SSF). Large-scale fading is affected by environmental features at the meter level, while small-scale fading fluctuates with wavelength. In millimeter-wave / centimeter-wave systems, estimating small-scale fading requires unrealistic millimeter-level positioning accuracy; therefore, small-scale fading is usually eliminated by taking the expected value, thus obtaining a channel gain dominated by large-scale fading. The corresponding average received energy is:

[0064] (2)

[0065] in, For signal bandwidth. Furthermore, in order to... Eliminate The influence of this ultimately defines the channel gain as...

[0066] (3)

[0067] This variable only characterizes the environment-related dependencies. However, the electromagnetic propagation mechanism is fundamentally governed by Maxwell's equations, which results in a bidirectional coupling relationship between channel gain and the environment, rather than a unidirectional one. Specifically, environmental parameters (material dielectric constant, geometric characteristics, scatterer configuration, etc.) not only respond to the channel gain, but the channel gain also acts as an information carrier to encode the physical propagation environment. In other words, the channel gain can be expressed as shown in equation (3). Conversely, the environment can also be represented as The function, that is, This bidirectional information coupling laid the theoretical foundation for the subsequently proposed joint reconstruction method of environment and channel gain.

[0068] Based on the above viewpoints, this invention aims to accurately acquire environmental information while precisely reconstructing channel gain. To explicitly capture the inherent coupling relationship between the two, this embodiment encodes the environmental information as a binary image: areas covered by buildings and other large obstructions are defined as impassable areas and assigned a value... Open areas such as streets are defined as passable areas, and the channel gain measured within these areas is... Linear scaling to range Based on this, channel fingerprinting is formally defined as the coupling between the environment and channel gain within the target area, rather than simply containing channel gain information. Channel fingerprinting can be specifically represented as:

[0069] (4)

[0070] and

[0071] (5)

[0072] Furthermore, assuming the region The space dimensions are To address the challenge of obtaining continuous location channel fingerprints in real-world scenarios, this embodiment employs a method based on granular parameters. Controlled uniform grid partitioning scheme, The spatial resolution of discretization is determined so that the target region is along shaft and The axis is simultaneously divided into equally spaced sections:

[0073] (6)

[0074] After discretization, the channel gain or environmental information corresponding to all possible UE locations within the target area can be arranged in spatial order as a two-dimensional tensor:

[0075] (7)

[0076] in, Indicates the first The system uses a grid, with each grid employing one sampling point. Therefore, by arranging environmental information and channel gain in spatial order and discretizing the region, this embodiment proposes a novel form of channel fingerprint representation.

[0077] II. Problem Statement

[0078] High-resolution channel fingerprints preserve channel state information with higher fidelity compared to low-resolution channel fingerprints. However, the cost of acquiring high-resolution channel fingerprints increases with the granularity parameter. The increase shows The secondary growth in resolution leads to a significant increase in sensor deployment costs, data acquisition overhead, and privacy risks. Therefore, in practical systems, only low-resolution channel fingerprints are typically obtained, making the inference of high-resolution channel fingerprints from low-resolution ones a critical engineering challenge. Based on this fundamental limitation, the core objective of this invention is to establish an efficient mapping from low-resolution channel fingerprint measurements to high-resolution channel fingerprints, thereby obtaining accurate estimates of high-resolution channel fingerprints. To characterize the relationship between high-resolution and low-resolution channel fingerprints, an upsampling factor is introduced. It satisfies:

[0079] (8)

[0080] As mentioned earlier, the coupling relationship between environmental information and channel gain allows them to provide each other with prior information during high-resolution prediction. This inherent information reciprocity enables the joint mapping from low resolution to high resolution to be formally represented as:

[0081] (9)

[0082] in, This represents the set of model parameters to be trained. The corresponding optimization objective is to find the optimal parameters that maximize the conditional log-likelihood.

[0083] (10)

[0084] Furthermore, within the aforementioned discretization modeling framework, each element in the channel fingerprint... All of these can be considered as pixel values ​​in an image. It should be noted that this low-resolution to high-resolution channel fingerprint mapping task is essentially an ill-conditioned inverse problem. Traditional feedforward neural networks often struggle to accurately reconstruct the detailed information of the channel fingerprint in this task. Variational autoencoders can effectively model the complex empirical distribution of the target data, capturing implicit generative priors by mapping the data distribution to a standard normal latent space. However, due to their random generation characteristics, variational autoencoders often produce overly smooth or inaccurate reconstruction results. Therefore, this embodiment introduces a conditional variational autoencoder, incorporating the low-resolution channel fingerprint as a conditional constraint into the model, thereby suppressing generation uncertainty while ensuring that the reconstruction results are consistent with existing high-resolution channel fingerprint measurements.

[0085] III. High-Resolution Channel Fingerprint Reconstruction Method Based on Conditional Variational Autoencoder

[0086] This embodiment proposes a model based on a conditional variational autoencoder to reconstruct a high-resolution channel fingerprint by incorporating a low-resolution channel fingerprint as conditional information. To simplify the notation, the following derivations will use... and replace and This represents high-resolution channel fingerprints and low-resolution channel fingerprints.

[0087] 1. Principle and Design of Conditional Variational Autoencoder

[0088] The basic framework of a variational autoencoder consists of two core parts: an encoder and a decoder. The encoder is responsible for processing the target data into a high-resolution channel fingerprint. Mapping to the latent variable space The decoder performs the opposite mapping process, reconstructing the original data from the latent space. However, in practical applications, the standard variational autoencoder framework often suffers from excessive randomness in the generated results, which can easily lead to overly blurry or inaccurate reconstructions.

[0089] To ensure that the reconstructed output meets the desired structural constraints and numerical accuracy requirements, this embodiment uses a conditional variational autoencoder structure to replace the traditional variational autoencoder, and distributes the low-resolution channel fingerprint data... Additional prior information is introduced into the model. In this way, the generation process is not only controlled by latent variables but also constrained by the known low-resolution channel fingerprint, thus significantly reducing generation uncertainty. Within the conditional variational autoencoder framework, the decoder output is typically modeled as a Gaussian distribution, whose conditional probability density function can be expressed as:

[0090] (11)

[0091] in, and These represent the mean and covariance matrices of the output distribution, respectively. In practical implementations, the covariance... They are usually set to a constant to simplify the model training and computation process.

[0092] However, directly addressing conditional probability In general, this solution is unsolvable. This is because the integral form... Analytical solutions are typically unavailable. To address this issue, the encoder introduces an auxiliary distribution. Used to approximate the true posterior distribution In this embodiment, the auxiliary distribution is assumed to be a Gaussian distribution, that is:

[0093] (12)

[0094] in, and Let represent the mean and covariance of the posterior distribution of the latent variables, respectively. Combining the optimization objective (10) given above, the training objective of the conditional variational autoencoder can be expressed as maximizing the conditional log-likelihood function:

[0095] (13)

[0096] in, The Kullback–Leibler divergence is defined as follows:

[0097] (14)

[0098] Furthermore, by applying Jensen's inequality, the conditional log-likelihood above can be lowered as follows:

[0099] (15)

[0100] The expression on the right is called the lower confidence bound.

[0101] Therefore, the original problem of maximizing the conditional log-likelihood can be equivalently transformed into maximizing the confidence lower bound, and its specific optimization objective function can be written as:

[0102] (16)

[0103] in, Let n represent the number of samples in a batch, with the subscript n indicating the nth sample. During the optimization process of equation (16), due to the latent variables... The sampling operation would block the backpropagation of the gradient, so this embodiment uses the reparameterization trick to solve this problem. Specifically, noise variables are first sampled from a standard normal distribution. The encoder then outputs the mean of the latent distribution. Covariance The latent variable sample was ultimately obtained in the following manner:

[0104] (17)

[0105] Furthermore, to make the KL divergence term computationally tractable, this embodiment assumes latent variables... The prior distribution of is a standard normal distribution. Under this assumption, the KL divergence can be simplified to the following analytical form:

[0106] (18)

[0107] in, Representing latent variables Dimensions and They represent and The elements in the corresponding dimension.

[0108] 2. Conditional Variational Autoencoder Network Structure

[0109] As mentioned earlier, the conditional variational autoencoder network structure proposed in this embodiment mainly consists of two core parts: an encoder and a decoder. Furthermore, to enhance feature representation capabilities and ensure spatial scale matching, several basic functional modules are introduced into the network, including a residual module, an upsampling module, and a downsampling module. The overall network structure diagram and the basic module structure diagram are shown in Figures 3 and 4, respectively. This represents the final high-resolution channel fingerprint result obtained from the reconstruction.

[0110] 1) Basic Modules: In order to build a deep, efficient and stable network structure, this embodiment designs and adopts three types of basic modules, which are used for feature extraction, scale transformation and deep information fusion, respectively.

[0111] The residual module is introduced to alleviate the vanishing gradient and model degradation problems commonly encountered during deep neural network training. This module introduces a skip connection between the input and output, enabling the network to learn identity mappings more easily, thereby improving training stability and convergence speed. Specifically, the residual module consists of two convolutional layers, two normalization layers, a non-linear activation function, and a skip connection structure. Its output is defined as follows:

[0112] (19)

[0113] in, , express Convolution operations are used to align channels when the number of input and output channels is inconsistent. , and They represent Convolution, batch normalization, and nonlinear activation operations are used. Here, the input feature tensor... , Output feature tensor Through the above design, the residual module effectively preserves the original input information while enhancing the nonlinear representation capability of features.

[0114] The upsampling module is used to improve the spatial resolution of the feature map and is a key component for achieving low-resolution to high-resolution reconstruction. This embodiment employs a structure of "convolution + pixel rearrangement + activation function" to enhance the model's learning ability and avoid information loss caused by traditional interpolation methods. Specifically, this module first upsamples the input features... Perform a convolution operation to expand the number of channels. Thus, intermediate feature mapping is obtained. Subsequently, a pixel rearrangement operation is used to remap the channel dimension information to the spatial dimension, thereby increasing the spatial size of the output feature map. The mapping relationship can be expressed as:

[0115] (20)

[0116] in, and In this way, each upsampling module can achieve a 2x increase in spatial resolution while maintaining high reconstruction accuracy.

[0117] The downsampling module reduces the spatial resolution of the feature map, thereby enabling multi-scale feature extraction and scale alignment. This module has a relatively simple structure, consisting of one convolutional layer and one activation function layer. By using strided convolution, the spatial size of the input feature map is reduced to 50% of its original size to accommodate the feature fusion requirements between different network branches.

[0118] 2) Encoder: To effectively incorporate low-resolution channel fingerprints as conditional information into the latent space modeling process, the encoder employs a dual-branch parallel structure. The first branch extracts feature information from the high-resolution channel fingerprint, while the second branch is specifically dedicated to extracting the feature representation of the low-resolution channel fingerprint. In the specific implementation, for... The branch will go through several downsampling modules to reduce its spatial resolution, making its feature dimensions similar to those of the previous branch. The branches maintain consistency to facilitate subsequent feature fusion. After passing through multiple residual modules and convolutional operations, the features extracted from the two branches are element-wise summed and then further input into subsequent convolutional layers to model the mean and variance of the latent variable space. Finally, the encoder outputs the parameters of the latent variable distribution. and Thus, the conditional posterior distribution is completed. The encoder structure is shown in the green area in Figure 3.

[0119] 3) Decoder: During the training phase, latent variables Sampling is performed using reparameterization techniques; however, during the inference phase, sampling is performed directly from the standard normal distribution. Sampling is performed during the sampling process. The latent features obtained from the sampling are... The latent features are then concatenated with the low-resolution channel fingerprint, explicitly injecting conditional information into the decoding process. It's important to note that the feature fusion method in the decoder differs from that in the encoding stage, with a greater emphasis on recovering high-level semantic information. The concatenated features are then passed sequentially through four cascaded residual modules to fully exploit the deep nonlinear correlation between the latent space and the conditional input. Simultaneously, a corresponding skip connection structure is introduced to alleviate the gradient vanishing problem in deep networks. Regarding spatial scale recovery, the latent features... The process requires several upsampling modules, combined with convolution operations, to gradually increase its spatial resolution to the target high-resolution size. Furthermore, the decoder employs a residual learning strategy: instead of directly predicting the final high-resolution channel fingerprint, it learns to predict its nearest-neighbor interpolation results. The residuals between the predicted residuals and the final reconstruction result is obtained by comparing the predicted residuals with the actual residuals. The results are obtained by adding them together. This strategy can effectively reduce the learning difficulty and further improve the reconstruction accuracy. The overall structure of the decoder is marked with a gold area in Figure 3.

[0120] This network, based on a conditional variational autoencoder, incorporates low-resolution channel fingerprints as conditions into the latent variable modeling and reconstruction process. This effectively constrains the latent space distribution while preserving random sampling capabilities, ensuring a high degree of consistency between the generated results and the input structure. Building upon this, the network fuses conditional information and latent variable features during the decoding stage through feature concatenation. This maximizes the preservation of complete information during reconstruction, further enhancing structure preservation and suppressing irrelevant noise. Subsequently, the network combines convolution and pixel rearrangement to perform upsampling, effectively avoiding checkerboard artifacts and enhancing high-frequency detail recovery. Simultaneously, a residual module is introduced to stabilize the training process and progressively refine texture information, thereby improving the overall stability and detail of the reconstruction.

[0121] IV. Implementation Results

[0122] To enable those skilled in the art to better understand the present invention, the performance results of this embodiment and existing methods under two specific system configurations are compared below.

[0123] First, the dataset used in this embodiment is described. The DPM dataset is used to train and validate the proposed model. This dataset contains 101 urban environment maps. In each map, the locations of 80 transmitters are given, and the corresponding channel gain distribution is generated based on electromagnetic propagation simulation. Figure 3 shows a typical sample example from this dataset, illustrating the urban environment layout and its corresponding channel characteristic distribution. In the simulation scenario set in this embodiment, the channel fingerprint includes both channel gain information and environmental information. The spatial scale parameters for high-resolution and low-resolution channel fingerprints are set as follows: and ,Right now, and , corresponding to one Reconstruction Task. It should be noted that the low-resolution channel fingerprints are not generated independently, but are obtained by uniformly downsampling their corresponding high-resolution channel fingerprints, thus ensuring strict consistency in the spatial structure of the sample pairs. Regarding sample partitioning, a total of 8080 sample pairs were generated and randomly divided into training and validation sets at a ratio of 4:1, containing 6464 training pairs and 1616 validation pairs respectively. Furthermore, to ensure that the channel fingerprint mapping task maintains numerical scale consistency with the image super-resolution problem, all channel fingerprint data were normalized in the experiment. Specifically, the min-max normalization method was used, with the expression:

[0124] (twenty one)

[0125] Through the above normalization operation, all channel fingerprint values ​​are mapped to... This interval improves the stability and convergence speed of model training.

[0126] Next, the experimental implementation details of this embodiment are given. All experiments were performed on an NVIDIA GeForce RTX3060 GPU (12GB VRAM), using PyTorch 2.3.0 as the deep learning framework, combined with CUDA 12.1 and cuDNN 8.8.1 for accelerated computation. During model training, the Adam optimizer was selected, along with the MultiStepLR learning rate scheduling strategy. The initial learning rate was set to... During training, the batch size is halved every 40 epochs, for a total of 200 epochs, with a batch size of 16. To comprehensively evaluate reconstruction quality, this embodiment employs four commonly used and standardized performance metrics: Mean Squared Error (MSE), Normalized Mean Squared Error (NMSE), Peak Signal-to-Noise Ratio (PSNR), and Structural Similarity Index (SSIM). MSE and NMSE quantify numerical reconstruction errors in the spatial domain; smaller values ​​indicate higher reconstruction accuracy. PSNR primarily reflects the model's ability to suppress noise and distortion; larger values ​​indicate better reconstruction quality. SSIM measures the consistency between the reconstructed result and the real channel fingerprint in terms of spatial structure from a structure-aware perspective; values ​​closer to 1 indicate higher structural similarity. It should be noted that to ensure comparability of evaluation results between different methods, all metrics are calculated based on 8-bit quantization.

[0127] Finally, a quantitative and qualitative comparison of the estimation performance of the conditional variational autoencoder (CDAE) method in this embodiment with existing methods is presented. The methods compared include nearest neighbor interpolation (NEAREST), bicubic interpolation (BICUBIC), and the SRGAN method proposed in the literature "Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network." Figure 5 shows a quantitative comparison of the performance of the CDAE method in this embodiment and the comparative methods on the same dataset under the considered wireless communication system. The results show that the CDAE method proposed in this embodiment significantly outperforms the comparative methods in all evaluation metrics. Specifically, compared with SRGAN and the other two traditional interpolation methods, the CDAE reduces the MSE and NMS metrics by 1–2 orders of magnitude, indicating a significant advantage in numerical reconstruction accuracy. In terms of PSNR, the Conditional Variational Autoencoder (CVA) improves upon SRGAN, NEAREST, and BICUBIC by 15.1425 dB, 17.0239 dB, and 16.8524 dB, respectively, demonstrating its superior performance in suppressing reconstruction errors and noise. Simultaneously, the CVA improves upon the superior BICUBIC method by 0.1155 in SSIM, indicating its outstanding performance in structure preservation. The significant reduction in NMSE further validates the effectiveness and robustness of the CVA in reconstructing high-resolution channel fingerprints from low-resolution ones. Figure 6 presents a qualitative comparison of the performance of the CVA in this embodiment with the comparative methods on the same dataset under the considered wireless communication system. The second row of the figure shows a magnified view to more intuitively observe the performance of each method in detail recovery. Visually, the reconstruction results of the CVA are almost indistinguishable from real high-resolution channel fingerprints, clearly recovering the spatial variation trends of environmental boundaries and channel gain. In contrast, SRGAN's reconstruction results exhibit noticeable mesh-like artifacts and noise, likely stemming from instability caused by the mismatch in learning progress between the generator and discriminator during adversarial training. Interpolation methods such as NEAREST and BICUBIC generally suffer from blurred boundaries and loss of structural details, which may introduce non-negligible errors in practical communication system applications. Combining quantitative and qualitative analysis results, it is evident that the proposed conditional variational autoencoder method significantly outperforms existing mainstream methods in terms of reconstruction accuracy, structural consistency, and visual quality.Furthermore, compared to methods that only model channel state information, this invention couples environmental information and channel state information in the same two-dimensional channel fingerprint, so that the reconstruction process is simultaneously subject to the combined effects of environmental constraints and channel state information. Thus, without introducing additional experimental complexity, it achieves a simultaneous improvement in the accuracy of channel state information reconstruction and the ability to perceive the environment.

[0128] Based on the same inventive concept, an embodiment of the present invention discloses a computer system including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the joint reconstruction method of environment and channel state information.

[0129] In a specific implementation, the computer system includes a processor, a communication bus, memory, and a communication interface. The processor can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. The communication bus may include a pathway for transmitting information between the aforementioned components. The communication interface, using any transceiver-like device, is used for communicating with other devices or communication networks. The memory can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), read-only optical disc (CD-ROM) or other optical disc storage, disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor via a bus. The memory can also be integrated with the processor.

[0130] The memory stores application code that executes the present invention and is controlled by the processor. The processor executes the application code stored in the memory to implement the refactoring method provided in the above embodiments. The processor may include one or more CPUs, or multiple processors, each of which may be a single-core processor or a multi-core processor. Here, "processor" may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0131] Based on the same inventive concept, this invention discloses a wireless communication system including a base station and multiple user terminals. The base station or its edge server is used to implement the steps of the aforementioned joint reconstruction method of environment and channel state information. The base station obtains a low-resolution geometric location map and channel state information of a specific area through environmental sensing equipment and measurement equipment. It then uses a decoder module in a pre-trained conditional variational autoencoder model to reconstruct a low-resolution channel fingerprint of the target area in real time from the measured low-resolution channel fingerprint.

[0132] It should be noted that relational terms such as "first" and "second" in this specification are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0133] In the embodiments provided in this application, it should be understood that the disclosed methods can be implemented in other ways without departing from the spirit and scope of this application. The current embodiments are merely exemplary examples and should not be considered limiting, nor should the specific content given limit the purpose of this application. For example, some features may be omitted or not implemented.

[0134] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for joint reconstruction of environment and channel state information, characterized in that, The steps include: the base station divides the target area into grids and collects wireless channel measurement data and environmental information in the target area through downlink detection and sensing; the environmental information in the target area is defined using a binary representation method and coupled with the channel state information as a mask; all grids in the area are arranged in spatial order to form a new type of two-dimensional grid channel fingerprint that simultaneously contains environmental information and channel state information. The target region is sampled using different spatial resolutions, the channel fingerprint is quantified and converted into a two-dimensional image format to obtain a high-resolution channel fingerprint and its paired low-resolution channel fingerprint. A conditional variational autoencoder network model is constructed, using the high-resolution channel fingerprint as the reconstruction target and the low-resolution channel fingerprint as the conditional input. The encoder maps the high-resolution channel fingerprint to the latent space, and the decoder reconstructs the high-resolution channel fingerprint under the given low-resolution channel fingerprint conditions, achieving end-to-end training and obtaining the optimal parameters of the model. The decoder in the trained model is used to achieve joint reconstruction of high-resolution channel state information and environment.

2. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, The step of collecting wireless channel measurement data and environmental information within the target area through downlink detection and sensing includes: dividing the target area into grids according to a preset spatial resolution; sending downlink detection signals to the spatial locations corresponding to each grid cell, and receiving the downlink detection signals by a receiver located in an open area; obtaining channel state information at the corresponding spatial locations based on the measurement results fed back by the receiver; performing environmental sensing grid by grid based on the downlink detection signals sent to the spatial locations corresponding to each grid cell; determining whether there are obstructions at the spatial locations corresponding to each grid cell based on the sensing results; and obtaining environmental occupancy status information corresponding to the spatial resolution.

3. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, The method of defining and distinguishing environmental information within the target area using binary representation and coupling it with channel state information as a mask includes: defining the area covered by obstructions as a non-passable area based on the environmental occupancy state information obtained from downlink detection and sensing, and assigning the corresponding grid cell value to... Open areas without obstructions are defined as passable areas, and the channel state information measured within these areas is linearly scaled and mapped to a preset interval. The binary-encoded environmental information and the scaled channel state information are arranged spatially on the same two-dimensional grid to form a channel fingerprint that characterizes both the environment and the channel characteristics. Where the channel fingerprint is... Corresponding to areas that are impassable.

4. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, In the process of sampling the target region using different spatial resolutions, the target region is divided into uniform grids, and the spatial resolution is determined by the granularity parameter. Control is performed, with the X and Y axes divided into...

1. Equally spaced intervals are obtained A 3D channel fingerprint, satisfying the following between high and low resolution: , These represent the granularity parameters of high-resolution and low-resolution channel fingerprints, respectively.

5. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, The objective of the conditional variational autoencoder network model is to find an optimal set of parameters that maximizes the conditional log-likelihood probability of the high-resolution channel fingerprint with respect to the low-resolution channel fingerprint. ;in, Indicates model parameters, and These represent the high-resolution channel fingerprint and the low-resolution channel fingerprint, respectively. The loss function used for model training includes the reconstruction loss calculated from the pixel-level error between the high-resolution true channel fingerprint and the reconstructed channel fingerprint, and the regularization term calculated from the Kullback-Leibler divergence between the approximate posterior distribution and the prior standard normal distribution of the latent variables. ;in, Represents latent spatial variables; Indicates to Solve for the expected value; This represents the auxiliary data distribution used to help fit the true data distribution; and These represent the true latent spatial variable distribution and the true high-resolution channel fingerprint prior distribution, respectively.

6. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, The encoder employs a dual-branch parallel structure. The first branch extracts feature information from the high-resolution channel fingerprint, while the second branch extracts feature representations from the low-resolution channel fingerprint. The first branch undergoes several downsampling modules to ensure that the feature dimensions are consistent with those of the second branch. After passing through multiple residual modules and convolution operations, the extracted features from both branches are summed element-wise and input into subsequent convolutional layers to model the mean and variance of the latent variable space. The decoder concatenates the sampled latent variable features with the low-resolution channel fingerprint and outputs the reconstructed high-resolution channel fingerprint through multiple residual modules and upsampling modules.

7. The method for joint reconstruction of environment and channel state information according to claim 6, characterized in that, The decoder sequentially passes the concatenated features through four cascaded residual modules to fully explore the deep nonlinear correlation between the latent space and the conditional input. At the same time, a corresponding skip connection structure is introduced to alleviate the gradient vanishing problem in deep networks. The decoder adopts a residual learning strategy to learn the residual between the predicted high-resolution channel fingerprint and the nearest neighbor interpolation result. The final reconstruction result is obtained by adding the predicted residual to the nearest neighbor interpolation result.

8. The method for joint reconstruction of environment and channel state information according to claim 1, characterized in that, During the model training phase, a simulated or measured urban environment dataset is used. The dataset includes multiple urban building floor plans, multiple transmitter locations, and corresponding channel state information. Based on this dataset, matching high- and low-resolution channel fingerprint sample pairs are generated for training the conditional variational autoencoder model. During the inference phase, the low-resolution channel fingerprint and the latent spatial variable samples obtained by random sampling from the standard normal distribution are input into the decoder module of the trained conditional variational autoencoder to achieve joint reconstruction of high-precision channel state information and environmental information.

9. A computer system, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the joint reconstruction method of environmental and channel state information as described in any one of claims 1 to 8.

10. A wireless communication system, comprising a base station and multiple user terminals, characterized in that, The base station or its edge server is used to implement the steps of the joint reconstruction method of environment and channel state information according to any one of claims 1 to 8.