Generative imaging method based on multi-view channel feature fusion
Through the generative imaging method of multi-view channel feature fusion, the multi-view channel encoder and point cloud diffusion model are used to solve the scatter imaging problem caused by changes in user equipment and base stations in the ISAC system, and efficient and flexible environmental perception and imaging are achieved.
Patent Information
- Application Number
- CN202510566097.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
In the integrated perception communication (ISAC), existing wireless communication systems are difficult to effectively deal with multi-view perception problems, especially when the number and position of user equipment and base stations are changed, it is impossible to accurately extract the environmental scatterer features and complete scatterer imaging.
The generative imaging method of multi-view channel feature fusion is adopted, and the shape and electromagnetic characteristics of the scatterer are extracted and reconstructed through the multi-view channel encoder and point cloud diffusion model, and the shape and electromagnetic characteristics of the scatterer are trained and optimized using the Transformer encoder and the shape-electromagnetic weighted diffusion loss function.
In the scenario of changing the number and location of users and base stations, high-quality scattering imaging is achieved, improving the flexibility and accuracy of environmental perception, and solving the limitations of traditional methods on fixed space measurement configuration.
Smart Images

Figure CN120451404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communications, and in particular to a generative imaging method based on multi-view channel feature fusion. Background Art
[0002] With the rapid development of wireless communication technology, Integrated Sensing and Communication (ISAC) will become a key technology for next-generation mobile communication systems. ISAC aims to leverage the widespread presence of wireless communication signals in space to achieve environmental perception, including the location, shape, and electromagnetic properties of environmental scatterers. This will enable a variety of emerging application scenarios, including the Internet of Vehicles, extended reality, and digital twins.
[0003] In uplink communication scenarios, the base station uses pilot signals sent by user equipment to obtain Channel State Information (CSI) between the user and the base station to support subsequent communication services. However, since the wireless signals transmitted by the user are scattered by the environment during propagation, the CSI obtained by the base station also includes the influence of scatterers. Therefore, the base station can use uplink CSI to extract information about scatterers in the environment, achieving passive environmental perception and imaging.
[0004] The environmental imaging task in an ISAC system is essentially an electromagnetic inverse scattering problem. Due to the varying occlusion relationships between scatterers under different incident wave illumination modes, the scattered signals received at different locations under the same illumination mode also vary significantly. To accurately perceive complex targets, it is necessary to jointly process the CSI from multiple transmit and receive perspectives in the wireless system to obtain a comprehensive observation of the environment, thereby addressing the inherent nonlinearity and pathological nature of the electromagnetic inverse scattering imaging problem. Currently, several research efforts have explored multi-perspective perception schemes for ISAC systems, verifying the feasibility of multi-perspective perception and the performance gains it brings. However, most of these current schemes employ traditional radar signal processing or compressed sensing algorithms, which are based on scattered signal propagation modeling, and some of which require approximate statistical models to describe scene prior information.
[0005] To overcome the reliance of traditional algorithms on forward modeling and statistical priors, applying artificial intelligence (AI) technology to environmental perception algorithm design is a promising research direction. Several studies have applied AI to solving electromagnetic inverse scattering problems under specific measurement configurations, beam optimization in ISAC systems, and environmental imaging tasks in single-base station radar mode, demonstrating the potential of AI in wireless perception. However, existing solutions are not designed for multi-perspective perception and struggle to cope with the frequent changes in the position and number of transmitter and receiver pairs in ISAC systems.
[0006] Unlike traditional AI solutions, generative AI technology can learn from complex data distributions and be controlled through conditional mechanisms or prompts, allowing for flexible application to a wide range of downstream tasks. With the increasing computing power of base stations, processing ISAC tasks based on generative models will also become a highly promising solution. Therefore, in ISAC systems, considering how to effectively extract the characteristics of environmental scatterers contained in multi-view channels when the number and location of user devices and base stations change, and controlling the generative model to complete the scatterer imaging task, is a highly challenging and practical issue. Summary of the Invention
[0007] The present invention aims to address the problem of utilizing uplink channels between multiple users and multiple base stations for scatterer imaging in uplink wireless communication scenarios. This approach utilizes uplink channel CSI (Critical Signal Indicator) estimated through pilot signals between multiple user-base station transceiver pairs for sensing. This approach is compatible with existing wireless communication systems and implements ISAC functionality.
[0008] The specific technical solution adopted by the present invention is as follows: a generative imaging method based on multi-view channel feature fusion, the method comprising the following steps:
[0009] S1. Each base station estimates the uplink CSI between the base station and each user through the pilot signal sent by the user;
[0010] S2. Upload all CSI data estimated by each base station to the central server and merge them into CSI of all transmitting and receiving perspectives;
[0011] S3. Build a multi-view channel encoder, embed the spatial location information of the user and base station into the channel features, extract the multi-view channel features and perform feature fusion to obtain the characteristics of the target scatterer;
[0012] S4. Build a point cloud diffusion model. The point cloud diffusion model controls the point cloud diffusion process through the characteristics of the target scatterer, and gradually reversely diffuses the noise point cloud into a meaningful scatterer point cloud.
[0013] S5. Use the shape-electromagnetic weighted diffusion loss function as the optimization objective to train the multi-view channel encoder and point cloud diffusion model. Use the trained model to process the CSI of all transmitting and receiving perspectives to obtain a point cloud that describes the shape and electromagnetic characteristics of the target scatterer, thereby achieving imaging of the scatterer.
[0014] Furthermore, the estimating of the uplink CSI between the base station and each user is specifically:
[0015] Single-view channel data is obtained by measuring the channel between each user and the base station or using a deterministic channel simulation model. The single-view channel data set includes a CSI matrix, base station coordinates, and user coordinates. The use of a deterministic channel simulation model to obtain single-view channel data includes: dividing the area of interest into pixels, obtaining the contrast of the pixelated scatterer including the scatterer shape and electromagnetic properties, and simulating the space-frequency domain channel between each base station and the user through the method of moments.
[0016] Furthermore, the embedding of the spatial position information of the user and the base station into the channel features specifically includes: flattening the CSI of each channel into a vector and mapping it to a single-view channel feature vector of uniform dimension through a fully connected layer, position encoding the base station coordinates and the user coordinates, obtaining the position vectors of the base station and the user and splicing them; and embedding the spliced position vector into the channel features through linear modulation.
[0017] Furthermore, the position encoding of the base station coordinates and the user coordinates is specifically as follows:
[0018] ξ(p)=[p;sin(2 0 πp); cos(2 0 πp);…;sin(2 dp-1 πp); cos(2 dp-1 πp)]
[0019] Map each coordinate component from a scalar to a (2d p +1)-dimensional vector, where p is the position vector, d p is the coding frequency number.
[0020] Furthermore, the linear modulation is specifically:
[0021]
[0022] where γ p and β p Both are 4·(2d p +1)→d model Fully connected layer, model width d model is the dimension of the channel feature vector, ξ b,u is the single view position vector.
[0023] Furthermore, the extraction of multi-view channel features and feature fusion are specifically as follows: a set of single-view channel feature vectors is input into a multi-view channel converter, the multi-view channel converter adopts a Transformer encoder structure, extracts the correlation between different view channel features through a multi-head self-attention layer, performs nonlinear feature transformation through a feedforward neural network, and then obtains multi-view channel features after output average pooling, inputs the multi-view channel features into two MLPs to predict the mean and variance respectively, and obtains the characteristics of the target scatterer after parameter resampling.
[0024] Furthermore, the point cloud diffusion model is specifically:
[0025] The original point cloud distribution is gradually diffused into noise distribution through a predefined noise addition process and is modeled as a Markov chain to obtain the distribution
[0026] Construct a learnable Markov chain based on the feature z output by the multi-view channel encoder to realize the diffusion process of gradually reversely diffusing the noise point cloud into the scatterer point cloud. The input point is distributed from the noise Sampling to approximate the distribution Right now
[0027]
[0028] And the transition probability is modeled as a Gaussian distribution of the following form,
[0029]
[0030] in is the noise variance at time step t during the back diffusion process, μ θ is the mean value estimated by the point cloud diffusion model with parameter θ, μ θ Has the following form,
[0031]
[0032] where ∈ θ A noise estimation network that progressively denoises noisy samples based on the noise scheme of the point cloud forward diffusion process;
[0033] The dimension of the point cloud is the sum of the shape dimension and the electromagnetic property dimension, where the electromagnetic property dimension includes the relative dielectric constant and conductivity of the scatterer; each point cloud is modeled as the result of multiple independent sampling of a point distribution, and the number of independent samplings is the total number of shape-electromagnetic points; the shape and electromagnetic parameters corresponding to the point cloud are controlled by the characteristics of the scatterer.
[0034] Furthermore, the diffusion loss function of the shape-electromagnetic weighting is specifically:
[0035]
[0036] in,
[0037]
[0038] Multi-view channel encoder The learnable neural network parameters are Represents the multi-view CSI and corresponding scatterer point cloud containing the position information of the transmitter and receiver in the training set, Expressed as a multi-view channel coding process, Noise Vector The subscripts s and EM represent the dimensions of the shape attribute and the electromagnetic attribute in the selection vector, respectively; γ s , γ EM , γ z are the weights of point cloud shape, electromagnetic properties, and regularization term respectively; β t is the noise variance, is the initial state of the point cloud, p w is the probability density function of the standard Gaussian distribution; the normalized flow model F φ It is composed of a series of affine coupled layers stacked to meet the requirements of bijection.
[0039] On the other hand, the present invention also provides a generative imaging device based on multi-view channel feature fusion, comprising a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it implements the generative imaging method based on multi-view channel feature fusion.
[0040] On the other hand, the present invention also provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the generative imaging method based on multi-view channel feature fusion.
[0041] The beneficial effects of the present invention are as follows: the multi-perspective channel encoder proposed in the present invention can effectively extract the shape and electromagnetic characteristics of scatterers in the scene according to the uplink CSI of multiple perspectives in a scene containing any number of users and base stations and the users and base stations are in any position, and complete the generative scatterer imaging in combination with the point cloud diffusion model, and the imaging quality can be improved as the number of perspectives increases. The multi-perspective channel encoder realizes the interaction of multi-perspective channel features through MVT, and performs feature aggregation through average pooling, fully integrating the intrinsic characteristics of multi-perspective channels, and well solves the limitation that the current inverse scattering imaging method based on deep learning is only applicable to fixed space measurement configurations. At the same time, we also propose a shape-electromagnetic weighted diffusion loss function based on the characteristics of the scatterer imaging task, and realize the generation of high-quality scatterer point clouds. The generative imaging method based on multi-perspective channel feature fusion proposed in the present invention significantly improves the flexibility and accuracy of environmental perception, and provides an efficient intelligent environmental perception solution for the design of future ISAC systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a schematic diagram of the ISAC system including multiple users and multiple base stations considered in this solution.
[0043] Figure 2 This is a schematic diagram of the architecture for scatterer point cloud reconstruction based on multi-view channel feature fusion. The left side shows the multi-view channel encoder structure and MVT network structure, and the right side shows the point cloud diffusion model for scatterer reconstruction.
[0044] Figure 3 It is the process of gradually reconstructing the scatterer point cloud from the noise distribution under the guidance of the scatterer characteristics by the diffusion model;
[0045] Figure 4 These are the reconstruction results of different scatterer targets under different numbers of viewing angles according to the present invention;
[0046] Figure 5 This is a graph showing the relationship between the number of users and the scatterer imaging accuracy under different numbers of base stations in the present invention;
[0047] Figure 6 Schematic diagram of a generative imaging device based on multi-view channel feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings.
[0049] In this embodiment, we consider Figure 1 In the two-dimensional ISAC scenario shown in the figure, a total of B base stations (BS) are deployed, each BS is equipped with N rAntennas are used, using a uniform linear array (ULA) with a spacing of half a wavelength, and there are U single-antenna user equipment (UE). It is assumed that the positions of the BS and UE are known. In the uplink communication scenario, the UE sends a pilot signal to the BS, which uses orthogonal frequency division multiplexing (OFDM) modulation and sends the pilot signal on a fixed K subcarriers. The transmitted signal is scattered by the scatterer and received by each BS. Each BS estimates the uplink channel between itself and the UE through the pilot. It is assumed that different UEs send pilot signals in different time slots to avoid interference between users. The goal of the ISAC system is to reconstruct the shape and electromagnetic properties of the scatterer based on the uplink multi-view channel data.
[0050] In this embodiment, B≤16, U≤32, N r =4, K=8, carrier center frequency f c =3 GHz, with a subcarrier spacing of Δf = 100 kHz. The number and location of base stations and user equipment (UEs) vary depending on the scatterers to be sensed in different scenarios. Specifically, base stations are randomly distributed within a range of 80 to 100 meters from the scene center, and user equipment (UEs) are randomly distributed within a range of 4 to 10 meters from the scene center. Scatterers are located within a 0.5 m × 0.5 m square region of interest (RoI) at the scene center.
[0051] This embodiment provides a generative imaging method based on multi-view channel feature fusion, which includes the following steps:
[0052] S1. Each BS estimates the N distance between it and each UE based on the pilot signal. r Single-view CSI matrix of antenna K carrier.
[0053] S2, all BSs upload the estimated single-view CSI to the central server and merge them to obtain B×U dimensions of N r ×K CSI matrices form the uplink multi-view channel.
[0054] S3. Build a multi-view channel encoder, where the learnable parameters are It extracts multi-view channel features and performs feature fusion based on the CSI data of each view and the corresponding UE and BS positions to obtain the feature z of the target scatterer.
[0055] S4. Build a point cloud diffusion model, where the learnable parameter is θ. Based on the characteristic z of the target scatterer, control the generation of the point cloud diffusion model to obtain a point cloud that describes the target scatterer's shape and electromagnetic properties, enabling imaging of the scatterer.
[0056] For step S1, we use the Method of Moments (MoM) to simulate the single-view channel between each UE and BS under specific scatterer targets and specific UE and BS deployment locations. The shape of the scatterer target is randomly sampled from the MNIST handwritten digit dataset and scaled to the size of the RoI. The relative dielectric constant ε is r ∈[1.5,2.5], conductivity σ∈[0,0.1](S / m), the locations of UE and BS. The specific steps are as follows:
[0057] S1.1 divides the RoI into D pixels evenly, and the pixel representation of the scatterer is obtained as This is similar to a two-channel image where Represent the relative dielectric constant and conductivity at each pixel respectively. At the frequency f corresponding to the kth subcarrier k The pixelated scatterer contrast is expressed as
[0058]
[0059] χ k Includes the shape and electromagnetic properties of the scatterer.
[0060] S1.2 calculates the uplink channel between any UE and BS by MoM. For the b-th BS and u-th UE consisting of a transceiver pair, the spatial channel on the k-th subcarrier is
[0061]
[0062] Among them, the channel from UE to RoI Channel from RoI to BS Both are completely determined by UE location and BS location Therefore, the scatterer target information x is completely contained in the scatterer response matrix Consider multiple fixed subcarrier frequencies f=[f1,…,f K ], then the spatial frequency domain channel between each UE and BS is
[0063]
[0064] in Therefore, the single-view channel data with BS and UE coordinates is obtained
[0065] For step S2, after merging all the single-view channel data obtained in step S1.2, the resulting multi-view channel data is the following set:
[0066]
[0067] In this case, we repeated steps S1 and S2 on samples from the MNIST dataset to synthesize 50,000 data points, which were divided into training, validation, and test sets in a ratio of 8:1:1 for subsequent model training, tuning, and testing.
[0068] For step S3, we need to encode the CSI of each view into a single-view channel feature and embed the corresponding transmitter-receiver pair position information. Then, through the multi-view channel converter, the channel features of all views are exchanged and nonlinearly transformed. Finally, the channel features of all views are aggregated to obtain the multi-view channel feature z. The specific steps are as follows:
[0069] S3.1. First, the CSIH of each view b,u Flatten to N r ·K-dimensional vector, then through a shared N r K→d model The fully connected layer f h , mapping it to a dimension of d model The single-view channel feature vector h b,u ,Right now
[0070]
[0071] In this case, the model width d model =256, B and U represent the number of base stations and users respectively.
[0072] S3.2. Use the following position coding format to encode the base station location and user location Mapped to a high-dimensional position vector,
[0073]
[0074] It maps each coordinate component from a scalar to a (2d p +1)-dimensional vector, in this case the encoding frequency d p = 10. The encoded base station and user position vectors are Splice the two to get the view position vector
[0075] S3.3. According to the viewing angle position vector ξ b,u , for the single-view channel feature vector h b,u The following form of linear modulation is performed to embed the spatial location information of the user and base station into the channel characteristics:
[0076]
[0077] where γ p and β p Both are 4·(2d p +1)→d model Fully connected layer.
[0078] S3.4. Group all single-view channel features after embedding view position information into a set Input multi-view channel converter MVTε, output Right now
[0079]
[0080] Among them, the structure of MVT is as follows Figure 2 As shown on the left, it uses a Transformer encoder structure, extracts the correlation between channel features from different perspectives through a multi-head self-attention layer, and performs nonlinear feature transformation through a feedforward neural network. In this case, the self-attention layer uses 8 attention heads, and the feedforward neural network structure is: fully connected - LeakyReLU activation - fully connected, with the intermediate hidden layer dimension of 2·d model Output The dimension is d model .
[0081] S3.5. Output Perform average pooling to obtain a dimension of d model The multi-view channel feature v, that is,
[0082]
[0083] S3.6. Pass the multi-view channel feature v through two MLPs to predict the mean μ respectively z and variance σ z , and then the characteristic z of the target scatterer is obtained by sampling through the parameter renormalization method,
[0084]
[0085] Used to control the generation of scatterers in the subsequent step S4. In this case, the structure of the MLP is: full connection - layer normalization - ReLU activation - full connection - layer normalization - ReLU activation - full connection, the intermediate hidden layer dimensions are 128 and 64 respectively, and the output dimension is d z = 128. Variance σ z Make predictions in the log domain.
[0086] We denote the neural network composed of steps S3.1-S3.6 as the multi-view channel encoder All learnable neural network parameters are recorded as The entire multi-view channel coding process can be summarized as follows:
[0087]
[0088] For step S4, point cloud is used to describe the scatterer target, and the point cloud is represented as a set Where M is the total number of shape-electromagnetic points, in this case M = 1000. The point i in the point cloud is characterized by normalized shape-electromagnetic parameters. For the two-dimensional scene in this case, the shape dimension d of the point cloud is s =2, and the electromagnetic properties of the scatterer include relative permittivity and conductivity. The electromagnetic property dimension of the point cloud is d EM = 2. Therefore, the scatterer point cloud consists of the following 4-dimensional points:
[0089]
[0090] It is normalized based on the xy coordinates of all points in the dataset and the mean and standard deviation of the relative permittivity and conductivity values. Modeled as a point distribution The result of M independent sampling is The shape and electromagnetic parameters of the point cloud are controlled by the scatterer feature z.
[0091] The specific scheme of step S4 is as follows:
[0092] S4.1 Define the forward diffusion process:
[0093] In the forward diffusion process, the original point cloud distribution gradually diffuses into the noise distribution through a predefined noise addition process. The forward diffusion process is modeled as a Markov chain in the following form, where the transition probability of each step adopts a Gaussian distribution.
[0094]
[0095] The time step t=1,...,T, T is the maximum time step in the diffusion process, β1,...,β T is a predefined noise variance plan used to control the rate of the diffusion process. In this case, T = 100, the noise variance β t It increases linearly from 0.0001 to 0.02. Further, according to the above formula, it can be derived that from the initial state Directly at any time step t Sampling method. Definition have
[0096]
[0097] in,
[0098] S4.2 Construct the reverse diffusion process:
[0099] According to Bayes' theorem, it can be derived from the forward diffusion process
[0100]
[0101] in,
[0102]
[0103] However, It is difficult to handle and it is impossible to obtain the Markov chain of the analytical form of the back diffusion process. Therefore, we construct a learnable Markov chain conditioned on the feature z to implement the back diffusion process, where the learnable parameter is θ and the input point is from the noise distribution Sampling to approximate the distribution Right now
[0104]
[0105] And the transition probability is modeled as a Gaussian distribution of the following form,
[0106]
[0107] in is the noise variance at time step t during the back diffusion process. Mean μ θ Has the following form,
[0108]
[0109] ∈ θ is a noise estimation network used to implement the point cloud diffusion model. In this case, ∈ θ It is composed of a series of concatsquash layers, and the operation of each layer is
[0110]
[0111] in, and For the noise estimation network ∈ θ The feature vectors and conditional vectors of the input and output of layer l ζ (t) is the time encoding of the diffusion model, which is in the form of The dimension d of the time encoding t =32. W1, W2, W3, b1, b2 are all learnable parameters, and σ is the Sigmoid activation function. ∈ θThe dimensions of each layer are 4-128-256-512-256-128-4, and LeakyReLU activation is used between every two layers.
[0112] S4.3 Construct training objectives:
[0113] To train the model to implement the back-diffusion process, we need to maximize the log-likelihood Where Θ represents all learnable parameters. For a single data sample pair For , the evidence lower bound (ELBO) of its log-likelihood function can be derived as
[0114]
[0115] (a) is because the reconstruction process of the diffusion model is only related to z, and uses a learnable prior distribution p φ (z) to approximate On this basis, according to the independence of points in the point cloud, the ELBO on all data samples can be derived as
[0116]
[0117] Specify the optimization objective as minimizing an upper bound on the negative log-likelihood:
[0118]
[0119] It can be further transformed into
[0120]
[0121] in is the KL divergence of two deterministic Gaussian distributions, which has nothing to do with the training parameters. It can be converted into KL divergence between two Gaussian distributions
[0122]
[0123] Item 3 Can be converted into the same form for merging. z It can be regarded as a regular term, which makes the distribution of the scatterer feature z have a certain smooth continuity in the high-dimensional latent space. φ (z) is constructed by normalizing the flow, which is obtained by a trainable bijective function F φ Map the Gaussian distribution to a complex distribution, so the target distribution can be obtained by variable substitution
[0124]
[0125] in p w is the probability density function of the standard Gaussian distribution. Therefore,
[0126]
[0127] Furthermore, we can improve the training efficiency by sampling any time step for gradient descent and omit the weight difference of different time steps. The simplified loss function is
[0128]
[0129] in Represents the multi-view CSI and corresponding scatterer point cloud containing the position information of the transmitter and receiver in the training set,
[0130] In order to facilitate the adjustment of the regularization term coefficient during the optimization process and balance the importance of the geometric shape dimension and electromagnetic property dimension of the scatterer point cloud during the reconstruction process, we simple Further modifications were made to introduce a shape-electromagnetic weighted diffusion loss function as the final optimization objective:
[0131]
[0132] in,
[0133]
[0134] Represents the multi-view CSI and corresponding scatterer point cloud containing the position information of the transmitter and receiver in the training set, The subscripts s and EM represent the first two and last two dimensions of the vector, respectively, representing the shape and electromagnetic properties of the scatterer. s , γ EM , γ z are the weights of point cloud shape, electromagnetic properties, and regularization term respectively. p w is the probability density function of the standard Gaussian distribution. Here, the normalized flow model F φ It is composed of a series of affine coupling layers stacked together to meet the requirements of bijection and facilitate the solution of the determinant of the Jacobian matrix.
[0135] Specifically, we randomly select multi-view channels and corresponding scatterer point clouds from the training set Randomly select time steps Randomly select noise vector First Input the multi-view channel encoder and obtain the scatterer features according to step S3 sampling The feature z and time step t are then fed into the noise estimation network ∈ θ , perform the gradient descent step
[0136]
[0137] Through back-propagation, the multi-view channel encoder parameters are And point cloud diffusion model parameters θ are optimized. In this case, we take the coefficient γ s =0.45,γ EM =0.05. Adam optimizer is used, and the initial learning rate is set to 10 -4 , every 10 5 After optimization, the learning rate is decayed to the original 0.8, and a total of 10 optimizations are performed. 6 Second-rate.
[0138] S4.4 Reasoning stage:
[0139] After training, the reconstruction effect and performance of the model can be tested on the test set. First, the multi-view channel encoder is used to extract the multi-view channel. The corresponding scatterer feature z,
[0140]
[0141] Then, we control each step of the back-diffusion model by z, so as to gradually reverse-diffusion the noise point cloud into a meaningful scatterer point cloud. (The 4-D coordinates of each point in the noise point cloud are independently sampled from the standard normal distribution.) Specifically, we first sample from the noise distribution Then for t=T,…,1, perform the following iterative sampling:
[0142]
[0143] Final output reconstruction point Perform the above sampling process M times in parallel to obtain the reconstructed point cloud Complete scatterer imaging.
[0144] Through computer simulation, it can be seen that under the control of the scatterer feature z extracted by the multi-view channel, the scatterer reconstruction process based on the point cloud diffusion model is as follows: Figure 3 This verifies that the proposed multi-view channel feature encoder can effectively extract the geometric shape, electromagnetic characteristics and other information of the scatterer from the multi-view channel data. The reconstruction results of some scatterers under different numbers of UEs and BSs are shown in Figure 4 As shown in the figure, the mean log-scale Chamfer distance (MLCD) of the reconstruction results of the test set is as follows Figure 5 This demonstrates that the proposed scatterer imaging scheme can effectively utilize the advantages of multi-view observation to improve perception accuracy.
[0145] In summary, the present invention designs a generative imaging method based on multi-perspective channel feature fusion. In an embodiment of the present invention, a multi-perspective channel encoder extracts scatterer features based on the multi-perspective CSI between multiple UEs and BSs, thereby controlling the point cloud diffusion model to complete the reconstruction of the scatterer shape and electromagnetic properties. The embodiment of the present invention can flexibly process channel data between UEs and BSs of any number and any position, and effectively extract scatterer features from multi-perspective observations to achieve high-quality reconstruction results, providing an efficient environmental imaging method for the design of future integrated perception and communication systems.
[0146] Corresponding to the aforementioned embodiment of a generative imaging method based on multi-view channel feature fusion, the present invention also provides an embodiment of a generative imaging device based on multi-view channel feature fusion.
[0147] See also Figure 6 An embodiment of the present invention provides a generative imaging device based on multi-view channel feature fusion, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a generative imaging method based on multi-view channel feature fusion in the above embodiment.
[0148] The embodiment of a generative imaging device based on multi-view channel feature fusion provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 6 As shown, it is a hardware structure diagram of any device with data processing capability where a generative imaging device based on multi-view channel feature fusion provided by the present invention is located. Figure 6 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0149] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0150] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.
[0151] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a generative imaging method based on multi-view channel feature fusion in the above embodiment is implemented.
[0152] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0153] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the generative imaging method based on multi-view channel feature fusion.
[0154] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered merely as exemplary, and the true scope and spirit of the present application are indicated by the claims.
[0155] It should be understood that the above general description and the detailed description that follows are exemplary and explanatory only and do not limit the present application. The present application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes may be made without departing from the scope of the present application. The scope of the present application is limited only by the appended claims.
Claims
1. A generative imaging method based on multi-view channel feature fusion, characterized in that: The method comprises the following steps: S1. Each base station estimates the uplink CSI between the base station and each user through the pilot signal sent by the user; S2. Upload all CSI data estimated by each base station to the central server and merge them into CSI of all transmitting and receiving perspectives; S3. Build a multi-view channel encoder, embed the spatial location information of the user and base station into the channel features, extract the multi-view channel features and perform feature fusion to obtain the characteristics of the target scatterer; S4. Build a point cloud diffusion model. The point cloud diffusion model controls the point cloud diffusion process through the characteristics of the target scatterer, and gradually reversely diffuses the noise point cloud into a meaningful scatterer point cloud. S5. Use the shape-electromagnetic weighted diffusion loss function as the optimization objective to train the multi-view channel encoder and point cloud diffusion model. Use the trained model to process the CSI of all transmitting and receiving perspectives to obtain a point cloud that describes the shape and electromagnetic characteristics of the target scatterer, thereby achieving imaging of the scatterer.
2. A generative imaging method based on multi-view channel feature fusion according to claim 1, characterized in that: The uplink CSI between the base station and each user is estimated as follows: Single-view channel data is obtained by measuring the channel between each user and the base station or using a deterministic channel simulation model. The single-view channel data set includes a CSI matrix, base station coordinates, and user coordinates. The use of a deterministic channel simulation model to obtain single-view channel data includes: dividing the area of interest into pixels, obtaining the contrast of the pixelated scatterer including the scatterer shape and electromagnetic properties, and simulating the space-frequency domain channel between each base station and the user through the method of moments.
3. The generative imaging method based on multi-view channel feature fusion according to claim 2, characterized in that: The embedding of the spatial position information of the user and the base station into the channel features specifically includes: flattening the CSI of each channel into a vector and mapping it to a single-view channel feature vector of uniform dimension through a fully connected layer, position encoding the base station coordinates and the user coordinates to obtain the position vectors of the base station and the user and splicing them; and embedding the spliced position vectors into the channel features through linear modulation.
4. The generative imaging method based on multi-view channel feature fusion according to claim 3, characterized in that: The position encoding of the base station coordinates and the user coordinates is specifically as follows: Map each coordinate component from a scalar to a (2d p +1)-dimensional vector where p is the position vector, d p is the coding frequency number.
5. The generative imaging method based on multi-view channel feature fusion according to claim 3, characterized in that: The linear modulation is specifically: where γ p and β p Both are 4·(2d p +1)→d model Fully connected layer, model width d model is the dimension of the channel feature vector, ξ b,u is the single view position vector.
6. The generative imaging method based on multi-view channel feature fusion according to claim 1, characterized in that: The method of extracting multi-view channel features and performing feature fusion is specifically as follows: a single-view channel feature vector set is input into a multi-view channel converter, the multi-view channel converter adopts a Transformer encoder structure, extracts the correlation between different view channel features through a multi-head self-attention layer, performs nonlinear feature transformation through a feedforward neural network, and then obtains multi-view channel features after output average pooling. The multi-view channel features are input into two MLPs to predict the mean and variance respectively, and the characteristics of the target scatterer are obtained after parameter resampling.
7. The generative imaging method based on multi-view channel feature fusion according to claim 1, characterized in that: The point cloud diffusion model is specifically: The original point cloud distribution is gradually diffused into noise distribution through a predefined noise addition process and is modeled as a Markov chain to obtain the distribution Construct a learnable Markov chain based on the feature z output by the multi-view channel encoder to realize the diffusion process of gradually reversely diffusing the noise point cloud into the scatterer point cloud. The input point is distributed from the noise Sampling to approximate the distribution Right now And the transition probability is modeled as a Gaussian distribution of the following form, in is the noise variance at time step t during the back diffusion process, μ θ is the mean value estimated by the point cloud diffusion model with parameter θ, μ θ Has the following form, where ∈ θ A noise estimation network that progressively denoises noisy samples based on the noise scheme of the point cloud forward diffusion process; The dimension of the point cloud is the sum of the shape dimension and the electromagnetic property dimension, where the electromagnetic property dimension includes the relative dielectric constant and conductivity of the scatterer; each point cloud is modeled as the result of multiple independent sampling of a point distribution, and the number of independent samplings is the total number of shape-electromagnetic points; the shape and electromagnetic parameters corresponding to the point cloud are controlled by the characteristics of the scatterer.
8. The generative imaging method based on multi-view channel feature fusion according to claim 1, characterized in that: The shape-electromagnetic weighted diffusion loss function is specifically: in, Multi-view channel encoder The learnable neural network parameters are Represents the multi-view CSI and corresponding scatterer point cloud containing the position information of the transmitter and receiver in the training set, Expressed as a multi-view channel coding process, Noise Vector The subscripts s and EM represent the dimensions of the shape attribute and the electromagnetic attribute in the selection vector, respectively; γ s , γ EM , γ z are the weights of point cloud shape, electromagnetic properties, and regularization term respectively; β t is the noise variance, is the initial state of the point cloud, p w is the probability density function of the standard Gaussian distribution; the normalized flow model F φ It is composed of a series of affine coupled layers stacked to meet the requirements of bijection.
9. A generative imaging device based on multi-view channel feature fusion, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, it implements a generative imaging method based on multi-view channel feature fusion according to any one of claims 1 to 8.
10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a generative imaging method based on multi-view channel feature fusion according to any one of claims 1 to 8 is implemented.