A Wi-Fi signal indoor imaging method based on improved VAE-MoE architecture
Through the improved VAE-MoE architecture, combined with WiFi signals and deep learning technology, the problem of insufficient precision and accuracy in indoor WiFi signal imaging is solved, efficient indoor image generation and real-time monitoring are achieved, and the accuracy of security monitoring and the performance of IoT devices are improved.
Patent Information
- Application Number
- CN202411932776.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing indoor Wi-Fi signal imaging technology has limited precision and accuracy in complex environments. Traditional methods have difficulty effectively dealing with the sparsity and noise interference of Wi-Fi signals, and their propagation characteristics are complex. Traditional image processing methods are unable to cope with challenges such as multipath effects and signal attenuation.
An improved VAE-MoE architecture is adopted, combining VAE modules, MoE modules and Self-Attention. WiFi signal RSSI data and RGB image data are obtained through path planning. Distance attenuation compensation and sparse signal processing are performed. The volume integral wave equation and Rytov approximate linearization model are constructed, and high-precision images are generated by combining deep learning model training.
It significantly improves the precision and accuracy of indoor imaging, can effectively remove noise and interference, enhance image detail retention and global structural consistency, improve the accuracy of indoor real-time monitoring and security monitoring, and provide low-cost indoor maps and object positioning information.
Smart Images

Figure CN119991840B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication technology, and in particular relates to a method for indoor imaging of WIFI signals based on an improved VAE-MoE architecture. Background Art
[0002] Real-time monitoring and analysis of indoor environments are crucial for applications such as smart homes, security surveillance, indoor navigation, and the Internet of Things. Traditional indoor detection methods, such as LiDAR, infrared detection, and ultrasonic detection, are often limited by expensive equipment and, in practical applications, by installation and environmental adaptability requirements. In contrast, indoor detection based on wireless communication signals (such as Wi-Fi) offers broad application prospects due to its wide applicability, low cost, and high signal flexibility.
[0003] Wi-Fi signals, as a mature wireless communication technology, are now widely used in fields such as indoor positioning. The rapid development of deep learning in recent years has led to the emergence of deep learning-based Wi-Fi signal imaging methods, which have achieved remarkable results in image recognition and processing. The feature extraction capabilities of deep learning can significantly improve the processing of Wi-Fi signal data, thereby enhancing the accuracy and precision of indoor imaging. However, the sparsity and complexity of Wi-Fi signals in practical applications, as well as the need to effectively preprocess and optimize the data to remove noise and interference, present challenges. Furthermore, the propagation characteristics of Wi-Fi signals are closely related to the physical environment. Combining physical models with deep learning methods to simulate signal propagation and improve the precision and accuracy of imaging results remains a hot topic. Existing Wi-Fi signal indoor imaging technologies are susceptible to interference in complex environments, such as multipath, wall obstruction, and signal attenuation. Traditional Wi-Fi signal received signal strength (RSSI) analysis methods are susceptible to interference, resulting in limited precision and accuracy in imaging results. Furthermore, the penetrating and reflective characteristics of Wi-Fi signals further complicate indoor imaging, making traditional image processing methods unable to fully address these challenges. To this end, developing a WIFI signal indoor imaging method based on the improved VAE-MoE architecture is the key to solving the above problems. Summary of the Invention
[0004] To solve the above technical problems, the purpose of the present invention is to provide a method for indoor Wi-Fi signal imaging based on an improved VAE-MoE architecture. By introducing a VAE module, a MoE module, and Self-Attention, the quality of generated images based on Wi-Fi signal RSSI data and RGB image data is successfully improved. Self-Attention effectively captures the long-range dependency information in the image, making the generated image more accurate.
[0005] The object of the present invention is achieved by comprising the following steps:
[0006] S100, Path Planning and Data Acquisition: A Wi-Fi transmitter is installed on one mobile device, and a Wi-Fi receiver is installed on another mobile device. The two mobile devices are located on opposite sides of the building to be tested and simultaneously move around the perimeter of the building to be tested in a clockwise or counterclockwise direction. During the movement, the Wi-Fi signal emitted by the Wi-Fi transmitter is received by the Wi-Fi receiver, and time-synchronized Wi-Fi signal RSSI data and the location data of the mobile device are obtained; an RGB camera is used to capture indoor RGB images;
[0007] S200, Data Preprocessing: Preprocess the raw Wi-Fi signal data, including distance attenuation compensation and sparse signal processing, to ensure data integrity and support precision imaging. It should be noted that Wi-Fi signal strength typically decreases with increasing signal propagation distance, especially in indoor environments where the presence of objects such as walls and furniture can cause signal attenuation. During Wi-Fi signal processing, the received signal may be interfered with by noise or other non-ideal factors, resulting in degraded signal quality. To ensure that the received Wi-Fi signal data is suitable for subsequent physical modeling and high-precision imaging tasks, distance attenuation compensation and sparse signal processing are effective methods for preprocessing the received RSSI data.
[0008] S300, Simulating Signal Propagation: Constructing a volume-integrated wave equation for the preprocessed data to establish the relationship between signal propagation and object distribution. This step introduces the volume-integrated wave equation, which considers the influence of all scatterers in the target area and provides the necessary physical basis for accurate modeling of signal propagation from the transmitter to the receiver and accurate imaging of Wi-Fi signals in the present invention.
[0009] S400, constructing a linearized physical model: Based on the volume-integrated wave equation, Rytov approximation linearization is introduced and physical modeling is performed by discretizing the volume-integrated wave equation and the RSSI simplified model. Based on the preliminary processing based on the volume-integrated wave equation, Rytov approximation is further introduced for linearization. The corresponding physical modeling process is completed by discretizing the volume-integrated wave equation and combining it with the RSSI simplified model.
[0010] By introducing Rytov linear approximation and sparse signal processing technology, the present invention can effectively extract useful features from the signal, remove noise and interference, and thus achieve more accurate indoor area image reconstruction;
[0011] S500, Deep Learning Model Training and Image Generation: This module extracts latent features from Wi-Fi signal RSSI data and RGB image data using an MLP, converting them into high-dimensional feature vectors of fixed dimensions. The VAE encoder module then encodes the multimodal input data, generating the mean and variance of the latent space. The MoE module then performs a weighted combination of the latent space. The generator module then generates an RGB image, which is then judged for authenticity using a discriminator module. The generator is then trained and optimized to output a high-precision image. This module effectively processes complex multimodal input data (Wi-Fi signal and image data), ensuring full data utilization and improving indoor imaging accuracy in deep learning, particularly in terms of detail preservation and global structural consistency.
[0012] Among them, in step S100, the mobile device can use a self-programmed unmanned vehicle, equipped with an antenna and precise positioning, to efficiently plan the route of the unmanned vehicle to collect WIFI signal RSSI data, collect multi-perspective data, improve data collection efficiency, and ensure data integrity; this self-programmed unmanned vehicle can use a commercially available unmanned vehicle, equipped with a positioning antenna, a PC interface, and draw the edge shape and perimeter of the unknown house to be measured in the form of lines on the PC interface, pre-set the planned route, and receive the driving path information of the two unmanned vehicles through the positioning antenna when the unmanned vehicle drives according to the planned route, ensuring that the positioning antenna is accurately aligned, and performing precise positioning on the PC interface to improve the accurate recording of WIFI signals and the high-quality generation of images.
[0013] The moving path of the mobile device can be that one mobile device is located on the south side of the building to be measured, and the other is located on the north side of the building to be measured. The two mobile devices move clockwise at the same time, covering different parts of the path to improve the collection efficiency.
[0014] Preferably, in step S100, the mobile device moves and collects data according to multiple layers of moving paths. Specifically, the initial layer moving path of the mobile device is 0.5 meters away from the outer wall of the building to be measured. The initial layer movement is completed by circling the building to be measured. Thereafter, a layer of moving path is set every 1 meter in the direction away from the initial layer, and each layer of moving path circles the building to be measured. After each layer of moving path is completed, each mobile device scans the target area again at the 0° starting point, 45° point, 90° point and 135° point of the layer path to collect wall-penetrating signal characteristics at different angles.
[0015] Preferably, in step S100 , the mobile device collects RSSI data and corresponding location data every time it moves 0.1 to 0.3 meters to ensure data integrity and consistency.
[0016] Preferably, the distance attenuation compensation in step S200 is specifically: measuring the original RSSI measurement value of the WIFI signal under unobstructed conditions and without any obstructions between the two unmanned vehicles, and obtaining a distance-related attenuation model of the signal in an open space; the formula is:
[0017] ,
[0018] in, : Received signal strength (RSSI) raw measurement value; :Derive distance-dependent power attenuation value based on obstacle-free path signal measurement; : Remove the signal after distance attenuation for subsequent modeling and imaging.
[0019] The distance attenuation formula is derived from the free space path loss formula:
[0020] ,
[0021] Where d is the distance the signal travels; is the signal wavelength; this formula is based on the free-space path loss model of wireless signal propagation, and its principles and formulas are public and widely used in communication theory.
[0022] These obstacle-free path data are used to build a reference model and estimate the distance attenuation portion, which generally does not include the characteristics caused by obstacles and only reflects the power loss of the signal in free space.
[0023] Preferably, the sparse signal processing in step S200 is to organize and improve the data before model building to provide high-quality input for subsequent high-definition modeling. This processing step combines physical modeling with signal optimization and is one of the core components of the entire imaging process. The formula is:
[0024] ,
[0025] Among them, TV(R): as the total variation of the target area, : local change of the i,jth point; Y: received signal strength vector; K: observation matrix, generated based on physical modeling; X: target sparse feature vector.
[0026] This optimization formula and technique originate from the basic theory of sparse signal processing and are classic methods in the field of compressed sensing. They are public and widely used. Total variation minimization is common in the fields of image processing and signal optimization. Related references include:
[0027] Rudin-Osher-Fatemi (ROF) model: a total variation method for image denoising;
[0028] The basic formula of compressed sensing:
[0029] ,
[0030] Preferably, the volume integral wave equation in step S300 is:
[0031] ,
[0032] Where E(r) is the electric field received at position r, indicating the actual received WIFI signal strength; E inc (r) is the incident electric field, that is, the signal propagated in the absence of obstacles; G(r,r') is the Green function, which describes the propagation characteristics of the electromagnetic wave from position r' to position r; O(r') is the physical property at position r' in the target area, such as the dielectric constant or reflectivity of the scatterer; E(r') is the electric field generated by each scatterer in the target area; dv' is the integral of each scattering point in the target area D.
[0033] The volume integral wave equation is a classic physical model for electromagnetic wave propagation and scattering. It combines the incident electric field and the scattered electric field caused by obstacles to describe the propagation of Wi-Fi signals. Through this equation, we can predict the signal strength received at a certain location, which is crucial for subsequent signal processing, physical modeling, and imaging tasks. A brief explanation of the above formula: The electric field E(r) is composed of two parts: the incident electric field E inc (r): The original transmitted signal, the ideal signal in the absence of obstacles; the scattered electric field (via the integral term): This portion of the electric field is caused by scatterers within the target area (such as walls and pillars). This term considers the influence of each scatterer within the target area and integrates the electric field contribution of each scattering point to ultimately obtain the composite scattered signal.
[0034] Preferably, the physical modeling in step S400 is specifically:
[0035] S401, perform Rytov approximation, the formula is:
[0036] ,
[0037] in, (r) is the phase change of the signal during propagation. This linearization process effectively reduces the nonlinear calculation in the signal propagation model and improves the modeling efficiency;
[0038] S402, establish a discretized volume integral wave equation, the formula is:
[0039] ,
[0040] in, is the phase change discrete vector, F is the propagation matrix, which represents the propagation relationship of the signal from the scattering source to the receiving point, and O is the physical characteristic vector of the target area. Through this discretization process, numerical optimization and calculation can be performed more efficiently, which facilitates subsequent data processing and image reconstruction.
[0041] The specific discretization process is:
[0042] ,
[0043] The target area D is divided into N discrete voxels, and the center point of each voxel is r n , the integral becomes the summation; Φ(r): phase change; g(r,r n ): Simplified propagation function; O(r n ): scattering properties of each voxel in the target area;
[0044] S403, RSSI simplified model. In order to combine the actual WIFI signal data, the RSSI simplified model is adopted. The signal model after volume integral wave equation and Rytov approximation is matched with the received signal strength RSSI of the WIFI device. The simplified model is expressed as: ,
[0045] Among them, P ryt Indicates the received signal strength change, F R is the real part of the propagation matrix, O R It is the real part of the physical properties of the target area, simplifies the signal strength representation, and combines it with the actual measurement data to improve the accuracy of signal modeling. By simplifying the model, the physical modeling results can be combined with the actual measured RSSI data to provide data input for subsequent deep learning algorithms, and ultimately generate high-precision indoor images.
[0046] To simplify calculations and improve processing efficiency, step S400 of the present invention uses the Rytov approximation to linearize the volume-integrated wave equation. The Rytov approximation simplifies the calculation process by converting the originally nonlinear wave equation into a linear form. During physical modeling, the volume-integrated wave equation is further discretized and converted into a matrix form for numerical calculation.
[0047] Step S400 introduces Rytov approximate linearization processing and discretized volume integral wave equation, combined with a simplified RSSI model, which can significantly improve the accuracy of WiFi signal propagation modeling in complex indoor environments and enhance the model's adaptability to wall-penetrating signals.
[0048] Preferably, in step S500, the physical model and the deep learning model are integrated to process the RSSI data and generate a high-precision image by constructing a neural network. The deep learning model training and image generation are specifically as follows:
[0049] S501. Extract potential features from RSSI data through MLP: The MLP network is used to enhance data representation. After convolution and pooling layers, the convolution kernel size is 3×1. After processing, a fully connected layer is designed for high-dimensional feature mapping. Finally, a fixed-length high-dimensional feature vector is output as the input for subsequent VAE and MoE.
[0050] The RGB image is processed through MLP to extract latent features: the input RGB image is reduced in resolution, the channel size is increased to 512 using successive convolutional blocks, each channel is normalized by mean and standard deviation, and LeakyReLU activation is performed on each layer. After the convolutional layer, it is normalized and converted into a latent space vector, which represents the spatial information and texture features in the image;
[0051] S502, VAE encoder module processing: Use the VAE encoder (Variational Autoencoder) to encode the multimodal features of the input WiFi signal RSSI data and RGB image data to generate the mean and variance of the latent space. The mean and standard deviation represent the Gaussian distribution of the latent space and are used to describe the distribution of the latent variable Z. This process uses the mean and variance as input, and the VAE can sample the latent space and obtain the latent vector Z. This step ensures that the generated latent vector is continuous and smooth.
[0052] Among them, the variational inference and reconstruction loss formula of the VAE encoder is:
[0053] Variational inference formula:
[0054] ,
[0055] Where, given input data x, learn a distribution of latent variables z, with the goal of maximizing the marginal log-likelihood; p(x|z) is the generative model (such as the decoder network); q(z|x) is the variational distribution (encoder network); D KL [.||.] is the Kullback-Leibler divergence, which measures the difference between the variational distribution and the true posterior distribution;
[0056] Reconstruction loss:
[0057] ,
[0058] The reconstruction loss in VAE usually uses mean square error (MSE) or cross entropy, depending on the nature of the data; x is the original image;
[0059] S503, MoE module processing: Dynamically select multiple expert networks to perform weighted combination of input latent vectors to generate more complex and diverse images. The specific process is:
[0060] S5031. Input latent space vector z': The input latent space vector z' is obtained from the VAE encoder and represents the feature representation of the WiFi signal RSSI data and the RGB image data;
[0061] S5032, MoE weighting: Automatically adjust each expert's contribution based on the input features. The formula is:
[0062] ,
[0063] The output of the above MoE is a weighted sum, w i (z') is the weight calculated by the gating network and satisfies ∑w i (z') = 1 (usually obtained through softmax activation); given the latent space vector z', the MoE module determines the weights of each expert through the expert selection mechanism; it is set to the output of the i-th expert network, and the weight {w i (z')} comes from a gating network (usually a small neural network); the generated latent vector {f i (z')};
[0064] It should be noted that the MoE (Mixture of Experts) module uses multiple expert networks to perform a weighted combination of latent space vectors to generate more complex and diverse images. Each expert network focuses on a certain part of the latent space and can extract different feature information. The weighted combination generates the final latent space vector, enhancing the expressiveness of image generation. The purpose of weighting is to automatically adjust the contribution of each expert based on the input features, making the generation process more flexible and targeted.
[0065] It should be noted that the Gating network determines which experts should participate in data processing by assigning weights to each expert, ensuring that the model can dynamically select the most appropriate expert for different input data. The Gating network works based on the softmax activation function to calculate the weight of each expert. In this way, the MoE module enhances the model's expressiveness and flexibility, enabling it to better adapt to diverse input data. It is defined as:
[0066] ,
[0067] Among them, g i (z) is the output of the gating network, indicating the importance of each expert;
[0068] S504, generator module processing: converting the potential vector z output by the VAE encoder and MoE module into an indoor area image;
[0069] S505, discriminator module processing: judge the authenticity of the generated image, evaluate the quality of the generated image, and provide feedback to the generator to optimize the generation result;
[0070] S506, output high-precision image: After the discriminator feedback optimization, output a high-precision image.
[0071] For training and evaluation, according to the unmanned vehicle path planning, sampling is performed at a fixed step size to obtain time-synchronized RSSI data and image datasets, obtain Wi-Fi data packets and images, align the images and Wi-Fi data packets, and divide the data into training sets, validation sets, and test sets.
[0072] Preferably, the specific process of the generator module in step S504 is:
[0073] S5041, fully connected layer processing: first, the latent vector z is mapped to a higher-dimensional feature space to ensure that the input latent vector can be fully processed in the generator;
[0074] S5042, Self-Attention Layer Processing: Through self-attention, the generator dynamically focuses on different parts of the image generation process, calculates the similarity between each part, and adjusts the weighting of each part. The self-attention mechanism is used to capture global dependencies in the image and improve the details of the generated image. Self-attention plays a key role in image generation, enhancing the generator's modeling of long-range dependencies and capturing global information in the image.
[0075] S5043, deconvolution layer (transposed convolution) processing: gradually decode the processed features into a high-precision image and restore the image's spatial structure;
[0076] S5044, output layer: Generates the final image through convolution. The output image contains the visual information of the indoor scene.
[0077] In step S504, the present invention introduces a self-attention mechanism to enhance the detail and consistency of the generated image. Self-attention captures global dependencies within the image, making the generated image more detailed and avoiding the problem of the generator focusing only on local information while ignoring global information.
[0078] Preferably, the specific process of the self-attention layer processing in step S5042 is:
[0079] (1) Query, key and value:
[0080] ,
[0081] The shape of the input feature matrix X is N×D, where N is the input feature vector and D is the dimension of each feature; Q 、W K , and W V is a learnable weight matrix corresponding to the conversion of query, key and value respectively;
[0082] (2) Calculation of attention weight:
[0083] ,
[0084] The above is to calculate the dot product between the query and the key to obtain the attention weight matrix, where is a scaling factor used to prevent the gradient from disappearing due to excessive dot product; the weight is then calculated using the softmax function, and the calculation formula is as follows:
[0085] ,
[0086] in is the attention weight of position i to position j, and the output is weight V;
[0087] (3) Weighted summation:
[0088] ,
[0089] The final output is the weighted sum of the values V by the attention weights, where α is the attention weight matrix and V is the value matrix.
[0090] Preferably, in step S505, the discriminator module includes multiple convolutional layers and self-attention layers to improve the accuracy and stability of the discriminator. The convolutional layers are responsible for extracting features from the image, while the self-attention layers help the discriminator capture the global structure in the image. The specific processing process of the discriminator module is as follows:
[0091] S5051, multi-layer convolution layer processing: extract features from the image, the convolution operation is expressed as:
[0092] ,
[0093] Among them, H (l) is the feature map of layer l, W (l) is the convolution kernel, b (l) is the bias term, * represents the convolution operation;
[0094] S5052, Self-Attention Layer Processing: Self-Attention is introduced. Self-Attention can help the discriminator better understand the long-range dependencies in the image, improve sensitivity to image details, and enhance the accuracy of image authenticity judgment. The calculation formula is: ,
[0095] Among them, Q, K and V are the matrices of query, key and value respectively, representing different transformations of input features, and softmax is used to calculate attention weights;
[0096] S5053, Fully Connected Layer processing: After processing by the convolutional layer and the self-attention layer, the discriminator flattens the feature map and inputs it into the fully connected layer to map the high-level features of the image to the final classification space (0 or 1), that is, to determine whether the image is real or generated, expressed as:
[0097] ,
[0098] Where W is the weight matrix, b is the bias term, and σ is the activation function (usually sigmoid), which is used to output a value between 0 and 1, indicating the true probability of the image.
[0099] S5054, output layer: outputs the true and false judgment results through the sigmoid activation function, 0 is the generated image, and 1 is the real image. The formula is as follows:
[0100] ,
[0101] Among them, p real Represents the probability that the input image is a real image. If p real If it is close to 1, it means the image is more likely to be real; if it is close to 0, it means the image is more likely to be generated.
[0102] The discriminator consists of multiple layers of convolutional layers and self-attention layers. The discriminator module is responsible for judging the authenticity of the generated image. Its goal is to determine whether an image comes from the real data distribution (i.e., the real image) or from the generator (i.e., the generated image). It participates in the image generation process as an auxiliary module, evaluates the quality of the generated image, and provides feedback to the generator to further optimize the generation results.
[0103] Preferably, in the step of generating an image output in S506, the activation function in the neural network architecture adopts LeakyReLU (leaky rectified linear unit) for encoding the image and K-divergence loss (KLDivergence Loss) to ensure that the latent space distribution conforms to the standard normal distribution; the optimizer adopts Adam optimizer for complex deep neural networks;
[0104] The Leak ReLU formula is as follows:
[0105] ,
[0106] Where, x: input value, : A small constant, usually between [0,1], with a common default value of 0.01;
[0107] The KL divergence loss formula is as follows:
[0108] ,
[0109] Where μ and σ are the mean and standard deviation output by the VAE encoder, and p(z) is the standard normal distribution;
[0110] The Adam optimizer formula is as follows:
[0111] ,
[0112] in, are model parameters, is the gradient momentum, v t is the gradient second moment estimate, is the learning rate, is a small constant to prevent division by zero.
[0113] After transforming the latent vector z (i.e., performing a transposed convolution and self-attention), the generator generates an image representing the indoor scene. This image generation not only takes into account the structural information in the latent space but also fully considers the image's global consistency and details, providing a more realistic image output.
[0114] Compared with the prior art, the present invention has the following technical effects:
[0115] 1. This invention deploys two groups of Wi-Fi signal devices and indoor cameras, collects RSSI information data using Wi-Fi signals, and combines physical modeling with deep learning technology for phased optimization, significantly improving data imaging accuracy. First, the volume integral wave equation and Rytov approximation theory are used to linearize and approximate the physical model. The raw Wi-Fi signal data is preprocessed, key features are extracted, and sparse signal processing is performed to remove noise and interference factors, laying the foundation for subsequent high-precision imaging. Secondly, a deep learning architecture is designed. Through a multilayer perceptron, a variational auto-encoder, and a mixture of experts, a self-attention mechanism is introduced. This generates high-precision indoor area images from data optimized by physical modeling, and further improves the performance of wall-penetrating signal modeling. During the Wi-Fi signal data acquisition process, the acquisition strategy is optimized through efficient path planning, effectively improving the integrity and accuracy of the data. This invention solves the indoor imaging problem in limited data environments by combining the physical model constructed using methods such as the volume integral wave equation and Rytov approximation theory with the neural network constructed using VAE and MoE modules and self-attention. The method has high technical innovation and practical application value.
[0116] 2. This invention introduces multimodal data processing to enhance the expressiveness of deep learning models: By using neural network architectures such as MLP, VAE modules, MoE modules, and Self-Attention, combined with RSSI data of WiFi signals and image data, it effectively processes complex multimodal input data, improving the expressiveness and detail retention of image generation;
[0117] 3. This invention improves the precision and accuracy of indoor area imaging by combining Wi-Fi signals with the neural network architecture in deep learning, especially the use of self-attention, which greatly improves the precision and accuracy of indoor area images;
[0118] 4. The processing and verification mechanism of the present invention further optimizes the quality of generated images by combining VAE modules and MSE modules, effectively removes noise during the generation process, and improves edge clarity and object recognition accuracy in images.
[0119] 5. This invention improves indoor real-time monitoring capabilities. By designing an integrity program, it can monitor indoor structures, personnel distribution, and other aspects in real time, and generate accurate images of indoor areas, greatly improving the performance and effectiveness of IoT devices in automated management and environmental perception.
[0120] 6. In terms of security monitoring capabilities, this invention utilizes RSSI data from Wi-Fi signals and deep learning models to penetrate walls and acquire indoor information, providing more accurate indoor object positioning, personnel tracking, and abnormal behavior detection for security monitoring. This technology can effectively overcome the limitations of traditional cameras in lighting, angles, and occlusion conditions, improving monitoring accuracy and reliability.
[0121] 7. The indoor images generated by the present invention are based on WIFI signals and deep learning technology. The imaging method is more flexible, low-cost and easy to deploy, providing accurate indoor maps and object positioning information, promoting the widespread application of intelligent perception and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0122] Figure 1 It is a schematic diagram of the process of the present invention;
[0123] Figure 2 Schematic diagram of the feature extraction by introducing the VAE module in the present invention;
[0124] Figure 3 Schematic diagram of the indoor image generation method performed by the generator module of the present invention;
[0125] Figure 4 Schematic diagram of a method for enhancing image precision and accuracy using a discriminator module according to the present invention;
[0126] Figure 5 Schematic diagram of the method for improving indoor image accuracy through self-attention mechanism;
[0127] Figure 6 This is an imaging effect diagram of the present invention. DETAILED DESCRIPTION
[0128] The present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention fall within the scope of protection of the present invention.
[0129] Example 1
[0130] As attached Figures 1 to 6 The embodiment shown in the figure is based on the improved VAE-MoE architecture WIFI signal indoor imaging method, which includes the following steps:
[0131] S100, Path Planning and Data Acquisition: A Wi-Fi transmitter is installed on one unmanned vehicle, and a Wi-Fi receiver is installed on another mobile device. The two unmanned vehicles are located on opposite sides of the building to be tested and move around the building in a clockwise or counterclockwise direction. During the movement, the Wi-Fi signal emitted by the Wi-Fi transmitter is received by the Wi-Fi receiver, and time-synchronized Wi-Fi signal RSSI data and the location data of the unmanned vehicle are obtained; an RGB camera is used to capture indoor RGB images;
[0132] S200, Data Preprocessing: Preprocess the raw Wi-Fi signal data, including distance attenuation compensation and sparse signal processing. In indoor environments, the presence of indoor structures such as walls and furniture can cause signal attenuation, and the received signal may be affected by noise or other non-ideal factors. Simultaneously, RGB images of the room are collected as subsequent training objects. Therefore, distance attenuation compensation and sparse signal processing are effective methods for preprocessing received RSSI data for Wi-Fi signal data.
[0133] S300, Simulating Signal Propagation: After the aforementioned processing, the input is a discretized vector. A volume integral wave equation is constructed for the preprocessed data to establish the relationship between signal propagation and object distribution. This equation comprehensively considers the influence of all scatterers within the target area, providing an accurate simulation of the signal propagation process from the transmitting source to the receiving point, and providing the necessary physical foundation for high-precision imaging of Wi-Fi signals.
[0134] S400, build a linearized physical model: Based on the volume-integrated wave equation, introduce Rytov approximate linearization and perform physical modeling by discretizing the volume-integrated wave equation and the RSSI simplified model;
[0135] S500, deep learning model training and image generation: Extract latent features from WiFi signal RSSI data and RGB image data through MLP, convert the WiFi signal RSSI data and RGB image data into high-dimensional feature vectors of fixed dimensions, then use the VAE encoder module to encode the multimodal input data to generate the mean and variance of the latent space. The MoE module then performs weighted combination of the latent space, and then generates an RGB image through the generator module. The discriminator module then determines the authenticity of the generated image, optimizes the generator through training, and finally generates a high-precision image.
[0136] Example 2
[0137] This embodiment is based on the improved VAE-MoE architecture WIFI signal indoor imaging method on the basis of Example 1. In step S100, the unmanned vehicle moves and collects data along multiple layers of moving paths. Specifically, the initial layer moving path of the unmanned vehicle is 0.5 meters away from the outer wall of the building to be measured. The initial layer movement is completed by circling the building to be measured. Thereafter, a layer of moving path is set every 1 meter in the direction away from the initial layer, and each layer of moving path circles the building to be measured. After each layer of moving path is completed, each mobile device scans the target area again at the 0° starting point, 45° point, 90° point and 135° point of the layer path to collect wall-penetrating signal characteristics at different angles.
[0138] Example 3
[0139] This embodiment is based on the improved VAE-MoE architecture WIFI signal indoor imaging method on the basis of Example 2. In step S100, the unmanned vehicle collects RSSI data and corresponding position data every time it moves 0.1 meter.
[0140] Example 4
[0141] This embodiment of the indoor Wi-Fi signal imaging method based on the improved VAE-MoE architecture is based on the third embodiment. The distance attenuation compensation in step S200 is specifically: under unobstructed conditions, with no obstructions between the two unmanned vehicles, the original RSSI measurement value of the Wi-Fi signal is measured to obtain a distance-dependent attenuation model of the signal in an open space. The formula is:
[0142] ,
[0143] in, : Received signal strength (RSSI) raw measurement value; :Derive distance-dependent power attenuation value based on obstacle-free path signal measurement; : Remove the signal after distance attenuation for subsequent modeling and imaging.
[0144] Example 5
[0145] This embodiment is based on the improved VAE-MoE architecture WIFI signal indoor imaging method based on Example 4. The sparse signal processing in step S200 is specifically: before model construction, the data is sorted and improved to provide high-quality input for subsequent high-definition modeling. The formula is:
[0146] ,
[0147] Among them, TV(R): as the total variation of the target area, : local change of the i,jth point; Y: received signal strength vector; K: observation matrix, generated based on physical modeling; X: target sparse feature vector.
[0148] Example 6
[0149] This embodiment is based on the improved VAE-MoE architecture WIFI signal indoor imaging method based on Example 5, step S300 volume integral wave equation:
[0150] ,
[0151] Where E(r) is the electric field received at position r, indicating the actual received WIFI signal strength; E inc (r) is the incident electric field, that is, the signal propagated in the absence of obstacles; G(r,r') is the Green function, which describes the propagation characteristics of the electromagnetic wave from position r' to position r; O(r') is the physical property at position r' in the target area, such as the dielectric constant or reflectivity of the scatterer; E(r') is the electric field generated by each scatterer in the target area; dv' is the integral of each scattering point in the target area D.
[0152] Example 7
[0153] This embodiment of the method for indoor WIFI signal imaging based on the improved VAE-MoE architecture is based on Example 6. The physical modeling in step S400 is specifically as follows:
[0154] S401, perform Rytov approximation, the formula is:
[0155] ,
[0156] in, (r) is the phase change of the signal during propagation; this linearization process effectively reduces the nonlinear calculation in the signal propagation model;
[0157] S402. Establish a discretized volume integral wave equation and convert it into a matrix form for numerical calculation. The calculation process is more efficient. The formula is:
[0158] ,
[0159] in, is the phase change discrete vector, F is the propagation matrix, which represents the propagation relationship of the signal from the scattering source to the receiving point, and O is the physical characteristic vector of the target area;
[0160] The specific discretization process is:
[0161] ,
[0162] The target area D is divided into N discrete voxels, and the center point of each voxel is r n , the integral becomes the summation; Φ(r): phase change; g(r,r n ): Simplified propagation function; O(r n ): scattering properties of each voxel in the target area;
[0163] S403, RSSI simplified model, the signal model after volume integral wave equation and Rytov approximation is matched with the received signal strength RSSI of the WIFI device, and the simplified model is expressed as: ,
[0164] Among them, P ryt Indicates the received signal strength change, F R is the real part of the propagation matrix, O R is the real part of the physical properties of the target area; this method provides data support for subsequent deep learning algorithms.
[0165] Example 8
[0166] This embodiment of the method for indoor WIFI signal imaging based on the improved VAE-MoE architecture is based on Example 7. The deep learning model training and image generation in step S500 are specifically as follows:
[0167] S501. Extract potential features from RSSI data through MLP: The MLP network is used to enhance data representation. After convolution and pooling layers, the convolution kernel size is 3×1. After processing, a fully connected layer is designed for high-dimensional feature mapping. Finally, a fixed-length high-dimensional feature vector is output as the input for subsequent VAE and MoE.
[0168] The RGB image is processed through MLP to extract latent features: the input RGB image is reduced in resolution, the channel size is increased to 512 using successive convolutional blocks, each channel is normalized by mean and standard deviation, and LeakyReLU activation is performed on each layer. After the convolutional layer, it is normalized and converted into a latent space vector, which represents the spatial information and texture features in the image;
[0169] S502, VAE encoder module processing: The VAE encoder consists of a convolutional layer, an MLP, and a fully connected layer. The final fully connected layer outputs the mean and standard deviation, which are generated by variational inference and sampled through reparameterization techniques. Specifically, the VAE encoder is used to encode the multimodal features of the RSSI data and RGB image data of the input Wi-Fi signal to generate the mean and variance of the latent space. The mean and standard deviation represent the Gaussian distribution of the latent space and are used to describe the distribution of the latent variable Z. The variational inference and reconstruction loss formula of the VAE encoder is:
[0170] Variational inference formula:
[0171] ,
[0172] Where, given input data x, learn a distribution of latent variables z, with the goal of maximizing the marginal log-likelihood; p(x|z) is the generative model (such as the decoder network); q(z|x) is the variational distribution (encoder network); D KL [.||.] is the Kullback-Leibler divergence, which measures the difference between the variational distribution and the true posterior distribution;
[0173] Reconstruction loss:
[0174] ,
[0175] The reconstruction loss in VAE usually uses mean squared error (MSE) or cross entropy, depending on the nature of the data; x is the original image; this term measures the signal reconstruction error, and the goal is to enable the generative model to reconstruct the original input data;
[0176] S503, MoE module processing: Dynamically select multiple expert networks to perform weighted combination of input latent vectors to generate more complex and diverse images. The specific process is:
[0177] S5031. Input latent space vector z': The input latent space vector z' is obtained from the VAE encoder and represents the feature representation of the WiFi signal RSSI data and the RGB image data. This vector contains the core information of the input data.
[0178] S5032, MoE weighting: Automatically adjust each expert's contribution based on the input features, making the generation process flexible and targeted. The formula is:
[0179] ,
[0180] The output of the above MoE is a weighted sum, w i (z') is the weight calculated by the gating network and satisfies ∑w i (z') = 1 (usually obtained through softmax activation); given the latent space vector z', the MoE module determines the weights of each expert through the expert selection mechanism; it is set to the output of the i-th expert network, and the weight {w i (z')} comes from a gating network; the generated latent vector {f i (z')};
[0181] The MoE module is used to enhance the latent space representation by dynamically selecting multiple expert networks to perform weighted combinations of the input latent vectors to generate more complex and diverse images.
[0182] S504, generator module processing: converting the potential vector z output by the VAE encoder and MoE module into an indoor area image;
[0183] S505, discriminator module processing: judge the authenticity of the generated image, evaluate the quality of the generated image, and provide feedback to the generator to optimize the generation result;
[0184] S506, output generated image: After being processed by the discriminator, a high-precision image is output.
[0185] Example 9
[0186] This embodiment is based on the improved VAE-MoE architecture WIFI signal indoor imaging method based on Example 8. In step S504, the generator module includes multiple deconvolution layers and convolution layers for decoding the latent space representation and generating a desired image. The specific processing process of the generator module is as follows:
[0187] S5041, fully connected layer processing: First, the latent vector z is mapped to a higher-dimensional feature space to ensure that the input latent vector can be fully processed in the generator, laying the foundation for subsequent image generation;
[0188] S5042, Self-Attention Layer Processing: Through Self-Attention, the generator dynamically focuses on different parts of the image generation process, calculates the similarity between each part, and adjusts the weights of each part; thus capturing long-range dependencies and global structure in the image;
[0189] S5043, deconvolution layer processing: gradually decode the processed features into a high-precision image and restore the image's spatial structure;
[0190] S5044, output layer: Generates the final image through convolution. The output image contains the visual information of the indoor scene and maintains the accuracy and clarity of the image.
[0191] Example 10
[0192] This embodiment of the method for indoor imaging of WIFI signals based on the improved VAE-MoE architecture is based on Example 9. The specific process of the self-attention layer processing in step S5042 is as follows:
[0193] (1) Query, key and value:
[0194] ,
[0195] The shape of the input feature matrix X is N×D, where N is the input feature vector and D is the dimension of each feature; Q 、W K , and W V is a learnable weight matrix corresponding to the conversion of query, key and value respectively;
[0196] (2) Calculation of attention weight:
[0197] ,
[0198] The above is to calculate the dot product between the query and the key to obtain the attention weight matrix, where is a scaling factor used to prevent the gradient from disappearing due to excessive dot product; the weight is then calculated using the softmax function, and the calculation formula is as follows:
[0199] ,
[0200] in is the attention weight of position i to position j, and the output is weight V;
[0201] (3) Weighted summation:
[0202] ,
[0203] The final output is the weighted sum of the values V by the attention weights, where α is the attention weight matrix and V is the value matrix.
[0204] Example 11
[0205] This embodiment of the method for indoor WIFI signal imaging based on the improved VAE-MoE architecture is based on the embodiment 10. The specific process of the discriminator module in step S505 is as follows:
[0206] S5051, multi-layer convolutional layer processing: Process the input image structure, texture and other information, and gradually extract deeper feature representations. The mathematical expression is:
[0207] ,
[0208] Among them, H (l) is the feature map of layer l, W (l) is the convolution kernel, b (l) is the bias term, * represents the convolution operation;
[0209] S5052, Self-Attention Layer Processing: Self-Attention is introduced to calculate the similarity between input features to adjust the weighted values of different features and improve the sensitivity to image details and structure. The calculation formula is: ,
[0210] Among them, Q, K and V are the matrices of query, key and value respectively, representing different transformations of input features, and softmax is used to calculate attention weights;
[0211] S5053, Fully Connected Layer Processing: After processing by the convolutional layer and the self-attention layer, the discriminator flattens the feature map and inputs it into the fully connected layer to map the high-level features of the image to the final classification space and judge the authenticity of the image, which is expressed as:
[0212] ,
[0213] Where W is the weight matrix, b is the bias term, and σ is the activation function (usually sigmoid), which is used to output a value between 0 and 1, indicating the true probability of the image.
[0214] S5054, output layer: outputs the true and false judgment results through the sigmoid activation function, 0 is the generated image, and 1 is the real image. The formula is as follows:
[0215] ,
[0216] Among them, p real Represents the probability that the input image is a real image. If p real If it is close to 1, it means that the image is more likely to be real; if it is close to 0, it means that the image is more likely to be generated;
[0217] S5055, Optimizer: Use the Adam optimizer to update the network parameters. The Adam optimizer accelerates convergence by adaptively adjusting the learning rate and avoids gradient explosion or gradient vanishing problems. The formula is as follows:
[0218] ,
[0219] in, are the parameters of the model, is the momentum of the gradient, v t is the second moment estimate of the gradient, is the learning rate, is a small constant to prevent division by zero.
[0220] Example 12
[0221] This embodiment of the indoor Wi-Fi signal imaging method based on the improved VAE-MoE architecture is based on Example 11. In step S506, the image generation step outputs a Leaky ReLU (rectified linear unit with leakage) activation function in the neural network architecture for image encoding and K-divergence loss (KLDivergence Loss) to ensure that the latent space distribution conforms to the standard normal distribution. The Adam optimizer is used for complex deep neural networks.
[0222] The Leak ReLU formula is as follows:
[0223] ,
[0224] Where, x: input value, : A small constant, usually between [0,1], with a common default value of 0.01;
[0225] The KL divergence loss formula is as follows:
[0226] ,
[0227] Where μ and σ are the mean and standard deviation output by the VAE encoder, and p(z) is the standard normal distribution;
[0228] The Adam optimizer formula is as follows:
[0229] ,
[0230] in, are model parameters, is the gradient momentum, v t is the gradient second moment estimate, is the learning rate, is a small constant to prevent division by zero.
Claims
1. A method for indoor imaging of WIFI signals based on an improved VAE-MoE architecture, characterized in that The following steps are involved: S100, Path Planning and Data Acquisition: A Wi-Fi transmitter is installed on one mobile device, and a Wi-Fi receiver is installed on another mobile device. The two mobile devices are located on opposite sides of the building to be measured and move around the perimeter of the building to be measured in a clockwise or counterclockwise direction. During the movement, the Wi-Fi signal emitted by the Wi-Fi transmitter is received by the Wi-Fi receiver, and time-synchronized Wi-Fi signal RSSI data and location data of the mobile devices are obtained; Use RGB camera to collect indoor RGB images; S200, data preprocessing: preprocessing the original WIFI signal data, including distance attenuation compensation and sparse signal processing; S300, simulating signal propagation: constructing a volume integral wave equation for the pre-processed data, and establishing the relationship between signal propagation and object distribution; S400, build a linearized physical model: Based on the volume-integrated wave equation, introduce Rytov approximate linearization and perform physical modeling by discretizing the volume-integrated wave equation and the RSSI simplified model; S500, deep learning model training and image generation: Extract latent features from WiFi signal RSSI data and RGB image data through MLP, convert the WiFi signal RSSI data and RGB image data into high-dimensional feature vectors of fixed dimensions, then use the VAE encoder module to encode the multimodal input data to generate the mean and variance of the latent space. The MoE module then performs weighted combination of the latent space, and then generates an RGB image through the generator module. The discriminator module then determines the authenticity of the generated image, optimizes the generator through training, and finally generates a high-precision image.
2. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 1 is characterized in that In step S100, the mobile device moves and collects data along multiple layers of moving paths. Specifically, the initial moving path of the mobile device is 0.5 meters away from the outer wall of the building to be measured. The initial movement is completed by circling the building to be measured. After that, a moving path is set every 1 meter in the direction away from the initial layer, and each moving path circles the building to be measured. After each moving path is completed, each mobile device scans the target area again at the 0° starting point, 45° point, 90° point and 135° point of the path to collect wall-penetrating signal characteristics at different angles.
3. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 2 is characterized in that In step S100 , the mobile device collects RSSI data and corresponding position data every time it moves 0.1 to 0.3 meters.
4. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 1 is characterized in that Step S200 distance attenuation compensation specifically involves measuring the original RSSI value of the Wi-Fi signal under unobstructed conditions with no obstructions between the two unmanned vehicles, and obtaining a distance-related attenuation model of the signal in an open space. The formula is: , in, : Received signal strength RSSI raw measurement value; :Derive distance-dependent power attenuation value based on obstacle-free path signal measurement; : Remove the signal after distance attenuation for subsequent modeling and imaging.
5. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 1 is characterized in that Step S200 sparse signal processing specifically involves organizing and improving the data before model building to provide high-quality input for subsequent high-definition modeling. The formula is: , Among them, TV(R): as the total variation of the target area, : local change of the i,jth point; Y: received signal strength vector; K: observation matrix, generated based on physical modeling; X: target sparse feature vector.
6. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 5 is characterized in that Step S300: Volume integral wave equation: , Where E(r) is the electric field received at position r, indicating the actual received WIFI signal strength; E inc (r) is the incident electric field, that is, the signal propagated in the absence of obstacles; G(r,r') is the Green function, which describes the propagation characteristics of the electromagnetic wave from position r' to position r; O(r') is the physical property at position r' in the target area, the dielectric constant or reflectivity of the scatterer; E(r') is the electric field generated by each scatterer in the target area; dv' is the integral of each scattering point in the target area D.
7. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 6 is characterized in that The physical modeling in step S400 is specifically as follows: S401, perform Rytov approximation, the formula is: , in, (r) is the phase change of the signal during propagation; S402, establish a discretized volume integral wave equation, the formula is: , in, is the phase change discrete vector, F is the propagation matrix, which represents the propagation relationship of the signal from the scattering source to the receiving point, and O is the physical characteristic vector of the target area; S403, RSSI simplified model, the signal model after volume integral wave equation and Rytov approximation is matched with the received signal strength RSSI of the WIFI device, and the simplified model is expressed as: , Among them, P ryt Indicates the received signal strength change, F R is the real part of the propagation matrix, O R It is the real part of the physical properties of the target area.
8. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 1 is characterized in that Step S500 deep learning model training and image generation specifically includes: S501. Extract potential features from RSSI data through MLP: The data representation is enhanced through the MLP network. The convolution layer and pooling layer are used. The convolution kernel size is 3×1. After processing, a fully connected layer is designed to perform high-dimensional feature mapping. Finally, a fixed-length high-dimensional feature vector is output as the input of the subsequent VAE and MoE. The RGB image is processed through MLP to extract latent features: the input RGB image is reduced in resolution, the channel size is increased to 512 using successive convolutional blocks, each channel is normalized by mean and standard deviation, and LeakyReLU activation is performed on each layer. After the convolutional layer, it is normalized and converted into a latent space vector, which represents the spatial information and texture features in the image; S502, VAE encoder module processing: Use the VAE encoder to encode the multimodal features of the RSSI data and RGB image data of the input WiFi signal to generate the mean and variance of the latent space. The mean and standard deviation represent the Gaussian distribution of the latent space and are used to describe the distribution of the latent variable Z. The variational inference and reconstruction loss formula of the VAE encoder is: Variational inference formula: , Among them, given the original image x as input data, learn the distribution of a latent variable z, the goal is to maximize the marginal log likelihood; p(x|z) is the generative model; q(z|x) is the variational distribution; D KL [.||.] is the Kullback-Leibler divergence, which measures the difference between the variational distribution and the true posterior distribution; Reconstruction loss: , The reconstruction loss in VAE uses mean square error MSE or cross entropy, depending on the nature of the data; x is the original image; S503, MoE module processing: Dynamically select multiple expert networks to perform weighted combination of input latent vectors to generate more complex and diverse images. The specific process is: S5031. Input latent space vector z': The input latent space vector z' is obtained from the VAE encoder and represents the feature representation of the WiFi signal RSSI data and the RGB image data; S5032, MoE weighting: Automatically adjust each expert's contribution based on the input features. The formula is: , The output of the above MoE is a weighted sum, w i (z') is the weight calculated by the gating network and satisfies ∑w i (z') = 1, which is obtained through softmax activation; given the latent space vector z', the MoE module determines the weight of each expert through the expert selection mechanism; it is set as the output of the i-th expert network, and the weight {w i (z')} comes from a gating network; the generated latent vector {f i (z')}; S504, generator module processing: converting the potential vector z output by the VAE encoder and MoE module into an indoor area image; S505, discriminator module processing: judge the authenticity of the generated image, evaluate the quality of the generated image, and provide feedback to the generator to optimize the generation result; S506. Output and generate a high-precision image: After evaluation and optimization by the discriminator, a high-precision image is output.
9. The indoor imaging method of WIFI signals based on the improved VAE-MoE architecture according to claim 8 is characterized in that The specific process of the generator module in step S504 is as follows: S5041, fully connected layer processing: first, the latent vector z is mapped to a higher-dimensional feature space to ensure that the input latent vector can be fully processed in the generator; S5042, Self-Attention Layer Processing: Through Self-Attention, the generator dynamically focuses on different parts of the image generation process, calculates the similarity between each part, and adjusts the weight of each part; S5043, deconvolution layer processing: gradually decode the processed features into a high-precision image and restore the image's spatial structure; S5044, output layer: Generates the final image through convolution. The output image contains the visual information of the indoor scene.
Citation Information
Patent Citations
WiFi sign language translation system and method based on deep learning
CN115188073A
Concept learning method, image generation method and related device
CN118247608A