Indoor layout modeling method based on deep learning and WIFI signals

By combining the deep learning method of WIFI signals and visual features, the problem that WIFI signals in the prior art is difficult to build high-precision three-dimensional indoor models, achieving more efficient feature fusion and high-resolution point cloud generation, significantly improving modeling accuracy and efficiency.

CN120219655APending Publication Date: 2025-06-27KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175696.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, WIFI signals are difficult to be used alone to build high-precision three-dimensional indoor models, and in complex scenarios or severe occlusion environments, geometric sensors may lose critical data, resulting in incomplete modeling results.

Method used

A deep learning-based method is adopted, combining WIFI signals and visual features, and the improved NeRF module and TorchSparse module are used for feature fusion and high-resolution point cloud generation, and the deep learning training weight is output, which significantly improves modeling accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and completeness of indoor scene modeling, supports signal prediction, expands application scenarios, and provides low-cost and high-root modeling solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219655A_ABST
    Figure CN120219655A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor layout modeling method based on deep learning and WIFI signals. The method comprises the steps of unmanned aerial vehicle path planning, data acquisition, data preprocessing, deep learning model training, three-dimensional model visualization and training model generation. According to the method, two unmanned aerial vehicles carrying WIFI signal equipment and an indoor RGB camera are utilized, RSSI signal data and RGB images are collected, data registration is carried out under the same coordinate system, and after processing is carried out through a feature extraction module and a three-dimensional reconstruction module in a deep learning model, a high-resolution three-dimensional point cloud can be output; according to the method, the problem that three-dimensional reconstruction of an indoor unknown scene is carried out only through WIFI signals is effectively solved, and the method has high innovativeness and wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a method for indoor layout modeling based on deep learning and WIFI signals. Background Art

[0002] With the continuous development of indoor positioning, wireless communication, and three-dimensional modeling technologies, indoor scene modeling based on WIFI signals has gradually become a research hotspot. Traditional indoor three-dimensional modeling methods usually rely on devices such as lidar and ultrasonic waves. Although these methods can generate high-precision three-dimensional models, the equipment cost is relatively high, and the requirements for equipment accuracy and operating environment are relatively strict. In addition, in complex scenarios or environments with severe occlusion, geometric sensors may lose key data, resulting in incomplete modeling results, thus limiting their application scope. In contrast, WIFI signals have the advantages of low cost, wide coverage, and the ability to penetrate obstacles, providing a new approach to solve the above problems.

[0003] However, the existing WIFI data is mainly applied to position estimation, human pose recognition, or signal propagation characteristic analysis, providing limited geometric information and unable to be used alone to construct a high-precision three-dimensional model. The sparsity and multipath effect of WIFI signals often lead to the lack of geometric information. In addition, in the process of fusing WIFI and visual data, there is a lack of a unified modeling framework, resulting in insufficient integrity and accuracy of scene modeling. These problems constitute the main challenges for indoor scene modeling through WIFI signals. Therefore, it is very necessary to develop a method for indoor layout modeling based on deep learning and WIFI signals that can solve the above problems. Summary of the Invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide a method for indoor layout modeling based on deep learning and WIFI signals, which combines WIFI signals and visual features, uses the sparse convolution operations of the improved NeRF module and TorchSparse module for feature fusion and high-resolution point cloud generation, and outputs the training weights of deep learning, significantly improving the modeling accuracy, efficiency, the ability of transfer learning, and the applicability of model reuse.

[0005] The purpose of the present invention is achieved as follows, including the following steps: S100. UAV path planning: Two UAVs are respectively located on opposite sides outside the building to be measured, set paths around the periphery of the building to be measured, set horizontal flight routes, vertical flight routes, and inclined flight routes to ensure coverage of the area to be measured, and maintain time synchronization; S200, Data Acquisition: One drone is equipped with a WIFI signal transmitter, and the other drone is equipped with a WIFI signal receiver. The receiver collects the real-time RSSI signal strength and GPS coordinates; the camera captures the indoor RGB image of the building to be measured, including depth information and camera position information; S300, Data Preprocessing: Establish a global coordinate system to unify the reference framework for GPS data and camera position coordinates, and then preprocess the RSSI data and RGB data; S400, Deep Learning Model Training: The feature extraction module uses an improved neural radiance field to perform high-frequency position encoding and implicit field modeling on WIFI and RGB features, extracts multi-modal features of geometry, texture, and signal, and forms high-quality features for the 3D reconstruction module to process. The features include 3D coordinates, RSSI data, and RGB color values; The 3D reconstruction module receives the latent features output by the feature extraction module, processes them through sparse convolution, upsamples layer by layer, and generates a 3D high-resolution point cloud after feature fusion; S500, 3D Model Visualization and Training Model Generation: Based on the high-resolution 3D point cloud, perform point cloud densification, generate a 3D model and visualize it, save the weights of the deep learning model and output.

[0006] Preferably, the specific horizontal flight route in step S100 is to set height levels and perform horizontal flight at different height levels. The two drones fly in horizontal parallel lines; the vertical flight route is to fly vertically along the outer wall, and set the diagonal points or specific observation points as the take-off points, and the points are distributed in groups of two; the inclined flight route is that the two drones fly along the building at an inclined angle of 30°, 45°, or 135°. The inclined routes of the two drones are parallel lines; when the three routes are executed, they all fly at the same speed and recording frequency.

[0007] Preferably, the drone carrying the WIFI signal receiver in step S200 is responsible for recording the GPS coordinates and RSSI data, and also records the timestamp for recording the data time and aligning the signal transmission information of the two drones. The drones need to maintain the same flight trajectory and speed, and the recording format is:

[0008] Where Timestamp: records the timestamp; x, y, z: GPS coordinates, indicating the position of the receiver in three-dimensional space; RSSI: signal strength; The camera captures the indoor RGB image and depth information, and records the three-dimensional position (x, y, z) of the camera. The recording format is as follows:

[0009] Among them, Camera_Position: three-dimensional coordinates of the camera, RGB_Image: captured color image, Depth_Map: corresponding depth map.

[0010] Preferably, the preprocessing of RSSI data and RGB data in step S300 is specifically as follows: S301. Establish a global coordinate system: Unify the reference frames of the indoor camera coordinates and the UAV coordinates to establish a global coordinate system. The UAV collects high-precision GPS coordinates and converts them to the global coordinate system through calibration. The GPS coordinate conversion is as follows: S3011. Convert geographical coordinates to a planar coordinate system: Use the ENU (East-North-Up) coordinate conversion formula to convert GPS coordinates to local three-dimensional planar coordinates (e, n, u). The formula is as follows:

[0011] Among them, R: radius of the earth, ∆latitude, ∆longitude: differences in latitude and longitude of the reference point, altitude0: altitude of the reference point; S3012. Align the indoor global coordinate system: Calculate the rotation and translation matrices through calibration points to convert ENU coordinates into the indoor global coordinate system. The formula is as follows:

[0012] Among them, T: coordinate transformation matrix; S302. RSSI data preprocessing: Normalize and denoise the collected RSSI data. Normalization maps the RSSI values to a unified range to make the data more friendly to the model. The formula is:

[0013] Among them, RSSI min : minimum value of RSSI, RSSI max : maximum value of RSSIz; RSSI data may be affected by multipath effects or obstacles in the environment. Denoising can eliminate small-scale fluctuations and make the signal smoother and more stable. The denoising processing formula is:

[0014] Among them, i: index of the current data point, k: half of the window size, usually set to k = 2 or k = 3; For the extreme values that still exist after denoising, perform threshold clipping:

[0015] Among them, : The RSSI value after denoising, i.e., the original signal strength; and : are the minimum and maximum signal strengths respectively; S303, RGB Image Preprocessing: The original data includes an RGB image and a depth map. The camera internal parameter matrix is used to undistort the RGB image and the depth map to ensure that the pixel points are consistent with the actual space, and the invalid values (such as 0 or NaN) in the depth map are filtered, and median filtering is used to remove noise; The specific process is as follows: S3031, Radial Distortion: Describes the radial offset of image pixel points. The formula is as follows:

[0016] where, (x, y): Normalized image plane coordinates; r 2 = x 2 + y 2 : The square of the radial distance from the point to the image center; k1, k2, k3: Radial distortion coefficients; Tangential Distortion: The offset of pixel points due to the tilt of the camera. The formula is as follows:

[0017] where, p1, p2: Tangential distortion coefficients; Pixel Coordinate Transformation: Converting normalized coordinates to actual pixel coordinates. The formula is as follows:

[0018] where, f x , f y : Focal length, c x , c y : Principal point; S3032, Depth Range Filtering: Set a reasonable depth range to filter invalid values. The formula is as follows:

[0019] where, D(u, v): Original depth value; D min , D max : The minimum and maximum effective measurement ranges of the depth sensor; Median Filtering for Denoising: Remove isolated noise points. The formula is as follows:

[0020] where, k: Half of the window size.

[0021] Preferably, the specific processing process of the feature extraction module in the S400 deep learning model training is: S4011. Location Encoding: The input features are extended from (x, y, z) to (x, y, z, RSSI, r, g, b), enhancing the joint modeling of multi-modal data for geometric and signal propagation characteristics; the output features are increased by (RSSI pred ), and the WIFI signal characteristics are incorporated into the 3D scene modeling; the input coordinates (x, y, z) are encoded by high-frequency sine and cosine functions to enhance the model's ability to capture geometric details. The formula is:

[0022] where L is the number of encoding frequency layers; S4012. Multi-modal Feature Concatenation: At the middle layer of the network, through a feature fusion module (such as feature concatenation or weighted sum), the RSSI and RGB features are jointly modeled with the coordinates after position encoding; the formula is:

[0023] The dimension of the concatenated feature vector is:

[0024] where "2L" is the position encoding dimension, "3" is the RGB color feature, and "1" is the RSSI feature; Through multi-modal feature concatenation, the input features are extended, enhancing the model's ability to model geometric information (position encoding), signal features (RSSI), and texture features (RGB); S4013. Implicit Field Modeling: Use a multi-layer perceptron (MLP) to model the input features, generate an implicit field in 3D space, and output the volume density and features of each point. The process is as follows: Input Features:

[0025] The features are used to extract high-dimensional features layer by layer through the MLP, and two branches are output at the last layer as follows:

[0026] where the occupancy probability indicates whether the point is occupied in 3D space and is used for geometric modeling. The feature representation features include signal features (RSSI) and texture features (RGB), which are used for the 3D reconstruction module to generate the final high-resolution point cloud; Implicit field modeling adds a processing branch for RSSI and RGB data features, strengthening the understanding of multi-modal data; The feature extraction module extracts multi-modal features of geometry, texture, and signals from multi-modal data to form high-quality features for the 3D reconstruction module to process. The features include 3D coordinates, RSSI data, and RGB color values.

[0027] Preferably, in the training of the deep learning model in step S400, the specific processing process of the 3D reconstruction module is as follows: S4021. Input data: Receive the data from the feature extraction module, including: {(x, y, z), σ, features}, which are coordinates, volume density, and feature vectors respectively; S4022. Sparse convolution operation: TorchSparse implements sparse convolution and sparse transposed convolution, directly performs feature extraction and aggregation on the sparse point cloud. Different from ordinary convolution that processes all spatial points, it reduces the computational cost; includes: sparse point cloud formatting, local feature extraction, global context aggregation; where the sparse convolution formula is:

[0028] Among them, : Updated feature vector, pi: Current point position, : Neighborhood points of point pi, : Convolution kernel weight function, representing the weight between neighborhood points, : Neighborhood point feature vector; S4023. Multi-layer progressive upsampling: The goal is to gradually generate a high-resolution dense point cloud from the sparse point cloud. Use sparse transposed convolution for progressive upsampling to generate a high-resolution point cloud, and at the same time retain the global context information in each layer of convolution, gradually refining the geometric details and signal characteristics, and finally output a high-resolution point cloud, including the 3D coordinates, volume density, and multi-modal features (RSSI and RGB) of each point; where the formula for sparse transposed convolution is:

[0029] Among them, : Feature vector of point pi after upsampling, : Set of points in the low-resolution space, : Transposed convolution kernel weight function, : Feature vector of low-resolution points; S4024. Feature fusion: Fuse the multi-modal features (RSSI and RGB data) output by the feature extraction module. The point cloud generated by the 3D reconstruction module contains geometric information, as well as texture and signal characteristics; after each layer of convolution, fuse the features of the current layer with the features output by the feature extraction module, and the fused features are used for the next layer of convolution and upsampling operations. The formula is as follows:

[0030] S4025. Output data: The 3D reconstruction module finally outputs a high-resolution point cloud, which is input into the signal prediction module to output the volume density and signal intensity distribution, achieving collaborative optimization of geometric modeling and signal modeling. The output is as follows:

[0031] Among them, (x, y, z): the three-dimensional spatial position of each point in the point cloud; : volume density; : scene texture features; : signal prediction intensity.

[0032] Preferably, in the feature extraction module, the loss function is optimized based on a specific scenario. The specific process is as follows: (1) Occupancy Loss: Optimize the geometric structure modeling of the 3D scene by the feature extraction module to make accurately describe the geometric distribution of the point cloud. The formula is as follows:

[0033] Among them, : the volume density predicted by the feature extraction module, : the true volume density, obtained by annotating the actual point cloud data; (2) Multimodal Feature Consistency Loss: Ensure that the feature extraction module has consistency and correlation in feature extraction of multimodal data (RSSI and RGB data). The formula is as follows:

[0034] Among them, : the multimodal features predicted by the feature extraction module, : the target features generated by multimodal fusion.

[0035] Preferably, in the 3D reconstruction module, the loss function is optimized for multimodal data. The specific process is as follows: (1) Signal Prediction Loss: Optimize the modeling of the WIFI signal intensity distribution by the 3D reconstruction module to make the predicted signal distribution more in line with the true signal propagation characteristics;

[0036] Among them, : the signal intensity predicted by the 3D reconstruction module, : the true RSSI data.

[0037] (2) Color Reconstruction Loss: Optimize the modeling of the point cloud texture information by the 3D reconstruction module to make the generated RGB color values more realistic.

[0038]

[0039] Among them, : The 3D reconstruction module predicts the point color value, : The true point cloud color value; (3) Dense point cloud sparsity constraint loss: In the process of dense point cloud, the sparsity of the point cloud is restricted. The formula is as follows:

[0040] Among them, : Sparsity threshold, usually set to a small value (such as 0.1). If the volume density of the point cloud exceeds the threshold, the loss is increased.

[0041] Preferably, the specific process of step S500 is as follows: S501 3D scene visualization: First, the output high-resolution point cloud needs to be densified to eliminate the discontinuity of the point cloud and improve the integrity of the scene. Voxel grid densification is used. The formula is as follows:

[0042] Among them, : The weight of point i; : The original feature in the point cloud ; Finally, a 3D model is presented in the software, including signal strength and color value coverage of the dense point cloud; S502 Training model weight output: To adapt to WIFI modeling in different scenarios and subsequent inference and transfer learning, it is necessary to save the weight outputs of the feature extraction module and the 3D reconstruction module for subsequent loading and optimization; the output weights of the feature extraction module mainly include the parameters of the sparse convolutional layer and the fully connected layer, which are used to extract multi-modal features from the input data; the output weights of the 3D reconstruction module include the parameters of the transposed convolutional layer, the feature fusion module, and the prediction branches (volume density, signal strength, color value), which are used to decode and generate a high-resolution point cloud from the latent representation of the feature extraction module; The saved model weights can be reused, and each repeated training can be continued indefinitely, which is convenient for transfer learning and applicable to different scenarios. Save the parameters of the feature extraction module and the 3D reconstruction module (such as.pth or.h5). Subsequently, the model weights can be loaded, new data can be input for prediction, or the weights can be fine-tuned according to different scenarios.

[0043] Preferably, the specific process of step S501 is as follows: S5011. Point cloud densification processing: To facilitate 3D scene modeling and rendering, it is necessary to convert the sparse point cloud output by the 3D reconstruction module into a dense point cloud to provide higher resolution and accuracy. Densification based on a voxel grid is adopted, projecting the point cloud into a 3D voxel grid. Each voxel is a small cube, and the grid resolution can be adjusted according to the scene size. The interpolation formula is as follows:

[0044] Among them, : The weight of point i; : The original features in the point cloud ; Output the dense point cloud, and the format is as follows:

[0045] Among them, the coordinates (x, y, z) correspond to the dense points in the voxel grid; S5012. 3D modeling formation: According to the coordinates (x, y, z) and volume density in the dense point cloud Reconstruct the 3D geometric structure, map the predicted signal strength RSSI pred onto the 3D scene surface to generate a signal coverage visualization effect; map (r, g, b) in the point cloud onto the 3D surface to render the scene; set the signal strength color coding to visualize the RSSI pred strength with the signal gradient; S5013. Visualization: Use common point cloud and 3D model visualization software (such as Open3D, Meshlab) for model visualization, and at the same time map the signal strength onto the color surface for rendering with color coding.

[0046] Compared with the prior art, the present invention has the following technical effects: 1. The present invention introduces multi-modal data fusion to improve data accuracy. The RSSI data is effectively fused with RGB images and depth information to construct a unified 3D scene modeling framework, making up for the deficiencies of single-modal technologies in spatial information modeling and significantly improving the accuracy and integrity of modeling; 2. The present invention is improved on the basis of the NeRF model, introduces WIFI signal characteristics into the modeling framework, combines sparse convolutional networks to achieve more efficient feature extraction and reconstruction, further optimizes the network structure, reduces the computational complexity, supports signal prediction, and expands the application scenarios; 3. Using WIFI signals as the data source has low cost and high adaptability. WIFI signals have good environmental penetration ability and show high adaptability in complex indoor environments such as areas with severe occlusion or insufficient lighting. It is a low-cost and highly robust modeling solution; 4. The present invention supports saving the trained deep learning model in a standard format (such as.pth), which facilitates subsequent model reuse and transfer learning. The output model weights can be quickly adapted to other scenarios and tasks, and have stronger scalability; 5. The present invention realizes the visualization of the signal coverage range, can intuitively understand the signal blind spots and hot spot positions, provides strong support for wireless network optimization and device layout, and expands the application dimension of three-dimensional modeling technology; 6. The indoor fine modeling implemented by the present invention provides a new option for security monitoring and military aspects. Through high-precision modeling of complex indoor environments and signal coverage evaluation, it provides global scene support. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic flow chart of the present invention; Figure 2 is a schematic diagram of three routes of the UAV path planning of the present invention; Figure 3 is a schematic plan view of the building to be measured of the present invention; Figure 4 is a schematic flow chart of the feature extraction module of the present invention generating potential features; Figure 5 is a schematic flow chart of the three-dimensional reconstruction module of the present invention generating high-resolution point clouds; Figure 6 is a schematic diagram of the three-dimensional model effect generated by the three-dimensional model visualization module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be further described below in conjunction with the embodiments and the drawings, but the present invention is not limited in any way. Any transformation or replacement based on the teachings of the present invention falls within the protection scope of the present invention.

[0049] Embodiment 1 As shown in the Figure 1 accompanying drawings, the indoor layout modeling method based on deep learning and WIFI signals in this embodiment includes the following steps: S100. UAV path planning: Two UAVs are respectively located on opposite sides outside the building to be measured, set paths around the periphery of the building to be measured, set horizontal flight routes, vertical flight routes and inclined flight routes to ensure coverage of the area of the building to be measured, and keep time synchronization; S200. Data acquisition: One UAV is equipped with a WIFI signal transmitter, and the other UAV is equipped with a WIFI signal receiver. The receiver collects the real-time RSSI signal strength and GPS coordinates; the camera collects the indoor RGB images of the building to be measured, including depth information and camera position information; S300. Data preprocessing: Establish a global coordinate system to unify the reference framework for GPS data and camera position coordinates, and then preprocess the RSSI data and RGB data; S400. Deep learning model training: The feature extraction module uses an improved neural radiance field to perform high-frequency position encoding and implicit field modeling on WIFI and RGB features, extracts multi-modal features of geometry, texture, and signal, and forms high-quality features for the 3D reconstruction module to process. The features include 3D coordinates, RSSI data, and RGB color values; The 3D reconstruction module receives the latent features output by the feature extraction module, processes them through sparse convolution, upsamples layer by layer, and generates a 3D high-resolution point cloud after feature fusion; S500. 3D model visualization and training model generation: Based on the high-resolution 3D point cloud, perform point cloud densification, generate a 3D model and visualize it, save the weights of the deep learning model and output.

[0050] Embodiment 2 As shown in the appendix Figure 2 In this embodiment, the indoor layout modeling method based on deep learning and WIFI signals is based on Embodiment 1. In step S100, the horizontal flight route specifically sets height levels, performs horizontal flight at different height levels, and two drones fly in horizontal parallel lines; the vertical flight route flies vertically along the outer wall, sets the diagonal points or specific observation points as the take-off points, and the points are distributed in groups of two; the inclined flight route is that two drones fly along the building at an inclined angle of 135°, and the inclined routes of the two drones are parallel lines; when the three routes are executed, they all fly at the same speed and recording frequency.

[0051] Embodiment 3 In this embodiment, the indoor layout modeling method based on deep learning and WIFI signals is based on Embodiment 2. In step S200, the drone carrying the WIFI signal receiver is responsible for recording the GPS coordinates and RSSI data, and at the same time records the timestamp, which is used to record the data time and align the signal transmission information of the two drones. The drones need to maintain the same flight trajectory and speed. The recording format is:

[0052] Among them, Timestamp: records the timestamp; x, y, z: GPS coordinates, indicating the position of the receiver in three-dimensional space; RSSI: signal strength; Multiple RGB cameras are installed in the scene to be modeled. The cameras capture indoor RGB images and depth information, record the three-dimensional position (x, y, z) of the cameras, and the recording format is as follows:

[0053] Among them, Camera_Position: the three-dimensional coordinates of the camera, RGB_Image: the captured color image, Depth_Map: the corresponding depth map.

[0054] Embodiment 4 The indoor layout modeling method based on deep learning and WIFI signals in this embodiment is based on Embodiment 3. The preprocessing of RSSI data and RGB data in step S300 is specifically as follows: S301. Establish a global coordinate system: Unify the reference frames of the indoor camera coordinates and the UAV coordinates to establish a global coordinate system. The high-precision GPS coordinates collected by the UAV are converted into the global coordinate system through calibration; the GPS coordinate conversion is as follows: S3011. Convert geographical coordinates to a plane coordinate system: Use the ENU coordinate conversion formula to convert GPS coordinates into local three-dimensional plane coordinates (e, n, u). The formula is as follows:

[0055] Among them, R: the radius of the earth, ∆latitude, ∆longitude: the difference in longitude and latitude of the reference point, altitude0: the height of the reference point; S3012. Align the indoor global coordinate system: Calculate the rotation and translation matrices through calibration points, and convert the ENU coordinates into the indoor global coordinate system. The formula is as follows:

[0056] Among them, T: the coordinate transformation matrix; S302. RSSI data preprocessing: Normalize and denoise the collected RSSI data. Normalization maps the RSSI value to a unified range. The formula is:

[0057] Among them, RSSI min : the minimum value of the RSSI value, RSSI max : the maximum value of RSSIz; After that, denoise. The denoising processing formula is:

[0058] Among them, i: the index of the current data point, k: half of the window size, usually set to k = 2 or k = 3; For the extreme values that still exist after denoising, perform threshold clipping: ; S303. RGB Image Preprocessing: The original data includes an RGB image and a depth map. The camera internal parameter matrix is used to undistort the RGB image and the depth map (including radial distortion and tangential distortion) to ensure that the pixel points are consistent with the actual space, filter out the invalid values in the depth map, and use median filtering to remove noise. The specific process is as follows: S3031. Radial Distortion: Describes the radial offset of image pixel points. The formula is as follows:

[0059] where (x, y) are the normalized image plane coordinates; r 2 = x 2 + y 2 is the square of the radial distance from the point to the image center; k1, k2, k3 are the radial distortion coefficients; Tangential Distortion: The offset of pixel points caused by the tilt of the camera. The formula is as follows:

[0060] where p1, p2 are the tangential distortion coefficients; Pixel Coordinate Transformation: Converts the normalized coordinates to actual pixel coordinates. The formula is as follows:

[0061] where f x , f y are the focal lengths, c x , c y are the principal points; S3032. Depth Range Filtering: Sets a reasonable depth range to filter out invalid values. The formula is as follows:

[0062] where D(u, v) is the original depth value; D min , D max are the minimum and maximum effective measurement ranges of the depth sensor; Median Filtering for Denoising: Removes isolated noise points. The formula is as follows:

[0063] where k is half of the window size.

[0064] Example 5 As shown in the appendix Figure 3 In this example, based on the deep learning and WIFI signal indoor layout modeling method, on the basis of Example 4, the specific processing process of the feature extraction module in the S400 deep learning model training is as follows: S4011. Location Encoding: The input features are extended from (x, y, z) to (x, y, z, RSSI, r, g, b); the output features are increased by RSSI pred , and the WIFI signal characteristics are integrated into the 3D scene modeling; the input coordinates (x, y, z) are encoded by high-frequency sine and cosine functions, and the formula is:

[0065] where L: the number of encoding frequency layers; S4012. Multimodal Feature Concatenation: In the middle layer of the network, through the feature fusion module, the RSSI and RGB features are jointly modeled with the coordinates after location encoding; the formula is:

[0066] The dimension of the concatenated feature vector is:

[0067] where, "2L": the location encoding dimension, "3": the RGB color feature, "1": the RSSI feature; S4013. Implicit Field Modeling: Through a multi-layer sparse convolutional network, multi-modal features are extracted to generate latent representations, including occupancy probability and feature representation. The sparse convolutional layer is designed with an input dimension of 2L + 4. Through the design of multi-layer sparse convolutional layers, the output dimension is 256, the activation function is ReLU, and the MLP further fuses the convolutional features to output the occupancy probability σ and the feature representation. Among them, the Sigmoid activation function is used for σ; the specific process is as follows: Input Features:

[0068] The features extract high-dimensional features layer by layer through the MLP, and two branches are output at the last layer as follows:

[0069] where, the occupancy probability : indicates whether the point is occupied in the 3D space and is used for geometric modeling. The feature representation features: contains signal features and texture features and is used for the 3D reconstruction module to generate the final high-resolution point cloud.

[0070] Embodiment 6 As shown in the appendix Figure 4 In this embodiment, based on the deep learning and WIFI signal indoor layout modeling method, on the basis of Embodiment 5, the specific processing process of the 3D reconstruction module in the S400 step of deep learning model training is: S4021. Input data: Receive data from the feature extraction module, including: {(x, y, z), σ, features}, which are the coordinates, volume density, and feature vector respectively; S4022. Sparse convolution operation: TorchSparse is used to implement sparse convolution and sparse transposed convolution, directly performing feature extraction and aggregation on the sparse point cloud; the sparse convolution formula is:

[0071] where, : Updated feature vector, pi: Position of the current point, : Neighborhood points of point pi, : Convolution kernel weight function, representing the weight between neighborhood points, : Feature vector of neighborhood points; S4023. Multi - layer progressive upsampling: The goal is to gradually generate a high - resolution dense point cloud from the sparse point cloud. Use sparse transposed convolution for progressive upsampling to generate a high - resolution point cloud, while retaining global context information in each layer of convolution, gradually refining geometric details and signal characteristics, and finally output a high - resolution point cloud containing the three - dimensional coordinates, volume density, and multi - modal features of each point; the formula for sparse transposed convolution is:

[0072] where, : Feature vector of point pi after upsampling, : Set of points in the low - resolution space, : Transposed convolution kernel weight function, : Feature vector of low - resolution points; S4024. Feature fusion: Fuse the multi - modal features output by the feature extraction module. The point cloud generated by the 3D reconstruction module contains geometric information, texture, and signal characteristics; after each layer of convolution, fuse the features of the current layer with the features output by the feature extraction module, and the fused features are used for the next layer of convolution and upsampling operations. The formula is as follows:

[0073] S4025. Output data: The 3D reconstruction module finally outputs a high - resolution point cloud and feeds it into the signal prediction module, outputting the volume density and signal intensity distribution. The output is as follows:

[0074] where, (x, y, z): Three - dimensional spatial position of each point in the point cloud; : Volume density; : Scene texture feature; : Signal prediction intensity.

[0075] Example 7 Based on the deep learning and WIFI signal indoor layout modeling method in this embodiment, on the basis of Embodiment 6, in the feature extraction module, the loss function is optimized based on a specific scenario. The specific process is as follows: (1) Occupancy Loss: Optimize the feature extraction module to model the three-dimensional scene geometric structure, so that accurately describe the geometric distribution of the point cloud. The formula is as follows:

[0076] Among them, : The volume density predicted by the feature extraction module, : The true volume density, obtained by annotating the actual point cloud data; (2) Multimodal feature consistency loss: The feature extraction module has consistency and correlation in the feature extraction of multimodal data. The formula is as follows:

[0077] Among them, : The multimodal features predicted by the feature extraction module, : The target features generated by multimodal fusion; After optimization, the Occupancy Loss in the feature extraction module more accurately describes the optimization of the geometric distribution of the point cloud; the multimodal feature consistency loss has consistency and correlation in the feature extraction of multimodal data.

[0078] Example 8 Based on the deep learning and WIFI signal indoor layout modeling method in this embodiment, on the basis of Embodiment 7, in the three-dimensional reconstruction module, the loss function is optimized for multimodal data. The specific process is as follows: (1) Signal prediction loss: Optimize the three-dimensional reconstruction module to model the WIFI signal strength distribution. The formula is:

[0079] Among them, : The signal strength predicted by the three-dimensional reconstruction module, : The true RSSI data; (2) Color reconstruction loss: Optimize the three-dimensional reconstruction module to model the point cloud texture information. The formula is:

[0080] Among them, : The predicted point color value by the three-dimensional reconstruction module, : The true point cloud color value; (3) Dense point cloud sparsity constraint loss: During the process of dense point cloud, the sparsity of the point cloud is restricted; the formula is as follows:

[0081] Among them, : Sparsity threshold, usually set to a small value (such as 0.1). If the volume density of the point cloud exceeds the threshold, the loss is increased; After optimization, the signal prediction loss in the 3D reconstruction module makes the predicted signal distribution more in line with the true signal propagation characteristics; the color reconstruction loss optimizes the point cloud texture modeling, and the RGB color values are more realistic; the dense point cloud sparsity constraint loss restricts the sparsity of the point cloud and avoids point cloud redundancy.

[0082] Embodiment 9 The indoor layout modeling method based on deep learning and WIFI signals in this embodiment is based on Embodiment 8. The specific process of step S500 is as follows: S501 3D scene visualization: First, densify the output high-resolution point cloud to eliminate the discontinuity of the point cloud and improve the integrity of the scene. Use voxel grid densification, and the formula is as follows:

[0083] Among them, : The weight of point i; : The original feature in the point cloud ; Finally, a 3D model is presented in the software, including the signal strength and the color value coverage of the dense point cloud; S502 Training model weight output: Save the weight outputs of the feature extraction module and the 3D reconstruction module; the weight outputs of the feature extraction module mainly include the parameters of the sparse convolutional layer and the fully connected layer, which are used to extract multi-modal features from the input data; the weight outputs of the 3D reconstruction module include the transposed convolutional layer, the feature fusion module, and the prediction branch parameters, which are used to decode and generate a high-resolution point cloud from the latent representation of the feature extraction module.

[0084] Embodiment 10 The indoor layout modeling method based on deep learning and WIFI signals in this embodiment is based on Embodiment 9. The specific process of step S501 is as follows: S5011 Point cloud densification processing: Adopt voxel grid-based densification, project the point cloud into a 3D voxel grid, each voxel is a small cube, and the grid resolution can be adjusted according to the scene size. The interpolation formula is as follows:

[0085] Among them, : The weight of point i; : The original feature in the point cloud ; Output dense point cloud in the following format:

[0086] Among them, the coordinates (x, y, z) correspond to the dense points in the voxel grid; S5012. 3D modeling and signal coverage visualization: For point cloud densification and 3D modeling, Open3D is used for implementation. The dense point cloud is converted into a 3D mesh model using the triangulation function of Open3D, and the poissonsurface reconstruction() method is used for surface reconstruction to generate a smooth 3D mesh. The 3D mesh model is combined with RSSI pred to perform the visualization of signal strength coverage.

[0087] S5013. RSSI pred Mapping visualization: Use Open3D for model visualization. The (r, g, b) in the point cloud is mapped to the 3D surface to render the scene; Use color mapping to map the signal strength of each point to the mesh surface, that is, the paintuniform color() method assigns colors to the mesh, and adjusts the colors according to the RSSI value, so as to realize the visualization of signal strength. Set low signal strength to be represented by blue and high signal strength to be represented by red, and the effect is more intuitive.

[0088] S5014. Result presentation and output: draw geometries() in Open3D visualizes the generated 3D model again (such as Figure 6 )), including the geometric structure of the building and the RSSI pred color mapping. More details of the model can be viewed by rotating and scaling. The generated 3D model is exported in a standard format (such as ply or obj) for convenient subsequent use or further processing.

Claims

1. A method for modeling indoor layout of WIFI signals based on deep learning, characterized in that The following steps are involved: S100, UAV path planning: Two UAVs are located on opposite sides of the building to be tested, and paths are set around the building to be tested. Horizontal flight routes, vertical flight routes and inclined flight routes are set to ensure coverage of the building area to be tested and maintain time synchronization; S200, data collection: one drone is equipped with a WIFI signal transmitter, and the other drone is equipped with a WIFI signal receiver. The receiver collects real-time RSSI signal strength and GPS coordinates; the camera collects the indoor RGB image of the building to be tested, including depth information and camera position information; S300, data preprocessing: establishing a global coordinate system, unifying the reference frame of GPS data and camera position coordinates, and then preprocessing RSSI data and RGB data; S400, deep learning model training: The feature extraction module uses an improved neural radiation field to perform high-frequency position encoding and implicit field modeling on WIFI and RGB features, extract multimodal features of geometry, texture, and signal, and form high-quality features for processing by the 3D reconstruction module. The features include 3D coordinates, RSSI data, and RGB color values. The 3D reconstruction module receives the potential features output by the feature extraction module, performs sparse convolution processing, upsamples layer by layer, and generates a 3D high-resolution point cloud after feature fusion; S500, 3D model visualization and training model generation: Based on high-resolution 3D point cloud, perform point cloud densification, generate 3D model and visualize it, save deep learning model weights and output them.

2. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that The horizontal flight route in step S100 is to set the height levels and perform horizontal flight at different height levels, with the two drones flying in horizontal parallel lines; the vertical flight route flies vertically along the outer wall, setting the diagonal point or specific observation point as the take-off point, and the points are distributed in groups of two; the inclined flight route is that the two drones fly along the building at an inclination angle of 30°, 45° or 135°, and the inclined routes of the two drones are parallel lines; the same flight speed and recording frequency are maintained during the execution of the three routes.

3. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that In step S200, the drone equipped with the WIFI signal receiver is responsible for recording the GPS coordinates and RSSI data, and also recording the timestamp, which is used to record the data time and align the signal transmission information of the two drones. The drones need to maintain synchronized flight trajectory and speed. The recording format is: Timestamp: record timestamp; x, y, z: GPS coordinates, indicating the position of the receiver in three-dimensional space; RSSI: signal strength; The camera captures the indoor RGB image and depth information, and records the three-dimensional position (x, y, z) of the camera in the following format: Among them, Camera_Position: camera 3D coordinates, RGB_Image: captured color image, Depth_Map: corresponding depth map.

4. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that Step S300 pre-processes the RSSI data and RGB data, specifically: S301, establish a global coordinate system: unify the indoor camera coordinates and the drone coordinates into a unified reference frame to establish a global coordinate system. The high-precision GPS coordinates collected by the drone are converted into the global coordinate system through calibration; the GPS coordinate conversion is as follows: S3011. Conversion of geographic coordinates to plane coordinate system: Use the ENU coordinate conversion formula to convert GPS coordinates to local three-dimensional plane coordinates (e, n, u). The formula is as follows: Where, R: radius of the earth, ∆latitude, ∆longitude: difference between longitude and latitude of the reference point, altitude0: altitude of the reference point; S3012, aligning the indoor global coordinate system: calculate the rotation and translation matrix through the calibration points, and transform the ENU coordinates into the indoor global coordinate system. The formula is as follows: Where, T: coordinate transformation matrix; S302, RSSI data preprocessing: normalize and remove noise from the collected RSSI data. Normalization maps the RSSI value to a uniform range. The formula is: Among them, RSSI min : Minimum RSSI value, RSSI max : RSSIz maximum value; Then remove the noise, the noise removal processing formula is: Where i: current data point index, k: half of the window size, usually set to k=2 or k=3; S303, RGB image preprocessing: The original data includes RBG image and depth map. The camera internal parameter matrix is ​​used to dedistort the RGB image and depth map to ensure that the pixels are consistent with the actual space, and the invalid values ​​in the depth map are filtered out, and the median filter is used to remove noise. The specific process is as follows: S3031, radial distortion: describes the radial displacement of image pixels. The formula is as follows: Where, (x, y): normalized image plane coordinates; r 2 =x 2 +y 2 : The square of the radial distance from the point to the image center; k1, k2, k3: radial distortion coefficients; Tangential distortion: The displacement of pixels due to camera tilt. The formula is as follows: Among them, p1, p2: tangential distortion coefficient; Pixel coordinate conversion: Normalized coordinates are converted to actual pixel coordinates. The formula is as follows: Among them, f x , f y : focal length, c x , c y : Main point; S3032, depth range filtering: set a reasonable depth range and filter invalid values. The formula is as follows: Where, D(u, v): original depth value; D min , D max : The minimum and maximum effective measurement range of the depth sensor; Median filter denoising: remove isolated noise points, the formula is as follows: Where k is half of the window size.

5. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that The specific processing process of the feature extraction module in the S400 deep learning model training is: S4011, position encoding: the input features are expanded from (x, y, z) to (x, y, z, RSSI, r, g, b); the output features are increased by RSSI pred , the WIFI signal characteristics are integrated into the 3D scene modeling; the input coordinates (x, y, z) are encoded by high-frequency sine and cosine functions, and the formula is: Where, L: number of coding frequency layers; S4012, multi-modal feature concatenation: In the middle layer of the network, the RSSI and RGB features are jointly modeled with the position-encoded coordinates through the feature fusion module; the formula is: The dimension of the feature vector after concatenation is: Among them, "2L": position encoding dimension, "3": RGB color feature, "1": RSSI feature; S4013, Implicit Field Modeling: Use a multi-layer perceptron to model the input features, generate a three-dimensional space implicit field, and output the volume density and features of each point. The process is: Input Features: The features are extracted layer by layer through MLP, and two branches are output in the last layer as follows: Among them, the occupancy probability : Indicates whether the point is occupied in the three-dimensional space, used for geometric modeling, feature representation features: contains signal features and texture features, used for the three-dimensional reconstruction module to generate the final high-resolution point cloud.

6. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that The specific processing process of the 3D reconstruction module in the deep learning model training in step S400 is: S4021, input data: receiving feature extraction module data, including: {(x, y, z), σ, features}, which are coordinates, volume density and feature vector respectively; S4022, sparse convolution operation: TorchSparse implements sparse convolution and sparse transposed convolution, and performs feature extraction and aggregation directly on sparse point clouds; the sparse convolution formula is: in, : updated feature vector, pi: current point position, : Neighborhood points of point pi, : Convolution kernel weight function, which represents the weight between neighborhood points. : Neighborhood point feature vector; S4023, multi-layer progressive upsampling: The goal is to gradually generate a high-resolution dense point cloud from a sparse point cloud, using sparse transposed convolution to gradually upsample and generate a high-resolution point cloud, while retaining global context information in each layer of convolution, gradually refining geometric details and signal characteristics, and finally outputting a high-resolution point cloud, including the three-dimensional coordinates, volume density, and multimodal features of each point; the formula for sparse transposed convolution is: in, : The feature vector of point pi after upsampling, : is a set of points in low-resolution space, : Transpose convolution kernel weight function, : feature vector of low-resolution points; S4024, feature fusion: Fusion of multimodal features output by the feature extraction module. The point cloud generated by the 3D reconstruction module contains geometric information, texture and signal characteristics. After each layer of convolution, the current layer features are fused with the features output by the feature extraction module. The fused features are used for the next layer of convolution and upsampling operations. The formula is as follows: S4025, output data: The 3D reconstruction module finally outputs a high-resolution point cloud, and adds a signal prediction module to output volume density and signal intensity distribution. The output is as follows: Among them, (x, y, z): the three-dimensional spatial position of each point in the point cloud; : Bulk density; : Scene texture features; : Signal prediction strength.

7. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 5 is characterized in that In the feature extraction module, the loss function is optimized based on a specific scenario. The specific process is: (1) Volume density loss: Optimize the feature extraction module to model the geometric structure of the 3D scene. Accurately describe the geometric distribution of the point cloud, the formula is as follows: in, : Volume density predicted by the feature extraction module, : True volume density, obtained by annotating actual point cloud data; (2) Multimodal feature consistency loss: The feature extraction module has consistency and correlation in the feature extraction of multimodal data. The formula is as follows: in, : Multimodal features predicted by the feature extraction module, : Multimodal fusion generates target features.

8. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 6 is characterized in that In the 3D reconstruction module, the loss function is optimized for multimodal data. The specific process is: (1) Signal prediction loss: Optimize the 3D reconstruction module to model the WIFI signal strength distribution. The formula is: in, : The 3D reconstruction module predicts signal intensity, : Real RSSI data; (2) Color reconstruction loss: Optimize the modeling of point cloud texture information in the 3D reconstruction module. The formula is: in, : The 3D reconstruction module predicts the point color value, : Real point cloud color value; (3) Sparsity constraint loss of dense point cloud: In the process of dense point cloud, the sparsity of point cloud is restricted; the formula is: in, : Sparsity threshold, usually set to a small value (such as 0.1). If the point cloud volume density exceeds the threshold, the loss will increase.

9. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 1 is characterized in that The specific process of step S500 is: S501, 3D scene visualization: First, the output high-resolution point cloud is densified to eliminate point cloud discontinuity and improve the integrity of the scene. Voxel grid densification is used. The formula is as follows: in, : The weight of point i; : Original features in point cloud ; Finally, the 3D model is presented in the software, including the signal strength and color value overlay of the dense point cloud; S502, training model weight output: save the weight output of the feature extraction module and the 3D reconstruction module; the output weight of the feature extraction module mainly includes the parameters of the sparse convolution layer and the fully connected layer, which are used to extract multimodal features from the input data; the output weight of the 3D reconstruction module includes the transposed convolution layer, the feature fusion module and the prediction branch parameters, which are used to decode and generate high-resolution point clouds from the potential representation of the feature extraction module.

10. The method for modeling indoor layout of WIFI signals based on deep learning according to claim 9 is characterized in that The specific process of step S501 is: S5011, point cloud densification processing: voxel grid-based densification is used to project the point cloud into a three-dimensional voxel grid. Each voxel is a small cube. The grid resolution can be adjusted according to the scene size. The interpolation formula is as follows: in, : The weight of point i; : Original features in point cloud ; Output dense point cloud in the following format: where the coordinates (x, y, z) correspond to dense points in the voxel grid; S5012, 3D modeling: Based on the coordinates (x, y, z) and volume density in the dense point cloud Reconstructing the 3D geometry will predict the signal strength RSSI pred Map to the 3D scene surface to generate signal coverage visualization; map (r, g, b) in the point cloud to the 3D surface to render the scene; set signal strength color coding to signal gradient to RSSI pred Visualize the intensity; S5013, Visualization: Model visualization using point cloud and 3D model visualization software, while mapping signal intensity to color surfaces for color-coded rendering.

Citation Information

Cited By

  • Wireless signal processing method

    CN121603966A