A RIS-Assisted Indoor Fingerprint Localization Method Based on Residual Neural Network
The method uses a single-antenna RIS system with a residual neural network to process CSI fingerprints for indoor positioning, addressing high configuration requirements and improving accuracy in complex indoor environments.
Patent Information
- Application Number
- CN202510650024.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing fingerprint positioning method based on intelligent reflection surfaces has high requirements for base station configuration, the positioning accuracy needs to be improved, and it is not robust enough in complex indoor environments.
Using the RIS-assisted indoor fingerprint positioning method based on residual neural network, a single-antenna base station and RIS are used to normalize the received signal matrix, a data set is constructed and the residual neural network is trained, and a three-dimensional position estimation is performed, and the signal propagation effect is enhanced by the reflection and regulation capabilities of RIS.
It realizes high-precision positioning under low configuration requirements, has good robustness and generalization capabilities, and can provide high-precision three-dimensional positioning estimation in complex indoor environments.
Smart Images

Figure CN120166526B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an indoor positioning technology, and more particularly to a RIS (Reconfigurable Intelligent Surface)-assisted indoor fingerprint positioning method based on a residual neural network. Background Art
[0002] Wireless positioning, as an important research direction in the field of array signal processing, aims to determine the spatial location or propagation direction of wireless signal sources. According to the different positioning environments, wireless positioning technologies are divided into outdoor and indoor positioning. The Global Navigation Satellite System (GNSS) can provide good positioning accuracy in outdoor environments, but in complex and dynamic indoor environments, its performance usually significantly degrades or even fails. Traditional indoor positioning technologies usually rely on the trilateration principle and calculate through measurement data of multiple transmitting nodes such as Time of Arrival (ToA), Angle of Arrival (AoA), and Received Signal Strength (RSS). However, these technologies are vulnerable to the effects of multipath and Non-Line of Sight (NLoS) signals and require precise time synchronization. In contrast, the fingerprint positioning method estimates the user's location by matching the pre-constructed fingerprint database, and has the advantages of low computational complexity, strong robustness, and high generalization ability.
[0003] In the field of wireless positioning, the Reconfigurable Intelligent Surface (RIS) has gradually received extensive attention due to its unique advantages. The advantages of RIS mainly include two aspects: on the one hand, it can reconstruct the line-of-sight path of the transmission link and provide additional channel degrees of freedom; on the other hand, it has a low-cost hardware cost. Specifically, RIS is a digitally controlled metasurface composed of a number of low-cost passive reflection elements, and the passive reflection elements therein can independently adjust the amplitude or phase of the incident signal and reflect the signal to improve the performance of wireless communication. When the direct link is blocked by obstacles, the intelligent reflection surface can construct an equivalent propagation link of BaseStation (BS)-reflective surface-User Equipment (UE), reducing or eliminating the influence of obstacles between the base station and the user. At the same time, the intelligent reflection surface can effectively utilize rich multipath information, providing higher reliability and accuracy, and effectively reducing the deployment and maintenance costs. Therefore, the fingerprint positioning method based on the intelligent reflection surface can achieve high-precision positioning in an indoor environment with dense obstacles, showing broad application potential.
[0004] Traditional fingerprint positioning algorithms do not consider RIS, and the positioning accuracy is vulnerable to the influence of obstacles. Based on this, the literature "T. Wu, C. Pan, Y. Pan, et al. Fingerprint-based mmWave positioning system aided by reconfigurable intelligent surface[J]. IEEE Wireless Communications Letters, 2023, 12(8):1379-1383." (Reconfigurable intelligent surface-aided fingerprint-based millimeter-wave positioning system, IEEE Wireless Communications Letters) proposed a residual convolutional network regression (RCNR) learning algorithm based on intelligent reflecting surface. This algorithm takes into account the presence of obstacles, uses the uplink, uses a uniform planar array as the base station to receive signals, and proposes a new spatio-temporal channel response vector (STCRV) as the positioning fingerprint, achieving good estimation results. However, this algorithm has high requirements for the configuration of the base station and requires the base station to be equipped with multiple antennas; moreover, the positioning accuracy of this algorithm still needs to be improved. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a RIS-assisted indoor fingerprint positioning method based on a residual neural network, which has low requirements for the configuration of the base station, high positioning accuracy, and good robustness.
[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows: A RIS-assisted indoor fingerprint positioning method based on a residual neural network, characterized in that this method is applicable to a downlink multipath transmission SISO millimeter-wave positioning system assisted by RIS. In this system, a base station with a single antenna, a RIS, and a user with a single antenna are set up. A three-dimensional coordinate system is established with the center position of the RIS. The RIS is deployed on the X-O-Z plane. The base station is located in the far-field area of the RIS, and the indoor environment is located in the Fresnel near-field area of the RIS. The RIS-user link includes a line-of-sight path and multiple non-line-of-sight paths. There is a scatterer on each non-line-of-sight path, and the base station-RIS link only includes a line-of-sight path; this method includes the following steps:
[0007] Step 1: Define the signal received by the user on the nth subcarrier at time slot t as y(n,t); then use the matrix composed of the signals received by the user on N subcarriers in T U time slots as the user's received signal matrix Y U . Y U contains the three-dimensional positions of the base station, RIS, user, and scatterer; where n = 1, 2,..., N, N represents the total number of subcarriers, and t = 1, 2,..., T U . TU Denote the total number of time slots as Y U has a dimension of N×T U ;
[0008] Step 2: Take Y U as the original fingerprint data; then normalize the real and imaginary parts of the original fingerprint data to obtain a real-valued matrix V, where the dimension of V is 2×N×T U , the first dimension represents the number of channels. The first channel is obtained by normalizing the real part of the original fingerprint data, and the second channel is obtained by normalizing the imaginary part of the original fingerprint data;
[0009] Step 3: Within the indoor environment range, adjust the user's three-dimensional position, and simultaneously record the user's true three-dimensional position after each adjustment as a label; then, following the processes of Step 1 and Step 2, obtain the corresponding real-valued matrix after each adjustment in the same manner as a sample; then form a dataset from the samples and labels obtained after multiple adjustments; then divide the dataset into a training set and a validation set;
[0010] Step 4: Use the training set to perform offline training on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; at the same time, use the validation set to monitor the training process and evaluate the network performance; after training, save the network weights with the optimal metrics on the validation set to obtain the trained residual neural network model;
[0011] Step 5: Conduct an online test in a near-field indoor positioning scenario. Obtain the corresponding real-valued matrix according to the process of Step 2 from the real-time received signal matrix of the user as a test sample; then input the test sample into the trained residual neural network model to output the predicted three-dimensional position of the user.
[0012] In the said Step 1, , where, () T represents the transpose operation, h BR (n) represents the channel of the base station - RIS link at the nth subcarrier, diag( ) represents constructing a diagonal matrix, w t represents the phase shift vector of the RIS at time slot t, h RU (n) represents the channel of the RIS - user link at the nth subcarrier, s t (n) represents the signal transmitted by the base station at the nth subcarrier at time slot t, z t (n) represents the zero-mean additive Gaussian noise in y(n,t).
[0013] The channel h BR (n) of the base station - RIS link at the nth subcarrier is modeled as , the channel hRU (n) Modeled as , where α BR represents the complex gain of the line-of-sight path in the base station - RIS link, j represents the imaginary part, represents the time difference of arrival of the line-of-sight path in the base station - RIS link, △f represents the sub-carrier frequency spacing of the OFDM signal, P B represents the three-dimensional position of the base station, s = 0, 1, 2, …, N s , N s represents the number of non-line-of-sight paths in the RIS - user link. When s = 0, α RU,s represents the complex gain of the line-of-sight path in the RIS - user link. When s = 1, 2, …, N s α RU,s represents the complex gain of the sth non-line-of-sight path in the RIS - user link. When s = 0, represents the time difference of arrival of the line-of-sight path in the RIS - user link. When s = 1, 2, …, N s When represents the time difference of arrival of the sth non-line-of-sight path in the RIS - user link. When s = 0, P s = P0, P0 represents the three-dimensional position of the user. When s = 1, 2, …, N s When P s represents the three-dimensional position of the scatterer on the sth non-line-of-sight path in the RIS - user link, e(P B ) represents the steering vector of the RIS in the line-of-sight path of the base station - RIS link. When s = 0, e(P s ) represents the steering vector of the RIS in the line-of-sight path of the RIS - user link. When s = 1, 2, …, N s When s ) represents the steering vector of the RIS in the sth non-line-of-sight path of the RIS - user link.
[0014] In the said step 2, the first channel V 1,:,: is , the second channel V 2,:,: is , where, Re{} represents taking the real part, Im{} represents taking the imaginary part, represents the two-norm operation of a vector or matrix.
[0015] The specific process of the said step 4 is as follows:
[0016] Step 4.1: Randomly divide all the training data in the training set into multiple batches, so that each batch contains batch size groups of training data. Among them, each group of training data contains a sample and its corresponding label;
[0017] Step 4.2: Take one batch from the training set, use all the training data in this batch as the input of the residual neural network, input it into the residual neural network for processing, and obtain the predicted 3D positions of the users corresponding to each group of training data in this batch;
[0018] Step 4.3: Calculate the mean squared error loss between the predicted 3D positions of the users and the labels; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of this batch on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm;
[0019] Step 4.4: Repeat the process of Step 4.2 to Step 4.3 until all batches of the training set have been used to train the residual neural network once;
[0020] Step 4.5: Randomly divide all the validation data in the validation set into multiple batches, with each batch containing batch size groups of validation data, where each group of validation data contains a sample and its corresponding label;
[0021] Step 4.6: Take one batch from the validation set; then input all the validation data in this batch into the residual neural network trained with the training set to obtain the predicted 3D positions of the users corresponding to each group of validation data in this batch; then calculate the mean squared error loss between the predicted 3D positions of the users and the labels;
[0022] Step 4.7: Repeat the process of Step 4.6 until all batches of the validation set have been processed once, and then calculate the average mean squared error loss of the validation set;
[0023] Step 4.8: Adopt an exponentially decreasing learning rate reduction strategy, repeat the process of Step 4.1 to Step 4.7, end after executing Num epochs in total, and save the network weights with the best validation set metrics to obtain the trained residual neural network model.
[0024] In the said Step 4.8, when adopting the exponentially decreasing learning rate reduction strategy, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original every M epochs.
[0025] In step 4, the residual neural network selects ResNet, which includes a basic convolution module, a deep convolution module with a residual structure, and a regression module. The basic convolution module is composed of a first convolutional layer, a first batch normalization layer, and a first PReLU activation function connected in sequence. The deep convolution module is composed of a first residual module, a second residual module, and a third residual module with the same structure connected in sequence. The regression module is composed of a Flatten layer and a fully connected layer connected in sequence. After passing the sample through the first convolutional layer, the first batch normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained. After passing the basic feature map through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained. After passing the deep feature map through the Flatten layer, a feature vector is obtained, and after passing the feature vector through the fully connected layer, the predicted three-dimensional position of the user corresponding to the sample is obtained.
[0026] The first residual module, the second residual module, and the third residual module each include a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map obtained by passing the feature map received by the corresponding residual module through the fourth convolutional layer is subjected to an element-wise addition operation with the feature map obtained by passing the feature map received by the corresponding residual module through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer in sequence. The feature map obtained after the element-wise addition operation is passed through the third PReLU activation function and used as the feature map output by the corresponding residual module.
[0027] The convolutional kernel size of the first convolutional layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the first residual module are 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the second residual module are 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the third residual module are 3×3, the stride is 4, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolutional kernel size of the fourth convolutional layer in the first residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel size of the fourth convolutional layer in the second residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel size of the fourth convolutional layer in the third residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The size of the sample is 2×N×T. U, the size of the basic feature map is 64×N×T U , the size of the deep feature map is 512×N / 8×T U / 8, the dimension of the feature vector is (8×N×T U )×1.
[0028] Compared with the prior art, the advantages of the present invention are as follows:
[0029] 1) The method of the present invention proposes a new type of CSI fingerprint, that is, the received signal matrix of the user obtained is used as the original fingerprint data, and the original fingerprint data contains angle and distance features related to coordinates with position information. This new type of CSI fingerprint can better improve the positioning accuracy.
[0030] 2) The method of the present invention utilizes the reflection and regulation ability of RIS for signals, can enhance the signal propagation effect, increase the signal path diversity, and thus improve the positioning accuracy.
[0031] 3) The method of the present invention uses a single-antenna base station, which does not reduce the positioning accuracy while reducing the complexity. In the simulation, compared with the RCNR method using a multi-antenna base station, the positioning accuracy is equivalent, which fully demonstrates the superiority of the method of the present invention.
[0032] 4) The method of the present invention regards indoor three-dimensional positioning as a regression problem, and uses a residual neural network to perform parameter learning for indoor three-dimensional positioning. The residual neural network is data-driven and does not depend on the signal model, so it is applicable to situations with complex models or harsh transmission environments.
[0033] 5) The method of the present invention has a small estimation error for the three-dimensional position of the user, and has strong generalization ability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic diagram of a RIS-assisted downlink multipath transmission SISO millimeter-wave positioning system applied to the method of the present invention;
[0035] Figure 2 is the overall implementation block diagram of the method of the present invention;
[0036] Figure 3 is the structural diagram of the residual neural network used in the method of the present invention;
[0037] Figure 4 is the structural diagram of the residual module in the residual neural network used in the method of the present invention;
[0038] Figure 5When the signal-to-noise ratio is 25 dB, and the total number of subcarriers and the total number of time slots are both 32, the predicted three-dimensional position scatter plots of the three users are obtained by using the trained residual neural network model to make 500 predictions on the three-dimensional positions of the three users respectively;
[0039] Figure 6 It is a schematic diagram showing the change of the cumulative distribution function (CDF) of the positioning error with the positioning error under different signal-to-noise ratios (SNR) obtained by using the method of the present invention;
[0040] Figure 7 It is a schematic diagram comparing the change of the cumulative distribution function (CDF) of the positioning error with the positioning error between the method of the present invention and the existing method. Detailed implementation manners
[0041] The present invention will be further described in detail below in conjunction with the embodiments with reference to the drawings.
[0042] A RIS-assisted indoor fingerprint positioning method based on a residual neural network proposed by the present invention is applicable to a downlink multipath transmission SISO (Single-Input Single-Output) millimeter-wave positioning system assisted by RIS. As Figure 1 shown, in this system, a single-antenna base station (BS), a RIS, and a single-antenna user (UE) are provided. A three-dimensional coordinate system is established with the center position of the RIS. The RIS is deployed on the X-O-Z plane, and the spacing d between the reflection elements in the RIS is half a wavelength , represents the signal wavelength. The base station is located in the far-field region of the RIS, and the indoor environment is located in the Fresnel near-field region of the RIS. The RIS-user link includes a line-of-sight (LOS) path and multiple non-line-of-sight (NLOS) paths. There is a scatterer on each non-line-of-sight path, and the base station-RIS link only includes a line-of-sight path. Among them, the three-dimensional positions of the base station and the RIS are known, and the three-dimensional positions of the user and the scatterer are unknown. Figure 1 in refers to the azimuth angle, which is defined as the angle between the projection of the incident direction on the X-O-Y plane and the positive direction of the X axis. refers to the elevation angle, which is defined as the angle between the incident direction and the positive direction of the Z axis. As Figure 2 shown, this method includes the following steps:
[0043] Step 1: Define the signal received by the user on the nth subcarrier at time slot t as y(n,t); then use the matrix composed of the signals received by the user on N subcarriers in T U time slots as the received signal matrix Y of the user U , , Y UIt includes the three-dimensional positions of the base station, RIS, user, and scatterer; where n = 1, 2, …, N, N represents the total number of subcarriers, and t = 1, 2, …, T U , T U represents the total number of time slots, and the dimension of Y U is N × T U .
[0044] In this embodiment, in step 1, , where ( ) T represents the transpose operation, h BR (n) represents the channel of the base station - RIS link at the nth subcarrier, diag( ) represents constructing a diagonal matrix, and w t represents the phase shift vector of the RIS at time slot t, and the dimension of w t is N R ×1, , correspondingly represents the phase shifts of the 1st, 2nd, …, Nth reflection elements of the RIS at time slot t, and N R represents the number of reflection elements included in the RIS, and N R = N R ×N x , and N z represents the number of reflection elements included in the RIS in the X-axis direction, and N x represents the number of reflection elements included in the RIS in the Z-axis direction, and h z (n) represents the channel of the RIS - user link at the nth subcarrier, s RU (n) represents the signal transmitted by the base station at the nth subcarrier at time slot t, and z t (n) represents the zero-mean additive Gaussian noise in y(n, t). t (n) represents the zero-mean additive Gaussian noise in y(n, t).
[0045] In this embodiment, the channel h BR (n) of the base station - RIS link at the nth subcarrier is modeled as , and the channel h RU (n) of the RIS - user link at the nth subcarrier is modeled as , where the dimensions of h BR (n) and h RU (n) are both N R ×1, α BR represents the complex gain of the line-of-sight path in the base station - RIS link, exp( ) represents the exponential function with the natural constant as the base, the natural constant is 2.71, and j represents the imaginary part, represents the time difference of arrival of the line-of-sight path in the base station - RIS link, △f represents the subcarrier frequency spacing of the OFDM signal, and P BRepresents the three-dimensional position of the base station, P B =[x B ,y B ,z B T , x B ,y B ,z B Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the base station, s = 0, 1, 2, …, N s , N s Represents the number of non-line-of-sight paths in the RIS-user link. When s = 0, α RU,s Represents the complex gain of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When α RU,s Represents the complex gain of the s-th non-line-of-sight path in the RIS-user link. When s = 0 Represents the time difference of arrival of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When Represents the time difference of arrival of the s-th non-line-of-sight path in the RIS-user link. Assume that the positioning distance is within the Fresnel zone of the RIS, that is , D represents the maximum aperture of the RIS, , , When s = 0, d R,s Represents the Euclidean distance between the user and the RIS, s = 1, 2, …, N s When d R,s Represents the Euclidean distance between the scatterer on the s-th non-line-of-sight path in the RIS-user link and the RIS. When s = 0, d RU,s Represents the Euclidean distance under the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When d RU,s Represents the Euclidean distance under the s-th non-line-of-sight path in the RIS-user link. c represents the speed of light, Represents the two-norm operation of a vector or matrix, P R Represents the three-dimensional position of the central reflection element in the RIS, P R =[x R ,y R ,z R T , x R ,y R ,z R Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the central reflection element in the RIS. When s = 0, P s =P0, P0 represents the three-dimensional position of the user, P0 = [x0, y0, z0] T , x0, y 0, z0 correspondingly represents the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the user, s = 1, 2, …, N s When P s represents the three-dimensional position of the scatterer on the s-th non-line-of-sight path in the RIS-user link, P s = [x s , y s , z s T , x s , y s , z s correspondingly represents the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the scatterer on the s-th non-line-of-sight path in the RIS-user link, represents the phase difference of the base station-RIS link, represents the phase difference of the RIS-user link, e(P B ) represents the steering vector of the RIS in the line-of-sight path in the base station-RIS link. When s = 0, e(P s ) represents the steering vector of the RIS in the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When e(P s ) represents the steering vector of the RIS in the s-th non-line-of-sight path in the RIS-user link, , given that the indoor environment is located in the Fresnel near-field region of the RIS, due to geometric relationships, e r (P B ) is described as , e r (P s ) is described as , r = 1, 2, …, N R , P r represents the three-dimensional position of the r-th reflecting element in the RIS.
[0046] Step 2: Since the obtained Y U for users with different three-dimensional positions is unique, Y U can be regarded as a new type of CSI (Channel State Information) fingerprint of the user, which can fully characterize the information of the indoor complex multipath environment. According to the definition of the fingerprint positioning algorithm, Y U can be used to estimate the three-dimensional position of the user. Take Y U as the original fingerprint data; then normalize the real part and imaginary part of the original fingerprint data to obtain a real-valued matrix V, where the dimension of V is 2 × N × T U , the first dimension represents the number of channels. The first channel is obtained by normalizing the real part of the original fingerprint data, and the second channel is obtained by normalizing the imaginary part of the original fingerprint data. The first channel V 1,:,: is , and the second channel V 2,:,: is , Re{} represents taking the real part, Im{} represents taking the imaginary part, represents the two-norm operation of a vector or matrix.
[0047] Step 3: Within the indoor environment range, adjust the user's three-dimensional position, and simultaneously record the user's true three-dimensional position after each adjustment, which is used as a label; then, following the processes of Step 1 and Step 2, obtain the corresponding real-value matrix after each adjustment in the same way, which is used as a sample; then form a dataset from the samples and labels corresponding to multiple adjustments; afterwards, divide the dataset into a training set and a validation set. In implementation, 80% of the dataset can form the training set, and 20% can form the validation set.
[0048] Step 4: Use the training set to perform offline training on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; at the same time, use the validation set to monitor the training process and evaluate the network performance; after the training is completed, save the network weights with the optimal metrics of the validation set to obtain the trained residual neural network model.
[0049] In this specific embodiment, the specific process of Step 4 is as follows:
[0050] Step 4.1: Randomly divide all the training data in the training set into multiple batches, such that each batch contains batch size groups of training data, where each group of training data contains a sample and its corresponding label.
[0051] Step 4.2: Take one batch from the training set, use all the training data in this batch as the input to the residual neural network, and input it into the residual neural network for processing to obtain the predicted three-dimensional position of the user corresponding to each group of training data in this batch.
[0052] Step 4.3: Calculate the mean square error (MSE) loss between the predicted three-dimensional position of the user and the label; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of this batch on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm.
[0053] Step 4.4: Repeat the process of Step 4.2 to Step 4.3 until all batches of the training set have been used to train the residual neural network once.
[0054] Step 4.5: Randomly divide all the validation data in the validation set into multiple batches, such that each batch contains batch size groups of validation data, where each group of validation data contains a sample and its corresponding label.
[0055] Step 4.6: Take one batch from the validation set; then input all the validation data in this batch into the residual neural network trained using the training set to obtain the predicted three-dimensional positions of the users corresponding to each group of validation data in this batch; then calculate the mean squared error loss between the predicted three-dimensional positions of the users and the labels.
[0056] Step 4.7: Repeat the process of Step 4.6 until all batches in the validation set have been processed once, and then calculate the average mean squared error loss of the validation set.
[0057] Step 4.8: Adopt an exponentially decreasing learning rate reduction strategy, repeat the process of Steps 4.1 to 4.7, end after executing Num (e.g., take Num = 60) epochs, and save the network weights with the best metrics on the validation set to obtain the trained residual neural network model, where when adopting the exponentially decreasing learning rate reduction strategy, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original every M epochs, and in the embodiment, take M = 10.
[0058] In this specific embodiment, the residual neural network selects ResNet, as Figure 3 shown, which includes a basic convolutional module, a deep convolutional module adopting a residual structure, and a regression module. The deep convolutional module can mine the deep features of the data, and the regression module realizes position regression. The basic convolutional module is composed of a first convolutional layer, a first batch normalization layer, and a first PReLU activation function connected in sequence. The deep convolutional module is composed of a first residual module, a second residual module, and a third residual module with the same structure connected in sequence. The regression module is composed of a Flatten (unfolding) layer and a fully connected layer connected in sequence; after passing the sample through the first convolutional layer, the first batch normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained; after passing the basic feature map through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained; after passing the deep feature map through the Flatten layer, a feature vector is obtained, and after passing the feature vector through the fully connected layer, the predicted three-dimensional position of the user corresponding to the sample is obtained.
[0059] In this specific embodiment, as Figure 4As shown, the first residual module, the second residual module, and the third residual module each include a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map obtained after passing the feature map received by the residual module where it is located through the fourth convolutional layer is element-wise added to the feature map obtained by sequentially passing the feature map received by the residual module where it is located through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer. The feature map obtained after the element-wise addition operation is passed through the third PReLU activation function and used as the feature map output by the residual module where it is located.
[0060] In this specific embodiment, the convolutional kernel size of the first convolutional layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the first residual module are 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the second residual module are 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the third residual module are 3×3, the stride is 4, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolutional kernel size of the fourth convolutional layer in the first residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel size of the fourth convolutional layer in the second residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel size of the fourth convolutional layer in the third residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512; the size of the sample is 2×N×T U , the size of the basic feature map is 64×N×T U , the size of the deep feature map is 512×N / 8×T U / 8, the dimension of the feature vector is (8×N×T U )×1.
[0061] In the basic convolutional module, the first convolutional layer performs preliminary feature extraction on the sample, and the feature map output by this process can be expressed by the following formula: F o,1 =ω1*F i,1 +b1, F o,1 represents the feature map output by the first convolutional layer, F i,1Denote the feature map of the input of the first convolutional layer, ω1 represents the weights of the convolutional kernel of the first convolutional layer, b1 represents the bias of the convolutional kernel of the first convolutional layer, and * represents the convolution operation; after the convolution operation of the first convolutional layer, the first batch normalization layer (Batch Normalization, BN) is used to accelerate network training; finally, the first PReLU (Parametric Rectified Linear Unit) activation function is adopted. Compared with the traditional ReLU (Rectified Linear Unit) activation function, the PReLU activation function can adaptively adjust the output when the input value is negative, rather than directly outputting zero, thereby alleviating the problem of neuron "death" that may occur in the negative value region of the ReLU activation function to a certain extent. The basic feature map obtained after passing through the basic convolution module can be expressed by the following formula: F b =max(0,BN(F o,1 ))+ρ×min(BN(F o,1 ),0), where F b represents the basic feature map, max( ) represents taking the maximum value, min( ) represents taking the minimum value, ρ represents the learnable parameter, and BN( ) represents the batch normalization operation.
[0062] There are certain limitations in the feature extraction ability of the shallow network, and it is difficult to meet the requirements of high-precision positioning. However, simply increasing the network depth will not only increase the training difficulty but also may cause problems such as gradient disappearance or degradation. Therefore, in the deep convolutional module, a residual learning strategy is introduced.
[0063] After the basic convolution module and the deep convolution module complete feature extraction, the regression module is responsible for the positioning task, that is, estimating the three-dimensional position of the user.
[0064] Step 5: Conduct an online test in the near-field indoor positioning scenario. Obtain the corresponding real-value matrix according to the process of Step 2 for the real-time obtained received signal matrix of the user and use it as a test sample; then input the test sample into the trained residual neural network model to output the predicted three-dimensional position of the user , which respectively represent the predicted X-axis coordinate position, predicted Y-axis coordinate position, and predicted Z-axis coordinate position of the user.
[0065] To further illustrate the feasibility and effectiveness of the method of the present invention, a simulation experiment is conducted on the method of the present invention.
[0066] The simulation parameters are shown in Table 1.
[0067] Table 1 Simulation Parameters
[0068]
[0069] Consider an indoor space with dimensions of 16×8×3m 3 In the indoor space, the length and width are divided at intervals of 0.2 meters. At the same time, to balance the dataset size, raw fingerprint data is only collected at heights of 0.2 meters, 0.4 meters, and 0.6 meters. Under signal-to-noise ratios of -5dB, 0dB, 5dB, 10dB, and 15dB, a total of 81×41×3×5 = 49815 raw fingerprint data are collected. After traversing 45 times, a total of 45×49815 = 2241675 raw fingerprint data are collected to form the dataset.
[0070] The number of Monte Carlo runs for the simulation experiment is 500.
[0071] Figure 5 A scatter plot of the predicted three-dimensional positions of three users is given. The prediction is made 500 times for each user's three-dimensional position using the trained residual neural network model when the signal-to-noise ratio is 25dB, and the total number of subcarriers and the total number of time slots are both 32. From Figure 5 it can be seen that the predicted three-dimensional positions of the users obtained by the method of the present invention are close to the true three-dimensional positions of the users, indicating that the method of the present invention can achieve good estimation.
[0072] Figure 6 A schematic diagram showing the variation of the cumulative distribution function (CDF) of the positioning error with respect to the positioning error is given for different signal-to-noise ratios (SNR) obtained by using the method of the present invention. From Figure 6 it can be seen that as the signal-to-noise ratio increases, the CDF rises significantly, demonstrating that the improvement of the signal-to-noise ratio has a significant impact on the algorithm.
[0073] Figure 7 A schematic diagram comparing the variation of the cumulative distribution function (CDF) of the positioning error with respect to the positioning error for the method of the present invention (ResNet) and existing methods is given. The existing methods include the RCNR method and the classic CNN method. The RCNR method is from the literature T. Wu, C. Pan, Y. Pan, et al. Fingerprint-based mmWave positioning system aided by reconfigurable intelligent surface[J]. IEEE Wireless Communications Letters, 2023, 12(8):1379-1383. (Reconfigurable intelligent surface-aided fingerprint-based millimeter-wave positioning system, IEEE Wireless Communications Letters). From Figure 7 it can be seen that the positioning performance of the method of the present invention is significantly better than that of the existing methods.
Claims
1. A RIS-assisted indoor fingerprint positioning method based on a residual neural network, characterized in that It includes the following steps: Step 1: Define the signal received by the user on the n-th subcarrier at time slot t as y(n,t); then use the matrix composed of the signals received by the user on N subcarriers in T U time slots as the user's received signal matrix Y U , where Y U contains the three-dimensional positions of the base station, RIS, user, and scatterers; where n = 1, 2, …, N, N represents the total number of subcarriers, and t = 1, 2, …, T U , and T U represents the total number of time slots, and the dimension of Y U is N×T U ; In the said step 1, , where, ( ) T represents the transpose operation, h BR (n) represents the channel of the base station - RIS link on the nth sub - carrier, diag( ) represents constructing a diagonal matrix, w t represents the phase - shift vector of the RIS at time slot t, h RU (n) represents the channel of the RIS - user link on the nth sub - carrier, s t (n) represents the signal transmitted by the base station on the nth sub - carrier at time slot t, z t (n) represents the zero - mean additive Gaussian noise in y(n,t); The channel \(h\) of the base station - RIS link at the \(n\)th sub - carrier BR is modeled as , and the channel \(h\) of the RIS - user link at the \(n\)th sub - carrier RU is modeled as , where \(\alpha\) BR represents the complex gain of the line - of - sight path in the base station - RIS link, \(j\) represents the imaginary part, represents the time - difference of arrival of the line - of - sight path in the base station - RIS link, \(\Delta f\) represents the sub - carrier frequency spacing of the OFDM signal, \(P\) B represents the three - dimensional position of the base station, \(s = 0,1,2,\cdots,N\) s , \(N\) s represents the number of non - line - of - sight paths in the RIS - user link. When \(s = 0\), \(\alpha\) RU,s represents the complex gain of the line - of - sight path in the RIS - user link. When \(s = 1,2,\cdots,N\) s , \(\alpha\) RU,s represents the complex gain of the \(s\)th non - line - of - sight path in the RIS - user link. When \(s = 0\), represents the time - difference of arrival of the line - of - sight path in the RIS - user link. When \(s = 1,2,\cdots,N\) s , represents the time - difference of arrival of the \(s\)th non - line - of - sight path in the RIS - user link. When \(s = 0\), \(P\) s =P0, \(P0\) represents the three - dimensional position of the user. When \(s = 1,2,\cdots,N\) s , \(P\) s represents the three - dimensional position of the scatterer on the \(s\)th non - line - of - sight path in the RIS - user link. \(e(P\) B ) represents the steering vector of the RIS in the line - of - sight path of the base station - RIS link. When \(s = 0\), \(e(P\) s ) represents the steering vector of the RIS in the line - of - sight path of the RIS - user link. When \(s = 1,2,\cdots,N\) s , \(e(P\) s ) represents the steering vector of the RIS in the \(s\)th non - line - of - sight path of the RIS - user link; Step 2: Take Y U as the original fingerprint data; then perform normalization processing on the real part and the imaginary part of the original fingerprint data to obtain a real-value matrix V, where the dimension of V is 2×N×T U , the first dimension represents the number of channels, the first channel is obtained by normalizing the real part of the original fingerprint data, and the second channel is obtained by normalizing the imaginary part of the original fingerprint data; In the said step 2, the first channel V 1,:,: is , and the second channel V 2,:,: is , where Re{} represents taking the real part, Im{} represents taking the imaginary part, represents the two-norm operation of a vector or matrix; Step 3: Within the indoor environment range, adjust the user's three-dimensional position, and record the user's true three-dimensional position after each adjustment as a label; then, following the processes of Step 1 and Step 2, obtain the corresponding real-value matrix after each adjustment in the same way as a sample; then form a data set from the samples and labels corresponding to multiple adjustments; after that, divide the data set into a training set and a validation set; Step 4: Use the training set to perform offline training on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; at the same time, use the validation set to monitor the training process and evaluate the network performance; after training is completed, save the network weights with the optimal metrics of the validation set to obtain a trained residual neural network model; Step 5: Conduct online testing in a near-field indoor positioning scenario. Obtain the corresponding real-value matrix from the received signal matrix of the user obtained in real time following the process of Step 2 as a test sample; then input the test sample into the trained residual neural network model to output the predicted three-dimensional position of the user.
2. The RIS-assisted indoor fingerprint positioning method based on the residual neural network according to claim 1, wherein The specific process of Step 4 is as follows: Step 4.1: Randomly divide all the training data in the training set into multiple batches, so that each batch contains batch size groups of training data. Among them, each group of training data contains a sample and its corresponding label; Step 4.2: Take one batch from the training set, and use all the training data in this batch as the input of the residual neural network, input it into the residual neural network for processing, and obtain the predicted three-dimensional position of the user corresponding to each group of training data in this batch; Step 4.3: Calculate the mean squared error loss between the predicted three-dimensional position of the user and the label; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of this batch on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; Step 4.4: Repeat the process of Step 4.2 to Step 4.3 until all batches of the training set have been trained on the residual neural network once; Step 4.5: Randomly divide all the validation data in the validation set into multiple batches, so that each batch contains batch size groups of validation data. Among them, each group of validation data contains a sample and its corresponding label; Step 4.6: Take one batch from the validation set; then input all the validation data in this batch into the residual neural network trained with the training set to obtain the predicted three-dimensional position of the user corresponding to each group of validation data in this batch; then calculate the mean squared error loss between the predicted three-dimensional position of the user and the label; Step 4.7: Repeat the process of Step 4.6 until all batches of the validation set have been processed once, and then calculate the average mean squared error loss of the validation set; Step 4.8: Adopt an exponentially decreasing learning rate reduction strategy, repeat the process of Step 4.1 to Step 4.7, end after executing Num epochs in total, and save the network weights with the optimal metrics of the validation set to obtain a trained residual neural network model.
3. The RIS-assisted indoor fingerprint positioning method based on a residual neural network according to claim 2, wherein In step 4.8, when using the exponential decay learning rate reduction strategy, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original value every M epochs.
4. The RIS-assisted indoor fingerprint positioning method based on a residual neural network according to claim 1, wherein In step 4, the residual neural network selects ResNet, which includes a basic convolution module, a deep convolution module with a residual structure, and a regression module. The basic convolution module is composed of a first convolutional layer, a first batch normalization layer, and a first PReLU activation function connected in sequence. The deep convolution module is composed of a first residual module, a second residual module, and a third residual module with the same structure connected in sequence. The regression module is composed of a Flatten layer and a fully connected layer connected in sequence. After passing the sample through the first convolutional layer, the first batch normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained. After passing the basic feature map through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained. After passing the deep feature map through the Flatten layer, a feature vector is obtained, and after passing the feature vector through the fully connected layer, the predicted three-dimensional position of the user corresponding to the sample is obtained.
5. The RIS-assisted indoor fingerprint positioning method based on a residual neural network according to claim 4, wherein The first residual module, the second residual module, and the third residual module all include a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map obtained by passing the feature map received by the residual module through the fourth convolutional layer is subjected to an element-wise addition operation with the feature map obtained by passing the feature map received by the residual module through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer in sequence. The feature map obtained after the element-wise addition operation is passed through the third PReLU activation function as the feature map output by the residual module.
6. The RIS-assisted indoor fingerprint positioning method based on a residual neural network according to claim 5, wherein The convolution kernel size of the first convolutional layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolution kernel sizes of the second and third convolutional layers in the first residual module are 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolution kernel sizes of the second and third convolutional layers in the second residual module are 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolution kernel sizes of the second and third convolutional layers in the third residual module are 3×3, the stride is 4, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolution kernel size of the fourth convolutional layer in the first residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolution kernel size of the fourth convolutional layer in the second residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolution kernel size of the fourth convolutional layer in the third residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512; the size of the sample is 2×N×T U , the size of the basic feature map is 64×N×T U , the size of the deep feature map is 512×N / 8×T U / 8, the dimension of the feature vector is (8×N×T U )×1.
Citation Information
Patent Citations
Large-scale MIMO fingerprint positioning method based on complex neural network
CN112995892A
Method for realizing RIS self-adaptive reconstruction of multi-user channel based on Vision Transform network
CN119675711A