RIS-assisted indoor fingerprint positioning method based on residual neural network

By introducing RIS assistive technology based on residual neural network in indoor positioning, RIS is used to reflect and regulate signals, and a new CSI fingerprint is constructed, which solves the problems of low positioning accuracy and high base station configuration requirements in the existing technology, and achieves high-precision and robust indoor three-dimensional positioning.

CN120166526AActive Publication Date: 2025-06-17NINGBO UNIV

Patent Information

Application Number
CN202510650024.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing indoor positioning technology has low positioning accuracy in complex environments, is susceptible to obstacles, and has high requirements for base station configuration.

Method used

Using RIS-assisted indoor fingerprint positioning method based on residual neural network, the signal is reflected and regulated through RIS, and a new CSI fingerprint is constructed to perform three-dimensional positioning estimation using a single-antenna base station and multiple non-sight paths in the RIS-user link.

Benefits of technology

It realizes high-precision indoor three-dimensional positioning under low base station configuration requirements, has good robustness and generalization capabilities, and has small positioning errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166526A_ABST
    Figure CN120166526A_ABST
Patent Text Reader

Abstract

The invention discloses an RIS-assisted indoor fingerprint positioning method based on a residual neural network, which is suitable for an RIS-assisted downlink multipath transmission SISO millimeter wave positioning system, and the system comprises a base station, an RIS and a user. The method is realized through the following steps: defining a received signal matrix of a user, including three-dimensional positions of a base station, an RIS, the user and a scatterer; normalizing the received signal matrix of the user to obtain a real value matrix as a sample; adjusting a three-dimensional position of a user in an indoor environment, recording a real three-dimensional position as a label, constructing a data set, and dividing the data set into a training set and a verification set; performing off-line training on the residual neural network by using the training set, optimizing network parameters and storing an optimal weight; in an actual positioning scene, processing a real-time received signal matrix, inputting the processed real-time received signal matrix into the trained model, and outputting a predicted three-dimensional position of a user; according to the invention, the RIS and the residual neural network are combined, the precision and robustness of indoor positioning are improved, and rapid positioning is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an indoor positioning technology, and more particularly to a RIS (Reconfigurable Intelligent Surface)-assisted indoor fingerprint positioning method based on a residual neural network. Background Art

[0002] Wireless positioning, as an important research direction in the field of array signal processing, aims to determine the spatial position or propagation direction of a wireless signal source. According to the different positioning environments, wireless positioning technologies are divided into outdoor and indoor positioning. The Global Navigation Satellite System (GNSS) can provide good positioning accuracy in outdoor environments, but in complex and dynamic indoor environments, its performance usually drops significantly or even fails. Traditional indoor positioning technologies usually rely on the trilateration principle and calculate through measurement data of multiple transmitting nodes such as Time of Arrival (ToA), Angle of Arrival (AoA), and Received Signal Strength (RSS). However, these technologies are vulnerable to the effects of multipath and Non-Line of Sight (NLoS) signals and require precise time synchronization. In contrast, the fingerprint positioning method estimates the user's position by matching a pre-constructed fingerprint database, and has advantages such as low computational complexity, strong robustness, and high generalization ability.

[0003] In the field of wireless positioning, the intelligent reflecting surface (RIS) has gradually attracted wide attention due to its unique advantages. The advantages of RIS mainly include two aspects: on the one hand, it can reconstruct the line-of-sight path of the transmission link and provide additional channel degrees of freedom; on the other hand, it has a low-cost hardware cost. Specifically, RIS is a digitally controlled metasurface composed of a number of low-cost passive reflecting elements, and the passive reflecting elements therein can independently adjust the amplitude or phase of the incident signal and reflect the signal to improve the performance of wireless communication. When the direct link is blocked by obstacles, the intelligent reflecting surface can construct an equivalent propagation link of Base Station (BS)-reflecting surface-User Equipment (UE) to reduce or eliminate the influence of obstacles between the base station and the user. At the same time, the intelligent reflecting surface can effectively utilize rich multipath information to provide higher reliability and accuracy, and effectively reduce the deployment and maintenance costs. Therefore, the fingerprint positioning method based on the intelligent reflecting surface can achieve high-precision positioning in an indoor environment with dense obstacles and shows broad application potential.

[0004] Traditional fingerprint positioning algorithms do not consider RIS, and the positioning accuracy is vulnerable to the influence of obstacles. Based on this, the literature T. Wu, C. Pan, Y. Pan, et al. Fingerprint-based mmWave positioning system aided by reconfigurable intelligent surface[J]. IEEE Wireless Communications Letters, 2023, 12(8):1379-1383. (Reconfigurable intelligent surface-aided fingerprint-based millimeter-wave positioning system, IEEE Wireless Communications Letters) proposed a residual convolutional network regression (RCNR) learning algorithm based on intelligent reflecting surface. This algorithm considers the existence of obstacles, adopts the uplink, uses a uniform planar array as the base station to receive signals, and proposes a new spatio-temporal channel response vector (STCRV) as the positioning fingerprint, obtaining good estimation results. However, this algorithm has high requirements for the configuration of the base station, requiring the base station to be equipped with multiple antennas; and the positioning accuracy of this algorithm still needs to be improved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a RIS-aided indoor fingerprint positioning method based on residual neural network, which has low requirements for the configuration of the base station, high positioning accuracy, and good robustness.

[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows: A RIS-aided indoor fingerprint positioning method based on residual neural network, characterized in that this method is applicable to a downlink multipath transmission SISO millimeter-wave positioning system aided by RIS. In this system, there is a base station with a single antenna, a RIS, and a user with a single antenna. A three-dimensional coordinate system is established with the center position of the RIS. The RIS is deployed on the X-O-Z plane. The base station is located in the far-field area of the RIS, and the indoor environment is located in the Fresnel near-field area of the RIS. The RIS-user link includes a line-of-sight path and multiple non-line-of-sight paths. There is a scatterer on each non-line-of-sight path. The base station-RIS link only includes a line-of-sight path; this method includes the following steps: Step 1: Define the signal received by the user on the nth subcarrier at time slot t as y(n,t); then use the matrix composed of the signals received by the user on N subcarriers in T U time slots as the user's received signal matrix Y U , Y U which contains the three-dimensional positions of the base station, RIS, user, and scatterer; where n = 1, 2,..., N, N represents the total number of subcarriers, and t = 1, 2,..., T U , T UDenote the total number of time slots as Y U has the dimension of N×T U ; Step 2: Take Y U as the original fingerprint data; then normalize the real and imaginary parts of the original fingerprint data to obtain a real-valued matrix V, where the dimension of V is 2×N×T U , the first dimension represents the number of channels, the first channel is obtained by normalizing the real part of the original fingerprint data, and the second channel is obtained by normalizing the imaginary part of the original fingerprint data; Step 3: Within the indoor environment range, adjust the user's three-dimensional position, and at the same time record the user's true three-dimensional position after each adjustment and use it as a label; then follow the processes of Step 1 and Step 2 to obtain the corresponding real-valued matrix after each adjustment in the same way and use it as a sample; then form a dataset with the samples and labels obtained after multiple adjustments; afterwards, divide the dataset into a training set and a validation set; Step 4: Use the training set to perform offline training on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; at the same time, use the validation set to monitor the training process and evaluate the network performance; save the network weights with the optimal validation set metrics after training to obtain the trained residual neural network model; Step 5: Conduct online testing in the near-field indoor positioning scenario, obtain the corresponding real-valued matrix according to the process of Step 2 for the received signal matrix of the user obtained in real time, and use it as a test sample; then input the test sample into the trained residual neural network model to output the predicted three-dimensional position of the user.

[0007] In the said Step 1, , where, ( ) T represents the transpose operation, h BR (n) represents the channel of the base station - RIS link at the nth subcarrier, diag( ) represents constructing a diagonal matrix, w t represents the phase shift vector of the RIS at time slot t, h RU (n) represents the channel of the RIS - user link at the nth subcarrier, s t (n) represents the signal transmitted by the base station at the nth subcarrier at time slot t, z t (n) represents the zero-mean additive Gaussian noise in y(n,t).

[0008] The channel h BR of the base station - RIS link at the nth subcarrier is modeled as , the channel h RU of the RIS - user link at the nth subcarrier is modeled as , where, α BRDenote the complex gain of the line-of-sight path in the base station-RIS link. j represents the imaginary part. Denote the time difference of arrival of the line-of-sight path in the base station-RIS link. △f represents the subcarrier frequency spacing of the OFDM signal. P B Denote the three-dimensional position of the base station, s = 0, 1, 2, …, N s , N s Denote the number of non-line-of-sight paths in the RIS-user link. When s = 0, α RU,s Denote the complex gain of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When s = 1, 2, …, N, α RU,s Denote the complex gain of the sth non-line-of-sight path in the RIS-user link. When s = 0, Denote the time difference of arrival of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When s = 1, 2, …, N, Denote the time difference of arrival of the sth non-line-of-sight path in the RIS-user link. When s = 0, P s = P0. P0 denotes the three-dimensional position of the user, s = 1, 2, …, N s When s = 1, 2, …, N, P s Denote the three-dimensional position of the scatterer on the sth non-line-of-sight path in the RIS-user link. e(P B ) denotes the steering vector of the RIS in the line-of-sight path in the base station-RIS link. When s = 0, e(P s ) denotes the steering vector of the RIS in the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When s = 1, 2, …, N, e(P s ) denotes the steering vector of the RIS in the sth non-line-of-sight path in the RIS-user link.

[0009] In step 2, the first channel V 1,:,: is , and the second channel V 2,:,: is , where Re{} represents taking the real part, Im{} represents taking the imaginary part, denotes the two-norm operation of a vector or matrix.

[0010] The specific process of step 4 is as follows: Step 4.1: Randomly divide all the training data in the training set into multiple batches, so that each batch contains batch size groups of training data. Among them, each group of training data contains a sample and its corresponding label; Step 4.2: Take one batch from the training set, use all the training data in this batch as the input of the residual neural network, input it into the residual neural network for processing, and obtain the predicted three-dimensional positions of the users corresponding to each group of training data in this batch; Step 4.3: Calculate the mean square error loss between the predicted three-dimensional positions of the users and the labels; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of this batch on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; Step 4.4: Repeat the process of Step 4.2 to Step 4.3 until all batches of the training set have been trained on the residual neural network once; Step 4.5: Randomly divide all the validation data in the validation set into multiple batches, so that each batch contains batch size groups of validation data, where each group of validation data contains a sample and its corresponding label; Step 4.6: Take one batch from the validation set; then input all the validation data in this batch into the residual neural network trained using the training set to obtain the predicted three-dimensional positions of the users corresponding to each group of validation data in this batch; then calculate the mean square error loss between the predicted three-dimensional positions of the users and the labels; Step 4.7: Repeat the process of Step 4.6 until all batches of the validation set have been processed once, and then calculate the average mean square error loss of the validation set; Step 4.8: Adopt an exponentially decaying learning rate reduction strategy, repeat the process of Step 4.1 to Step 4.7, end after executing Num epochs in total, and save the network weights with the best validation set metrics to obtain the trained residual neural network model.

[0011] In the said Step 4.8, when adopting the exponentially decaying learning rate reduction strategy, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original every M epochs.

[0012] In step 4, the residual neural network selects ResNet, which includes a basic convolutional module, a deep convolutional module with a residual structure, and a regression module. The basic convolutional module is composed of a first convolutional layer, a first batch normalization layer, and a first PReLU activation function connected in sequence. The deep convolutional module is composed of a first residual module, a second residual module, and a third residual module with the same structure connected in sequence. The regression module is composed of a Flatten layer and a fully connected layer connected in sequence. After the sample passes through the first convolutional layer, the first batch normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained. After the basic feature map passes through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained. After the deep feature map passes through the Flatten layer, a feature vector is obtained, and after the feature vector passes through the fully connected layer, the predicted three-dimensional position of the user corresponding to the sample is obtained.

[0013] Each of the first residual module, the second residual module, and the third residual module includes a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map obtained by passing the feature map received by the corresponding residual module through the fourth convolutional layer is subjected to an element-wise addition operation with the feature map obtained by passing the feature map received by the corresponding residual module through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer in sequence. The feature map obtained after the element-wise addition operation is passed through the third PReLU activation function as the feature map output by the corresponding residual module.

[0014] The convolutional kernel size of the first convolutional layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the first residual module are 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the second residual module are 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the third residual module are 3×3, the stride is 4, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolutional kernel size of the fourth convolutional layer in the first residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel size of the fourth convolutional layer in the second residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel size of the fourth convolutional layer in the third residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The size of the sample is 2×N×T. U, the size of the basic feature map is 64×N×T U , the size of the deep feature map is 512×N / 8×T U / 8, the dimension of the feature vector is (8×N×T U )×1

[0015] Compared with the prior art, the advantages of the present invention are as follows: 1) The method of the present invention proposes a new type of CSI fingerprint, that is, taking the received signal matrix of the user as the original fingerprint data, and the original fingerprint data contains angle and distance features related to coordinates with position information. This new type of CSI fingerprint can better improve the positioning accuracy.

[0016] 2) The method of the present invention utilizes the reflection and regulation ability of RIS for signals, can enhance the signal propagation effect, increase the signal path diversity, and thus improve the positioning accuracy.

[0017] 3) The method of the present invention uses a single-antenna base station, which reduces the complexity without reducing the positioning accuracy. In the simulation, compared with the RCNR method using a multi-antenna base station, the positioning accuracy is comparable, which fully demonstrates the superiority of the method of the present invention.

[0018] 4) The method of the present invention regards indoor three-dimensional positioning as a regression problem, and uses a residual neural network to learn the parameters of indoor three-dimensional positioning. The residual neural network is data-driven and does not depend on the signal model, so it is applicable to the situation where the model is complex or the transmission environment is harsh.

[0019] 5) The method of the present invention has a small estimation error for the three-dimensional position of the user, and has strong generalization ability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram of a RIS-assisted downlink multipath transmission SISO millimeter-wave positioning system applied to the method of the present invention; Figure 2 is the overall implementation block diagram of the method of the present invention; Figure 3 is the structural diagram of the residual neural network used in the method of the present invention; Figure 4 is the structural diagram of the residual module in the residual neural network used in the method of the present invention; Figure 5 is the predicted three-dimensional position scatter plot obtained by making 500 predictions on the three-dimensional positions of three users respectively using the trained residual neural network model when the signal-to-noise ratio is 25 dB, and the total number of subcarriers and the total number of time slots are both 32; Figure 6Schematic diagram of the cumulative distribution function (CDF) varying with the positioning error at different signal-to-noise ratios (SNRs) obtained by using the method of the present invention; Figure 7 Schematic diagram for comparing the cumulative distribution function (CDF) varying with the positioning error between the method of the present invention and the existing method. Detailed implementation manners

[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0022] A RIS-assisted indoor fingerprint positioning method based on a residual neural network proposed by the present invention is applicable to a downlink multipath transmission SISO (Single-Input Single-Output) millimeter-wave positioning system assisted by RIS. As Figure 1 shown, in this system, there is a single-antenna base station (BS), a RIS, and a single-antenna user (UE). A three-dimensional coordinate system is established with the center position of the RIS. The RIS is deployed on the X-O-Z plane, and the spacing d between the reflection elements in the RIS is half a wavelength , denotes the signal wavelength. The base station is located in the far-field region of the RIS, and the indoor environment is located in the Fresnel near-field region of the RIS. The RIS-user link includes a line-of-sight (LOS) path and multiple non-line-of-sight (NLOS) paths. There is a scatterer on each NLOS path, and the base station-RIS link only includes one LOS path. Among them, the three-dimensional positions of the base station and the RIS are known, and the three-dimensional positions of the user and the scatterer are unknown. Figure 1 In , refers to the azimuth angle, which is defined as the angle between the projection of the incident direction on the X-O-Y plane and the positive direction of the X-axis, Figure 2 refers to the elevation angle, which is defined as the angle between the incident direction and the positive direction of the Z-axis. As shown, this method includes the following steps: U Step 1: Define the signal received by the user on the nth subcarrier at time slot t as y(n,t); then use the matrix composed of the signals received by the user on N subcarriers in T U time slots as the received signal matrix Y of the user. Y U contains the three-dimensional positions of the base station, the RIS, the user, and the scatterer. Among them, n = 1, 2,..., N, where N represents the total number of subcarriers, and t = 1, 2,..., T U , T U represents the total number of time slots, and the dimension of Y U is N×T U .

[0023] In this embodiment, in step 1, , where, ( ) T represents the transpose operation, h BR (n) represents the channel of the base station - RIS link at the nth sub - carrier, diag( ) represents constructing a diagonal matrix, w t represents the phase - shift vector of the RIS at time slot t, and the dimension of w t is N R ×1, , correspondingly represents the phase - shifts of the 1st, 2nd, …, N R th reflecting elements of the RIS at time slot t, N R represents the number of reflecting elements included in the RIS, and N R = N x ×N z , N x represents the number of reflecting elements included in the RIS in the X - axis direction, and N z represents the number of reflecting elements included in the RIS in the Z - axis direction, h RU (n) represents the channel of the RIS - user link at the nth sub - carrier, s t (n) represents the signal transmitted by the base station at the nth sub - carrier at time slot t, and z t (n) represents the zero - mean additive Gaussian noise in y(n,t).

[0024] In this embodiment, the channel h BR (n) of the base station - RIS link at the nth sub - carrier is modeled as , and the channel h RU (n) of the RIS - user link at the nth sub - carrier is modeled as , where the dimensions of h BR (n) and h RU (n) are both N R ×1, α BR represents the complex gain of the line - of - sight path in the base station - RIS link, exp( ) represents the exponential function with the natural constant as the base, the natural constant is 2.71, and j represents the imaginary part, represents the time - difference of arrival of the line - of - sight path in the base station - RIS link, △f represents the sub - carrier frequency spacing of the OFDM signal, and P B represents the three - dimensional position of the base station, and P B = [x B , y B , z B T , x B , y B , z B ​Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the base station, s = 0, 1, 2, …, N s , N s represents the number of non-line-of-sight paths in the RIS-user link. When s = 0, α RU,s represents the complex gain of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When α RU,s represents the complex gain of the sth non-line-of-sight path in the RIS-user link. When s = 0 represents the time difference of arrival of the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When represents the time difference of arrival of the sth non-line-of-sight path in the RIS-user link. Assuming the positioning distance is within the Fresnel zone of the RIS, that is , D represents the maximum aperture of the RIS, , , When s = 0, d R,s represents the Euclidean distance between the user and the RIS, s = 1, 2, …, N s When d R,s represents the Euclidean distance between the scatterer on the sth non-line-of-sight path in the RIS-user link and the RIS. When s = 0, d RU,s represents the Euclidean distance under the line-of-sight path in the RIS-user link, s = 1, 2, …, N s When d RU,s represents the Euclidean distance under the sth non-line-of-sight path in the RIS-user link. c represents the speed of light represents the two-norm operation of a vector or matrix, P R represents the three-dimensional position of the central reflection element in the RIS, P R = [x R , y R , z R T , x R , y R , z R Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the central reflection element in the RIS. When s = 0, P s = P0, P0 represents the three-dimensional position of the user, P0 = [x0, y0, z0] T , x0, y 0, z0 Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the user, s = 1, 2, …, N s When P s represents the three-dimensional position of the scatterer on the sth non-line-of-sight path in the RIS-user link, P s ​=[x s ,y s ,z s T ,x s ,y s ,z s Correspondingly represent the X-axis coordinate position, Y-axis coordinate position, and Z-axis coordinate position of the scatterer on the sth non-line-of-sight path in the RIS-user link, represent the phase difference of the base station-RIS link, represent the phase difference of the RIS-user link, e(P B ) represents the steering vector of the RIS in the line-of-sight path in the base station-RIS link. When s = 0, e(P s ) represents the steering vector of the RIS in the line-of-sight path in the RIS-user link. When s = 1, 2, …, N s When e(P s ) represents the steering vector of the RIS in the sth non-line-of-sight path in the RIS-user link, , considering that the indoor environment is located in the Fresnel near-field region of the RIS, from the geometric relationship, e r (P B ) is described as , and e r (P s ) is described as , r = 1, 2, …, N R , P r represents the three-dimensional position of the rth reflection element in the RIS.

[0025] Step 2: Since the Y U obtained for users with different three-dimensional positions is unique, Y U can be regarded as a new type of CSI (Channel State Information) fingerprint of the user. This fingerprint can fully characterize the information of the indoor complex multipath environment. According to the definition of the fingerprint positioning algorithm, Y U can be used to estimate the three-dimensional position of the user. Take Y U as the original fingerprint data; then normalize the real part and the imaginary part of the original fingerprint data to obtain a real-valued matrix V. Among them, the dimension of V is 2×N×T U , the first dimension represents the number of channels. The first channel is obtained by normalizing the real part of the original fingerprint data, and the second channel is obtained by normalizing the imaginary part of the original fingerprint data. The first channel V 1,:,: is , and the second channel V 2,:,: is , Re{} represents taking the real part, and Im{} represents taking the imaginary part, ​Represents the two-norm operation of a vector or matrix.

[0026] Step 3: Within the indoor environment range, adjust the user's three-dimensional position, and simultaneously record the user's true three-dimensional position after each adjustment as a label; then, following the processes of Step 1 and Step 2, obtain the corresponding real-valued matrix after each adjustment in the same way as a sample; then form a dataset from the samples and labels obtained after multiple adjustments; afterwards, divide the dataset into a training set and a validation set. During implementation, 80% of the dataset can form the training set, and 20% can form the validation set.

[0027] Step 4: Use the training set to perform offline training on the residual neural network, and iteratively optimize the network parameters through the backpropagation algorithm; simultaneously use the validation set to monitor the training process and evaluate the network performance; after the training is completed, save the network weights with the optimal metrics of the validation set to obtain the trained residual neural network model.

[0028] In this specific embodiment, the specific process of Step 4 is as follows: Step 4.1: Randomly divide all the training data in the training set into multiple batches, such that each batch contains batch size groups of training data, where each group of training data contains a sample and its corresponding label.

[0029] Step 4.2: Take one batch from the training set, use all the training data in this batch as the input to the residual neural network, and input it into the residual neural network for processing to obtain the predicted three-dimensional position of the user corresponding to each group of training data in this batch.

[0030] Step 4.3: Calculate the mean-square error (MSE) loss between the predicted three-dimensional position of the user and the label; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of the residual neural network with this batch, and iteratively optimize the network parameters through the backpropagation algorithm.

[0031] Step 4.4: Repeat the processes of Step 4.2 to Step 4.3 until all batches of the training set have trained the residual neural network once.

[0032] Step 4.5: Randomly divide all the validation data in the validation set into multiple batches, such that each batch contains batch size groups of validation data, where each group of validation data contains a sample and its corresponding label.

[0033] Step 4.6: Take one batch from the validation set; then input all the validation data in this batch into the residual neural network trained using the training set to obtain the predicted three-dimensional positions of the users corresponding to each group of validation data in this batch; and then calculate the mean squared error loss between the predicted three-dimensional positions of the users and the labels.

[0034] Step 4.7: Repeat the process of Step 4.6 until all batches of the validation set have been processed once, and then calculate the average mean squared error loss of the validation set.

[0035] Step 4.8: Adopt an exponentially decaying learning rate reduction strategy, repeat the process of Steps 4.1 to 4.7, and end after executing Num (e.g., take Num = 60) epochs. Then save the network weights with the best metrics on the validation set to obtain the trained residual neural network model. When adopting the exponentially decaying learning rate reduction strategy, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original every M epochs. In the embodiment, take M = 10.

[0036] In this specific embodiment, the residual neural network selects ResNet, as Figure 3 shown. It includes a basic convolution module, a deep convolution module with a residual structure, and a regression module. The deep convolution module can mine the deep features of the data, and the regression module realizes position regression. The basic convolution module is composed of a first convolution layer, a first batch normalization layer, and a first PReLU activation function connected in sequence. The deep convolution module is composed of the first residual module, the second residual module, and the third residual module with the same structure connected in sequence. The regression module is composed of a Flatten (unfolding) layer and a fully connected layer connected in sequence. After passing the sample through the first convolution layer, the first batch normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained. After passing the basic feature map through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained. After passing the deep feature map through the Flatten layer, a feature vector is obtained, and after passing the feature vector through the fully connected layer, the predicted three-dimensional position of the user corresponding to the sample is obtained.

[0037] In this specific embodiment, as Figure 4As shown, the first residual module, the second residual module, and the third residual module each include a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map obtained by passing the feature map received by the residual module through the fourth convolutional layer is element-wise added to the feature map obtained by passing the feature map received by the residual module through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer in sequence. The feature map obtained after the element-wise addition operation is passed through the third PReLU activation function and used as the feature map output by the residual module.

[0038] In this specific embodiment, the convolutional kernel size of the first convolutional layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the first residual module are 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the second residual module are 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel sizes of the second convolutional layer and the third convolutional layer in the third residual module are 3×3, the stride is 4, the padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolutional kernel size of the fourth convolutional layer in the first residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolutional kernel size of the fourth convolutional layer in the second residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolutional kernel size of the fourth convolutional layer in the third residual module is 1×1, the stride is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512; the size of the sample is 2×N×T. U , the size of the basic feature map is 64×N×T. U , the size of the deep feature map is 512×N / 8×T. U / 8, the dimension of the feature vector is (8×N×T. U )×1.

[0039] In the basic convolutional module, the first convolutional layer performs preliminary feature extraction on the sample. The feature map output in this process can be represented by the following formula: F. o,1 =ω1*F. i,1 +b1, F. o,1 represents the feature map output by the first convolutional layer, F. i,1Denote the feature map of the input of the first convolutional layer, ω1 represents the weights of the convolutional kernel of the first convolutional layer, b1 represents the bias of the convolutional kernel of the first convolutional layer, and * represents the convolution operation; after the convolution operation of the first convolutional layer, the first batch normalization layer (Batch Normalization, BN) is used to accelerate network training; finally, the first PReLU (Parametric Rectified Linear Unit) activation function is adopted. Compared with the traditional ReLU (Rectified Linear Unit) activation function, the PReLU activation function can adaptively adjust the output when the input value is negative instead of directly outputting zero, thus alleviating the problem of neuron "death" that may occur in the negative value region of the ReLU activation function to a certain extent. The basic feature map obtained after passing through the basic convolutional module can be expressed by the following formula: F b =max(0, BN(F o,1 )) + ρ × min(BN(F o,1 ), 0), where F b represents the basic feature map, max( ) represents taking the maximum value, min( ) represents taking the minimum value, ρ represents the learnable parameter, and BN( ) represents the batch normalization operation.

[0040] There are certain limitations in the feature extraction ability of the shallow network, making it difficult to meet the requirements of high-precision positioning. However, simply increasing the network depth will not only increase the training difficulty but also may cause problems such as gradient disappearance or degradation. Therefore, in the deep convolutional module, a residual learning strategy is introduced.

[0041] After the basic convolutional module and the deep convolutional module complete feature extraction, the regression module is responsible for the positioning task, that is, estimating the three-dimensional position of the user.

[0042] Step 5: Conduct an online test in the near-field indoor positioning scenario. Obtain the corresponding real-value matrix from the received signal matrix of the user obtained in real time according to the process of Step 2 and use it as a test sample; then input the test sample into the trained residual neural network model to output the predicted three-dimensional position of the user , corresponding to the predicted X-axis coordinate position, predicted Y-axis coordinate position, and predicted Z-axis coordinate position of the user.

[0043] To further illustrate the feasibility and effectiveness of the method of the present invention, a simulation experiment is conducted on the method of the present invention.

[0044] The simulation parameters are shown in Table 1.

[0045] Table 1 Simulation Parameters

[0046] Consider an indoor space with dimensions of 16×8×3 m 3 The length and width are divided at intervals of 0.2 m. To balance the dataset size, raw fingerprint data is only collected at heights of 0.2 m, 0.4 m, and 0.6 m. At signal-to-noise ratios of -5 dB, 0 dB, 5 dB, 10 dB, and 15 dB, a total of 81×41×3×5 = 49815 raw fingerprint data are collected. After traversing 45 times, a total of 45×49815 = 2241675 raw fingerprint data are collected to form the dataset.

[0047] The number of Monte Carlo runs for the simulation experiment is 500.

[0048] Figure 5 A scatter plot of the predicted three-dimensional positions of three users is given when the signal-to-noise ratio is 25 dB, the total number of subcarriers and the total number of time slots are both 32, and the trained residual neural network model is used to make 500 predictions for the three-dimensional positions of the three users. From Figure 5 it can be seen that the predicted three-dimensional positions of the users obtained by the method of the present invention are close to the true three-dimensional positions of the users, indicating that the method of the present invention can achieve good estimation.

[0049] Figure 6 A schematic diagram showing the variation of the cumulative distribution function (CDF) of the positioning error with the positioning error at different signal-to-noise ratios (SNR) obtained by using the method of the present invention is given. From Figure 6 it can be seen that as the signal-to-noise ratio increases, the CDF rises significantly, demonstrating that the improvement of the signal-to-noise ratio has a significant impact on the algorithm.

[0050] Figure 7 A schematic diagram comparing the variation of the cumulative distribution function (CDF) of the positioning error with the positioning error between the method of the present invention (ResNet) and existing methods is given. The existing methods include the RCNR method and the classical CNN method. The RCNR method is from the literature T. Wu, C. Pan, Y. Pan, et al. Fingerprint-based mmWave positioning system aided by reconfigurable intelligent surface[J]. IEEE Wireless Communications Letters, 2023, 12(8):1379-1383. (Reconfigurable intelligent surface-aided fingerprint-based millimeter-wave positioning system, IEEE Wireless Communications Letters). From Figure 7 it can be seen that the positioning performance of the method of the present invention is significantly better than that of existing methods.

Claims

1. A RIS-assisted indoor fingerprint positioning method based on residual neural network, characterized in that The following steps are involved: Step 1: Define the signal received by the user on the nth subcarrier at time slot t as y(n,t); then transform T U The matrix composed of the signals received by the user on N subcarriers in a time slot is used as the user's received signal matrix Y U , Y U includes the three-dimensional positions of base stations, RIS, users, and scatterers; where n=1,2,…,N, N represents the total number of subcarriers, and t=1,2,…,T U , T U Indicates the total number of time slots, Y U The dimension is N×T U ; Step 2: Y U As the original fingerprint data; then the real and imaginary parts of the original fingerprint data are normalized to obtain a real-valued matrix V, where the dimension of V is 2×N×T U , the first dimension represents the number of channels, the first channel is the real part of the original fingerprint data after normalization, and the second channel is the imaginary part of the original fingerprint data after normalization; Step 3: Adjust the user's 3D position within the indoor environment, and record the user's real 3D position after each adjustment as a label; then follow the process of steps 1 and 2 to obtain the corresponding real-valued matrix after each adjustment in the same way and use it as a sample; then form a data set with the samples and labels obtained after multiple adjustments; then divide the data set into a training set and a validation set; Step 4: Use the training set to train the residual neural network offline, and iteratively optimize the network parameters through the back propagation algorithm; at the same time, use the validation set to monitor the training process and evaluate the network performance; after the training is completed, save the network weights with the best validation set indicators to obtain the trained residual neural network model; Step 5: Perform an online test in a near-field indoor positioning scenario, obtain the corresponding real-valued matrix of the user's received signal matrix obtained in real time according to the process of step 2, and use it as a test sample; then input the test sample into the trained residual neural network model to output the user's predicted three-dimensional position.

2. According to claim 1, a RIS-assisted indoor fingerprint positioning method based on residual neural network is characterized in that In the step 1, ,in,( ) T represents the transpose operation, h BR (n) represents the channel of the base station-RIS link at the nth subcarrier, diag( ) represents the construction of a diagonal matrix, w t represents the phase shift vector of RIS at time slot t, h RU (n) represents the channel of the RIS-user link at the nth subcarrier, s t (n) represents the signal transmitted by the base station on the nth subcarrier at time slot t, z t (n) represents zero-mean additive Gaussian noise in y(n,t).

3. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 2 is characterized in that The base station-RIS link is on the channel h of the nth subcarrier. BR (n) Modeled as , the RIS-user link has a channel h at the nth subcarrier RU (n) Modeled as , where α BR represents the complex gain of the line-of-sight path in the base station-RIS link, j is the imaginary part, represents the arrival time difference of the line-of-sight path in the base station-RIS link, △f represents the subcarrier frequency spacing of the OFDM signal, and P B represents the three-dimensional position of the base station, s=0,1,2,…,N s , N s represents the number of non-line-of-sight paths in the RIS-user link, α when s=0 RU,s represents the complex gain of the line-of-sight path in the RIS-user link, s=1,2,…,N s Time α RU,s represents the complex gain of the sth non-line-of-sight path in the RIS-user link, when s=0 represents the arrival time difference of the line-of-sight path in the RIS-user link, s=1,2,…,N s hour represents the arrival time difference of the sth non-line-of-sight path in the RIS-user link. When s=0, P s =P0, P0 represents the user's three-dimensional position, s=1,2,…,N s Time s represents the three-dimensional position of the scatterer on the sth non-line-of-sight path in the RIS-user link, e(P B ) represents the steering vector of RIS in the line-of-sight path in the base station-RIS link. When s=0, e(P s ) represents the steering vector of RIS in the line-of-sight path in the RIS-user link, s=1,2,…,N s When e(P s ) represents the steering vector of the RIS in the sth non-line-of-sight path in the RIS-user link.

4. A RIS-assisted indoor fingerprint positioning method based on residual neural network according to any one of claims 1 to 3, characterized in that In step 2, the first channel V 1,:,: for , the second channel V 2,:,: for , where Re{} means taking the real part, and Im{} means taking the imaginary part. Represents the bi-norm operation of a vector or matrix.

5. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 1 is characterized in that The specific process of step 4 is as follows: Step 4.1: Randomly divide all training data in the training set into multiple batches, so that each batch contains batch size groups of training data, where each group of training data contains a sample and a corresponding label; Step 4.2: Take one of the batches of the training set, use all the training data in this batch as the input of the residual neural network, input it into the residual neural network for processing, and obtain the predicted 3D position of the user corresponding to each set of training data in this batch; Step 4.3: Calculate the mean square error loss of the user's predicted 3D position and label; then use the Adam optimizer to train the network parameters of the residual neural network, complete the training of the residual neural network for this batch, and iteratively optimize the network parameters through the back propagation algorithm; Step 4.4: Repeat the process from step 4.2 to step 4.3 until all batches of the training set have trained the residual neural network once; Step 4.5: Randomly divide all the verification data in the verification set into multiple batches, so that each batch contains batch size groups of verification data, where each group of verification data contains a sample and the corresponding label; Step 4.6: Take one of the batches of the validation set; then input all the validation data in this batch into the residual neural network trained with the training set to obtain the predicted 3D position of the user corresponding to each set of validation data in this batch; then calculate the mean square error loss of the predicted 3D position of the user and the label; Step 4.7: Repeat the process of step 4.6 until all batches of the validation set have been processed once, and then calculate the average mean square error loss of the validation set; Step 4.8: Use an exponentially decreasing learning rate reduction strategy, repeat the process from step 4.1 to step 4.7, and end after executing Num epochs in total. Save the network weights with the best validation set indicators to obtain the trained residual neural network model.

6. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 5 is characterized in that In step 4.8, when the exponentially decreasing learning rate reduction strategy is adopted, the initial value of the learning rate is set to 0.001, and the learning rate is reduced to 80% of the original value every M epochs.

7. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 1 is characterized in that In the step 4, the residual neural network uses ResNet, which includes a basic convolution module, a deep convolution module using a residual structure, and a regression module. The basic convolution module is composed of a first convolution layer, a first normalization layer, and a first PReLU activation function connected in sequence, the deep convolution module is composed of a first residual module, a second residual module, and a third residual module with the same structure connected in sequence, and the regression module is composed of a Flatten layer and a fully connected layer connected in sequence; after the sample passes through the first convolution layer, the first normalization layer, and the first PReLU activation function in sequence, a basic feature map is obtained; after the basic feature map passes through the first residual module, the second residual module, and the third residual module in sequence, a deep feature map is obtained; after the deep feature map passes through the Flatten layer, a feature vector is obtained, and after the feature vector passes through the fully connected layer, a predicted three-dimensional position of the user corresponding to the sample is obtained.

8. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 7 is characterized in that The first residual module, the second residual module, and the third residual module all include a second convolutional layer, a second batch normalization layer, a second PReLU activation function, a third convolutional layer, a third batch normalization layer, a third PReLU activation function, and a fourth convolutional layer. The feature map received by the residual module is passed through the fourth convolutional layer to obtain a feature map, and the feature map received by the residual module is sequentially passed through the second convolutional layer, the second batch normalization layer, the second PReLU activation function, the third convolutional layer, and the third batch normalization layer to obtain a feature map. The feature map obtained after the element-by-element addition operation is passed through the third PReLU activation function as the feature map output by the residual module.

9. The RIS-assisted indoor fingerprint positioning method based on residual neural network according to claim 8 is characterized in that The convolution kernel size of the first convolution layer is 3×3, the stride is 1, the padding is 1, the number of input channels is 2, and the number of output channels is 64. The convolution kernel size of the second and third convolution layers in the first residual module is 3×3, the stride is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolution kernel size of the second and third convolution layers in the second residual module is 3×3, the stride is 2, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolution kernel size of the second and third convolution layers in the third residual module is 3×3, the stride is 4, The padding is 1, the number of input channels is 256, and the number of output channels is 512. The convolution kernel size of the fourth convolution layer in the first residual module is 1×1, the step size is 1, the padding is 1, the number of input channels is 64, and the number of output channels is 128. The convolution kernel size of the fourth convolution layer in the second residual module is 1×1, the step size is 1, the padding is 1, the number of input channels is 128, and the number of output channels is 256. The convolution kernel size of the fourth convolution layer in the third residual module is 1×1, the step size is 1, the padding is 1, the number of input channels is 256, and the number of output channels is 512; the size of the sample is 2×N×T U , the size of the basic feature map is 64×N×T U , the size of the deep feature map is 512×N / 8×T U / 8, the dimension of the feature vector is (8×N×T U )×1.

Citation Information

Patent Citations

  • Large-scale MIMO fingerprint positioning method based on complex neural network

    CN112995892A

  • Channel estimation method for passive intelligent reflection surface based on deep learning

    CN113179232A

  • Method for realizing RIS self-adaptive reconstruction of multi-user channel based on Vision Transform network

    CN119675711A

  • Method and apparatus communication in cooperative wireless communication systems

    US20240236878A1

  • Channel estimation technique for reconfigurable intelligent surfaces

    WO2023218079A1

Cited By

  • RIS-assisted wireless positioning method

    CN120603047A

  • A RIS-assisted wireless positioning method

    CN120603047B

  • Multi-environment adaptive visible light indoor positioning method, system and device and storage medium

    CN120891463A