A user positioning method
Through the STAR-RIS-assisted positioning method, the neural network is used to estimate the user's position and predict the trajectory, which solves the positioning problem of obstructed visual channel, realizes high-precision user positioning and trajectory prediction, and improves the reliability of the communication system.
Patent Information
- Application Number
- CN202310521802.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-05-10
AI Technical Summary
The existing positioning algorithms cannot accurately locate user locations when the line-of-sight channel is blocked or the signal is weak, and cannot effectively predict user trajectory, which is costly and has strong dependence on the line-of-sight channel.
The STAR-RIS-assisted positioning method is adopted to estimate the absolute position of the user relative to STAR-RIS by using a neural network, and the user's subsequent motion trajectory is predicted through the channel state matrix, and the position estimation and prediction are combined with the convolutional neural network and the fully connected network.
It improves user positioning accuracy, solves the problem of position estimation under no direct connection channel, and realizes high-quality prediction of user trajectory, providing communication guarantee for subsequent mobile users.
Smart Images

Figure CN116528166B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communications and relates to a user positioning method. Background Art
[0002] With the advancement of digital electronics and electromagnetic metamaterials, reconfigurable intelligent surfaces (RIS) have been designed and implemented. Theoretical derivation and experimental verification have shown that they can improve energy efficiency, spectral efficiency, and positioning accuracy. Because RIS are passive, they can be used as intelligent reflectors to mitigate blocking and shadowing effects, expand communication coverage, and construct new cascade channels to address the problem of blocked direct paths. Recently, a new type of RIS, the Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface (STAR-RIS), has emerged. It is considered a promising bridge for connecting users on either side of the system. Unlike traditional RIS, STAR-RIS can simultaneously perform two different functions, such as reflection and transmission. This means that users on both sides of the STAR-RIS system can be served simultaneously without interfering with each other, for example, by using mutually orthogonal carriers for signal transmission. Furthermore, these two functions can be implemented through various methods, such as energy splitting, time switching, and mode switching.
[0003] At present, traditional positioning algorithms mainly include positioning algorithms based on GPS / Beidou, template matching positioning algorithms based on received signal strength, and positioning algorithms based on received signal arrival time (ToA). Their defects are as follows: (1) the problem of being unable to locate when the user is blocked by buildings, that is, the line-of-sight channel is blocked or the received signal is weak; (2) the problem of being unable to accurately locate the user's direct location and being unable to distinguish between users in similar areas; (3) the problem of relying on three base stations based on the arrival time of the received signal, and then calculating the distance between the base station and the device based on the arrival time of the signal, and then using geometric knowledge to calculate the position of the device, which leads to high cost and high dependence on the line-of-sight channel; (4) the existing algorithm does not predict the user trajectory based on the received time slot signal. Summary of the Invention
[0004] The present invention provides a user positioning method to solve the problem that when the line-of-sight channel from the base station to the user is blocked, accurate positioning cannot be achieved, and a single neural network cannot be used to simultaneously estimate and predict the positioning of users on both sides.
[0005] The present invention is achieved through the following technical solutions.
[0006] A user positioning method uses a neural network to estimate the user's absolute position relative to a simultaneously transmitting and reflecting reconfigurable smart surface, and predicts the user's subsequent motion trajectory based on the channel state matrix of the previous time slot. The method includes the following steps:
[0007] S1: First, STAR-RIS is introduced into the MISO communication system and suspended on the yoz plane, with the coordinates of its center point being p S =(x S ,y S , z S ), the transmission matrix of STAR-RIS is Φ t , the reflection matrix is Φ r , are all diagonal matrices, where the nth (n∈N) unit of STAR-RIS can be expressed as
[0008] The energy splitting parameter of STAR-RIS is ε t With ε r , representing the energy coefficients for transmission and reflection, respectively, and satisfying ε t 2 +ε r 2 =1, the channel between STAR-RIS and the transmission side user is h t,2 , and the user channel on the reflection side is h r,2 , the channel between the base station and STAR-RIS is H1, and the transmission side cascade channel from the base station to the transmission side user can be expressed as The user cascade channel on the reflection side is
[0009] The system uses orthogonal frequency division multiplexing to complete signal coding and modulation, and the channel occupies K subcarriers.
[0010] Base station equipped with N r antennas, the transmission side user and the reflection side user are equipped with a single antenna, and the position of the transmission side user is p t =(x t ,y t , z t ), the position of the user on the reflecting side is p r =(x r ,y r , z r ).
[0011] The network label is l′=(R, φ, θ, t, r), where R is the reference position of STAR-RIS (the center point of STAR-RIS p S ) to the user, φ is the Euclidean distance from the user to p SThe pitch angle, θ is the angle between the user and the x-axis. S The azimuth angle, t and r are not 1 at the same time. When t = 1 and r = 0, it means that the user is located on the transmission side of STAR-RIS. Conversely, when t = 0 and r = 1, it means that the user is located on the reflection side of STAR-RIS. Where R, φ, θ satisfy the following formula:
[0012] (x S -x r ,y S -y r , z S -z r )=(Rcosφcosθ, Rcosφsinθ, Rsinφ)
[0013] (x S -x t ,y S -y t , z S -z t )=(-Rcosφcosθ, Rcosφsinθ, Rsinφ)
[0014] When the user moves, the channel state information matrix data of each user is continuously collected for T time slots, and each time slot is Where c is the speed of light, v is the user's moving speed, and f c is the carrier frequency, so the dimension of the channel state matrix of each user in the data set is T×N r ×K, vertical splicing by time slot;
[0015] The N of the above matrix T time slots sample The channel state information matrix H t 、H r and label l′ together constitute the dataset and where N sample is the total number of samples in the dataset;
[0016] Furthermore, the users on the transmission side and the users on the reflection side do not interfere with each other, and the channel state matrix of the uplink is completed at the base station end. and The data set is used as the training network data set and is divided into training set, verification set and test set according to 7:1:2. The labels correspond to the channel state matrix in the data set.
[0017] S2: In each channel state information matrix in the data set, the transmitting antenna m (m∈N r ) in time slot t(t∈T) on K subcarriers is Pass it Complete the normalization of the vector and obtain the normalized channel amplitude at this time
[0018] The channel vector is then converted to Converted into a complex-valued channel vector, normalized for each transmit antenna m and subcarrier k, and concatenated into a normalized dimension of T×N r ×K channel state information matrix H csi ;
[0019] Furthermore, the original vector The amplitude angle of the channel vector is converted into a complex-valued channel state matrix H csi After the normalization process, the channel state matrix has a more statistical distribution and retains the original phase angle information.
[0020] S3: Pre-train the channel feature extraction network CsiNet, which builds a neural network with multiple convolution layers. Each convolution layer includes a convolution layer, a batch normalization layer, an activation function (GeLU layer), a pooling layer (max pooling), and three fully connected layers (FC) and an activation layer structure.
[0021] The network completes H by extracting the spatial position information from the channel state matrix. csi The nonlinear mapping to l = (R, φ, θ), where l is the first three items of l′, and the output of the CsiNet network is The training goal of the network is to make the output of CsiNet Continuously fit the label l, and its specific training includes the following steps:
[0022] S3-1: The loss function used by the CsiNet network is:
[0023]
[0024] S3-2: The last layer FC output of the network The feature vector mapped by the penultimate layer FC is defined as the channel position information feature fc2;
[0025] S3-3: When the loss value of two adjacent rounds of training The difference is less than 1e for 5 consecutive times -5 Indicates that network training is complete and the network parameters are stored, otherwise continue training.
[0026] Furthermore, the CsiNet network learning rate is Y, and the stochastic gradient descent method based on batch learning is used, and the optimization method is set to the Adam algorithm.
[0027] S4: Train the overall network CsiNet-Former, where the prediction network is the Former and the label value of the entire network is l′. The specific training steps include:
[0028] S4-1: Set the channel gain parameter to h gain (t), where a ave (t) is the state matrix H of each channel in each time slot csi The average amplitude of (t), It is used to normalize in the time dimension and then multiply it with the user normalization parameter ∈. The overall normalization method is as follows:
[0029]
[0030] In the above formula a UE =Max(a ave (t)), is a ave (t) is the maximum channel gain in time slot T, and α and β are the channel gains in the entire data set. UE The maximum and minimum values of ∈ determine the relative gain of each channel gain in the entire data set;
[0031] S4-2: The collected T time slot channel state matrix sequence is divided into T tr 、T ad and T pr Three parts, the overall timing division satisfies T tr +T ad +T pr =T.
[0032] Where T tr +T ad The time slots are used to pre-train the CsiNet-Former network, T ad The channel state matrix label l′ of the time slot is used to adjust the parameters of the trained network’s prediction ability;
[0033] T pr This is the prediction part, using T tr +T ad The corresponding l′ of the part calculates the prediction loss function. During the operation, the entire network dynamically inputs and the sliding window predicts the user's trajectory;
[0034] S4-3: T normalized by step S2 tr +T ad The channel state matrix of the time slot is input into the pre-trained CsiNet to complete the channel spatial state feature fc2 extraction and obtain its rough estimated user location information. These two vectors and h obtained in S4-1 gain(t) horizontal splicing is x t , input to the encoder of the Former network;
[0035] Among them, T pr Part of the channel state matrix is input to the CsiNet network to obtain the channel space state feature fc2 and the rough estimated user location information The encoder that is not input to the Former network, T pr Some parts are unknown to the Former network and need to be predicted;
[0036] S4-4: The encoder of the Former network takes the input T tr +T ad x t Encoding to hidden state, the encoder is composed of FC, and the attention mechanism is introduced to further extract the relationship between sequences;
[0037] S4-5: The decoder of Former completes the hidden state decoding, and its output is The dimension is (T ad +T pr )×5, including the real location estimate T ad Partial and predicted position estimate T pr part;
[0038] Furthermore, Former consists of a three-layer encoder and a two-layer decoder, both of which are fully connected layers. The encoder is responsible for converting the input x t Mapped to a hidden state (an implicit nonlinear mathematical expression containing channel spatial information), the decoder then maps the hidden state to a fine estimate of the user position Including the true position estimate T ad Partial and predicted position estimate T pr part.
[0039] By L Former Complete the loss function calculation and use it to train and update network parameters, L Former The expression is as follows:
[0040]
[0041] where l and is the label value and output of the CsiNet network described in S3 to roughly estimate the user location, while l′ and The label value and output of the CsiNet-Former network are used to estimate the user location;
[0042] The first part is used to correct the estimation and prediction capabilities of the Former network, and the second part is used to correct the feature extraction and coarse user position estimation capabilities of the CsiNet network;
[0043] S4-6: When the loss value of two adjacent rounds of training The difference is less than 1e for 10 consecutive times -5 Indicates that network training is complete and the network parameters are stored, otherwise continue training.
[0044] The network learning rate is γ′, and the stochastic gradient descent method based on batch learning is used. The learning rate parameter changes with the round, and the optimization method is set to the Adam algorithm.
[0045] Furthermore, the attention mechanism is adopted in the CsiNet-Former network to map the hidden state into three matrices Q (query), K (key) and V (value). By searching the query idea, the attention of the calculated key-value pair is defined as This value indicates the value of each key-value pair, which is a linear combination of V. The weight is generated according to the similarity of Q and K to obtain the probability value of the key-value pair to be queried. The higher the value, the more valuable the key-value pair is to be queried. The calculation is as follows:
[0046]
[0047] Where d is the input dimension, that is, x t Dimension;
[0048] Defined as is the criterion of the sparsity of Q, q i is the i-th column of Q, L is the mapping dimension, and the probability p(k j |q i )=1 / L. The key-value pairs with larger Wasserstein distance have greater query value. By only querying the key-value pairs with greater value, Q becomes a sparse matrix, simplifying the computational complexity.
[0049] In particular, the network adopts a multi-head self-attention mechanism, that is, it learns different behaviors based on the same attention mechanism, and then combines the different behaviors as knowledge to obtain the output result through an FC mapping;
[0050] The introduction of mask attention in the decoder will make the available hidden layer, T tr +T ad The mask of the sequence part is set to 1, and the rest of T pr The sequence is 0, which means that the subsequent T can only be completed through the previous hidden state prUser location prediction, but cannot use T pr (Channel state matrix to be predicted) completes the preamble output.
[0051] S5: Calculate the user's absolute position according to The t′ and r′ parameters in the equation are used to estimate the R′, where if t′>r′, then it is -R′, otherwise it is R′, and the estimated p is further obtained. t ′=(x t ′,y t ′,z t ′) and p r ′=(x r ′,y r ′,z r ′) is shown in the following formula:
[0052]
[0053] The advantages of the present invention are: Compared to existing STAR-RIS-assisted user positioning methods, the STAR-RIS-assisted bilateral user positioning method provided herein utilizes nonlinear mapping within a neural network to improve user positioning accuracy. Leveraging STAR-RIS's simultaneous transmission and reflection capabilities, it solves the problem of estimating user positions in the absence of direct channels. Furthermore, by using a sliding window input channel state information matrix, user spatial position information can be estimated and predicted. This predictive capability ensures high-quality communication for subsequent mobile users. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0055] Figure 1 is a schematic diagram of a communication scenario according to an embodiment of the present invention;
[0056] Figure 2 2 is a schematic diagram showing a comparison of normalized channel state matrices according to an embodiment of the present invention;
[0057] Figure 3 Schematic diagram of a specific network structure of a neural network according to an embodiment of the present invention;
[0058] Figure 4 1 is a schematic diagram of the overall timing division of an embodiment of the present invention;
[0059] Figure 5 is an overall training flow chart of an embodiment of the present invention;
[0060] Figure 6 The embodiment of the present invention has different t 2 Error curve of the value;
[0061] Figure 7 It is a schematic diagram of the overall operation flow of the present invention. DETAILED DESCRIPTION
[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0063] The implementation steps of the present invention are as follows:
[0064] S1: Introduce STAR-RIS into the MISO communication system and hang it on the yoz plane with the center coordinates of p S =[x S ,y S , z S ];
[0065] Using the yoz plane as the dividing interface, users on the same side as the base station are defined as reflection-side users, and users on the opposite side are defined as transmission-side users. There are obstacles in the direct connection channel between the reflection-side users on the transmission side and the base station.
[0066] The transmission matrix of STAR-RIS is Φ t , the reflection matrix is Φ r , are diagonal matrices, and their energy splitting parameter is ε t With ε r , representing the energy coefficients for transmission and reflection, respectively, and satisfying ε t 2 +ε r 2 =1, the nth (n∈N) unit of STAR-RIS can be expressed as The overall communication system Figure 1 As shown, the number of STAR-RIS elements in the communication system of this example is N=144.
[0067] In this example, ε is generated respectively. t 2 = 0.9 and ε t 2 = 0.5, two sets of data sets, and Randomly generated and satisfying The channel between STAR-RIS and the user on the transmission side is h t,2 , and the user channel on the reflection side is h r,2 , the channel between the base station and STAR-RIS is H1.
[0068] The transmission side cascade channel from the base station to the transmission side user can be expressed as The user cascade channel on the reflection side is
[0069] in and The expression is:
[0070]
[0071]
[0072] where d t / r,2 (in meters) and ρ t / r,2 (Assume ) are the distance and path loss between the user and the base station on the transmission and reflection sides, respectively;
[0073] λ is the wavelength of the carrier frequency f c The reciprocal of h t,2 and h r,2 θ in the expression t / r,2 and φ t / r,2 are the azimuth and elevation angles of the users on the transmission side and reflection side relative to STAR-RIS, respectively, α x (θ t / r,2 ,φ t / r,2 ) and α z (φ t / r,2 ) is the steering vector;
[0074] Since the system uses orthogonal frequency division multiplexing to complete signal encoding, the channel occupies K = 64 subcarriers;
[0075] In the MISO communication system of this example, the base station is equipped with N r = 100 antennas, the user is equipped with a single antenna, and the phase angles of the transmission and reflection elements of STAR-RIS are randomly set within the range set in this example (in terms of p S The position of the user on the transmission or reflection side is randomly generated within a radius of 30m from the center as p t =[x t ,y t , z t ] and p r =[x r ,y r , zr ];
[0076] The network label is l′=(R, φ, θ, t, r), where R is the reference position of STAR-RIS (the center point of STAR-RIS p S ) to the user, φ is the Euclidean distance from the user to p S The pitch angle, θ is the angle between the user and the x-axis. S The azimuth angle, t and r are not 1 at the same time. When t = 1 and r = 0, it means that the user is located on the transmission side of STAR-RIS. Conversely, when t = 0 and r = 1, it means that the user is located on the reflection side of STAR-RIS. Where R, φ, θ satisfy the following formula:
[0077] (x S -x r ,y S -y r , z S -z r )=(Rcosφcosθ, Rcosφsinθ, Rsinφ)
[0078] (x S -x t ,y S -y t , z S -z t )=(-Rcosφcosθ, Rcosφsinθ, Rsinφ)
[0079] When the user moves, in this example, the channel state information data of each user is continuously collected for T=150 time slots, each time slot is Where c is the speed of light, v∈(0,1] is the user's moving speed, and f c =2.8GHz is the carrier frequency, so the dimension of each user channel state matrix in the data set is T×N r ×K, vertical splicing by time slot;
[0080] N of the above T time slots sample The channel state information matrix H t 、H r and label l′ together constitute the dataset and where N sample is the total number of samples in the dataset;
[0081] Furthermore, the users on the transmission side and the users on the reflection side do not interfere with each other, and the channel state matrix of the uplink is completed at the base station end. and The data set is used as the training network data set and is divided into training set, verification set and test set according to 7:1:2. The labels correspond to the channel state matrix in the data set.
[0082] In each channel state information matrix in the data set, the transmitting antenna m(m∈N r ) in time slot t(t∈T) on K subcarriers is Pass it Complete the normalization of the vector and obtain the normalized channel amplitude at this time
[0083] The channel vector is then converted to Converted into a complex-valued channel vector, normalized for each transmit antenna m and subcarrier k, and concatenated into a normalized dimension of T×N r ×K channel state information matrix H csi ;
[0084] Furthermore, the original vector The amplitude angle of the channel vector is converted into a complex-valued channel state matrix H csi After the normalization process, the channel state matrix has a more statistical distribution and retains the original phase angle information. The normalized channel state matrix is as follows: Figure 2 shown.
[0085] In actual network training, the channel state matrix is divided into two channels according to the real and imaginary parts, that is, the single user channel finally input to the network training becomes T×2×N r ×K;
[0086] In this example, when training the channel feature extraction network CsiNet, the time feature has no effect on the training results. The channel state matrix H input to CsiNet csi The dimension is 2×N r ×K.
[0087] S2: Pre-trained channel feature extraction network CsiNet, constructing a neural network with multiple convolution operation layers. Each convolution operation layer includes a convolution layer, a batch normalization layer, an activation function (GeLU layer), a pooling layer (max pooling), and a 3-layer fully connected layer (FC) and an activation layer structure. The CsiNet network structure diagram is shown in the figure. Figure 3 As shown;
[0088] The network completes H by extracting the spatial position information from the channel state matrix. csiThe nonlinear mapping to l = (R, φ, θ), where l is the first three items of l′, and the output of the CsiNet network is The training goal of the network is to make the output of CsiNet Continuously fit the label l, and its specific training includes the following steps:
[0089] S3-1: The loss function used by the CsiNet network is:
[0090]
[0091] S3-2: The last layer FC output of the network The feature vector mapped by the penultimate layer FC is defined as the channel position information feature fc2;
[0092] S3-3: When the loss value of two adjacent rounds of training The difference is less than 1e for 5 consecutive times -5 Indicates that network training is complete and the network parameters are stored, otherwise continue training.
[0093] Furthermore, the CsiNet network learning rate is γ = 0.005, and the stochastic gradient descent method based on batch learning is used, and the optimization method is set to the Adam algorithm.
[0094] S4: Training the overall network CsiNet-Former. The overall network structure diagram is as follows Figure 4 As shown, the prediction network is Former, the label value of the entire network is l′, and its specific training includes the following steps:
[0095] S4-1: Set the channel gain parameter to h gain (t), where a ave (t) is the state matrix H of each channel in each time slot csi The average amplitude of (t), It is used to normalize in the time dimension and then multiply it with the user normalization parameter ∈. The overall normalization method is as follows:
[0096]
[0097] In the above formula a UE =Max(a ave (t)), is a ave (t) is the maximum channel gain in time slot T, and α and β are the channel gains in the entire data set. UE The maximum and minimum values of ∈ determine the relative gain of each channel gain in the entire data set;
[0098] S4-2: The collected T time slot channel state matrix sequence is divided into Ttr 、T ad and T pr The time slot division diagram is as follows: Figure 5 As shown, satisfying T tr +T ad +T pr =T.
[0099] Where T tr +T ad The time slots are used to pre-train the CsiNet-Former network, T ad The channel state matrix label l′ of the time slot is used to adjust the parameters of the trained network’s prediction ability;
[0100] T pr This is the prediction part, using T tr +T ad The corresponding l′ of the part calculates the prediction loss function. During the operation, the entire network dynamically inputs and the sliding window predicts the user's trajectory;
[0101] S4-3: T normalized by step S2 tr +T ad The channel state matrix of the time slot is input into the pre-trained CsiNet to complete the channel spatial state feature fc2 extraction and obtain its rough estimated user location information. These two vectors and h obtained in S4-1 gain (t) horizontal splicing is x t , input to the encoder of the Former network;
[0102] Among them, T pr Part of the channel state matrix is input to the CsiNet network to obtain the channel space state feature fc2 and the rough estimated user location information The encoder that is not input to the Former network, T pr Some parts are unknown to the Former network and need to be predicted;
[0103] S4-4: The encoder of the Former network takes the input T tr +T ad x t Encoding to hidden state, the encoder is composed of FC, and the attention mechanism is introduced to further extract the relationship between sequences;
[0104] S4-5: The decoder of Former completes the hidden state decoding, and its output is The dimension is (T ad +T pr )×5, including the real location estimate T ad Partial and predicted position estimate Tpr part;
[0105] Furthermore, Former consists of a three-layer encoder and a two-layer decoder, both of which are fully connected layers. The encoder is responsible for converting the input x t Mapped to a hidden state (an implicit nonlinear mathematical expression containing channel spatial information), the decoder then maps the hidden state to a fine estimate of the user position Including the true position estimate T ad Partial and predicted position estimate T pr part.
[0106] And through L Former Complete the loss function calculation and use it to train and update network parameters, L Former The expression is as follows:
[0107]
[0108] where l and is the label value and output of the CsiNet network described in S3 to roughly estimate the user location, while l′ and The label value and output of the CsiNet-Former network are used to estimate the user location;
[0109] The first part is used to correct the estimation and prediction capabilities of the Former network, and the second part is used to correct the feature extraction and coarse user position estimation capabilities of the CsiNet network;
[0110] S4-6: When the loss value of two adjacent rounds of training The difference is less than 1e for 10 consecutive times -5 Indicates that network training is complete and the network parameters are stored, otherwise continue training.
[0111] The network learning rate is γ′=0.005, and the stochastic gradient descent method based on batch learning is used. The learning rate parameter changes with the round, and the learning rate decreases to 0.5 times the original value in each round. The optimization method is set to the Adam algorithm.
[0112] Furthermore, the attention mechanism is adopted in the CsiNet-Former network to map the hidden state into three matrices Q (query), K (key) and V (value). By searching the query idea, the attention of the calculated key-value pair is defined as This value indicates the value of each key-value pair, which is a linear combination of V. The weight is generated according to the similarity of Q and K to obtain the probability value of the key-value pair to be queried. The higher the value, the more valuable the key-value pair is to be queried. The calculation is as follows:
[0113]
[0114] Where d is the input dimension, that is, x t Dimension;
[0115] The query probability of each key-value pair is defined as q i is the i-th column of Q, L is the mapping dimension, and the probability p(k j |q i )=1 / L. The key-value pairs with larger Wasserstein distance have greater query value. By only querying the key-value pairs with greater value, Q becomes a sparse matrix, simplifying the computational complexity.
[0116] In particular, the network adopts a multi-head self-attention mechanism, that is, it learns different behaviors based on the same attention mechanism, and then combines the different behaviors as knowledge to obtain the output result through an FC mapping;
[0117] The introduction of mask attention in the decoder will make the available hidden layer, T tr +T ad The mask of the sequence part is set to 1, and the rest of T pr The sequence is 0, which means that the subsequent T can only be completed through the previous hidden state pr User location prediction, but cannot use T pr (Channel state matrix to be predicted) completes the preamble output.
[0118] S5: Calculate the user's absolute position according to The t′ and r′ parameters in the equation are used to estimate the R′, where if t′>r′, then it is -R′, otherwise it is R′, and the estimated p is further obtained. t ′=(x t ′,y t ′,z t ′) and p r ′=(x r ′,y r ′,z r ′) is shown in the following formula:
[0119]
[0120] The entire network can be deployed in a STAR-RIS-assisted uplink MISO communication system, and user position estimation and trajectory prediction can be completed at the base station end.
[0121] For different ε t 2 The error curves are compared with those of the conventional method, and the results are as follows Figure 6 shown.
[0122] pass Figure 6 It can be seen that the method of this embodiment is used to complete the estimation and prediction of user location information. t 2 = 0.9, the error converges stably to 0.01, while in ε t 2 = 0.5, the error converges to 0.04. This error primarily stems from CisNet-Former's inability to distinguish which side the user belongs to. Therefore, by adjusting the STAR-RIS splitting coefficient, the system can ensure that users can be distinguished between the transmission and reflection sides. The test results in the test set show a bilateral error of approximately 0.75m and a unilateral error of approximately 0.034m.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A user positioning method, characterized in that: A neural network is used to estimate the user's absolute position relative to the simultaneously transmitting and reflecting reconfigurable smart surface, and the user's subsequent motion trajectory is predicted based on the channel state matrix of the previous time slot. The process includes the following steps: S1: First, STAR-RIS is introduced into the multi-input single-output communication system and suspended on the yoz plane with the center coordinates of p S =(x S ,y S ,z S ), the transmission matrix of STAR-RIS is Φ t , the reflection matrix is Φ r , are all diagonal matrices, where the nth (n∈N) unit of STAR-RIS can be expressed as The energy splitting parameter of STAR-RIS is ε t With ε r , representing the energy coefficients for transmission and reflection, respectively, and satisfying ε t 2 +ε r 2 =1, the channel between STAR-RIS and the transmission side user is h t,2 , and the channel between the user on the reflection side is h r,2 , the channel between the base station and STAR-RIS is H1, and the transmission side cascade channel from the base station to the transmission side user can be expressed as The user cascade channel on the reflection side is The system uses orthogonal frequency division multiplexing to complete signal coding and modulation, and the channel occupies K subcarriers; Base station equipped with N r antennas, the transmission side user and the reflection side user are equipped with a single antenna, and the position of the transmission side user is p t =(x t ,y t ,z t ), the position of the user on the reflecting side is p r =(x r ,y r ,z r ); The network label is l'=(R,φ,θ,t,r), where R is the Euclidean distance from the STAR-RIS reference position to the user, and φ is the distance from the user to p S The pitch angle, θ is the angle between the user and the x-axis. S The azimuth angle, t and r are not 1 at the same time. When t = 1 and r = 0, it means that the user is on the transmission side of STAR-RIS. Conversely, when t = 0 and r = 1, it means that the user is on the reflection side of STAR-RIS. R, φ, θ satisfy the following equation: ( x S -x r ,y S -y r ,z S -z r )=(Arcs φcos θ,Arcs φsin θ,Rsin φ) ( x S -x t ,y S -y t ,z S -z t )=(-Rcos φcos θ,Rcos ϕsin θ,Rsin φ) When the user moves, the channel state information matrix data of each user is continuously collected for T time slots, and each time slot is Where c is the speed of light, v is the user's moving speed, and f c is the carrier frequency, so the dimension of each user channel state matrix in the data set is T×N r ×K, vertical splicing by time slot; The N of the above matrix T time slots sample The channel state information matrix H t 、H r and label l' together constitute the data set and where N sample is the total number of samples in the dataset; S2: In each channel state information matrix in the data set, the transmitting antenna m (m∈N r ) in time slot t(t∈T) on K subcarriers is Pass it Complete the normalization of the vector and obtain the normalized channel amplitude at this time The channel vector is then converted to Converted into a complex-valued channel vector, normalized for each transmit antenna m and subcarrier k, and concatenated into a normalized dimension of T×N r ×K channel state information matrix H csi ; S3: Pre-train the channel feature extraction network CsiNet, building a neural network with multiple convolution layers. Each convolution layer includes a convolution layer, a batch normalization layer, an activation function, a pooling layer, and a three-layer fully connected layer and an activation layer structure. The network completes H by extracting the spatial position information from the channel state matrix. csi To the nonlinear mapping of l = (R, φ, θ), where l is the first three items of l', the output of the CsiNet network is The training goal of the network is to make the output of CsiNet Continuously fit the label l; S4: Train the overall network CsiNet-Former, where the prediction network is Former and the label value of the entire network is l'; S5: Calculate the user's absolute position according to The t' and r' parameters in the estimated R' are set, where if t'>r', it is -R', otherwise it is R', and the estimated p is further obtained. t '=(x t ',y t ',z t ') and p r '=(x r ',y r ',z r ') is shown as follows:
2. A user positioning method according to claim 1, characterized in that: In the S1, the transmission side user and the reflection side user do not interfere with each other, and the uplink channel state matrix is completed at the base station end. and The data set is used as the training network data set and is divided into training set, verification set and test set according to 7:1:
2. The labels correspond to the channel state matrix in the data set.
3. A user positioning method according to claim 1, characterized in that: The original vector in S2 The amplitude angle of the channel vector is converted into a complex-valued channel state matrix H csi After the normalization process, the channel state matrix has a more statistical distribution and retains the original phase angle information.
4. A user positioning method according to claim 1, characterized in that: The S3 specific training includes the following steps: S3-1: The loss function used by the CsiNet network is: S3-2: The last layer FC output of the network The feature vector mapped by the penultimate layer FC is defined as the channel position information feature fc2; S3-3: When the loss value of two adjacent rounds of training The difference is less than 1e for 5 consecutive times -5 Indicates that network training is completed and the network parameters are stored, otherwise continue training; The learning rate of the network CsiNet is γ, the stochastic gradient descent method based on batch learning is used, and the optimization method is set to the Adam algorithm.
5. A user positioning method according to claim 1, characterized in that: The specific training of S4 includes the following steps: S4-1: Set the channel gain parameter to h gain (t), where a ave (t) is the state matrix H of each channel in each time slot csi The average amplitude of (t), It is used to normalize in the time dimension and then multiplied by the user normalization parameter ∈. The overall normalization method is as follows: In the above formula a UE =Max(a ave (t)), is a ave (t) is the maximum channel gain in time slot T, and α and β are the channel gains in the entire data set. UE The maximum and minimum values of ∈ determine the relative gain of each channel gain in the entire data set; S4-2: The collected T time slot channel state matrix sequence is divided into T tr 、T ad and T pr Three parts, meet T tr +T ad +T pr =T, where T tr +T ad The time slots are used to pre-train the CsiNet-Former network, T ad The channel state matrix label l' of the time slot is used to adjust the parameters of the prediction ability of the training network; T pr This is the prediction part, using T tr +T ad The corresponding l' part calculates the prediction loss function. During the operation, the entire network dynamically inputs and the sliding window predicts the user's trajectory; S4-3: T normalized by step S2 tr +T ad The channel state matrix of the time slot is input into the pre-trained CsiNet to complete the channel spatial state feature fc2 extraction and obtain its rough estimated user location information. These two vectors and h obtained in S4-1 gain (t) horizontal splicing is x t , input to the encoder of the Former network; Among them, T pr Part of the channel state matrix is input to the CsiNet network to obtain the channel space state feature fc2 and the rough estimated user location information The encoder that is not input to the Former network, T pr Some parts are unknown to the Former network and need to be predicted; The Former consists of a three-layer encoder and a two-layer decoder, both of which are fully connected layers. The encoder is responsible for converting the input x t Mapped to hidden state, the decoder then maps the hidden state to a fine estimate of the user position Including the true position estimate T ad Partial and predicted position estimate T pr part; S4-4: The encoder of the Former network takes the input T tr +T ad x t Encoding to hidden state, the encoder is composed of FC, and the attention mechanism is introduced to further extract the relationship between sequences; S4-5: The decoder of Former completes the hidden state decoding, and its output is The dimension is (T ad +T pr )×5, including the real location estimate T ad Partial and predicted position estimate T pr part, and through L Former Complete the loss function calculation and use it to train and update network parameters, L Former The expression is as follows: where l and is the label value and output of the CsiNet network described in S3 to roughly estimate the user location, while l' and The label value and output of the CsiNet-Former network are used to estimate the user location; The first part is used to correct the estimation and prediction capabilities of the Former network, and the second part is used to correct the feature extraction and coarse user position estimation capabilities of the CsiNet network; S4-6: When the loss value of two adjacent rounds of training The difference is less than 1e for 10 consecutive times -5 Indicates that network training is completed and the network parameters are stored. Otherwise, training continues. The network learning rate is γ', and the stochastic gradient descent method based on batch learning is used. The learning rate parameter changes with the round, and the optimization method is set to the Adam algorithm.
6. According to the user positioning method of claim 1, the attention mechanism is adopted in the CsiNet-Former network, the hidden state is mapped into three matrices Q, K and V, and the attention of the calculated key-value pairs is defined as This value indicates the value of each key-value pair, which is a linear combination of V. The weight is generated according to the similarity of Q and K to obtain the probability value of the key-value pair to be queried. The higher the value, the more valuable the key-value pair is to be queried. The calculation is as follows: Where d is the input dimension, that is, x t Dimension; The query probability of each key-value pair is defined as q i is the i-th column of Q, L is the mapping dimension, and the probability p(k j ∣q i )=1 / L. The key-value pairs with larger Wasserstein distance have greater query value. By only querying the key-value pairs with greater value, Q becomes a sparse matrix, simplifying the computational complexity. The network adopts a multi-head self-attention mechanism, that is, it learns different behaviors based on the same attention mechanism, and then combines the different behaviors as knowledge to obtain the output result through an FC mapping; The introduction of mask attention in the decoder will make the available hidden layer, T tr +T ad The mask of the sequence part is set to 1, and the rest of T pr The sequence is 0, which means that the subsequent T can only be completed through the previous hidden state pr User location prediction, but cannot use T pr Complete the previous output.
Citation Information
Patent Citations
Optimized sparse antenna activation reconfigurable intelligent surface auxiliary communication method
CN112564752A
Unmanned aerial vehicle auxiliary ground communication method based on trajectory and phase joint optimization
CN114980169A