Three-dimensional indoor fingerprint positioning algorithm based on joint feature extraction network under assistance of multiple intelligent reflecting surfaces

By constructing a cascaded channel frequency response matrix of angle domain and time delay domain information and decoupling channel parameters, combined with an attention-enhanced joint feature extraction network, the problem of positioning instability caused by channel parameter coupling under intelligent reflector assistance is solved, and high-precision three-dimensional indoor positioning is achieved.

CN121486969APending Publication Date: 2026-02-06SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511742043.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

With the assistance of intelligent reflective surfaces, there is deep coupling between various channel parameters in existing fingerprint algorithms, leading to unstable positioning and insufficient accuracy.

Method used

By constructing a cascaded channel frequency response matrix containing information in the angle domain and time delay domain, decoupling the channel parameters using the inverse discrete Fourier transform, reshaping them into a three-dimensional tensor, and then using an attention-enhanced joint feature extraction network for localization, the feature learning ability is improved by utilizing the attention mechanism and residual connections.

Benefits of technology

High-precision positioning was achieved in complex multipath environments. By decoupling channel parameters and attention enhancement networks, the stability and accuracy of positioning were improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486969A_ABST
    Figure CN121486969A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional indoor fingerprint positioning algorithm based on a joint feature extraction network under the assistance of multiple intelligent reflecting surfaces. In the existing indoor positioning technology, fingerprints constructed under the assistance of most intelligent reflecting surfaces have the problem of multipath information deep coupling, which is very unfavorable for realizing accurate positioning. Therefore, the invention provides a joint angle time delay power matrix fingerprint, the fingerprint is obtained by carrying out two-dimensional inverse Fourier transform on a cascade channel frequency response matrix under the assistance of an original intelligent reflecting surface, and the angle, time delay and power information in a multipath environment is fully decoupled; and the method has higher capability of representing the propagation environment. Afterwards, the invention provides a joint feature extraction network positioning algorithm, which introduces residual connection to alleviate the problem of gradient decline caused by network depth deepening, and enhances the feature expression ability of the network through two attention mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of wireless positioning, and particularly relates to a three-dimensional indoor fingerprint positioning method based on a joint feature extraction network under the assistance of multiple intelligent reflecting surfaces. BACKGROUND

[0002] Precise wireless positioning is becoming increasingly critical in many location-based services, such as car navigation, smart cities, Internet of Things, etc. These wireless positioning technologies can generally be divided into two categories, one is the traditional two-step positioning method, and the other is the fingerprint feature-based positioning method. Among them, the two-step positioning method requires line-of-sight links to dominate, and in a complex multipath environment of non-line-of-sight links, its performance will be greatly deteriorated, and due to the need to estimate channel parameters, the complexity of these algorithms is often high. Therefore, in order to achieve stable and efficient positioning in a complex multipath environment, more and more research has turned to fingerprint feature-based positioning algorithms.

[0003] However, due to the obstruction of obstacles in the indoor environment, the direct link between the base station and the user may be cut off, which results in that the traditional fingerprint algorithm cannot achieve stable and high-precision positioning. In order to solve the above problem, more and more scholars introduce intelligent reflecting surfaces to establish virtual line-of-sight links to replace the cut-off direct links, and because the base station to intelligent reflecting surface and intelligent reflecting surface to user channels are cascaded, the new channel state response type fingerprint constructed contains more channel features than the traditional channel state response type fingerprint. However, there is a deep mutual coupling between each channel parameter in the fingerprint constructed under the assistance of the existing intelligent reflecting surface, which is very unfavorable for subsequent fingerprint feature matching. SUMMARY

[0004] In order to solve the above technical problems, the application proposes a three-dimensional indoor fingerprint positioning method based on a joint feature extraction network under the assistance of multiple intelligent reflecting surfaces. First, a cascaded channel frequency response matrix containing angle domain information and delay domain information is constructed, and then the angle domain and delay domain parameters in the cascaded channel frequency response matrix are decoupled by inverse discrete Fourier transform to obtain a joint angle-delay power matrix fingerprint. After remodeling the joint angle-delay power matrix into a three-dimensional tensor, the three-dimensional joint angle-delay power matrix fingerprint is sent into the proposed attention-enhanced joint feature extraction network to realize positioning. The network uses the attention mechanism to better learn comprehensive features and improve positioning accuracy, and the problem of gradient disappearance is alleviated by introducing residual connection.

[0005] The three-dimensional indoor fingerprint positioning algorithm based on a joint feature extraction network under the assistance of multiple intelligent reflecting surfaces comprises the following steps:

[0006] Step 1: Establish a geographic model of a three-dimensional indoor positioning system assisted by multiple intelligent reflective surfaces;

[0007] Step 2: Perform channel modeling on the links from the base station to the i-th smart reflector and from the i-th smart reflector to the user, and simultaneously obtain the cascaded channel frequency response matrix assisted by multiple smart reflectors;

[0008] Step 3: Use the phase-shift discrete Fourier transform matrix and the discrete Fourier transform matrix to extract angle and time delay features from the base station to the i-th smart reflector, respectively, to obtain the channel model from the base station to the i-th smart reflector after feature extraction; and use the discrete Fourier transform matrix to extract time delay features from the channel from the i-th smart reflector to the user, to obtain the channel model from the i-th smart reflector to the user after feature extraction.

[0009] Step 4: Substitute the two channel models obtained in Step 3 into the cascaded channel frequency response matrix in Step 2 to obtain the joint angle delay response matrix. Extract the average power of the joint angle delay response matrix and normalize it to obtain the joint angle delay power matrix. Finally, reshape the joint angle delay power matrix into a three-dimensional tensor.

[0010] Step 5: In the offline phase, user coordinates and joint angle delay power matrix fingerprints are collected at set intervals to establish an offline fingerprint database; the offline fingerprint database is fed into the proposed joint feature extraction network for feature learning, and the trained network parameters are stored.

[0011] Step 6: In the online positioning stage, unknown user coordinates and the corresponding joint angle delay power matrix fingerprint are collected and fed into the pre-trained joint feature extraction network to achieve positioning.

[0012] Furthermore, the specific features of step 1 include the following steps:

[0013] Consider an indoor multipath transmission positioning system assisted by multiple smart reflectors. The direct link between the base station and the user is severed by an obstacle, so I smart reflectors are deployed to assist positioning, where I ≥ 1. The positioning system has one base station equipped with a uniform linear array parallel to the y-axis, with M array elements. Each smart reflector is equipped with a uniform planar array, with N array elements. R =N y N z N y N is the number of array elements parallel to the y-axis. z The number of array elements is parallel to the z-axis; the user has configured a single omnidirectional antenna; considering multipath transmission of indoor signals, N is set. s Several scatterers are distributed indoors; the array spacing of all arrays is denoted as d. rThe base station transmits N subcarriers to the user through multiple smart reflective panels. c The sampling interval is T s Positioning is achieved using orthogonal frequency division multiplexing signals.

[0014] Furthermore, step 2 specifically includes the following steps:

[0015] Step 2.1: The channel modeling from the base station to the i-th smart reflector is as follows: Among them, i=1,2,…I,α i Let k be the complex gain of the link from the base station to the i-th smart reflector. i This is the normalized delay of the link. τ i This refers to the latency of the link. This indicates a rounding operation, f(k) i () represents the frequency response vector caused by the normalized delay of this link. (·) T Represents the transpose of a matrix or vector. This represents the array response vector of the antenna array on the smart reflector of the link. and Let represent the array response vectors of the antenna array on the smart reflector in the vertical direction and the array response vectors in the horizontal direction, respectively. θ i This represents the pitch angle in the angle of arrival of the link. d represents the azimuth angle in the angle of arrival of this link. r λ represents the element spacing of the antenna array. c Indicates the signal wavelength. This represents the Kronecker product, (·). H e represents the conjugate transpose of a matrix or vector. B (β i ) is the array response vector of the antenna array on the base station of this link. β i Indicates the departure angle of the link;

[0016] Step 2.2, the channel modeling from the i-th smart reflector to the user is as follows: Where s = 0, 1, ..., N s ;α′ i,s e R (ψ i,s ,φ i,s ), k′ i,s and f(k′) i,sLet represent the complex gain of the s-th link from the i-th smart reflector to the user, the array response vector of the antenna array on the smart reflector, the normalized delay, and the frequency response vector caused by the normalized delay, respectively.

[0017] Step 2.3, Combine H i,BR and H i,BR The cascaded channel frequency response matrix is ​​obtained, i.e. in, The average power of the transmitted orthogonal frequency division multiplexed signal. w is the phase shift vector of the i-th smart reflector. i,k This represents the phase shift of the k-th element of the antenna array on the smart reflector, where k = 1, 2, ..., N. R , diag(·) represents the diagonalization operation on a vector.

[0018] Furthermore, step 3 specifically includes the following steps:

[0019] Step 3.1: First, define an N×N dimensional DFT matrix. and N×N dimensional phase-shift DFT matrix [X] l,k Let X represent the element in the l-th row and k-th column of a matrix X; the angle domain and time delay domain information of the cascaded channel are decoupled from the frequency response matrix of the cascaded channel using a two-dimensional discrete Fourier inverse transform, which is expressed as:

[0020]

[0021] in, This represents the M×M dimensional phase-shifted DFT matrix after the conjugate transpose operation. N represents the result after the conjugation operation. c ×N c The DFT matrix is ​​2D, and G represents the frequency response matrix of the decoupled cascaded channel. i,BR G is the channel from the base station to the i-th smart reflector after decoupling the angle domain and time delay domain information. i,RU It is the channel from the i-th smart reflector to the user after decoupling the time delay domain information. (·) * Represents the conjugate of a matrix or vector;

[0022] Step 3.2, G i,BR Expressed as follows:

[0023]

[0024] So, after extracting the angular features using the phase-shift discrete Fourier transform... It is given by the following formula:

[0025]

[0026] in, Let x represent a scalar of the independent variable, [x] k Let x represent the k-th element of a vector x, where l, k = 0, ..., M-1; as M → ∞, according to F M Asymptotic properties of (x), F M (x) has a peak M only at integer indices, and the values ​​at other indices tend to 0; and because Then if and only if At that time, F M (x) has a peak value M, at which point the corresponding value is... There is a peak value The values ​​at the remaining indices are close to 0; therefore, Approximately:

[0027]

[0028] Where δ(·) represents the Dirac impulse function;

[0029] Step 3.3: Extracting time delay features after Discrete Fourier Transform It is given by the following formula:

[0030]

[0031] in, When N c →∞, Approximately:

[0032]

[0033] Step 3.4, based on the obtained and [G i,BR ] k,r Rewritten as:

[0034]

[0035] Among them, [X] k,r Let X represent the element in the k-th row and r-th column of a matrix X, where k = 0, ..., MN. c -1, r = 0, ..., N R -1; [G] i,RU ] r,k Rewritten

[0036]

[0037] Among them, [X] r,kLet r represent the element in the r-th row and k-th column of a matrix X, where r = 0, ..., N. R -1, k = 0, ..., N c -1.

[0038] Furthermore, step 4 specifically includes the following steps:

[0039] Step 4.1: Combine the two channel models G obtained in Step 3. i,BR and G i,RU Substituting these values ​​into the cascaded channel frequency response matrix of step 2, we obtain the joint angle delay response matrix, where the elements in the l-th row and k-th column of the joint angle delay response matrix are represented by the following equation:

[0040]

[0041] Where l = 0, ..., MN c -1, k = 0, ..., N c -1, This represents the total channel gain at a specific location. Where ⊙ represents the Hadamard product;

[0042] Step 4.2: Since positioning requires stable channel information, the joint angular delay power matrix is ​​defined as follows:

[0043]

[0044] Where E{·} denotes the desired operation;

[0045] Step 4.3: For the joint angular delay power matrix P g Normalization is performed to obtain P g ,Right now Among them, ||·|| F Let Frobenius norm be the matrix; then, Reconstructed into a three-dimensional tensor, i.e. Its three dimensions respectively indicate the departure angle of the link from the base station to the smart reflector, the latency of the link from the base station to the smart reflector, and the latency of the link from the smart reflector to the user, [V] v,h,n This represents the (v,h,n)th element of a certain quantity V.

[0046] Furthermore, step 5 specifically includes the following steps:

[0047] Step 5.1: Consider a 16m×12m×5m 3D indoor positioning scene. In the offline phase, reference points are collected at intervals of 0.5m×0.5m×0.25m, and the fingerprints and 3D coordinates corresponding to these reference points are saved. The training sample set consists of A 3D joint angle delay power matrix samples V, i.e. For the i-th training sample V i The corresponding output label is the continuous three-dimensional coordinates of the sample. Therefore, the complete training dataset is represented as

[0048] Step 5.2: Feed the training dataset into the joint feature extraction network for offline training. The joint feature extraction network consists of three parts: a joint feature base module, an attention enhancement layer module, and a regression module. The joint feature base module uses two parallel branches with asymmetric convolutional kernels to extract the time-delay domain features of the base station-to-smart reflector link and the smart reflector-to-user link, respectively, and fuses the different features. Then, pooling layers are used to downsample the feature maps, while residual blocks help the network learn deeper features. The feature values ​​after the convolutional layers are given by the following formula:

[0049] V (k) =Act(BN(W (k) *V (k-1) +b (k) )),

[0050] Among them, V (k-1) V represents the input of the k-th convolutional layer. (k) W represents the output of the k-th convolutional layer. (k) and b (k) V represents the weights and biases of the convolutional kernel in the k-th layer, respectively. (0) =V represents the original input, (*) is the convolution operation, BN(·) represents the batch normalization operation, and Act(·) is the activation function; the feature values ​​after the residual block are represented as:

[0051] Y out =Act(F(V) (1) )+F d (V (1) )),

[0052] Among them, Y out V represents the output of the residual block. (1) F represents the output after the first convolutional layer. d For quick connection, F(·) represents the feature values ​​after passing through two convolutional layers;

[0053] Step 5.3: The feature map after passing through the joint feature base module is input into the attention enhancement layered module. In addition to convolutional blocks, residual blocks and pooling layers, two attention mechanisms are introduced in the attention enhancement layered module to realize feature extraction, namely the convolutional block attention module mechanism and the self-attention mechanism. The attention enhancement residual block is obtained by combining the attention mechanism with the residual connection.

[0054] Step 5.4: After the first two modules, the extracted deep features are fed into the regression module, which consists of a global average pooling layer, a Dropout layer, and a fully connected layer without an activation function. The global average pooling layer compresses the multi-dimensional feature maps into one-dimensional feature vectors by calculating the global average of each feature map. Following the global average pooling layer is the Dropout layer, which acts as a regularization mechanism, randomly dropping neurons to mitigate the risk of overfitting. Finally, the feature vector processed by the Dropout layer is fed into a fully connected layer without an activation function, which directly outputs the final estimated three-dimensional coordinates. The loss function used in the proposed joint feature extraction network is the mean squared error function, given by the following equation:

[0055]

[0056] Where θ is the set of trainable parameters of the network, and ||·|| denotes the Euclidean norm of the matrix; the proposed joint feature extraction network is trained on the training dataset. The training is performed offline in a supervised manner, and the network parameters θ are updated using the backpropagation algorithm and gradient descent method; the trained network is then stored.

[0057] Furthermore, step 6 specifically includes the following steps:

[0058] Step 6 specifically includes the following steps:

[0059] During the online localization phase, 500 unknown user coordinates and their corresponding joint angle delay power matrix fingerprints are collected from the localization area. These fingerprints are then fed into a pre-trained joint feature extraction network to achieve localization. The joint feature extraction network obtains the final user coordinate estimate by minimizing the error between the predicted and true coordinates.

[0060] The beneficial effects of this invention are as follows:

[0061] 1. The angle and time delay domain information in the original cascaded channel frequency response matrix were decoupled by using the phase-shifted discrete Fourier transform matrix and the discrete Fourier transform matrix, respectively, which can more clearly reveal the characteristics of the propagation environment;

[0062] 2. A localization algorithm based on attention-enhanced joint feature extraction network is designed. The network is composed of multiple residual blocks and attention-enhanced residual blocks stacked together. The residual connections are beneficial for training deep networks, and the attention mechanism can enhance the feature extraction and representation capabilities of the network, thereby achieving high-precision localization performance. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0064] Figure 1 This is a schematic diagram of a three-dimensional indoor positioning scene assisted by multiple intelligent reflective surfaces established in this invention;

[0065] Figure 2 This is a simplified structural diagram of the attention-enhanced joint feature extraction network of the present invention;

[0066] Figure 3 This is a simplified diagram of the attention-enhanced residual block structure in the attention-enhanced joint feature extraction network of the present invention. Detailed Implementation

[0067] To make the objectives and technical solutions of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be described more clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0068] Step 1: Establish a geographic model for a 3D indoor positioning system assisted by multiple intelligent reflective surfaces:

[0069] (1.1) As Figure 1 As shown, consider an indoor multipath transmission positioning system assisted by multiple intelligent reflectors. The direct link between the base station and the user is severed by obstacles. Therefore, I (I≥1) intelligent reflectors are deployed to assist positioning, i=1,2,…I. Assume the positioning system has one base station with a uniform linear array parallel to the y-axis, containing M array elements; each intelligent reflector is equipped with a uniform planar array, containing N array elements. R =N y N z N y N z Let N be the number of array elements parallel to the y-axis and z-axis, respectively; the user has configured a single omnidirectional antenna; considering multipath transmission of indoor signals, assume there are N... s Several scatterers are distributed indoors. For simplicity, it is assumed that the array spacing of all arrays is denoted as d. r The base station transmits N subcarriers to the user through multiple smart reflective panels. c The sampling interval is T sPositioning is achieved using orthogonal frequency division multiplexing signals.

[0070] Step 2: Perform channel modeling on the links from the base station to the i-th smart reflector and from the i-th smart reflector to the user, and simultaneously obtain the cascaded channel frequency response matrix assisted by multiple smart reflectors:

[0071] (2.1) The channel from the base station to the i-th smart reflector can be modeled as follows: Where, α i k is the complex gain of the link. i For normalized delay, τ i For time delay, Indicates rounding operation. (·) T Represents the transpose of a matrix or vector. The array response vector representing the smart reflector. θ i , λ represents the pitch and azimuth angles in the angle of arrival of this link, respectively. c Indicates the signal wavelength. This represents the Kronecker product, (·). H e represents the conjugate transpose of a matrix or vector. B (β i ) is the array response vector of the base station. β i This indicates the departure angle of the link.

[0072] (2.2) Similarly, the channel from the i-th smart reflector to the base station can be modeled as follows: Where, α′ i,s e R (ψ i,s ,φ i,s ), k′ i,s and f(k′) i,s The meaning of ) and the channel model H from the base station to the i-th smart reflector i,BR α in i , k i and f(k) i () have the same meaning.

[0073] (2.3) Combined H i,BR and H i,BR The cascaded channel frequency response matrix is ​​obtained, i.e. in, The average power of the transmitted orthogonal frequency division multiplexed signal. It is the phase shift vector of the i-th smart reflective surface, and diag(·) represents the diagonalization operation on the vector.

[0074] Step 3: Extract the angle features of the channel from the base station to the i-th smart reflector using the phase-shifted discrete Fourier transform matrix, and extract the delay features of the channel from the base station to the i-th smart reflector and the channel from the i-th smart reflector to the user using the discrete Fourier transform matrix.

[0075] (3.1) First, define the discrete DFT matrix. and Discrete Phase Shift (DFT) Matrix [X] l,k Let X represent the element in the l-th row and k-th column of matrix X. The inverse two-dimensional discrete Fourier transform can be used to decouple the angle and time delay domain information of the cascaded channel from its frequency response matrix, which can be expressed as:

[0076]

[0077] in, (·) * Represents the conjugate of a matrix or vector.

[0078] (3.2) Further, G can be... i,BR Expressed as follows:

[0079]

[0080] So, after extracting the angular features using the phase-shift discrete Fourier transform... It is given by the following formula:

[0081]

[0082] in, [x] k Let x represent the k-th element of vector x, where l, k = 0, ..., M-1. As M → ∞, according to F... M Asymptotic properties of (x), F M (x) has a peak M only at integer indices; values ​​at other indices tend to 0. Also, because... Then if and only if At that time, F M (x) has a peak value M, at which point the corresponding value is... There is a peak value The values ​​at the remaining indices are close to 0. Therefore, It can be approximated as:

[0083]

[0084] Where δ(·) represents the Dirac impulse function.

[0085] (3.3) Similarly, after extracting the time delay features via discrete Fourier transform... It is given by the following formula:

[0086]

[0087] in, When N c →∞, Similarly, it can be approximated as:

[0088]

[0089] (3.4) Based on the obtained and [G i,BR ] k,r It can be rewritten as:

[0090]

[0091] Where k = 0, ..., MN c -1, r = 0, ..., N R -1. Similarly, [G i,RU ] r,k Can be rewritten as

[0092]

[0093] Where r = 0, ..., N R -1, k = 0, ..., N c -1.

[0094] Step four: Substitute the two channel models obtained in (3) into the cascaded channel frequency response matrix in step (2) to obtain the joint angle delay response matrix. Extract the average power of the joint angle delay response matrix and normalize it to obtain the joint angle delay power matrix. Finally, reshape the joint angle delay power matrix into a three-dimensional tensor:

[0095] (4.1) Substitute the two channel models obtained in step (3) into the cascaded channel frequency response matrix in step (2) to obtain the joint angular delay response matrix, where the elements in the l-th row and k-th column of the matrix are represented by the following formula:

[0096]

[0097] Where l = 0, ..., MN c -1, k = 0, ..., N c -1, This represents the total channel gain at a specific location. ⊙ represents the Hadamard product.

[0098] (4.2) Since positioning requires stable channel information, the joint angle delay power matrix is ​​defined as follows:

[0099]

[0100] Where E{·} represents the desired operation.

[0101] (4.3) For the joint angular delay power matrix P g Normalization process is performed to obtain Right now Among them, ||·|| F Let f(x) represent the Frobenius norm of the matrix. Then, ... Reconstructed into a three-dimensional tensor, i.e. Its three dimensions respectively indicate the departure angle of the link from the base station to the smart reflector, the latency of the link from the base station to the smart reflector, and the latency of the link from the smart reflector to the user.

[0102] Step 5: In the offline phase, user coordinates and joint angular delay power matrix fingerprints are collected at set intervals to establish an offline fingerprint database. This offline fingerprint database is then fed into the proposed joint feature extraction network for feature learning, and the trained network is stored.

[0103] (5.1) Consider a 16m × 12m × 5m 3D indoor positioning scene. During the offline phase, reference points are acquired at intervals of 0.5m × 0.5m × 0.25m, and the fingerprints and 3D coordinates corresponding to these reference points are saved. Then, the training samples consist of A 3D joint angle delay power matrix samples V, i.e. For the i-th training sample V i The corresponding output label is the continuous three-dimensional coordinates of the sample. Therefore, the complete training dataset is represented as

[0104] (5.2) The training dataset is fed into the proposed joint feature extraction network for offline training, such as... Figure 2 As shown, the network consists of three parts: a joint feature base module, an attention-enhanced hierarchical module, and a regression module. The joint feature base module uses two parallel branches with asymmetric convolutional kernels to extract delay-domain features from the base station to the smart reflector link and from the smart reflector to the user link, respectively. These different features are then fused to form a more comprehensive representation. Pooling layers are then used to downsample the feature maps, while residual blocks help the network learn deeper features. The feature values ​​after the convolutional layers are given by the following formula:

[0105] V (k) =Act(BN(W (k) *V (k-1)+b (k) )),

[0106] Among them, V (k-1) W represents the input of the k-th convolutional layer. (k) and b (k) These represent the weights and biases of the convolutional kernel in the k-th layer, respectively.

[0107] V (0) =V represents the original input, (*) is the convolution operation, BN(·) represents the batch normalization operation, and Act(·) is the activation function. The feature values ​​after the residual block can be expressed as:

[0108] Y out =Act(F(V) (1) )+F d (V (1) )),

[0109] Among them, Y out F represents the output of the residual block. d For quick connection, F(·) represents the feature values ​​after passing through two convolutional layers.

[0110] (5.3) The feature map after passing through the joint feature base module is input into the attention enhancement layered module. Two attention mechanisms are introduced in the attention enhancement layered module to achieve more comprehensive feature extraction: the convolutional block attention module mechanism and the self-attention mechanism. The attention-enhanced residual block is obtained by combining the attention mechanism with residual connections, and its structure is as follows: Figure 3 As shown.

[0111] (5.4) After the first two modules, the extracted deep features are fed into a global average pooling layer. This layer compresses the multi-dimensional feature maps into a one-dimensional feature vector by calculating the global average of each feature map. Following the global average pooling layer is a Dropout layer, which acts as a regularization mechanism, randomly dropping neurons to mitigate the risk of overfitting. Finally, the feature vector processed by the Dropout layer is fed into a fully connected layer without an activation function, which directly outputs the final estimated three-dimensional coordinates. The loss function used in the proposed network is the mean squared error function, given by the following equation:

[0112]

[0113] Where θ is the set of trainable parameters of the network, and ||·|| denotes the Euclidean norm of the matrix. The proposed network is trained on the training dataset. The training was performed offline in a supervised manner, and the network parameters θ were updated using backpropagation and gradient descent. The trained network was then stored.

[0114] Step six: In the online localization stage, unknown user coordinates and the corresponding joint angle delay power matrix fingerprint are collected and fed into the pre-trained joint feature extraction network to achieve localization.

[0115] (6.1) In the online localization stage, 500 unknown user coordinates and their corresponding joint angle delay power matrix fingerprints are collected in the localization area. These fingerprints are then fed into a pre-trained joint feature extraction network to achieve localization. This network obtains the final user coordinate estimate by minimizing the error between the predicted and true coordinates.

[0116] This invention utilizes two-dimensional discrete Fourier inverse transform to decouple cascaded channel parameters. The time delay information from the decoupled intelligent reflector to the user link is directly related to the user location. Then, the proposed joint feature extraction network is used to achieve positioning. This network refines the input fingerprint, extracts more comprehensive joint features, and uses an attention residual learning mechanism to enhance feature representation capabilities, achieving high-precision positioning results.

Claims

1. A three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces, characterized in that, Includes the following steps: Step 1: Establish a geographic model of a three-dimensional indoor positioning system assisted by multiple intelligent reflective surfaces; Step 2: Perform channel modeling on the links from the base station to the i-th smart reflector and from the i-th smart reflector to the user, and simultaneously obtain the cascaded channel frequency response matrix assisted by multiple smart reflectors; Step 3: Use the phase-shift discrete Fourier transform matrix and the discrete Fourier transform matrix to extract angle and time delay features from the base station to the i-th smart reflector, respectively, to obtain the channel model from the base station to the i-th smart reflector after feature extraction; and use the discrete Fourier transform matrix to extract time delay features from the channel from the i-th smart reflector to the user, to obtain the channel model from the i-th smart reflector to the user after feature extraction. Step 4: Substitute the two channel models obtained in Step 3 into the cascaded channel frequency response matrix in Step 2 to obtain the joint angle delay response matrix. Extract the average power of the joint angle delay response matrix and normalize it to obtain the joint angle delay power matrix. Finally, reshape the joint angle delay power matrix into a three-dimensional tensor. Step 5: In the offline phase, user coordinates and joint angle delay power matrix fingerprints are collected at set intervals to establish an offline fingerprint database; the offline fingerprint database is fed into the proposed joint feature extraction network for feature learning, and the trained network parameters are stored. Step 6: In the online positioning stage, unknown user coordinates and the corresponding joint angle delay power matrix fingerprint are collected and fed into the pre-trained joint feature extraction network to achieve positioning.

2. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 1, characterized in that, The specific features of step 1 include the following steps: Consider an indoor multipath transmission positioning system assisted by multiple smart reflectors. The direct link between the base station and the user is severed by an obstacle, so I smart reflectors are deployed to assist positioning, where I ≥ 1. The positioning system has one base station equipped with a uniform linear array parallel to the y-axis, with M array elements. Each smart reflector is equipped with a uniform planar array, with N array elements. R =N y N z N y N is the number of array elements parallel to the y-axis. z The number of array elements is parallel to the z-axis; the user has configured a single omnidirectional antenna; considering multipath transmission of indoor signals, N is set. s Several scatterers are distributed indoors; the array spacing of all arrays is denoted as d. r The base station transmits N subcarriers to the user through multiple smart reflective panels. c The sampling interval is T s Positioning is achieved using orthogonal frequency division multiplexing signals.

3. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 2, characterized in that, Step 2 specifically includes the following steps: Step 2.1: The channel modeling from the base station to the i-th smart reflector is as follows: Among them, i=1,2,…I,α i Let k be the complex gain of the link from the base station to the i-th smart reflector. i This is the normalized delay of the link. τ i This refers to the latency of the link. This indicates a rounding operation, f(k) i () represents the frequency response vector caused by the normalized delay of this link. (·) T Represents the transpose of a matrix or vector. This represents the array response vector of the antenna array on the smart reflector of the link. and Let represent the array response vectors of the antenna array on the smart reflector in the vertical direction and the array response vectors in the horizontal direction, respectively. θ i This represents the pitch angle in the angle of arrival of the link. d represents the azimuth angle in the angle of arrival of this link. r λ represents the element spacing of the antenna array. c Indicates the signal wavelength. This represents the Kronecker product, (·). H e represents the conjugate transpose of a matrix or vector. B (β i ) is the array response vector of the antenna array on the base station of this link. β i Indicates the departure angle of the link; Step 2.2, the channel modeling from the i-th smart reflector to the user is as follows: Where s = 0, 1, ..., N s ;α′ i,s e R (ψ i,s ,φ i,s ), k′ i,s and f(k′) i,s Let represent the complex gain of the s-th link from the i-th smart reflector to the user, the array response vector of the antenna array on the smart reflector, the normalized delay, and the frequency response vector caused by the normalized delay, respectively. Step 2.3, Combine H i,BR and H i,BR The cascaded channel frequency response matrix is ​​obtained, i.e. in, The average power of the transmitted orthogonal frequency division multiplexed signal. w is the phase shift vector of the i-th smart reflector. i,k This represents the phase shift of the k-th element of the antenna array on the smart reflector, where k = 1, 2, ..., N. R , diag(·) represents the diagonalization operation on a vector.

4. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 3, characterized in that, Step 3 specifically includes the following steps: Step 3.1: First, define an N×N dimensional DFT matrix. and N×N dimensional phase-shift DFT matrix [X] l,k Let X represent the element in the l-th row and k-th column of a matrix X; the angle domain and time delay domain information of the cascaded channel are decoupled from the frequency response matrix of the cascaded channel using a two-dimensional discrete Fourier inverse transform, which is expressed as: in, This represents the M×M dimensional phase-shifted DFT matrix after the conjugate transpose operation. N represents the result after the conjugation operation. c ×N c The DFT matrix is ​​2D, and G represents the frequency response matrix of the decoupled cascaded channel. i,BR G is the channel from the base station to the i-th smart reflector after decoupling the angle domain and time delay domain information. i,RU It is the channel from the i-th smart reflector to the user after decoupling the time delay domain information. (·) * Represents the conjugate of a matrix or vector; Step 3.2, G i,BR Expressed as follows: So, after extracting the angular features using the phase-shift discrete Fourier transform... It is given by the following formula: in, Let x represent a scalar of the independent variable, [x] k Let x represent the k-th element of a vector x, where l, k = 0, ..., M-1; as M → ∞, according to F M Asymptotic properties of (x), F M (x) has a peak M only at integer indices, and the values ​​at other indices tend to 0; and because cosβ i ∈(-1,1), 0≤k≤M-1, then if and only if At that time, F M (x) has a peak value M, at which point the corresponding value is... There is a peak value The values ​​at the remaining indices are close to 0; therefore, Approximately: Where δ(·) represents the Dirac impulse function; Step 3.3: Extracting time delay features after Discrete Fourier Transform It is given by the following formula: in, When N c →∞, Approximately: Step 3.4, based on the obtained and [G i,BR ] k,r Rewritten as: Among them, [X] k,r Let X represent the element in the k-th row and r-th column of a matrix X, where k = 0, ..., MN. c -1, r = 0, ..., N R -1; [G] i,RU ] r,k Rewritten Among them, [X] r,k Let r represent the element in the r-th row and k-th column of a matrix X, where r = 0, ..., N. R -1, k = 0, ..., N c -1.

5. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 4, characterized in that, Step 4 specifically includes the following steps: Step 4.1: Combine the two channel models G obtained in Step 3. i,BR and G i,RU Substituting these values ​​into the cascaded channel frequency response matrix of step 2, we obtain the joint angle delay response matrix, where the elements in the l-th row and k-th column of the joint angle delay response matrix are represented by the following equation: Where l = 0, ..., MN c -1, k = 0, ..., N c -1, This represents the total channel gain at a specific location. Where ⊙ represents the Hadamard product; Step 4.2: Since positioning requires stable channel information, the joint angular delay power matrix is ​​defined as follows: Where E{·} denotes the desired operation; Step 4.3: For the joint angular delay power matrix P g Normalization process is performed to obtain Right now Among them, ||·|| F Let Frobenius norm be the matrix; then, Reconstructed into a three-dimensional tensor, i.e. Its three dimensions respectively indicate the departure angle of the link from the base station to the smart reflector, the latency of the link from the base station to the smart reflector, and the latency of the link from the smart reflector to the user, [V] v,h,n This represents the (v,h,n)th element of a certain quantity V.

6. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 5, characterized in that, Step 5 specifically includes the following steps: Step 5.1: Consider a 16m×12m×5m 3D indoor positioning scene. In the offline phase, reference points are collected at intervals of 0.5m×0.5m×0.25m, and the fingerprints and 3D coordinates corresponding to these reference points are saved. The training sample set consists of A 3D joint angle delay power matrix samples V, i.e. For the i-th training sample V i The corresponding output label is the continuous three-dimensional coordinates of the sample. Therefore, the complete training dataset is represented as Step 5.2: Feed the training dataset into the joint feature extraction network for offline training. The joint feature extraction network consists of three parts: a joint feature base module, an attention enhancement layer module, and a regression module. The joint feature base module uses two parallel branches with asymmetric convolutional kernels to extract the time-delay domain features of the base station-to-smart reflector link and the smart reflector-to-user link, respectively, and fuses the different features. Then, pooling layers are used to downsample the feature maps, while residual blocks help the network learn deeper features. The feature values ​​after the convolutional layers are given by the following formula: V (k) =Act(BN(W (k) *V (k-1) +b (k) )), Among them, V (k-1) V represents the input of the k-th convolutional layer. (k) W represents the output of the k-th convolutional layer. (k) and b (k) V represents the weights and biases of the convolutional kernel in the k-th layer, respectively. (0) =V represents the original input, (*) is the convolution operation, BN(·) represents the batch normalization operation, and Act(·) is the activation function; the feature values ​​after the residual block are represented as: Y out =Act(F(V (1) )+F d (V (1) )), Among them, Y out V represents the output of the residual block. (1) F represents the output after the first convolutional layer. d For quick connection, F(·) represents the feature values ​​after passing through two convolutional layers; Step 5.3: The feature map after passing through the joint feature base module is input into the attention enhancement layered module. In addition to convolutional blocks, residual blocks and pooling layers, two attention mechanisms are introduced in the attention enhancement layered module to realize feature extraction, namely the convolutional block attention module mechanism and the self-attention mechanism. The attention enhancement residual block is obtained by combining the attention mechanism with the residual connection. Step 5.4: After the first two modules, the extracted deep features are fed into the regression module, which consists of a global average pooling layer, a Dropout layer, and a fully connected layer without an activation function. The global average pooling layer compresses the multi-dimensional feature maps into one-dimensional feature vectors by calculating the global average of each feature map. Following the global average pooling layer is the Dropout layer, which acts as a regularization mechanism, randomly dropping neurons to mitigate the risk of overfitting. Finally, the feature vector processed by the Dropout layer is fed into a fully connected layer without an activation function, which directly outputs the final estimated three-dimensional coordinates. The loss function used in the proposed joint feature extraction network is the mean squared error function, given by the following equation: Where θ is the set of trainable parameters of the network, and ||·|| denotes the Euclidean norm of the matrix; the proposed joint feature extraction network is trained on the training dataset. The training is performed offline in a supervised manner, and the network parameters θ are updated using the backpropagation algorithm and gradient descent method; the trained network is then stored.

7. The three-dimensional indoor fingerprint localization algorithm based on a joint feature extraction network assisted by multiple intelligent reflective surfaces according to claim 6, characterized in that, Step 6 specifically includes the following steps: During the online localization phase, 500 unknown user coordinates and their corresponding joint angle delay power matrix fingerprints are collected from the localization area. These fingerprints are then fed into a pre-trained joint feature extraction network to achieve localization. The joint feature extraction network obtains the final user coordinate estimate by minimizing the error between the predicted and true coordinates.