5G Indoor Positioning Method Based on SSB Signal Spatiotemporal Feature Enhancement and Attention Mechanism
Patent Information
- Application Number
- CN202511329047.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-09-17
AI Technical Summary
[0005]尽管取得了这些重大进展,但整个基于人工智能的本地化领域仍然面临着阻碍大规模现实部署的关键瓶颈
[0073] 1. By using a TSA module, this invention effectively mitigates signal degradation caused by multipath effects and improves the robustness of fingerprint features; the use of a time-weighted method in the fingerprint construction process reduces transient interference and improves positioning accuracy in dynamic environments.
Smart Images

Figure CN121037774B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of wireless communication and indoor positioning technology, specifically a 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism. Background Technology
[0002] The rapid proliferation of the Internet of Things (IoT) and the large-scale deployment of 5G mobile networks have spurred an exponential increase in demand for high-precision indoor positioning across various scenarios, including smart homes, smart healthcare, industrial automation, and commercial complexes. Market research predicts that the global indoor positioning market will exceed $40 billion by 2026, with a compound annual growth rate of 32%. Location-based services are a core application in the 5G era, and their efficiency in complex indoor environments highly depends on positioning accuracy and real-time performance. Although Global Navigation Satellite Systems (GNSS) achieve meter-level accuracy outdoors, their performance degrades significantly indoors due to non-line-of-sight (NLOS) propagation and multipath effects. This can lead to significant signal attenuation, positioning errors exceeding 10 meters, or even complete failure. For example, in multi-story buildings or underground facilities, GNSS accuracy may drop to 30 meters, failing to meet the stringent requirements of critical tasks such as emergency response and asset tracking.
[0003] Artificial intelligence-based fingerprinting, particularly with the use of rich channel state information (CSI), has revolutionized indoor positioning, offering a powerful alternative to traditional methods. Initial efforts focused on conventional machine learning, employing classic models such as K-Nearest Neighbors (KNN) and Random Forests (RFFP) to achieve sub-meter accuracy. Simultaneously, fingerprint features themselves evolved from simple metrics to complex structures, such as the Angular Delay Domain Channel Power Matrix (ADCPM), which captures detailed path information. However, this first wave of technology was fundamentally limited by its reliance on manual feature engineering, a persistent trade-off between performance and efficiency, and a high sensitivity to environmental changes.
[0004] To overcome these shortcomings, the field turned to deep learning (DL), which leverages its ability to automatically extract features from high-dimensional data. This paradigm began with deep neural networks (DNNs) such as DeepFi and quickly evolved into the widespread use of convolutional neural networks (CNNs), which excel at handling CSI matrices like those in images. Subsequently, more sophisticated architectures using LSTM and attention mechanisms emerged, pushing the boundaries of accuracy. Furthermore, methods such as transfer learning and semi-supervised learning can alleviate the reliance on extensively labeled data, providing new avenues for cross-scenario generalization. Therefore, intelligent fingerprint positioning technology for 5G / IoT terminals has significant research value. Theoretically, it promotes the deep integration of wireless sensing and artificial intelligence.
[0005] Despite these significant advances, the entire field of AI-based localization still faces key bottlenecks hindering large-scale real-world deployment. The main challenges today are: data dependence, computational cost, and poor generalization ability. Finally, the competing need to balance model performance, deployment cost, and environmental robustness remains a core challenge for the future of this field.
[0006] To this end, this invention proposes a 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism. Summary of the Invention
[0007] The purpose of this invention is to provide a 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism, which can effectively reduce signal degradation caused by multipath effect and fundamentally improve the robustness of fingerprint features; enhance the dynamic expressiveness of features by fusing historical channel information, and achieve a dual improvement in feature representation and construction efficiency; and improve the generalization ability of the model by dynamically weighting complex multipath signals using channel attention.
[0008] According to a first aspect of the present invention, in order to achieve the above-mentioned objective, the present invention provides the following technical solution: a 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism, comprising the following steps:
[0009] The channel state information data of 5G New Radio is received, and a three-level adaptive filter is designed to preprocess the channel state information data.
[0010] Based on the preprocessed channel state information data, a time-series-angle-delay channel frequency amplitude fingerprint matrix that integrates historical channel information is constructed.
[0011] An attention-based augmented residual network model is constructed and trained to obtain the trained augmented residual network model. The input of the augmented residual network model is the fingerprint matrix, and the output is the position coordinates.
[0012] The preprocessed fingerprint data collected online is input into the trained enhanced residual network model for location estimation, thus completing indoor fingerprint localization.
[0013] Furthermore, the three-stage adaptive filter includes an amplitude processing module and a phase processing module, as detailed below:
[0014] (21) The amplitude processing module performs data preprocessing, as follows:
[0015] (21.1) Use the Hampel identifier to remove outliers from the channel state information subcarrier amplitude sequence;
[0016] (21.2) Smooth each subcarrier sequence using a wavelet filter based on an improved wavelet threshold function;
[0017] (21.3) Use a cross-correlation detector to eliminate outliers in the channel state information data;
[0018] (22) The phase processing module performs data preprocessing, as follows:
[0019] (22.1) Use the phase expansion algorithm to eliminate periodic jumps in the original phase data;
[0020] (22.2) A linear regression model based on the minimum mean square error criterion is used to compensate for the error in the expanded phase;
[0021] (22.3) Phase outlier processing is performed using the phase TSA module, which is the same as the amplitude TSA method.
[0022] Furthermore, in step (21.3), a cross-correlation detector is used to eliminate outliers in the channel state information data, as follows:
[0023] For the nth CSI sample at a given location, the average correlation coefficient of all samples at that point is calculated as follows:
[0024]
[0025] Where M is the total number of CSI samples collected at a single node, affecting the sampling density of the positioning system; Cov(·) measures the linear correlation between samples, capturing the spatial features of CSI; σ(n) and σ(m) are A n and A m The standard deviation of represents the amplitude dispersion. If (Predefined threshold) The system will classify abnormal sample A n Replace with the mean of the sample matrix, i.e.
[0026] Furthermore, the linear regression model based on the minimum mean square error criterion in step (22.2) is used to compensate for the error in the expanded phase, and the specific formula is as follows:
[0027]
[0028] in To correct the phase, φ(k) is the unpacking phase, d(k) represents the index of the k-th pilot subcarrier in the OFDM symbol, and K represents the total number of pilot subcarriers. This compensation isolates the hardware-induced bias from the true spatial channel characteristics, thereby generating distinguishable phase waveform fingerprints at different locations for subsequent feature extraction.
[0029] Furthermore, based on the preprocessed channel state information data, a time-series-angle-delay channel frequency amplitude fingerprint matrix incorporating historical channel information is constructed, as follows:
[0030] (51) Let F be the channel frequency response matrix of the angle delay from the k-th user to the base station. k for:
[0031]
[0032] Where V is the phase-shifted DFT matrix, expressed as:
[0033]
[0034] In the formula (·) H It is the conjugate transpose. It is the channel frequency response matrix;
[0035] (52) Due to F k The computational burden of complex operations used for positioning in the matrix, the angular delay generated by the Hadamard inner product, and the channel power P k ,Right now:
[0036]
[0037] Where E{·} denotes the expected value. This represents the Hadama dot product; then, by taking P... k The square root of the element is used to derive the Angular Delay Channel Frequency Amplitude (ADCFA) matrix C. k :
[0038]
[0039] (53) An exponential decay weighting function is used to perform time-weighted fusion of the current and historical angle-delay channel frequency amplitude matrices to form a time-series-angle-delay channel frequency amplitude fingerprint matrix:
[0040]
[0041] In the formula t c It is the current timestamp, t n is the timestamp of the nth historical fingerprint, λ is the control attenuation rate, set to 0.1; the current angle-delay channel frequency amplitude matrix C k The time-angle-delay channel frequency-amplitude fingerprint matrix T is formed by fusing with historical fingerprints. k :
[0042]
[0043] Where N is the sliding window size, storing the latest N angle-delay channel frequency amplitude matrices, and normalization ensures that the sum of the weights is 1;
[0044] (54) To avoid redundant calculations, the recursive calculation formula is as follows:
[0045]
[0046] Where T k ′ is the updated time-angle-delay channel frequency amplitude fingerprint matrix, w′ and C k ′ is the latest weight and angle-delay channel frequency amplitude matrix.
[0047] Furthermore, the enhanced residual network model includes residual blocks, pooling blocks, attention-enhanced residual blocks, and a fully connected network, wherein the attention-enhanced residual block structure integrates convolutional layers for local information and attention-enhanced convolutional layers for global context.
[0048] Furthermore, during the training phase, the enhanced residual network model is trained by minimizing the L2 norm loss function between the predicted and true locations:
[0049]
[0050] Where ||·||2 is the L2 norm, Z n T is the true location of the nth sample. n It is in Z n The processed time-angle-delay channel frequency amplitude fingerprint matrix, Y n It's an estimated location.
[0051] Furthermore, the trained enhanced residual network model is used for location estimation, as follows:
[0052] (81) Extract shallow features I0 from the input fingerprint matrix T through a convolutional layer:
[0053] I0 = Conv(T) (8)
[0054] Where Conv(·) represents a convolutional layer;
[0055] (82) A series of cascaded residual blocks RB and pooling blocks PB are used to fuse local information in I0 and gradually reduce the size and dimension of the matrix;
[0056] The pooling block increases the number of channels through a convolutional layer, then reduces the spatial dimension through average pooling, and finally restores the original number of channels through another convolutional layer;
[0057] For Q cascaded residual blocks and pooling blocks, the output of the Qth pooling block is:
[0058] I Q =PB Q (RB Q (…PB1(RB1(I0))…)) (9)
[0059] In obtaining spatially sampled features I Q Then, the stacked residual blocks RB continue to extract additional local information;
[0060] To capture global context information, every e residual blocks RB are replaced with attention-enhanced residual blocks AARB; if e = 1, then I Q After using only attention-enhanced residual blocks (AARB), and with l blocks after the last pooling block (PB), the output feature map of the (Q+l)th block is:
[0061]
[0062] (83) Finally, the deep feature I after processing L blocks Q+L Flattened into vector F f :
[0063] F f =Flatten(I Q+L (11)
[0064] Vector F f The input is fed into a three-layer fully connected network FCN, and the fully connected network will... f Mapping to two-dimensional coordinates yields the predicted position Y:
[0065] Y = FCN(F f (12).
[0066] According to a second aspect of the present invention, the present invention provides a 5G indoor positioning system based on SSB signal spatiotemporal feature enhancement and attention mechanism, for implementing the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism described in the first aspect, comprising:
[0067] The data preprocessing module is used to receive channel state information data from the 5G New Radio interface and to design a three-level adaptive filter to preprocess the channel state information data.
[0068] The fingerprint construction module is used to construct a time-series-angle-delay channel frequency amplitude fingerprint matrix that integrates historical channel information based on preprocessed channel state information data.
[0069] The model training module is used to construct and train an attention-based augmented residual network model to obtain the trained augmented residual network model. The input of the augmented residual network model is a fingerprint matrix, and the output is position coordinates.
[0070] The online positioning module is used to input the pre-processed fingerprint data collected online into the trained enhanced residual network model for location estimation, thereby completing indoor fingerprint positioning.
[0071] According to a second aspect of the present invention, a terminal device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism described in the first aspect.
[0072] This invention has at least the following beneficial effects:
[0073] 1. By using a TSA module, this invention effectively mitigates signal degradation caused by multipath effects and improves the robustness of fingerprint features; the use of a time-weighted method in the fingerprint construction process reduces transient interference and improves positioning accuracy in dynamic environments.
[0074] 2. The meticulously designed preprocessing module (TSA) of this invention effectively improves the quality of the original CSI fingerprint; it proposes a T-ADCFA fingerprint matrix, which enhances robustness and adaptability through time-weighted fusion, solves the problem of transient interference, supports high-precision positioning in 5G-A / 6G systems, and also has the characteristics of low complexity.
[0075] 3. This invention can overcome the performance degradation problem caused by data distribution drift in traditional methods, and is more suitable for constantly changing indoor scenes in the real world.
[0076] 4. This invention enhances the dynamic expressiveness of features by fusing historical channel information, thereby improving both feature representation and construction efficiency; and improves the generalization ability of the model by dynamically weighting complex multipath signals using channel attention.
[0077] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0078] Figure 1 This is a flowchart illustrating the positioning method described in this invention;
[0079] Figure 2 This is an architecture diagram of the three-stage adaptive filter (TSA) in this invention;
[0080] Figure 3This is a roadmap for the CSI-based fingerprint matrix improvement in this invention;
[0081] Figure 4 This is an architecture diagram of the enhanced residual network model in this invention;
[0082] Figure 5 This is an architecture diagram of the attention-enhanced residual block in this invention;
[0083] Figure 6 This is the experimental platform of the present invention, wherein (a) is a physical diagram of the 5G downlink signal acquisition and communication module; and (b) is a 5G signal sampling test platform.
[0084] Figure 7 This is the display interface of the smart terminal APP of the present invention. (a) is for channel response parameter selection, and (b) is for channel response data analysis.
[0085] Figure 8 This is a schematic diagram of an office experimental scenario in this invention, wherein (a) is a real-life image of the office scenario; and (b) is a fingerprint collection modeling diagram of the office scenario.
[0086] Figure 9 This is a schematic diagram of the corridor experimental scenario in this invention, wherein (a) is a real-life image of the corridor scenario; and (b) is a fingerprint collection modeling diagram of the corridor scenario.
[0087] Figure 10 This is the amplitude Hampel anomaly detection map (subcarrier index 3) in the three-level adaptive filter (TSA) of this invention, where (a) is the unprocessed CSI sample; and (b) is the CSI sample after detection processing.
[0088] Figure 11 This is a schematic diagram of amplitude wavelet filtering in the three-stage adaptive filter (TSA) of this invention;
[0089] Figure 12 This invention relates to the outlier removal of the amplitude cross-correlation detector in the three-stage adaptive filter (TSA).
[0090] Figure 13 This is a phase diagram based on linear correction of the present invention, where (a) is the original phase; (b) is the phase unfolding; and (c) is the CSI sample after linear correction.
[0091] Figure 14 This is a comparison chart of the amplitude distribution of the five fingerprint feature matrices of this invention (office scene), where (a) is the amplitude distribution of the T-ADCFA matrix, (b) is the amplitude distribution of the ADCFP matrix, (c) is the amplitude distribution of the ADCFA matrix, (d) is the amplitude distribution of the ADCPM matrix, and (e) is the amplitude distribution of the ADCAM matrix.
[0092] Figure 15 The CDF values of various localization algorithms under different training rounds are shown, where (a) is 100 rounds, (b) is 200 rounds, (c) is 300 rounds, and (d) is 400 rounds.
[0093] Figure 16 These are the CDF diagrams of the AARES algorithm of this invention in different scenarios;
[0094] Figure 17 This is a comparison of the localization performance of AARES in different scenarios and training rounds.
[0095] Figure 18 These are CDFs for different matching algorithms, (a) in an office setting and (b) in a corridor setting.
[0096] Figure 19 This is the CDF of the five fingerprint feature matrices of this invention;
[0097] Figure 20 This is a comparison chart of positioning errors under different training rounds of the present invention;
[0098] Figure 21 This is a comparison of the positioning overhead of different fingerprint matrices in this invention, (a) is a comparison of time consumption, and (b) is a comparison of storage overhead.
[0099] Figure 22 This is a schematic diagram illustrating the localization performance of the AARES network under different training rounds of the present invention. Detailed Implementation
[0100] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0101] Please see Figures 1-22 This invention provides a technical solution: a 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism, comprising:
[0102] S1. Receive channel state information data from 5G New Radio and design a three-level adaptive filter to preprocess the channel state information data;
[0103] Raw CSI data is susceptible to multipath effects, noise interference, and clock asynchrony. Therefore, this embodiment designs a three-stage adaptive filter (TSA) preprocessing module, including an amplitude processing module and a phase processing module. The specific steps are as follows:
[0104] S11: Amplitude processing module (see...) Figure 2 )
[0105] Hampel identifier: An enhanced Hampel identifier is used to remove anomalous CSI samples. It detects coarse-grained outliers according to the following criteria:
[0106]
[0107] in μ represents the amplitude of the i-th sample. i It is the local median within the sliding window, σ i γ is the median absolute deviation (MAD), and γ is an adjustable threshold coefficient (default value γ = 3);
[0108] Wavelet filter: An improved wavelet threshold function is used to smooth each subcarrier sequence to reduce time jitter;
[0109] Cross-correlation detector: To minimize the impact of signal interference on system performance, the cross-correlation detector eliminates outliers in the CSI samples of the positioning area. For the nth CSI sample at a given location, the average correlation coefficient of all samples at that point is calculated as follows:
[0110]
[0111] Where M is the total number of CSI samples collected at a single node, affecting the sampling density of the positioning system; Cov(·) measures the linear correlation between samples, capturing the spatial features of CSI; σ(n) and σ(m) are A n and A m The standard deviation represents the amplitude dispersion; if (Predefined threshold) The system will classify abnormal sample A n Replace with the mean of the sample matrix, i.e.
[0112] S12: Phase processing module (see...) Figure 2 )
[0113] In practical communication systems, the sampling frequency offset (SFO) and carrier frequency offset (CFO) caused by transceiver clock asynchrony introduce a large number of errors into the original CSI phase measurement. The uncorrected phase exhibits nonlinear spatial distortion and lacks consistent waveform similarity between positioning points, which hinders the effective construction of fingerprint feature space. To solve this problem, the linear transformation (LT) module is applied to reduce the impact of SFO and CFO.
[0114] Phase expansion algorithm: The π phase expansion algorithm eliminates the periodic jumps in the original phase data and maps the values to a continuous, monotonically increasing domain;
[0115] Linear Transformation Module: A linear regression model based on the minimum mean square error (MMSE) criterion, used to compensate for errors in the expanded phase. The specific formula is as follows:
[0116]
[0117] in To correct the phase, φ(k) is the unpacking phase, d(k) represents the index of the k-th pilot subcarrier in the OFDM symbol, and K represents the total number of pilot subcarriers. This compensation isolates the hardware-induced deviation from the real spatial channel characteristics, thereby generating distinguishable phase waveform fingerprints at different locations for subsequent feature extraction.
[0118] Phase TSA module: Phase outlier handling, same as amplitude TSA method;
[0119] S2. Based on the preprocessed channel state information data, construct a time-series-angle-delay channel frequency amplitude fingerprint matrix that integrates historical channel information;
[0120] This embodiment extracts Channel State Information (CSI) data from a USRP B210 receiver that integrates a Spartan-6XC6SLX150 FPGA and an AD9361 RFIC. The process of estimating Channel State Information (CSI) from a 5G NR Synchronization Signal Block (SSB) includes three main steps: SSB detection, demodulation reference signal (DMRS) extraction, and channel estimation.
[0121] S21. Let the locally generated DMRS pilot sequence be X, and the received DMRS symbol be Y. By applying the minimum mean square error (MMSE) principle to channel estimation, the CSI is estimated as follows:
[0122]
[0123] S22. For example Figure 3 As shown, based on the ADCPM framework, this embodiment simplifies the meta-fingerprint matrix by eliminating unitary matrix decomposition and applies square root transformation to generate the ADCFA matrix, reducing the computational complexity to O(n). 2 logn);
[0124] First, the channel frequency response matrix (CFR)H k Transformed to the angular domain, this produces the angular time delay channel frequency response matrix (ADCFR)F. k ,Right now:
[0125]
[0126] The transformation matrix V is defined as follows:
[0127]
[0128] In the formula (·) H Represents the conjugate transpose of a matrix;
[0129] Considering F k The computational burden of complex operations used for positioning is reduced by the Hadamard inner product, which generates the angular delay frequency channel power (ADCFP) matrix P. k ,Right now:
[0130]
[0131] Subsequently, by taking P k The square root of the element is used to derive the angle-delay channel frequency amplitude (ADCFA) matrix C. k :
[0132] In dynamic environments, changes in NLOS caused by motion can reduce positioning accuracy. To mitigate transient interference, the time-weighted method employs an exponentially decaying weight function:
[0133]
[0134] Where t c It is the current timestamp, t n It is the timestamp of the nth historical fingerprint, λ (usually 0.1) controls the decay rate, and the current ADCFA matrix C k The time-angle-delay channel frequency amplitude fingerprint matrix (T-ADCFA) is formed by fusing with historical fingerprints. k :
[0135]
[0136] Where N is the sliding window size, storing the latest N ADCFA matrices, and normalization ensures that the sum of the weights is 1; to avoid redundant calculations, the recursive calculation formula is:
[0137]
[0138] Where T k ′ is the updated T-ADCFA matrix, w′ and C k ′ is the latest weight and ADCFA matrix;
[0139] S3. Construct and train an enhanced residual network model based on an attention mechanism to obtain the trained enhanced residual network model, wherein the input of the enhanced residual network model is a fingerprint matrix and the output is position coordinates;
[0140] This embodiment proposes a segment localization system that uses an attention-enhanced ResNet (AARES) model for fingerprint localization, such as... Figure 4 As shown, the network consists of four main parts: residual blocks (RBs), pooling blocks (PBs), attention-enhanced residual blocks (AARBs), and fully connected networks (FCNs).
[0141] During the training phase, the augmented residual network model is trained by minimizing the L2 norm loss function between the predicted and true locations.
[0142]
[0143] Where ||·||2 is the L2 norm, Z n T is the true location of the nth sample. n It is in Z n The processed time-angle-delay channel frequency amplitude fingerprint matrix, Y n It is an estimated location;
[0144] S4. Input the preprocessed fingerprint data collected online into the trained augmented residual network model to estimate the location and complete indoor fingerprint localization;
[0145] The input to the enhanced residual network model is the fingerprint matrix T obtained after TSA data preprocessing and T-ADCFA feature extraction, and the output is the estimated location Y. The processing flow is as follows:
[0146] S41. First, shallow features I0 are extracted from the input T through a convolutional layer:
[0147] I0 = Conv(T) (26)
[0148] Where Conv(·) represents a convolutional layer;
[0149] S42. Next, to handle the complexity of CSI in real-world environments, a series of cascaded RB and PB blocks are used to fuse local information in I0, and the size and dimension of the matrix are gradually reduced. The PB block increases the number of channels through convolutional layers, then reduces the spatial dimension through average pooling (AveP), and finally restores the original number of channels through another convolutional layer; for Q cascaded RB and PB blocks, the output of the Qth PB is:
[0150] I Q =PB Q (RB Q (…PB1(RB1(I0))…)) (27)
[0151] In obtaining spatially sampled features I QThen, the stacked RB blocks continue to extract additional local information;
[0152] To capture global context information, every e RB blocks are replaced with AARB blocks. If e = 1, then I Q If only AARB is used after the last PB, and there are l blocks after the last PB, then the output feature map of the (Q+l)th block is:
[0153]
[0154] AARB structure (such as) Figure 5 (As shown) It integrates convolutional layers for local information and attention-enhanced convolutional layers (AAConv) for global context. AAConv considers all input pixels and dynamically generates weights for each output pixel, thereby expanding the receptive field to the entire input and enhancing global feature extraction.
[0155] S43. Finally, the deep feature I after processing L blocks. Q+L Flattened into vector F f :
[0156] F f =Flatten(I Q+L (29)
[0157] Vector F f The input is fed into a three-layer fully connected network (FCN), which will... f Mapping to two-dimensional coordinates yields the predicted position Y:
[0158] Y = FCN(F f (30).
[0159] The technical solution of the present invention will be further described below with reference to specific embodiments:
[0160] This embodiment collected fingerprint data in a typical office building, such as... Figure 6 As shown in (a), a software-defined radio-based test platform was established for signal sampling, recording, and CSI acquisition. The platform uses USRP, specifically the USRPB210 receiver, which integrates a Spartan-6XC6SLX150 FPGA and an AD9361 RFIC. This configuration allows for the reception of signals with carrier frequencies from 70MHz to 6GHz and bandwidths up to 56MHz.
[0161] like Figure 6As shown in Figure (b), the 5G signal sampling test bench adopts a modular integrated design. Its core components are housed in a shielded aluminum alloy chassis (dimensions: 300×200×20mm). An omnidirectional broadband antenna (frequency range: 2400~2700MHz, gain: 5dBi) is mounted on the top for receiving RF signals. To ensure long-term stable operation, an active cooling system is integrated into the chassis base. This system includes two 8025 ball bearing fans (adjustable speed: 1500~3000RPM) and a copper-aluminum composite heat sink, maintaining the operating temperature within an industrial-grade range of -10℃ to +55℃. The communication interface supports dual-mode connectivity: a gigabit Ethernet port (compliant with IEEE 802.3ab) for wired data transmission and a dual-band wireless router (2.4 / 5GHz, 802.11ac) for wireless connectivity. In this experiment, the wireless mode is used to mitigate multipath interference introduced by physical cabling.
[0162] The collected 5G signals are processed through SSB detection, DMRS extraction, and channel estimation to obtain the CSI. The gNB parameters are configured according to the 5G NR protocol TS 38.104, and signals from specific operators are selected for analysis. In order to cooperate with the commercial deployment of specific operators in the n41 band (2.5GHz ISM band), the receiver is configured with a center frequency of 2565MHz (center frequency point: 504990), a subcarrier spacing of 30kHz, and a receiving bandwidth of 10MHz to capture the main lobe energy of the SSB.
[0163] like Figure 7 (a) Channel response is performed using signals in frequency band 41. The signal processing link consists of three stages: synchronization signal block detection, demodulation reference signal (DMRS) extraction, and MMSE channel estimation. CSI acquisition is completed under NLOS propagation conditions, strictly following the 3GPP Case C networking specification. Figure 7 (b) Displays the SSB signal acquisition parameters for frame number 262. The top of the channel response interface of the smart terminal APP displays information such as the frequency band, PCI, gain and delay of the acquired signal, the middle displays the CSI information of the signal, and the bottom can set the data tag and acquisition time for convenient subsequent processing.
[0164] The test environment consisted of a single 5G NR gNB operated by a specific carrier. The outdoor base station was deployed inside an office building. Field tests were conducted in two different indoor scenarios, as detailed below. Figure 8 , Figure 9 As shown:
[0165] Office Scene: This scene takes place on the fourth floor and includes an 8x12 meter rectangular office area furnished with numerous chairs, desks, and a lectern. Figure 8 (a)); such as Figure 8As shown in (b), a grid of 42 sampling points is predefined, consisting of 28 training points (red) and 14 test points (green). The training point in the upper left corner is designated as the origin (1,1).
[0166] Corridor Scene: Located on the 3rd floor, this is a 5x14 semi-open corridor, partially obstructed by a central wall. The area has multiple entrances and experiences regular pedestrian traffic, making it a more complex environment than the office. Figure 9 (a) A total of 92 points were established, of which 72 were used for training (red) and 20 were used for testing (green). Figure 9 (b)); The spacing between adjacent training points is maintained at 1.0m. The perpendicular distance between each test point and the line connecting the two nearest training points is 0.5m. The training point in the upper left corner is set to (1,1), and the positions of all other points are derived from this geometric layout.
[0167] In both office and corridor scenarios, test points were intentionally distributed between training (reference) points and did not overlap with them. At each acquisition point, CSI samples were continuously acquired for 30 seconds. A set of 20 stable samples selected from the middle of the acquisition period was retained at each location. All training and test data were collected on the same day to ensure channel consistency.
[0168] Four metrics are used to evaluate performance: mean absolute error (MAE), root mean square error (RMSE), standard deviation of error (STDE), and maximum error (MAXE).
[0169] MAE is defined as follows:
[0170]
[0171] RMSE is defined as:
[0172]
[0173] Furthermore, the cumulative distribution function (CDF) was used to perform statistical analysis on CSI data and positioning errors. The CDF describes the probability distribution of random variables and is an important tool for evaluating the effectiveness of positioning systems. By calculating the 1-σ (68.27%) and 2-σ (95.45%) error intervals, the degree of positioning error coverage was quantified, providing valuable insights into error distribution characteristics and system enhancement opportunities.
[0174] like Figures 10-22 As shown, the specific experimental results are as follows:
[0175] 1. TSA module effectiveness: via Hampel filter ( Figure 10Effectively detects and corrects amplitude anomalies, reducing local amplitude variance by 66.8% and overall dataset amplitude noise by 57.14% (standard deviation decreased from 4.27 dB to 1.83 dB); time-frequency joint wavelet filtering ( Figure 11 Suppressing distortion caused by frequency-selective fading, reducing the peak standard deviation of the subcarrier dimension by 60%–71.4%, and achieving an average noise reduction of 50% across the entire frequency band; based on cross-correlation anomaly detection ( Figure 12 (Accuracy 92.3%) and mean correction improved the CSI matrix cosine similarity by 19.7%, reducing the mean square error of positioning by 37.5%; phase processing ( Figure 13 Phase jumps and SFO / CFO effects are eliminated through π phase expansion and linear compensation models, enhancing phase continuity and anomaly distinguishability; Figure 18 The comparison of error CDF curves shows that the TSA module significantly improves the accuracy and robustness of various positioning algorithms: the average error decay rate across algorithms reaches 12.3±3.8%, with TSA+CNN achieving a cumulative probability of 60.4% at the 3-meter error threshold (an improvement of 10.5 percentage points over the baseline), TSA+SVR achieving an 11.7% improvement at the 4-meter threshold, and TSA+WKNN reducing the 1-σ error boundary by 8.0%. Even in dynamic corridor scenarios (with severe multipath interference), an average error decay of 14.5±2.8% is still achieved. These results verify that TSA effectively enhances the performance of different positioning architectures by detecting outliers and suppressing phase jumps / carrier frequency offset disturbances. In summary, the TSA module effectively optimizes amplitude stability, frequency offset resistance, correlation robustness, and phase reliability.
[0176] 2. Validity of the T-ADCFA fingerprint matrix: Figure 14 The channel amplitude distribution shows that, compared to the ADCFP and ADCPM matrices (weak path amplitude differential is finite, total amplitude is four orders of magnitude higher) and the ADCAM matrix (angular domain differential difference), the T-ADCFA matrix exhibits the clearest amplitude difference between weak and strong paths, while recording the minimum channel amplitude. This characteristic significantly improves the calculation weight of weak paths, increases the overall processing speed, and demonstrates stronger robustness to transient interference in dynamic environments. ADCPM and ADCAM, due to their complex calculations and insufficient performance, have lower overall practicality than the T-ADCFA and ADCFA matrices. Error distribution ( Figure 19 The accuracy ranking is T-ADCFA > ADCFA > ADCFP > ADCAM > ADCPM, with a 3m error threshold CDF value of 74.7% (significantly higher than ADCPM's 33.6%). Training analysis ( Figure 20Further analysis shows that all fingerprint matrices overfit after reaching their peak at 200 rounds, but T-ADCFA consistently maintains its lead (with a difference of over 1.63m from the worst fingerprint matrix), and also possesses optimal stability (standard deviation 0.1097m) and convergence speed. Cost analysis ( Figure 21 The results show that the T-ADCFA matrix achieves the optimal balance between efficiency and environmental adaptability. In corridor scenarios, the increased reference points (115% more) lead to a surge in resource consumption (average generation time +107%). While ADCFP is the fastest in office scenarios (1.11s), its corridor generation time increases by 113%. ADCAM / ADCPM, on the other hand, has the highest computational cost due to its complex decomposition (over 3.4s in corridors). In contrast, T-ADCFA's office generation time (1.27s) is close to optimal, with only a 96% increase in corridor time, while maintaining the lowest storage overhead. Therefore, T-ADCFA is the most efficient and reliable solution for accurate office positioning.
[0177] 3. Performance advantages of AARES localization: The network architecture adopts spatial downsampling (depth features reduced by 1 / 16) and four layers of AAConv (5×5 convolutional kernels, 32 feature channels), trained with ReLU activation and Adam optimizer. Figure 15 Experimental results show that the model exhibits a steeper upward trend in the cumulative error distribution (CDF) (with a higher probability of small error intervals of 1-2m), consistently outperforming HiLoc, CNN, SVR, and WKNN. Quantitative analysis reveals that its MAE after 400 epochs is as low as 1.5442m, a 46.2% reduction in error compared to WKNN, and it consistently leads at different epochs (e.g., exceeding HiLoc by 7.47% after 100 epochs). The error gap stabilizes at 0.32m during training, with the attention mechanism effectively suppressing overfitting. After 300 epochs, it enters the superlinear optimization stage (MAE further decreases by 3.9%), significantly outperforming the linear decay of traditional methods. Figure 16 , 17 The experimental results show that the office scenario has higher accuracy due to its simpler layout (MAE = 1.5442m after 400 rounds), while the corridor scenario consistently has a higher MAE of 18.6% ± 2.3% due to repetitive features, and the optimization speed slows down significantly with each round (16.7% improvement in the first 200 rounds, only 6.3% improvement in the 300-400 rounds); from Figure 22 In terms of positioning performance and positioning time, increasing the training rounds from 100 to 500 reduced the average positioning error from 1.8784m to 1.4845m. However, the inference time increased from 1.2546s to 1.5139s. The improved accuracy reflects a more thorough feature learning process, while the extended duration indicates increased computational complexity during inference. For practical deployment, this requires careful calibration of the training time to balance the competitive requirements of high accuracy and real-time performance in specific application scenarios.
[0178] In summary, the TSA method of this invention can effectively suppress signal distortion caused by multipath effects, the new spatiotemporal fingerprint matrix T-ADCFA significantly reduces computational complexity while enhancing feature discriminability, and the ResNet-based localization algorithm effectively mines deep features and achieves higher accuracy. Therefore, this invention provides a robust, efficient and high-precision solution for complex indoor localization challenges.
[0179] Example 2:
[0180] This embodiment provides a 5G indoor positioning system based on SSB signal spatiotemporal feature enhancement and attention mechanism, used to implement the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism described in Embodiment 1, including:
[0181] The data preprocessing module is used to receive channel state information data from the 5G New Radio interface and to design a three-level adaptive filter to preprocess the channel state information data.
[0182] The fingerprint construction module is used to construct a time-series-angle-delay channel frequency amplitude fingerprint matrix that integrates historical channel information based on preprocessed channel state information data.
[0183] The model training module is used to construct and train an attention-based augmented residual network model to obtain the trained augmented residual network model. The input of the augmented residual network model is a fingerprint matrix, and the output is position coordinates.
[0184] The online positioning module is used to input the pre-processed fingerprint data collected online into the trained enhanced residual network model for location estimation, thereby completing indoor fingerprint positioning.
[0185] Example 3:
[0186] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor loads and executes the computer program, it adopts the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism described in Embodiment 1.
[0187] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server, and the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include input / output devices, network access devices, and buses.
[0188] Furthermore, the processor can be a central processing unit (CPU). Of course, depending on the actual use, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be used. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it in this regard.
[0189] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0190] For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances. When an element is referred to as being "assembled on," "mounted on," "fixed to," or "set on" another element, it may be directly on the other element or there may be an intermediate element present. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be an intermediate element present. The terms "vertical," "horizontal," "upper," "lower," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible embodiments.
[0191] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0192] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
Claims
1. A 5G indoor positioning method based on SSB signal space-time feature enhancement and attention mechanism, characterized in that, Includes the following steps: The channel state information data of 5G New Radio is received, and a three-level adaptive filter is designed to preprocess the channel state information data. The three-stage adaptive filter includes an amplitude processing module and a phase processing module. The amplitude processing module performs data preprocessing, as follows: Use the Hampel identifier to remove outliers from the channel state information subcarrier amplitude sequence; Smooth each subcarrier sequence using a wavelet filter based on an improved wavelet threshold function; Outliers in the channel state information data are eliminated using a cross-correlation detector, as follows: For a given position, the first The average correlation coefficient of all samples at a given CSI point is calculated as follows: in, It is the total number of CSI samples collected at a single node, which affects the sampling density of the positioning system; Measure the linear correlation between samples to capture CSI spatial features; and yes and The standard deviation represents the amplitude dispersion. The system will identify abnormal samples. Replace with the mean of the sample matrix, i.e. ; The phase processing module performs data preprocessing, as follows: Use the phase unrolling algorithm to eliminate periodic jumps in the original phase data; A linear regression model based on the minimum mean square error criterion is used to compensate for errors in the expanded phase. The specific formula is as follows: in To correct the phase, It is the unpacking phase. Indicating the OFDM symbol, the first Index of each pilot subcarrier Indicates the total number of pilot subcarriers; Phase outlier processing is performed using the phase TSA module, which is the same as the amplitude TSA method. Based on the preprocessed channel state information data, a time-series-angle-delay channel frequency amplitude fingerprint matrix integrating historical channel information is constructed, as follows: Record No. The angle delay channel frequency response matrix from the user to the base station for: in, It is the phase-shifted DFT matrix, expressed as: In the formula It is the conjugate transpose. It is the channel frequency response matrix; because The computational burden of complex operations used for positioning in the matrix, the angular delay generated by the Hadamard inner product, and the frequency channel power. ,Right now: in, This indicates the calculation of mathematical expectation. Indicates the Hadama inner product; Afterwards, by taking The square root of the element is used to derive the angle-delay channel frequency amplitude matrix. : ; An exponentially decaying weighting function is used to perform time-weighted fusion of the current and historical angle-delay channel frequency amplitude matrices to form a time-series-angle-delay channel frequency amplitude fingerprint matrix. In the formula It is the current timestamp. It is the first A historical fingerprint timestamp To control the attenuation rate, it is set to 0.1; Current angle-delay channel frequency amplitude matrix. This is fused with historical fingerprints to form a time-series-angle-delay channel frequency-amplitude fingerprint matrix. : in It's a sliding window size, storing the latest data. The angle-delay channel frequency amplitude matrix is normalized to ensure that the sum of the weights is... ; To avoid redundant calculations, the recursive calculation formula is as follows: in It is an updated time-angle-delay channel frequency amplitude fingerprint matrix. and It is the latest weight and angle-delay channel frequency amplitude matrix; An attention-based augmented residual network model is constructed and trained to obtain the trained augmented residual network model. The input of the augmented residual network model is the fingerprint matrix, and the output is the position coordinates. The preprocessed fingerprint data collected online is input into the trained enhanced residual network model for location estimation, thus completing indoor fingerprint localization.
2. The 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism according to claim 1, characterized in that: The enhanced residual network model includes residual blocks, pooling blocks, attention-enhanced residual blocks, and a fully connected network, wherein the attention-enhanced residual block structure integrates convolutional layers for local information and attention-enhanced convolutional layers for global context.
3. The 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism according to claim 2, characterized in that: During the training phase, the enhanced residual network model minimizes the difference between the predicted and actual locations. Training with norm loss function: in yes Norm, It is the first The true location of each sample Is The processed time-angle-delay channel frequency amplitude fingerprint matrix, It's an estimated location.
4. The 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism according to claim 3, characterized in that: The trained enhanced residual network model is used for location estimation, as detailed below: (41) The fingerprint matrix is processed through a convolutional layer. Extracting shallow features : in Indicates a convolutional layer; (42) A series of cascaded residual blocks RB and pooling blocks PB are used for fusion. Local information within the matrix, and gradually reduce the size and dimension of the matrix; The pooling block increases the number of channels through a convolutional layer, then reduces the spatial dimension through average pooling, and finally restores the original number of channels through another convolutional layer; for Each cascaded residual block and pooling block, the first... The output of each pooling block is: Obtaining spatially sampled features Then, the stacked residual blocks RB continue to extract additional local information; In order to capture global context information, each Each residual block RB is replaced with an attention-enhanced residual block AARB; if ,but Then only attention-enhanced residual blocks (AARB) are used, followed by the last pooling block (PB). The ( ) block, then the ( ) The output feature maps of each block are: (43) Finally, after Deep features after block processing Flattened into a vector : vector The input is fed into a three-layer fully connected network (FCN), and the fully connected network will... Mapping to two-dimensional coordinates to obtain the predicted position. : 。 5. A 5G indoor positioning system based on SSB signal spatiotemporal feature enhancement and attention mechanism, used to implement the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism as described in any one of claims 1 to 4, characterized in that, include: The data preprocessing module is used to receive channel state information data from the 5G New Radio interface and to design a three-level adaptive filter to preprocess the channel state information data. The fingerprint construction module is used to construct a time-series-angle-delay channel frequency amplitude fingerprint matrix that integrates historical channel information based on preprocessed channel state information data. The model training module is used to construct and train an attention-based augmented residual network model to obtain the trained augmented residual network model. The input of the augmented residual network model is a fingerprint matrix, and the output is position coordinates. The online positioning module is used to input the pre-processed fingerprint data collected online into the trained enhanced residual network model for location estimation, thereby completing indoor fingerprint positioning.
6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the 5G indoor positioning method based on SSB signal spatiotemporal feature enhancement and attention mechanism as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device for carrying out indoor positioning by adopting multi-head attention mechanism neural network
CN117395598A
Vehicle positioning method and device based on fingerprint positioning, and electronic equipment
CN119767406A