Indoor positioning method and system

By combining lightweight neural networks and the MUSIC algorithm, the problems of multipath interference and phase jitter in indoor positioning are solved, achieving high-precision and stable indoor positioning, especially in complex environments such as airports.

CN120928284APending Publication Date: 2025-11-11CHINA ELECTRONICS TECH GRP NO 7 RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510962882.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing indoor positioning technologies face intractable technical challenges in dealing with multipath interference, phase jitter, and the lack of multi-antenna fusion, resulting in low and unstable positioning accuracy.

Method used

The IQ signals acquired by multiple antennas are enhanced by a lightweight neural network. The azimuth and elevation angles are calculated using the MUSIC algorithm. Multi-channel feature fusion is performed by combining the covariance matrix and attention mechanism to suppress multipath interference and eliminate phase jitter.

Benefits of technology

It significantly improves the accuracy and stability of indoor positioning, achieves centimeter-level real-time positioning capability, and enhances positioning reliability in dynamic and complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928284A_ABST
    Figure CN120928284A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor positioning method and system, and the method comprises the following steps: carrying out the enhancement processing of original IQ in-phase orthogonal signals collected by multiple antennas through a lightweight neural network, and outputting an anti-multipath interference enhanced IQ signal matrix; and constructing a covariance matrix based on the enhanced signal matrix, and calculating an azimuth angle and a pitch angle by using a MUSIC multiple signal classification algorithm. According to the invention, multipath interference can be effectively suppressed, phase jitter is eliminated, multichannel feature fusion is realized, and the accuracy of indoor positioning is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of indoor positioning technology, and specifically to an indoor positioning method and system. Background Technology

[0002] In modern airport operations and management, accurate and real-time indoor positioning technology is crucial for the efficient scheduling and management of various resources within the airport. Airports contain not only large-scale baggage handling equipment, cargo fleets, and boarding bridges—all non-powered—but also mobile terminal devices used for passenger guidance and security. Traditional positioning methods (such as GPS) struggle to meet the demands for high-precision, large-area real-time positioning in indoor environments due to signal obstruction and attenuation.

[0003] Current mainstream indoor positioning methods, such as fingerprint matching schemes based on Received Signal Strength Indicator (RSSI), have low implementation costs, but their positioning accuracy is highly susceptible to dynamic environmental changes. For example, changes in personnel movement or the location of goods can cause the pre-built fingerprint database to become invalid, requiring frequent updates and maintenance, and making it difficult to meet the requirements for high-precision real-time positioning. While schemes based on Time of Arrival (TOA) or Time Difference of Arrival (TDOA) can improve accuracy, they rely on stringent time synchronization conditions. In environments with significant signal multipath effects, the time delay estimation error increases dramatically, and the equipment costs are high.

[0004] While Bluetooth positioning technology based on Angle of Arrival (AoA) has demonstrated advantages in multipath resistance in recent years, it still suffers from the following drawbacks in practical applications: First, indoor multipath reflections lead to high signal coherence, causing severe distortion of the pseudospectral spectrum in the core MUSIC algorithm, with frequent false peaks, significantly weakening angular resolution. Second, the hardware-acquired I / Q signals are susceptible to noise, clock jitter, and channel offset, resulting in unstable I / Q phase data that is easily interfered with, further increasing the fluctuation in angle estimation. Third, traditional algorithms process data from each antenna channel independently, ignoring the spatiotemporal correlation of multiple antennas and failing to fully exploit the joint information gain of the array signals. These shortcomings collectively limit the practicality and reliability of indoor positioning systems in dynamic and complex scenarios.

[0005] Therefore, there is an urgent need for an indoor positioning method that can effectively suppress multipath interference, eliminate phase jitter, and achieve multi-channel feature fusion. Summary of the Invention

[0006] The main objective of this invention is to provide an indoor positioning method and system that aims to solve the technical problems of multipath interference, phase jitter, and lack of multi-antenna fusion in the prior art.

[0007] To achieve the above objectives, the first aspect of the present invention provides an indoor positioning method, comprising the following steps:

[0008] Step S100: Enhance the original IQ in-phase orthogonal signals acquired by multiple antennas using a lightweight neural network to output an enhanced IQ signal matrix that resists multipath interference;

[0009] Step S200: Construct a covariance matrix based on the enhanced signal matrix, and calculate the azimuth and elevation angles using the MUSIC multi-signal classification algorithm.

[0010] Preferably, in step S100, the step of enhancing the original IQ in-phase quadrature signals acquired by multiple antennas using a lightweight neural network includes:

[0011] Step S110: Normalize the raw IQ signals acquired by the multiple antennas and output the normalized IQ matrix;

[0012] Step S120: Input the normalized IQ matrix into a lightweight two-dimensional convolutional network to extract spatial features between multiple antennas and output a spatial feature map;

[0013] Step S130: Input the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporal enhanced features;

[0014] Step S140: Input the temporal enhancement features into the attention mechanism layer, perform multi-channel feature weighted fusion, and output the weighted fused features;

[0015] Step S150: Based on the weighted fusion features, perform linear mapping and combination to reconstruct the output enhanced IQ signal matrix.

[0016] Preferably, in step S110, the step of normalizing and preprocessing the raw IQ signals acquired by multiple antennas and outputting the normalized IQ matrix includes:

[0017] Step S111: Construct the original IQ data matrix based on the original IQ signals acquired by multiple antennas.

[0018] In the formula, N represents the total number of antennas, M represents the number of sampling points per antenna, and X... n,m Represents the complex IQ value of the m-th sampling point of the n-th antenna (n = 1, 2, ..., N; m = 1, 2, ..., M);

[0019] Step S112: Based on the original IQ data matrix X, calculate the average signal amplitude of each antenna n. The calculation formula is as follows:

[0020]

[0021] In the formula, μ n This represents the average signal amplitude across M sampling points for each antenna n;

[0022] Step S113: For each element X in the complex matrix X n,m Normalization is performed to obtain the normalized IQ matrix. The calculation formula is:

[0023]

[0024] In the formula, This represents the normalized complex signal value of the nth antenna at the mth sampling point.

[0025] Preferably, in step S120, the step of inputting the normalized IQ matrix into a lightweight two-dimensional convolutional network to extract spatial features between multiple antennas and output a spatial feature map includes:

[0026] Step S121: Take the normalized IQ matrix as input, perform spatial dimension sliding convolution using a convolution kernel of size 3×3 and stride of 1, and output the convolution result matrix;

[0027] Step S122: Apply the ReLU activation function to the convolution result matrix to output the spatial feature map. The calculation formula is as follows:

[0028]

[0029] In the formula, * denotes the convolution operation, W1 is the convolution kernel parameter, and b1 is the bias term. Let F1 ∈ R be the normalized IQ matrix, ReLU(·) be the ReLU activation function, and F1 ∈ R. N×M×C The three-dimensional tensor representing the spatial feature map is denoted by C, where C is the number of channels, N is the total number of antennas, and M is the number of sampling points per antenna.

[0030] Preferably, in step S130, the step of inputting the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporally enhanced features includes:

[0031] Step S131: Apply one-dimensional dilated convolution to the temporal feature vector corresponding to each antenna n in the spatial feature map. The calculation formula is as follows:

[0032]

[0033] In the formula, This represents the output feature vector of the nth antenna in the lth layer, where l represents the network layer number, and d l The expansion rate of the l-th layer is d. l =2 l-1 ReLU(·) is the ReLU activation function, and the kernel size is k. Represents the timing characteristics of the nth antenna, Conv1Dcausal For causal convolution operations, left padding is used to ensure that the output length is consistent with the input.

[0034] Step S132: Perform a residual connection operation on the dilated convolution output. After stacking L layers of processing sequentially, the output is a temporal augmentation feature.

[0035] In the formula, F2∈R N×M×C This represents the output of the timing enhancement feature, where N is the total number of antennas, M is the number of sampling points, and C is the number of feature channels. This represents the final feature vector of the nth antenna after processing through L layers.

[0036] Preferably, in step S140, the step of inputting the temporal enhancement features into the attention mechanism layer, performing multi-channel feature weighted fusion, and outputting weighted fused features includes:

[0037] Step S141: Map the temporal enhancement feature F2 to the query matrix Q, key matrix K, and value matrix V respectively, using the following formula:

[0038] Q = F2W Q

[0039] K = F2W K

[0040] V = F2W V

[0041] In the formula, W Q To query the trainable weight matrix of matrix Q, W K W is the trainable weight matrix of the key matrix K. V Let V be the trainable weight matrix of the value matrix V;

[0042] Step S142: Based on the query matrix Q and the key matrix K, calculate the attention weight matrix using the following formula:

[0043]

[0044] In the formula, A represents the attention weight matrix, and d k Let Q represent the key matrix dimension, Q represent the query matrix, and K represent the key matrix dimension. T represents the transpose of the key matrix, and softmax(·) represents the normalized exponential function;

[0045] Step S143: Perform weighted fusion of the value matrix V using the attention weight matrix A, and output the weighted fusion feature:

[0046] F3 = AV

[0047] In the formula, F3 represents the weighted fusion feature.

[0048] Preferably, in step S150, the step of reconstructing the output enhanced IQ signal matrix by performing linear mapping and combination based on the weighted fusion features includes:

[0049] Step S151: Perform a two-way linear mapping on the last channel dimension of the weighted fusion feature to generate the real part matrix and the imaginary part matrix. The calculation formula is as follows:

[0050]

[0051] In the formula, f3(m,n) represents the feature vector of the m-th antenna and the n-th sampling point in the F3 indoor positioning system, I m,n w represents the real part signal value of the m-th antenna at the n-th sampling point after reconstruction. r Let b be the weight vector mapped to the real part. r The bias term for the real part mapping; Q m,n w represents the imaginary part signal value of the m-th antenna at the n-th sampling point after reconstruction. i Let b be the weight vector of the imaginary part mapping. i The bias term for the mapping of the imaginary part; w r ,w i and b r ,b i These are trainable parameters;

[0052] Step S152: Combine the real part matrix and the imaginary part matrix to generate the enhanced IQ signal matrix, calculated using the following formula:

[0053]

[0054] In the formula, Let j represent the complex IQ signal enhanced by the m-th antenna at the n-th sampling point, where j represents the imaginary unit. This represents the final output enhanced IQ signal matrix.

[0055] Preferably, in step S200, the step of constructing a covariance matrix based on the enhanced signal matrix and calculating the azimuth and elevation angles using the MUSIC multiple signal classification algorithm includes:

[0056] Step S210: Divide the time snapshot samples based on the enhanced IQ signal matrix, and calculate the output signal covariance matrix by autocorrelation averaging;

[0057] Step S220: Perform eigenvalue decomposition on the signal covariance matrix, separate and output the noise subspace matrix corresponding to the smallest eigenvalue;

[0058] Step S230: Based on the noise subspace and planar array model, calculate the two-dimensional pseudospectral function that characterizes the signal intensity at different incident angles;

[0059] Step S240: Based on the two-dimensional pseudospectral function, perform a global peak search in the azimuth and elevation parameter space, and output the estimated values ​​of the target's azimuth and elevation angles.

[0060] A second aspect of the present invention provides an indoor positioning system, comprising:

[0061] The signal enhancement module enhances the original IQ in-phase orthogonal signals acquired by multiple antennas through a lightweight neural network, and outputs an enhanced IQ signal matrix that resists multipath interference.

[0062] The angle calculation module constructs a covariance matrix based on the enhanced signal matrix and uses the MUSIC multi-signal classification algorithm to calculate the azimuth and elevation angles.

[0063] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the indoor positioning method as described in the first aspect.

[0064] A fourth aspect of the present invention provides a computer-readable storage medium storing an indoor positioning processing program, which, when executed by a processor, implements the steps of the indoor positioning method as described in the first aspect.

[0065] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0066] This invention provides an indoor positioning method and system. The indoor positioning method uses a lightweight neural network to jointly enhance the original IQ signals from multiple antennas. During the signal input stage, it deeply suppresses multipath reflection interference and hardware phase jitter, generating a highly consistent enhanced IQ signal matrix. Then, it uses the MUSIC algorithm to construct a highly robust spatial spectrum model based on the enhanced IQ signal matrix. By accurately eliminating spurious direction-of-arrival components through the noise subspace, it ultimately solves the technical problems of pseudo-spectral distortion caused by multipath interference in angle calculation in complex indoor scenarios, covariance inaccuracy caused by phase jitter, and the weakening of spatiotemporal correlation due to the isolation of multi-antenna information. This significantly improves the estimation accuracy and dynamic stability of azimuth and elevation angles, providing reliable positioning capabilities for indoor positioning scenarios.

[0067] Furthermore, the indoor positioning method of this invention also eliminates hardware gain differences through normalization preprocessing, improving the phase comparability of multi-channel signals; extracts spatial correlation features through a two-dimensional convolutional network, strengthening the spatial consistency expression of array signals; models long-term temporal dependencies through a temporal convolutional network, enhancing the capture of signal continuity and periodic patterns; dynamically weights multi-channel features through an attention mechanism, increasing the weight of effective components in interference environments; ensures the integrity of signal waveform structure and phase fidelity through dual-path linear mapping and complex reconstruction; suppresses noise in snapshot samples and improves the robustness of spatial correlation measurement through statistical averaging optimization of the covariance matrix; accurately suppresses pseudo-spectral peak interference caused by multipath through orthogonal projection of the noise subspace; and ensures the global optimality of angle estimation in dynamic scenes through a global spectral peak search strategy. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort. The realization of the purpose, functional features, and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings.

[0069] Figure 1 A flowchart of an indoor positioning method provided in an embodiment of the present invention;

[0070] Figure 2 A flowchart of a method for enhancing an IQ signal matrix using a lightweight neural network, provided in an embodiment of the present invention;

[0071] Figure 3 This is a schematic diagram of the network structure of the L-MDA lightweight neural network provided in an embodiment of the present invention;

[0072] Figure 4 A flowchart of a method for normalizing raw IQ signals provided in an embodiment of the present invention;

[0073] Figure 5 A flowchart of a method for extracting spatial features from a normalized IQ matrix provided in an embodiment of the present invention;

[0074] Figure 6 This is a flowchart of a method for enhancing temporal features based on spatial feature maps, provided in an embodiment of the present invention.

[0075] Figure 7 This is a flowchart of a method for multi-channel feature weighted fusion using an attention mechanism provided in an embodiment of the present invention.

[0076] Figure 8A flowchart of a method for reconstructing an enhanced IQ signal matrix based on enhanced fusion features, provided in an embodiment of the present invention;

[0077] Figure 9 This is a flowchart of a method for calculating azimuth and elevation angles based on an enhanced signal matrix, provided in an embodiment of the present invention.

[0078] Figure 10 A schematic diagram of an indoor positioning system provided in an embodiment of the present invention;

[0079] Figure 11 This is a schematic diagram of an indoor positioning computer device provided in an embodiment of the present invention. Detailed Implementation

[0080] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0081] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0082] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the meaning of "and / or" throughout the text includes three parallel solutions; for example, "A and / or B" includes solution A, solution B, or a solution where both A and B are satisfied simultaneously. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0083] Definitions:

[0084] AoA: Angle of Arrival, refers to the physical direction angle at which an electromagnetic wave or signal arrives at a receiving antenna array, and is a core parameter in wireless positioning systems. It calculates the spatial orientation of the signal source relative to the receiving array by measuring the phase difference or time difference of the signal across different antenna elements, including the azimuth (horizontal azimuth) and elevation (pitch) angle. In indoor positioning scenarios, AoA technology utilizes multi-antenna collaboration to capture the signal's direction of arrival, achieving target location estimation. It offers advantages such as tag-free operation and low latency, but is also susceptible to multipath interference and noise.

[0085] IQ signal: In-Phase and Quadrature (IQ) signal. An IQ signal is a complex signal representing the physical characteristics of electromagnetic waves, composed of mutually orthogonal in-phase (I) and quadrature (Q) components. The in-phase component directly reflects the amplitude of the original carrier wave, while the quadrature component is a copy of the carrier wave with a 90° phase shift. Together, they provide a complete description of the signal's amplitude, frequency, and phase information. In wireless communication systems, IQ signals are down-converted to convert radio frequency signals to baseband processing, providing underlying data support for positioning technologies such as direction-of-arrival (DOA) estimation. Their phase stability directly affects positioning accuracy.

[0086] Bluetooth AoA technology: An indoor positioning technology based on Bluetooth 5.1, typically consisting of a fixed Bluetooth base station (Locator) equipped with a multi-antenna array and mobile Bluetooth beacons. The beacons periodically broadcast packets containing Continuous Tone Extension (CTE). The Bluetooth base station, through a radio frequency switch, sequentially collects IQ samples from each antenna according to a predetermined switching mode. A typical AoA process is as follows: the beacon node continuously transmits signals carrying CTEs, which are received by the multi-antenna array. The collected IQ phase information is then sent to the positioning server via an IQ sampler. The server then performs angle calculations (azimuth and elevation) and coordinate estimation. The Bluetooth protocol stack can configure parameters such as CTE length and antenna switching mode using the Host command to ensure that the receiver collects I / Q data that meets specifications. Overall, a Bluetooth AoA system includes beacons, antenna arrays, IQ samplers, and positioning servers, encompassing the entire process from physical signal acquisition to angle measurement and positioning calculation.

[0087] MUSIC: Multiple Signal Classification. The MUSIC algorithm is a high-resolution direction-of-arrival (DOA) estimation method based on subspace decomposition. Its core principle is to perform eigenvalue decomposition on the covariance matrix of the received signal, dividing the observation space into a signal subspace and a noise subspace, and constructing a spatial spectral function using the orthogonality between the signal steering vector and the noise subspace. The target azimuth and elevation angles can be calculated by searching for the spectral peak positions. Under ideal conditions, it can achieve super-resolution angle estimation, but in real multipath environments, it suffers from pseudo-spectral distortion defects that require additional correction.

[0088] Existing Bluetooth AoA indoor positioning methods based on the MUSIC algorithm have several problems: First, indoor multipath reflections cause severe distortion of the MUSIC pseudospectrum, which easily leads to false spectral peaks and reduces the accuracy of angle estimation; second, the IQ phase acquired by the hardware is easily affected by noise and clock jitter, resulting in large fluctuations in the estimation results; finally, each antenna channel is processed independently, lacking spatiotemporal context fusion, and failing to make full use of the joint information of multiple antennas, thus restricting the overall performance of the positioning system.

[0089] Therefore, the main objective of this invention is to provide an indoor positioning method and system that aims to solve the technical problems of multipath interference, phase jitter, and lack of multi-antenna fusion in the prior art.

[0090] To achieve the above objectives, refer to Figures 1 to 10 The first aspect of this invention provides an indoor positioning method, comprising the following steps:

[0091] Step S100: Enhance the original IQ in-phase orthogonal signals acquired by multiple antennas using a lightweight neural network to output an enhanced IQ signal matrix that resists multipath interference;

[0092] Step S200: Construct a covariance matrix based on the enhanced signal matrix, and calculate the azimuth and elevation angles using the MUSIC multi-signal classification algorithm.

[0093] For details, please refer to Figure 1In a specific embodiment of the present invention, a Bluetooth beacon periodically broadcasts signal packets containing a continuous tone extended CTE. An N-element Bluetooth base station Locator deployed at the airport polls multiple antenna channels through a radio frequency switch in a preset switching mode. The signal acquisition module synchronously captures the signal through an N-element antenna array, obtaining the original IQ data of M sampling points of each antenna. After mixing and demodulation, an N×M-dimensional original IQ complex matrix containing phase difference information is generated. This matrix is ​​input into an L-MDA lightweight neural network module, which extracts inter-antenna correlation features through spatial feature extraction, models long-term dependence of CTE signals through temporal convolution, and fuses multi-channel information through an attention mechanism to achieve joint denoising and phase calibration, outputting an enhanced IQ matrix resistant to multipath interference. Subsequently, the angle calculation module constructs a covariance matrix based on the enhanced matrix, uses the MUSIC algorithm to perform eigenvalue decomposition to separate the signal / noise subspace, constructs a two-dimensional pseudo-spectrum of azimuth and elevation through the orthogonality of the noise subspace, and searches for spectral peaks. The direction corresponding to the position of the spectral peak is the azimuth and elevation angle of the signal. Combined with known positioning geometry, the beacon coordinates are calculated, achieving centimeter-level real-time positioning of airport equipment. It should be noted that the L-MDA (Lightweight Multi-channel Denoising & Alignment) in this embodiment is a lightweight neural network architecture specifically designed for multi-antenna signal processing. It achieves joint denoising and phase calibration by extracting spatial features from multiple antennas through collaborative convolutional layers, capturing long-term signal dependencies through temporal convolutional layers, and fusing multi-channel information through an attention mechanism. Its core innovation lies in achieving spatial-temporal feature alignment of the original IQ signal with extremely low computational cost, suppressing multipath reflection components, and eliminating hardware-introduced phase jitter, providing robust enhanced signal input for subsequent high-precision angle calculation.

[0094] Understandably, this invention enhances the collaborative framework of IQ signals and MUSIC algorithm through neural network preprocessing, which can significantly suppress pseudo-spectral distortion caused by multipath reflection and enhance angle resolution accuracy. At the same time, it overcomes the random jitter of IQ phase acquired by hardware and improves the stability of angle estimation. Through the multi-antenna spatiotemporal feature fusion mechanism, it achieves centimeter-level positioning accuracy in the dynamic environment of airports, greatly enhancing the system robustness in complex scenarios.

[0095] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, the L-MDA lightweight neural network module can be replaced with a lightweight Transformer architecture with similar functions to perform multi-channel signal enhancement, or the ESPRIT algorithm can be used to replace MUSIC for angle calculation. Alternatively, a network preprocessing module can be deployed on an embedded FPGA to optimize real-time performance through hardware acceleration.

[0096] Preferably, in step S100, the step of enhancing the original IQ in-phase quadrature signals acquired by multiple antennas using a lightweight neural network includes:

[0097] Step S110: Normalize the raw IQ signals acquired by the multiple antennas and output the normalized IQ matrix;

[0098] Step S120: Input the normalized IQ matrix into a lightweight two-dimensional convolutional network to extract spatial features between multiple antennas and output a spatial feature map;

[0099] Step S130: Input the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporal enhanced features;

[0100] Step S140: Input the temporal enhancement features into the attention mechanism layer, perform multi-channel feature weighted fusion, and output the weighted fused features;

[0101] Step S150: Based on the weighted fusion features, perform linear mapping and combination to reconstruct the output enhanced IQ signal matrix.

[0102] For details, please refer to Figure 2 and Figure 3 In a specific embodiment of the present invention, a lightweight neural network enhancement is performed on the original IQ data of N×M dimensions through a five-stage pipeline architecture. The L-MDA lightweight neural network includes: (1) a two-dimensional CNN layer (lightweight two-dimensional convolutional network) for extracting multi-antenna spatial features; (2) a one-dimensional TCN layer (temporal convolutional network) for capturing temporal correlations by using causal convolution with a doubling of the dilation rate layer by layer; and (3) an attention layer (multi-head attention mechanism layer) for weighted fusion of multi-channel features. The specific workflow is as follows: First, the original IQ signals acquired by multiple antennas are normalized by amplitude mean to eliminate channel gain differences and output a normalized IQ matrix in the complex domain. Then, the normalized IQ matrix is ​​input into a lightweight two-dimensional convolutional network, and a 3×3 convolutional kernel is used to slide across the spatial dimension to extract the correlation features of the multi-antenna signals, generating a spatial feature map containing channel dimension expansion. Next, the feature map is processed by a temporal convolutional network. By stacking causal convolutional layers with progressively increasing dilation rates, long temporal dependencies are captured and residual connections are added to output temporally smoothed temporal enhancement features. Then, the temporal enhancement features are input into a multi-head attention mechanism layer to dynamically calculate the weights between channels and fuse multi-scale information to generate a context-aware fusion feature tensor, which is output as a weighted fusion feature. Finally, the weighted fusion features are separated and reconstructed through a linear layer to form a complex-form enhanced IQ signal matrix.

[0103] Understandably, this invention eliminates antenna hardware gain deviations through normalization preprocessing, ensuring phase alignment references for multiple antennas; through the cascaded design of two-dimensional convolution and temporal convolution, it simultaneously optimizes spatial feature resolution and temporal signal smoothness, effectively suppressing phase jumps caused by multipath interference; and through the combination of attention mechanisms and linear reconstruction, it fully integrates multi-channel joint information to reconstruct an enhanced IQ signal with high coherence. In summary, this embodiment significantly improves the data quality of the MUSIC algorithm input, resulting in lower measured azimuth fluctuation ranges in strong reflection environments like airports, and higher positioning stability compared to traditional methods.

[0104] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, replacing two-dimensional convolution with depthwise separable convolution can reduce computational load, or using bidirectional LSTM to replace temporal convolutional networks can enhance temporal modeling capabilities. Alternatively, sparse weight constraints can be introduced into the attention mechanism layer to improve real-time performance, or linear reconstruction can be replaced with direct transformation in the complex domain to simplify the operation process. It should be stated that any processing chain including signal normalization, spatial feature extraction, temporal enhancement, multi-channel fusion, and signal reconstruction, if its core objective is to generate anti-interference enhanced IQ signals, falls under the equivalent technical solutions of this claim.

[0105] Preferred, Reference Figure 4 In step S110, the step of normalizing and preprocessing the raw IQ signals acquired by multiple antennas and outputting the normalized IQ matrix includes:

[0106] Step S111: Construct the original IQ data matrix based on the original IQ signals acquired by multiple antennas.

[0107] In the formula, N represents the total number of antennas, M represents the number of sampling points per antenna, and X... n,m Represents the complex IQ value of the m-th sampling point of the n-th antenna (n = 1, 2, ..., N; m = 1, 2, ..., M);

[0108] Step S112: Based on the original IQ data matrix X, calculate the average signal amplitude of each antenna n. The calculation formula is as follows:

[0109]

[0110] In the formula, μ n This represents the average signal amplitude across M sampling points for each antenna n;

[0111] Step S113: For each element X in the complex matrix X n,m Normalization is performed to obtain the normalized IQ matrix. The calculation formula is:

[0112]

[0113] In the formula, This represents the normalized complex signal value of the nth antenna at the mth sampling point.

[0114] For details, please refer to Figure 4 In a specific embodiment of the present invention, the data structure is organized after the raw IQ signal is acquired by a multi-antenna base station array. First, a complex matrix X with dimensions N rows and M columns is constructed, where N represents the total number of antennas (set to 8 in this embodiment), and M represents the number of sampling points per antenna (set to 8 in this embodiment). Matrix elements X... n,m This corresponds to the complex signal value at the m-th sampling time of the n-th antenna. An example input matrix is ​​shown below (primarily demonstrating the data flow and dimensional changes of each module, without detailing the specific numerical calculations of each element): Subsequently, each antenna n is processed independently, and the mean signal amplitude μ of all M sampling points of that antenna is calculated. n , The specific operation involves taking the modulus value of each point in the nth row of data and accumulating them to calculate the arithmetic mean. For example, for the eight complex elements of the first antenna: 28+54j, -62+40j, ..., calculate their absolute values ​​and then average them. Finally, the core normalization operation is performed, which normalizes each complex element X in the original antenna matrix. n,m Divide by the mean amplitude μ calculated by this antenna n , The final output is a normalized matrix. at this time The shape remains 8×8, with the real and imaginary parts serving as two-channel input networks, each with a size of 8×8×2. This preprocessing process is completed in an embedded processor, which can effectively reduce the amplitude variance of each channel signal, eliminate hardware inconsistencies, and establish a phase alignment reference for subsequent processing.

[0115] Understandably, this invention ensures the integrity of spatiotemporal relationships by constructing a structured data matrix; accurately captures channel characteristic differences through single-antenna amplitude mean calculation; and effectively eliminates hardware gain deviations through element-by-element complex normalization. This ensures that the signals from each antenna are on a unified dimensional benchmark, providing a standardized data foundation for subsequent spatial feature extraction and avoiding interference from channel differences on neural network feature alignment.

[0116] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, they can replace the mean amplitude with the root mean square value to adapt to high dynamic signal environments, or introduce a sliding window mean update mechanism while maintaining independent processing of each channel; or modify μ n Apply amplitude limiting constraints to prevent outlier interference.

[0117] Preferably, in step S120, the step of inputting the normalized IQ matrix into a lightweight two-dimensional convolutional network to extract spatial features between multiple antennas and output a spatial feature map includes:

[0118] Step S121: Take the normalized IQ matrix as input, perform spatial dimension sliding convolution using a convolution kernel of size 3×3 and stride of 1, and output the convolution result matrix;

[0119] Step S122: Apply the ReLU activation function to the convolution result matrix to output the spatial feature map. The calculation formula is as follows:

[0120]

[0121] In the formula, * denotes the convolution operation, W1 is the convolution kernel parameter, and b1 is the bias term. Let F1 ∈ R be the normalized IQ matrix, ReLU(·) be the ReLU activation function, and F1 ∈ R. N×M×C The three-dimensional tensor representing the spatial feature map is denoted by C, where C is the number of channels, N is the total number of antennas, and M is the number of sampling points per antenna.

[0122] For details, please refer to Figure 5 In a specific embodiment of the present invention, the spatial feature extraction process is as follows: First, the normalized 8×8-dimensional complex matrix is ​​input into a lightweight two-dimensional convolutional network. → Consider it as H×W×C in =8×8×2, using a 3×3 pixel convolution kernel to slide across the matrix space dimension, with a kernel stride of 1 pixel to ensure the output size matches the input; the convolution kernel parameters and bias terms are optimized during training. After performing a linear weighted summation on the input matrix, the intermediate output matrix maintains an 8×8 dimension but implicitly adds a feature channel dimension; subsequently, a linear rectified function is applied to this matrix, setting negative elements to zero while keeping positive values ​​unchanged, generating a 3D feature map with dimensions of 8×8×C, where C is the number of channels determined by the number of convolution kernels. In this embodiment, C is set to 16. The final output feature map is... The output feature map significantly enhances the spatial correlation between multiple antennas, providing structured input for subsequent time series modeling.

[0123] Understandably, this invention effectively captures the phase difference features of adjacent antenna signals through dense sliding operations of fixed-size convolution kernels, enhancing the spatial correlation of array signals. By using nonlinear activation of a linear rectified function, it suppresses abnormal negative responses in the feature map caused by hardware noise, thereby improving the sparsity and robustness of feature representation. In actual testing in airport equipment positioning scenarios, the spatial feature map output by this module effectively improves the signal-to-noise ratio of subsequent angle calculation modules and reduces the false positioning rate caused by reflections from metal shelves.

[0124] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, the kernel size can be adjusted to 5×5 pixels to expand the receptive field, or a parameterized linear rectified function can be used to replace the standard function to achieve nonlinear adaptive control; or the stride can be set to 2 pixels to reduce the feature map size.

[0125] Preferably, in step S130, the step of inputting the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporally enhanced features includes:

[0126] Step S131: Apply one-dimensional dilated convolution to the temporal feature vector corresponding to each antenna n in the spatial feature map. The calculation formula is as follows:

[0127]

[0128] In the formula, This represents the output feature vector of the nth antenna in the lth layer, where l represents the network layer number, and d l The expansion rate of the l-th layer is d. l =2 l-1 ReLU(·) is the ReLU activation function, and the kernel size is k. Represents the timing characteristics of the nth antenna, Conv1D causal For causal convolution operations, left padding is used to ensure that the output length is consistent with the input.

[0129] Step S132: Perform a residual connection operation on the dilated convolution output. After stacking L layers of processing sequentially, the output is a temporal augmentation feature.

[0130] In the formula, F2∈R N×M×C This represents the output of the timing enhancement feature, where N is the total number of antennas, M is the number of sampling points, and C is the number of feature channels. This represents the final feature vector of the nth antenna after processing through L layers.

[0131] The ReLU activation function (Rectified Linear Unit) is one of the most widely used non-linear activation functions in deep learning. Its core function is to introduce non-linear characteristics into neural networks, enabling models to learn and fit complex data patterns and features.

[0132] For details, please refer to Figure 6In a specific embodiment of the present invention, the previously output 8×8×16 dimensional spatial feature map is input into a temporal convolutional network. For ease of discussion later, C is set to 16 channels, and the temporal convolutional network is set to 3 layers. ReLU is applied element-wise to F1: F1←max(F1,0), with the size remaining unchanged at 8×8×16. For the 8-step time sequence of each antenna n, three layers of dilated causal convolution are performed in parallel, maintaining the number of channels at 16, kernel width k=3, and dilation rates d=1, 2, 4 respectively. Residual connections are added after each layer and ReLU is applied. Specifically: the first layer uses a one-dimensional convolutional kernel of size 3, performs uninterrupted sliding operations in the temporal dimension with a stride of 1, and ensures that the input and output lengths are consistent by left-padding with zero values; the second to third layers progressively set the dilation rate to 2 times and 4 times, respectively, to achieve exponentially expanded receptive field coverage; the output of each layer is processed by a linear rectified function and then superimposed with the original input through a residual connection across layers to suppress gradient decay in the deep network. This can be expressed by the following formula:

[0133] Layer 1 (d=1):

[0134] Layer 2 (d=2):

[0135] Layer 3 (d=4):

[0136] in:

[0137] Layer 1: Dilation rate = 1 → Convolutional kernel covers [t-1, t, t+1]

[0138] Layer 2: Expansion rate = 2 → Covers [t-2, t, t+2]

[0139] Layer 3: Expansion rate = 4 → Covers [t-4, t, t+4]

[0140] The final output combines the outputs of all antennas n=1,…,8 and outputs a timing enhancement feature with dimensions maintained at 8×8×16. The phase abrupt change points caused by multipath interference are smoother than before processing, providing temporal consistency assurance for subsequent attention fusion.

[0141] Understandably, this invention, through its exponentially expanded convolutional layer design, allows the network's receptive field to multiply with the number of layers, significantly improving its ability to model long temporal dependencies and solving the inherent problem of gradient vanishing in long sequences in traditional recurrent networks. By using causal convolution combined with a left-padding mechanism, it ensures that the output sequence is of equal length to the input and does not leak future information, meeting real-time processing requirements. The introduction of residual connections avoids optimization degradation in deep networks and ensures efficient cross-layer feature reuse. Compared to traditional solutions that independently process each antenna channel, this embodiment can fully exploit the phase continuity of the signal in the temporal domain, resulting in lower trajectory jitter in airport baggage cart positioning scenarios and higher stability in predicting the pitch angle of high-speed moving targets. This temporal enhancement mechanism provides a temporally smooth and context-aware feature foundation for subsequent multi-channel information fusion, improving the dynamic trajectory tracking accuracy of the positioning system.

[0142] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, they can use gated convolutional units to replace standard convolutional kernels to enhance the selectivity of temporal modeling, or use bidirectional temporal convolutional networks to break through the unidirectional limitation of causal convolution; or adopt non-linear growth patterns such as the Fibonacci sequence in the dilation rate setting, or introduce learnable scaling weights in the residual connections.

[0143] Preferably, in step S140, the step of inputting the temporal enhancement features into the attention mechanism layer, performing multi-channel feature weighted fusion, and outputting weighted fused features includes:

[0144] Step S141: Map the temporal enhancement feature F2 to the query matrix Q, key matrix K, and value matrix V respectively, using the following formula:

[0145] Q = F2W Q

[0146] K = F2W K

[0147] V = F2W V

[0148] In the formula, W Q To query the trainable weight matrix of matrix Q, W K W is the trainable weight matrix of the key matrix K. V Let V be the trainable weight matrix of the value matrix V;

[0149] Step S142: Based on the query matrix Q and the key matrix K, calculate the attention weight matrix using the following formula:

[0150]

[0151] In the formula, A represents the attention weight matrix, and d kLet Q represent the key matrix dimension, Q represent the query matrix, and K represent the key matrix dimension. T represents the transpose of the key matrix, and softmax(·) represents the normalized exponential function;

[0152] Step S143: Perform weighted fusion of the value matrix V using the attention weight matrix A, and output the weighted fusion feature:

[0153] F3 = AV

[0154] In the formula, F3 represents the weighted fusion feature.

[0155] For details, please refer to Figure 7 In a specific embodiment of the present invention, the previously output 8×8×16 temporal enhancement feature F2 is first flattened into 64×16 (64 spatial locations, dimension 16), and then mapped in parallel to the query, key, and value feature spaces through three sets of independent trainable weight matrices. The spatial dimension remains unchanged during the mapping process, but the feature representation capability is expanded; then, the similarity correlation matrix between the query vector and the key vector is calculated. After standardization adjustment, a dynamic weight distribution is generated using a normalized exponential function. This distribution significantly enhances effective signal components (such as line-of-sight paths) and suppresses multipath interference components (such as reflections from metal shelves). Finally, this weight matrix is ​​used to perform spatiotemporal adaptive weighted aggregation of the value space features, outputting the fused weighted fusion features.

[0156] Understandably, this invention overcomes the limitations of traditional fixed templates through a dynamic weight allocation mechanism, significantly enhancing effective signal components and deeply suppressing multipath reflection interference; it strengthens phase consistency through channel fusion, effectively eliminating sampling point offsets caused by receiver clock jitter; and it ensures signal continuity through temporal correlation modeling, maintaining trajectory stability in high-speed moving scenarios. The overall design enhances feature discrimination capabilities, establishing robust feature representations for high-precision positioning, and exhibits excellent adaptability, particularly in complex indoor environments.

[0157] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, in the feature mapping stage, grouped convolution can be used to replace fully connected layers to compress the number of parameters, or low-rank approximation can be used to optimize the efficiency of matrix operations; sparse constraints can be introduced in the weight calculation stage to improve real-time performance, or cosine similarity can be used to enhance the rationality of high-dimensional space; or a gating mechanism can be combined in the fusion output part to control the information flow, or residual paths can be added to retain the original features.

[0158] Preferably, in step S150, the step of reconstructing the output enhanced IQ signal matrix by performing linear mapping and combination based on the weighted fusion features includes:

[0159] Step S151: Perform a two-way linear mapping on the last channel dimension of the weighted fusion feature to generate the real part matrix and the imaginary part matrix. The calculation formula is as follows:

[0160]

[0161] In the formula, f3(m,n) represents the feature vector of the m-th antenna and the n-th sampling point in the F3 indoor positioning system, I m,n w represents the real part signal value of the m-th antenna at the n-th sampling point after reconstruction. r Let b be the weight vector mapped to the real part. r The bias term for the real part mapping; Q m,n w represents the imaginary part signal value of the m-th antenna at the n-th sampling point after reconstruction. i Let b be the weight vector of the imaginary part mapping. i The bias term for the mapping of the imaginary part; w r ,w i and b r ,b i These are trainable parameters;

[0162] Step S152: Combine the real part matrix and the imaginary part matrix to generate the enhanced IQ signal matrix, calculated using the following formula:

[0163]

[0164] In the formula, Let j represent the complex IQ signal enhanced by the m-th antenna at the n-th sampling point, where j represents the imaginary unit. This represents the final output enhanced IQ signal matrix.

[0165] For details, please refer to Figure 8 In a specific embodiment of the present invention, for the 16-dimensional feature vector f3(n,m) at each spatial location (n,m), a weighted transformation is performed on the feature vector at each spatiotemporal location using independently trainable parameters, generating a real component matrix along the way. Another approach generates the imaginary component matrix. The real and imaginary matrices are then combined element by element to form a complex signal. The final enhanced complex matrix is ​​obtained as follows:

[0166]

[0167] The reconstructed enhanced IQ matrix has the same dimensions as the original acquired signal. While retaining the multi-channel feature discrimination power, it accurately restores the physical waveform structure of the signal, ensuring that the output can be directly input into the traditional MUSIC algorithm module.

[0168] Understandably, this invention achieves three technological breakthroughs through a dual-path linear mapping and complex combination mechanism: First, the accurate restoration of feature vectors to physical signals ensures the integrity of phase information and avoids distortion during the conversion between the feature domain and the signal domain; second, the real-virtual part separation mapping design inherits the anti-interference characteristics of neural networks, enabling the reconstructed signal to maintain a low waveform distortion rate in a multipath environment; third, the dimensional consistency output eliminates algorithm adaptation costs and improves the integration efficiency of the positioning system.

[0169] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, they can use complex convolutional layers to replace dot product operations to achieve direct complex mapping, or introduce hyperbolic tangent activation functions to constrain the dynamic range of signals; or they can merge the mapping of real and imaginary parts into a single-path complex linear transformation to improve efficiency.

[0170] Preferably, in step S200, the step of constructing a covariance matrix based on the enhanced signal matrix and calculating the azimuth and elevation angles using the MUSIC multiple signal classification algorithm includes:

[0171] Step S210: Divide the time snapshot samples based on the enhanced IQ signal matrix, and calculate the output signal covariance matrix by autocorrelation averaging;

[0172] Step S220: Perform eigenvalue decomposition on the signal covariance matrix, separate and output the noise subspace matrix corresponding to the smallest eigenvalue;

[0173] Step S230: Based on the noise subspace and planar array model, calculate the two-dimensional pseudospectral function that characterizes the signal intensity at different incident angles;

[0174] Step S240: Based on the two-dimensional pseudospectral function, perform a global peak search in the azimuth and elevation parameter space, and output the estimated values ​​of the target's azimuth and elevation angles.

[0175] For details, please refer to Figure 9 In a specific embodiment of the present invention, the incident angle is estimated by inputting an enhanced IQ signal matrix, specifically including:

[0176] Step S210: Divide the time snapshot samples based on the enhanced IQ signal matrix, and calculate the output signal covariance matrix by autocorrelation averaging;

[0177] Specifically, this involves enhancing the IQ signal matrix. Divide the samples into L=M snapshots along the time dimension. Each snapshot sample Given the sampling vectors (L=M) of N antennas at a single sampling time, calculate the signal covariance matrix R:

[0178]

[0179] In the formula, (·) H Represents the conjugate transpose, R∈C N×N Represents the covariance matrix. Let L represent the l-th snapshot sample, and L represent the total number of snapshot samples.

[0180] Step S220: Perform eigenvalue decomposition on the signal covariance matrix, separate and output the noise subspace matrix corresponding to the smallest eigenvalue;

[0181] Specifically, this involves performing eigenvalue decomposition on the signal covariance matrix R to separate the signal subspace V. s With noise subspace V n Output spatial feature vector set V n :

[0182] R = VΛV H

[0183] In the formula, Λ represents the diagonal eigenvalue matrix, and V = [V s V n ] represents the eigenvector matrix, V s V represents the signal subspace. n This represents the noise subspace (composed of eigenvectors with ND minimum eigenvalues, where D is the number of incident signals).

[0184] Step S230: Based on the noise subspace and planar array model, calculate the two-dimensional pseudospectral function that characterizes the signal intensity at different incident angles;

[0185] Specifically, based on spatial feature vector V n Given the array steering vector a(φ,θ), calculate the two-dimensional pseudospectral function:

[0186]

[0187] In the formula, a(φ,θ) represents the array steering vector, (·) H Indicates conjugate transpose; elements of a(φ,θ) are a n = Where k is the wave number, d n φ represents the position coordinates of the nth antenna; φ represents the azimuth angle, and θ represents the elevation angle.

[0188] Step S240: Based on the two-dimensional pseudospectral function, perform a global peak search in the azimuth and elevation parameter space, and output the estimated values ​​of the target's azimuth and elevation angles.

[0189] Specifically: within the angle range [φ] min ,φ max ]×[θ min ,θmax The peak position of the pseudospectral function P(φ,θ) is determined by a global search within the domain, using the following formula:

[0190]

[0191] The final result is: azimuth angle and pitch angle

[0192] Understandably, this invention suppresses single-point sampling noise through statistical averaging of the covariance matrix, ensuring the robustness of spatial correlation measurement; it effectively suppresses pseudo-spectral peaks caused by multipath propagation through noise subspace projection operation, improving the angle resolution accuracy of the real signal; and it avoids local optimization traps through a global peak search strategy, ensuring the stability of angle estimation in dynamic environments.

[0193] Based on the above technical solutions, those skilled in the art can make corresponding equivalent improvements according to specific application scenarios. For example, in the covariance calculation stage, a robust estimation method can be used to replace the arithmetic mean to enhance the ability to resist outliers, or shrinkage regularization can be used to optimize the matrix ill-conditioned problem in small sample scenarios; or the noise subspace extraction can be improved to a feature selection adaptive algorithm to dynamically determine the number of signal sources; or interpolation scanning technology can be introduced in the pseudospectral generation stage to improve angular resolution, or compressed sensing can be used to reconstruct the sparse spatial spectrum to reduce the amount of computation; or the peak search process can be combined with heuristic optimization algorithms to accelerate localization.

[0194] refer to Figure 10 A second aspect of the present invention provides an indoor positioning system, comprising:

[0195] The signal enhancement module enhances the original IQ in-phase orthogonal signals acquired by multiple antennas through a lightweight neural network, and outputs an enhanced IQ signal matrix that resists multipath interference.

[0196] The angle calculation module constructs a covariance matrix based on the enhanced signal matrix and uses the MUSIC multi-signal classification algorithm to calculate the azimuth and elevation angles.

[0197] refer to Figure 11 The third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the indoor positioning method as described in the first aspect.

[0198] A fourth aspect of the present invention also provides a computer-readable storage medium storing an indoor positioning processing program, which, when executed by a processor, implements the steps of the indoor positioning method as described in any of the above embodiments.

[0199] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the indoor positioning system, connecting various parts of the entire indoor positioning system's processing and operating devices via various interfaces and lines.

[0200] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the indoor positioning system by running or executing the computer programs and / or modules stored in the memory and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0201] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0202] This invention provides an indoor positioning method and system. The indoor positioning method uses a lightweight neural network to jointly enhance the original IQ signals from multiple antennas. During the signal input stage, it deeply suppresses multipath reflection interference and hardware phase jitter, generating a highly consistent enhanced IQ signal matrix. Then, it uses the MUSIC algorithm to construct a highly robust spatial spectrum model based on the enhanced IQ signal matrix. By accurately eliminating spurious direction-of-arrival components through the noise subspace, it ultimately solves the technical problems of pseudo-spectral distortion caused by multipath interference in angle calculation in complex indoor scenarios, covariance inaccuracy caused by phase jitter, and the weakening of spatiotemporal correlation due to the isolation of multi-antenna information. This significantly improves the estimation accuracy and dynamic stability of azimuth and elevation angles, providing reliable positioning capabilities for indoor positioning scenarios.

[0203] Furthermore, the indoor positioning method of this invention also eliminates hardware gain differences through normalization preprocessing, improving the phase comparability of multi-channel signals; extracts spatial correlation features through a two-dimensional convolutional network, strengthening the spatial consistency expression of array signals; models long-term temporal dependencies through a temporal convolutional network, enhancing the capture of signal continuity and periodic patterns; dynamically weights multi-channel features through an attention mechanism, increasing the weight of effective components in interference environments; ensures the integrity of signal waveform structure and phase fidelity through dual-path linear mapping and complex reconstruction; suppresses noise in snapshot samples and improves the robustness of spatial correlation measurement through statistical averaging optimization of the covariance matrix; accurately suppresses pseudo-spectral peak interference caused by multipath through orthogonal projection of the noise subspace; and ensures the global optimality of angle estimation in dynamic scenes through a global spectral peak search strategy.

[0204] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. Equivalent structural transformations made using the description and drawings of the present invention, or direct / indirect applications in other related technical fields, are all included within the scope of patent protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. An indoor positioning method, characterized in that, Includes the following steps: Step S100: Enhance the original IQ in-phase orthogonal signals acquired by multiple antennas using a lightweight neural network to output an enhanced IQ signal matrix that resists multipath interference; Step S200: Construct a covariance matrix based on the enhanced signal matrix, and calculate the azimuth and elevation angles using the MUSIC multi-signal classification algorithm.

2. The indoor positioning method as described in claim 1, characterized in that, In step S100, the step of enhancing the original IQ in-phase quadrature signals acquired by multiple antennas using a lightweight neural network includes: Step S110: Normalize and preprocess the raw IQ signals acquired by the multiple antennas, and output the normalized IQ matrix; Step S120: Input the normalized IQ matrix into a lightweight two-dimensional convolutional network to extract spatial features between multiple antennas and output a spatial feature map; Step S130: Input the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporal enhanced features; Step S140: Input the temporal enhancement features into the attention mechanism layer, perform multi-channel feature weighted fusion, and output the weighted fused features; Step S150: Based on the weighted fusion features, perform linear mapping and combination to reconstruct the output enhanced IQ signal matrix.

3. The indoor positioning method as described in claim 2, characterized in that, In step S110, the step of normalizing and preprocessing the raw IQ signals acquired by multiple antennas and outputting the normalized IQ matrix includes: Step S111: Construct the original IQ data matrix based on the original IQ signals acquired by multiple antennas. In the formula, N represents the total number of antennas, M represents the number of sampling points per antenna, and X... n,m Represents the complex IQ value of the m-th sampling point of the n-th antenna (n = 1, 2, ..., N; m = 1, 2, ..., M); Step S112: Based on the original IQ data matrix X, calculate the average signal amplitude of each antenna n. The calculation formula is as follows: In the formula, μ n This represents the average signal amplitude across M sampling points for each antenna n; Step S113: For each element X in the complex matrix X n,m Normalization is performed to obtain the normalized IQ matrix. The calculation formula is: In the formula, This represents the normalized complex signal value of the nth antenna at the mth sampling point.

4. The indoor positioning method as described in claim 3, characterized in that, In step S120, the step of inputting the normalized IQ matrix into a lightweight two-dimensional convolutional network, extracting spatial features between multiple antennas, and outputting a spatial feature map includes: Step S121: Take the normalized IQ matrix as input, perform spatial dimension sliding convolution using a convolution kernel of size 3×3 and stride of 1, and output the convolution result matrix; Step S122: Apply the ReLU activation function to the convolution result matrix to output the spatial feature map. The calculation formula is as follows: In the formula, * denotes the convolution operation, W1 is the convolution kernel parameter, and b1 is the bias term. Represents the normalized IQ matrix, and ReLU(·) is the ReLU activation function. The three-dimensional tensor representing the spatial feature map is denoted by C, where C is the number of channels, N is the total number of antennas, and M is the number of sampling points per antenna.

5. The indoor positioning method as described in claim 4, characterized in that, In step S130, the step of inputting the spatial feature map into a temporal convolutional network to enhance its temporal correlation and output temporally enhanced features includes: Step S131: Apply one-dimensional dilated convolution to the temporal feature vector corresponding to each antenna n in the spatial feature map. The calculation formula is as follows: In the formula, This represents the output feature vector of the nth antenna in the lth layer, where l represents the network layer number, and d l The expansion rate of the l-th layer is d. l =2 l-1 ReLU(·) is the ReLU activation function, and the kernel size is k. Represents the timing characteristics of the nth antenna, Conv1D causal For causal convolution operations, left padding is used to ensure that the output length is consistent with the input. Step S132: Perform a residual connection operation on the dilated convolution output. After stacking L layers of processing sequentially, the output is a temporal augmentation feature. In the formula, F2∈R N×M×C This represents the output of the timing enhancement feature, where N is the total number of antennas, M is the number of sampling points, and C is the number of feature channels. This represents the final feature vector of the nth antenna after processing through L layers.

6. The indoor positioning method as described in claim 5, characterized in that, In step S140, the step of inputting the temporal enhancement features into the attention mechanism layer, performing multi-channel feature weighted fusion, and outputting weighted fused features includes: Step S141: Map the temporal enhancement feature F2 to the query matrix Q, key matrix K, and value matrix V respectively, using the following formula: Q=F2W Q K = F2W K V=F2W V In the formula, W Q To query the trainable weight matrix of matrix Q, W K W is the trainable weight matrix of the key matrix K. V Let V be the trainable weight matrix of the value matrix V; Step S142: Based on the query matrix Q and the key matrix K, calculate the attention weight matrix using the following formula: In the formula, A represents the attention weight matrix, and d k Let Q represent the key matrix dimension, Q represent the query matrix, and K represent the key matrix dimension. T represents the transpose of the key matrix, and softmax(·) represents the normalized exponential function; Step S143: Perform weighted fusion of the value matrix V using the attention weight matrix A, and output the weighted fusion feature: F3 = AV In the formula, F3 represents the weighted fusion feature.

7. The indoor positioning method as described in claim 6, characterized in that, In step S150, the step of reconstructing the output enhanced IQ signal matrix by performing linear mapping and combination based on the weighted fusion features includes: Step S151: Perform a two-way linear mapping on the last channel dimension of the weighted fusion feature to generate the real part matrix and the imaginary part matrix. The calculation formula is as follows: In the formula, f3(m,n) represents the feature vector of the m-th antenna and the n-th sampling point in the F3 indoor positioning system, I m,n w represents the real part signal value of the m-th antenna at the n-th sampling point after reconstruction. r Let b be the weight vector mapped to the real part. r The bias term for the real part mapping; Q m,n w represents the imaginary part signal value of the m-th antenna at the n-th sampling point after reconstruction. i Let b be the weight vector of the imaginary part mapping. i The bias term for the mapping of the imaginary part; w r ,w i and b r ,b i These are trainable parameters; Step S152: Combine the real part matrix and the imaginary part matrix to generate the enhanced IQ signal matrix, calculated using the following formula: In the formula, Let j represent the complex IQ signal enhanced by the m-th antenna at the n-th sampling point, where j represents the imaginary unit. This represents the final output enhanced IQ signal matrix.

8. The indoor positioning method according to any one of claims 1 to 7, characterized in that, In step S200, the step of constructing a covariance matrix based on the enhanced signal matrix and calculating the azimuth and elevation angles using the MUSIC multiple signal classification algorithm includes: Step S210: Divide the time snapshot samples based on the enhanced IQ signal matrix, and calculate the output signal covariance matrix by autocorrelation averaging; Step S220: Perform eigenvalue decomposition on the signal covariance matrix, separate and output the noise subspace matrix corresponding to the smallest eigenvalue; Step S230: Based on the noise subspace and planar array model, calculate the two-dimensional pseudospectral function that characterizes the signal intensity at different incident angles; Step S240: Based on the two-dimensional pseudospectral function, perform a global peak search in the azimuth and elevation parameter space, and output the estimated values ​​of the target's azimuth and elevation angles.

9. An indoor positioning system, characterized in that, include: The signal enhancement module enhances the original IQ in-phase orthogonal signals acquired by multiple antennas through a lightweight neural network, and outputs an enhanced IQ signal matrix that resists multipath interference. The angle calculation module constructs a covariance matrix based on the enhanced signal matrix and uses the MUSIC multi-signal classification algorithm to calculate the azimuth and elevation angles.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the indoor positioning method as described in any one of claims 1 to 8.