Self-supervised ultra-wideband positioning method based on improved double-delay depth deterministic strategy
By introducing a self-supervised UWB localization method based on an improved dual-delay deep deterministic strategy, a spatiotemporal feature encoding layer with a dual-critic network structure and a self-attention mechanism is introduced. This solves the model generalization problem of traditional deep learning methods in UWB localization and achieves high-precision UWB localization.
Patent Information
- Application Number
- CN202511240272.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-02
AI Technical Summary
Existing deep learning methods for UWB localization rely on costly labeled data and suffer from unacceptable ranging error correction models when the deployment environment changes. The model performance deteriorates significantly when the environment layout changes, new obstacles are added, or the anchor node position is changed. These are the technical problems existing in the current technology.
A self-supervised ultrawideband localization method employing an improved dual-delay deep deterministic strategy is proposed. This method introduces a spatiotemporal feature encoding layer based on a self-attention mechanism through a dual-commenter network structure, a delayed update mechanism, and an action-noise self-supervised ultrawideband localization method. By capturing long-range dependencies through a dynamic weight allocation mechanism, a spatiotemporal correlation feature mapping across timestamps is established, which significantly improves learning stability and model accuracy.
This method effectively solves the problem of Q-value overestimation in traditional deep reinforcement learning algorithms, significantly improves learning stability and model accuracy, and is particularly suitable for handling noise and multipath effects in CIR data, achieving high-precision UWB localization.
Smart Images

Figure CN121099418A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of ultra-wideband positioning, and particularly relates to a self-supervised ultra-wideband positioning method based on an improved double-delay depth deterministic strategy. BACKGROUND
[0002] Ultra-wideband (UWB) can theoretically achieve centimeter-level positioning accuracy due to its wide bandwidth (>500MHz) and extremely short pulse duration (about 2ns), providing a feasible solution for indoor precise positioning. However, in actual application scenarios, especially in non-line-of-sight (NLOS) conditions, UWB positioning faces an intolerable ranging error problem. When the signal passes through obstacles such as walls and furniture, signal delay, multipath effect and energy attenuation occur, resulting in ranging errors of tens of centimeters or even more in some complex environments, which severely restricts the large-scale application of UWB technology in high-precision scenarios.
[0003] With the rapid development of artificial intelligence technology, deep learning methods have been introduced into the field of UWB ranging error correction, showing significant advantages. Neural network-based methods can learn the complex mapping relationship between signal features and ranging errors from channel impulse response (CIR) data, effectively identifying and correcting errors under NLOS conditions, and improving positioning accuracy to the centimeter level. This data-driven approach is more adaptable and robust than traditional signal processing methods, opening up new avenues for improving UWB positioning system performance. Typical practices such as the hierarchical graph learning model proposed by Gu et al. construct a base station-tag heterogeneous graph structure and use node embedding technology to explicitly encode the geometric relationship between spatial coordinates and ranging values, improving positioning accuracy by 41% compared to traditional methods in non-line-of-sight scenarios. Existing mature models have shown that the neighborhood information aggregation strategy based on graph attention mechanism (GAT) can effectively correct the ranging bias caused by multipath effects, while the cross-layer graph pooling operation can extract collaborative positioning features of multi-tag systems.
[0004] However, the key bottleneck of existing deep learning methods lies in the model training process. Traditional supervised learning methods require the collection of a large amount of labeled UWB ranging data, which usually relies on high-precision positioning systems as ground truth references. These systems are not only costly and complex to set up, but also require professional operation. The more serious challenge is that the trained model often faces serious generalization problems in the deployment environment - when the environment layout changes, new obstacles are added, or anchor node positions are changed, the model performance significantly decreases, requiring re-collection of data and training, which greatly limits the practicality and scalability of the system.
[0005] Reinforcement learning (RL) is a machine learning method that learns optimal decisions through environment interaction without relying on static labeled data, although reinforcement learning has been preliminarily explored in the fields of signal processing and wireless communication, its application in the field of UWB positioning error correction is still in its early stages. SUMMARY
[0006] The purpose of the present application is to provide a self-supervised ultra-wideband positioning method based on an improved double-delay deep deterministic strategy, which effectively solves the Q value overestimation problem in traditional deep reinforcement learning algorithms by introducing a space-time feature encoding layer based on a self-attention mechanism through a double critic network structure, a delay update mechanism and action noise regularization, significantly improving learning stability and model accuracy. This method is particularly suitable for processing noise and multipath effects in CIR data, enabling high-precision UWB positioning.
[0007] The technical solution to achieve the purpose of the present application is:
[0008] A self-supervised ultra-wideband positioning method based on an improved double-delay deep deterministic strategy, comprising the following steps:
[0009] Obtain channel impulse response data CIR and preprocess it;
[0010] Construct an improved double-delay deep deterministic strategy model for UWB ranging error correction, which is based on a double-delay deep deterministic strategy gradient model, designs a double critic network structure, introduces a space-time feature encoding layer based on a self-attention mechanism, captures long-range dependencies through a dynamic weight distribution mechanism, and establishes a cross-timestamp space-time correlation feature mapping; Introduce a target network delay update mechanism to make the update frequency of the target network lower than that of the critic network and the actor network, regularize the action noise generated by the strategy, and prevent the strategy from using the overestimated area of the Q function;
[0011] Generate pseudo-labels using an extended Kalman filter, dynamically adjust model parameters and learning strategies for iterative optimization, and obtain positioning results.
[0012] In the preferred technical solution, preprocessing includes:
[0013] Complex IQ sampling is performed on the channel impulse response data CIR, and the complex IQ sampling array is converted to an RSSI sampling array: I is the real part, Q is the imaginary part, and the RSSI array is clipped to a certain number of samples, focusing on the signal features near the first path;
[0014] Min-Max normalization on the cropped CIR: CIRnorm = (CIR-min(CIR)) / (max(CIR)-min(CIR));
[0015] CIRnorm is the normalized CIR data, min(CIR) and max(CIR) are the minimum and maximum values of the CIR data, respectively;
[0016] The normalized CIR data is filtered and dimensionally reduced to eliminate the influence of hardware transmission power difference and environmental distance change on the absolute signal strength.
[0017] In the preferred technical solution, the improved double-delay deep deterministic strategy model includes an Actor network and a Critic network. The Actor network serves as a policy generator and directly analyzes the time-domain features of the channel impulse response to dynamically output an error correction quantity that adapts to channel distortion. The Critic network plays the role of a value evaluator and builds an implicit quality evaluation system by analyzing the physical consistency of historical correction effects and motion trajectories. Coupled optimization is achieved through self-supervised signals formed by environmental feedback: the Actor network generates a correction strategy based on the current channel state, and the Critic network evaluates the effectiveness of the strategy through a reward function derived by a trajectory smoothing module, thereby guiding the parameter update direction.
[0018] In the preferred technical solution, the Actor network includes a current Actor network and a target Actor network, and the Critic network module includes two current Critic networks Q1 and Q2 and two target Critic networks. When calculating the target Q value, the minimum value of the outputs of the two current Critic networks is selected.
[0019] The Actor network and the Critic network have similar structures, including an input layer that receives preprocessed CIR data. The Critic network additionally inputs the action a output by the Actor network. Subsequently, the input is captured by a feature extraction layer composed of convolutional layers to capture local features of the data, and then input into a transformer encoder layer to capture global dependencies between features. Finally, through a fully connected layer, the Actor network and the Critic network output actions a and action values Q, respectively.
[0020] The Critic network module obtains the appropriate action value y through the mathematical expression: y = r + γ·min(Q1(s',π(s')),Q2(s',π(s'))).
[0021] Where r is the immediate reward, γ is the discount factor, s' is the next state, π is the policy network, and Q1 and Q2 are the outputs of the two current Critic networks.
[0022] In a preferred technical solution, the action noise regularization generated by the strategy comprises:
[0023] A clipping noise is added to the output a of the Actor network to obtain a':
[0024] a' = clip (π (s') + clip (ε, -c, c), a min ,a max )
[0025] Wherein, clip () is a clipping function, ε is random noise, c is noise amplitude limit, a min ,a max are the upper and lower bounds of action respectively.
[0026] In a preferred technical solution, the iterative optimization comprises:
[0027] The ranging information corrected by reinforcement learning is processed by extended Kalman filter, and state estimation and prediction are performed in combination with historical observation data. The filtered data are smoothed by a circular buffer to generate pseudo labels.
[0028] The pseudo label data are input into a smoothing buffer for processing, and a ranging reference for reward calculation is obtained through the Euclidean distance conversion ranging value link. Based on the converted ranging value and the actual ranging related quantity, the reward is calculated by a reward function, fed back to an experience replay buffer to provide an optimization signal to drive the improved double-delay deep deterministic policy model to iterate continuously, forming a closed-loop optimization.
[0029] In a preferred technical solution, the generation of pseudo labels comprises:
[0030] The current Actor network determines the range correction amount a, and the distance Δ t measured at the current time t, and the combination of the two obtains the corrected range estimate Δ t ' = Δ t -a.
[0031] Then, Δ t ' is processed by extended Kalman filter, and in the process of converting Δ t ' to position p EKF,t =EKF(Δ t '), EKF() is extended Kalman filter, first based on the dynamic change model of "range to position" to make prediction, using the position state at the last time to estimate the position at the current time, while considering the influence of process noise; Then, combined with the observation equation, the uncertainty of the prediction and the uncertainty of the observation are balanced through the Kalman gain, and the observation information is used to correct the prediction result to obtain a more accurate position; Abstract range information is mapped into position coordinates to realize the conversion from range to position.
[0032] Preferably, the driving improved double-delay deep deterministic policy model is iterated continuously to form a closed-loop optimization, which comprises:
[0033] Position p EKF,t The length of the ring buffer C is stored for smoothing, when the buffer is full, new data is added and the oldest data is removed, and the average position of the middle position of the buffer is calculated, and the middle position index Ensure the association with the data before and after the buffer, and the average position formula x i , y i is the position coordinate;
[0034] The average position p avg,m is converted into the corrected range estimate Δ avg,m , and the reward function Quantifies the correction effect, and Δ' m is the range estimate of the middle position m after being corrected by the Actor network, and Δ' m , Δ avg,m is closer, the reward is higher, which guides the agent to optimize the range correction, the reward provides feedback for the action of the agent, and the experience data for data training is obtained from the experience replay buffer batch sampling training, which helps the agent to learn a better range correction strategy.
[0035] The application also discloses a self-supervised ultra-wideband positioning system based on an improved double-delay deep deterministic policy, which comprises:
[0036] A data acquisition and preprocessing module is used to acquire channel impulse response data CIR and perform preprocessing;
[0037] The reinforcement learning module is used for constructing an improved double-delay deep deterministic policy model to correct UWB ranging error, wherein the improved double-delay deep deterministic policy model is based on a double-delay deep deterministic policy gradient model, a double-critic network structure is designed, a space-time feature encoding layer based on a self-attention mechanism is introduced, a long-range dependency relationship is captured through a dynamic weight distribution mechanism, and a time-space correlation feature mapping across timestamps is established; a target network delay updating mechanism is introduced, so that the updating frequency of the target network is lower than that of the critic network and the actor network, the action noise generated by the strategy is regularized, and the strategy is prevented from using the overestimated area of the Q function;
[0038] The self-supervised learning module is used for generating pseudo labels by using an extended Kalman filter, dynamically adjusting model parameters and learning strategies for iterative optimization, and obtaining a positioning result.
[0039] The application also discloses a computer storage medium, which stores a computer program, and the computer program is executed by a computer to realize the self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy.
[0040] Compared with the prior art, the application has the following advantages:
[0041] 1. The improved double-delay deep deterministic policy gradient algorithm is applied to the UWB ranging error correction field for the first time, and through key technologies such as a double critic network structure, a delay update mechanism and action noise regularization, the Q value overestimation problem in the traditional deep reinforcement learning algorithm is effectively solved, and the learning stability and model precision are significantly improved. The method is particularly suitable for processing noise and multipath effects in CIR data, and provides a new technical path for high-precision UWB positioning.
[0042] 2. The hierarchical extraction of time domain features is realized through a multi-level convolution structure, and the local pulse features and multipath interference modes in channel distortion are peeled off layer by layer; then a space-time feature encoding layer (Transformer encoder layer) based on a self-attention mechanism is introduced, and the long-range dependence relationship is captured through a dynamic weight distribution mechanism, thereby breaking through the limitation of the local receptive field of the traditional convolution operation. This "local-global" collaborative feature learning paradigm, while maintaining the local feature resolution, establishes a time-space correlation feature mapping across timestamps, and significantly improves the representation ability of the model to complex multipath propagation patterns.
[0043] 3. A self-supervised reinforcement learning training method without labeled data is innovatively proposed, which automatically generates iteratively improved pseudo-labels by combining trajectory predictability and extended Kalman filtering, and completely solves the core pain point of the traditional supervised learning method that relies on a large amount of labeled data. This mechanism enables the system to learn and adapt autonomously in actual deployment environments, significantly reduces deployment costs and technical barriers, and lays a foundation for the wide application of UWB positioning technology. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 A flowchart of the self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy;
[0045] Figure 2 A block diagram of the self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy;
[0046] Figure 3 A whole flowchart of the self-supervised ultra-wideband positioning system based on the improved double-delay deep deterministic policy;
[0047] Figure 4 A laboratory scene diagram;
[0048] Figure 5Intelligent robot graph. DETAILED DESCRIPTION
[0049] The principle of the present application is to apply the improved double-delay deep deterministic policy gradient algorithm to the field of UWB ranging error correction, introduce the spatio-temporal feature encoding layer based on the self-attention mechanism through key technologies such as the double critic network structure, delayed update mechanism and action noise regularization, capture long-range dependencies through the dynamic weight distribution mechanism, and establish the spatio-temporal correlation feature mapping across timestamps; effectively solve the Q value overestimation problem in the traditional deep reinforcement learning algorithm, and significantly improve the learning stability and model precision. This method is particularly suitable for processing noise and multipath effects in CIR data, and provides a new technical path for high-precision UWB positioning.
[0050] EMBODIMENT
[0051] As shown in Figure 1 A self-supervised ultra-wideband positioning method based on an improved double-delay deep deterministic policy, comprising the following steps:
[0052] Obtain channel impulse response data CIR and preprocess it;
[0053] Construct an improved double-delay deep deterministic policy model for UWB ranging error correction, which is based on a double-delay deep deterministic policy gradient model, designs a double critic network structure, introduces a spatio-temporal feature encoding layer based on the self-attention mechanism, captures long-range dependencies through the dynamic weight distribution mechanism, and establishes the spatio-temporal correlation feature mapping across timestamps; introduce the target network delay update mechanism, so that the update frequency of the target network is lower than that of the critic network and the actor network, regularize the action noise generated by the policy, and prevent the policy from using the overestimated area of the Q function;
[0054] Generate pseudo-labels using extended Kalman filtering, dynamically adjust model parameters and learning strategies for iterative optimization, and obtain the positioning result.
[0055] Another embodiment is a computer storage medium having a computer program stored thereon, wherein the computer executes the computer program to realize the self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy of any one of the above embodiments.
[0056] As shown in Figure 2 The present embodiment mainly includes the following contents:
[0057] 1) UWB signal feature analysis and preprocessing
[0058] The propagation characteristics of UWB signals under multipath and NLOS conditions are studied, and the features related to ranging error contained in the CIR data are analyzed. An effective CIR data preprocessing method is designed, including complex IQ sampling array conversion to RSSI sampling array, data cropping and normalization processing, to reduce data dimension and noise influence. The optimal preprocessing parameter settings such as sampling point number, normalization method and filtering technique are explored to provide high-quality input features for subsequent reinforcement learning models.
[0059] 2) Double-delay deep deterministic policy gradient reinforcement learning framework design
[0060] Based on the double-delay deep deterministic policy gradient algorithm, a reinforcement learning framework suitable for UWB error correction is designed, with the state space (preprocessed CIR data) and action space (ranging error correction value) defined explicitly. A double critic network structure is designed to reduce Q value overestimation and improve learning stability. A delayed update mechanism is implemented to ensure that the policy network is updated based on accurate value evaluation. A target policy smoothing technique is designed to prevent the policy from using the overestimated area of the Q function. A reward function suitable for UWB ranging error correction is constructed to effectively guide the agent to learn the correct correction behavior.
[0061] 3) Development and verification of self-supervised learning mechanism
[0062] A self-supervised learning mechanism based on trajectory predictability and filter smoothing is developed, which generates pseudo-labels using the continuity and smoothness of mobile target trajectories. An extended Kalman filter is implemented for state estimation and trajectory smoothing, effectively reducing noise influence. A circular buffer smoothing processing mechanism is designed to improve trajectory estimation accuracy through multi-frame data fusion. An iterative improvement process without labeled data is established to automatically improve the quality of pseudo-labels as learning progresses. Multi-environment adaptability experiments are designed, including different obstacle layouts and anchor node position changes, to comprehensively verify the performance of the algorithm under varying environmental conditions.
[0063] Specifically, to address the problems of high dimensionality, strong noise interference and large information redundancy of UWB raw data, an efficient data preprocessing scheme is researched and designed, including complex IQ sampling data conversion, signal feature extraction, data dimensionality reduction and denoising methods. The complex IQ sampling array is converted to an RSSI sampling array: I is the real part and Q is the imaginary part. The RSSI array is cropped to a certain number of samples, such as 150, focusing on the signal features near the first path, reducing redundant data, and thus improving model efficiency while ensuring accuracy.
[0064] The minimum-maximum normalization is performed on the cropped channel impulse response data CIR: CIRnorm = (CIR-min(CIR)) / (max(CIR)-min(CIR)). The normalized value range is compressed to [0, 1]. The normalized CIR data is subjected to noise filtering and dimension reduction processing, eliminating the influence of hardware transmission power difference and environmental distance change on the absolute signal strength, so that the model focuses on the relative characteristics such as signal-to-noise ratio (SNR) and peak shape. The sensitivity and recognition ability of the reinforcement learning algorithm to complex environmental signal characteristics are improved, laying a solid foundation for the accuracy and stability of the subsequent error correction model.
[0065] An improved double-delay deep deterministic policy model is constructed for UWB ranging error correction. The improved double-delay deep deterministic policy model is based on a double-delay deep deterministic policy gradient model, including an Actor network and a Critic network.
[0066] The Actor network includes a current Actor network and a target Actor network; the Critic network module includes two current Critic networks Q1 and Q2 and two target Critic networks. Two critic networks Q1 and Q2 are introduced, and the minimum value of the outputs of the two networks is selected when calculating the target Q value, effectively reducing the overestimation problem of Q value and improving the learning stability.
[0067] The Actor network and the Critic network have similar structures, including an input layer that receives preprocessed CIR data. The Critic network also needs to add the action a output by the Actor network as input. Then the input is input into the feature extraction layer composed of convolutional layers to capture local features of the data. Then the transformer encoder layer is input to capture the global dependency between features. Finally, through the fully connected layer, the Actor network and the Critic network output the action a and the action value Q respectively.
[0068] The Critic network module obtains the appropriate action value y = r + γ·min(Q1(s', π(s')), Q2(s', π(s'))) through mathematical expression, where r is the immediate reward, γ is the discount factor, s' is the next state, π is the policy network, Q1 and Q2 are the outputs of the two current Critic networks. This design is particularly suitable for UWB ranging error correction tasks, because the noise and multipath effect in CIR data can easily lead to uncertainty in value evaluation, and the double-critic structure can provide more conservative and stable value estimation.
[0069] Meanwhile, to ensure the stability of the training process, a target network delay update mechanism is introduced, the update frequency of the target network is lower than that of the critic network and the actor network, and the parameter is updated through a soft update method: θ'←τθ+(1-τ)θ', wherein θ is the parameter of the target network, θ' is the updated parameter, and τ is a very small constant (for example, 0.001). The delay update mechanism makes the target value more stable and reduces the oscillation in the training process, which is particularly important for the UWB error correction task which needs accurate estimation.
[0070] In addition, during the training process, the action generated by the strategy (the output of the actor network) is added with a clipping noise a' = clip (π (s') + clip (ε, -c, c), a min ,a max ), wherein clip is a clipping function, ε is a random noise, c is a noise amplitude limit, and a min ,a max are the upper and lower bounds of the action respectively. This technique prevents the strategy from using the deterministic error in the critic network, enhances the exploration ability and generalization performance, and enables the system to better adapt to different UWB signal conditions.
[0071] A spatio-temporal feature encoding layer based on a self-attention mechanism (Transformer encoder layer) is introduced to capture long-range dependencies through a dynamic weight distribution mechanism, breaking through the limitations of the local receptive field of traditional convolutional operations. This "local-global" collaborative feature learning paradigm maintains the local feature resolution while establishing a cross-time-stamp spatio-temporal correlation feature mapping, significantly improving the model's representation ability for complex multipath propagation patterns. It provides feature support with both discriminability and robustness for subsequent error correction decisions.
[0072] Specifically, the method for establishing a cross-time-stamp spatio-temporal correlation feature mapping includes:
[0073] The action generated by the strategy uses a cross-attention mechanism to mine potential temporal dependency patterns to obtain attention calculation results.
[0074] A temporal contrast model is constructed based on the attention calculation results, a strongly generalized embedding for temporal modeling is obtained through a self-supervised method, the embedding obtained by temporal modeling is taken as a temporal feature, and a graph convolution network is used for spatial modeling to obtain a cross-time-stamp spatio-temporal correlation feature mapping result.
[0075] Specifically, the cross-attention mechanism is used to mine potential temporal dependency patterns, which includes:
[0076] S11: The input of each related module is represented as H∈R L×d , wherein R is a real number, L and d are the length and dimension of the input respectively, and the initial hidden state H0 = X p, X p For single-head attention, the input sequence H is projected by three projection matrices to obtain query Q, key K and value V for the time-series feature block;
[0077] Query segment Q i and key segment K j The correlation measure c ij between them is:
[0078]
[0079] Where is the dot product operation between two matrices of the same size;
[0080] Each query segment Q i The correlation measure c i1 ,c i2 ,…,c in will be normalized by the Softmax function to obtain the aggregation weight
[0081] The output of the attention module is obtained by concatenating all outputs Y i The calculation process is:
[0082] H p = Concat(Y1,…,Y m )S12: Calculate the time-series reconstruction loss function L R :
[0083]
[0084] Where N R represents the total number of reconstructions, represents the true value of the reconstruction, represents the time-series cross-modeling result of the reconstruction;
[0085] The goal of the time-series reconstruction loss function is to minimize the overall loss.
[0086] Specifically, spatial domain modeling by graph convolution network includes:
[0087] Calculate the learned adaptive adjacency matrix A k :
[0088] A k = Softmax(Relu(X c (X c ) T ))
[0089] Where Relu is an activation function;
[0090] Get the predefined adjacency matrix As The adaptive adjacency matrix A k and A s are combined to define a multi-directional multi-graph spatial convolution, A s including a forward spatial adjacency matrix P f and a backward adjacency matrix P b The calculation process of the spatial domain modeling is as follows:
[0091]
[0092] wherein H l-1 is a hidden layer state of an l-1 layer, is a spatial domain modeling result, W s1 , W s2 and W s3 are linear transformation parameters;
[0093] The final enhanced result is obtained using a residual connection and batch normalization
[0094] The loss function L spa of the spatial domain modeling stage is as follows:
[0095]
[0096] wherein N s represents a number of points participating in the spatial domain loss calculation, is an output result of the spatial domain modeling, is a label value corresponding to the spatial domain modeling.
[0097] Based on the self-supervised learning mechanism of target motion trajectory continuity and smoothness, the system fully utilizes the data characteristics and spatio-temporal correlation of the system itself to autonomously generate high-quality pseudo labels. Combined with filtering methods such as extended Kalman filter (EKF) and circular buffer smoothing mechanism, the system performs state estimation and prediction based on historical observation data, effectively reduces noise interference in trajectory estimation, and improves trajectory prediction accuracy. Further, the system establishes an iterative optimization process without additional manual annotation data. As the model is gradually trained and iterated, the system autonomously improves the quality of pseudo labels and gradually improves system performance. At the same time, the project will also explore environmental change detection and adaptive adjustment mechanism, research how to perceive the changes of indoor environment (such as obstacle position adjustment, anchor node layout change, etc.), and dynamically adjust the model parameters and learning strategy, so that the system can still quickly adapt to environmental changes and maintain high-precision positioning performance without the need for re-collection of artificial annotation data or complete re-training.
[0098] In another embodiment, a self-supervised ultra-wideband positioning system based on an improved double-delay deep deterministic policy includes:
[0099] A data acquisition and preprocessing module acquires channel impulse response (CIR) data and performs preprocessing;
[0100] A reinforcement learning module constructs an improved double-delay deep deterministic policy model for UWB ranging error correction. The improved double-delay deep deterministic policy model is based on a double-delay deep deterministic policy gradient model, designs a double-critic network structure, introduces a spatio-temporal feature encoding layer based on a self-attention mechanism, captures long-range dependencies through a dynamic weight distribution mechanism, establishes a cross-time-stamp spatio-temporal correlation feature mapping, introduces a target network delay update mechanism, makes the update frequency of the target network lower than that of the critic network and the actor network, regularizes the action noise generated by the policy, and prevents the policy from using the overestimation area of the Q function.
[0101] A self-supervised learning module generates pseudo labels using an extended Kalman filter, dynamically adjusts model parameters and learning strategies for iterative optimization, and obtains positioning results.
[0102] In one possible embodiment, the UWB positioning system takes multi-technology cooperation as the core, constructs a complete process from data acquisition to closed-loop optimization, and realizes accurate ranging and error correction. The following detailed process is described: First, data acquisition and preprocessing are performed. UWB devices are used to acquire channel impulse response (CIR) raw data, which is the basis for subsequent processing. The collected CIR data is preprocessed, such as filtering and denoising, to regularize the data form and prepare for subsequent modules.
[0104] The preprocessed CIR data is input into the UWB ranging system, and the initial ranging related quantities, i.e., channel impulse response data (CIR) and initial ranging estimate value are output. This is used as the input of the improved double-delay deep deterministic policy model, and the error correction process is officially started.
[0105] Enter the reinforcement learning core process, using the Actor-Critic framework. The current Actor network (μ(CIR|θ)) takes CIR as input, generates actions (related to error correction quantities, i.e., a related output) based on the current Actor network parameter θ, and is represented by the formula μ(CIR|θ). Its output is used for error correction to realize the correction of the initial ranging value. The target Actor network of the target Actor network parameter is updated according to the soft update rule (τ actor is a soft update coefficient that controls the update amplitude), which provides a stable target reference for the Actor network. Its output also participates in the error correction logic of to assist the stable iteration of the reinforcement learning process.
[0106] Actor network outputs (including CIR, a and subsequent associated rewards R, etc.) are stored in the experience replay buffer ((CIR, a, R)), and then small batches of data (denoted as small batch b) are extracted by random sampling for Critic network training, breaking data correlation and improving training stability.
[0107] For Critic network training, the current Critic network takes CIR and a as input, and is based on parameters to evaluate the action value and output Q value, and its loss function is (B is the amount of small batch data, y i is the output of the target value function), i is the iteration number, and the network parameters are updated by minimizing the loss to learn the action value evaluation. The parameters of the target Critic network are updated according to soft update (τ critic is the soft update coefficient) for calculating the target value (R i is the output of the reward function, and γ is the discount factor to balance the weight of current and future rewards), which provides a supervision signal for the current Critic network. Based on the output of the Critic network, the Actor network is updated using the sampling policy gradient, and the objective function is This objective is maximized by gradient ascent to optimize the action generation policy of the Actor network, so that the error correction is more in line with the system requirements.
[0108] Then comes the trajectory processing and pseudo-label generation link. The ranging information corrected by reinforcement learning participates in trajectory prediction, and the output data are input to the extended Kalman filter processing. The extended Kalman filter formula (state prediction ), covariance prediction Kalman gain state update covariance update where f is the state transition function, h is the observation function, Q k represents the system noise covariance matrix, F k is the state transition function at time k, H k represents the Jacobian matrix of the observation function, and z k is the measurement value), the trajectory data are filtered and denoised, and output to the circular buffer for smoothing. The filtered data are smoothed by the circular buffer to generate pseudo-labels (p EKF ), which provide a reference benchmark for the system for reward calculation and model optimization closed loop.
[0109] Finally, the reward calculation and closed-loop optimization link. The pseudo-label data (p EKF ) is input into the smoothing buffer, further processed and output p avg , and the Euclidean distance conversion to range value link (using the Euclidean distance formula to convert the position information into range value Δ avg ), to obtain the range reference for reward calculation. Based on the converted range value Δ avg and the actual range related quantities of the system, the reward R is calculated by the reward function, fed back to the experience replay buffer, providing optimization signals for the reinforcement learning module, driving the Actor-Critic network to continue iteration, achieving gradual improvement of system ranging accuracy, forming a closed-loop optimization process.
[0110] The UWB ranging system process is based on UWB data acquisition and preprocessing, using the initial output of the UWB ranging system, implementing error correction through the Actor-Critic reinforcement learning framework, combining trajectory prediction, Kalman filtering, buffer smoothing to build pseudo-label and reward mechanism, forming a closed-loop process of "acquisition-preprocessing-ranging-reinforcement learning optimization-filtering smoothing-reward feedback", each link cooperates through the formula algorithm and parameter update rule, continuously iterates to improve the ranging accuracy, and adapts to the precise ranging demand in complex scenarios.
[0111] The implementation of the self-supervised learning mechanism based on the continuity and smoothness of the target motion trajectory includes the following steps:
[0112] First, the actor determines the range correction amount a, and the target actor generates the correction amount e t , and the corrected range estimate Δ t ' = Δ t -a, t is the time, and here Δ t is the original range related basic data. Then use Extended Kalman Filter (EKF) to process Δ t ', EKF is a method of state estimation by linear approximation of state equation and observation equation. In the process of converting Δ t ' to position p EKF,t =EKF(Δ t '), it first predicts the position at the current time based on the dynamic change model of "range to position" (state equation), estimates the position at the current time using the position state at the last time, and considers the influence of process noise (such as environmental disturbance, model error); Then, through the Kalman gain, the uncertainty of the prediction and the uncertainty of the observation are weighed, and the prediction result is corrected by the observation information, and finally a more accurate position is obtained. In this way, EKF maps the abstract range information into calculable and processable position coordinates, paving the way for the subsequent process, realizing the effective conversion from range to position.
[0113] After that, position p EKF,t Smoothing is performed on a circular buffer C of length N (N is an odd number, affecting the smoothing effect and adaptability to different motions; a large buffer provides high accuracy on straight paths but is prone to lag when handling maneuvers). When the buffer is full, new data (including p) is processed. EKF,t and related Δ t ′、CIR t a, CIR t (Auxiliary data) is added and the oldest data is removed, then the average position of the middle position in the buffer is calculated. Middle position index. To ensure correlation with data before and after the buffer, the average position is calculated using the following formula:
[0114] x i y i These are location coordinates, averaged to reduce noise impact.
[0115] Then take the average position p avg,m Inverse conversion to corrected range estimate Δ avg,m (by calculating to anchor point a) n (Euclidean distance implementation), reward function Quantitative correction effect, Δ′ m The range estimate for the intermediate position m after correction by the Actor network, Δ m ′ and Δ avg,m The closer the range, the higher the reward, guiding the agent to optimize the range correction (i.e., adjust 'a'). In reinforcement learning, this reward provides feedback for the agent's actions (range correction 'a'). Because updating the network with single-step samples is inefficient, batch sampling containing CIRs is performed from the experience replay buffer. m a m R m (CIR m Training with empirical data (such as training data) improves stability and avoids overfitting. The entire process forms a self-supervised closed loop of "range correction → position transformation → smoothing and noise reduction → reward generation → policy optimization", which helps the agent learn a better range correction strategy.
[0116] like Figure 4 , 5 As shown, a variety of typical indoor environmental scenarios were set up for comprehensive experimental verification and analysis to evaluate the accuracy, stability and environmental adaptability of the developed method under different environmental conditions, providing reliable technical basis and experience reference for practical applications.
[0117] The laboratory is equipped with the following experimental equipment:
[0118] Link Track-P UWB sensor several; RS-Helios-5515 32-line laser radar ×1; intelligent robot ×1; Nvidia Jetson Orin Nano ×1, NVIDIA GeForce RTX 4090. Can carry out experiments, build datasets, and train models. In addition, the laboratory also has perfect data analysis and processing software, which helps to analyze the collected data in depth.
[0119] Experiments show that compared with the traditional single-modal feature extraction scheme, the architecture improves the feature discrimination in dynamic industrial scenes by 57%, laying a solid foundation for millimeter-level error correction.
[0120] The above embodiments are preferred embodiments of the present application, but the embodiments of the present application are not limited by the above embodiments, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application should be equivalent replacement methods, and are included in the protection scope of the present application.
Claims
1. A self-supervised ultra-wideband positioning method based on an improved double-delay deep deterministic policy, characterized in that, The method comprises the following steps: obtaining channel impulse response data CIR and preprocessing; an improved double-delay deep deterministic policy model is constructed for UWB ranging error correction, the improved double-delay deep deterministic policy model is based on a double-delay deep deterministic policy gradient model, a double-critic network structure is designed, a space-time feature encoding layer based on a self-attention mechanism is introduced, a dynamic weight distribution mechanism is used to capture long-range dependencies, and a time-space correlation feature mapping across timestamps is established; a target network delay update mechanism is introduced, so that the update frequency of the target network is lower than that of the critic network and the actor network, the action noise generated by the strategy is regularized, and the overestimation area of the strategy using the Q function is prevented; an extended Kalman filter is used to generate pseudo labels, and model parameters and learning strategies are dynamically adjusted and iteratively optimized to obtain a positioning result.
2. The self-supervised ultra-wideband positioning method based on improved double-delay deep deterministic policy of claim 1, wherein, The preprocessing comprises: Complex IQ sampling is performed on the channel impulse response data CIR, and the complex IQ sampling array is converted into an RSSI sampling array: I is the real part, Q is the imaginary part, and the RSSI array is clipped to a certain number of samples, focusing on the signal characteristics near the first path. minimum-maximum normalization is performed on the cropped CIR: CIRnorm=(CIR-min(CIR)) / (max(CIR)-min(CIR)); CIRnorm is the normalized CIR data, and min(CIR) and max(CIR) are the minimum value and the maximum value of the CIR data, respectively; noise filtering and dimensionality reduction are performed on the normalized CIR data to eliminate the influence of hardware transmission power difference and environmental distance change on the absolute signal strength.
3. The modified double-delay depth-deterministic policy based self-supervised ultra-wideband positioning method according to claim 1, characterized in that, The improved double-delay deep deterministic policy model comprises an actor network and a critic network, the actor network acts as a policy generator, directly analyzes the time-domain features of the channel impulse response, and dynamically outputs an error correction quantity adapted to channel distortion; the critic network plays the role of a value evaluator, constructs an implicit quality evaluation system by analyzing the physical consistency of historical correction effects and motion trajectories, and realizes coupled optimization through self-supervised signals formed by environment feedback: the actor network generates a correction strategy based on the current channel state, the critic network evaluates the effectiveness of the strategy through a reward function derived by a trajectory smoothing module, and then guides the parameter update direction.
4. The self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy of claim 3, wherein, The actor network comprises a current actor network and a target actor network; the critic network module comprises two current critic networks Q1 and Q2 and two target critic networks, and the minimum value of the outputs of the two current critic networks is selected when the target Q value is calculated; The actor network and the critic network have similar structures, comprising an input layer that receives the preprocessed CIR data, the critic network additionally inputs the action a output by the actor network, then inputs a feature extraction layer composed of convolutional layers to capture local features of the data, then inputs a transformer encoder layer to capture global dependencies between features, and finally inputs a fully connected layer to make the actor network and the critic network output actions a and action values Q, respectively; The critic network module obtains a suitable action value y through mathematical expression: y=r+γ·min(Q1(s',π(s')),Q2(s',π(s'))) where r is the immediate reward, g is the discount factor, s' is the next state, p is the policy network, Q1 and Q2 are the outputs of the two current Critic networks.
5. The modified double-delay depth-deterministic policy based self-supervised ultra- wideband positioning method according to claim 4, characterized in that, The action noise regularization generated by the policy includes: Adding a clipping noise to the output a of the Actor network to obtain a': a' = clip(π(s') + clip(ε, -c, c), a min ,a max ) where clip() is a clipping function, ε is random noise, c is a noise amplitude limit, a min max are the upper and lower action bounds, respectively. 6. The self-supervised ultra-wideband positioning method based on improved double-delay deep deterministic policy of claim 4, wherein, The iterative optimization includes: The ranging information corrected by the reinforcement learning is processed by an extended Kalman filter, state estimation and prediction are performed in combination with historical observation data, the filtered data are smoothed by a circular buffer to generate pseudo labels; The pseudo label data are input into a smoothing buffer for processing, a ranging reference for reward calculation is obtained through a Euclidean distance conversion ranging value link, rewards are calculated based on the converted ranging value and the actual ranging related quantity by a reward function, and the rewards are fed back to an experience replay buffer to provide an optimization signal to drive continuous iteration of the improved double-delay deep deterministic policy model to form a closed-loop optimization.
7. The self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy of claim 6, wherein, The generation of the pseudo labels includes: The current Actor network determines the range correction a, the distance Δ measured at the current time t t , and the corrected range estimate Δ t ′ = Δ t -a; Then Δ t ′ is processed by extended Kalman filter t ′ to position p EKF,t = EKF(Δ t ′) where EKF() is extended Kalman filter. Firstly, the position state at the current time is predicted based on the dynamic model of "range → position", and the process noise is considered. Secondly, the predicted result is revised by the observation information through Kalman gain, and the more accurate position is obtained. Finally, the abstract range information is mapped to position coordinates, and the conversion from range to position is realized.
8. The self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy of claim 7, wherein, The continuous iteration of the improved double-delay deep deterministic policy model to form a closed-loop optimization includes: Position p EKF,t The length of the circular buffer C is N. When the buffer is full, new data is added and the oldest data is removed. The average position of the middle position of the buffer is calculated again. The index of the middle position is The average position formula is associated with the data before and after the buffer x i , y i is the position coordinate; The average position p is converted back to a revised range estimate Δ avg,m The revised range estimate Δ is converted back to a modified position p avg,m The reward function The quantified correction effect, Δ' m The revised range estimate Δ' for the intermediate position m is modified by the Actor network m , Δ avg,m The closer, the higher the reward, guiding the agent to optimize the range correction. The reward provides feedback for the agent's actions. Since single-step sample updates the network inefficiently, the experience data is trained in batches from the experience replay buffer, helping the agent learn a better range correction strategy.
9. A self-supervised ultra-wideband positioning system based on an improved double-delay depth-deterministic policy, characterized in that, It includes: A data acquisition and preprocessing module acquires channel impulse response data CIR and performs preprocessing; A reinforcement learning module constructs an improved double-delay deep deterministic policy model for UWB ranging error correction, the improved double-delay deep deterministic policy model is based on a double-delay deep deterministic policy gradient model, a double-critic network structure is designed, a spatio-temporal feature encoding layer based on a self-attention mechanism is introduced, a dynamic weight distribution mechanism is used to capture long-range dependencies, a cross-timestamp spatio-temporal correlation feature mapping is established, a target network delay update mechanism is introduced, the update frequency of the target network is lower than that of the critic network and the actor network, the action noise regularization generated by the policy is regularized, and the overestimation area of the policy using the Q function is prevented; A self-supervised learning module generates pseudo labels by using an extended Kalman filter, iteratively optimizes model parameters and learning strategies, and obtains positioning results.
10. A computer storage medium having stored thereon a computer program, characterized in that The computer executes the computer program to implement the self-supervised ultra-wideband positioning method based on the improved double-delay deep deterministic policy according to any one of claims 1-8.
Citation Information
Patent Citations
Ultra-wideband indoor positioning method based on deep attention mechanism and geometric information
CN116170746A
Transform-based diffusion diagram attention network traffic flow prediction method
CN116504060A
Numerical sensor credibility self-evaluation method and system based on motion constraint Transform
CN117634555A
UWB ranging error compensation method and device based on cascade residual attention network
CN118915036A
Unmanned aerial vehicle air combat confrontation target tracking method based on deep reinforcement learning
CN119180844A
Cited By
Method for jointly determining taxi scheduling strategy and charging station pricing strategy
CN121391344A