Fusion positioning method based on reinforcement learning and particle filtering, medium and equipment
By using a fusion positioning method combining reinforcement learning and particle filtering, the standard deviation of UWB ranging is adaptively adjusted and the base station combination is optimized. This solves the accuracy and stability problems of UWB positioning systems in complex environments, achieving a fusion positioning effect with high accuracy and high stability.
Patent Information
- Application Number
- CN202411600404.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In complex environments, the positioning accuracy and stability of UWB positioning systems are affected by multipath effects and non-line-of-sight errors. Traditional particle filter algorithms cannot adaptively adjust the ranging standard deviation, resulting in insufficient robustness of fusion positioning systems.
A reinforcement learning algorithm is introduced to adaptively select the standard deviation of UWB ranging, and the effective number of particles is optimized through a particle filter fusion positioning model to select the optimal base station combination. Pedestrian trajectory estimation is performed in conjunction with an inertial sensor module to achieve highly scalable and highly stable fusion positioning.
It improves the accuracy and robustness of the fusion positioning system, reduces the cost of base station deployment, expands the scope of application, and reduces positioning deviations caused by environmental changes.
Smart Images

Figure CN119469151B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of wireless positioning technology, and particularly relates to a fusion positioning method based on reinforcement learning and particle filtering. BACKGROUND
[0002] In indoor wireless positioning technology, UWB technology has great advantages in anti-environmental noise signal interference ability, coverage range and positioning accuracy due to its high time resolution and large bandwidth. In recent years, UWB positioning schemes are more widely used in large indoor scenes such as factories and mines, and have obvious performance advantages.
[0003] UWB technology can improve positioning accuracy to some extent relying on its own hardware characteristics, but in actual application, it will be affected by multipath effect, non-line-of-sight error, etc., causing the positioning error to increase and the positioning stability to decrease. In order to improve the stability and accuracy of the UWB positioning system, the prior art proposes a variety of optimization schemes, one is to adjust the geometric configuration of the base station layout, add a sufficient number of base station nodes, so that the UWB positioning system selects three base station nodes with the smallest ranging error to complete the three-side positioning, greatly increasing the cost of base station node layout; two is to deploy at least three UWB base station nodes on the basis of three-side positioning, and use relevant algorithms to identify and weaken non-line-of-sight measurement error to optimize the final positioning result. When the obstacle completely blocks the UWB transmission signal, the number of valid base stations received by the system is less than 3, three-side positioning cannot be performed, and the UWB positioning system will fail, so pure UWB positioning technology is difficult to meet the demand of low cost and high performance for actual indoor positioning. Three is to further improve the applicability of the UWB positioning system to the positioning scene, reduce the cost of densely laying UWB base stations, and enhance the stability of the positioning system. A PDR-assisted UWB fusion positioning method is proposed, which can perform fusion positioning through a particle filtering algorithm only with ranging information of ≤3 base stations. This method uses the high-precision ranging of UWB to correct the cumulative error of PDR in real time on one hand, and uses passive positioning of pedestrian dead reckoning to solve the continuous positioning of the UWB positioning system under the condition of less than the necessary number of base stations on the other hand. The positioning accuracy and adaptability to the environment of the fusion of the two are higher than those of a single positioning method, which effectively reduces the base station layout cost of the UWB positioning system and expands the applicability of the UWB positioning system.
[0004] When PDR is used to fuse positioning with the ranging information of two UWB base stations, the ranging standard deviation between the UWB tag and the effective base station is a key parameter of the particle filter algorithm. In the traditional particle filter, the ranging standard deviation is calibrated as a fixed value through prior experiments. However, in actual positioning, the ranging accuracy between the UWB tag and the UWB base station is affected by various factors, including but not limited to multipath effect, signal attenuation, and hardware performance fluctuation. The fixed ranging error cannot adapt to the changes in complex environments, directly affecting the selection of effective particles in the particle filter algorithm, thereby affecting the positioning accuracy and reducing the robustness of the fusion positioning system. In addition, the fusion accuracy of the ranging information of any two base stations with PDR is different. How to adaptively select the base station data with the lowest ranging error can significantly improve the performance of the fusion positioning system. Therefore, how to realize a fusion positioning method with high scalability and high stability in complex environments such as electromagnetic interference and multipath effect is a problem that needs to be overcome. SUMMARY
[0005] The present application aims to provide a fusion positioning method with high scalability and high stability in complex environments such as electromagnetic interference and multipath effect. The specific technical solutions are as follows:
[0006] The present application provides a fusion positioning method based on reinforcement learning and particle filtering, comprising the following steps:
[0007] S1: Arrange UWB base stations in a certain scene, integrate the UWB module and the inertial sensor module as a UWB tag on a moving body, give the initial position of the UWB tag, and select any two effective UWB base stations in communication with the UWB tag as UWB base station one and UWB base station two;
[0008] S2: Initialize the particle filter fusion positioning model parameters of UWB and Pedestrian Dead Reckoning (PDR), wherein the particle filter fusion positioning model parameters include particle initialization, particle state at time t, and ranging vector of the UWB tag and the two UWB base stations; the initial position of the UWB tag corresponds to the particle filter fusion positioning model parameters at time t=0;
[0009] S3: Input the particle state at time t-1 into the particle filter system state prediction model to predict the particle state at time t;
[0010] S4: Based on the particle state at time t, introduce a reinforcement learning algorithm to train the particle filter system state prediction model to obtain a trained filter prediction model that can adaptively select the UWB ranging standard deviation;
[0011] S5: According to the trained filter prediction model, adaptively select the UWB ranging standard deviation at time t to update the particle weight w at time t t, to obtain the updated particle weight w' t ;
[0012] S6: determining a resampling threshold of the particle filter, if the effective particle number N eff is less than the resampling threshold N th , performing particle resampling according to the updated particle weight distribution to obtain the resampled particle weight w" t , and entering S7; otherwise, setting w" t =w' t and directly entering S7;
[0013] S7: outputting the positioning result of the UWB tag at time t according to the particle weight w" t and the particle state at time t in S2 by using the particle filter algorithm;
[0014] S8: setting t=t+1, and filtering the optimal two UWB base station combinations as the new UWB base station one and UWB base station two based on the principle of minimum effective particle number according to the reinforcement learning, and returning to S2.
[0015] Optionally, in S2, initializing the particle filter fusion positioning model parameter based on the initial position of the UWB tag specifically includes:
[0016] S2.1, generating a uniformly distributed particle set around the initial position of the UWB tag, and constructing a particle filter fusion positioning model;
[0017] S2.2, initializing the particle filter fusion positioning model parameter of the Pedestrian Dead Reckoning (PDR);
[0018] In S3, the specific expression of the particle filter system state prediction model is:
[0019]
[0020] wherein: is the particle state at time t; H is the state transition matrix of the particle filter system state prediction model; j is the particle number; is the particle state at time t-1; B is the input matrix of the particle filter system state prediction model.
[0021] Optionally, S4 includes:
[0022] S4.1, constructing a reward function of the reinforcement learning based on the total number of effective particles in the particle set, and the specific expression is as follows:
[0023]
[0024] wherein: R[s q ][a j] is a reward function; N is the total number of particles in the particle set; s q is a particle state, s q m , m is determined by the number of positioning times of model training; a p is a standard action set in the reinforcement learning algorithm, a p n , n is determined according to the tolerance setting in the set; is the particle weight of the jth particle at time t;
[0025] S4.2, the reward of the action in each state is calculated by the reward function, and the action with the highest reward is selected as the UWB ranging standard deviation at time t.
[0026] Optionally, the determination process of the standard action set a p is specifically: determining the upper limit and the lower limit of the ranging error between any two UWB base stations and UWB tags in the dynamic environment; selecting a suitable tolerance to set an arithmetic sequence set according to the upper limit and the lower limit, and taking the arithmetic sequence set as the standard action set a p in the reinforcement learning.
[0027] Optionally, the S5 includes:
[0028] S5.1, based on the particle state at time t obtained by prediction in S2, the probability of the ranging vector O t corresponding to the particle state at time t is calculated;
[0029] S5.2, according to the probability of the ranging vector O t corresponding to the particle state at time t and the adaptively selected UWB ranging standard deviation, the ranging probability corresponding to the jth particle at time t is calculated;
[0030] S5.3, according to the ranging probability obtained in S5.2 and the particle weight at time t-1, the updated particle weight at time t is updated.
[0031] Optionally, in S5.1, the probability of the ranging vector O t corresponding to the particle state at time t is calculated The specific calculation formula is as follows:
[0032]
[0033] In S5.2, the specific calculation formula of the ranging probability corresponding to the jth particle at time t is as follows:
[0034]
[0035] The specific calculation formula of S5.3 is as follows:
[0036]
[0037] Wherein: i is the UWB base station number, i = 1, 2; is the ranging value of the UWB tag Tag at the t step and the i UWB base station, σ r is the adaptive selected UWB ranging standard deviation; is the X coordinate of the i UWB base station, is the Y coordinate of the i UWB base station; x t is the X coordinate of the UWB tag at t time, y t is the Y coordinate of the UWB tag at t time.
[0038] Optionally, in S7, the specific calculation formula of the positioning result at t time is as follows:
[0039]
[0040] Wherein: is the output particle state at t time, and the positioning result of the UWB tag at t time;
[0041] In S8, the two optimal UWB base station combinations based on reinforcement learning are selected as the UWB base station one and the UWB base station two, which are:
[0042] According to the ranging values between the UWB tag and each UWB base station, the reward functions corresponding to each UWB base station are calculated, and the reward functions corresponding to each UWB base station are sorted, and the two UWB base stations corresponding to the smallest two reward functions are selected as the new base station one and the base station two.
[0043] Optionally, in S6, the determination of the resampling threshold of the particle filter specifically includes:
[0044] The resampling threshold set is set, which is specifically:
[0045] N th ={0.2N, 0.3N, 0.4N, 0.5N, 0.6N, 0.7N, 0.8N, 0.9N, N};
[0046] According to the experimental data of the particle filter, the value corresponding to the state with the minimum average positioning error in the resampling threshold set is selected as the resampling threshold N th , and is set as the hyperparameter of the reinforcement learning algorithm part.
[0047] The application further provides a readable storage medium, which stores computer program instructions, and the computer program instructions realize the fusion positioning method based on reinforcement learning and particle filtering when executed by a processor.
[0048] The application further provides an electronic device, which comprises at least one processor, at least one memory, and computer program instructions stored in the memory, and the computer program instructions realize the fusion positioning method based on reinforcement learning and particle filtering when executed by the processor.
[0049] The application establishes the relationship between the UWB ranging error and the effective particle number of the particle filter fusion, and further establishes the relationship between the UWB ranging error and the positioning accuracy of the fusion system, and by introducing the reinforcement learning algorithm, the UWB ranging standard deviation that can make the effective particle number closest to the real condition is adaptively selected as the key parameter of the particle filter fusion in the fusion positioning, so that the accuracy and robustness of the fusion positioning system are improved.
[0050] The optimal two UWB base station combinations are screened based on the reinforcement learning, so that the adaptive selection of any two base station combinations is realized, and the accuracy of the fusion positioning system is improved.
[0051] By determining the resampling threshold of the particle filter, it is determined whether to perform resampling, so that the update of the particle weight is balanced in two extreme cases of particle degradation and uniform distribution, the optimization ability and robustness of the fusion positioning system are improved, the positioning deviation caused by the environmental change is reduced, the requirement for the UWB base station geometric shape layout is reduced, the cost is further reduced, and the application range of the fusion positioning system is further expanded.
[0052] In addition to the objects, features and advantages described above, the application has other objects, features and advantages. The application will be further described below with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0053] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application, illustrate the preferred embodiments of the application and assist in the explanation of the application. In the drawings:
[0054] Figure 1 is a technical roadmap of the fusion positioning method based on reinforcement learning and particle filtering in the embodiments of the application;
[0055] Figure 2 is a position schematic diagram of a UWB base station and a UWB tag in the embodiments of the application;
[0056] Figure 3is a positioning accuracy comparison chart of the fusion positioning method (RL-PF) based on reinforcement learning and particle filtering in the embodiments of the present application and the results of the nonlinear least squares method (NonL-LS), the extended Kalman filter method (EKF), pure UWB positioning, pure PDR positioning, and the real trajectory, i.e., ground truth data (GD). DETAILED DESCRIPTION
[0057] The embodiments of the present application are described in detail below with reference to the accompanying drawings, but the present application can be implemented in various different ways as limited and covered by the claims.
[0058] In an embodiment, as shown in Figure 1 and Figure 2 , several UWB base stations are arranged in any indoor scene, and in the experimental verification of the present application, four UWB base stations are arranged in a 6x11 indoor space, numbered A, B, C, and D. The obstacles in the indoor space are not drawn in the scene diagram, which does not affect the implementation of the algorithm described in the present application. It is intended to start from any initial position and walk in the direction shown by the dashed line in the figure. The deployment of the base station can ensure that the tag receives ranging information from not less than two UWB base stations at any position, determine a suitable resampling threshold, introduce reinforcement learning to adaptively select the ranging standard deviation, update the particle weight by optimizing the effective particle number, and improve the accuracy and robustness of the positioning system.
[0059] A fusion positioning method based on reinforcement learning and particle filtering includes the following steps:
[0060] (1) Initialize the particle filter fusion positioning model parameters, such as X0, h t ,s t ,r t 1 ,r t 2 , etc.
[0061] S1: Arrange UWB base stations in a certain scene, integrate UWB modules and inertial sensor modules as UWB tags on a moving body; give an initial position of the UWB tag, and select any two effective UWB base stations in communication with the UWB tag as UWB base station one and UWB base station two;
[0062] S2: Initialize the particle filter fusion positioning model parameters of UWB and Pedestrian Dead Reckoning (PDR), wherein: the particle filter fusion positioning model parameters include particle initialization, particle state at time t, and ranging vector of the UWB tag and two UWB base stations; the initial position of the UWB tag corresponds to the particle filter fusion positioning model parameters at time t=0 when the iteration number is 0;
[0063] In S2, initializing the particle filter fusion positioning model parameter based on the initial position of the UWB tag specifically includes:
[0064] S2.1, generating a uniformly distributed particle set around the initial position of the UWB tag to construct a particle filter fusion positioning model;
[0065] S2.2, initializing the particle filter fusion positioning model parameter of pedestrian dead reckoning (PDR), specifically:
[0066] 1) Particle initialization, generating a uniformly distributed particle set around the initial position of the tag (here, the tag refers to the integrated module of UWB and micro-inertial navigation, which is responsible for PDR positioning, and UWB is responsible for distance measurement) (here, the surrounding can be the range of the positioning area), and the initial particle set is
[0067] N is the total number of particles, and the initial weight of the jth particle is
[0068] 2) The state of the particle in the initial state is represented as: X0=[x0,y0,0,0] T , (x0,y0) represents the coordinate position of the initial state UWB tag Tag.
[0069] 3) Let the system state of the particle at the tth time be represented as: X t =[x t ,y t ,h t ,s t ] T ; wherein (x t ,y t ) is the coordinate of the UWB tag at the tth time, h t and s t are the heading and step length estimated by the pedestrian dead reckoning (PDR) on the UWB tag Tag at the tth time. The particle set P t at the tth time is: where N is the total number of particles. is the system state vector, is the weight of the jth particle at the tth time.
[0070] 4) The ranging vector O t =[r t 1 ,r t 2 ] T of the UWB tag Tag and two UWB base stations Anchor in the initial state is: t 1 and r t 2respectively represent the ranging values of the UWB tag Tag and two base stations at the t th moment;
[0071] (ii) inputting the input parameters into the particle filter prediction model to perform state prediction;
[0072] S3: inputting the particle state at the t-1 th moment into the particle filter system state prediction model to predict the particle state at the t th moment;
[0073] In S3, the specific expression of the particle filter system state prediction model is:
[0074]
[0075] wherein: is the particle state at the t th moment; H is the state transition matrix of the particle filter system state prediction model; j is the particle number; is the particle state at the t-1 th moment, and the state at the t-1 th step is (x t-1 ,y t-1 ) is the coordinate of the UWB tag Tag at the t-1 th step, h t-1 and s t-1 respectively represent the heading and step length at the t-1 th step; B is the input matrix of the particle filter system state prediction model;
[0076]
[0077] wherein: Δh represents the zero-mean Gaussian noise in the heading estimation, Δs represents the zero-mean Gaussian noise in the step length estimation, and and are disturbed by the zero-mean Gaussian noises Δh and Δs, respectively; σ h and σ s respectively represent the standard deviations of the heading h t and the step length s t .
[0078] (iii) introducing a reinforcement learning algorithm to perform model training, such as adaptively selecting σ r , the effective particle N eff ; obtaining a set resampling threshold N th , and performing particle resampling of the particle filter;
[0079] S4: based on the particle state at the t th moment, introducing a reinforcement learning algorithm to train the particle filter system state prediction model to obtain a trained filter prediction model capable of adaptively selecting the UWB ranging standard deviation;
[0080] Based on the Q-learning reinforcement learning algorithm, which does not rely on the environment to build a model, the algorithm aims to select the action with the highest reward in each state, and we use Q-learning to obtain the optimal particle weight distribution in each iteration in order to obtain better positioning accuracy.
[0081] S4 comprises:
[0082] ①First, build the reward function of Q-learning, which is a key point of reference for reinforcement learning. As known from the above subsection, the performance of the fusion positioning system is affected by the particle weight distribution, which in turn is affected by the observation vector standard deviation σ r , that is, UWB ranging, an intuitive assumption is to use the distribution of particle weights to evaluate the effect σ r . Therefore, the total number of effective particles N eff is set as the reward set.
[0083] S4.1, based on the total number of effective particles in the particle set, the reward function of reinforcement learning is constructed, and the specific expression is as follows:
[0084]
[0085] Where: R[s q ][a j ] is the reward function; N is the total number of particles in the particle set; s q is the particle state, s q ={s1,s2,...s m}, the value of m is determined by the number of positioning times of model training; a p is the standard action set in the reinforcement learning algorithm, a p ={a1,a2,...a n}, the value of n is determined according to the setting of the tolerance in the set; is the particle weight of the jth particle at time t;
[0086] ②In the reward function R[s q ][a j ] of the above ①, reinforcement learning is aimed at selecting the action with the highest reward in each state, and the idea of introducing the reinforcement learning algorithm is to select the state s q with the optimal particle weight in the action set (special component of the reinforcement learning algorithm) a p σ r .
[0087] ③a pDetermination: Fusion positioning, which integrates ranging information from dual UWB base stations and tags and extrapolates from pedestrian trajectories, is used. In dynamic environments (including real-world environments with various interferences such as line-of-sight and non-line-of-sight), the upper and lower limits of the ranging error between any two UWB base stations and tags are determined. Based on these limits, appropriate tolerances are selected and set as an arithmetic sequence set. This arithmetic sequence set is used as the standard action set a in model-free Q-learning for reinforcement learning. p ={a1,a2,...a n The value of n depends on the tolerance setting in the set. A smaller tolerance indicates higher accuracy, depending on the positioning accuracy requirements. The specific process for determining the standard action set 'a' is as follows: In a dynamic environment, determine the upper and lower limits of the ranging error between any two UWB base stations and UWB tags; based on the upper and lower limits, select appropriate tolerances to set as an arithmetic sequence set, and use this arithmetic sequence set as the standard action set 'a' in reinforcement learning. p .
[0088] ④ State s q Determination of: In the particle filter fusion algorithm, the ranging error closest to the real state is selected. The more accurate the positioning result, the better. Introducing a reinforcement learning training model allows the positioning system to adaptively select the action with the highest reward in each state. Therefore, state s q Set as the set of localization results at each step during training; the larger the training sample, the more accurate the results.
[0089] S4.2 Calculate the reward for each action in each state using a reward function, and take the action with the highest reward as the standard deviation of UWB ranging at time t.
[0090] S5: Based on the trained filtering prediction model, adaptively select the UWB ranging standard deviation at time t to update the particle weight w at time t. t The updated particle weights w' are obtained. t ;
[0091] S5 includes:
[0092] S5.1 Based on the particle state at time t predicted in S2, calculate the ranging vector O corresponding to the particle state at time t. t The probability of;
[0093] The ranging vector O corresponding to the particle state at time t t probability The specific calculation formula is as follows:
[0094]
[0095] S5.2, Based on the ranging vector O corresponding to the particle state at time t. tthe probability of the UWB ranging and the adaptive selected UWB ranging standard deviation, to calculate the ranging probability of the jth particle at time t;
[0096] the ranging probability of the jth particle at time t The specific calculation formula is as follows:
[0097]
[0098] S5.3, according to the ranging probability obtained in S5.2 and the particle weight updated at time t-1, the updated particle weight at time t is obtained,
[0099] The specific calculation formula is as follows:
[0100]
[0101] Wherein: i is the UWB base station number, i = 1, 2; r t i is the ranging value of the UWB tag Tag and the ith UWB base station at the tth step, σ r is the adaptive selected UWB ranging standard deviation; is the X coordinate of the ith UWB base station, is the Y coordinate of the ith UWB base station; x t is the X coordinate of the UWB tag at time t, y t is the Y coordinate of the UWB tag at time t.
[0102] It can be known that the weight of the particle at the tth step The influence factor of the weight of the particle at the tth step Therefore, the accurate The update of the particle weight is very important, which is a function of the UWB ranging standard deviation σ r The UWB ranging accuracy will be seriously affected by the NLOS condition, and it is difficult to accurately determine σ r In the existing particle filter, σ r is usually set as a fixed value, and the influence of the environment cannot be adaptive. This algorithm references reinforcement learning to adaptively select σ r , improves the positioning performance of the traditional particle filter, and makes it more robust to the environment.
[0103] (Four), the updated particle and weight are obtained, and the particle filter algorithm is used to output the positioning result; the optimal two UWB base station combinations are selected based on reinforcement learning to obtain the latest ranging vector; return to S1 to perform iterative cycle positioning.
[0104] S6: determining a resampling threshold of the particle filter, if the effective particle number is less than the resampling threshold, performing particle resampling according to the updated particle weight distribution to obtain a resampled particle weight w t , and entering S7; otherwise, setting w t = w' t and directly entering S7;
[0105] From the above ①, the premise of selecting the optimal effective particle number is that N eff <N th , that is, selecting a suitable resampling threshold can improve the performance of the system, if the weight of any particle at step t is uniformly distributed, that is, , then the effective particle number N eff is the maximum value N (the total number of initialized particles), if only one particle has a total weight of 1 and other particles have a weight of 0, then the effective particle number N eff is the minimum value and is 1, therefore, in order to select a suitable threshold N th to control particle resampling.
[0106] The determination of the resampling threshold of the particle filter specifically includes:
[0107] Setting a resampling threshold set, specifically:
[0108] N th ={0.2N, 0.3N, 0.4N, 0.5N, 0.6N, 0.7N, 0.8N, 0.9N, N};
[0109] According to the experimental data of the particle filter, the value corresponding to the minimum average positioning error in the state is selected as the resampling threshold N th , and is set as a hyperparameter of the reinforcement learning algorithm part.
[0110] The specific calculation process of the average positioning error is as follows:
[0111] Before Q-learning learning, an arbitrary initialized Q table is used to store the reward of each action in a state, and the learning process is based on the Bellman equation:
[0112]
[0113] Wherein: Q is a cumulative reward function for estimating the quality of state-action combination, Q(s, a) represents the prior cumulative reward of a in the state s, represents the posterior reward, R(s, a) is the immediate reward of the state s and action a, Q(s', a') is the cumulative reward of the next state s' through performing the action based on the current (s, a) combination, the variable max(Q(s', a')) represents the selected maximum achievable Q(s', a'), a is the learning rate which determines the value of the current information pair, generally set to 0.5, and g is a discount factor balancing the current and future rewards, and is set to 0.9 on the system;
[0114] In addition, in order to make the reinforcement learning training model eventually converge to obtain the minimum error value, a PE (Positioning error) is defined in the training model, which is represented by a positioning error set PE[k], wherein the value of k is consistent with N th The number of values in the set is consistent. represents the RL-PF (reinforcement learning + particle filter) positioning estimation coordinate value, (x tp ,y tp ) is the accurate coordinate of the test point.
[0115] The accuracy is evaluated by the maximum permissible error MPE (Maximum Permissible Error) and the standard deviation of the error STDE (Standard Deviation of the Error), as shown in the following table. In the positioning process, the positioning accuracy obtained by different resampling thresholds under different UWB base station combinations obtained by training is shown in Table 1. When N th is 0.7N, the positioning accuracy is the best.
[0116] Table 1 Positioning accuracy evaluation table obtained by different resampling thresholds
[0117]
[0118] S7: According to the particle weight w t and the particle state at time t in S2, the particle filter algorithm is used to output the positioning result of the UWB tag at time t;
[0119] S8: Let t = t + 1, and based on the reinforcement learning, the optimal two UWB base station combinations are selected as the new UWB base station one and UWB base station two according to the minimum effective particle number principle, and the process returns to S2.
[0120] The specific calculation formula of the positioning result at time t is as follows:
[0121]
[0122] Wherein: is the output particle state at time t, and the positioning result of the UWB tag at time t;
[0123] In S8, the two optimal UWB base station combinations are selected based on reinforcement learning as UWB base station one and UWB base station two, specifically:
[0124] The reward functions corresponding to each UWB base station are calculated based on the ranging values between the UWB tag and each UWB base station, and the reward functions corresponding to each UWB base station are sorted, and the two UWB base stations corresponding to the smallest two reward functions are selected as the new base station one and base station two.
[0125] Calculate the output effective particles
[0126]
[0127] Sort the effective particles Select the two UWB base stations corresponding to the smallest two in the set.
[0128] According to the N th value selected by reinforcement learning, the two UWB base stations (base stations C and D) with the smallest ranging error are adaptively selected to participate in the fusion positioning method based on reinforcement learning and particle filtering (RL-PF) and the nonlinear least squares method (Non-L-LS), extended Kalman filter method (EKF), pure UWB positioning, pure PDR positioning, and real trajectory, i.e., ground truth data (Ground truth Data: GD). The positioning accuracy graph is shown in Figure 3 The accuracy and robustness of the particle filter adaptive multi-source fusion algorithm based on reinforcement learning of the embodiment are the best.
[0129] The embodiment also provides a readable storage medium, characterized in that a computer program instruction is stored thereon, and the computer program instruction is executed by a processor to realize the fusion positioning method based on reinforcement learning and particle filtering
[0130] It should be noted that the apparatus embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the apparatus embodiments provided by the present application in the drawings represent that there is a communication connection between the modules, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.
[0131] The embodiment also includes an electronic device, comprising at least one processor, at least one memory, and computer program instructions stored in the memory for causing the processor to perform the fusion positioning method based on reinforcement learning and particle filtering as described above.
[0132] The computer program can be segmented into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the electronic device.
[0133] The electronic device can be a mobile phone, a desktop computer, a notebook computer, a palm computer, a cloud server, and other computing devices. The electronic device can include, but is not limited to, a processor, a memory. For example, the electronic device can also include an input / output device, a network access device, a bus, etc.
[0134] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is the control center of the electronic device, which connects all parts of the electronic device through various interfaces and lines.
[0135] The memory can be used to store the computer program and / or modules, and the processor realizes the computer program by running or executing the computer program and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.
[0136] The modules / units integrated in the electronic device can be stored in a computer readable storage medium if they are realized in the form of software function units and sold or used as independent products. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a readable storage medium. When the processor executes the computer program, the steps of the above-mentioned various method embodiments can be realized. The computer program includes computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the contents included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0137] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A fusion positioning method based on reinforcement learning and particle filtering, characterized in that, Comprising the following steps: S1: arranging UWB base stations in a certain scene, integrating UWB modules and inertial sensor modules as UWB tags on a moving body; given the initial position of the UWB tag, selecting any two effective UWB base stations in communication with the UWB tag as UWB base station one and UWB base station two; S2: initializing the particle filter fusion positioning model parameters of UWB and Pedestrian Dead Reckoning (PDR), wherein: the particle filter fusion positioning model parameters include particle initialization, particle state at time t, and ranging vectors of the UWB tag and the two UWB base stations; the initial position of the UWB tag corresponds to the particle filter fusion positioning model parameters at time t=0; S3: inputting the particle state at time t-1 into the particle filter system state prediction model to predict the particle state at time t; S4: based on the particle state at time t, introducing a reinforcement learning algorithm to train the particle filter system state prediction model to obtain a trained filter prediction model capable of adaptively selecting UWB ranging standard deviation; S5: adaptively selecting the UWB ranging standard deviation at time t according to the trained filtering prediction model to update the particle weight w at time t t , to obtain the updated particle weight w' t ; S6: determine the resampling threshold of particle filter, if the effective particle number N eff is less than the resampling threshold N th , then perform particle resampling according to the updated particle weight distribution to obtain the resampled particle weight w" t , and enter S7; otherwise, let w" t = w' t directly enter S7; S7: according to the particle weight w t and the particle state at time t in S2 adopts the particle filter algorithm to output the positioning result of the UWB tag at time t; S8: let t=t+1, and based on reinforcement learning, filter the optimal two UWB base station combinations as new UWB base station one and UWB base station two according to the minimum effective particle number principle, and return to S2.
2. The fusion positioning method based on reinforcement learning and particle filtering according to claim 1, characterized in that, In S2, initializing the particle filter fusion positioning model parameters based on the initial position of the UWB tag specifically includes: S2.1: generating a uniformly distributed particle set around the initial position of the UWB tag to construct a particle filter fusion positioning model; S2.2: initializing the particle filter fusion positioning model parameters of Pedestrian Dead Reckoning (PDR); In S3, the specific expression of the particle filter system state prediction model is: wherein: is the particle state at time t; H is the state transition matrix of the particle filter system state prediction model; j is the particle number; is the particle state at time t-1; B is the input matrix of the particle filter system state prediction model.
3. The fusion positioning method based on reinforcement learning and particle filtering according to claim 2, characterized in that, S4 includes: S4.1: constructing a reward function for reinforcement learning based on the total number of effective particles in the particle set, and the specific expression is as follows: Where: R[s] q ][a j ] represents the reward function; N represents the total number of particles in the particle set; s q In particle state, s q ={s1,s2,...s} m The value of m depends on the number of localization attempts during model training; a p For the standard action set in reinforcement learning algorithms, a p ={a1,a2,...a n The value of n depends on the tolerance setting in the set; Let be the particle weight of the j-th particle at time t; S4.2: using the reward function to calculate the reward of each action in each state, and selecting the action with the highest reward as the selected UWB ranging standard deviation at time t.
4. The fusion positioning method based on reinforcement learning and particle filtering according to claim 3, characterized in that, The standard action set a p The determination process is as follows: determining the upper limit and lower limit of the ranging error between any two UWB base stations and UWB tags in a dynamic environment; selecting a suitable tolerance according to the upper and lower limit values as an arithmetic sequence set, and taking the arithmetic sequence set as the standard action set a in reinforcement learning p .
5. The fusion positioning method based on reinforcement learning and particle filtering according to claim 4, characterized in that, S5 includes: S5.1, based on the particle state at time t predicted in S2, calculate the ranging vector O corresponding to the particle state at time t t the probability that S5.2, the ranging vector O corresponding to the particle state at time t t the probability of the ranging vector O corresponding to the particle state at time t is calculated according to the ranging vector O corresponding to the particle state at time t, the probability of the ranging standard deviation of the adaptive selected UWB, and the ranging probability corresponding to the jth particle at time t. S5.3: updating the particle weight at time t according to the ranging probability obtained in S5.2 and the particle weight at time t-1.
6. The fusion positioning method based on reinforcement learning and particle filtering according to claim 5, characterized in that, In S5.1, the ranging vector O corresponding to the particle state at time t t The specific calculation formula of the probability The specific calculation formula is as follows: In S5.2, the ranging probability corresponding to the jth particle at time t The specific calculation formula is as follows: The specific calculation formula of S5.3 is as follows: Wherein: i is the UWB base station number, i = 1, 2; is the ranging value of the tth step UWB tag and the ith UWB base station, and σ r is the adaptive selected UWB ranging standard deviation; is the X coordinate of the ith UWB base station, is the Y coordinate of the ith UWB base station; x t is the X coordinate of the UWB tag at t time, y t is the Y coordinate of the UWB tag at t time.
7. The fusion positioning method based on reinforcement learning and particle filtering according to claim 6, characterized in that, In S7, the specific calculation formula of the positioning result at time t is as follows: wherein: is the particle state at time t and the positioning result of the UWB tag at time t; In S8, filtering the optimal two UWB base station combinations as UWB base station one and UWB base station two based on reinforcement learning is specifically: Calculate the reward function corresponding to each UWB base station based on the ranging values between the UWB tag and each UWB base station, sort the reward functions corresponding to each UWB base station, and select the two UWB base stations corresponding to the smallest two reward functions as the new base station one and base station two.
8. The fusion positioning method based on reinforcement learning and particle filtering according to any one of claims 1 to 7, characterized in that, In S6, determining the resampling threshold of particle filtering specifically includes: Setting a resampling threshold set, specifically: N th = {0.2N, 0.3N, 0.4N, 0.5N, 0.6N, 0.7N, 0.8N, 0.9N, N}; According to the experimental data of the particle filter, the value corresponding to the minimum average positioning error in the resampling threshold set is selected as the resampling threshold N th , and is set as the hyperparameter of the reinforcement learning algorithm part.
9. A readable storage medium, characterized by, The computer program instructions stored thereon, when executed by a processor, implement the fusion positioning method based on reinforcement learning and particle filtering according to any one of claims 1-7.
10. An electronic device, comprising: Comprising: The at least one processor, the at least one memory, and the computer program instructions stored in the memory, when executed by the processor, perform the fusion positioning method based on reinforcement learning and particle filtering as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Pedestrian track plotting assisted UWB positioning method
CN114111802A
Robot navigation control method and system based on partial observable reinforcement learning
CN114911157A