Beidou satellite positioning method and system based on lightweight multi-agent reinforcement learning

By employing a BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning, and through dynamic pruning and hard pruning fine-tuning of deep neural networks, the real-time performance and accuracy issues of cloud-based AI models in GNSS positioning are resolved, achieving efficient and real-time satellite positioning results.

CN120928402APending Publication Date: 2025-11-11GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511066118.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing cloud-based AI-driven GNSS positioning models are insufficient to meet the stringent requirements of high-frequency real-time positioning, and suffer from data transmission latency, high network dependence, and privacy and security issues. Furthermore, existing model compression techniques exhibit significant performance degradation in GNSS positioning.

Method used

A BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning is adopted. Through dynamic mask pruning and hard pruning fine-tuning, the deep neural network model is optimized, and the model is deployed on the microcontroller of the GNSS receiver in a lightweight manner. Combined with the positioning enhancement algorithm of deep reinforcement learning and the multi-view environment perception module, the satellite positioning accuracy is improved.

Benefits of technology

It achieves efficient and real-time BeiDou satellite positioning on GNSS receivers, significantly improving positioning accuracy and stability in complex urban environments, reducing storage footprint and computing costs, and meeting the 1Hz real-time positioning requirement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928402A_ABST
    Figure CN120928402A_ABST
Patent Text Reader

Abstract

The invention discloses a Beidou satellite positioning method and system based on lightweight multi-agent reinforcement learning, and the method comprises the steps: obtaining satellite data through a positioning enhancement algorithm of deep reinforcement learning based on the measurement information of current environment observation and the state information of historical behaviors; performing dynamic pruning training on the deep neural network model, and outputting the deep neural network model after parameter rarefaction; trimming and fine-tuning the deep neural network model after parameter rarefaction to obtain a deep neural network model after hard pruning and fine-tuning; and deploying satellite data to the deep neural network model after hard pruning and fine tuning to obtain a Beidou satellite positioning result. According to the method, the positioning result precision of the satellite can be improved through dynamic mask pruning and hard pruning fine tuning after training. The Beidou satellite positioning method and system based on lightweight multi-agent reinforcement learning can be widely applied to the technical field of satellite positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite positioning technology, and in particular to a BeiDou satellite positioning method and system based on lightweight multi-agent reinforcement learning. Background Technology

[0002] Global Navigation Satellite Systems (GNSS), as an indispensable infrastructure of modern society, are widely used in transportation, surveying and mapping, agriculture, national defense, and many other fields. Navigation companies and research institutions worldwide are actively exploring the deep integration of AI and GNSS technologies. AI-based GNSS positioning technology can learn the complex characteristics of environmental noise models using measurement data, possessing anti-noise interference and adaptive learning capabilities, showing great potential in improving GNSS positioning in complex urban areas. However, the actual deployment and real-time efficient use of AI-based algorithms on positioning receivers face many challenges. For example, cloud-based AI-driven GNSS positioning models have become a research hotspot. This model utilizes the powerful computing capabilities of cloud computing to process massive amounts of GNSS data, enabling complex model training and optimization. However, this deployment method has inherent limitations: data transmission latency makes it difficult to guarantee real-time performance (especially for high-frequency real-time positioning applications at least 1Hz), high dependence on network connectivity limits its application in environments without or with weak networks, and user privacy and security issues are increasingly prominent. These factors collectively make it difficult for cloud-based AI models to meet the stringent practical requirements of high-frequency real-time positioning.

[0003] Therefore, despite the promising future of side-end AI navigation chips, there is currently a lack of mature and stable lightweight model solutions to support the application and deployment of AI+GNSS positioning models. Although model compression techniques (such as pruning, quantization, and knowledge distillation) have made significant progress in computer vision and natural language processing, their application in AI+GNSS positioning models is still in the exploratory stage. The real-time and high-precision requirements of GNSS positioning, as well as the need for robustness in dynamic and complex environments, mean that simple model compression may lead to a sharp decline in performance. Summary of the Invention

[0004] To address the aforementioned technical problems, the present invention aims to provide a BeiDou satellite positioning method and system based on lightweight multi-agent reinforcement learning, which can improve the accuracy of satellite positioning results through dynamic mask pruning and fine-tuning by hard pruning after training.

[0005] The first technical solution adopted in this invention is: a BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning, comprising the following steps:

[0006] Based on current environmental observation measurement information and historical behavior state information, satellite data is acquired through a deep reinforcement learning-based positioning enhancement algorithm;

[0007] Dynamic pruning training is performed on a deep neural network model to output a deep neural network model with sparsed parameters.

[0008] The deep neural network model with sparsified parameters is pruned and fine-tuned to obtain the deep neural network model after hard pruning and fine-tuning.

[0009] Satellite data is deployed onto a hard-pruned, finely tuned deep neural network model to obtain BeiDou satellite positioning results.

[0010] Furthermore, the step of acquiring satellite data based on measurement information from current environmental observations and state information from historical behavior, using a deep reinforcement learning-based positioning enhancement algorithm, specifically includes:

[0011] Acquire measurement information and historical behavior status information of the current environment observation. The measurement information includes the satellite pseudorange residual, line-of-sight vector, elevation angle and carrier-to-noise ratio at the current moment. The status information includes the historical behavior information of the historical trajectory in the last N time steps.

[0012] The measurement information of the current environment observation is processed by feature extraction to obtain the current environment feature observation vector;

[0013] By extracting state information of historical behavior through long short-term memory networks, important temporal information in historical observation data can be obtained.

[0014] Satellite data is obtained by correcting the current environmental feature observation vector and important temporal information in historical observation data using a near-end strategy optimization method.

[0015] Furthermore, the step of dynamically pruning the deep neural network model during training and outputting a deep neural network model with sparsified parameters specifically includes:

[0016] A method for screening neurons by pruning percentile is designed. The value sequences of neurons above the pruning percentile are retained and amplified by 1 / (1-p) times, while the value sequences of neurons below the pruning percentile are set to zero, thus generating a dynamic mask for neurons.

[0017] Determine the global pruning rate based on the cosine-rising model;

[0018] Using the global pruning rate as input, the output is the real-time pruning rate of neurons in each local layer. Combined with the neuron association parameters of each layer, the neuron pruning rate of each layer of the current deep neural network is determined.

[0019] The neuron features of the current deep neural network are input into the neuron value evaluator to obtain the value sequence of the neuron. The neuron features include the group norm of the associated weights of the neuron, the group norm of the associated weight gradients, the mask value of the neuron, and the mean activation value of the neuron.

[0020] Based on the value sequence of neurons and the pruning rate of neurons in each layer of the current deep neural network, combined with the pruning rate percentile selection method, the weights of neurons in the active state are updated by backpropagation, and the deep neural network model with sparsed parameters is output.

[0021] Furthermore, the specific expression for calculating the global pruning rate is as follows:

[0022]

[0023] In the above formula, P pops P represents the real-time global pruning rate. target The target global pruning rate is represented by `iter`, the current training iteration number is represented by `total_iter`, and the total number of training iterations is represented by `total_iter`.

[0024] Furthermore, the specific expression for calculating the neuron pruning rate of each network layer of the current deep neural network is as follows:

[0025]

[0026] In the above formula, P i M represents the pruning rate of neurons in each local layer in real time. total P represents the total number of parameters. pops D represents the real-time global pruning rate. i C represents the number of associated parameters of the remaining neurons in the current i-th layer. i This represents the number of associated parameters of neurons in the i-th layer after pruning, where n represents the number of network layers, and d represents the number of associated parameters. M This represents the ratio of the total number of parameters to the number of parameters associated with neurons.

[0027] Furthermore, the step of pruning and fine-tuning the parameter-sparsed deep neural network model to obtain the hard-pruned and fine-tuned deep neural network model specifically includes:

[0028] A pruning executor is built using the Torch-Pruning library;

[0029] Set the global pruning rate as the target pruning rate, determine the number of prunings based on the target pruning rate, perform several coarse-grained fine-tuning processes on the parameter-sparsed deep neural network model, and obtain the historical moving average of the model for each coarse-grained fine-tuning.

[0030] When the historical moving average of the coarse-grained fine-tuning model satisfies the preset descent stopping condition and the preset stability condition, the coarse-grained fine-tuned deep neural network model is output.

[0031] The coarse-grained fine-tuning deep neural network model is iteratively fine-tuned until the preset convergence condition is met, and the hard-pruned fine-tuned deep neural network model is output.

[0032] Furthermore, the preset convergence conditions specifically include the absolute value of error condition, the first reward state condition, the second reward state condition, and the value loss condition, wherein:

[0033] The absolute value error condition means that the condition is met when the absolute value of the current error is less than or equal to a preset error threshold.

[0034] The first return status condition indicates that the exponential moving average return value is greater than or equal to a preset status threshold and the current return value is greater than or equal to twice its exponential moving average, and the condition is met.

[0035] The second return state condition indicates that the exponential moving average value is greater than or equal to a preset state threshold and the current value is greater than or equal to twice its exponential moving average, and the condition is met.

[0036] The value loss condition means that the condition is met when the value loss is less than or equal to a preset value loss threshold.

[0037] Furthermore, the step of deploying satellite data onto a hard-pruned, fine-tuned deep neural network model to obtain BeiDou satellite positioning results specifically includes:

[0038] The microcontroller unit receives satellite data, collects and integrates the satellite data, and obtains the stored satellite data;

[0039] The stored satellite data is used to perform a pseudorange-based point positioning calculation task to obtain the satellite data after positioning calculation.

[0040] The satellite data after positioning calculation is format converted and preprocessed, and positioning correction is performed to obtain the initial BeiDou satellite positioning result;

[0041] The input to the hard-pruned, finely tuned deep neural network model is precisely calibrated to obtain the BeiDou satellite positioning results.

[0042] The second technical solution adopted in this invention is: a BeiDou satellite positioning system based on lightweight multi-agent reinforcement learning, comprising:

[0043] The first module is used to acquire satellite data based on measurement information from current environmental observations and state information from historical behavior, through a deep reinforcement learning-based positioning enhancement algorithm.

[0044] The second module is used to perform dynamic pruning training on the deep neural network model and output the deep neural network model with sparsified parameters.

[0045] The third module is used to prune and fine-tune the deep neural network model after parameter sparsification, so as to obtain the deep neural network model after hard pruning and fine-tuning.

[0046] The fourth module is used to deploy satellite data onto a hard-pruned and finely tuned deep neural network model to obtain BeiDou satellite positioning results.

[0047] The beneficial effects of the method and system of this invention are as follows: This invention acquires satellite data through a positioning enhancement algorithm based on measurement information from current environmental observations and state information from historical behavior, using deep reinforcement learning. This data is then used to dynamically prune and train a deep neural network model, outputting a parameter-sparsed deep neural network model. This model can synchronously iterate the global pruning rate and the dynamic pruning threshold of each network layer as the policy action network learns, thereby calculating the sparsity of each network layer. The parameter-sparsed deep neural network model is further pruned and fine-tuned to obtain a hard-pruned, fine-tuned deep neural network model. Ultimately, while achieving the global pruning goal, it ensures the smooth convergence of the policy network and the model accuracy. Finally, the satellite data is deployed to the hard-pruned, fine-tuned deep neural network model to obtain BeiDou satellite positioning results, thus improving the accuracy of satellite positioning. Attached Figure Description

[0048] Figure 1 This is a flowchart of the steps of the BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning in this invention;

[0049] Figure 2 This is a structural block diagram of the BeiDou satellite positioning system based on lightweight multi-agent reinforcement learning, as described in this invention.

[0050] Figure 3 This is a flowchart illustrating the steps of BeiDou satellite positioning provided in a specific embodiment of the present invention;

[0051] Figure 4 This is a schematic diagram of a localization enhancement algorithm based on deep reinforcement learning provided in a specific embodiment of the present invention;

[0052] Figure 5 This is a schematic diagram of the workflow design of the dynamic pruning training module provided in a specific embodiment of the present invention;

[0053] Figure 6 This is a schematic diagram of the joint optimization process design provided in a specific embodiment of the present invention;

[0054] Figure 7This is a schematic diagram of weighted data transformation provided in a specific embodiment of the present invention;

[0055] Figure 8 This is a schematic diagram of the deployment procedure design process provided in a specific embodiment of the present invention. Detailed Implementation

[0056] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.

[0057] First, it's important to note that deploying AI models directly onto the microcontroller chip (MCU) of a GNSS receiver—a process known as edge deployment—is becoming increasingly mainstream. This approach offers significant advantages, avoiding data transmission delays and potential privacy breaches, while also promising to significantly improve real-time positioning performance while reducing costs. Therefore, AI-based GNSS positioning technology can learn the complex characteristics of environmental noise models using measurement data, exhibiting noise resistance and adaptive learning capabilities. This demonstrates immense potential for improving GNSS positioning in complex urban areas.

[0058] However, existing technologies have the following problems:

[0059] 1) Pruning typically relies on importance criteria (such as the L1 norm and L2 norm of weights) to determine which channels or parameters should be pruned. Although this method can handle general structural pruning problems for any architecture such as CNN, RNN, GNN, and Transformer, the performance loss of the model after structured pruning is significant due to the simplistic evaluation system.

[0060] 2) While the binary masking scheme improves the model's generalization performance and accuracy during training, it does not fundamentally change the model's topology and fails to achieve the goals of reducing storage usage and accelerating inference.

[0061] 3) Although the pruning strategy is gradual, the gradual approach is flawed. The mathematical model has an excessively high pruning rate in the early stages, which can easily remove important neurons prematurely. Furthermore, the training becomes increasingly unstable as the target pruning rate increases. Additionally, the granularity of the pruning rate is not fine enough, and it fails to analyze the specific redundancy of parameters at each layer.

[0062] 4) Unstructured pruning techniques were applied to the DRL model, and a high compression ratio was achieved through iterative hard threshold pruning. However, the irregular and sparse weight matrix structure generated by this technique is difficult to utilize the hardware acceleration capabilities of general-purpose platforms, which significantly limits its deployment efficiency on practical inference platforms.

[0063] Based on this, this invention proposes a model compression method based on an AI+GNSS BeiDou high-precision positioning algorithm, aiming to achieve model lightweighting while maintaining positioning accuracy. For existing platforms and AI+GNSS algorithms, an innovative and efficient model lightweighting scheme is proposed and validated. Specifically, an adaptive KF algorithm based on multi-agent reinforcement learning is used, employing structured pruning and efficient training process control strategies to achieve stable model training and maintain key positioning accuracy. Structured pruning directly alters the neural network structure by physically removing grouping parameters (such as convolutional kernels, layers, or channels), offering significant advantages in reducing storage consumption and computational costs, and is widely applicable as it does not rely on specific hardware or software acceleration. The innovative and efficient training process control, through a series of optimization designs, ensures stable model training and maintains key positioning accuracy.

[0064] First, it needs to be explained that, as Figure 3 As shown, this invention provides a model compression method based on an AI+GNSS BeiDou high-precision positioning algorithm, aiming to achieve model lightweighting while maintaining positioning accuracy. The invention mainly comprises two parts: first, the design of a system for building and pruning a positioning enhancement algorithm based on multi-agent deep reinforcement learning; and second, the design of a joint optimization process for training, evaluation, and deployment.

[0065] Reference Figure 1 This invention provides a BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning, which includes the following steps:

[0066] S100: Based on measurement information from current environmental observations and state information from historical behavior, satellite data is acquired through a deep reinforcement learning-based positioning enhancement algorithm.

[0067] Specifically, the process involves acquiring measurement information from current environmental observations and state information from historical behavior. The measurement information includes the satellite pseudorange residual, line-of-sight vector, elevation angle, and carrier-to-noise ratio at the current moment. The state information includes historical behavior information from the last N time steps of the historical trajectory. Feature extraction processing is performed on the measurement information from current environmental observations to obtain the current environmental feature observation vector. Data extraction of the state information from historical behavior is performed using a long short-term memory network to obtain important temporal information from the historical observation data. The current environmental feature observation vector and the important temporal information from the historical observation data are corrected using a near-end strategy optimization method to obtain satellite data.

[0068] In this embodiment, the localization enhancement algorithm is developed based on the traditional Kalman filter localization algorithm. It uses the localization algorithm calculated by the traditional Kalman filter as the initial point, and through feature extraction as input and the output of the deep reinforcement learning model, the localization correction result is obtained, achieving the localization enhancement effect. We first built a localization enhancement algorithm based on deep reinforcement learning on a PC using the PyTorch framework. The deep reinforcement learning-based localization enhancement algorithm mainly consists of a multi-view environment perception module and a localization correction strategy learning module, such as... Figure 4 As shown.

[0069] To reduce the impact of perception differences on noise modeling in complex urban environments, a multi-view environmental perception module was designed to improve the model's ability to perceive the current environment and its data utilization. We considered measurement information from current environmental observations and state information from historical behavior. Current measurement information includes a feature observation vector composed of the current satellite pseudorange residual RES, line-of-sight vector LOS, elevation angle Ei, and carrier-to-noise ratio C / N0. Historical state information includes historical behavior information from the last N time steps. A feature extraction module was constructed to extract current environmental measurement information. A Long Short-Term Memory (LSTM) module was used to extract important temporal information from historical observation data; a multilayer perceptron was constructed to output measurement information. Behavioral information was output through the LSTM network, and measurement information was output through the multilayer perceptron network. Finally, the measurement information was input into a positioning correction strategy learning module to learn a positioning correction strategy.

[0070] It should be noted that a localization correction policy learning module was constructed based on Proximal Policy Optimization (PPO). PPO consists of an actor network and an evaluator network. The actor network outputs the agent's actions, while the evaluator network approximates the value of the current state and guides the actor network's learning. By combining importance sampling and pruning the objective function, the magnitude of policy updates is limited, thereby reducing the interference of outlier data points on the model's policy training. This allows the model to learn a more stable correction policy.

[0071] S200: Perform dynamic pruning training on the deep neural network model and output the deep neural network model with sparsified parameters.

[0072] Specifically, a pruning percentile neuron selection method is designed, which retains and amplifies the value sequences of neurons above the pruning percentile by 1 / (1-p) times, and sets the value sequences of neurons below the pruning percentile to zero, generating a dynamic neuron mask. The global pruning rate is determined based on a cosine ascending model. Using the global pruning rate as input, the real-time pruning rates of neurons in each local layer are output. Combined with the neuron association parameters of each layer, the neuron pruning rates of each network layer in the current deep neural network are determined. The neuron features of the current deep neural network are input to a neuron value evaluator to obtain the neuron value sequence. These neuron features include the group norm of the neuron's association weights, the group norm of the association weight gradients, the neuron's mask value, and the mean of the neuron's activation values. Based on the neuron value sequence and the neuron pruning rates of each network layer in the current deep neural network, combined with the pruning percentile neuron selection method, the weights of activated neurons are updated via backpropagation, and the parameter-sparsed deep neural network model is output.

[0073] In this embodiment, the dynamic training module is integrated into the deep reinforcement learning model. It achieves parameter sparsity by dynamically updating the mask in real time to zero out neurons during training. The output model with sparsified parameters is then input into the hard pruning fine-tuning module. The dynamic training module includes three core functions: dynamic mask generation, neuron value evaluation, and iterative calculation of the dynamic pruning rate. Figure 5 As shown.

[0074] 1) Dynamic mask generation;

[0075] A mask matrix is ​​generated based on the current DNN structure. A value of 1 indicates that the neuron value is retained during the current training process, while 0 indicates the opposite. This dynamic mask is positioned after the activation values ​​of each layer in the model inference process, similar to dropout in regularization methods. Its main function is to prevent overfitting and achieve sparsity of model parameters during backpropagation. This is based on the current pruning rate P of the network layers. i The neuron value sequence S value As an importance criterion, the pruning rate P i Weights above the percentile are retained and amplified by a factor of 1 / (1-p) (Inverted Dropout) to maintain the total expected value of the layer's output; weights below the percentile are set to 0, and during backpropagation, only the weights of neurons in the active state are updated. This module outputs a dynamic mask for neurons, containing a set of masks for each layer of neurons {mask1, mask2, ... mask}. n}∈Mask. This module takes the current value sequence of neurons in each layer as input and outputs the dynamic mask of the neurons.

[0076] 2) Iterative calculation of layer sparsity;

[0077] Traditional non-progressive pruning rates cause significant training fluctuations during deep reinforcement learning model training, failing to balance compression effectiveness and model performance. This invention designs a progressive iterative pruning algorithm that gradually increases the global pruning rate over time and allocates local pruning rates to each network layer based on the number of parameters, thus balancing compression effectiveness and model performance.

[0078] The mathematical model for the global pruning rate is designed to be cosine increasing, with the aim of making the global pruning rate more linear during training, thereby enabling the model to learn a more stable correction strategy.

[0079]

[0080] Among them, P pops P represents the real-time global pruning rate. target The target global pruning rate is 0.5, meaning a pruning rate of 0.5 is applied, removing 50% of the global parameters. `iter` represents the current training iteration number, and `total_iter` represents the total number of training iterations.

[0081] The pruning rate of each local network layer is designed based on different parameter allocations. The neuron association parameters of each layer are observed; the association parameter quantity is defined as follows:

[0082] C i =M i-1 *M i +M i *M i+1

[0083] Among them, C i M represents the number of parameters associated with neurons in the i-th layer; i This represents the output dimension of the i-th layer of the network, for example, M. i-1 Let represent the output dimension of the (i-1)th layer of the network. The mathematical model for the local network pruning rate is designed with the real-time global pruning rate P as the input. pops Output the real-time pruning rate P of neurons in each local layer. i .

[0084] Other definitions are as follows:

[0085]

[0086] In the above formula, M total is the total number of parameters, n is the number of network layers, in is the dimension of the input layer, and out is the dimension of the output layer.

[0087]

[0088] ΔC i =P i *(Ci )

[0089]

[0090] D i =C i *mask i

[0091]

[0092] Where, d M ΔC is the scaling factor between the total number of parameters and the number of neuron-related parameters. i D represents the change in the number of association parameters of neurons in the i-th layer after pruning. i This represents the number of neuron association parameters remaining in the current i-th layer. This indicates that the pruning rate of neurons in the current layer is related to the amount of remaining parameters in that layer. Combining the above expressions, the formula for calculating the pruning rate of neurons in each layer is obtained as follows:

[0093]

[0094] 3) Neuron value assessment;

[0095] Pruning typically relies on importance criteria (such as the L1 norm and L2 norm of weights) to determine which channels or parameters should be pruned. Some neurons in a DNN are redundant to the overall architecture; therefore, evaluating the value of neurons in the current work to set importance criteria is crucial for pruning. This invention integrates four key features of the current neuron into a neuron value evaluator (NVE) to learn and output a neuron value sequence S. value Features include the group norm of the neuron's association weights. Group norm of associated weight gradient The neuron's mask value i The mean activation value of this neuron

[0096] The specific definitions are as follows:

[0097]

[0098] in, Output neuron value sequence S value L refers to a linear layer, and GELU refers to a nonlinear activation function. Gaussian Error Linear Unit is an improved smoothed ReLU activation function. The pruning rate P of the network layer is then used. i Using the neuron value sequence as the importance criterion, the pruning rate P iFor values ​​above the percentile, the weights are retained and amplified by a factor of 1 / (1-p) (Inverted Dropout) to maintain the total expected value of the layer's output; values ​​below the percentile are set to 0, and only the weights of neurons in the active state are updated during backpropagation. The input features of this module include the magnitude of the associated weights of the neuron, the magnitude of the associated gradient, the mask value (activation state), the magnitude of the activation value, and the output neuron value.

[0099] S300. Prune and fine-tune the deep neural network model after parameter sparsification to obtain the hard-pruned and fine-tuned deep neural network model.

[0100] Specifically, the Torch-Pruning library is used to construct a pruning executor; the global pruning rate is set as the target pruning rate, and the number of pruning operations is determined according to the target pruning rate. Several coarse-grained fine-tuning operations are performed on the parameter-sparsed deep neural network model, and the historical moving average of the model is obtained for each coarse-grained fine-tuning. When the historical moving average of the coarse-grained fine-tuning model meets the preset descent stopping condition and the preset stability condition, the coarse-grained fine-tuned deep neural network model is output; the coarse-grained fine-tuned deep neural network model is iteratively fine-tuned until the preset convergence condition is met, and the hard-pruned fine-tuned deep neural network model is output.

[0101] In this embodiment, to ensure the model meets the resource and computing power limitations of the MCU system, a simplified model capable of real-time inference needs to be deployed on the MCU. The hard pruning and fine-tuning module prunes and fine-tunes the parameter-sparsed model trained through dynamic pruning. Unlike dynamic pruning, this module alters the model's topology, reducing storage footprint and accelerating inference. The hard pruning and fine-tuning module includes the construction of the pruning executor and the design of the fine-tuning strategy.

[0102] 1) Construction of the pruning actuator;

[0103] We use the Torch-Pruning library as a support library for building pruning executors. Torch-Pruning (TP) is a structured pruning framework that supports structured pruning of various deep neural networks. Unlike torch.nn.utils.prune, which sets parameters to zero using masks, Torch-Pruning utilizes a dynamic computation graph to group and remove coupled parameters. By building pruning executors, the number of pruning parameters is fine-tuned according to the pruning strategy.

[0104] 2) Design of fine-tuning strategies;

[0105] The goal of the fine-tuning strategy is to maximize the retention of accuracy after pruning. Specifically, this involves retraining the model a certain number of times after each parameter reduction step by the pruning executor. The strategy consists of three steps: setting the number of pruning steps and the number of basic fine-tuning steps based on the target pruning rate; determining the initial convergence condition after completing the basic coarse-grained fine-tuning; and finally, fine-tuning to achieve optimal performance.

[0106] First, set the global pruning rate P of the dynamic pruning system. target To achieve the target pruning rate, increase the number of pruning cycles based on different target pruning rates, for example, P. target =0.5, then prune in 5 stages, with each stage involving 10 coarse-grained fine-tuning adjustments. P target =0.9, then it is trimmed in 9 stages, with each stage involving 10 coarse-grained fine-tuning adjustments.

[0107] Collect the historical moving average r of the model returns from 10 coarse-grained fine-tuning iterations. ema As an indicator. When r ema If both the performance degradation condition and the stability condition are met, the model is considered to have met the preliminary convergence condition. Here, "preliminary" means that the model may have reached a relatively stable state, but further fine-grained adjustments may still be needed. The judgment conditions are mainly divided into two categories:

[0108] Decline Stopping Conditions: This set of conditions primarily focuses on whether the exponential moving average (EMA) of value and return in the PPO algorithm shows a downward trend. This condition is true when the EMA of the current batch is less than or equal to the EMA of the previous batch. This indicates that the value estimate may be stabilizing or has begun to decline.

[0109] Stability Conditions: These conditions focus on the stability and policy performance during model training. This condition is true when the EMA standard deviation of the current batch's value is less than or equal to a preset standard deviation threshold. This indicates that the model's value estimation has become sufficiently stable with low volatility.

[0110] The combination of these conditions helps to identify early in the training process whether the model has entered a relatively stable phase, thus allowing for consideration of fine-tuning or other optimizations in the next phase.

[0111] Once the initial convergence conditions are met, the model is fine-tuned. Each fine-tuning check verifies whether the absolute model performance metrics are met. If they are met, the accuracy requirements are satisfied, and the fine-tuned model weight file is output. The convergence criteria for fine-tuning are based on the following four core metrics to determine whether convergence has been achieved.

[0112] Error absolute value condition: This condition is met when the absolute value of the current error is less than or equal to a preset error threshold. This ensures that the model's predictions on the primary objective are sufficiently accurate.

[0113] Return State Conditions: These conditions consist of two parts. First, the exponentially moving average return must be greater than or equal to a preset state threshold, indicating that the model has accumulated sufficient return experience. Second, the current return must be greater than or equal to twice its exponentially moving average, suggesting a significant improvement in the current model performance.

[0114] Value State Condition: Similar to the return state condition, this condition also consists of two parts. First, the exponentially moving average value must be greater than or equal to a preset state threshold, indicating that the model's estimation of future value has reached a certain level. Second, the current value must be greater than or equal to twice its exponentially moving average, reflecting a significant improvement in the current model's value prediction.

[0115] Value loss condition: This condition is met when the value loss is less than or equal to a preset value loss threshold. This ensures that the model's learning in value assessment has become stable.

[0116] The system will only determine that fine-grained tuning has converged when all of the above conditions are met simultaneously. This means that the model has reached the expected optimization level in multiple dimensions, including error, reward, value assessment, and value loss.

[0117] S400 deploys satellite data onto a hard-pruned, finely tuned deep neural network model to obtain BeiDou satellite positioning results.

[0118] Specifically, the microcontroller unit receives satellite data, collects and integrates the satellite data to obtain stored satellite data; it then performs a pseudorange-based single-point positioning calculation task on the stored satellite data to obtain the positioned satellite data; it performs format conversion and preprocessing on the positioned satellite data, and performs positioning correction to obtain the initial BeiDou satellite positioning result; finally, it inputs the hard-pruned and finely tuned deep neural network model for precise correction to obtain the BeiDou satellite positioning result.

[0119] In this embodiment, the joint optimization process includes three steps: training, evaluation, and deployment. Specifically, this invention integrates a dynamic pruning module into the training process to obtain a parameter-sparse model weight file, which is then input into a hard pruning fine-tuning module for optimization to obtain a high-performance and lightweight model. Finally, the model is deployed to the vehicle-mounted MCU platform. The joint optimization process is designed as follows: Figure 6 As shown.

[0120] Training process design: At the start of training, a dynamic mask (MASK) initially set to all 1s is generated; in the i-th iteration of the reinforcement learning PPO algorithm, the sequence of neuron value sequences is collected from the input neuron value evaluator (NVE); the dynamic global pruning rate P is calculated based on the current iteration number. pops Calculate the pruning rate P of neurons in each i-th layer. i Take the current pruning rate P of neurons in each layer. i Percentiles are used as a threshold to update the dynamic mask. Applying the mask, neurons with values ​​less than the threshold are set to zero to complete the current dynamic pruning operation. This process is repeated iteratively until training ends. After training, a sparse weight model file is obtained, which is then input into the hard pruning fine-tuning module. After optimization, a high-performance and lightweight model is obtained.

[0121] Furthermore, we collect datasets for testing. We divide the collected datasets into training and testing sets. After the model undergoes a training process on the training set, a model parameter file is obtained. The performance of the model is then judged based on the test results on the testing set. Once the expected performance is achieved, we can determine which model parameters will be used for subsequent deployment.

[0122] Therefore, TensorFlow Lite for Micro was adopted as the support library for deploying deep learning on the MCU. TensorFlow Lite for Micro is a deep learning library that supports MCUs; it is not limited to any processor vendor and supports various processor architectures. Furthermore, by reducing the model and binary file size based on TensorFlow Lite, it further enables MCUs to run deep learning models. The model conversion process is as follows: Figure 7 As shown.

[0123] In the MCU, call TFLM library functions to build an isomorphic model to the previously trained model, and run the model for inference. The MCU program flow design is as follows: Figure 8 As shown. This invention employs an RTOS (Real-Time Operating System) (taking FreeRTOS as an example), which is particularly suitable for fields with high real-time requirements, such as positioning and correction. The processing task flow forms a directed acyclic graph, and the correct operation of the program can be controlled through a message queue.

[0124] Upon receiving satellite data input from the front-end processing chip, the microcontroller unit (MCU) immediately initiates data processing task one. The core function of this task is to efficiently collect and integrate the received raw satellite data and securely store it in a global shared memory area to ensure data accessibility and consistency. After data normalization, the system sends processing instructions to the pseudorange-based point positioning calculation task via message queue 1. Upon receiving the signal from message queue 1, this calculation task immediately begins its positioning calculation process.

[0125] Subsequently, data processing task two receives and processes data from message queues 2 and 3. The key responsibility of this task is to convert and preprocess the localization data to conform to the input format requirements of the Deep Reinforcement Learning (DRL) model for localization correction. The formatted data is then transmitted to the localization correction task via message queue 4. The localization correction task, based on the initial localization results, uses the DRL model to precisely correct the location information, ultimately outputting optimized and calibrated high-precision location information.

[0126] In summary, this invention proposes a model compression method based on multi-agent reinforcement learning for neuron value learning and dynamic pruning, aiming to achieve model lightweighting while maintaining accuracy. Specifically, for unreasonable pruning strategies, during PPO training, a global pruning rate is set that grows synchronously with the learning of the policy action network. Dynamic pruning thresholds for each network layer are calculated synchronously based on the sparsity of each layer, ultimately achieving the global target pruning rate while ensuring model performance and smooth convergence of the policy network. For a single importance criterion, a sub-network based on weight magnitude, gradient magnitude, activation value, and activation state is constructed to learn neuron values. The learned neuron values ​​can more accurately assess the decision contribution of each neuron. For a single training optimization strategy, a joint model optimization method is designed, combining dynamic masked pruning training and hard pruning fine-tuning strategies, improving computational efficiency while maximizing network accuracy. This invention fills the gap in the design of neuron value importance criteria.

[0127] Therefore, the present invention differs from the prior art in the following ways:

[0128] 1) This invention presents a novel dynamic pruning strategy that iteratively calculates the sparsity of each network layer by synchronously iterating the global pruning rate and the dynamic pruning threshold of each network layer as the policy action network learns. Ultimately, this achieves the global pruning objective while ensuring the smooth convergence of the policy network and the accuracy of the model.

[0129] 2) This invention presents a novel neuron value learning method that overcomes the limitations of traditional single-dimensional evaluation thresholds. The learned neuron values ​​can more accurately assess the decision contribution of each neuron.

[0130] 3) The joint model optimization method designed in this embodiment combines dynamic mask pruning during training with hard pruning fine-tuning after training, which maximizes network accuracy while reducing storage usage and accelerating inference.

[0131] Compared with the prior art, the present invention has the following advantages:

[0132] 1) This paper proposes and implements a stable, lightweight model design scheme for AI+GNSS algorithms. Considering the characteristics of existing GNSS hardware platforms and AI+GNSS algorithms, this research focuses on the Adaptive Kalman Filter (AKF) algorithm based on multi-agent reinforcement learning. We will delve into and apply advanced model compression techniques, particularly structured pruning and efficient flow control strategies, to achieve stable model training and maintain key positioning accuracy. Through meticulous design, the embodiments of this invention aim to ensure that while significantly reducing the number of model parameters, its positioning correction capability remains unaffected or even optimized.

[0133] 2) Achieving ultra-low storage footprint and high real-time inference performance: Based on the proposed lightweight scheme, this embodiment of the invention successfully deploys the AI+GNSS model onto the microcontroller chip of a GNSS receiver. We anticipate achieving an extremely low storage footprint of less than 2MB, which will greatly expand the application scope of the AI+GNSS model on resource-constrained edge devices. Simultaneously, the model's inference speed on the microcontroller will be consistently below 20ms, fully meeting the stringent requirements of 1Hz offline real-time positioning enhancement, providing a hardware foundation for high-frequency, low-latency real-time positioning applications.

[0134] 3) Significantly improves positioning accuracy and stability in complex urban scenarios. Through rigorous testing and verification, this invention aims to demonstrate that, compared to traditional standard GNSS receivers, the AI+GNSS model of this invention will improve positioning accuracy and stability by approximately 32.0% in different urban scenarios. This significant improvement will effectively address interference from complex environments such as urban canyon effects and multipath effects on GNSS signals, thereby providing more reliable and higher-precision positioning services and bringing revolutionary breakthroughs to applications such as intelligent driving, urban management, and personal navigation.

[0135] Reference Figure 2 The BeiDou satellite positioning system based on lightweight multi-agent reinforcement learning includes:

[0136] The first module 201 is used to acquire satellite data based on measurement information from current environmental observations and state information from historical behavior, through a deep reinforcement learning-based positioning enhancement algorithm.

[0137] The second module 202 is used to perform dynamic pruning training on the deep neural network model and output the deep neural network model with sparsified parameters.

[0138] The third module 203 is used to prune and fine-tune the deep neural network model after parameter sparsification to obtain the deep neural network model after hard pruning and fine-tuning.

[0139] Module 4, 204, is used to deploy satellite data onto a hard-pruned, finely tuned deep neural network model to obtain BeiDou satellite positioning results.

[0140] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0141] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning, characterized in that, Includes the following steps: Based on current environmental observation measurement information and historical behavior state information, satellite data is acquired through a deep reinforcement learning-based positioning enhancement algorithm; Dynamic pruning training is performed on a deep neural network model to output a deep neural network model with sparsed parameters. The deep neural network model with sparsified parameters is pruned and fine-tuned to obtain the deep neural network model after hard pruning and fine-tuning. Satellite data is deployed onto a hard-pruned, finely tuned deep neural network model to obtain BeiDou satellite positioning results.

2. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 1, characterized in that, The step of acquiring satellite data based on measurement information from current environmental observations and state information from historical behavior, using a deep reinforcement learning-based positioning enhancement algorithm, specifically includes: Acquire measurement information and historical behavior status information of the current environment observation. The measurement information includes the satellite pseudorange residual, line-of-sight vector, elevation angle and carrier-to-noise ratio at the current moment. The status information includes the historical behavior information of the historical trajectory in the last N time steps. The measurement information of the current environment observation is processed by feature extraction to obtain the current environment feature observation vector; By extracting state information of historical behavior through long short-term memory networks, important temporal information in historical observation data can be obtained. Satellite data is obtained by correcting the current environmental feature observation vector and important temporal information in historical observation data using a near-end strategy optimization method.

3. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 2, characterized in that, The step of dynamically pruning the deep neural network model and outputting a deep neural network model with sparsed parameters specifically includes: A method for screening neurons by pruning percentile is designed. The value sequences of neurons above the pruning percentile are retained and amplified by 1 / (1-p) times, while the value sequences of neurons below the pruning percentile are set to zero, thus generating a dynamic mask for neurons. Determine the global pruning rate based on the cosine-rising model; Using the global pruning rate as input, the output is the real-time pruning rate of neurons in each local layer. Combined with the neuron association parameters of each layer, the neuron pruning rate of each layer of the current deep neural network is determined. The neuron features of the current deep neural network are input into the neuron value evaluator to obtain the value sequence of the neuron. The neuron features include the group norm of the associated weights of the neuron, the group norm of the associated weight gradients, the mask value of the neuron, and the mean activation value of the neuron. Based on the value sequence of neurons and the pruning rate of neurons in each layer of the current deep neural network, combined with the pruning rate percentile selection method, the weights of neurons in the active state are updated by backpropagation, and the deep neural network model with sparsed parameters is output.

4. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 3, characterized in that, The specific expression for calculating the global pruning rate is as follows: In the above formula, P pops P represents the real-time global pruning rate. target The target global pruning rate is represented by `iter`, the current training iteration number is represented by `total_iter`, and the total number of training iterations is represented by `total_iter`.

5. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 4, characterized in that, The specific expressions for calculating the neuron pruning rate of each layer of the current deep neural network are as follows: In the above formula, P i M represents the pruning rate of neurons in each local layer in real time. total P represents the total number of parameters. pops D represents the real-time global pruning rate. i C represents the number of associated parameters of the remaining neurons in the current i-th layer. i This represents the number of associated parameters of neurons in the i-th layer after pruning, where n represents the number of network layers, and d represents the number of associated parameters. M This represents the ratio of the total number of parameters to the number of parameters associated with neurons.

6. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 5, characterized in that, The step of pruning and fine-tuning the parameter-sparsed deep neural network model to obtain the hard-pruned and fine-tuned deep neural network model specifically includes: A pruning executor is built using the Torch-Pruning library; Set the global pruning rate as the target pruning rate, determine the number of prunings based on the target pruning rate, perform several coarse-grained fine-tuning processes on the parameter-sparsed deep neural network model, and obtain the historical moving average of the model for each coarse-grained fine-tuning. When the historical moving average of the coarse-grained fine-tuning model satisfies the preset descent stopping condition and the preset stability condition, the coarse-grained fine-tuned deep neural network model is output. The coarse-grained fine-tuning deep neural network model is iteratively fine-tuned until the preset convergence condition is met, and the hard-pruned fine-tuned deep neural network model is output.

7. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 6, characterized in that, The preset convergence conditions specifically include the absolute value of error condition, the first reward state condition, the second reward state condition, and the value loss condition, wherein: The absolute value error condition means that the condition is met when the absolute value of the current error is less than or equal to a preset error threshold. The first return status condition indicates that the exponential moving average return value is greater than or equal to a preset status threshold and the current return value is greater than or equal to twice its exponential moving average, and the condition is met. The second return state condition indicates that the exponential moving average value is greater than or equal to a preset state threshold and the current value is greater than or equal to twice its exponential moving average, and the condition is met. The value loss condition means that the condition is met when the value loss is less than or equal to a preset value loss threshold.

8. The BeiDou satellite positioning method based on lightweight multi-agent reinforcement learning according to claim 7, characterized in that, The step of deploying satellite data onto a hard-pruned, fine-tuned deep neural network model to obtain BeiDou satellite positioning results specifically includes: The microcontroller unit receives satellite data, collects and integrates the satellite data, and obtains the stored satellite data; The stored satellite data is used to perform a pseudorange-based point positioning calculation task to obtain the satellite data after positioning calculation. The satellite data after positioning calculation is format converted and preprocessed, and positioning correction is performed to obtain the initial BeiDou satellite positioning result; The input to the hard-pruned, finely tuned deep neural network model is precisely calibrated to obtain the BeiDou satellite positioning results.

9. A BeiDou satellite positioning system based on lightweight multi-agent reinforcement learning, characterized in that, Includes the following modules: The first module is used to acquire satellite data based on measurement information from current environmental observations and state information from historical behavior, through a deep reinforcement learning-based positioning enhancement algorithm. The second module is used to perform dynamic pruning training on the deep neural network model and output the deep neural network model with sparsified parameters. The third module is used to prune and fine-tune the deep neural network model after parameter sparsification, so as to obtain the deep neural network model after hard pruning and fine-tuning. The fourth module is used to deploy satellite data onto a hard-pruned and finely tuned deep neural network model to obtain BeiDou satellite positioning results.

Citation Information

Cited By

  • LSTM reasoning method based on channel cutting and gating configuration and electronic equipment

    CN121212207A