Multi-source data fusion combined positioning system and method based on neural network
The neural network-based multi-source data fusion system addresses GPS accuracy issues in urban canyons and indoors by integrating GPS, IMU, and WiFi sensors with reinforcement learning, achieving enhanced positioning precision and robustness.
Patent Information
- Application Number
- CN202510382261.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing technology has insufficient positioning accuracy in urban canyons or indoor environments. Traditional multi-source data fusion methods have failed to fully utilize the complementarity of sensor data, and rely on artificial design fusion rules to lack adaptive optimization capabilities and cannot adapt to dynamic environments.
A multi-source data fusion combined positioning system based on neural networks is adopted, including sensor modules, sub-network modules and main network modules. Through a layered reinforcement learning mechanism, different types of data are processed using GPS sub-network, IMU sub-network and WiFi sub-network, and positioning results are generated through deep neural networks.
It improves the positioning accuracy of vehicles in urban canyons or indoor environments, enhances the robustness and real-time response capabilities of the system, reduces dependence on a single data source, and achieves efficient modular learning and improvement of positioning accuracy.
Smart Images

Figure CN120313616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle positioning. Specifically, the present invention relates to a multi-source data fusion combined positioning system and method based on a neural network. Background Art
[0002] In recent years, GPS (Global Positioning System) technology has been widely used. However, traditional GPS positioning technology is prone to signal occlusion and reflection in urban canyons or indoor environments, resulting in a decrease in positioning accuracy. Therefore, in the prior art, a multi-source data fusion method has been proposed to improve positioning accuracy. However, most multi-source data fusion methods adopt simple weighting or filtering methods, and fail to fully utilize the complementarity of sensor data (such as the global nature of GPS, the continuity of IMU, and the local reference of WiFi), resulting in insufficient positioning accuracy in dynamic and complex environments. At the same time, technologies such as visual SLAM are prone to failure in low-light or dynamic scenarios, and UWB requires pre-deployment of base stations and high costs, which also bring inconvenience to multi-source data fusion. In addition, traditional methods rely on manually designed fusion rules and lack the ability of adaptive optimization.
[0003] Patent CN109613581A discloses a GPS positioning system, a GPS positioning method, and a GPS positioning terminal. By making a pure radial movement of a cooperative target and observing the deflection angle between the point track detected by the radar and the GPS point track, the angle error can be corrected. The delay in sending the GPS point track only has an error in the radial distance, so that the radar 0° and true north angle errors can be accurately corrected; by making a pure tangential movement of the cooperative target, two arcs centered on the station center should be formed on the radar monitoring operation terminal interface, and the distance difference between the two arcs is the radar distance error to be corrected. The delay in sending the GPS point track only has a front-back error in the angle and does not affect the correction of the radar distance accuracy.
[0004] However, the technology disclosed in the above patent relies on a specific cooperative target to make a pure radial or pure tangential movement, which means that it may not be suitable for dynamic environments or scenarios where the target movement trajectory cannot be controlled. For example, in actual road traffic, it is difficult for moving targets such as vehicles to move in a predetermined manner, and it is impossible to meet the positioning accuracy requirements in urban canyons or indoor environments. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multi-source data fusion combined positioning system and method based on a neural network for the deficiencies of the prior art, so as to achieve the following objectives: improve vehicle positioning accuracy, especially in urban canyons or indoor environments.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: A multi-source data fusion and combined positioning system based on a neural network, comprising a sensor module, a sub-network module, and a main network module; the sensor module is connected to the sub-network module; the sub-network module is connected to the main network module; wherein, the sub-network module is used to process different types of data sources and cooperate with the main network through a hierarchical reinforcement learning mechanism for position estimation.
[0007] Preferably, the sub-network module includes a GPS sub-network, an IMU sub-network, and a WiFi sub-network; the input ends of the main network module are respectively connected to the GPS sub-network, the IMU sub-network, and the WiFi sub-network, and the output ends of the sensor module are respectively connected to the GPS sub-network, the IMU sub-network, and the WiFi sub-network; wherein, the GPS sub-network uses a long short-term memory network to process sequential GPS data; the IMU sub-network adopts a recurrent neural network or a long short-term memory network to process inertial measurement unit data; the WiFi sub-network uses a feed-forward neural network to analyze WiFi signals.
[0008] Preferably, the sensor module outputs sensor data to the sub-network module; the sub-network module processes the sensor data input by the sensor module and outputs the processed data to the main network module; the main network module receives the feature vectors from all sub-network modules as inputs and generates the final positioning result, i.e., the longitude and latitude coordinates of the vehicle, after comprehensively processing this information through a deep neural network model.
[0009] Preferably, the sub-network module includes a sub-network state space S i , a sub-network action space A i and a sub-network policy D i ; wherein, the sub-network state space S i is a feature set of the sensor data output by the sensor module; the sub-network action space Ai defines the set of actions that each sub-network may perform, and the sub-network policy Di defines the probability distribution of selecting a certain action ai ∈ Ai under the given state si ∈ Si;
[0010] The main network module includes a main network state space S m , a main network action space A m and a main network policy D m ; the main network state space Sm is composed of the data features output by all sub-network modules; the main network action space Am represents the longitude and latitude coordinates of the finally predicted vehicle; the main network policy Dm defines the probability distribution of selecting a certain action am ∈ Am under the given comprehensive state sm ∈ Sm.
[0011] Preferably, the sub-network module further includes a sub-network reward function Ri (s t ,a t ) and the sub-network value function V i (s i );The main network module further includes a main network reward function R m (s m ,a m ) and the main network value function V m (s m );
[0012] The sub-network reward function described above:
[0013] R i (s t ,a t ) = β × R i (local) (s t ,a t ) + (1 - β) × R m (s t ,a t );
[0014] Among them, s t is the state of the sub-network module at time step t, a t is the action of the sub-network module at time step t, R i (local) (s t ,a t ) is the reward function of the sub-network module itself, R m (s t ,a t ) is the manifestation of the main network reward function R m (s m ,a m ) in the sub-network module, and β is a tuning parameter used to adjust the weight of the reward function R i (local) (s t ,a t ) of the sub-network module itself and the manifestation Rm(st, at) of the main network reward function Rm(sm, am) in the sub-network module; The system automatically adjusts the weight parameters of the sub-network and the main network according to the feedback of the reward function using the gradient ascent method until a predetermined performance standard or convergence condition is reached; The weight updates of the sub-network and the main network are performed independently, but are kept consistent through a shared reward mechanism;
[0015] The main network reward function described above:
[0016] R m (s m ,am ) = -‖a m -a true ‖ 2 ;
[0017] where s m is the comprehensive state of the main network module, a m is the action of the main network module, a true is the true position coordinate of the vehicle, ‖a m -a true ‖ is the Euclidean distance between a m and a true ;
[0018] The sub-network value function:
[0019] V i (s i ) = E(∑ t=0 ∞ γ t ×R i (s i t , a i t )|s i 0 = s i );
[0020] where E is the expected value, γ is the discount factor, t is the time, s i 0 is the state of the sub-network module at the initial moment, s i t is the state of the sub-network module at time t, a i t is the action of the sub-network module at time t, R i (s i t , a i t ) is the reward function of the sub-network module at time step t;
[0021] The main network value function V m (s m ) = E(∑ t=0 ∞ γ t ×R m (s m t , a m t )|s m 0 = s m );
[0022] Among them, E is the expected value, γ is the discount factor, t is the time, s m 0 is the state of the main network module at the initial moment, s m t is the comprehensive state of the main network module at time t, a m t is the action of the main network module at time t, R m (s m t , a m t ) is the reward function of the main network module at time step t.
[0023] Preferably, the weights of the GPS sub-network, the IMU sub-network, and the WiFi sub-network are updated according to the reward functions R1(s t , a t ) of the GPS sub-network, the reward function R2(s t , a t ) of the IMU sub-network, and the reward function R3(s t , a t ) of the WiFi sub-network; the weight of the main network module is updated according to the reward function R m (s m , a m ).
[0024] Preferably, the state space S1 of the GPS sub-network includes the intensity of the GPS signal and the number of satellites; the action space A1 of the GPS sub-network includes the longitude and latitude coordinate information of the vehicle; the policy D1 of the GPS sub-network defines the probability that the GPS sub-network outputs the longitude and latitude coordinate information of the vehicle to the main network module under the state of the given intensity of the GPS signal and the number of satellites;
[0025] The reward function of the GPS sub-network:
[0026] R1(s t , a t ) = β × R1 (local) (s t , a t ) + (1 - β) × R m (s t , a t );
[0027] Among them, s t is the state of the GPS sub-network at time step t, a t is the action of the GPS sub-network at time step t, R1 (local) (st , a t ) is the reward function of the GPS sub-network itself, R m (s t , a t ) is the main network reward function R m (s m , a m ) reflected in the GPS sub-network, β is to adjust the reward function R1 of the GPS sub-network itself (local) (s t , a t ) weight and the main network reward function R m (s m , a m ) reflected in the GPS sub-network R m (s t , a t ) weight parameter;
[0028] The value function of the GPS sub-network:
[0029] V1(s1) = E(∑ t=0 ∞ γ t ×R1(s1 t , a1 t )|s1 0 = s1);
[0030] Where E is the expected value, γ is the discount factor, t is the time, s1 0 is the state of the GPS sub-network at the initial moment, s1 t is the state of the GPS sub-network at time t, a1 t is the action of the GPS sub-network at time t, R1(s1 t , a1 t ) is the reward function of the GPS sub-network at time step t.
[0031] 8. A multi-source data fusion and combined positioning system based on a neural network according to claim 6, characterized in that: the state space S2 of the IMU sub-network includes vehicle acceleration and vehicle angular velocity; the action space A2 of the IMU sub-network includes vehicle heading and speed information; the policy D2 of the IMU sub-network defines the probability that the IMU sub-network outputs vehicle heading and speed information to the main network module under the state of given vehicle acceleration and vehicle angular velocity;
[0032] The reward function of the IMU sub-network:
[0033] R2(s t , a t) = β × R2 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t );
[0034] Where s t is the state of the IMU sub - network at time step t, a t is the action of the IMU sub - network at time step t, R2 (local) (s t ,a t ) is the reward function of the IMU sub - network itself, R m (s t ,a t ) is the manifestation of the main network reward function R m (s m ,a m ) in the IMU sub - network, and β is the weight parameter that adjusts the weight of the reward function R2 (local) (s t ,a t ) of the IMU sub - network itself and the weight of the manifestation R m (s m ,a m ) of the main network reward function R m (s t ,a t ) in the IMU sub - network;
[0035] The value function of the IMU sub - network:
[0036] V2(s2) = E(∑ t=0 ∞ γ t ×R2(s2 t ,a2 t )|s2 0 = s2);
[0037] Where E is the expected value, γ is the discount factor, t is time, s2 0 is the state of the IMU sub - network at the initial moment, s2 t is the state of the IMU sub - network at time t, a2 t is the action of the IMU sub - network at time t, and R2(s2 t ,a2 t ) is the reward function of the IMU sub - network at time step t.
[0038] Preferably, the state space S3 of the WiFi sub-network includes the WiFi signal strength and the AP access point information; the action space A3 of the WiFi sub-network includes the position information of the vehicle relative to the WiFi access point; the policy D3 of the WiFi sub-network defines the probability that the WiFi sub-network outputs the position information of the vehicle relative to the WiFi access point to the main network module under the state of the given WiFi signal strength and AP access point information;
[0039] The reward function of the WiFi sub-network:
[0040] R3(s t ,a t ) = β × R3 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t );
[0041] Where s t is the state of the WiFi sub-network at time step t, a t is the action of the WiFi sub-network at time step t, R3 (local) (s t ,a t ) is the reward function of the WiFi sub-network itself, R m (s t ,a t ) is the embodiment of the main network reward function R m (s m ,a m ) in the WiFi sub-network, and β is the weight parameter that adjusts the weight of the reward function R3 (local) (s t ,a t ) of the WiFi sub-network itself and the weight of the embodiment R m (s m ,a m ) of the main network reward function R m (s t ,a t ) in the WiFi sub-network;
[0042] The value function of the WiFi sub-network:
[0043] V3(s3) = E(∑ t=0 ∞ γ t × R3(s3 t ,a3 t ) | s3 0 = s3);
[0044] Among them, E is the expected value, γ is the discount factor, t is the time, and s3 0 is the state of the WiFi sub-network at the initial moment, and s3 t is the state of the WiFi sub-network at time t, and a3 t is the action of the WiFi sub-network at time t, and R3(s3 t , a3 t ) is the reward function of the WiFi sub-network at time step t.
[0045] Meanwhile, this application also proposes a multi-source data fusion combined positioning method based on a neural network. The method includes the following steps:
[0046] Step 1, the sensor module outputs sensor data to the sub-network module;
[0047] Step 2, the sub-network module processes the sensor data input by the sensor module and outputs the processed data to the main network module;
[0048] Step 3, the main network module determines the longitude and latitude coordinates of the vehicle according to the data input by the sub-network module;
[0049] Step 4, calculate the reward function of the sub-network module and the reward function of the main network module;
[0050] Step 5, calculate the gradient of the reward function of the sub-network module with respect to the weight and the gradient of the reward function of the main network module with respect to the weight;
[0051] Step 6, update the weights of the sub-network module and the main network module;
[0052] Step 7, determine whether the convergence condition is satisfied. If not, go to Step 1. If so, end.
[0053] The positioning system and method based on a neural network of the present invention have the following advantages:
[0054] (1) The positioning system and method based on a neural network of the present invention can improve the positioning accuracy of a vehicle in an urban canyon or an indoor environment.
[0055] (2) By fusing multi-source data and applying the fast processing ability of a neural network, the system of the present invention can respond in real time, is applicable to a dynamic environment, and at the same time, the system of the present invention can provide more accurate position information than a single data source. Meanwhile, the fusion of multi-source data can reduce the dependence on a single data source and enhance the robustness of the system.
[0056] (3) By decomposing complex tasks into subtasks, the present invention enables the sub-network module and the main network module to focus on their respective tasks. Meanwhile, through the shared reward signal and policy, the behaviors of the sub-network module and the main network module are coordinated, thereby achieving efficient and modular learning of the entire system. This coordination mechanism allows the system to optimize the overall performance of the system while maintaining flexibility and scalability.
[0057] (4) In the present invention, each sub-network in the sub-network module has its own reward function, which can then motivate it to accurately interpret the corresponding sensor data; the main network module integrates the outputs of the sub-network module into an accurate positioning result through the main network reward function.
[0058] (5) In the present invention, the complex positioning task is decomposed into two levels through the concept of hierarchical reinforcement learning, making the learning process more modular, facilitating the understanding and optimization of each part, reducing the complexity of the overall design, and improving the learning efficiency and processing ability of the system.
[0059] (6) The sub-network module and the main network module in the present invention have their own reward functions. This hierarchical design allows each network to focus on its own task. Meanwhile, through effective coordination using the reward function, the goals of the sub-network module and the main network module are consistent while being distinguishable. The sub-network module focuses on the quality of feature extraction, and the main network module focuses on the accuracy of position estimation.
[0060] (7) The system framework of the present invention is an end-to-end system framework. The weights of the sub-network module and the weights of the main network module can be automatically adjusted through the feedback of the reward function, improving the learning efficiency.
[0061] (8) The present invention uses a neural network to learn the complex relationships between different data sources, realizes the deep fusion of data through the neural network, and improves the accuracy and reliability of positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] This specification includes the following drawings, and the shown contents are respectively:
[0063] Figure 1 is the logical structure block diagram of a multi-source data fusion combined positioning system based on a neural network according to the present invention;
[0064] Figure 2 is the schematic diagram of a multi-source data fusion combined positioning system based on a neural network according to the present invention;
[0065] Figure 3 is the flowchart of a multi-source data fusion combined positioning method based on a neural network according to the present invention.
[0066] Description of the reference numerals: 1. Sensor module; 2. Sub-network module; 3. Main network module; 21. Global Positioning System (GPS) sub-network; 22. Inertial Measurement Unit (IMU) sub-network; 23. Wireless Fidelity (WiFi) sub-network. Detailed implementation manners
[0067] The following will, with reference to the accompanying drawings, further elaborate on the detailed implementation manners of the present invention through the description of the embodiments, aiming to help those skilled in the art have a more complete, accurate, and in-depth understanding of the inventive concept and technical solution of the present invention, and facilitate its implementation.
[0068] Figure 1 is a logical structure block diagram of a multi-source data fusion and combined positioning system based on a neural network according to the present invention, including a sensor module 1, a sub-network module 2, and a main network module 3. The sensor module 1 is connected to the sub-network module 2, and the sub-network module 2 is connected to the main network module 3. The sub-network module 2 is used to process different types of data sources and cooperate with the main network module 3 through a hierarchical reinforcement learning mechanism to perform accurate position estimation.
[0069] The sensor module 1 outputs sensor data to the sub-network module 2; the sub-network module 2 processes the sensor data input by the sensor module 1 and outputs the processed data to the main network module 3; the main network module 3 receives the feature vectors from all sub-network modules 2 as inputs and generates the final positioning result, that is, the longitude and latitude coordinates of the vehicle, after comprehensively processing this information through a deep neural network model.
[0070] Among them, the sub-network module 1 includes a GPS (Global Positioning System) sub-network 21, an IMU (Inertial Measurement Unit) sub-network 22, and a WiFi (Wireless Fidelity) sub-network 23. The input ends of the main network module are respectively connected to the GPS sub-network 21, the IMU sub-network 22, and the WiFi sub-network 23. The output ends of the sensor module are respectively connected to the GPS sub-network 21, the IMU sub-network 22, and the WiFi sub-network 23.
[0071] In the present invention, the complex positioning task is decomposed into two levels through the concept of hierarchical reinforcement learning, making the learning process more modular, facilitating the understanding and optimization of each part, reducing the complexity of the overall design, and improving the learning efficiency and processing ability of the system. In addition, considering the timeliness and possible noise of GPS data, the GPS sub-network 21 can adopt a long short-term memory network (LSTM) or a gated recurrent unit (GRU). In this specific embodiment, the long short-term memory network (LSTM) is taken as an example; considering the time series characteristics of IMU data, the IMU sub-network 22 can adopt a recurrent neural network (RNN) or a long short-term memory network (LSTM). In this specific embodiment, the recurrent neural network (RNN) is taken as an example; considering that the WiFi signal is relatively simple, the WiFi sub-network 23 can adopt a feedforward neural network (FNN).
[0072] In the present invention, the sub-network module 2 includes a sub-network state space S i , a sub-network action space A i and a sub-network policy D i ; where the sub-network state space S i of the sub-network module 2 i (a i |s i ) = P(a i |s i ), defines the probability distribution of selecting a certain action ai ∈ Ai under the given state si ∈ Si. This policy can be deterministic or random, and is selected according to different requirements of network design, increasing the adaptability of the algorithm.
[0073] The main network module 3 includes a main network state space S m , a main network action space A m and a main network policy D m ; where the main network state space S m of the main network module 3 m (a m |s m ) = P(a m |s m ), defines the probability distribution of selecting a certain action am ∈ Am under the given comprehensive state sm ∈ Sm. This policy can be deterministic or random, and is selected according to different requirements of network design, increasing the adaptability of the algorithm.
[0074] Specifically, the state space S1 of the GPS sub-network 21 includes the intensity of the Global Positioning System (GPS) signal and the number of satellites. The action space A1 of the GPS sub-network 21 includes the longitude and latitude coordinate information of the vehicle. The policy D1 of the GPS sub-network 21 defines the probability that the GPS sub-network 21 outputs the longitude and latitude coordinate information of the vehicle to the main network module 3 under the state of the given intensity of the Global Positioning System (GPS) signal and the number of satellites.
[0075] The state space S2 of the IMU sub-network 22 includes the vehicle acceleration and the vehicle angular velocity. The action space A2 of the IMU sub-network 22 includes the heading and speed information of the vehicle. The policy D2 of the IMU sub-network 22 defines the probability that the IMU sub-network 22 outputs the heading and speed information of the vehicle to the main network module 3 under the state of the given vehicle acceleration and the vehicle angular velocity.
[0076] The state space S3 of the WiFi sub-network 23 includes the Wireless Fidelity (WiFi) signal strength and the AP access point information; the action space A3 of the WiFi sub-network 23 includes the position information of the vehicle relative to the Wireless Fidelity (WiFi) access point; the policy D3 of the WiFi sub-network 23 defines the probability that the WiFi sub-network 23 outputs the position information of the vehicle relative to the Wireless Fidelity (WiFi) access point to the main network module 3 under the state of the given Wireless Fidelity (WiFi) signal strength and the AP access point information.
[0077] The present invention can reduce the dependence on a single data source through the fusion of multi-source data, enhance the robustness of the system. At the same time, the present invention realizes the deep fusion of data through a neural network by learning the complex relationships between different data sources, and thus can improve the accuracy and reliability of positioning.
[0078] In the present invention, the sub-network module 2 further includes a sub-network reward function R i (s t , a t ) and a sub-network value function V i (s i ); the main network module 3 further includes a main network reward function R m (s m , a m ) and a main network value function V m (s m ); if the sub-network module 2 is optimized only based on its reward function, the information important to the main network module 3 will be ignored. By adding a term related to the main network module 3 to the sub-network module 2, the main network module 3 can guide the sub-network module 2 to focus on the aspects more important to the overall task. Therefore, the sub-network reward function R i (st , a t ) can take the following form:
[0079] R i (s t , a t ) = β × R i (local) (s t , a t ) + (1 - β) × R m (s t , a t );
[0080] Among them, s t is the state of sub-network module 2 at time step t, a t is the action of sub-network module 2 at time step t, R i (local) (s t , a t ) is the reward function of sub-network module 2 itself, R m (s t , a t ) is the manifestation of the main network reward function R m (s m , a m ) in sub-network module 2, β is a tuning parameter used to adjust the weight of the reward function R i (local) (s t , a t ) of sub-network module 2 itself and the manifestation R m (s m , a m ) of the main network reward function R m (s t , a t ) in sub-network module 2; The system uses the gradient ascent method to automatically adjust the weight parameters of the sub-network and the main network according to the feedback of the reward function until a predetermined performance standard or convergence condition is reached; The weight updates of the sub-network and the main network are performed independently, but are kept consistent through a shared reward mechanism.
[0081] The reward function of the main network module 3 is the core of the entire system, directly reflecting the ultimate goal of the positioning task, that is, positioning accuracy. Therefore, the main network reward function R m (s m , a m ) can adopt a reward function in the following form:
[0082] R m (s m , a m ) = -‖a m - atrue ‖ 2 ;
[0083] This reward function encourages accurate positioning by minimizing the squared distance between the predicted position and the true position. The negative value of the squared distance is used as the reward, meaning that the closer to the true position, the higher the reward. Here, s m is the comprehensive state of the main network module 3, which is formed by concatenating the feature vectors of all sub-networks in the sub-network module 2. a m is the action of the main network module 3, that is, the predicted longitude and latitude coordinates. a true is the true position coordinates of the vehicle. ‖a m −a true ‖ is the Euclidean distance (L2 norm) between a m and a true .
[0084] Sub-network value function: V i (s i ) = E(∑ t=0 ∞ γ t ×R i (s i t , a i t )|s i 0 = s i );
[0085] The sub-network value function represents the expected cumulative reward that can be obtained starting from a certain state s i and following the policy D i . Here, E is the expected value, γ is the discount factor used to balance the importance of the current reward and future rewards, t is the time, s i 0 is the state of the sub-network module 2 at the initial time, s i t is the state of the sub-network module 2 at time t, a i t is the action of the sub-network module 2 at time t, and R i (s i t , a i t ) is the reward function of the sub-network module 2 at time step t.
[0086] Main network value function: V m (s m ) = E(∑ t=0 ∞ γ t ×R m (sm t , a m t )|s m 0 = s m );
[0087] The main network value function represents the expected cumulative reward that can be obtained starting from a certain comprehensive state s m and following policy D m , where E is the expected value, γ is the discount factor used to balance the importance of current and future rewards, t is time, s m 0 is the state of the main network module 3 at the initial moment, and s m t is the comprehensive state of the main network module 3 at time t, a m t is the action of the main network module 3 at time t, and R m (s m t , a m t ) is the reward function of the main network module 3 at time step t.
[0088] Specifically, the reward function of the GPS sub-network 21:
[0089] R1(s t , a t ) = β × R1 (local) (s t , a t ) + (1 - β) × R m (s t , a t );
[0090] where s t is the state of the GPS sub-network 21 at time step t, a t is the action of the GPS sub-network 21 at time step t, R1 (local) (s t , a t ) is the reward function of the GPS sub-network 21 itself, and R m (s t , a t ) is the manifestation of the main network reward function R m (s m , a m ) in the GPS sub-network 21, and β is a tuning parameter used to adjust the reward function R i (local) (s t , a t)The weights and the main network reward function R m (s m ,a m ) The weights; The system automatically adjusts the weights of the sub-network and the main network according to the feedback of the reward function using the gradient ascent method until a predetermined performance standard or convergence condition is reached; The weight updates of the sub-network and the main network are independent, but are kept consistent through a shared reward mechanism.
[0091] Value function of the GPS sub-network 21:
[0092] V1(s1) = E(∑ t=0 ∞ γ t ×R1(s1 t ,a1 t )|s1 0 = s1);
[0093] Where E is the expected value, γ is the discount factor, t is the time, s1 0 is the state of the GPS sub-network 21 at the initial time, s1 t is the state of the GPS sub-network 21 at time t, a1 t is the action of the GPS sub-network 21 at time t, R1(s1 t ,a1 t ) is the reward function of the GPS sub-network 21 at time step t.
[0094] Reward function of the IMU sub-network 22:
[0095] R2(s t ,a t ) = β × R2 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t );
[0096] Where s t is the state of the IMU sub-network 22 at time step t, a t is the action of the IMU sub-network 22 at time step t, R2 (local) (s t ,a t ) is the reward function of the IMU sub-network 22 itself, R m (s t ,a t ) is the main network reward function R m (s m ,a m) Embodied in the IMU sub-network 22, β is a tuning parameter used to adjust the self-reward function R2 of the IMU sub-network 22 (local) (s t ,a t ) of the weight with the main network reward function R m (s m ,a m ) embodied in the IMU sub-network 22 R m (s t ,a t ).
[0097] Value function of the IMU sub-network 22:
[0098] V2(s2) = E(∑ t=0 ∞ γ t ×R2(s2 t ,a2 t )|s2 0 = s2);
[0099] Where E is the expected value, γ is the discount factor, t is the time, s2 0 is the state of the IMU sub-network 22 at the initial moment, s2 t is the state of the IMU sub-network 22 at time t, a2 t is the action of the IMU sub-network 22 at time t, R2(s2 t ,a2 t ) is the reward function of the IMU sub-network 22 at time step t.
[0100] Reward function of the WiFi sub-network 23:
[0101] R3(s t ,a t ) = β × R3 (local) (s t ,a t )+(1 - β) × R m (s t ,a t );
[0102] Where s t is the state of the WiFi sub-network 23 at time step t, a t is the action of the WiFi sub-network 23 at time step t, R3 (local) (s t ,a t ) is the self-reward function of the WiFi sub-network 23, R m (s t ,a t ) is the main network reward function Rm (s m ,a m ) in the WiFi sub-network 23; β is an adjustment parameter for adjusting the self-reward function R3 of the WiFi sub-network 23 (local) (s t ,a t ) The weight and the main network reward function R m (s m ,a m ) in the WiFi sub-network 23 is embodied as R m (s t ,a t ) weight.
[0103] Value function of the WiFi sub-network 23:
[0104] V3(s3) = E(∑ t=0 ∞ γ t ×R3(s3 t ,a3 t )|s3 0 = s3);
[0105] where E is the expected value, γ is the discount factor, t is the time, s3 0 is the state of the WiFi sub-network 23 at the initial moment, s3 t is the state of the WiFi sub-network 23 at time t, a3 t is the action of the WiFi sub-network 23 at time t, R3(s3 t ,a3 t ) is the reward function of the WiFi sub-network 23 at time step t. That is, in the present invention, each sub-network in the sub-network module 2 has its own reward function, which can then motivate it to accurately interpret the corresponding sensor data. The main network module 3 integrates the output of the sub-network module 2 into an accurate positioning result through the main network reward function; and this hierarchical design can enable each network to focus on its own task, while effectively coordinating through the reward function, so that the goals of the sub-network module 2 and the main network module 3 are consistent, while being distinguishable. The sub-network module 2 focuses on the quality of feature extraction, and the main network module 3 focuses on the accuracy of position estimation.
[0106] Figure 2 is the schematic diagram of a multi-source data fusion and combined positioning system based on a neural network according to the present invention. Combining Figure 1 and Figure 2 the working principle of the present invention is introduced as follows:
[0107] In the present invention, the sensor module 1 outputs sensor data to the sub-network module 2. The sub-network module 2 processes the sensor data input by the sensor module 1 and outputs the processed data to the main network module 3. Then, the main network module 3 determines the longitude and latitude coordinates of the vehicle according to the data input by the sub-network module 2. Specifically, the sensor module 1 outputs information including the intensity of the Global Positioning System (GPS) signal and the number of satellites to the GPS sub-network 21, the sensor module 1 outputs information including the vehicle acceleration and the vehicle angular velocity to the IMU sub-network 22, and the sensor module 1 outputs information including the Wireless Fidelity (WiFi) signal strength and the AP access point information to the Wireless Fidelity (WiFi) signal strength and the AP access point information; the GPS sub-network 21 processes the GPS data and outputs the longitude and latitude coordinate information of the vehicle to the main network module 3, the IMU sub-network 22 processes the IMU data and outputs the heading and speed information of the vehicle to the main network module 3, and the WiFi sub-network 23 processes the WiFi data and outputs the position information of the vehicle relative to the Wireless Fidelity (WiFi) access point to the main network module 3; then the main network module 3 synthesizes the data input by the GPS sub-network 21, the IMU sub-network 22 and the WiFi sub-network 23 to determine the longitude and latitude coordinates of the vehicle. That is, the positioning system and method based on neural network of the present invention can improve the positioning accuracy of the vehicle in urban canyons or indoor environments.
[0108] In the process of determining the longitude and latitude coordinates of the vehicle as described above, the sub-network module 2 calculates its reward function R i (s t ,a t ) and the gradient ▽ i (s t ,a t ) of its reward function R θi R i (s t ,a t ) with respect to the weights, that is, the GPS sub-network 21, the IMU sub-network 22 and the WiFi sub-network 23 will respectively calculate their reward functions R1(s t ,a t ), R2(s t ,a t ), R3(s t ,a t ) and the gradients ▽ t t t t t t θ1 t t θ2 R2(s t , a t ), ▽ θ3 R3(s t , a t ), and at the same time, the main network module 3 calculates its reward function R m (s m , a m ) and its gradient ▽ m R m (s m , a θm R m (s m , a m ) with respect to the weights; subsequently, the weights θ1 of the GPS sub-network 21, the weights θ2 of the IMU sub-network 22, the weights θ3 of the WiFi sub-network 23, and the weights θ m。
[0109] of the main network module 3 are updated respectively by the gradient ascent method; that is, the system framework of the present invention is an end-to-end system framework, and the weights of the sub-network module 2 and the main network module 3 can be automatically adjusted through the feedback of the reward function, improving the learning efficiency; after the weights θ m of the main network module 3 are updated, it is determined whether the convergence condition is satisfied. If it is satisfied, the update ends. If it is not satisfied, it returns to the stage where the sensor module 1 outputs sensor data to the sub-network module 2 and continues to iterate and upgrade, adding details of the training parameters: learning rate α = 0.001, discount factor γ = 0.99; β(t) = max(0.2, 0.8×e -0.001t ); Convergence condition: The validation set error < 2 meters for 10 consecutive epochs. t is the current training step. This strategy can not only quickly optimize the local tasks of the sub-network in the initial stage of training, but also gradually increase the attention to the global tasks in the later stage of training, while avoiding the risk brought by too low β.
[0110] The convergence condition therein is that the average longitude and latitude error on the validation set is less than 2 meters within 10 consecutive cycles. In the above process, the weights of the GPS sub-network 21, the IMU sub-network 22, and the WiFi sub-network 23 are updated according to the reward function R1(s t , a t ) of the GPS sub-network 21, the reward function R2(s t , a t ) of the IMU sub-network 22, and the reward function R3(s t , a t ) of the WiFi sub-network 23 respectively; the weights of the main network module 3 are updated according to the reward function R m (s m , a m ) of the main network module 3.
[0111] Specifically, the formulas for updating the weights θ1 of the GPS sub-network 21, the weights θ2 of the IMU sub-network 22, and the weights θ3 of the WiFi sub-network 23 respectively by the gradient ascent method can be expressed in the following form:
[0112] θ i (k+1) = θ i (k) + α × (β × ▽ θi R i (local) (s t ,a t ) + (1 - β) × ▽ θi R m (s m ,a m ));
[0113] Among them, θ i (k) is the weight of the sub-network i at the k-th iteration, θ i (k+1) is the weight of the sub-network i at the (k + 1)-th iteration, α is the learning rate used to control the weight update step size, β is the weight parameter that adjusts the weight of the reward function R i (local) (s t ,a t ) of the sub-network i and the manifestation R m (s m ,a m ) of the reward function R m (s t ,a t ) of the main network in the sub-network i, ▽ θi R i (local) (s t ,a t ) is the gradient of the reward function R i (local) (s t ,a t ) of the sub-network i with respect to the weight θ i , ▽ θi R m (s m ,a m ) is the gradient of the reward function R m (s m ,a m ) of the main network module 3 with respect to θ iThe gradient, where i takes 1, 2, or 3. When i = 1, it represents the GPS sub-network 21; when i = 2, it represents the IMU sub-network 22; when i = 3, it represents the WiFi sub-network 23. Update the weights θ of the main network module 3 by the gradient ascent method. m The formula can be expressed in the following form;
[0114] θ m (k+1) = θ m (k) + α × ▽ θi R m (s m ,a m );
[0115] Among them, θ m (k) is the weight of the main network module at the k-th iteration, θ m (k+1) is the weight of the main network module at the (k + 1)-th iteration, α is the learning rate used to control the weight update step size, and ▽ θi R m (s m ,a m ) is the reward function R of the main network module 3 m (s m ,a m ) with respect to θ m .
[0116] In addition, by fusing multi-source data, the present invention applies the fast processing ability of the neural network to enable the system to respond in real time and be applicable to dynamic environments. At the same time, the system of the present invention can provide more accurate position information than a single data source. Moreover, by decomposing complex tasks into subtasks, the sub-network module 2 and the main network module 3 each focus on their own tasks, and at the same time, the behavior of the sub-network module 2 and the main network module 3 is coordinated through shared reward signals and policies, thereby realizing the efficient and modular learning of the entire system. This coordination mechanism allows the system to optimize the overall performance of the system while maintaining flexibility and scalability.
[0117] In addition, in this specific embodiment, the sensor module 1 includes a GPS sensor, an IMU sensor, and a WiFi sensor.
[0118] Figure 3 is a flowchart of a multi-source data fusion and combined positioning method based on a neural network according to the present invention. The method includes the following steps:
[0119] Step 1, the sensor module 1 outputs sensor data to the sub-network module 2.
[0120] Step 2, the sub-network module 2 processes the sensor data input by the sensor module 1 and outputs the processed data to the main network module 3.
[0121] Step 3, the main network module 3 determines the longitude and latitude coordinates of the vehicle based on the data input by the sub-network module 2.
[0122] Step 4, calculate the reward function of the sub-network module 2 and the reward function of the main network module 3.
[0123] Step 5, calculate the gradient of the reward function of the sub-network module 2 with respect to the weights and the gradient of the reward function of the main network module 3 with respect to the weights.
[0124] Step 6, update the weights of the sub-network module 2 and the weights of the main network module 3.
[0125] Step 7, determine whether the convergence condition is satisfied. If not, go to Step 1; if so, end.
[0126] In summary, the present invention has the following technical effects:
[0127] (1) In the present invention, the complex positioning task is decomposed into two levels, namely the sub-network and the main network, through the concept of hierarchical reinforcement learning. The sub-network is responsible for feature extraction, and the main network is responsible for position estimation, improving the modularity and scalability of the system, making the learning process more modular, facilitating the understanding and optimization of each part, reducing the complexity of the overall design, and improving the learning efficiency and processing ability of the system.
[0128] (2) The present invention can reduce the dependence on a single data source through the fusion of multi-source data, enhancing the robustness of the system. At the same time, the present invention uses neural networks to learn the complex relationships between different data sources, and through the deep fusion of multi-source data such as GPS, IMU, and WiFi by deep neural networks, significantly improves the positioning accuracy, especially showing excellent performance in environments with complex signals (such as urban canyons and indoors).
[0129] (3) In the present invention, each sub-network in the sub-network module 2 has its own reward function, which can then motivate it to accurately interpret the corresponding sensor data. The main network module 3 integrates the output of the sub-network module 2 into an accurate positioning result through the main network reward function; and this hierarchical design allows each network to focus on its own task, while effectively coordinating through the reward function, making the goals of the sub-network module 2 and the main network module 3 consistent while being distinguishable. The sub-network module 2 focuses on the quality of feature extraction, and the main network module 3 focuses on the accuracy of position estimation.
[0130] (4) Shared reward mechanism: By sharing the main network reward function, the optimization directions of the sub-network and the main network are coordinated to ensure that while the sub-network optimizes its own tasks, it also takes into account the overall positioning accuracy.
[0131] (5) Respectively update the weights θ1 of the GPS sub-network 21, the weights θ2 of the IMU sub-network 22, the weights θ3 of the WiFi sub-network 23, and the weights θ of the main network module 3 through the gradient ascent method m . That is, the system framework of the present invention is an end-to-end system framework. The weights of the sub-network module 2 and the main network module 3 can be automatically adjusted through the feedback of the reward function, improving the learning efficiency.
[0132] (6) In addition, by fusing multi-source data and applying the fast processing ability of the neural network, the system of the present invention can respond in real time, is applicable to dynamic environments. At the same time, the system of the present invention can provide more accurate position information than a single data source. Moreover, by decomposing complex tasks into sub-tasks, the sub-network module 2 and the main network module 3 each focus on their own tasks, and at the same time, the behavior of the sub-network module 2 and the main network module 3 is coordinated through shared reward signals and strategies, thereby realizing the efficient and modular learning of the entire system. This coordination mechanism allows the system to optimize the overall performance of the system while maintaining flexibility and scalability.
[0133] The present invention also conducts comparative experiments with traditional methods for verification, and the verification results are shown in Table 1:
[0134]
[0135]
[0136] Table 1
[0137] As can be seen from the table, compared with traditional methods, the positioning error of the present invention in urban canyons is controlled within 1.8 meters, and the response speed is within 50 ms; the indoor positioning error is controlled within 2 meters, and the response speed is within 50 ms; there is a significant improvement compared with traditional methods.
[0138] The present invention has been described exemplarily in combination with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention; or without improvement, the above concept and technical solution of the present invention are directly applied to other occasions, they are all within the protection scope of the present invention.
Claims
1. A multi-source data fusion and combined positioning system based on a neural network, characterized in that: It includes a sensor module, a sub-network module, and a main network module; the sensor module is connected to the sub-network module; the sub-network module is connected to the main network module; wherein, the sub-network module is used to process different types of data sources and cooperate with the main network through a hierarchical reinforcement learning mechanism for position estimation.
2. The multi-source data fusion and combined positioning system based on a neural network according to claim 1, wherein: The sub-network module includes a GPS sub-network, an IMU sub-network, and a WiFi sub-network; the input ends of the main network module are respectively connected to the GPS sub-network, the IMU sub-network, and the WiFi sub-network, and the output ends of the sensor module are respectively connected to the GPS sub-network, the IMU sub-network, and the WiFi sub-network; wherein, the GPS sub-network uses a long short-term memory network to process sequential GPS data; the IMU sub-network adopts a recurrent neural network or a long short-term memory network to process inertial measurement unit data; the WiFi sub-network uses a feedforward neural network to analyze WiFi signals.
3. The multi-source data fusion and combined positioning system based on a neural network according to claim 2, wherein: The sensor module outputs sensor data to the sub-network module; the sub-network module processes the sensor data input by the sensor module and outputs the processed data to the main network module; the main network module receives the feature vectors from all sub-network modules as inputs and comprehensively processes this information through a deep neural network model to generate the final positioning result, that is, the longitude and latitude coordinates of the vehicle.
4. A multi-source data fusion and combination positioning system based on a neural network according to claim 3, wherein: The described sub-network module includes a sub-network state space S i , a sub-network action space A i and a sub-network policy D i ; among them, the described sub-network state space S i is a feature set of the sensor data output by the described sensor module; the sub-network action space Ai defines the set of actions that each sub-network may execute, and the sub-network policy Di defines the probability distribution of selecting a certain action ai∈Ai under the given state si∈Si; The described main network module includes a main network state space S m , a main network action space A m , and a main network policy D m ; the main network state space Sm is composed of the data features output by all sub-network modules; the main network action space Am represents the finally predicted vehicle longitude and latitude coordinates; the main network policy Dm defines the probability distribution of selecting a certain action am ∈ Am under a given comprehensive state sm ∈ Sm.
5. The multi-source data fusion and combination positioning system based on a neural network according to claim 4, characterized in that: The sub-network module further includes a sub-network reward function R i (s t , a t ) and a sub-network value function V i (s i ); The main network module further includes a main network reward function R m (s m , a m ) and a main network value function V m (s m ); The sub-network reward function: R i (s t ,a t ) = β × R i (local) (s t ,a t ) + (1 - β) × R m (s t ,a t ); where s t is the state of the sub-network module at time step t, a t is the action of the sub-network module at time step t, R i (local) (s t , a t ) is the reward function of the sub-network module itself, R m (s t , a t ) is the manifestation of the main network reward function R m (s m , a m ) in the sub-network module. β is a tuning parameter used to adjust the weight of the reward function R i (local) (s t , a t ) of the sub-network module itself and the weight of the manifestation Rm(st, at) of the main network reward function Rm(sm, am) in the sub-network module; The system automatically adjusts the weight parameters of the sub-network and the main network according to the feedback of the reward function using the gradient ascent method until a predetermined performance standard or convergence condition is reached; The weight updates of the sub-network and the main network are performed independently, but are kept consistent through a shared reward mechanism; The main network reward function: R m (s m ,a m )=-‖a m -a true ‖ 2 ; where s m is the comprehensive state of the main network module, a m is the action of the main network module, a true is the true position coordinate of the vehicle, ‖a m −a true ‖ is the Euclidean distance between a m and a true ; The sub-network value function: V i (s i )=E(∑ t=0 ∞ γ t ×R i (s i t ,a i t )|s i 0 =s i ); where E is the expected value, γ is the discount factor, t is the time, s i 0 is the state of the sub-network module at the initial moment, s i t is the state of the sub-network module at time t, a i t is the action of the sub-network module at time t, R i (s i t , a i t ) is the reward function of the sub-network module at time step t; The main network value function V m (s m ) = E(∑ t=0 ∞ γ t × R m (s m t , a m t ) | s m 0 = s m ); where E is the expected value, γ is the discount factor, t is the time, s m 0 is the state of the main network module at the initial moment, s m t is the comprehensive state of the main network module at time t, a m t is the action of the main network module at time t, R m (s m t , a m t ) is the reward function of the main network module at time step t.
6. The multi-source data fusion and combined positioning system based on a neural network according to claim 5, characterized in that: The weights of the GPS sub-network, the IMU sub-network, and the WiFi sub-network are updated according to the reward functions R1(s t , a t ), R2(s t , a t ), and R3(s t , a t ) of the GPS sub-network, the IMU sub-network, and the WiFi sub-network respectively; the weight of the main network module is updated according to the reward function R m (s m , a m ) of the main network module.
7. The multi-source data fusion and combined positioning system based on a neural network according to claim 6, characterized in that: The state space S1 of the GPS sub-network includes the intensity of the GPS signal and the number of satellites; the action space A1 of the GPS sub-network includes the longitude and latitude coordinate information of the vehicle; the policy D1 of the GPS sub-network defines the probability that the GPS sub-network outputs the longitude and latitude coordinate information of the vehicle to the main network module in a state given the intensity of the GPS signal and the number of satellites. The reward function of the GPS sub-network: R1(s t ,a t ) = β × R1 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t ); where s t is the state of the GPS sub-network at time step t, a t is the action of the GPS sub-network at time step t, R1 (local) (s t , a t ) is the reward function of the GPS sub-network itself, R m (s t , a t ) is the main network reward function R m (s m , a m ) as reflected in the GPS sub-network, and β is a weight parameter that adjusts the weight of the GPS sub-network's own reward function R1 (local) (s t , a t ) and the weight of the reflection R m (s m , a m ) of the main network reward function R m (s t , a t ) in the GPS sub-network; The value function of the GPS sub-network: V1(s1) = E(∑ t=0 ∞ γ t ×R1(s1 t , a1 t )|s1 0 = s1); where E is the expected value, γ is the discount factor, t is the time, s1 0 is the state of the GPS sub-network at the initial moment, s1 t is the state of the GPS sub-network at time t, a1 t is the action of the GPS sub-network at time t, R1(s1 t , a1 t ) is the reward function of the GPS sub-network at time step t.
8. The multi-source data fusion and combined positioning system based on a neural network according to claim 6, characterized in that: The state space S2 of the IMU sub-network includes the vehicle acceleration and the vehicle angular velocity; the action space A2 of the IMU sub-network includes the heading and speed information of the vehicle; the policy D2 of the IMU sub-network defines the probability that the IMU sub-network outputs the heading and speed information of the vehicle to the main network module in a state given the vehicle acceleration and the vehicle angular velocity. The reward function of the IMU sub-network: R2(s t ,a t ) = β × R2 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t ); where s t is the state of the IMU sub-network at time step t, a t is the action of the IMU sub-network at time step t, R2 (local) (s t , a t ) is the reward function of the IMU sub-network itself, R m (s t , a t ) is the main network reward function R m (s m , a m ) reflected in the IMU sub-network, and β is a weight parameter that adjusts the weight of the IMU sub-network's own reward function R2 (local) (s t , a t ) and the weight of the reflection R m (s m , a m ) of the main network reward function R m (s t , a t ) in the IMU sub-network; The value function of the IMU sub-network: V2(s2) = E(∑ t=0 ∞ γ t ×R2(s2 t , a2 t )|s2 0 = s2); Among them, E is the expected value, γ is the discount factor, t is the time, and s2 0 is the state of the IMU sub-network at the initial moment, s2 t is the state of the IMU sub-network at time t, a2 t is the action of the IMU sub-network at time t, R2(s2 t , a2 t ) is the reward function of the IMU sub-network at time step t.
9. The multi-source data fusion and combination positioning system based on a neural network according to claim 6, characterized in that: The state space S3 of the WiFi sub-network includes the WiFi signal strength and the AP access point information; the action space A3 of the WiFi sub-network includes the position information of the vehicle relative to the WiFi access point; the policy D3 of the WiFi sub-network defines the probability that the WiFi sub-network outputs the position information of the vehicle relative to the WiFi access point to the main network module in a state given the WiFi signal strength and the AP access point information. The reward function of the WiFi sub-network: R3(s t ,a t ) = β × R3 (local) (s t ,a t ) + (1 - β) × R m (s t ,a t ); where s t is the state of the WiFi sub-network at time step t, a t is the action of the WiFi sub-network at time step t, R3 (local) (s t , a t ) is the reward function of the WiFi sub-network itself, R m (s t , a t ) is the main network reward function R m (s m , a m ) reflected in the WiFi sub-network, β is to adjust the weight of the reward function R3 (local) (s t , a t ) of the WiFi sub-network itself and the weight parameter of the weight of the reflection R m (s m , a m ) of the main network reward function R m (s t , a t ) in the WiFi sub-network; The value function of the WiFi sub-network: V3(s3) = E(∑ t=0 ∞ γ t ×R3(s3 t , a3 t )|s3 0 = s3); where E is the expected value, γ is the discount factor, t is the time, s3 0 is the state of the WiFi sub-network at the initial moment, s3 t is the state of the WiFi sub-network at time t, a3 t is the action of the WiFi sub-network at time t, R3(s3 t , a3 t ) is the reward function of the WiFi sub-network at time step t.
10. A multi-source data fusion and combined positioning method based on a neural network applicable to any one of claims 1 to 9, characterized in that: The method includes the following steps: Step 1, the sensor module outputs sensor data to the sub-network module; Step 2, the sub-network module processes the sensor data input by the sensor module and outputs the processed data to the main network module; Step 3, the main network module determines the latitude and longitude coordinates of the vehicle based on the data input by the sub-network module; Step 4, calculate the reward function of the sub-network module and the reward function of the main network module; Step 5, calculate the gradient of the reward function of the sub-network module with respect to the weights and the gradient of the reward function of the main network module with respect to the weights; Step 6, update the weights of the sub-network module and the weights of the main network module; Step 7, determine whether the convergence condition is satisfied. If not, go to Step 1. If so, end.
Citation Information
Patent Citations
GPS positioning system, GPS positioning method and GPS positioning terminal
CN109613581A