Geomagnetic vector matching navigation method based on deep reinforcement learning and related device
By employing a geomagnetic vector matching navigation method based on deep reinforcement learning, the rotation and translation errors in geomagnetic contour matching navigation are corrected using standard geomagnetic contour lines, thereby improving positioning accuracy and enhancing the robustness and adaptability of the navigation system.
Patent Information
- Application Number
- CN202511262383.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing geomagnetic contour matching navigation methods only correct translational errors and ignore rotational errors, resulting in insufficient positioning accuracy.
A geomagnetic vector matching navigation method based on deep reinforcement learning is adopted. By acquiring the measured geomagnetic information of multiple track points, reference track points are determined from the standard geomagnetic contour lines. Rotational and translational deviations are calculated, and position adjustments are made to ensure that the error correction process has clear physical meaning and reliable reference basis.
It improves positioning accuracy, solves the problem of insufficient positioning accuracy caused by single-dimensional correction, and enhances the robustness and adaptability of the navigation system.
Smart Images

Figure CN120778121B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of navigation, and more specifically, to a geomagnetic vector matching navigation method and related apparatus based on deep reinforcement learning. Background Technology
[0002] Geomagnetic matching navigation, as an autonomous navigation technology, demonstrates significant advantages in scenarios with satellite signal shielding and unclear visual features due to its independence from external signal sources. However, current navigation methods based on geomagnetic contour matching achieve positioning by making the geomagnetic field at the matching location as close as possible to the characteristics of geomagnetic measurements. But in terms of error correction, geomagnetic contour matching algorithms only correct for translational errors, neglecting the potential impact of rotational errors on the overall positioning effect, resulting in insufficient positioning accuracy. Summary of the Invention
[0003] To overcome at least one deficiency in the prior art, this application provides a geomagnetic vector matching navigation method and related device based on deep reinforcement learning. It is used to perform correction based on standard geomagnetic contour lines by using multiple measured geomagnetic information as evaluation criteria, ensuring that the error correction process has clear physical meaning and reliable reference basis. By simultaneously correcting translation error and rotation error, it solves the problem of insufficient positioning accuracy caused by single-dimensional correction in traditional geomagnetic contour matching methods, thereby improving positioning accuracy.
[0004] In a first aspect, this application provides a geomagnetic vector matching navigation method based on deep reinforcement learning, the method comprising:
[0005] Obtain multiple waypoints;
[0006] Based on the measured geomagnetic information of the multiple track points, multiple reference track points corresponding to the multiple track points are determined from standard geomagnetic contour lines, wherein the distance between each reference track point and its corresponding track point satisfies a preset constraint condition.
[0007] Based on the multiple waypoints and the multiple reference waypoints, the rotational deviation and translational deviation are obtained;
[0008] The positions of the multiple waypoints are adjusted based on the rotational and translational deviations to obtain updated waypoints.
[0009] Secondly, this application provides a geomagnetic vector matching navigation device based on deep reinforcement learning, the device comprising:
[0010] The data acquisition module is used to acquire multiple waypoints;
[0011] The data correction module is used to determine multiple reference track points corresponding to the multiple track points from standard geomagnetic contour lines based on the measured geomagnetic information of the multiple track points, wherein the distance between each reference track point and the corresponding track point satisfies a preset constraint condition.
[0012] The data correction module is also used to obtain rotational deviation and translational deviation based on the plurality of waypoints and the plurality of reference waypoints;
[0013] The data correction module is also used to adjust the positions of the multiple track points according to the rotational deviation and translational deviation, so as to obtain updated multiple track points.
[0014] Thirdly, this application provides a storage medium storing a computer program that, when executed by a processor, implements the deep reinforcement learning-based geomagnetic vector matching navigation method.
[0015] Thirdly, this application provides an electronic device, which includes a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements the deep reinforcement learning-based geomagnetic vector matching navigation method.
[0016] Compared with the prior art, this application has the following beneficial effects:
[0017] This application provides a geomagnetic vector matching navigation method and related device based on deep reinforcement learning. The electronic device acquires multiple waypoints and, based on the measured geomagnetic information of these waypoints, determines multiple reference waypoints corresponding to the multiple waypoints from standard geomagnetic contour lines. The distance between each measured geomagnetic information point and its corresponding waypoint satisfies preset constraints. Rotational and translational deviations are obtained based on the multiple waypoints and reference waypoints. The positions of the multiple waypoints are adjusted according to the rotational and translational deviations to obtain updated waypoints. By introducing multiple measured geomagnetic information points from standard geomagnetic contour lines as reference benchmarks, the error correction process has clear physical meaning and reliable reference basis. Furthermore, by simultaneously correcting translational and rotational errors, the method solves the problem of insufficient positioning accuracy caused by single-dimensional correction in traditional methods, thereby improving positioning accuracy. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 One of the flowcharts for the geomagnetic vector matching navigation method based on deep reinforcement learning provided in the embodiments of this application;
[0020] Figure 2 This is a schematic diagram of the sliding window provided in an embodiment of this application;
[0021] Figure 3 The second flowchart illustrates the geomagnetic vector matching navigation method based on deep reinforcement learning provided in this application embodiment.
[0022] Figure 4 A schematic diagram of the structure of a geomagnetic vector matching navigation device based on deep reinforcement learning provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application (hereinafter referred to as "the embodiments") clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0026] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0027] In the description of this application, it should be noted that the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0028] Based on the above statement, as introduced in the background technology, the geomagnetic contour matching algorithm only corrects for translation errors, while ignoring the potential impact of rotation errors on the overall positioning effect, resulting in insufficient positioning accuracy of the algorithm.
[0029] Based on the discovery of the aforementioned technical problems, the following technical solutions are proposed through creative effort to solve or improve these problems. It should be noted that the deficiencies in the solutions of the prior art are the result of practical experience and careful research. Therefore, the discovery process of the aforementioned problems and the solutions proposed in the embodiments of this application below should be considered as contributions made to this application during the inventive process, and should not be construed as technical content known to those skilled in the art.
[0030] Therefore, this embodiment provides a geomagnetic vector matching navigation method based on deep reinforcement learning. For example... Figure 1 As shown, the method includes:
[0031] S1, obtain multiple waypoints.
[0032] S2, based on the measured geomagnetic information of multiple track points, determines multiple reference track points corresponding to the multiple track points from standard geomagnetic contour lines.
[0033] Among them, the distance between each measured geomagnetic information and the corresponding track point satisfies the preset constraint conditions.
[0034] S3, based on multiple track points and multiple reference track points, obtains the rotational deviation and translational deviation.
[0035] S4. Adjust the positions of multiple track points based on rotation and translation deviations to obtain updated track points.
[0036] Thus, multiple measured geomagnetic information is used as evaluation criteria for correction based on standard geomagnetic contour lines, ensuring that the error correction process has clear physical meaning and reliable reference basis. By simultaneously correcting translation and rotation errors, the insufficient positioning accuracy caused by single-dimensional correction in traditional geomagnetic contour matching methods is solved, thereby improving positioning accuracy.
[0037] It should be understood that the electronic device implementing the deep reinforcement learning-based geomagnetic vector matching navigation method provided in this embodiment can be, but is not limited to, a mobile terminal, tablet computer, laptop computer, server, mobile vehicle, etc. The mobile vehicle can be, but is not limited to, drones, autonomous vehicles, unmanned ground vehicles, unmanned underwater vehicles, etc.
[0038] To make the solution provided in this embodiment clearer, a drone is used as the electronic device implementing the method, and the steps of the method shown in the figure are described in detail below. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. Furthermore, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowchart, or remove one or more operations from the flowchart. See also... Figure 1 The method includes:
[0039] S1, obtain multiple waypoints.
[0040] During the above steps, the UAV simultaneously utilizes multiple navigation methods, including an inertial navigation system and a geomagnetic navigation system. During flight, the UAV collects navigation data generated by the inertial navigation system at a fixed frequency and simultaneously acquires measured geomagnetic information from the geomagnetic navigation device. The inertial navigation system generates a series of coordinate points through built-in sensors; each coordinate point represents a waypoint during the UAV's flight, and these waypoints, marked on a map, represent the UAV's trajectory over a specific time period. Meanwhile, the geomagnetic navigation device synchronously records measured geomagnetic information, including the geomagnetic field strength and direction. In practical applications, because the external environment may interfere with the measured geomagnetic information, magnetic interference compensation processing is required to ensure its accuracy and reliability.
[0041] like Figure 2 As shown, in this embodiment, to improve the efficiency of waypoint correction, the UAV does not use all waypoints for correction simultaneously, but selectively selects a portion of waypoints as the analysis object. Specifically, the size of the sliding window is set to... This means extracting a continuous sequence. Five trackpoints are captured in the sliding window in the image. Therefore, the UAV performs geomagnetic matching-based correction only on these five trackpoints within the sliding window's coverage area each time. After correction, it outputs the last trackpoint within the sliding window. The updated position of the last trackpoint (i.e., the end trackpoint at the end) is taken as the correction result for the current moment. This updated last trackpoint then becomes the new correction trackpoint.
[0042] It's important to note that the multiple waypoints within the sliding window can be divided into two categories: corrected waypoints and recently added waypoints. Corrected waypoints are those that have already undergone correction, while recently added waypoints, located at the end of the sliding window, represent those requiring focused correction. As the drone continues to move, the sliding window continuously slides forward, incorporating the newest waypoints while removing the oldest, ensuring the correction process remains focused on the current flight status.
[0043] S2, based on the measured geomagnetic information of multiple track points, determines multiple reference track points corresponding to the multiple track points from standard geomagnetic contour lines.
[0044] It should be understood that standard geomagnetic isolines refer to a pre-constructed set of isoline data reflecting the geomagnetic field distribution characteristics of a specific region. Similar to contour lines, each isoline connects points with the same geomagnetic field strength or direction characteristics, thus forming a complete geomagnetic field distribution map. Therefore, any point on a geomagnetic isoline has the same geomagnetic field. These isolines are generated based on high-precision geomagnetic measurements or simulations and can accurately describe the geomagnetic field strength and direction information at different locations within the region. Geomagnetic isolines can be represented in two-dimensional or three-dimensional form and typically include the isoline distribution of multiple components of the geomagnetic field (such as the X, Y, and Z components).
[0045] It should also be understood that the distance between each measured geomagnetic information point and its corresponding track point satisfies a preset constraint. This preset constraint can be the point on the geomagnetic isopleths that is closest to the track point. Of course, it can also be any point on the geomagnetic isopleths within a preset range of distances from the track point. Furthermore, it should be understood that the geomagnetic isopleths are generated from a pre-constructed standard geomagnetic vector reference map, which reflects the actual spatial distribution characteristics of the three components of the magnetic field (i.e., the X, Y, and Z directional components) within the navigation area. Therefore, this means that geomagnetic isopleths exist on the X, Y, and Z directional components.
[0046] Taking the nearest point as an example, the UAV uses the track point provided by the inertial navigation system and the real-time collected measured geomagnetic information to match the measured geomagnetic information with the contour lines in the standard geomagnetic vector reference map to determine the geomagnetic contour lines corresponding to the measured geomagnetic information, and then determines the point closest to the track point from the geomagnetic contour lines as a reference track point.
[0047] Here, the reference waypoint is represented as The nearest neighbor points on the isopleths of the three geomagnetic components represent the values of the three components respectively. , , The final reference waypoint can then be calculated using the following formula. :
[0048]
[0049] in, , , The weighting coefficients for the three components are used to adjust the importance of different directional components. Therefore, this weighting process can comprehensively consider the distribution characteristics of each component of the magnetic field, thereby more accurately determining the position of the reference track point.
[0050] In this embodiment, for the first A sliding window is used, and the objective function for optimization is set as follows:
[0051]
[0052] Therefore, the waypoint update method can be expressed as:
[0053]
[0054] in, It's the learning rate. yes right gradient, It is the first The first sliding window One waypoint Is with The corresponding reference waypoint.
[0055] Based on the above description of the reference waypoints in the embodiments, the following will continue to discuss... Figure 1 Step S3 will be explained below:
[0056] S3, based on multiple track points and multiple reference track points, obtains the rotational deviation and translational deviation.
[0057] As an optional implementation, step S3 may include:
[0058] S3-1, based on each track point and its corresponding reference track point, obtain the angle between two points.
[0059] S3-2, based on the statistical characteristics of multiple included angles, the rotational deviation is obtained.
[0060] During the above steps, the UAV represents the waypoint and reference waypoint as vectors and calculates the angular relationship between them using the formula for the angle between the two vectors. Therefore, a series of angle values are obtained, reflecting the directional differences between the waypoint and the corresponding reference waypoint. To synthesize the errors of multiple waypoints, the UAV can adjust the mean of these angles using a second update rate to obtain the rotational deviation. This can be understood as summing all angle values and dividing by the total number of angles to obtain a rotational deviation in the form of an expected value. This rotational deviation describes the overall rotational error of the waypoint relative to the reference waypoint within the entire sliding window.
[0061] For example, for the first A sliding window, representing the waypoints within it as... The corresponding reference waypoint is represented as Therefore, rotational deviation The expression is:
[0062]
[0063]
[0064] In the formula, This represents the average of multiple included angles. Rotation angle deviation The second update rate, Indicates the first The iteration. This can be understood as, for the... The track points within a sliding window can be corrected not only once, but also multiple times.
[0065] Based on the above description of rotational deviation in the embodiments, step S3 further includes:
[0066] S3-3, based on the rotational deviation, obtain the translational deviation.
[0067] For the calculation of translational deviation, this embodiment incorporates the rotation angle deviation from the above embodiments. Therefore, as an optional implementation, the UAV can obtain a rotation matrix based on the rotation deviation, wherein the rotation matrix uses the indicated rotation deviation as the rotation angle; apply the rotation matrix to multiple waypoints to obtain adjusted waypoints; compare the adjusted waypoints with multiple reference waypoints to obtain the expected deviation between the waypoints and the reference waypoints; and adjust the expected deviation using a first update rate to obtain the translational deviation.
[0068] For example, also for the first A sliding window, translation deviation The calculation method is as follows:
[0069]
[0070] In the formula, Indicates the first The rotation matrix for the next iteration. This indicates rotational deviation.
[0071]
[0072] In the formula, This indicates the expected deviation.
[0073]
[0074] In the formula, Translational deviation The first update rate.
[0075] Based on the above embodiments explaining rotational and translational deviations, the following will continue to discuss... Figure 1 Step S4 will be explained below:
[0076] S4. Adjust the positions of multiple track points based on rotation and translation deviations to obtain updated track points.
[0077] In this embodiment, both the rotational deviation and the translational deviation are considered. Therefore, as an optional implementation of step S4, for each waypoint, the UAV can apply a rotation matrix based on the rotational deviation to the waypoint to obtain the rotational angle adjustment; the waypoint is then corrected using the rotational angle adjustment and the translational deviation to obtain an updated waypoint.
[0078] For example, the track points of the sliding window can be iteratively corrected according to the following formula to obtain the first... The updated waypoints :
[0079]
[0080] In the formula, Indicates the first The waypoints are updated for the [number]th time. This can be understood as, if the waypoints continue to be updated, then the [number]th [time] ... The next updated waypoint is .
[0081] In practice, it was found that although simultaneously correcting translation and rotation errors can solve the problem of insufficient positioning accuracy caused by single-dimensional correction in traditional geomagnetic contour matching methods, the correction effect depends on the actual measurement accuracy of the geomagnetic field and the corresponding accuracy of the geomagnetic contour distribution. Therefore, its robustness to disturbances is not strong, and it is prone to getting trapped in local optima, leading to a decrease in positioning accuracy. Therefore, after correcting for rotation and translation errors, a comparative analysis with standard geomagnetic information and measured geomagnetic information can be performed to evaluate the quality of the current correction result and provide a basis for subsequent correction strategies. In this way, after a certain number of iterations, a more accurate positioning result can be obtained. This embodiment needs to consider three factors, such as... Figure 3 As shown, the method also includes:
[0082] S5 acquires updated standard geomagnetic information and measured geomagnetic information for the locations of multiple track points.
[0083] S6 calculates the first difference information between multiple standard geomagnetic information and multiple measured geomagnetic information.
[0084] During the above steps, the first step is to read the standard geomagnetic information corresponding to all track points within the current sliding window from the geomagnetic vector reference map. This standard geomagnetic information originates from a pre-constructed geomagnetic vector reference map, which reflects the geomagnetic field strength and direction characteristics at each location within the region. Simultaneously, the actual measured geomagnetic information for the same set of track points also needs to be acquired. This measured geomagnetic information is collected in real-time by the equipment during its movement and, after necessary magnetic interference compensation processing, is used to characterize the actual geomagnetic field characteristics of the current location.
[0085] To quantify the degree of difference between these two sets of information, this embodiment introduces mean square error, which is obtained by performing a component-by-component difference square operation on the standard geomagnetic information and the measured geomagnetic information of each track point, and averaging the results, thereby obtaining the mean square error of the geomagnetic vector in each direction of the track sequence within the sliding window, which serves as the first difference information.
[0086] For example, the geomagnetic vector values of the current sliding window corrected track sequence are read from the geomagnetic vector reference map. And calculate the corresponding geomagnetic vector measurement value. Mean square error .
[0087]
[0088] In the formula, , , They represent the first The updated trackpoint's location has standard geomagnetic information in the X, Y, and Z directions. Similarly, , , Indicates the first The measured geomagnetic information of each track point in the X, Y, and Z directions.
[0089] S7, obtain the single standard geomagnetic information of the location of the last track point among the updated multiple track points, and the single measured geomagnetic information of the location of the last track point among the multiple track points.
[0090] S8 calculates the second difference information between a single standard geomagnetic information and a single measured geomagnetic information.
[0091] In this implementation, for the last track point in the sliding window, its corresponding standard geomagnetic information and measured geomagnetic information need to be acquired separately. The purpose is to focus on the matching accuracy of the current position, thereby providing a more specific reference for subsequent navigation strategies. During the above steps, the UAV reads the standard geomagnetic information and measured geomagnetic information of the last track point before and after the update from the geomagnetic vector reference map, and directly calculates the geomagnetic vector error through the difference between the two, as the second difference information between them.
[0092] For example, continuing to assume that there are a total of Given several trajectory points, the expression for the second difference information is:
[0093]
[0094] In the formula, This indicates the second difference information. Represents a single standard geomagnetic information, It represents a single measured geomagnetic information.
[0095] S9. Based on the last track point, the first difference information, and the second difference information among the updated multiple track points, determine the subsequent correction strategy for the updated multiple track points.
[0096] To address this, this embodiment introduces an update strategy model to determine whether further correction of waypoints is needed. The UAV can input the last waypoint, the first difference information, and the second difference information from the updated multiple waypoints into the pre-trained update strategy model to obtain the subsequent correction strategy for the updated multiple waypoints.
[0097] This can be understood as follows: the updated policy model is obtained through deep reinforcement learning techniques. This updated policy model can use the state feature vector of the last of the multiple updated waypoints and related difference information to generate subsequent correction policies.
[0098] For example, considering that the update strategy model is a neural network-based model, it is necessary to construct a state feature vector for the current iteration cycle, which is then input into the update strategy model. This state feature vector consists of the updated last track point in the sliding window, its corresponding geomagnetic vector error, and the root mean square error of the geomagnetic vectors of multiple track points within the sliding window in each direction. These data reflect the current estimation accuracy, local geomagnetic matching deviation, and global geomagnetic matching deviation of the UAV, respectively. The UAV integrates the above three state components into a single state feature vector, which is then input into the update strategy model, which predicts the subsequent correction strategy.
[0099] In this embodiment, the updated policy model is trained using a Q-network. It should be understood that a Q-network is a neural network structure based on deep reinforcement learning, comprising an input layer, two hidden layers, and an output layer. The number of neurons in the input layer is consistent with the dimension of the state feature vector, ensuring complete reception of all information within the state feature vector. Each hidden layer is a fully connected layer, using the ReLU function as the activation function to extract the nonlinear relationships of the state features. The output layer is also a fully connected layer, with the number of output neurons consistent with the dimension of the action space, corresponding to possible action selections. For example, action 1 indicates continuing correction along the current iteration gradient direction; action 0 indicates terminating correction; and action -1 indicates backtracking to the previous estimated position and terminating correction. This can be understood as follows: as long as the updated policy model outputs action 1, it continues to correct the track points within the sliding window until the maximum number of iterations is reached.
[0100] Thus, after correcting multiple waypoints within the current sliding window at least once, the corrected position of the last waypoint in the sliding window is retained. If the UAV continues to fly, the waypoint for the next moment is estimated based on inertial navigation sensor data and at least some of the corrected waypoints within the current sliding window. Then, the sliding window is moved to add the waypoint for the next moment to the sliding window (the corrected waypoint at the head of the sliding window is removed). Finally, the waypoints within the current sliding window are corrected again.
[0101] For the training process of the above-mentioned update policy model, the Q-network can adopt an ε-greedy policy as the action selection policy. This policy uses probability... Randomly select an action, with probability Select the action with the highest Q-value estimated by the current Q-network. After selecting the action, record the current state, the selected action, and the corresponding Q-value estimate, and calculate the state feature vector for the next time step based on the selected action. Then, input this state feature vector back into the Q-network to obtain the Q-value estimate vector for the next time step, and select the maximum value from it as the optimal Q-value estimate for the next state.
[0102] Furthermore, combining the reward signal of the current action, the target value of the Q-value of the current state is calculated using the Bellman equation of reinforcement learning. The reward function is designed based on the distance change between the estimated position and the actual position of the last waypoint: a positive reward of +1 is given when the estimated distance is shorter than the previous state; a penalty of -1 is applied if the distance increases; and a reward of 0 is given if the distance remains unchanged. In this way, the model is guided to gradually optimize the navigation strategy and improve matching accuracy.
[0103] Based on the above explanation, the loss function of the Q-network can be the mean squared error function, and the parameters of the Q-network can be updated and optimized through the backpropagation algorithm. The corresponding expression is as follows:
[0104]
[0105] In the formula, It is the reward value for the current action. It is a discount factor, between 0 and 1. The next state All possible actions The maximum Q-value estimate in the model.
[0106]
[0107] In the formula, The current state Q-value estimation, It is the target Q value. This refers to the batch size.
[0108] It should be noted that in practical applications, in addition to using Q-networks for training, improved Q-network architectures or other high-performance deep reinforcement learning algorithms (such as the PPO algorithm) can be introduced to further enhance model performance. Therefore, the above methods can not only effectively address navigation needs in complex environments, but also continuously improve the robustness and adaptability of the algorithm through continuous learning.
[0109] Furthermore, in practice, it was found that during geomagnetic vector matching navigation, there might be a significant deviation between the initial estimated position and the actual position at the beginning. This deviation could be caused by various factors, such as large cumulative errors in the inertial navigation system or a complex geomagnetic environment. To ensure effective convergence of subsequent fine-tuning steps, it is necessary to perform preliminary correction on the position of the initial track estimation point during the initial matching.
[0110] Specifically, within the preset distance at the start of flight, the UAV calculates the weighted mean square error (MSE) between the initial standard geomagnetic information (i.e., the theoretical geomagnetic vector value read from the standard geomagnetic reference map) corresponding to the initial waypoint within the current sliding window and its corresponding measured initial geomagnetic information. It should be understood that the MSE weighted value is a comprehensive evaluation index reflecting the potential deviation between the current initial waypoint and the true location. If the MSE weighted value is greater than a preset threshold, it indicates that these initial waypoints may deviate from a reasonable range and require further adjustment.
[0111] When the weighted mean square error exceeds a set threshold, a traditional geomagnetic matching method, such as magnetic profile matching (MAGCOM), can be used to coarsely locate the initial track points. It should be noted that although the traditional geomagnetic matching method has lower accuracy, it can quickly determine a relatively reasonable region globally, thereby updating the initial estimated position to this region.
[0112] The geomagnetic vector matching navigation method based on deep reinforcement learning (referred to as geomagnetic navigation) provided in this embodiment was tested in a 2-hour actual flight test together with satellite navigation and inertial navigation. It was found that the trajectory of geomagnetic navigation is very close to that of satellite navigation, with a root mean square error of 46m and a terminal point matching error of 12.3m. In contrast, the root mean square error of conventional inertial navigation is 4.587km. Therefore, it has good matching and positioning results.
[0113] Based on the same inventive concept as the deep reinforcement learning-based geomagnetic vector matching navigation method provided in this embodiment, this embodiment also provides a deep reinforcement learning-based geomagnetic vector matching navigation device. This device includes at least one software functional module that can be stored in a memory or embedded in an electronic device. The processor in the electronic device executes the executable module stored in the memory. For example, the software functional modules and computer programs included in this device. Please refer to... Figure 4 Functionally, the device may include:
[0114] The data acquisition module is used to acquire multiple waypoints;
[0115] The data correction module is used to determine multiple reference track points corresponding to multiple track points from standard geomagnetic contour lines based on the measured geomagnetic information of multiple track points. The distance between each reference track point and its corresponding track point satisfies preset constraints.
[0116] The data correction module is also used to obtain rotational and translational deviations based on multiple waypoints and multiple reference waypoints.
[0117] The data correction module is also used to adjust the positions of multiple track points based on rotational and translational deviations to obtain updated track points.
[0118] Optionally, the data correction module is also specifically used for:
[0119] Based on each track point and its corresponding reference track point, the angle between the two points is obtained;
[0120] Based on the statistical characteristics of multiple included angles, the rotational deviation is obtained;
[0121] The translational deviation is obtained from the rotational deviation.
[0122] Optionally, the data correction module is also specifically used for:
[0123] Based on the rotation deviation, the rotation matrix is obtained, where the rotation matrix uses the rotation deviation shown as the rotation angle;
[0124] Applying the rotation matrix to multiple waypoints yields adjusted waypoints.
[0125] The adjusted multiple waypoints are compared with multiple reference waypoints to obtain the expected deviation between the multiple waypoints and the multiple reference waypoints;
[0126] The expected deviation is adjusted using the first update rate to obtain the translation deviation.
[0127] Optionally, the data correction module is also specifically used for:
[0128] The rotational deviation is obtained by adjusting the mean between multiple angles using the second update rate.
[0129] Optionally, the data correction module is also specifically used for:
[0130] For each waypoint, a rotation matrix based on rotation deviation is applied to the waypoint to obtain the rotation angle adjustment.
[0131] The track points are corrected using rotation angle adjustment and translation deviation to obtain updated track points.
[0132] Optionally, the data correction module is also used for:
[0133] Obtain updated standard geomagnetic information for the locations of multiple track points, as well as multiple measured geomagnetic information for the locations of multiple track points;
[0134] Calculate the first difference information between multiple standard geomagnetic information and multiple measured geomagnetic information;
[0135] Obtain single standard geomagnetic information of the location of the last track point among multiple updated track points, and single measured geomagnetic information of the location of the last track point among multiple track points;
[0136] Calculate the second difference information between a single standard geomagnetic information and a single measured geomagnetic information;
[0137] Based on the last waypoint, the first difference information, and the second difference information in the updated multiple waypoints, the subsequent correction strategy for the updated multiple waypoints is determined.
[0138] Optionally, the data correction module is also specifically used for:
[0139] The last waypoint, the first difference information, and the second difference information from the updated multiple waypoints are input into a pre-trained update policy model to obtain the subsequent correction policy for the updated multiple waypoints.
[0140] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0141] It should also be understood that if the above embodiments are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0142] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. This storage medium stores a computer program, which, when executed by a processor, implements the deep reinforcement learning-based geomagnetic vector matching navigation method provided in this embodiment. The storage medium can be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0143] This embodiment provides an electronic device that implements a geomagnetic vector matching navigation method based on deep reinforcement learning. For example... Figure 5 As shown, the electronic device may include a processor 22 and a memory 21. The memory 21 stores a computer program, and the processor reads and executes the computer program corresponding to the above-described embodiments in the memory 21 to implement the geomagnetic vector matching navigation method based on deep reinforcement learning provided in this embodiment.
[0144] See also Figure 5 The electronic device also includes a communication unit 23. The memory 21, processor 22 and communication unit 23 are electrically connected to each other directly or indirectly through system bus 24 to realize data transmission or interaction.
[0145] The memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principles, used to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, volatile memory, non-volatile memory, memory drive, etc.
[0146] In some embodiments, the volatile memory may be random access memory (RAM); in some embodiments, the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.; in some embodiments, the storage drive may be a disk drive, solid-state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or a combination thereof.
[0147] The communication unit 23 is used to send and receive data over a network. In some embodiments, the network may include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.
[0148] The processor 22 may be an integrated circuit chip with signal processing capabilities, and may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor described above may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC) computer, or a microprocessor, or any combination thereof.
[0149] Understandable. Figure 5The structure shown is for illustrative purposes only. Electronic devices may also have more advanced features. Figure 5 Showing more or fewer components, or having with Figure 5 The different configurations shown. Figure 5 The components shown can be implemented using hardware, software, or a combination thereof.
[0150] It should be understood that the apparatus and methods disclosed in the above embodiments can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0151] The above descriptions are merely various embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A geomagnetic vector matching navigation method based on deep reinforcement learning, characterized in that, The method includes: Obtain multiple waypoints; Based on the measured geomagnetic information of the multiple track points, multiple reference track points corresponding to the multiple track points are determined from standard geomagnetic contour lines, wherein the distance between each reference track point and its corresponding track point satisfies a preset constraint condition. Based on the multiple waypoints and the multiple reference waypoints, the rotational deviation and translational deviation are obtained; The positions of the multiple track points are adjusted based on the rotational and translational deviations to obtain updated track points. Obtain multiple standard geomagnetic information and multiple measured geomagnetic information of the locations of the updated multiple track points; Calculate the first difference information between the plurality of standard geomagnetic information and the plurality of measured geomagnetic information; Obtain a single standard geomagnetic information of the location of the last track point among the updated multiple track points, and a single measured geomagnetic information of the location of the last track point among the multiple track points; Calculate the second difference information between the single standard geomagnetic information and the single measured geomagnetic information; The last waypoint, the first difference information, and the second difference information from the updated multiple waypoints are input into a pre-trained update strategy model to obtain the subsequent correction strategy for the updated multiple waypoints. The update strategy model is a neural network structure based on deep reinforcement learning and is obtained through deep reinforcement learning technology.
2. The geomagnetic vector matching navigation method based on deep reinforcement learning according to claim 1, characterized in that, Based on the multiple waypoints and the multiple reference waypoints, rotational deviations and translational deviations are obtained, including: The angle between two points is obtained based on each stated track point and its corresponding reference track point; The rotational deviation is obtained based on the statistical characteristics of the multiple included angles; The translational deviation is obtained based on the rotational deviation.
3. The geomagnetic vector matching navigation method based on deep reinforcement learning according to claim 2, characterized in that, The translational deviation is obtained based on the rotational deviation, including: Based on the rotation deviation, a rotation matrix is obtained, wherein the rotation matrix uses the indicated rotation deviation as the rotation angle; The rotation matrix is applied to the plurality of waypoints to obtain the adjusted plurality of waypoints; The adjusted multiple waypoints are compared with the multiple reference waypoints to obtain the expected deviation between the multiple waypoints and the multiple reference waypoints; The expected deviation is adjusted using the first update rate to obtain the translation deviation.
4. The geomagnetic vector matching navigation method based on deep reinforcement learning according to claim 2, characterized in that, The rotational deviation is obtained based on the statistical characteristics of the multiple included angles, including: The rotational deviation is obtained by adjusting the mean value among the multiple included angles using a second update rate.
5. The geomagnetic vector matching navigation method based on deep reinforcement learning according to claim 1, characterized in that, The positions of the multiple waypoints are adjusted based on the rotational and translational deviations to obtain updated waypoints, including: For each of the track points, a rotation angle adjustment is obtained by applying a rotation matrix based on the rotation deviation to the track point. The track point is corrected using the rotation angle adjustment and the translation deviation to obtain an updated track point.
6. A geomagnetic vector matching navigation device based on deep reinforcement learning, characterized in that, The device includes: The data acquisition module is used to acquire multiple waypoints; The data correction module is used to determine multiple reference track points corresponding to the multiple track points from standard geomagnetic contour lines based on the measured geomagnetic information of the multiple track points, wherein the distance between each reference track point and the corresponding track point satisfies a preset constraint condition. The data correction module is also used to obtain rotational deviation and translational deviation based on the plurality of waypoints and the plurality of reference waypoints; The data correction module is also used to adjust the positions of the multiple track points according to the rotation deviation and translation deviation to obtain updated multiple track points; The data correction module is also used to acquire multiple standard geomagnetic information of the locations of the updated multiple track points and multiple measured geomagnetic information of the locations of the multiple track points; Calculate the first difference information between the plurality of standard geomagnetic information and the plurality of measured geomagnetic information; Obtain a single standard geomagnetic information of the location of the last track point among the updated multiple track points, and a single measured geomagnetic information of the location of the last track point among the multiple track points; Calculate the second difference information between the single standard geomagnetic information and the single measured geomagnetic information; The last waypoint, the first difference information, and the second difference information from the updated multiple waypoints are input into a pre-trained update strategy model to obtain the subsequent correction strategy for the updated multiple waypoints. The update strategy model is a neural network structure based on deep reinforcement learning and is obtained through deep reinforcement learning technology.
7. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the geomagnetic vector matching navigation method based on deep reinforcement learning as described in any one of claims 1-5.
8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the geomagnetic vector matching navigation method based on deep reinforcement learning as described in any one of claims 1-5.
Citation Information
Patent Citations
Geomagnetic-assisted inertial navigation and laser velocimeter combined navigation method
CN116412820A