A Visual Localization Method for Auxiliary Airport Unmanned Vehicles Based on Meta-Learning
By adopting a meta-learning-based auxiliary visual positioning method in driverless cars, using RGB images collected by on-board cameras for scene coordinate regression, the problem of insufficient positioning of unmanned vehicles in airports and other environments is solved, and high-precision pose estimation and assisted driving are achieved.
Patent Information
- Application Number
- CN202210777001.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-04
AI Technical Summary
The existing driverless cars are not able to meet the needs of real-life applications due to lighting changes and scene changes in airports and other environments.
Using an auxiliary visual positioning method based on meta-learning, RGB images are collected through the vehicle-mounted camera, and scene coordinate regression is performed using the random sampling consistency method optimized by meta-learning to realize the position prediction of unmanned vehicles.
In special environments such as GPS signal occlusion, high-precision unmanned parking posture estimation is achieved, with position accuracy errors within 5cm and 5°, effectively assisting the driving of unmanned vehicles.
Smart Images

Figure CN115615446B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of visual positioning, and specifically relates to a visual positioning method for assisting an airport driverless vehicle based on meta-learning. Background Art
[0002] With the progress of society and the continuous development of the automotive industry, driverless vehicles have emerged as the times require and have become a focus of major car manufacturers for a while, and have also become one of the research hotspots in the current field of artificial intelligence. As a relatively closed environment, the airport is one of the scenarios where driverless technology is most likely to be realized first. From a system perspective, a driverless system includes an environmental perception system, a behavior decision-making system, and a vehicle control system. Among them, the positioning module and the steering control module are two key modules of the driverless vehicle. At present, traditional methods are still the main technologies supporting the implementation of driverless vehicles. However, traditional methods have poor robustness against problems such as light changes and scene changes, and it is difficult to meet the application requirements of driverless vehicles in real-world scenarios.
[0003] The positioning methods of driverless vehicles can be divided into various methods such as map information matching positioning based on the global positioning system, magnetic induction, inertial navigation, vision, and lidar. Different positioning methods can be adopted according to the application scenarios of the driverless vehicle. Among them, the method based on GPS is an absolute pose estimation method. This method uses GPS to locate the vehicle. The advantages of the positioning method based on GPS are that it can perform continuous positioning all-weather, centimeter-level positioning can be achieved using differential GPS, and it is suitable for global positioning; the disadvantages are that it is easily affected by environmental changes, and GPS signals will be blocked in high-rise buildings, tunnels, and indoor environments. For driverless vehicles in the airport, the limitations of repositioning based on GPS are obvious. When satellite signals are blocked or delayed, GPS will not work or become inaccurate. At the same time, autonomous driving technology requires reliable and high-precision camera position and orientation estimation. Therefore, this patent focuses on the technology of visual positioning to assist driverless vehicles, uses visual sensors to construct a deep neural network, and realizes accurate estimation of the end-to-end pose of the driverless vehicle. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention aims to provide a visual positioning method for assisting an airport driverless vehicle based on meta-learning, which can be used to implement auxiliary positioning and navigation technologies such as GPS signal occlusion in the airport driverless environment, and can meet the low-cost and accurate navigation performance of general in-vehicle cameras.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] A visual positioning method for assisting an airport driverless vehicle based on meta-learning, characterized in that the specific steps are as follows:
[0007] The camera acquisition unit carried by the S11 airport unmanned vehicle acquires RGB images, and performs scene coordinate regression on the acquired RGB images through the random sample consensus method optimized based on meta-learning, obtaining the pose prediction of the central point of the RGB image, thereby obtaining the pose of the unmanned vehicle equipped with the camera.
[0008] The above-mentioned random sample consensus method optimized based on meta-learning includes:
[0009] S11-1 Scene coordinate regression, predicting the 3D scene coordinates y corresponding to each pixel i of the 2D RGB image through the basic convolutional neural network ConvNeXt i (ω), where ω is the parameter of the neural network ConvNeXt model, and through this model, the 2D pixel point (x i , y i ) to the 3D scene coordinates: p is the coordinate of the 2D pixel point (x i , y i ) in the pixel coordinate system, P C is the coordinate of the point in the camera coordinate system, P W is the coordinate of the point in the world coordinate system, ω is the depth of the point, K is the internal parameter matrix of the camera, R CW and is the pose transformation from the world coordinate system to the camera coordinate system;
[0010] Use f(I; ω) to represent the mapping from the 2D image camera coordinate system to the scene coordinate system, where I represents a given image, ω is the parameter that the ConvNeXt neural network needs to learn, and optimize the learnable parameter ω by minimizing the expected pose loss of the final estimate on the training set
[0011]
[0012] In the formula, f * represents the true value of the pose of image I. In order to continuously train and optimize the parameter ω of the neural network through the gradient descent method, take the derivative of the parameter ω, and the partial derivative of the above formula is:
[0013]
[0014] Through the first step of using the neural network to achieve the prediction of scene coordinates, that is, the mapping from 2D image coordinate points to the 3D scene coordinate system, it prepares for the next robust pose optimization;
[0015] S11-2 First, obtain a large number of hypotheses h, where each hypothesis h depends on the parameters of the corresponding scene coordinates. The image I and the scene coordinate prediction Y define a dense correspondence set C over all image pixels i. As the first step of robust pose optimization, randomly select the smallest subset C of M correspondences RGB ={(p i ,y i )|y i ∈Y}, 0 ≤ i < M, and each C j corresponds to a camera pose hypothesis h j . Use a pose solver to recover it, i.e.:
[0016] h j = g(C j )
[0017] After that, select from the large number of collected hypotheses. The principle for evaluating each hypothesis is:
[0018]
[0019] In the formula, the function s(·) is an evaluation function, indicating the performance of each hypothesis model;
[0020] After continuously iteratively selecting the optimal hypothesis model, it is necessary to optimize the optimal model. Use all scene coordinates to optimize the optimal model:
[0021]
[0022] Continuously iterate the above robust pose optimization process to finally obtain a camera pose matrix with very high accuracy;
[0023] S11-3 Meta-learning optimization training. Through meta-learning, the two training processes are fused to achieve a fast pose model estimation process while ensuring the accuracy of the neural network's prediction of scene coordinates;
[0024] S21 Obtain the current camera pose through visual positioning of the images captured by the on-vehicle camera of the airport unmanned vehicle. Among them, the coordinate regression method based on the differentiable random sample consensus algorithm of meta-learning has very high accuracy, and the pose accuracy error is within 5 cm and 5°. It is an effective auxiliary means for assisting the unmanned vehicle during driving;
[0025] As a further improvement of the present invention, in the selection of the random hypothesis model in step S11-2, since the input used is an RGB image, the pose solver for recovery is:
[0026] p i = Kh -1 y i
[0027] where K is the camera calibration matrix, or the internal calibration parameters of the camera. Using this relationship, bundle adjustment in the perspective PnP solver recovers the camera pose from at least 4 2D-3D correspondences to obtain a unique solution: C j ≥4. In practice, 4 pairs of image-to-scene coordinates are randomly selected for prediction.
[0028] As a further improvement of the present invention, the evaluation function s(·) in the hypothesis selection in step S11-2 is determined according to the applied scenario and input as:
[0029]
[0030] where the function r(·) measures the residual between the pose parameter h and the scene coordinates. If the residual is less than the inner threshold τ, Ι[·] is calculated as 1. The residual function r(·) is specifically expressed as:
[0031] r(y i , h) = ||p i - Kh -1 y i ||.
[0032] As a further improvement of the present invention, the specific method for implementing the fusion training of steps S11-1 and S11-2 in step S11-3 is to use the MAML method to divide the training process into an inner loop and an outer loop, specifically as follows:
[0033] For model initialization, introduce MAML to initialize the model parameters and make the model have a better initial performance. Divide the training process of the neural network into an inner loop and an outer loop. The inner loop implements the training of the basic model functions, and the outer loop improves the training of the model generalization. Similarly, divide the training set into two parts for the training of the two loops;
[0034] Among them, the inner loop is to implement the mapping of scene coordinates in step S11-1, realizing the mapping from the 2D coordinates of the camera to the 3D scene coordinates. The optimization of the inner loop is:
[0035]
[0036] The specific parameter iteration process of the inner loop is expressed as:
[0037] ω' = ω - μ▽ ω L(f(I, ω))
[0038] where ω' is the optimal parameter of the inner loop iteration, ω is the initialized parameter, μ is the learning rate in the inner loop training process, and ▽ ω L(f ω ) represents the gradient of the inner loop to implement the scene coordinate regression function.
[0039] For the outer loop, it is mainly the robust pose optimization process of step S11-2. Therefore, the optimization process of the inner loop is as follows:
[0040]
[0041] Regarding the transfer of inner loop and outer loop parameters, the specific parameter transfer method is to perform meta-update or meta-optimization before sampling the next batch of tasks. By training the inner loop, the optimal parameter ω' is found. Then, the gradient of ω' for each inner loop is calculated, and the randomly initialized parameter ω is updated through the gradient update method. This enables the randomly initialized parameter ω to find an initial parameter that is more approximate to the target task, so that when training the next batch of tasks, many gradient steps do not need to be taken.
[0042] The whole process is described as follows:
[0043]
[0044] In the formula, ω is the initialized parameter, η is the learning rate hyperparameter of the outer loop, is the gradient of the number ω of the robust pose optimization hypothesis model optimization function for the outer loop.
[0045] As a further improvement of the present invention, when training the camera positioning model of the airport unmanned vehicle, data augmentation is applied to the input RGB images. We randomly adjust the brightness and contrast of the input images within the range of ±10%. We randomly rotate the images, the ground truth scene coordinates and the camera poses within the range of ±30°. We randomly rescale the images within the ranges of 66% and 150% and adjust the focal length accordingly.
[0046] As a further improvement of the present invention, the visual positioning model integrates the two-step training process into a one-step training through the meta-learning method, greatly saving the training cost.
[0047] Aiming at the problem that the current general positioning and navigation system fails to position or has low accuracy due to signal occlusion and other problems in environments such as airport interiors, the present invention proposes a visual positioning method for assisting airport unmanned vehicles based on meta-learning. The implementation of this method is divided into two steps. In the first step, the RGB images obtained from the on-vehicle camera of the unmanned vehicle are used to train through the visual positioning method proposed by the present invention to obtain a usable pose prediction model of the on-vehicle camera. In the second step, when the positioning effect of other positioning and navigation systems of the unmanned vehicle fails or is poor in special environments, the RGB images obtained from the on-vehicle camera are used to output the position and attitude data in real time through the trained model, so as to assist the positioning and navigation of the airport unmanned vehicle.
[0048] For the positioning and navigation problem of airport auxiliary unmanned vehicles, the present invention proposes a deep learning vision positioning method based on meta-learning, which can output position and attitude data in real time according to the RGB images obtained by the camera, effectively solve the problem of low positioning accuracy without GPS in special cases, and at the same time can achieve navigation positioning on the in-vehicle camera platform, and can be widely applied to the positioning and navigation service of auxiliary airport unmanned vehicles. Brief Description of the Drawings
[0049] Figure 1 It is the overall flowchart of a vision positioning method for auxiliary airport unmanned vehicles based on meta-learning of the present invention;
[0050] Figure 2 It is a schematic diagram of the scenario regression vision positioning algorithm for optimizing and integrating the differentiable random sample consensus algorithm by meta-learning. Detailed Embodiment
[0051] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0052] A vision positioning and navigation method for auxiliary airport unmanned vehicles based on meta-learning of the present invention, as Figure 1 shown, the present invention is an indoor positioning method that fuses vision and inertia, and includes the following steps:
[0053] S11 The camera acquisition unit carried by the airport unmanned vehicle acquires RGB images, and performs scene coordinate regression on the acquired RGB images through the random sample consensus method optimized by meta-learning to obtain the pose prediction of the center point of the RGB image, thereby obtaining the pose of the unmanned vehicle equipped with the camera.
[0054] The schematic diagram of the scenario regression vision positioning algorithm for optimizing and integrating the differentiable random sample consensus algorithm by meta-learning is as Figure 2 shown, the random sample consensus method optimized by meta-learning includes:
[0055] S11-1 Scene coordinate regression, predicting the 3D scene coordinates y i (ω) corresponding to each pixel i of the two-dimensional RGB image through the basic convolutional neural network ConvNeXt, where ω is the parameter of the neural network ConvNeXt model, and realizing the conversion from the two-dimensional pixel point (x i , y i ) of the image to the 3D scene coordinates through this model: p is the coordinate of the two-dimensional pixel point (x i , y i ) in the pixel coordinate system, P C is the coordinate of the point in the camera coordinate system, P W is the coordinate of the point in the world coordinate system, ω is the depth of the point, K is the internal parameter matrix of the camera, R CW and It is the pose transformation from the world coordinate system to the camera coordinate system.
[0056] Let f(I; ω) represent the mapping from the 2D image camera coordinate system to the scene coordinate system, where I represents a given image and ω is the parameter that the ConvNeXt neural network needs to learn. The learnable parameter ω is optimized by minimizing the expected pose loss of the final estimate on the training set.
[0057]
[0058] In the formula, f * represents the ground truth of the pose of image I. In order to continuously optimize the parameter ω of the neural network through gradient descent, the derivative of the parameter ω is calculated. The partial derivative of the above formula is:
[0059]
[0060] By using the neural network in the first step, the prediction of the scene coordinates is realized, that is, the mapping from the 2D image coordinate points to the 3D scene coordinate system, which prepares for the next robust pose optimization.
[0061] S11-2 First, a large number of hypotheses are obtained. The image I and the scene coordinate prediction Y define a dense correspondence set C on all image pixels i. As the first step of the robust pose optimization, we randomly select the smallest subset C of M correspondences RGB ={(p i , y i ) | y i ∈Y}, 0 ≤ i < M. Each C j corresponds to a camera pose hypothesis h j , and we use a pose solver to recover it, that is:
[0062] h j = g(C j )
[0063] The relationship between the 2D image coordinate points and the 3D scene coordinate system can be represented as follows:
[0064] p i = Kh -1 y i
[0065] In the formula, K is the camera calibration matrix, or the internal calibration parameter of the camera. Using this relationship, bundle adjustment in the perspective PnP solver recovers the camera pose from at least 4 2D-3D correspondences to obtain a unique solution: C j ≥ 4. In practice, 4 pairs of image-to-scene coordinates are randomly selected for prediction. The input used in the invention is an RGB image, so the pose solution recovery is:
[0066]
[0067] After that, a large number of collected hypotheses are selected, and the evaluation principle for each hypothesis is as follows:
[0068]
[0069] In the formula, the function s(·) is the evaluation function, indicating the performance of each hypothesis model. The evaluation function s(·) is determined according to the applied scenario and input as:
[0070]
[0071] In the formula, the function r(·) measures the residual between the pose parameter h and the scene coordinates. If the residual is less than the inner threshold τ, Ι[·] is calculated as 1. The residual function r(·) is specifically expressed as:
[0072] r(y i ,h)=||p i -Kh -1 y i ||
[0073] After continuously iteratively selecting the optimal hypothesis model, the optimal model needs to be optimized using all scene coordinates:
[0074]
[0075] If you want to obtain a pose prediction model through an end-to-end training method, the entire process needs to be differentiable. Therefore, the above optimization process is made differentiable through the Gauss-Newton method:
[0076]
[0077] In the formula is the pseudo-inverse of the Jacobian matrix J i of the previous residual function vector r(y r . Specifically, the Jacobian matrix J r consists of the following partial derivatives:
[0078]
[0079] After differentiability, the entire model can be continuously iteratively trained, and the above robust pose optimization process is continuously iterated. Finally, a camera pose matrix with high accuracy can be obtained.
[0080] S11 - 3 - Meta - learning optimization training. Generally, if only the process of robust pose optimization is used to train the model, the obtained effect is very poor. This is because the model requires a process of predicting scene coordinates by the neural network in the first half of the initialization. Therefore, the training needs to be divided into two steps. Although this can prevent the robust pose estimation from affecting the neural network scene coordinate mapping process, the training process is rather cumbersome and requires a large amount of time for training. Therefore, the present invention enables the fusion of the two training processes through meta - learning, achieving a fast pose model estimation process while ensuring the accuracy of the neural network in predicting scene coordinates.
[0081] The specific method for realizing the fusion training of steps S11 - 1 and S11 - 2 is to use the MAML method to divide the training process into an inner loop and an outer loop, as follows:
[0082] For model initialization, introduce MAML to initialize the model parameters and enable the model to have a good initial performance. Divide the training process of the neural network into an inner loop and an outer loop. The inner loop realizes the training of the basic model functions, and the outer loop improves the generalization training of the model. Similarly, divide the training set into two parts for the training of the two loops;
[0083] Among them, the inner loop is to implement the mapping of scene coordinates in step S11 - 1, realizing the mapping from the 2D camera coordinates to the 3D scene coordinates. The optimization of the inner loop is as follows:
[0084]
[0085] The specific parameter iteration process of the inner loop is expressed as:
[0086] ω' = ω - μ▽ ω L(f(I, ω))
[0087] In the formula, ω' is the optimal parameter of the inner - loop iteration, ω is the initialization parameter, μ is the learning rate in the inner - loop training process, and ▽ ω L(f ω ) represents the gradient of the inner - loop implementation of the scene coordinate regression function.
[0088] For the outer loop, it is mainly the robust pose optimization process of step S11 - 2. Therefore, the optimization process of the inner loop is as follows:
[0089]
[0090] Regarding the transfer of inner loop and outer loop parameters, the specific parameter transfer method is to perform meta-update or meta-optimization before sampling the next batch of tasks. By training the inner loop, the optimal parameter ω' is found. Then, the gradient of ω' for each inner loop is calculated, and the randomly initialized parameter ω is updated by the gradient update method. This enables the randomly initialized parameter ω to find an initial parameter that is more approximate to the target task, so that when training the next batch of tasks, not many gradient steps need to be taken;
[0091] The whole process is described as follows:
[0092]
[0093] In the formula, ω is the initialized parameter, and η is the learning rate hyperparameter of the outer loop. For the robust pose optimization of the outer loop, assume the gradient of the number ω' of the model optimization function;
[0094] S21 obtains the current pose of the camera by visual positioning of the images captured by the on-vehicle camera of the airport unmanned vehicle. Among them, the coordinate regression method of the differentiable random sample consensus algorithm based on meta-learning has very high accuracy, and the pose accuracy error can be within 5 cm and 5°. It can be used as an effective auxiliary means for the unmanned vehicle to drive.
[0095] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.
Claims
1. A visual positioning method for an auxiliary airport unmanned vehicle based on meta - learning, characterized in that, the specific steps are as follows: S11 The camera acquisition unit carried by the airport unmanned vehicle acquires RGB images, and performs scene coordinate regression on the acquired RGB images through the random sample consensus method optimized by meta - learning to obtain the pose prediction of the center point of the RGB image, thereby obtaining the pose of the unmanned vehicle equipped with the camera. The above - mentioned random sample consensus method optimized by meta - learning includes: S11-1 Scene coordinate regression, predicting the 3D scene coordinates y corresponding to each pixel i of the two-dimensional RGB image through the basic convolutional neural network ConvNeXt i (ω), where ω are the parameters of the neural network ConvNeXt model, and through this model, the two-dimensional pixel points (x i , y i ) of the image are mapped to the 3D scene coordinates: p is the coordinate of the two-dimensional pixel point (x i , y i ) in the pixel coordinate system, P C is the coordinate of the point in the camera coordinate system, P W is the coordinate of the point in the world coordinate system, ω is the depth of the point, K is the internal parameter matrix of the camera, R CW and are the pose transformations from the world coordinate system to the camera coordinate system; Use f(I; ω) to represent the mapping from the 2D image camera coordinate system to the scene coordinate system, where I represents a given image, and ω is the parameter that the ConvNeXt neural network needs to learn. Optimize the learnable parameter ω by minimizing the expected pose loss l estimated finally on the training set: where f * represents the true value of the pose of image I. In order to continuously optimize the parameter ω of the neural network through gradient descent and take the derivative of parameter ω, the partial derivative of the above formula is as follows: Through the first step, use the neural network to realize the prediction of scene coordinates, that is, the mapping from 2D image coordinate points to the 3D scene coordinate system, to prepare for the next robust pose optimization; S11-2 First, obtain a large number of hypotheses h, where each hypothesis h depends on the parameters of the corresponding scene coordinates. The image I and the scene coordinate prediction Y define a dense set of correspondences C over all image pixels i. As the first step of robust pose optimization, randomly select the smallest subset C of M correspondences RGB ={(p i , y i ) | y i ∈Y}, 0 ≤ i < M, and each C j corresponds to a camera pose hypothesis h j , and use a pose solver to recover it, that is: h j = g(C j ) After that, a large number of hypotheses collected are selected, and the judgment principle for each hypothesis is: The s(·) function in the formula is an evaluation function, indicating the performance of each hypothesis model; After continuously iteratively selecting the optimal hypothesis model, the optimal model needs to be optimized, and all scene coordinates are used to optimize the optimal model: Continuously iterate the above - mentioned robust pose optimization process, and finally obtain a camera pose matrix with high accuracy; S11 - 3 Meta - learning optimization training, through meta - learning, fuse the two training processes, and on the basis of ensuring the accuracy of the neural network predicting scene coordinates, realize a fast pose model estimation process; S21 Obtain the pose of the current camera through visual positioning of the images taken by the on - vehicle camera of the airport unmanned vehicle. Among them, the coordinate regression method of the differentiable random sample consensus algorithm based on meta - learning has very high accuracy, and the pose accuracy error is within 5 cm and 5°, which is an effective auxiliary means for assisting the unmanned vehicle during driving.
2. A visual positioning method for an auxiliary airport unmanned vehicle based on meta - learning according to claim 1, characterized in that, In the selection of the random hypothesis model in step S11 - 2, since the input used is an RGB image, the pose solver is: p i = Kh -1 y i where K is the camera calibration matrix, or the camera intrinsic calibration parameters. Using this relationship, bundle adjustment in the perspective PnP solver recovers the camera pose from at least 4 2D-3D correspondences, giving a unique solution: C j ≥ 4. In practice, 4 pairs of image-to-scene coordinates are randomly selected for prediction.
3. A visual positioning method for an auxiliary airport unmanned vehicle based on meta - learning according to claim 2, characterized in that, The evaluation function s(·) in the hypothesis selection in step S11 - 2 is determined according to the applied scene and input as: In the formula, the function r(·) measures the residual between the pose parameter h and the scene coordinates. If the residual is less than the inner - layer threshold τ, then Ι[·] is calculated as 1, where the residual function r(·) is specifically expressed as:
4. A visual positioning method for an auxiliary airport unmanned vehicle based on meta - learning according to claim 3, characterized in that, The specific method for realizing the fusion training of steps S11 - 1 and S11 - 2 in step S11 - 3 is to use the MAML method to divide the training process into an inner loop and an outer loop, specifically as follows: For model initialization, MAML is introduced to initialize the model parameters and enable the model to have a better initial performance. The training process of the neural network is divided into an inner loop and an outer loop. The inner loop realizes the training of the basic model functions, and the outer loop improves the training of the model generalization. Similarly, the training set is also divided into two parts for the training of the two loops; Among them, the inner loop is to implement the mapping of scene coordinates in step S11-1, realizing the mapping from the 2D camera coordinates to the 3D scene coordinates. The optimization of the inner loop is as follows: The specific parameter iteration process of the inner loop is expressed as: Where ω' is the optimal parameter for the inner-loop iteration, ω is the initialization parameter, and μ is the learning rate for the inner-loop training process. represents the gradient of the inner-loop implementation of the scene coordinate regression function; For the outer loop, it is mainly the robust pose optimization process of step S11-2. Therefore, the optimization process of the inner loop is: Regarding the transfer of parameters between the inner loop and the outer loop, the specific parameter transfer method is to perform meta-update or meta-optimization before sampling the next batch of tasks. By training the inner loop, the optimal parameter ω' is found. Then, the gradient of ω' for each inner loop is calculated, and the randomly initialized parameter ω is updated by the gradient update method. This enables the randomly initialized parameter ω to find an initial parameter that is more approximate to the target task, so that when training the next batch of tasks, many gradient steps do not need to be taken; The whole process is described as follows: where ω is the initialization parameter, η is the learning rate hyperparameter of the outer loop, is the gradient of the number ω' of the robust pose optimization hypothesis model optimization function for the outer loop.
Citation Information
Patent Citations
Deep reinforcement learning improvement method based on adaptive hyper-parameters
CN113269322A
Foresight scene depth estimation method based on self-supervised learning
CN113313732A