A vision-guided reinforcement learning method for autonomous asteroid detection and mapping

Through visually guided reinforcement learning combined with neural radiation field model, the problem of low efficiency of autonomous detection and map construction in asteroid exploration is solved, efficient autonomous detection and map construction is achieved, and the autonomy and map construction quality of the detector are improved.

CN117570967BActive Publication Date: 2025-09-02BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311604287.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-09-02
Estimated Expiration
2043-11-28

AI Technical Summary

Technical Problem

In the existing asteroid exploration technology, autonomous detection and map construction efficiency are low. The existing methods cannot effectively integrate three-dimensional map reconstruction and autonomous detection strategies, lack evaluation of reconstruction quality and consistency, and deep reinforcement learning is slow to learn when high-dimensional image input.

Method used

The visually guided reinforcement learning method is adopted, combined with the neural radiation field model as a three-dimensional map reconstruction device, and the detector trajectory is optimized by estimating the global uncertainty, and the reinforcement learning model is designed to improve the efficiency of autonomous detection and map construction.

Benefits of technology

It realizes independent detection and efficient mapping during the asteroid orbiting process, improves the autonomy and mapping quality of the detector, and has the advantages of efficient computing and fixed memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117570967B_ABST
    Figure CN117570967B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for autonomous asteroid exploration and mapping using visually guided reinforcement learning, which belongs to the field of deep space exploration technology and includes the following steps: (1) acquiring the detector state and the asteroid surface image; (2) using the acquired detector state and the asteroid surface image to update the parameters of a neural radiation field model, mapping the asteroid surface using the neural radiation field model and estimating the mapping uncertainty; (3) using a reinforcement learning model to make decisions on detector flight adjustments based on the detector state, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model; and (4) evaluating the reinforcement learning decision effect and optimizing the reinforcement learning model parameters. This method achieves simultaneous asteroid surface mapping and flyby decision-making by fusing the neural radiation field model and the reinforcement learning model, which can improve the autonomous capability of the detector during the asteroid flyby process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep space exploration technology, and specifically relates to a method for autonomous asteroid detection and mapping using visually guided reinforcement learning. Background Art

[0002] Asteroid exploration is a crucial tool for understanding the origins of the solar system and life itself. Therefore, since the 1990s, space exploration activities aimed at the scientific exploration or on-orbit disposal of asteroids have been increasing, becoming a hot topic in deep space exploration. According to my country's overall deep space exploration plan, starting in 2025, my country will launch multiple asteroid exploration missions, including surface sampling. This will require detailed mapping and analysis of asteroid surface morphology.

[0003] Due to the long distances, irregular shapes, and limited pre-mission knowledge of asteroids, current exploration relies primarily on a repeated, iterative process of orbiting the asteroid from far to close range, using various instruments and techniques to gradually map the asteroid. For example, Japan's Hayabusa2 spent two months collecting images while orbiting and gradually approaching the asteroid Ryugu. Using structure-from-motion reconstruction, it generated a global three-dimensional model of the asteroid to support a rough screening of landing areas. Similarly, the European Space Agency's Rosetta mission mapped the surface of Comet 67P / Churyumov–Gerasimenko over a period of several months.

[0004] Because asteroid orbits require extensive, around-the-clock manpower for deployment, oversight, and planning, supporting the probe's gradual approach to the asteroid and conducting observations, the exploration cycle is long and inefficient. Consequently, major space agencies around the world, such as NASA, ESA, and JAXA, are currently focused on improving the autonomous capabilities of asteroid exploration. However, the highly uncertain and complex process of orbiting and mapping asteroids remains a significant challenge. This paper proposes a novel vision-based autonomous asteroid mapping method that simultaneously implements mapping quality assessment and decision-making by fusing neural radiation fields with reinforcement learning models, enhancing the level of autonomy in the asteroid orbit process.

[0005] Autonomous exploration and mapping have become a hot topic in some ground robot scenarios, and are known as active simultaneous mapping and localization technology (or ASLAM). ASLAM is the task of actively planning the robot's path, building a map, and positioning within it. Due to the complexity of the problem, existing methods mainly focus on how to actively plan the exploration path and build a map, while using other means for positioning. Therefore, vision-based ASLAM is often referred to as visual exploration. ASLAM robots can build maps on their own without human interaction, thus simplifying the setup and construction of navigation systems in many applications. Currently, autonomous exploration and the establishment of a three-dimensional map representation of the entire environment are often considered difficult because the computational complexity increases dramatically with the size of the exploration area and the accuracy of the map. Therefore, existing technologies mainly focus on the construction of two-dimensional maps. However, as a prerequisite for in-situ exploration such as attachment to the asteroid surface and sampling, the construction of a three-dimensional map of the asteroid surface is necessary to ensure the safe implementation of the mission.

[0006] For autonomous asteroid mapping, existing technologies and solutions typically do not build a 3D map within the iterative and optimization loop. Instead, they score probe trajectories based on known 3D models of the asteroid and then update the strategy based on the score. This is because commonly used 3D mapping methods, such as structure from motion (SfM) and multi-view geometry (MVS), are time-consuming and require feature point matching and adjustment of newly added images with the original images, which cannot meet the efficiency requirements of components within the policy iteration loop. Therefore, existing methods for scoring probe trajectories primarily evaluate image quality and novelty by calculating the relative position of the sun elevation angle, textureless model patches, and probe at a single viewpoint, and then summing the results across multiple viewpoints along the trajectory to be evaluated. This autonomous exploration strategy evaluation method, which fails to integrate the 3D mapping process, does not directly assess information gain. Specifically, it focuses on evaluating the coverage and completeness of the map, lacking an assessment of the quality and consistency of the scene reconstruction. Therefore, one of the issues addressed by this technology is how to use neural radiance fields as a 3D map reconstructor and integrate them into the autonomous exploration strategy optimization process, providing direct constraints on reconstruction quality and consistency during autonomous exploration.

[0007] On the other hand, deep reinforcement learning is a possible solution for optimizing autonomous exploration and mapping strategies and is one of the most actively researched areas in the field of visual ASLAM. However, due to the lack of a clear goal in visual ASLAM, the agent receives very limited rewards in most states of the environment. For example, completing an entire exploration process and then receiving a single feedback based on the detection result. However, in deep reinforcement learning, this results in slow learning, especially when the agent requires high-dimensional image input. Existing solutions tend to focus on maximizing the coverage and completeness of the acquired image. Summary of the Invention

[0008] The present invention proposes to encourage the agent to access the state with high uncertainty and optimal reconstruction effect, that is, by evaluating the global uncertainty when the agent is in the process of mapping, it selects the next viewpoint that is most conducive to helping map construction.

[0009] In summary, this paper proposes a method for autonomous asteroid detection and mapping using visually guided reinforcement learning, and designs a reinforcement learning model that integrates a neural radiation field model as a three-dimensional map reconstructor. It provides an autonomous visual exploration framework for asteroid flyby and accompanying flyby detection. By synchronously estimating global uncertainty when reconstructing the three-dimensional map, the neural radiation field model is used to provide information gain rewards for the optimization of reinforcement learning strategies. The proposed mapping method based on the neural radiation field model not only supports the optimization of intelligent agents, but also has the advantages of efficient computing and fixed memory usage. It can also meet the needs of map reconstruction in other common visual detection and ASLAM scenarios.

[0010] The purpose of this invention is to provide a method for autonomous asteroid exploration and mapping using vision-guided reinforcement learning, to fill the gaps in existing technology and research, and to achieve an improvement in the level of autonomy of the probe during the asteroid flyby.

[0011] An embodiment of the present invention provides a method for autonomous asteroid detection and mapping using visually guided reinforcement learning, comprising the following steps:

[0012] The first step is to obtain the status of the probe and the image of the asteroid surface;

[0013] The second step is to use the acquired detector status and asteroid surface images to update the neural radiation field model parameters, use the neural radiation field model to map the asteroid surface and estimate the uncertainty of the mapping;

[0014] The third step is to use a reinforcement learning model to make decisions on the probe's flight adjustments based on the probe's status, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model.

[0015] The fourth step is to evaluate the effectiveness of reinforcement learning decisions and optimize the reinforcement learning model parameters.

[0016] Furthermore, in the second step, the acquired detector status and asteroid surface images are used to update the parameters of the neural radiation field model. The neural radiation field model is used to map the asteroid surface and estimate the uncertainty of the mapping, including:

[0017] Calculate the camera parameters for imaging the asteroid surface based on the probe's position in the asteroid coordinate system;

[0018] Determine the sampling light according to the camera parameters and the asteroid surface image when imaging the asteroid surface image, and perform sampling on the sampling light to obtain the sampling points;

[0019] Dividing the coordinates of the sampling points into a plurality of grids with different resolutions, and fusing the grid vertex features to represent the features of the sampling points;

[0020] According to the characteristics of the sampling points, a volume rendering algorithm is used to estimate the color and uncertainty of the light, and a back-propagation method is used to adjust the mesh vertex characteristics and neural radiation field model parameters.

[0021] Furthermore, in the third step, a reinforcement learning model is used to make decisions on probe flight adjustments based on the probe status, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model, including:

[0022] Build a reinforcement learning model with action decision-making and action value estimation capabilities;

[0023] The probe state, asteroid surface image, and mapping uncertainty estimated by the neural radiation field model are jointly input into the reinforcement learning model to make decisions in the action space.

[0024] Furthermore, in the fourth step, evaluating the reinforcement learning decision effect and optimizing the reinforcement learning model parameters include:

[0025] Based on the decision-making of the reinforcement learning model, the detector status and asteroid surface images are further obtained;

[0026] Based on the newly acquired detector status and asteroid surface images, the neural radiation field model parameters are updated to obtain updated mapping results and their uncertainty.

[0027] Based on the uncertainty of the updated mapping results, the quality of the reinforcement learning model decision is evaluated, and the reinforcement learning model parameters are adjusted using the backpropagation algorithm.

[0028] Furthermore, in the second step, the camera parameters for imaging the asteroid surface image are calculated based on the position of the probe in the asteroid coordinate system. First, the three orthogonal bases of the direction vectors pointing from the current position of the probe (x, y, z) to the origin are calculated. The forward vector f, the right vector r, the up vector up, and the rotation matrix R during imaging are expressed as follows:

[0029]

[0030]

[0031] up=r×f (3)

[0032]

[0033] Among them, f x 、f y 、f z are the components of f in the x-axis, y-axis, and z-axis directions, r x 、r y 、r z They are the components of r in the x-axis, y-axis, and z-axis directions, up x 、up y up z are the components of up in the x-axis, y-axis, and z-axis directions respectively; the imaging relationship of the asteroid surface image is expressed using the following formula:

[0034]

[0035] Where K is the intrinsic parameter matrix of the imaging camera carried by the detector, [R|t] is the extrinsic parameter matrix of the imaging camera carried by the detector, where t = (x, y, z)'; (u, v) is the pixel coordinate of the image point (X, Y, Z) projected onto the image.

[0036] Furthermore, in the second step, the sampling light is determined according to the camera parameters and the asteroid surface image when the asteroid surface image is imaged. The implementation method of sampling on the sampling light to obtain the sampling point is as follows: for each pixel coordinate (u, v) on each image, according to all three-dimensional points that satisfy the non-homogeneous equation (5), it is a light ray starting from the optical center of the imaging camera. The light ray equation is r(t) = o + td, where o is the starting point of the light ray, d is the direction of the light ray, and t is the propagation distance of the light ray. The light direction is expressed as d = (θ, Φ), where θ, Φ are the azimuths corresponding to the direction vectors. In the light ray r(t) = o + td, the value range of t is [t n ,t f ] uniform or random sampling is performed in the interval to obtain N sampling points, each sampling point is expressed in the form of (X, Y, Z, θ, Φ), which represents its three-dimensional coordinates and observation direction.

[0037] Furthermore, in the second step, the coordinates of the sampling points are divided into multiple grids of different resolutions, and the grid vertex features are fused to represent the features of the sampling points. The implementation method is to divide the space into multiple grids of different resolutions according to a preset interval, and the division method is to divide the space into a combination of cubic grids with side lengths of preset intervals. According to the coordinates (X, Y, Z) of the sampling points, the index of the grids of different resolutions can be queried, and then the coordinates of the vertices of the grid containing the sampling points can be obtained, and the grid vertex features can be queried. The query method for the grid vertex features at one resolution is to hash the grid vertex coordinates to obtain the hash value corresponding to the vertex coordinates, and obtain the value of the grid vertex feature at the resolution by accessing the memory address corresponding to the hash value; use the linear interpolation method to calculate the features of the sampling points in the grid of the resolution; and obtain the features of the sampling points by splicing the features from the grids of different resolutions.

[0038] Furthermore, in the second step, based on the features of the sampling points, a volume rendering algorithm is used to estimate the light color, object distribution and uncertainty, and a back propagation method is used to adjust the mesh vertex features and the neural radiation field model parameters. The implementation method is to input the acquired sampling point features into the neural radiation field model using a multi-layer perceptron F with a weight parameter Θ Θ In the fitting process, the color c = (r, g, b), volume density σ and uncertainty β of the three-dimensional position (X, Y, Z) of the asteroid reconstruction area along the light direction d = (θ, Φ) are observed, and the multi-layer perceptron F Θ Denoted as:

[0039] (r,g,b,σ,β)=F Θ (f(X,Y,Z,θ,Φ)) (6)

[0040] Among them, f(X, Y, Z, θ, Φ) is the concatenation of the sampling point feature and the light direction d = (θ, Φ). The neural radiation field model uses the color value of the asteroid surface image as supervision and adjusts the multi-layer perceptron F in the neural radiation field model. Θ The parameter Θ makes the multilayer perceptron F Θ It can fit the relationship between the three-dimensional position (X, Y, Z) in the direction d = (θ, Φ) and the observed color c, volume density σ, and uncertainty β; in this step, the neural radiation field model uses volume rendering to fuse the color, volume density, and uncertainty estimation of different samples on the sampling light. The sampling point light r(t) = o + td starts from o and moves along the direction d. The volume density σ of the three-dimensional position (X, Y, Z) represents the probability that the light r(t) terminates at this point. The multi-layer perceptron F in the neural radiation field model is converted into Θ The color, transparency and uncertainty of the query at the i-th sampling point are recorded as ci=(ri,gi,bi), σ i and βi The neural radiance field model generates color estimates by fusing sampled values ​​through volume rendering.

[0041]

[0042] Among them, δ in formula (7) i =t i+1 -t i , is the distance between two adjacent sampling points, T i is the ray traveling from the near boundary to the sampling point t i Estimation of the cumulative transmittance:

[0043]

[0044] The parameter tuning process of the neural radiation field model uses back propagation and stochastic gradient descent. A batch of N pixels are randomly selected from the input training set image, and the light sets corresponding to these pixels are calculated. The color estimation values ​​of these pixels are calculated according to formulas (5)(7)(8): The second-order moment of the residual between the pixel color C(r) and the true pixel color, and the optimization of the uncertainty estimation using the Bayesian estimation principle, are used as the loss function for the back propagation optimization of the neural radiation field model parameters:

[0045]

[0046] Among them, r i It represents the i-th sampling light in the batch of pixels. The loss function has the function of optimizing the color, density distribution and uncertainty distribution simultaneously. By using the neural radiation field model to calculate the color, density and uncertainty of the mesh vertices in the reconstructed area, a three-dimensional mesh with color, density and uncertainty can be obtained. This mesh is stored in the form of voxels, thereby completing the mapping of the asteroid and obtaining the three-dimensional map of the asteroid and the uncertainty of the three-dimensional map.

[0047] Furthermore, the implementation method of constructing a reinforcement learning model with value estimation function in the third step is to construct a reinforcement learning model with an action space dimension of 2, and the dimensions of the action space are set to the change in the orbital altitude of the spacecraft relative to the previous time step and the change in the current orbital inclination of the spacecraft; the orbital altitude h is the distance between the spacecraft and the asteroid, and the orbital inclination θ represents the angle between the current probe position and the line connecting the origin and the xOy plane in the asteroid body coordinate system, and its value range is [-π,π]; the reinforcement learning model includes a decision maker and a critic. The role of the decision maker is to generate the optimal action for the current environmental state, and the work of the critic part is to predict the long-term return of the action selected by the decision maker.

[0048] Furthermore, in the third step, the detector state, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model are input into the decision maker of the reinforcement learning model, and the decision-making in the action space is implemented as follows: the mapping uncertainty estimated by the neural radiation field model is encoded into an image with the same perspective and size as the asteroid surface image, and is input into the convolutional neural network of the decision maker as an additional dimension of the asteroid surface image; the convolutional neural network of the decision maker contains 4 convolutional layers, each convolutional layer is combined with a Squeeze-Extract layer, a BatchNorm layer, a pooling layer, and a nonlinear layer to obtain a high-dimensional feature vector, which is concatenated with the vector of the detector state, and the output action is calculated by a multi-layer perceptron. The output action dimension is 2, representing a set of actions in the action space.

[0049] The above technical solution of the present invention overcomes the disadvantages of poor autonomy and low mapping efficiency during the asteroid flyby process, and realizes efficient, rapid and autonomous asteroid detection and mapping, with the following beneficial technical effects:

[0050] (1) A vision-guided autonomous asteroid mapping model is proposed, and a reinforcement learning model that integrates a neural radiation field model as a three-dimensional map reconstructor is designed, providing an autonomous visual exploration framework for the asteroid flyby and companion exploration process.

[0051] (2) A mapping method that supports iterative optimization of intelligent agents in visual detection and ASLAM tasks is proposed. This method is based on the neural radiation field model and has the advantages of efficient computing and fixed memory usage. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is the overall flow chart of the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of this application more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present invention.

[0054] like Figure 1 As shown, an embodiment of the present invention provides a method for autonomous asteroid detection and mapping using visually guided reinforcement learning. The method may include:

[0055] S110: Obtain the status of the probe and the image of the asteroid surface;

[0056] S120: Update the parameters of the neural radiation field model using the acquired detector status and the asteroid surface image, map the asteroid surface using the neural radiation field model, and estimate the uncertainty of the mapping;

[0057] S130: Using a reinforcement learning model to make decisions on probe flight adjustments based on the probe state, asteroid surface images, and mapping uncertainty estimated by the neural radiation field model;

[0058] S140: Evaluate the effectiveness of reinforcement learning decisions and optimize reinforcement learning model parameters.

[0059] The above-mentioned embodiment method collects the detector status including position and attitude and the asteroid surface image; calculates the camera intrinsic and extrinsic parameters during imaging based on the detector status and the asteroid surface image; and uses the asteroid surface image and the camera intrinsic and extrinsic parameters during imaging as input to construct a neural radiation field model to achieve reconstruction of the asteroid surface environment structure and estimation of uncertainty; thereby, the estimated uncertainty is jointly input into the reinforcement learning model with the detector status and the asteroid surface image to achieve decision-making on adjusting the detector flight state; further, the reinforcement learning model evaluates the decision by changes in uncertainty and other constraints, thereby optimizing the reinforcement learning model parameters.

[0060] The method in this embodiment is a method for autonomous asteroid exploration and mapping using visually guided reinforcement learning. Experimental verification results in a Unity3D simulation environment built based on 3D models of asteroids targeted by multiple international missions demonstrate that the method in this embodiment can provide a decision-making basis for autonomous exploration with mapping as the goal during the flyby and orbit phases of asteroid exploration.

[0061] In some embodiments, in S120, updating the neural radiation field model parameters based on the acquired detector status and the asteroid surface image includes:

[0062] The camera parameters for imaging the asteroid surface are calculated based on the probe's position in the asteroid coordinate system. The implementation is to first calculate the three orthogonal bases of direction vectors pointing from the probe's current position (x, y, z) to the origin. The forward vector f, right vector r, up vector up, and rotation matrix R are expressed as:

[0063]

[0064]

[0065] u=r×f (3)

[0066]

[0067] Among them, f x、f y 、f z are the components of f in the x-axis, y-axis, and z-axis directions, r x 、r y 、r z They are the components of r in the x-axis, y-axis, and z-axis directions, up x 、up y up z are the components of up in the x-axis, y-axis, and z-axis directions respectively. The imaging relationship of the asteroid surface is expressed using the following formula:

[0068]

[0069] Where K is the intrinsic parameter matrix of the imaging camera carried by the detector, [R|t] is the extrinsic parameter matrix of the imaging camera carried by the detector, where t = (x, y, z)'; (u, v) is the pixel coordinate of the image point (X, Y, Z) projected onto the image.

[0070] According to the camera parameters and the asteroid surface image when the asteroid surface image is imaged, the sampling light is determined, and sampling is performed on the sampling optics to obtain the sampling points. The implementation method is that for each pixel (u, v) on each image, according to all three-dimensional points that satisfy the non-homogeneous equation (5), it is a light ray starting from the optical center of the imaging camera, and the light ray equation is r(t) = o + td. Where o is the starting point of the light ray, d is the direction of the light ray, and t is the propagation distance of the light ray. The direction of the light ray can be expressed as d = (θ, Φ), where θ, Φ are the azimuth angles corresponding to the direction vector. In the light ray r(t) = o + td, the value range of t is [t n ,t f ] uniform or random sampling is performed in the interval to obtain N sampling points. Each sampling point can be expressed in the form of (X, Y, Z, θ, Φ), which represents its three-dimensional coordinates and observation direction.

[0071] The coordinates of the sampling points are divided into a plurality of grids of different resolutions, and the features of the grid vertices are fused to represent the features of the sampling points. The implementation method is to divide the space into a plurality of grids of different resolutions according to a preset interval. The division method is to divide the space into a combination of cubic grids with side lengths of preset intervals. According to the coordinates (X, Y, Z) of the sampling point, the index of the grid contained in the grids of different resolutions can be queried, and then the coordinates of the vertices of the grid containing it can be obtained, and the features of the grid vertices can be queried. The query method for the vertex features of the grid at one resolution is to hash the coordinates of the grid vertices to obtain the hash value corresponding to the vertex coordinates, and obtain the value of the vertex features of the grid at the resolution by accessing the memory address corresponding to the hash value. The features of the sampling points in the grid of the resolution are calculated using the linear interpolation method. The features of the sampling points are obtained by splicing the features from the grids of different resolutions.

[0072] Based on the characteristics of the sampling points, a volume rendering algorithm is used to estimate the light color, object distribution and uncertainty, and a back propagation method is used to adjust the mesh vertex characteristics and neural radiation field model parameters. The implementation method is to input the obtained sampling point features into the neural radiation field model using a multi-layer perceptron F with a weight parameter Θ Θ In the 3D position (X, Y, Z) of the asteroid reconstruction region, the observed color c = (r, g, b), volume density σ and uncertainty β along the light direction d = (θ, Φ) are fitted. Multilayer Perceptron F Θ It can be written as:

[0073] (r,g,b,σ,β)=F Θ (f(X,Y,Z,θ,Φ)) (6)

[0074] Where f(X, Y, Z, θ, Φ) is the concatenation of the sampling point feature and the light direction d = (θ, Φ). The neural radiation field model uses the color value of the asteroid surface image as supervision and adjusts the multi-layer perceptron F in the neural radiation field model. Θ The parameter Θ makes the multilayer perceptron F Θ It can fit the relationship between the three-dimensional position (X, Y, Z) in the direction d = (θ, Φ) and the observed color c, volume density σ, and uncertainty β. In this step, the neural radiation field model uses volume rendering to fuse the color, volume density, and uncertainty estimates of different samples on the sampling light. In the discretized volume rendering theory, the sampling point light r(t) = o + td starts from o and travels along the direction d. The volume density σ of the three-dimensional position (X, Y, Z) represents the probability that the light r(t) terminates at this location. The multi-layer perceptron F in the neural radiation field model is used. Θ The color, transparency and uncertainty of the query at the i-th sampling point are recorded as c i =(r i ,gi ,b i ),σ i and β i The neural radiance field model generates color estimates by fusing sampled values ​​through volume rendering.

[0075]

[0076] Among them, δ in formula (7) i =t i+1 -t i , is the distance between two adjacent sampling points, T i is the ray traveling from the near boundary to the sampling point t i Estimation of the cumulative transmittance:

[0077]

[0078] The parameter tuning process of the neural radiation field model uses back propagation and stochastic gradient descent. A batch of N pixels are randomly selected from the input training set image, and the light sets corresponding to these pixels are calculated. The color estimation values ​​of these pixels are calculated according to formulas (5)(7)(8): The second-order moment of the residual between the pixel color C(r) and the true pixel color, and the optimization of the uncertainty estimation using the Bayesian estimation principle, are used as the loss function for the back propagation optimization of the neural radiation field model parameters:

[0079]

[0080] Among them, r i represents the i-th sampled ray in the batch of pixels. This loss function simultaneously optimizes the color, density, and uncertainty distributions. By using a neural radiance field model to calculate the color, density, and uncertainty of the mesh vertices in the reconstructed area, a 3D mesh with color, density, and uncertainty can be obtained. This mesh can be stored as voxels, completing the mapping of the asteroid and obtaining a 3D map of the asteroid and its uncertainty.

[0081] In some embodiments, in S130, the decision on probe flight adjustment using a reinforcement learning model based on the probe state, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model is implemented by the following method:

[0082] Construct a reinforcement learning model with value estimation capabilities. The implementation method is to construct a reinforcement learning model with an action space dimension of 2, where the dimensions of the action space are set to the change in the spacecraft's orbital altitude and the spacecraft's current orbital inclination relative to the previous time step. The orbital altitude h is the distance between the spacecraft and the asteroid. The orbital inclination θ represents the angle between the current probe position and the line connecting the origin and the xOy plane in the asteroid's coordinate system, and its value range is [-π,π]. The reinforcement learning model used in this technology can be constructed using various reinforcement learning algorithms, including deep deterministic policy gradient algorithms. The reinforcement learning model includes a decision maker and a critic. The decision maker's role is to generate the optimal action for the current environmental state, while the critic's role is to predict the long-term reward of the action selected by the decision maker.

[0083] The probe state, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model are input to the decision maker of the reinforcement learning model, which then makes a decision within the action space. The mapping uncertainty estimated by the neural radiation field model is encoded into an image with the same viewing angle and size as the asteroid surface image. This is then fed into the decision maker's convolutional neural network as an additional dimension to the asteroid surface image. The decision maker's convolutional neural network consists of four convolutional layers, each of which is combined with a squeeze-extract layer, a batch-normalized layer, a pooling layer, and a nonlinear layer to produce a high-dimensional feature vector. This high-dimensional feature vector is concatenated with the probe state vector, and the output action is calculated using a multilayer perceptron. The output action has a dimension of 2, representing a set of actions within the action space.

[0084] In some embodiments, in S140, evaluating the reinforcement learning decision effect and optimizing the reinforcement learning model parameters are achieved by the following method:

[0085] Based on the reinforcement learning model's decisions, the probe's state and an image of the asteroid's surface are further acquired. This is accomplished by applying the actions output by the decision maker to a probe in a simulation or a physical probe. This process then acquires the probe's state after a predetermined period of operation, along with an image of the asteroid's surface acquired by the probe during that period.

[0086] Based on the newly acquired detector state and asteroid surface image, the neural radiation field model parameters are updated, and an updated map construction result and the uncertainty of the map construction result are obtained. The implementation method is to update the neural radiation field model parameters according to the method for updating the neural radiation field model parameters in S120 based on the newly acquired detector state and asteroid surface image, and obtain an updated map construction result and the uncertainty of the map construction result.

[0087] The quality of the reinforcement learning model's decisions is evaluated based on the uncertainty of the updated map, and the model parameters are adjusted using a backpropagation algorithm. The difference between the average uncertainty of the updated map and the average uncertainty of the previous map is used as the average reduction in the uncertainty of the map. This reduction serves as a reward for the reinforcement learning model. The larger the reward, the higher the quality of the reinforcement learning model's decisions. The reinforcement learning model calculates the gradient of backpropagation and updates the model parameters with the goal of maximizing this reward.

[0088] The above embodiment provides an autonomous visual exploration framework for asteroid flyby and flyby exploration, forming a reusable loop. By simultaneously estimating global uncertainty during 3D map reconstruction, the neural radiance field model provides information gain rewards for optimizing reinforcement learning strategies. The proposed mapping method based on the neural radiance field model not only supports agent optimization but also offers the advantages of efficient computation and fixed memory usage. It can also meet the map reconstruction requirements of other common visual detection and ASLAM scenarios.

[0089] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0091] The contents not described in detail in the specification of the present invention belong to the prior art known to those skilled in the art.

[0092] The above embodiments are provided for the purpose of describing the present invention only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention are intended to be within the scope of the present invention.

Claims

1. A method for autonomous asteroid detection and mapping using visually guided reinforcement learning, characterized in that: The following steps are involved: The first step is to obtain the status of the probe and the image of the asteroid surface; The second step is to use the acquired detector status and asteroid surface images to update the neural radiation field model parameters, use the neural radiation field model to map the asteroid surface and estimate the uncertainty of the mapping; The third step is to use a reinforcement learning model to make decisions on the probe's flight adjustments based on the probe's status, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model. The fourth step is to evaluate the effectiveness of reinforcement learning decisions and optimize the reinforcement learning model parameters.

2. The method according to claim 1, wherein: In the second step, the acquired detector status and asteroid surface images are used to update the parameters of the neural radiation field model. The neural radiation field model is used to map the asteroid surface and estimate the uncertainty of the mapping, including: Calculate the camera parameters for imaging the asteroid surface based on the probe's position in the asteroid coordinate system; Determine the sampling light according to the camera parameters and the asteroid surface image when imaging the asteroid surface image, and perform sampling on the sampling light to obtain the sampling points; Dividing the coordinates of the sampling points into a plurality of grids with different resolutions, and fusing the grid vertex features to represent the features of the sampling points; According to the characteristics of the sampling points, a volume rendering algorithm is used to estimate the color and uncertainty of the light, and a back-propagation method is used to adjust the mesh vertex characteristics and neural radiation field model parameters.

3. The method according to claim 1, wherein: In the third step, a reinforcement learning model is used to make decisions on probe flight adjustments based on the probe status, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model, including: Build a reinforcement learning model with action decision-making and action value estimation capabilities; The probe state, asteroid surface image, and mapping uncertainty estimated by the neural radiation field model are jointly input into the reinforcement learning model to make decisions in the action space.

4. The method according to claim 1, wherein: In the fourth step, the reinforcement learning decision effect is evaluated and the reinforcement learning model parameters are optimized, including: Based on the decision-making of the reinforcement learning model, the detector status and asteroid surface images are further obtained; Based on the newly acquired detector status and asteroid surface images, the neural radiation field model parameters are updated to obtain updated mapping results and their uncertainty. Based on the uncertainty of the updated mapping results, the quality of the reinforcement learning model decision is evaluated, and the reinforcement learning model parameters are adjusted using the backpropagation algorithm.

5. The method according to claim 2, characterized in that The second step is to calculate the camera parameters for imaging the asteroid surface based on the probe's position in the asteroid coordinate system. First, the three orthogonal bases of the direction vectors pointing from the probe's current position (x, y, z) to the origin are calculated. The forward vector f, right vector r, up vector up, and rotation matrix R are expressed as follows: up=r×f (3) Among them, f x 、f y 、f z are the components of f in the x-axis, y-axis, and z-axis directions, r x 、r y 、r z They are the components of r in the x-axis, y-axis, and z-axis directions, up x 、up y up z are the components of up in the x-axis, y-axis, and z-axis directions respectively; the imaging relationship of the asteroid surface image is expressed using the following formula: Where K is the intrinsic parameter matrix of the imaging camera carried by the detector, [R|t] is the extrinsic parameter matrix of the imaging camera carried by the detector, where t = (x, y, z)'; (u, v) is the pixel coordinate of the image point (X, Y, Z) projected onto the image.

6. The method according to claim 5, characterized in that In the second step, the sampling light is determined according to the camera parameters and the asteroid surface image when the asteroid surface image is imaged. The implementation method of sampling on the sampling light to obtain the sampling point is as follows: for each pixel coordinate (u, v) on each image, according to all three-dimensional points that satisfy the non-homogeneous equation (5), it is a light ray starting from the optical center of the imaging camera. The light ray equation is r(t) = o + td, where o is the starting point of the light ray, d is the direction of the light ray, and t is the propagation distance of the light ray. The light direction is expressed as d = (θ, Φ), where θ, Φ are the azimuths corresponding to the direction vectors. In the light ray r(t) = o + td, the value range of t is [t n ,t f ] uniform or random sampling is performed in the interval to obtain N sampling points, each sampling point is expressed in the form of (X, Y, Z, θ, Φ), which represents its three-dimensional coordinates and observation direction.

7. The method according to claim 6, characterized in that In the second step, the coordinates of the sampling points are divided into multiple grids of different resolutions, and the grid vertex features are fused to represent the features of the sampling points. The implementation method is to divide the space into multiple grids of different resolutions according to a preset interval. The division method is to divide the space into a combination of cubic grids with side lengths of preset intervals. According to the coordinates (X, Y, Z) of the sampling points, the index of the grids of different resolutions can be queried, and then the coordinates of the vertices of the grid containing it can be obtained, and the grid vertex features can be queried; the query method for the grid vertex features at one resolution is to hash the grid vertex coordinates to obtain the hash value corresponding to the vertex coordinates, and obtain the value of the grid vertex feature at the resolution by accessing the memory address corresponding to the hash value; use the linear interpolation method to calculate the features of the sampling points in the grid of the resolution; obtain the features of the sampling points by splicing the features from the grids of different resolutions.

8. The method according to claim 7, characterized in that In the second step, based on the features of the sampling points, the volume rendering algorithm is used to estimate the light color, object distribution and uncertainty, and the back propagation method is used to adjust the mesh vertex features and the neural radiation field model parameters. The implementation method is to input the acquired sampling point features into the neural radiation field model using a multi-layer perceptron F with a weight parameter Θ. Θ In the fitting process, the color c = (r, g, b), volume density σ and uncertainty β of the three-dimensional position (X, Y, Z) of the asteroid reconstruction area along the light direction d = (θ, Φ) are observed, and the multi-layer perceptron F Θ Denoted as: (r,g,b,σ,β)=F Θ (f(X,Y,Z,θ,Φ)) (6) Among them, f(X, Y, Z, θ, Φ) is the concatenation of the sampling point feature and the light direction d = (θ, Φ). The neural radiation field model uses the color value of the asteroid surface image as supervision and adjusts the multi-layer perceptron F in the neural radiation field model. Θ The parameter Θ makes the multilayer perceptron F Θ It can fit the relationship between the three-dimensional position (X, Y, Z) in the direction d = (θ, Φ) and the observed color c, volume density σ, and uncertainty β; in this step, the neural radiation field model uses volume rendering to fuse the color, volume density, and uncertainty estimation of different samples on the sampling light. The sampling point light r(t) = o + td starts from o and moves along the direction d. The volume density σ of the three-dimensional position (X, Y, Z) represents the probability that the light r(t) terminates at this point. The multi-layer perceptron F in the neural radiation field model is converted into Θ The color, transparency and uncertainty of the query at the i-th sampling point are recorded as ci=(ri,gi,bi), σ i and β i The neural radiance field model generates color estimates by fusing sampled values ​​through volume rendering. Among them, δ in formula (7) i =t i+1 -t i , is the distance between two adjacent sampling points, T i is the ray traveling from the near boundary to the sampling point t i Estimation of the cumulative transmittance: The parameter tuning process of the neural radiation field model uses back propagation and stochastic gradient descent. A batch of N pixels are randomly selected from the input training set image, and the light sets corresponding to these pixels are calculated. The color estimation values ​​of these pixels are calculated according to formulas (5)(7)(8): The second-order moment of the residual between the pixel color C(r) and the true pixel color, and the optimization of the uncertainty estimation using the Bayesian estimation principle, are used as the loss function for the back propagation optimization of the neural radiation field model parameters: Among them, r i It represents the i-th sampling light in the batch of pixels. The loss function has the function of optimizing the color, density distribution and uncertainty distribution simultaneously. By using the neural radiation field model to calculate the color, density and uncertainty of the mesh vertices in the reconstructed area, a three-dimensional mesh with color, density and uncertainty can be obtained. This mesh is stored in the form of voxels, thereby completing the mapping of the asteroid and obtaining the three-dimensional map of the asteroid and the uncertainty of the three-dimensional map.

9. The method according to claim 3, characterized in that The implementation method for constructing a reinforcement learning model with value estimation function in the third step is to construct a reinforcement learning model with an action space dimension of 2, where the dimensions of the action space are set to the change in the orbital altitude of the spacecraft relative to the previous time step and the change in the current orbital inclination of the spacecraft; the orbital altitude h is the distance between the spacecraft and the asteroid, and the orbital inclination θ represents the angle between the current probe position and the line connecting the origin and the xOy plane in the asteroid's coordinate system, and its value range is [-π,π]; the reinforcement learning model includes a decision maker and a critic. The role of the decision maker is to generate the optimal action for the current environmental state, while the job of the critic is to predict the long-term return of the action selected by the decision maker.

10. The method according to claim 9, characterized in that In the third step, the detector state, the asteroid surface image, and the mapping uncertainty estimated by the neural radiation field model are input into the decision maker of the reinforcement learning model, and the decision-making process in the action space is implemented as follows: the mapping uncertainty estimated by the neural radiation field model is encoded into an image with the same perspective and size as the asteroid surface image, and is input into the convolutional neural network of the decision maker as an additional dimension of the asteroid surface image; the convolutional neural network of the decision maker contains 4 convolutional layers, each of which is combined with a Squeeze-Extract layer, a BatchNorm layer, a pooling layer, and a nonlinear layer to obtain a high-dimensional feature vector. The high-dimensional feature vector is concatenated with the vector of the detector state, and the output action is calculated through a multi-layer perceptron. The output action dimension is 2, representing a set of actions in the action space.