Voxel-feedback multi-satellite cooperative asteroid detection path planning method

By combining active disturbance rejection control and voxel mapping with reinforcement learning, the stability and efficiency issues of visual full-coverage detection of asteroid surfaces were solved, achieving efficient and full-coverage detection under complex gravitational fields.

CN121947799APending Publication Date: 2026-05-01SHANGHAI AEROSPACE CONTROL TECH INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI AEROSPACE CONTROL TECH INST
Filing Date
2025-12-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve efficient and comprehensive visual detection on asteroid surfaces, especially in complex gravitational fields where it is difficult to maintain a stable baseline distance and attitude for the probe. Furthermore, traditional methods cannot balance burnup, time, and coverage effectiveness.

Method used

The system employs a dual-detector combination with a binocular stereo vision camera and active disturbance rejection control. By using an extended state observer and nonlinear state error feedback, it maintains a stable baseline and attitude between the detectors and utilizes voxel mapping and reinforcement learning models for path planning, thereby achieving efficient coverage of the asteroid's surface.

Benefits of technology

Maintaining a stable baseline for the detector under complex gravitational fields, accurately characterizing the irregularities on the asteroid surface, and achieving low-time, high-quality, full-coverage visual detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121947799A_ABST
    Figure CN121947799A_ABST
Patent Text Reader

Abstract

The invention discloses a voxel feedback-based multi-satellite cooperative asteroid detection path planning method. The method comprises the following steps of: 1, gathering and combining double detectors near an asteroid based on active-disturbance-rejection control to form a stereoscopic vision module; 2, the stereoscopic vision module carries out maneuvering, carries out detection on the surface of the asteroid, carries out stereoscopic perception on the surface of the asteroid, and continuously constructs and updates a voxel map through the obtained perception information; 3, according to the state of the detector and asteroid surface voxel map feedback information, a reinforcement learning model is used for flight path planning and attitude control adjustment of the detector, and a reinforcement learning decision is achieved; and step 4, evaluating a reinforcement learning decision effect through a reward function, continuously adjusting and optimizing reinforcement learning model parameters, and finally achieving efficient and high-quality coverage of the asteroid surface. According to the invention, surface vision full-coverage path planning of the asteroid is realized by the binocular stereo vision camera combined by the double detectors based on active-disturbance-rejection control.
Need to check novelty before this filing date? Find Prior Art

Description

A Voxel Feedback Multi-Star Cooperative Asteroid Exploration Path Planning Method Technical Field

[0001] This invention relates to a voxel-feedback multi-satellite collaborative asteroid exploration path planning method, belonging to the field of multi-spacecraft collaborative exploration decision planning technology. Background Technology

[0002] The environment surrounding asteroids is characterized by weak gravitational fields, unstructured environments, and complex lighting, placing high demands on the autonomy and intelligence of spacecraft for achieving full visual coverage of their surfaces. Current methods for visually covering asteroids primarily rely on single spacecraft carrying stereo vision cameras. However, these sensors require relatively close proximity to acquire information, and the irregular gravitational fields near asteroids further complicate matters, posing challenges to the safety of the spacecraft.

[0003] Existing methods propose combining two detectors into a stereo vision camera for orbital exploration using nonlinear model predictive control. However, model predictive control requires the construction of a relatively accurate disturbance model. In actual deep space exploration, due to the irregularity of the asteroid's gravitational field, this disturbance model is difficult to construct accurately, which affects the precise control of the baseline distance between the two spacecraft. Furthermore, traditional asteroid visual full-coverage path planning mainly relies on artificially designed flyby orbits for exploration, which cannot effectively balance multiple objectives such as low fuel consumption, short time, and good coverage.

[0004] Therefore, this invention proposes a method for path planning of asteroid surface visual full coverage using a binocular stereo vision camera based on a dual-detector combination with active disturbance rejection control. Summary of the Invention

[0005] The technical problem solved by this invention is to overcome the shortcomings of existing technologies and provide a method for path planning of asteroid surface visual full coverage using a binocular stereo vision camera based on a dual-detector combination with active disturbance rejection control. The technical solution of this invention is: a multi-planet collaborative asteroid exploration path planning method with voxel feedback, comprising: Step 1, dual detectors assembling near the asteroid based on active disturbance rejection control to form a stereo vision module; Step 2, the stereo vision module maneuvering to explore the asteroid surface, performing stereo perception of the asteroid surface, and continuously constructing and updating a voxel map using the acquired perception information; Step 3, based on the detector status and asteroid surface voxel map feedback information, using a reinforcement learning model to plan the detector's flight path and adjust its attitude control, realizing reinforcement learning decision-making; Step 4, evaluating the effect of the reinforcement learning decision-making through a reward function, continuously adjusting and optimizing the reinforcement learning model parameters, ultimately achieving efficient and high-quality coverage of the asteroid surface.

[0006] Furthermore, the active disturbance rejection control method is used to control two detectors equipped with monocular visible light cameras to form a stable binocular stereo vision module, thereby enabling the acquisition of stereo information on the asteroid surface. During the control process, the two detectors need to maintain a stable distance and attitude to form a stereo vision baseline. In this process, one detector acts as a guide and the other as a follower. The guide's line of sight is aligned with the center of mass, and the follower is controlled to maintain a stable baseline distance and relative attitude to ensure overlapping fields of view and stereo imaging.

[0007] Furthermore, the controller of the active disturbance rejection control includes a position loop and an attitude loop. The design of the guide detector position loop specifically includes: (1) establishing a dynamic model of the controlled object. (1) Among them, Indicates gravitational field perturbation, For light pressure, For detector control force, Includes unmodeled perturbations and coupling effects; the subscript L indicates the leader detector. Represents the gravitational constant; This represents the position vector of the Guide probe in the asteroid-fixed coordinate system. This represents the second derivative of the vector at that position. for The magnitude of the vector; (2) Design the extended state observer ESO: (2) Among them, For detector position estimation, For detector velocity estimation, For the total perturbation estimate, the dot above the head represents the first derivative; , , ESO gain coefficient; It is the control gain coefficient, The detector reference position vector in the x-axis component, The x-axis component of the detector control output is the acceleration; the y-axis and z-axis directions are also solved in the same way; (3) Design the nonlinear state error feedback NLSEF: (3) Among them, , , This is an intermediate variable representing the tracking error; This represents the component of the desired reference trajectory in the x-axis direction. The first derivative of the reference trajectory in the x-axis direction; This represents the control quantity calculated by NLSEF; This represents the proportional gain coefficient, which adjusts the system's response to position errors. It is the differential gain coefficient, which adjusts the system's response strength to speed errors; It is a nonlinear function used for error signals Shaping is performed, assigning large gains to large errors and small gains to small errors, in order to improve the dynamic performance and anti-interference capability of the control system. This refers to or , It is a power exponent, controlling the nonlinear shape of the function. The width of the linear interval determines the threshold from the linear segment with small error to the nonlinear segment with large error in the function; the y-axis and z-axis are also solved in the same way; (4) Design the final control quantity. Considering the limited thrust capability of the detector, anti-saturation design is carried out. (4) Among them , This indicates the x-component of the detector's output control quantity after considering control capabilities. Refers to thrust saturation, It is the saturated maximum thrust; the y-axis and z-axis can also be solved in the same way.

[0008] Furthermore, the guide detector attitude loop design specifically includes: (1) establishing the attitude dynamics model of the controlled object: (5) Among them, Here is the detector's rotational inertia matrix; The angular velocity of the guide probe relative to the orbital system; Angular acceleration; To control the torque; Gravity gradient torque; (2) Design an extended state observer for other disturbance moments: (6) Among them, This is the estimated vector of the detector's attitude four elements. This is the first derivative of the estimated vector; To measure angular velocity The estimated vector; For the estimation of total disturbance moment; A four-element vector for measuring or referencing attitude; , , The observer gain coefficient; (2) The angular velocity of the guide probe relative to the orbital system; (3) Design of the nonlinear state error feedback (NLSEF): (7) Among them, This is a suitable portion of the error for the four elements; ; Represents the error rotation amount required for the detector to rotate from the current estimated attitude to the target attitude, a four-element vector expressed in body coordinates; This represents the detector's desired target attitude vector; This represents the proportional gain, corresponding to the feedback coefficient of the attitude error. This represents the differential gain, corresponding to the feedback coefficient of the angular velocity error; This represents the calculated initial control torque. The tracking control torque needs to be subtracted from the estimated total disturbance torque vector. , It is the target angular velocity vector of the detector. For the estimated detector angular velocity vector, This represents the detector's angular velocity vector error.

[0009] Furthermore, the follower probe's position loop design specifically includes: the follower's active disturbance rejection control (ADRC) control structure is the same as the guide's, but the reference input is time-varying. Therefore, to maintain the relative position with the guide, it is necessary to solve for the follower's target position within the asteroid's fixed linkage. (8) Among them, This is the attitude transformation matrix from the probe's body coordinate system to the asteroid's fixed coordinate system; The relative position of the two detectors represents the distance between the follower and the leader. This indicates the target position of the follower probe in the asteroid-fixed coordinate system. The position of the Guide probe in the asteroid fixed coordinate system.

[0010] Furthermore, the follower detector attitude ring design specifically includes:

[0011] The vector.

[0012] Furthermore, the continuous construction and updating of the voxel map using the acquired sensory information specifically involves: (1) constructing a voxel map of the asteroid surface to detect and characterize the irregular three-dimensional surface; M represents the voxel map, and the voxel map uses an octree data structure to store space occupancy information; the voxel map divides the asteroid surface space into cubic grid cells. In the initial state, the three-dimensional environment of the asteroid surface is completely unknown, and all voxels... , It is the collection of all voxels on the surface of an asteroid; when voxels Occupancy probability Reaching the set upper boundary At that time, voxels were considered occupied, i.e. , It is the collection of occupied voxels detected by the detector during the detection process, when Below the lower boundary When the area is idle, it indicates a low-value area that can be accessed. , It is the set of voxels detected by the detector but whose occupancy probability is below the threshold; (2) Perform voxel map update and set the initial voxel map occupancy probability. With a prior value of 0.5, the stereo vision module can only detect a portion of the voxels on the asteroid's surface with each movement. The probability map of voxels on the asteroid's surface is derived by continuously maneuvering the detector and updating it using Bayes' theorem. The update formula for the voxel occupancy probability follows Bayes' theorem: (10) Among them, This is the prior probability, initially set to 0.5, indicating that the environment is completely unknown; The observation likelihood probability represents the probability that a voxel detected by the perceptron is observed. This represents the posterior probability of a voxel being occupied after it has been detected. This represents the false sensing rate of the sensor carried by the detector, and is generally a constant value. The prior probability is the absence of the target; the observation likelihood probability is significantly affected by the distance between the detector and the detection surface region, and has the following relationship: (11) Among them, among them, This refers to the detector's detection capability range, that is, the furthest detection distance of the sensor on the detector. For coefficients, This indicates that the voxel occupancy probability varies with the distance of the probe from the voxels on the asteroid surface. The probability of occupancy increases as the probe gets closer to the asteroid surface, but the range of voxels detected decreases as the field of view is fixed. When the probe moves away from the asteroid surface, the occupancy rate tends to be 0. The stereo perception module continuously refreshes the detected voxel probability map through formulas (10)-(11) during the orbital maneuver detection process, and then judges the target exploration point of the next maneuver cycle through the distribution of voxel maps, and uses it as the navigation target.

[0013] Furthermore, the use of reinforcement learning models for the planning and adjustment of the detector's flight path and attitude control requires the design of the detector's action space, state space, reward function, evaluation network, and action network: 1) Action Space (12) Among them, Three accelerations These are the accelerations of the probe relative to the asteroid's body coordinate system, and these accelerations are subject to a maximum thrust constraint. Three torques These are the detector attitude torques, with the maximum torque being... 2) State space (13) Among them, These represent the probe's position and velocity in the asteroid's fixed coordinate system, respectively. The four elements of posture; 3) Reward function covering rewards: (This is a map of the voxel log-probability of asteroid surfaces.) (14) Energy consumption bonus: (15) Among them, It is the time interval between each decision step; collision penalty: (16) It is a constant value for the collision penalty; mission completion reward: (17) This refers to the reward value given when the asteroid's surface coverage exceeds a certain threshold; the total reward function. : (18) Among them, , , , is a coefficient.

[0014] In a second aspect, the present invention also proposes a non-volatile storage medium comprising: a computer program product, wherein the method is executed when the computer program product is executed.

[0015] Thirdly, the present invention also proposes a computer program product, which includes a computer program that, when executed by a processor, implements the method described.

[0016] The advantages of this invention compared with the prior art are: (1) This invention proposes a dual-detector baseline distance stabilization method based on active disturbance rejection control. Active disturbance rejection control uses an extended state observer (ESO) to estimate and compensate for the total disturbance inside and outside the system (including gravitational uncertainty caused by irregular gravitational fields, model coupling effects, etc.) as an extended state in real time. This makes it naturally strong in suppressing complex disturbances. Compared with model predictive control, it can quickly eliminate disturbances without an accurate model, which is beneficial to maintaining a stable baseline between the two detectors under complex irregular gravitational fields. Moreover, the algorithm structure is relatively simple, the computational complexity is low, and it is conducive to solving optimization problems online.

[0017] (2) This invention proposes to use voxel maps to characterize the spatial uncertainty of asteroid surfaces, which can more accurately characterize the uncertainty of the irregular spatial environment of asteroid surfaces compared with planar grid maps.

[0018] (3) This invention proposes a relationship between the observation likelihood probability of the voxel occupancy of the stereo vision module and the distance to the asteroid surface, as shown in Equation (11), which can characterize the detection capability of the stereo vision module and finely update the voxel map to reasonably guide the probe's detection orbit maneuver. Attached Figure Description

[0019] Figure 1 is a schematic diagram of the octree structure; Figure 2 is a plot of the likelihood occupancy probability curve of voxel observations on the asteroid surface; Figure 3 is a schematic diagram of baseline stabilization control of the dual detectors on the asteroid surface; Figure 4 is a flowchart of the dual detector stereo vision detection process. Detailed Description of Embodiments: The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0020] This invention provides a voxel-feedback multi-star collaborative asteroid exploration path planning method. It utilizes active disturbance rejection control (ADCC) to control two detectors equipped with monocular visible light cameras to form a stable binocular stereo vision camera, enabling the acquisition of stereo information about the asteroid surface. This approach is advantageous for asteroid exploration from a greater distance, ensuring detector safety and reducing the impact of irregular gravitational fields near the asteroid on the quality of information acquisition. Compared to model predictive control, ADCC can quickly eliminate complex disturbances such as unknown interferences in complex gravitational scenarios without requiring an accurate model, and is more conducive to maintaining a stable baseline between the two detectors under complex and irregular gravitational fields.

[0021] Considering that asteroids are usually irregular three-dimensional shapes, the grid maps used in traditional UAV location and environment detection can only represent the detection level of two-dimensional planar environment. To overcome this deficiency, a voxel map-based method is proposed to characterize the detection level of irregular three-dimensional surfaces of asteroids, which can achieve fine characterization of the irregular shape of asteroids.

[0022] Based on this, considering that traditional spacecraft exploration mainly relies on human orbit planning and design, which cannot be precise and cannot achieve optimal fuel consumption, time, and coverage effect, this invention proposes to construct a voxel map of the asteroid surface to overcome this deficiency. During the exploration process, the probability entropy gain of the asteroid surface environment is updated based on the voxel map, and feedback is used to guide the combination of multiple detectors to perform orbital maneuvers, thereby achieving a highly efficient and high-quality asteroid surface visual full coverage mission with low time consumption, low fuel consumption, and low repetition coverage.

[0023] As shown in Figure 4, this invention proposes a voxel-feedback multi-star collaborative asteroid exploration path planning method, including the following steps: Step 1: The two detectors, based on active disturbance rejection control (ADRC), assemble near the asteroid to form a stereo vision module. The two spacecraft use lasers to perform fine spacing measurements and maintain a stable observation baseline spacing based on ADRC to form a binocular stereo vision for sensing the asteroid surface. ADRC uses the Extended State Observer (ESO) to treat the total disturbances inside and outside the system (including gravitational uncertainties caused by irregular gravitational fields, model coupling effects, etc.) as an extended state for real-time estimation and compensation.

[0024] The Auto-Disturbance Rejection Control (ADRC) controller is designed as follows: The core of ADRC includes an Extended State Observer (ESO), a Tracking Differentiator (TD), and a Nonlinear State Error Feedback (NLSEF). During control, the two detectors need to maintain a stable distance and stable attitude to form a stereo vision baseline. In this process, one detector acts as a guide and the other as a follower. The guide's line of sight is aligned with the center of mass, and the follower maintains a stable baseline distance and relative attitude to ensure overlapping fields of view and stereo imaging, as shown in Figure 3.

[0025] Design a dual-loop ADRC controller, consisting of a position loop and an attitude loop.

[0026] I. Establishing the position loop design of the guide detector (1) Establishing the dynamic model of the controlled object (1) Among them, Indicates gravitational field perturbation, For light pressure, For detector control force, It includes unmodeled perturbations and coupling effects. Represents the gravitational constant; This represents the position vector of the guide probe in the Earth's inertial coordinate system. This represents the second derivative of the vector at that position, i.e., the acceleration; for The modulus of a vector.

[0027] L represents the guide probe. The dynamic model of the controlled object is defined in the asteroid fixed coordinate system.

[0028] (2) Design of Extended State Observer (ESO) (2) Among them, For location estimation; For speed estimation; For the total disturbance estimate; , , This is the ESO gain coefficient. It is the control gain coefficient, For the detector reference position vector x-axis component, The detector controls the output, i.e., the acceleration. The y-axis and z-axis directions can also be solved in the same way.

[0029] (3) Nonlinear Error Feedback (NLSEF) Design: (3) Among them, , , This is an intermediate variable representing the tracking error; This represents the component of the desired reference trajectory in the x-axis direction. The first derivative of the reference trajectory in the x-axis direction; This represents the control quantity calculated by NLSEF. This represents the proportional gain coefficient, which adjusts the system's response to position errors. It is the differential gain coefficient, which adjusts the system's response strength to speed errors; It is a nonlinear function used for error signals Shaping is performed, assigning large gains to large errors and small gains to small errors, in order to improve the dynamic performance and anti-interference capability of the control system. This refers to or , It is a power exponent, controlling the nonlinear shape of the function. The width of the linear interval determines the threshold between the linear segment with small errors and the nonlinear segment with large errors in the function. The y-axis and z-axis can be solved in the same way.

[0030] (4) The final control quantity design takes into account the limited thrust capability of the space probe and carries out anti-saturation design. (4) Among them ; This indicates the x-component of the detector's output control quantity after considering control capabilities. Refers to thrust saturation, This is the maximum saturation thrust. The solutions for the y-axis and z-axis can also be obtained in the same way.

[0031] II. Guide Detector Attitude Loop (1) Establishing the Attitude Dynamics Model of the Controlled Object (5) Among them, Here is the detector's rotational inertia matrix; The angular velocity of the guide probe relative to the orbital system; Angular acceleration; To control the torque; Gravity gradient torque; For other perturbation moments. The attitude dynamics model of the controlled object is defined in the L-guide system relative to the asteroid's fixed system.

[0032] (2) Design of extended state observer (6) This is the estimated vector of the detector's attitude four elements. This is the first derivative of the estimated vector; To measure angular velocity The estimated vector; For the estimation of total disturbance moment; A four-element vector for measuring or referencing attitude; , , The observer gain coefficient; The angular velocity of the guide probe relative to the orbital system.

[0033] (3) Nonlinear state error feedback (NLSEF) (7) Among them, This represents the error rotation amount required for the detector to rotate from its current estimated attitude to the target attitude, expressed in body coordinates as a four-element vector component. , It is the target angular velocity vector of the detector. For the estimated detector angular velocity vector, This refers to the detector's angular velocity vector error. This represents the detector's desired target attitude vector; This represents the proportional gain, corresponding to the feedback coefficient of the attitude error. This represents the differential gain, corresponding to the feedback coefficient of the angular velocity error; This represents the calculated initial control torque. The tracking control torque needs to be subtracted from the estimated total disturbance torque vector. .

[0034] III. Follower Probe Position Loop: The control structure of the follower's ADRC is the same as that of the guide, meaning the follower's position loop design is identical to the guide's. However, the reference input is time-varying. To maintain the relative position with the guide, the core task is to solve for the target position of the follower within the asteroid's fixed linkage.

[0035] (8) Among them, This is the attitude transformation matrix from the probe's body coordinate system to the asteroid's fixed coordinate system; The relative position of the two detectors represents the distance between the follower and the leader. To track the target position of the probe in the asteroid fixed coordinate system, The position of the Guide probe in the asteroid fixed coordinate system.

[0036] IV. The follower detector's attitude loop has the same attitude control structure as the facilitator (i.e., the same design as the facilitator's attitude loop), and the follower's reference attitude is...

[0037] (9) Among them, The current pose vector of the guide. This represents the relative line-of-sight vector between the follower and the guide. From a perspective calculate. , Let be the vector from the target point to the leader and the follower.

[0038] Step 2: The stereo vision module maneuvers to explore the asteroid surface, performing stereo perception and continuously constructing and updating a voxel map using the acquired sensory information. 1) Voxel Map Definition and Representation: Constructing a voxel map of the asteroid surface allows for the exploration and representation of its irregular three-dimensional surface. M represents the voxel map, which uses an octree data structure (as shown in Figure 1) to store space occupancy information. The voxel map divides the asteroid surface space into cubic grid cells. Initially, the three-dimensional environment of the asteroid surface is completely unknown, and all voxels... , It is the collection of all voxels on the surface of an asteroid. When voxels Occupancy probability Reaching the set upper boundary At that time, voxels were considered occupied, i.e. , It is the collection of occupied voxels detected by the detector during the detection process, when Below the lower boundary When the area is idle, it indicates a low-value area that can be accessed. , It is the set of voxels detected by the detector but whose occupancy probability is below the threshold.

[0039] 2) Voxel map update: Set the initial voxel map occupancy probability. Assuming a prior value of 0.5, the stereo vision module can only detect a portion of the asteroid's surface voxels with each step of movement. The voxel probability map is derived by continuously maneuvering the detector and updating it using Bayes' theorem.

[0040] The update formula for the voxel occupancy probability follows Bayes' theorem.

[0041] (10) Among them, This is the prior probability, initially set to 0.5, indicating that the environment is completely unknown; The observation likelihood represents the probability that a voxel detected by the perceptron will be observed. This represents the posterior probability of a voxel being occupied after it has been detected. This represents the false sensing rate of the sensor carried by the detector, and is generally a constant value. The prior probability is that the target does not exist.

[0042] The observation likelihood probability is significantly affected by the distance between the detector and the detection surface area.

[0043] (11) Among them, This refers to the detector's detection capability range, that is, the furthest detection distance of the sensor on the detector. For coefficients; This indicates that the voxel occupancy probability varies with the distance of the probe from the voxels on the asteroid surface. The probability of detection increases as the probe gets closer to the asteroid's surface, but the range of voxels detected decreases because the field of view is fixed. When the probe moves away from the asteroid's surface, the probability of detection approaches zero, as shown in Figure 2.

[0044] During the orbital maneuvering detection process, the dual-detector stereo perception module continuously refreshes the detected voxel probability map through formulas (10)-(11), and then uses the voxel map distribution to determine the target exploration point for the next maneuvering cycle and uses it as the navigation target.

[0045] Step 3: Multi-Star Collaborative Asteroid Exploration Path Planning Based on Voxel Map Feedback. Since the two-star collaborative system is a single module, a single-agent reinforcement learning method is used for trajectory planning. This requires designing the probe's action space, state space, reward function, evaluation network, and action network.

[0046] 1) Action Space (12) Among them, Three accelerations These are the accelerations of the probe relative to the asteroid's body coordinate system, and these accelerations are subject to a maximum thrust constraint. Three torques These are the detector attitude torques, with the maximum torque being... .

[0047] 2) State Space (13) Among them, These represent the velocity and position of the probe in the asteroid's fixed coordinate system; The four elements of posture; This is a voxel log probability map of the asteroid surface.

[0048] 3) Reward function covers rewards: (14) Energy consumption bonus: (15) It is the time interval for each decision step.

[0049] Collision penalty: (16) It is a constant value for collision penalty.

[0050] Task completion reward: (17) It is the reward value given when the surface coverage of an asteroid exceeds a certain threshold.

[0051] Overall reward function: (18) Among them, , , , is a coefficient.

[0052] Step 4: Evaluate the effectiveness of reinforcement learning decisions through a reward function, continuously adjust and optimize the parameters of the reinforcement learning model, and ultimately achieve efficient and high-quality coverage of the asteroid surface.

[0053] For the action network and evaluation network, a typical multilayer fully connected neural network can be selected.

[0054] Training phase: Initialize the state of multiple detectors, network, and environmental state parameters. The detectors update their state by interacting with the asteroid exploration environment. The network parameters of the multi-agent reinforcement learning algorithm model are updated according to the reward function. When the average reward of the detector exploration model stabilizes within a certain range and no longer increases, training is stopped and the model is saved.

[0055] Testing phase: Set the initial position and velocity of multiple detectors and the initial conditions of the voxel map. Using the trained multi-spacecraft full-coverage decision planning model, each detector obtains observation information of the environment through its local perception as the input of the decision network, and the output is the control action taken by the spacecraft. When the maximum detection time threshold is reached or the environmental uncertainty represented by the voxel map decreases to the set threshold, the mission is determined to end.

[0056] The parts of this invention not described in detail are common knowledge to those skilled in the art.

Claims

1. A multi-star collaborative asteroid exploration path planning method with voxel feedback, characterized in that... include: Step 1: The two detectors are assembled near the asteroid based on self-disturbance rejection control to form a stereo vision module; Step 2: The stereo vision module maneuvers to explore the asteroid surface, performing stereo perception and continuously building and updating a voxel map using the acquired sensory information. Step 3: Based on the probe's status and feedback information from the asteroid surface voxel map, a reinforcement learning model is used to plan the probe's flight path and adjust its attitude control, achieving reinforcement learning decision-making. Step 4: The effect of the reinforcement learning decision-making is evaluated through a reward function, and the parameters of the reinforcement learning model are continuously adjusted and optimized to ultimately achieve efficient and high-quality coverage of the asteroid surface.

2. The method for multi-star collaborative asteroid exploration path planning with voxel feedback according to claim 1, characterized in that: The active disturbance rejection control method is used to control two detectors equipped with monocular visible light cameras to form a stable binocular stereo vision module, enabling the acquisition of stereo information on the surface of the asteroid. During the control process, the two detectors need to maintain a stable distance and attitude to form a stereo vision baseline. In this process, one detector acts as a guide and the other as a follower. The guide's line of sight is aligned with the center of mass, and the follower is controlled to maintain a stable baseline distance and relative attitude to ensure overlapping fields of view and stereo imaging.

3. The method for multi-star collaborative asteroid exploration path planning with voxel feedback according to claim 2, characterized in that: The controller of active disturbance rejection control includes a position loop and an attitude loop. The design of the position loop of the guide detector specifically includes: (1) establishing a dynamic model of the controlled object. (1) Among them, Indicates gravitational field perturbation, For light pressure, For detector control force, Includes unmodeled perturbations and coupling effects; the subscript L indicates the leader detector. Represents the gravitational constant; This represents the position vector of the Guide probe in the asteroid-fixed coordinate system. This represents the second derivative of the vector at that position. for The magnitude of the vector; (2) Design the extended state observer ESO: (2) Among them, For detector position estimation, For detector velocity estimation, For the total perturbation estimate, the dot above the head represents the first derivative; 、 、 ESO gain coefficient; It is the control gain coefficient, The detector reference position vector in the x-axis component, The x-axis component of the detector control output is the acceleration; the y-axis and z-axis directions are also solved in the same way; (3) Design the nonlinear state error feedback NLSEF: (3) Among them, , 、 This is an intermediate variable representing the tracking error; This represents the component of the desired reference trajectory in the x-axis direction. The first derivative of the reference trajectory in the x-axis direction; This represents the control quantity calculated by NLSEF; This represents the proportional gain coefficient, which adjusts the system's response to position errors. It is the differential gain coefficient, which adjusts the system's response strength to speed errors; It is a nonlinear function used for error signals Shaping is performed, assigning large gains to large errors and small gains to small errors, in order to improve the dynamic performance and anti-interference capability of the control system. This refers to or 、 It is a power exponent, controlling the nonlinear shape of the function. The width of the linear interval determines the threshold from the linear segment with small error to the nonlinear segment with large error in the function; the y-axis and z-axis are also solved in the same way; (4) Design the final control quantity. Considering the limited thrust capability of the detector, anti-saturation design is carried out. (4) Among them , This indicates the x-component of the detector's output control quantity after considering control capabilities. Refers to thrust saturation, It is the saturated maximum thrust; the y-axis and z-axis can also be solved in the same way.

4. The multi-star cooperative asteroid exploration path planning method with voxel feedback according to claim 3, characterized in that: The guide detector attitude loop design specifically includes: (1) establishing the attitude dynamics model of the controlled object: (5) Among them, Here is the detector's rotational inertia matrix; The angular velocity of the guide probe relative to the orbital system; Angular acceleration; To control the torque; Gravity gradient torque; (2) Design an extended state observer for other disturbance moments: (6) Among them, This is the estimated vector of the detector's attitude four elements. This is the first derivative of the estimated vector; To measure angular velocity The estimated vector; For the estimation of total disturbance moment; A four-element vector for measuring or referencing attitude; 、 、 The observer gain coefficient; (2) The angular velocity of the guide probe relative to the orbital system; (3) Design of the nonlinear state error feedback (NLSEF): (7) Among them, This is a suitable portion of the error for the four elements; ; Represents the error rotation amount required for the detector to rotate from the current estimated attitude to the target attitude, a four-element vector expressed in body coordinates; This represents the detector's desired target attitude vector; This represents the proportional gain, corresponding to the feedback coefficient of the attitude error. This represents the differential gain, corresponding to the feedback coefficient of the angular velocity error; This represents the calculated initial control torque. The tracking control torque needs to be subtracted from the estimated total disturbance torque vector. 、 It is the target angular velocity vector of the detector. For the estimated detector angular velocity vector, This represents the detector's angular velocity vector error.

5. The multi-star cooperative asteroid exploration path planning method with voxel feedback according to claim 4, characterized in that: The follower probe's position loop design specifically includes: the follower's active disturbance rejection control (ADRC) control structure is the same as the guide's, but the reference input is time-varying. Therefore, to maintain the relative position with the guide, it is necessary to solve for the follower's target position within the asteroid's fixed linkage. (8) Among them, This is the attitude transformation matrix from the probe's body coordinate system to the asteroid's fixed coordinate system; The relative position of the two detectors represents the distance between the follower and the leader. The target position of the follower probe in the asteroid fixed coordinate system. The position of the Guide probe in the asteroid fixed coordinate system.

6. The multi-star cooperative asteroid exploration path planning method with voxel feedback according to claim 5, characterized in that: The follower detector attitude ring design specifically includes: The vector.

7. The method for multi-star collaborative asteroid exploration path planning with voxel feedback according to claim 1, characterized in that: The method of continuously constructing and updating voxel maps using the acquired sensory information specifically includes: (1) constructing a voxel map of the asteroid surface to detect and characterize the irregular three-dimensional surface; M represents the voxel map, and the voxel map uses an octree data structure to store the space occupancy information; the voxel map divides the asteroid surface space into cubic grid units. In the initial state, the three-dimensional environment of the asteroid surface is completely unknown, and all voxels , It is the collection of all voxels on the surface of an asteroid; when voxels Occupancy probability Reaching the set upper boundary At that time, voxels are considered occupied, i.e. , It is the collection of occupied voxels detected by the detector during the detection process, when Below the lower boundary When the area is idle, it indicates a low-value area that can be accessed. , It is the set of voxels detected by the detector but whose occupancy probability is below the threshold; (2) Perform voxel map update and set the initial voxel map occupancy probability. With a prior value of 0.5, the stereo vision module can only detect a portion of the voxels on the asteroid's surface with each movement. The probability map of voxels on the asteroid's surface is derived by continuously maneuvering the detector and updating it using Bayes' theorem. The update formula for the voxel occupancy probability follows Bayes' theorem: (10) Among them, This is the prior probability, initially set to 0.5, indicating that the environment is completely unknown; The observation likelihood probability represents the probability that a voxel detected by the perceptron is observed. This represents the posterior probability of a voxel being occupied after it has been detected. This represents the false sensing rate of the sensor carried by the detector, and is generally a constant value. The prior probability is the absence of the target; the observation likelihood probability is significantly affected by the distance between the detector and the detection surface region, and has the following relationship: (11) Among them, among them, This refers to the detector's detection capability range, that is, the furthest detection distance of the sensor on the detector. For coefficients, This indicates that the voxel occupancy probability varies with the distance of the probe from the voxels on the asteroid surface. The probability of occupancy increases as the probe gets closer to the asteroid surface, but the range of voxels detected decreases as the field of view is fixed. When the probe moves away from the asteroid surface, the occupancy rate tends to be 0. The stereo perception module continuously refreshes the detected voxel probability map through formulas (10)-(11) during the orbital maneuver detection process, and then judges the target exploration point of the next maneuver cycle through the distribution of voxel maps, and uses it as the navigation target.

8. The method for multi-star collaborative asteroid exploration path planning with voxel feedback according to claim 1, characterized in that: The use of reinforcement learning models for the flight path and attitude control planning and adjustment of the probe requires the design of the probe's action space, state space, reward function, evaluation network, and action network: 1) Action Space (12) Among them, Three accelerations These are the accelerations of the probe relative to the asteroid's body coordinate system, and these accelerations are subject to a maximum thrust constraint. Three torques These are the detector attitude torques, with the maximum torque being... 2) State space (13) Among them, These represent the probe's position and velocity in the asteroid's fixed coordinate system, respectively. The four elements of posture; 3) Reward function covering rewards: (This is a map of the voxel log probability of asteroid surfaces.) (14) Energy consumption bonus: (15) Among them, It is the time interval between each decision step; collision penalty: (16) It is a constant value for the collision penalty; mission completion reward: (17) This refers to the reward value given when the asteroid's surface coverage exceeds a certain threshold; the total reward function. : (18) Among them, 、 、 、 is a coefficient.

9. A non-volatile storage medium, characterized in that, include: A computer program product that, when executed, performs the method described in any one of claims 1 to 8.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.