Autonomous hidden monitoring robot system based on multi-modal sensing

By fusing multimodal sensors and using an improved CE-GPPO reinforcement learning model, a robust environment model is constructed, which solves the problem of incomplete perception by robots in complex environments, achieves physical rationality and policy stability in path planning, and enhances the autonomy and reliability of covert monitoring tasks.

CN121348747APending Publication Date: 2026-01-16杨兆佳
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511477992.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing mobile monitoring robots suffer from incomplete perception in complex environments, insufficient robustness in localization, lack of covert constraints in path planning, and unstable strategy optimization, making it difficult to perform covert monitoring tasks in high-risk areas.

Method used

By employing multimodal sensor fusion and anti-interference preprocessing, a robust environment model is constructed. Combined with hyperbolic hierarchical dynamics modeling and an improved CE-GPPO reinforcement learning model, a monitoring trajectory that meets the requirements of concealment, security, and robustness is generated.

Benefits of technology

It improves the robot's autonomy and task execution reliability in complex and disturbing environments, solves the problem of incomplete perception, and enhances the physical rationality and strategy stability of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121348747A_ABST
    Figure CN121348747A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an autonomous hidden monitoring robot system based on multi-mode perception. The system comprises a multi-modal sensor group, an anti-interference data preprocessing unit, a positioning and environment modeling unit, a multi-modal feature fusion module, a hierarchical consistency modeling unit, a hidden path planning unit and a motion control unit. A hyperbolic hierarchical dynamic modeling framework is introduced into the system to generate dynamic smooth cost and hierarchical consistency cost, and constraint is provided for a hidden path planning process; an improved CE-GPPO reinforcement learning model is adopted for hidden path planning, a shaping reward function and a gradient preserving regulation and control mechanism are combined, and an optimal track meeting the concealment and safety requirements is generated. According to the method, the autonomy and task execution reliability of the robot in a complex interference environment are effectively improved, and autonomous robust hidden monitoring can be realized in the complex environment and under the interference condition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to an autonomous concealment monitoring robot system based on multi-modal perception. BACKGROUND

[0002] With the rapid development of artificial intelligence, sensor fusion and robot control technology, autonomous mobile robots are increasingly widely used in military reconnaissance, environmental monitoring, disaster search and rescue, industrial inspection and other scenarios. Existing mobile monitoring robots mostly rely on single or limited modal perception methods, such as using only visual cameras or laser radars for environment modeling. Such methods are easily affected by insufficient light, occlusion, noise interference and dangerous factors in complex environments, resulting in incomplete environment perception, which in turn affects the positioning accuracy and effectiveness of path planning.

[0003] On the other hand, existing robot path planning and control methods are usually based on traditional graph search algorithms or shallow reinforcement learning methods, which lack the ability to model complex risk factors and are difficult to balance concealment, safety and passability at the same time. In the concealment monitoring task, the robot needs to perform low-noise, weak-signal exposure movement in high-risk areas while avoiding excessive exposure of electromagnetic radiation, acoustic features and thermal imaging. However, existing methods often only optimize path length or obstacle avoidance performance, and do not consider the dynamic constraints of concealment and hierarchical consistency, which can easily produce trajectory mutations, excessively strong exposed signals or unreasonable path deviations across classes.

[0004] In addition, existing strategy optimization methods rely on a single reward signal during training, making it difficult to converge stably in complex tasks. Especially when there are out-of-boundary gradients or policy entropy anomalies, problems such as entropy collapse, insufficient exploration or slow convergence can easily occur. These deficiencies limit the autonomy and robustness of robots in concealment monitoring tasks. SUMMARY

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies in terms of incomplete perception, insufficient robustness in localization, lack of covert constraints in path planning, and unstable strategy optimization, and to provide an autonomous covert monitoring robot system based on multimodal perception. This system achieves comprehensive acquisition and standardized processing of visual, point cloud, acoustic, thermal field, and hazard factor information through multimodal sensor fusion and anti-interference preprocessing; it constructs an environmental model with uncertainty estimation using robust factor graph optimization and multi-source redundant odometry; based on this, it introduces a hyperbolic hierarchical dynamics modeling framework, combining dynamic smoothing constraints with hierarchical consistency regularization to ensure the temporal continuity and hierarchical logical rationality of potential trajectories; further, it generates monitoring trajectories that meet the requirements of covertness, safety, and robustness through an improved CE-GPPO reinforcement learning model combined with a shaping reward function and gradient preservation control mechanism. This invention effectively improves the robot's autonomy and task execution reliability in complex interference environments, and solves the shortcomings of existing technologies in perception, modeling, and strategy optimization.

[0006] This invention provides an autonomous covert monitoring robot system based on multimodal perception. The system includes: a multimodal sensor group, an anti-interference data preprocessing unit, a localization and environment modeling unit, a multimodal feature fusion module, a hierarchical consistency modeling unit, a covert path planning unit, a motion control unit, and a robot.

[0007] The multimodal sensor array includes a visible light camera and a low-light camera for acquiring environmental images; a 3D lidar and millimeter-wave radar for acquiring spatial point clouds and structural depth information; an array microphone for acquiring environmental noise and sound source signals; an infrared thermal imager and a contact temperature sensor for acquiring temperature field distribution; and gas sensors, particulate matter sensors, electromagnetic radiation intensity detectors, and radiation dose rate sensors for detecting hazardous factors. By acquiring data from the target environment using the above multimodal sensor array, raw sensor data containing visual, geometric, acoustic, thermal field, and hazardous factor information is obtained.

[0008] The anti-interference data preprocessing unit is electrically connected to the multimodal sensor group. In the presence of electromagnetic interference, it performs clock synchronization, time-domain / frequency-domain denoising and adaptive interference suppression on the raw sensor data, and outputs a standardized multimodal observation sequence.

[0009] The localization and environmental modeling unit is connected to the anti-interference data preprocessing unit. Based on the standardized multimodal observation sequence, it uses the motion estimation constraints provided by the multi-source redundant odometry to generate a factor map. It introduces a robust kernel function to optimize the factor map and obtain a robust factor map. SLAM calculation is performed on the robust factor map to generate an environmental model with uncertainty estimation.

[0010] The multimodal feature fusion module is used to perform cross-modal alignment and feature fusion of standardized multimodal observation sequences and environmental models, and outputs a risk cost raster map that includes concealment risk, accessibility and target threat.

[0011] The hierarchical consistency modeling unit embeds and models standardized multimodal observation sequences through a hyperbolic hierarchical dynamics modeling framework to obtain dynamic smoothing cost constraints and hierarchical consistency cost constraints.

[0012] The hidden path planning unit establishes the CE-GPPO model and performs risk-sensitive decision-making and constraint trajectory optimization on the risk cost grid. During the optimization process, the advantage estimation and policy update of the CE-GPPO model are guided by dynamic smoothing cost constraints and hierarchical consistency cost constraints to generate hidden safe trajectories. The construction process of the CE-GPPO model is as follows: establish the PPO model, introduce gradient preservation and two-sided regulation mechanisms to optimize the pruning method of out-of-bounds gradients in the PPO model, and construct the CE-GPPO model.

[0013] The motion control unit generates control commands based on a concealed safety trajectory and sends them to the actuators.

[0014] Furthermore, the process of obtaining dynamic smoothing cost constraints and hierarchical consistency cost constraints through the hyperbolic hierarchical dynamics modeling framework specifically includes the following steps:

[0015] Step B1: Introduce the hyperbolic manifold space and tangent space. Embed the standardized multimodal observation sequence into the hyperbolic manifold space through post-constraint mapping to obtain the latent variable sequence. In this embedding process, logarithmic mapping is used to project the standardized multimodal observation sequence from the hyperbolic manifold space to the tangent space for parametric modeling. Then, exponential mapping is used to map the tangent space result back to the hyperbolic manifold space, thereby ensuring that the latent variable sequence is strictly within the hyperbolic manifold space and has geometric consistency for subsequent statistical modeling. Based on the conditional probability relationship between the latent variable sequence and the standardized multimodal observation sequence, define the observation likelihood distribution.

[0016] Step B2: Establish a first-order Markov dynamics model on the sequence of latent variables to form a hyperbolic dynamics prior for the entire trajectory; this prior characterizes the temporal smoothness and physical consistency of the latent trajectory under the manifold structure and is used as a temporal constraint in joint optimization.

[0017] Step B3: For scene-level label-observation sample pairs in the standardized multimodal observation sequence, calculate the shortest path distance between corresponding nodes of the scene-level label-observation sample pair on the taxonomy graph, and calculate the geodesic distance between latent variables of the latent variable sequence of the scene-level label-observation sample pair in the hyperbolic manifold space; construct a hierarchical consistency regularization based on the difference between the two distances.

[0018] Step B4: Based on the observation likelihood distribution, hyperbolic dynamics prior, and hierarchical consistency regularization, a joint optimization objective is constructed. This objective function balances the contributions of the three factors while ensuring that the latent variable sequence conforms to the observation data distribution and satisfies temporal smoothness and hierarchical structure consistency. During the optimization process, the Riemann optimization method is used to jointly solve the hyperparameters of the latent variable sequence and the observation model, dynamic model, and hierarchical regularization model in the hyperbolic hierarchical dynamics modeling framework, thereby obtaining the latent trajectory solution that satisfies temporal smoothness and hierarchical consistency.

[0019] Step B5: Based on the potential trajectory solution, construct dynamic smoothing cost constraints and hierarchical consistency cost constraints.

[0020] Furthermore, the CE-GPPO model, in its process of generating covert security trajectories, specifically includes the following steps:

[0021] Step S1: Perform gridded modeling of the hidden risk factors, accessibility information and threat source locations contained in the risk cost grid diagram, and initialize the robot's current position, posture and path constraints to form the initial state space for planning;

[0022] Step S2: Construct the state and action space for path planning based on the initial state space. The state includes local environmental risk, occlusion factor, and passability probability. The actions include forward movement, turning, and obstacle avoidance trajectory segments. Generate an initial policy distribution using the CE-GPPO model and obtain a candidate trajectory set through interactive sampling with the environment. Calculate the immediate reward of the candidate trajectories in the candidate trajectory set based on the shaped reward function, providing a basis for policy gradient updates in steps S3 and S4. The shaped reward function is composed of the risk reward reflecting concealment, passability, and threat measurement in the risk cost grid, and the dynamic smoothing cost and hierarchical consistency cost output in step B5.

[0023] Step S3: Preset the clipping interval, calculate the update ratio of candidate trajectories in the candidate trajectory set, compare the update ratio with the clipping interval, determine whether there are out-of-bounds trajectories, and obtain the out-of-bounds trajectories; introduce a gradient stopping mechanism to decouple the update of out-of-bounds trajectories, so that the gradients of out-of-bounds trajectories that were not removed in the clipping process still participate in the update.

[0024] Step S4: Based on the out-of-bounds direction of the out-of-bounds trajectory, introduce two-sided control coefficients to explicitly control the policy entropy of the CE-GPPO model. This achieves dynamic adjustment while preserving the out-of-bounds gradient, further avoiding the entropy collapse or entropy explosion problems caused by traditional pruning, and obtaining the policy distribution of entropy control and gradient correction. The two-sided control coefficients include exploration enhancement coefficients and convergence enhancement coefficients.

[0025] Step S5: Based on the strategy distribution of entropy regulation and gradient correction, perform multiple rounds of iterative sampling and optimization on the candidate trajectory set to screen out hidden and safe trajectories.

[0026] By adopting the above solution, the beneficial effects achieved by the present invention are as follows:

[0027] First, this invention achieves comprehensive acquisition and standardized processing of visual, point cloud, acoustic, thermal field and hazard factor information through multimodal sensor fusion and anti-interference preprocessing. This improves the perception integrity and robustness in complex environments with low light, electromagnetic interference and noise, solves the problem of incomplete information caused by single sensor being easily limited by environmental conditions in the prior art, and enhances the robot's ability to continuously perceive and stably model in concealed environments.

[0028] Secondly, this invention utilizes robust factor graph optimization and multi-source redundant odometer to construct an environmental model with uncertainty estimation. Combined with a hyperbolic hierarchical dynamics modeling framework, it introduces dynamic smoothing constraints and hierarchical consistency regularization into potential trajectory modeling, thereby achieving temporal continuity and hierarchical logical rationality of the trajectory. This enhances the ability to suppress abrupt changes and cross-class shifts during path planning, solves the problems of non-smooth trajectories and logical inconsistencies in existing path planning, and strengthens the physical rationality and semantic consistency of the planning results.

[0029] Finally, this invention introduces an improved CE-GPPO reinforcement learning model in the concealed path planning stage. By combining the shaping reward function and gradient preservation control mechanism, it achieves stable convergence of the policy under risk avoidance and concealment constraints, improves the diversity and stability of generated trajectories, and solves the problems of insufficient exploration, entropy collapse and slow convergence in traditional reinforcement learning in concealed monitoring tasks. It enhances the robot's ability to execute concealed, safe and robust trajectories in complex interference environments, thereby significantly improving the reliability and concealment of task execution. Attached Figure Description

[0030] Figure 1 This is the average shaping reward convergence curve proposed in Example 5;

[0031] Figure 2 This is the entropy curve of the strategy proposed in Example 5;

[0032] Figure 3 This is the δ distribution histogram proposed in Example 5. Detailed Implementation

[0033] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0034] Example 1: This invention provides an autonomous covert monitoring robot system based on multimodal perception. The system includes: a multimodal sensor group, an anti-interference data preprocessing unit, a localization and environment modeling unit, a multimodal feature fusion module, a hierarchical consistency modeling unit, a covert path planning unit, a motion control unit, and a robot.

[0035] The multimodal sensor array includes a visible light camera and a low-light camera for acquiring environmental images; a 3D lidar and millimeter-wave radar for acquiring spatial point clouds and structural depth information; an array microphone for acquiring environmental noise and sound source signals; an infrared thermal imager and a contact temperature sensor for acquiring temperature field distribution; and gas sensors, particulate matter sensors, electromagnetic radiation intensity detectors, and radiation dose rate sensors for detecting hazardous factors. By acquiring data from the target environment using the above multimodal sensor array, raw sensor data containing visual, geometric, acoustic, thermal field, and hazardous factor information is obtained.

[0036] The anti-interference data preprocessing unit is electrically connected to the multimodal sensor group. In the presence of electromagnetic interference, it performs clock synchronization, time-domain / frequency-domain denoising and adaptive interference suppression on the raw sensor data, and outputs a standardized multimodal observation sequence.

[0037] The localization and environmental modeling unit is connected to the anti-interference data preprocessing unit. Based on the standardized multimodal observation sequence, it uses the motion estimation constraints provided by the multi-source redundant odometry to generate a factor map. It introduces a robust kernel function to optimize the factor map and obtain a robust factor map. SLAM calculation is performed on the robust factor map to generate an environmental model with uncertainty estimation.

[0038] Multi-source redundant odometry includes IMU pre-integration, wheel speed odometry, and visual odometry;

[0039] The multimodal feature fusion module is used to perform cross-modal alignment and feature fusion of standardized multimodal observation sequences and environmental models, and outputs a risk cost raster map that includes concealment risk, accessibility and target threat.

[0040] The hierarchical consistency modeling unit embeds and models standardized multimodal observation sequences through a hyperbolic hierarchical dynamics modeling framework to obtain dynamic smoothing cost constraints and hierarchical consistency cost constraints for subsequent path planning and strategy optimization. The hyperbolic hierarchical dynamics modeling framework includes an observation model (step B1), a dynamics model (step B2), and a hierarchical regularization model (step B3). The hyperbolic hierarchical dynamics modeling framework is constructed by introducing observation likelihood modeling, hyperbolic dynamics prior methods, and hierarchical consistency regularization methods, and combining them with post-constraint mapping and Riemann optimization methods.

[0041] The concealed path planning unit establishes a CE-GPPO model, performs risk-sensitive decision-making and constraint trajectory optimization on a risk-cost grid graph, and guides the advantage estimation and strategy update of the CE-GPPO model through dynamic smoothing cost constraints and hierarchical consistency cost constraints during the optimization process, generating a concealed safe trajectory. The concealed safe trajectory is the optimal trajectory that satisfies concealment constraints and safety requirements. The construction process of the CE-GPPO model is as follows: establish a PPO model, introduce gradient preservation and two-sided regulation mechanisms to optimize the pruning method of out-of-bounds gradients in the PPO model, and construct the CE-GPPO model.

[0042] The motion control unit, based on a concealed safety trajectory, generates control commands and sends them to the actuators; the control commands include:

[0043] The robot's speed is dynamically adjusted so that it can reduce its speed when approaching risky areas or threat sources, thereby reducing the exposure of its acoustic and dynamic features.

[0044] Optimize and control movement posture and joint angles to avoid significant thermal signals or visual exposure caused by excessive movements.

[0045] The control drive mechanism and sensor execution unit operate intermittently or at low power to reduce electromagnetic radiation and acoustic signal interference, ensuring the concealment of the monitoring task.

[0046] While ensuring path tracking accuracy, the vehicle's attitude is dynamically adjusted based on the environmental model and risk cost grid map, enabling covert monitoring and real-time data acquisition of the target environment.

[0047] Through the above controls, the robot can achieve a combination of low-noise movement, weak signal exposure, and continuous monitoring while executing a covert and safe trajectory, thereby meeting the dual requirements of stealth and task monitoring.

[0048] Example 2, based on Example 1, describes the process of obtaining dynamic smoothing cost constraints and hierarchical consistency cost constraints using a hyperbolic hierarchical dynamics modeling framework. The specific steps include:

[0049] Step B1: Introduce the hyperbolic manifold space and tangent space. Embed the standardized multimodal observation sequence into the hyperbolic manifold space through post-constraint mapping to obtain the latent variable sequence. In this embedding process, logarithmic mapping is used to project the standardized multimodal observation sequence from the hyperbolic manifold space to the tangent space for parametric modeling. Then, exponential mapping is used to map the tangent space result back to the hyperbolic manifold space, thereby ensuring that the latent variable sequence is strictly within the hyperbolic manifold space and has geometric consistency for subsequent statistical modeling. Based on the conditional probability relationship between the latent variable sequence and the standardized multimodal observation sequence, an observation likelihood distribution is defined to characterize the explanatory power of the latent variables on the observation data.

[0050] Step B2: Establish a first-order Markov dynamics model on the sequence of latent variables to form a hyperbolic dynamics prior for the entire trajectory; this prior characterizes the temporal smoothness and physical consistency of the latent trajectory under the manifold structure and is used as a temporal constraint in joint optimization.

[0051] Step B3: For scene-level label-observation sample pairs in the standardized multimodal observation sequence, calculate the shortest path distance between corresponding nodes of the scene-level label-observation sample pair on the taxonomy graph, and calculate the geodesic distance between latent variables of the latent variable sequence of the scene-level label-observation sample pair in the hyperbolic manifold space; construct a hierarchical consistency regularization based on the difference between the two distances to constrain the latent variable sequence to maintain a geometric structure consistent with the scene-level label in the latent space;

[0052] Step B4: Based on the observation likelihood distribution, hyperbolic dynamics prior, and hierarchical consistency regularization, a joint optimization objective is constructed. This objective function balances the contributions of the three factors while ensuring that the latent variable sequence conforms to the observation data distribution and satisfies temporal smoothness and hierarchical structure consistency. During the optimization process, the Riemann optimization method is used to jointly solve the hyperparameters of the latent variable sequence and the observation model, dynamic model, and hierarchical regularization model in the hyperbolic hierarchical dynamics modeling framework, thereby obtaining the latent trajectory solution that satisfies temporal smoothness and hierarchical consistency.

[0053] Step B5: Based on the latent trajectory solution, construct dynamic smoothing cost constraints and hierarchical consistency cost constraints. The dynamic smoothing cost is used to suppress non-physical abrupt changes in the latent trajectory, ensuring the temporal continuity and physical rationality of the trajectory on the hyperbolic manifold. The hierarchical consistency cost is used to maintain the hierarchical semantic consistency between the latent trajectory and the taxonomy graph structure, avoiding path deviations that do not conform to hierarchical logic during cross-class transitions. The above two types of costs are output as constraint quantities for subsequent path planning and strategy optimization modules to call.

[0054] Dynamic smoothing cost:

[0055] ;

[0056] in, Indicates the time step index. This represents the cost of dynamic smoothing, used to penalize drastic changes in potential trajectories and encourage smoothness and continuity over time. , Indicates the latent variable at time t. and The value of , Indicates hyperbolic geodesic distance;

[0057] Cost of hierarchical consistency:

[0058] ;

[0059] in, This represents the cost of hierarchical consistency. This represents a set of observation samples with scene-level annotations. express The corresponding target-side latent variable point set, Indicates to The difference between the taxonomy map distance and the geodesic distance in the latent space is calculated for the pair of start-end latent variables, and the sum of squares is performed; this ensures that the geometric relationship of the trajectory in the latent space is consistent with the hierarchical semantic relationship of the taxonomy map.

[0060] Example 3, based on Example 2, describes the process of generating a covert security trajectory using the CE-GPPO model, which specifically includes the following steps:

[0061] Step S1: Perform gridded modeling of the hidden risk factors, accessibility information and threat source locations contained in the risk cost grid diagram, and initialize the robot's current position, posture and path constraints to form the initial state space for planning;

[0062] Step S2: Construct the state and action space for path planning based on the initial state space. The state includes local environmental risk, occlusion factor, and passability probability. Actions include forward movement, turning, and obstacle avoidance trajectory segments. Generate an initial policy distribution using the CE-GPPO model and obtain a candidate trajectory set through interactive sampling with the environment. Calculate the immediate reward of the candidate trajectories in the candidate trajectory set based on the shaped reward function, providing a basis for policy gradient updates in steps S3 and S4. The shaped reward function consists of the risk reward reflecting concealment, passability, and threat measurement in the risk cost grid, and the dynamic smoothing cost and hierarchical consistency cost output in step B5. The immediate reward reflects the risk cost performance of the candidate trajectories in the current environment. These data are stored in the experience pool for subsequent iterative updates.

[0063] Reward function after reshaping:

[0064] ;

[0065] in, The reward after the transformation is the original risk reward minus the weighted terms of the two types of constraint costs, so that the learned strategy simultaneously takes into account concealment / safety, time smoothness and hierarchical consistency. This indicates that the risk reward comes from the risk cost grid; , This represents the cost weighting coefficient;

[0066] Step S3: Preset the clipping interval, calculate the update ratio of candidate trajectories in the candidate trajectory set, compare the update ratio with the clipping interval, determine whether there are out-of-bounds trajectories, and obtain the out-of-bounds trajectories; introduce a gradient stopping mechanism to decouple the update of out-of-bounds trajectories, so that the gradients of out-of-bounds trajectories that were not removed in the clipping process still participate in the update.

[0067] Step S4: Based on the out-of-bounds direction of the out-of-bounds trajectory, introduce two-sided adjustment coefficients to explicitly control the policy entropy of the CE-GPPO model. This achieves dynamic adjustment while preserving the out-of-bounds gradient, further avoiding the entropy collapse or entropy explosion problems caused by traditional pruning, and obtaining the policy distribution of entropy regulation and gradient correction. The two-sided adjustment coefficients include exploration enhancement coefficients and convergence enhancement coefficients. In the early stage of training, the exploration enhancement coefficient is increased to amplify the exploration signal of the CE-GPPO model and encourage the generation of diverse trajectories. In the later stage of training, the convergence enhancement coefficient is increased to amplify the convergence signal of the CE-GPPO model, so that the policy can stabilize quickly.

[0068] Gradient update formula:

[0069] ;

[0070] in, Indicates parameters The gradient operator (used to update policy parameters) is the optimization direction that explicitly controls the policy entropy. This represents the optimization objective of the CE-GPPO model. Expressing expectations, Indicates the number of candidate trajectories. Indicates the candidate trajectory index. This indicates the number of decision-making time steps contained in the trajectory. Indicates the weight adjustment term; Indicates parameters The current policy (action distribution) is represented. Indicates the first The trajectory in the first The selected action, Indicates the first The sequence of actions performed before the step, This represents the current perception and environmental representation, which is the observation jointly constituted by the risk cost raster and the ontology state; Indicates the advantage estimate;

[0071] ;

[0072] in, This represents the convergence enhancement coefficient, which applies to out-of-bounds conditions on the left side. This represents the exploration enhancement coefficient, which applies to out-of-bounds conditions on the right side. Indicates the clipping threshold. It represents the importance sampling ratio, which is the ratio of the probability of the current policy to that of the old policy on a certain action;

[0073] Step S5: Based on the strategy distribution of entropy regulation and gradient correction, perform multiple rounds of iterative sampling and optimization on the candidate trajectory set to screen out hidden and safe trajectories.

[0074] Example 4, based on Example 2, describes the process of generating a covert security trajectory using the CE-GPPO model, which includes the following steps:

[0075] Step R1: Perform gridded modeling of the hidden risk factors, accessibility information and threat source locations contained in the risk cost grid diagram, and initialize the robot's current position, posture and path constraints to form the initial state space for planning;

[0076] Step R2: Construct the state and action space for path planning based on the initial state space. The state includes local environmental risk, occlusion factor, and passability probability. Actions include forward movement, turning, and obstacle avoidance trajectory segments. Generate an initial policy distribution using the CE-GPPO model and obtain a candidate trajectory set through interactive sampling with the environment. Calculate the immediate reward of the candidate trajectories in the candidate trajectory set based on the reward function, providing a basis for policy gradient updates in steps S3 and S4. The reward function consists of risk rewards reflecting concealment, passability, and threat measurement in the risk cost grid. The immediate reward reflects the risk cost performance of the candidate trajectories in the current environment. These data are stored in the experience pool for subsequent iterative updates.

[0077] Step R3: Preset the clipping interval, calculate the update ratio of candidate trajectories in the candidate trajectory set, compare the update ratio with the clipping interval, determine whether there are out-of-bounds trajectories, and obtain the out-of-bounds trajectories; introduce a gradient stopping mechanism to decouple the update of out-of-bounds trajectories, so that the gradients of out-of-bounds trajectories that were not removed in the clipping process still participate in the update.

[0078] Step R4: Based on the out-of-bounds direction of the out-of-bounds trajectories, introduce two-sided adjustment coefficients to explicitly control the policy entropy of the CE-GPPO model. This achieves dynamic adjustment while preserving out-of-bounds gradients, further avoiding entropy collapse or entropy explosion problems caused by traditional pruning, resulting in a policy distribution of entropy regulation and gradient correction. The two-sided adjustment coefficients include exploration enhancement coefficients and convergence enhancement coefficients. In the early stages of training, increasing the exploration enhancement coefficient amplifies the exploration signal of the CE-GPPO model, encouraging the generation of diverse trajectories. In the later stages of training, increasing the convergence enhancement coefficient amplifies the convergence signal of the CE-GPPO model, enabling the policy to stabilize quickly.

[0079] Step R5: Based on the strategy distribution of entropy regulation and gradient correction, perform multiple rounds of iterative sampling and optimization on the candidate trajectory set to screen out hidden and safe trajectories.

[0080] Example 5, according to Figure 1 , Figure 2 , Figure 3 This embodiment is based on Embodiment 3. In this embodiment,

[0081] The concealed path planning unit establishes a CE-GPPO model and performs risk-sensitive decision-making and constraint trajectory optimization on the risk cost grid. During the optimization process, the advantage estimation and strategy update of the CE-GPPO model are guided by dynamic smoothing cost constraints and hierarchical consistency cost constraints to generate concealed safe trajectories.

[0082] In this embodiment, during the policy generation stage, the hierarchical consistency modeling unit constrains the potential trajectory using a hyperbolic hierarchical dynamics modeling framework, resulting in a dynamics smoothing cost of 0.18 and a hierarchical consistency cost of 0.12. The CE-GPPO model established by the hidden path planning unit converges after 500 iterations, with a final average value of 0.67 for the shaped reward function, which is about 63% higher than the 0.41 of the traditional PPO method. Under the condition of a pruning threshold of 0.2, the gradient preservation and two-sided regulation mechanism stabilize the policy entropy around 0.89, avoiding the entropy collapse phenomenon.

[0083] Figure 1 The average shaping reward convergence curve for the CE-GPPO model after 500 iterations; Figure 1 In the diagram, the horizontal axis represents the number of training iterations (0-500), and the vertical axis represents the average shaping reward. The solid blue line represents the curve of change with iterations, and the dashed blue line represents the initial reference of 0.41 and the target reference of 0.67.

[0084] Figure 2 The policy entropy curve for the CE-GPPO model after 500 iterations; Figure 2In the diagram, the horizontal axis represents training iterations (number of iterations), ranging from 0 to 500; the vertical axis represents policy entropy; the solid blue line indicates the change in policy entropy with iterations; the dashed blue line represents the steady-state reference of 0.89.

[0085] Figure 3 The δ distribution histograms of the CE-GPPO model at different stages (early, middle, and late) during training. Figure 3 In the middle, the horizontal axis represents the importance sampling ratio δ; the vertical axis represents the probability density; the legend indicates the early, middle, and late stages.

[0086] The motion control unit generates control commands based on a concealed safety trajectory and sends them to the actuators.

[0087] During the control execution phase, the motion control unit issues dynamic control commands based on the concealed safety trajectory:

[0088] When the robot is within 2 meters of the noise source interference device, its speed automatically decreases from 0.8 m / s to 0.35 m / s, and the noise power level decreases by 28%.

[0089] When passing near the heat source simulator, the joint amplitude is reduced by 22%, and the surface heat exposure signal is reduced to the range of background thermal noise.

[0090] In the electromagnetic interference zone, the sensor acquisition frequency automatically switches to an intermittent low-power mode, reducing electromagnetic radiation exposure by 41%.

[0091] Ultimately, the robot completed a full-coverage inspection of the warehouse area within a total task time of 26 minutes;

[0092] The results showed that the trajectory smoothness index (mean square velocity change rate) decreased by 35%, the concealment exposure probability decreased from 0.21 in the comparison system to 0.08, and the mission success rate increased from 82% to 97%.

[0093] The present invention and its embodiments have been described above. This description is not restrictive. The accompanying drawings are only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.

Claims

1. An autonomous covert monitoring robot system based on multi-modal perception, characterized by: The method comprises the following steps: An anti-interference data preprocessing unit acquires a standardized multi-modal observation sequence; A positioning and environment modeling unit generates a factor graph based on the standardized multi-modal observation sequence and motion estimation constraints provided by a multi-source redundant odometry, introduces a robust kernel function to optimize the factor graph, obtains a robust factor graph, and performs SLAM calculation on the robust factor graph to generate an environment model; Multi-modal feature fusion is performed on the standardized multi-modal observation sequence and the environment model to perform cross-modal alignment and feature fusion, and a risk cost grid map is output; A hierarchical consistency modeling unit embeds and models the standardized multi-modal observation sequence through a hyperbolic hierarchical dynamics modeling framework to obtain dynamics smoothness cost constraints and hierarchical consistency cost constraints; A concealed path planning unit establishes a CE-GPPO model, performs risk-sensitive decision-making and constraint trajectory optimization on the risk cost grid map, and guides the advantage estimation and policy update of the CE-GPPO model through the dynamics smoothness cost constraints and the hierarchical consistency cost constraints during the optimization process to generate a concealed safe trajectory; A motion control unit generates control instructions based on the concealed safe trajectory and sends the control instructions to an execution mechanism.

2. The autonomous covert monitoring robot system based on multi-modal perception of claim 1, wherein: The construction process of the CE-GPPO model comprises the following steps:

3. The autonomous covert monitoring robot system based on multi-modal perception of claim 1, wherein: Step B1: A hyperbolic manifold space and a tangent space are introduced, the standardized multi-modal observation sequence is embedded into the hyperbolic manifold space through a post-constraint mapping to obtain a latent variable sequence, and an observation likelihood distribution is defined based on the conditional probability relationship between the latent variable sequence and the standardized multi-modal observation sequence; Step B2: A first-order Markov dynamics model is established on the latent variable sequence to form a hyperbolic dynamics prior for the entire trajectory; Step B3: For scene level label-observation sample pairs in the standardized multi-modal observation sequence, the shortest path distance between the nodes of the scene level label-observation sample pairs is calculated on the taxonomy graph, and the geodesic distance between the latent variable sequences of the scene level label-observation sample pairs is calculated in the hyperbolic manifold space; a hierarchical consistency regularization is constructed based on the difference between the two distances; Step B4: Based on the observation likelihood distribution, the hyperbolic dynamics prior, and the hierarchical consistency regularization, a joint optimization target is constructed to optimize the hyperbolic hierarchical dynamics modeling framework; during the optimization process, the Riemann optimization method is used to jointly solve the latent variable sequence and the hyperparameters of the hyperbolic hierarchical dynamics modeling framework to obtain a latent trajectory solution; Step B5: Based on the latent trajectory solution, the dynamics smoothness cost constraints and the hierarchical consistency cost constraints are constructed. In the embedding process of step B1, the standardized multi-modal observation sequence is projected from the hyperbolic manifold space to the tangent space for parameterized modeling through logarithmic mapping, and the tangent space result is mapped back to the hyperbolic manifold space through exponential mapping to ensure that the latent variable sequence is in the hyperbolic manifold space.

4. The autonomous covert monitoring robot system based on multi-modal perception of claim 3, wherein: ​ 5. The autonomous covert monitoring robot system based on multi-modal perception as claimed in claim 2, wherein: The process of generating a hidden safe trajectory by the CE-GPPO model specifically comprises the following steps: Step S1: performing raster modeling on the risk cost grid map, and initializing the current position, pose and path constraint conditions of the robot to form an initial state space; Step S2: constructing a state and action space for path planning based on the initial state space; generating an initial policy distribution by the CE-GPPO model, and sampling to obtain a candidate trajectory set; introducing a reward function after shaping to calculate the immediate reward of the candidate trajectories in the candidate trajectory set, thereby providing a basis for policy gradient updating in steps S3 and S4; Step S3: presetting a clipping interval, calculating the update ratio of the candidate trajectories, comparing the update ratio with the clipping interval to determine whether there is an out-of-bound trajectory, and obtaining the out-of-bound trajectory; introducing a gradient stopping mechanism to decouple the update of the out-of-bound trajectory; Step S4: according to the out-of-bound direction of the out-of-bound trajectory, introducing a double-sided regulation coefficient to explicitly control the policy entropy of the CE-GPPO model, thereby obtaining a policy distribution with entropy regulation and gradient correction; Step S5: based on the policy distribution with entropy regulation and gradient correction, performing multiple rounds of iterative sampling and optimization on the candidate trajectory set, and screening out a hidden safe trajectory.

6. The autonomous covert monitoring robot system based on multi-modal perception as claimed in claim 5, wherein: The reward function after shaping is composed of a risk reward reflecting the concealment, passability and threat metric in the risk cost grid, and the dynamic smoothness cost constraint and the hierarchical consistency cost constraint output in step B5.

7. The autonomous covert monitoring robot system based on multi-modal perception as claimed in claim 5, wherein: The double-sided regulation coefficient includes an exploration enhancement coefficient and a convergence enhancement coefficient.

Citation Information

Cited By

  • Unmanned vehicle approaching reconnaissance method and system based on dynamic shielding path planning

    CN121977579A

  • A method and system for unmanned vehicle close-range reconnaissance based on dynamic masking path planning

    CN121977579B