Self-adaptive visual navigation method, system and equipment for dynamic pedestrian environment
Through the adaptive visual navigation method, the neural network and dynamic map update module are used to solve the problem of pedestrian behavior simulation in the dynamic environment of the navigation system, and efficient and safe navigation is achieved.
Patent Information
- Application Number
- CN202510415069.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
AI Technical Summary
Existing navigation technologies are difficult to effectively cope with the complexity in dynamic environments, especially in densely populated public places, which cannot truly simulate pedestrian behavior and environmental interaction, resulting in unsatisfactory results of the navigation system in real scenarios.
The environment scene images are obtained through the robot, the pre-trained neural network model is used to detect and predict pedestrians, and the three-dimensional coordinates of feature points are restored in combination with point cloud data, the initial map is established, and the dynamic map update module and the autonomous hybrid path planning module are realized to realize adaptive visual navigation.
It improves the navigation efficiency and safety of the robot in complex dynamic environments, can truly simulate pedestrian behavior, avoid collisions with pedestrians, and improves the performance of the navigation system in dynamic environments.
Smart Images

Figure CN120351933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an adaptive visual navigation method, system, and device for a dynamic pedestrian environment. Background Art
[0002] In recent years, with the rapid development of robot technology, autonomous navigation has become one of the core issues in intelligent robot applications. Especially in the fields of smart home and service robots, robots that can autonomously navigate in dynamic environments have broad application prospects, such as providing auxiliary services, item transfer, obstacle avoidance, etc. for household users in daily life.
[0003] Currently, most of the existing navigation technologies on the market focus on path planning and target recognition in static environments. Traditional visual navigation systems mostly ignore the complexity in dynamic environments and cannot effectively cope with the influence of dynamic factors in the environment. Although multi-sensor fusion solutions can alleviate this influence to a certain extent, the additional cost increase hinders the popularization of related navigation systems. In the real world, especially in densely populated public places such as shopping malls, train stations, etc., navigation robots need to handle the complex dynamic changes brought by pedestrian movement. This makes the existing static navigation technologies difficult to meet the needs in real life and unable to efficiently complete tasks.
[0004] In addition, existing pedestrian dynamic simulators, such as iGibson-Social Navigation (iGibson-SN), Isaac Sim, and HabiCrowd, although attempt to simulate pedestrian behavior, their models are usually too simplified to truly reproduce the diverse behaviors of pedestrians and their interactions with the environment. Moreover, the scenarios of these simulators are limited, and pedestrian behaviors are mostly fixed trajectories, making it difficult to truly reproduce complex dynamic environments, thus resulting in unsatisfactory effects of the navigation system in real scenarios.
[0005] Therefore, developing a system that can truly simulate pedestrian behavior and perform adaptive visual navigation in dynamic and complex scenarios has become an urgent problem to be solved in the field of autonomous navigation. Summary of the Invention
[0006] Based on the technical problems existing in the background art, the present invention proposes an adaptive visual navigation method, system, and device for a dynamic pedestrian environment, significantly improving the navigation efficiency and safety of robots in complex dynamic environments.
[0007] The adaptive visual navigation method for a dynamic pedestrian environment proposed by the present invention includes: The robot acquires the environmental scene image to obtain the point cloud data of the environment, and inputs the environmental scene image into the visual positioning and navigation framework, which includes a pedestrian detection and prediction module, a visual positioning module for dynamic scenes, a dynamic map update module, and an autonomous hybrid path planning module; The pedestrian detection and prediction module uses a pre-trained neural network model to perform real-time detection on pedestrians and associated items in the environmental scene image to obtain a mask, and predicts the future positions of pedestrians based on this; The visual positioning module for dynamic scenes extracts features based on the input environmental scene image, combines the point cloud data to restore the three-dimensional coordinates of the feature points, establishes an initial map, and in each subsequent frame, filters the extracted feature points according to the mask result and matches them with the feature points in the initial map to estimate the pose information of each frame; The dynamic map update module generates a semantic map using the received pose information, local map point cloud information, semantic information, and pedestrian position information, dynamically integrates the future positions of pedestrians into the semantic map, and updates it using the maximum pooling method of temporal variation; The autonomous hybrid path planning module ensures that the robot reaches the target position based on the updated semantic map and the target position, based on adaptive path planning.
[0008] Furthermore, the pedestrian detection and prediction module realizes the detection of dynamic pedestrians and their associated objects through a two-stream branch network. The two branch networks simultaneously use the features of the backbone network to calculate semantic information and optical flow information, and finally the prediction head combines these two types of information to output the masks of dynamic pedestrians and associated objects 。
[0009] Furthermore, the pose information of each frame of the robot is solved by means of nonlinear optimization, and the solution formula is as follows: ; where, represents the homogeneous coordinates of the 2D image after filtering through the mask is a function that projects 3D points onto a 2D image, represents the camera internal parameter matrix, represents the external parameter matrix of the camera, that is, the pose information of the robot for each frame to be solved, represents the homogeneous coordinates of the 3D feature points in the initial map.
[0010]
[0010] Furthermore, in the pedestrian detection and prediction module, the tracking of pedestrians and the prediction of their motion trajectories and the point cloud of dynamic obstacles are carried out through Kalman filtering. The prediction result formula of the pedestrian position at the next moment is as follows: ; where, They are respectively the pedestrian prediction results at a moment, is the pedestrian detection results at a moment, the state transition matrix, and is the process noise.
[0011] The prediction result formula of the point cloud of pedestrians and their associated objects is as follows: ; wherein, and are respectively and the point cloud positions at moments, represents and the position transformation matrix of the point cloud under the change of the point cloud positions at moments.
[0012] Furthermore, the generation process of the semantic map is as follows: Label the objects and pedestrians obtained by the robot according to their categories in the global map, and at the same time project the obstacle point cloud onto a 2D plane to obtain an obstacle map, and then generate a semantic map.
[0013] Furthermore, when updating the semantic map, when a pedestrian moves out of the field of view or cannot be detected, the corresponding obstacle will still be retained in the semantic map and will only be updated when this part is observed again, or until it is removed after exceeding the time threshold ; when the robot encounters an impassable area, a new obstacle area will be generated in this area and gradually expand until the robot changes the planned path.
[0014] Furthermore, in the autonomous hybrid path planning module, the path planning stage is specifically as follows: The robot first calculates the global optimal path through a heuristic search algorithm; When the static obstacle does not block the global path, use a faster time elastic band algorithm to plan the local path in order to avoid collisions with dynamic obstacles in a timely manner; When a new change occurs in the scene, such as the original global path being blocked by a static obstacle, use the heuristic search algorithm again to update the global optimal path; When the robot is within the range of the target position, in order to stop more precisely at a suitable peripheral position of the target object, the robot switches to a slightly more time-consuming fast marching algorithm, and ensures that the robot accurately reaches the area near the target object by calculating the optimal action on the gradient field.
[0015] Furthermore, in order to complete the training of the network module of the adaptive visual navigation method, as well as the testing and parameter adjustment of the overall method framework, based on the existing static home simulator, a fully automatic dynamic pedestrian simulator that can be quickly deployed was constructed by introducing a dynamic digital human and designing a dynamic digital human action algorithm. The specific usage process is as follows: Set the scene configuration, robot position, target position, and pedestrian pose parameters according to requirements. This dynamic pedestrian simulator can quickly generate diverse pedestrian postures according to requirements, and model the pedestrian postures through a parameterized 3D human body model, dynamically adjusting the pedestrian postures in the global coordinate system and the local coordinate system.
[0016] When collecting training data, all configurations of the dynamic pedestrian simulator can be randomly set within a reasonable range, including the robot's initial position, map information, target information, dynamic pedestrian behavior, etc., and it can run automatically. The collected data such as RGBD observations are used to complete the training of the network module of the adaptive visual navigation method. The overall adaptive visual navigation method is also first tested on the dynamic pedestrian simulator and relevant parameters are adjusted to avoid the complexity of parameter adjustment in the actual scenario.
[0017] An adaptive visual navigation system for a dynamic pedestrian environment, characterized by comprising: It includes an image acquisition module and a visual positioning and navigation framework. The visual positioning and navigation framework includes a pedestrian detection and prediction module, a visual positioning module for dynamic scenarios, a dynamic map update module, and an autonomous hybrid path planning module; The image acquisition module is used to obtain the environmental scene image through the robot, obtain the point cloud data of the environment, and input the environmental scene image into the visual positioning and navigation framework; The pedestrian detection and prediction module uses a pre-trained neural network model to perform real-time detection on pedestrians and associated items in the environmental scene image to obtain a mask, and accordingly predicts the future positions of pedestrians; The visual positioning module for dynamic scenarios extracts features based on the input environmental scene image, combines with the point cloud data to restore the three-dimensional coordinates of the feature points, establishes an initial map, and in subsequent frames, filters the restored feature points according to the mask and matches them with the feature points in the initial map to estimate the pose of each frame; The dynamic map update module uses the received pose information, local map point cloud information, semantic information (mask), and pedestrian position information to generate a semantic map, dynamically integrates the future positions of pedestrians into the semantic map, and updates it using the maximum pooling method with temporal changes; The autonomous hybrid path planning module ensures that the robot reaches the target position based on the updated semantic map and the target position, based on adaptive path planning.
[0018] Further, the pedestrian detection and prediction module realizes the detection of dynamic pedestrians and their associated objects through a two-stream branch network. The two branch networks simultaneously utilize the features of the backbone network to calculate semantic information and optical flow information. Finally, the prediction head combines these two types of information to output the masks of dynamic pedestrians and associated objects. 。
[0019] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the adaptive visual navigation method described above is implemented.
[0020] The advantages of the adaptive visual navigation method, system, and device for a dynamic pedestrian environment provided by the present invention are as follows: A complete adaptive visual navigation framework is constructed, and a low-cost and efficient navigation method for a dynamic pedestrian environment that only relies on visual sensors is proposed. The detection of dynamic pedestrians and associated objects is achieved through a two-stream network, and then a visual positioning module that can operate robustly in a dynamic scene is constructed. After introducing the pedestrian behavior prediction and dynamic map update strategy, the navigation system can maintain a high success rate and safety in a changing scene and avoid collisions with pedestrians. The adaptive hybrid path planning strategy improves the efficiency and accuracy of navigation. At the same time, a dynamic pedestrian simulator with a diverse pedestrian behavior model is constructed, which can simulate more realistic pedestrian behaviors and can be used to assist in improving the performance of various navigation systems in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic structural diagram of the present invention; Figure 2 is a schematic diagram of the dynamic pedestrian simulator. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] Hereinafter, the technical solutions of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0023] As Figure 1 and 2 shown, the adaptive visual navigation method for a dynamic pedestrian environment proposed by the present invention includes: The robot acquires an environmental scene image to obtain the point cloud data of the environment: The robot uses an RGBD visual sensor to acquire the RGB image and depth image in the scene, aligns them, and then obtains the point cloud data of the environment.
[0024] Input the environmental data obtained by the robot into the visual positioning and navigation framework: The visual positioning and navigation framework includes a pedestrian detection and prediction module, a visual positioning module for dynamic scenes, a dynamic map update module, and an autonomous hybrid path planning module: Pedestrian detection and prediction module: Used to detect pedestrians and related items in the scene in real time through a pre-trained neural network model in a dynamic pedestrian environment, obtain the corresponding mask results, and transmit the mask results to the visual positioning module. Then, obtain the pose information of each frame from the visual positioning module and perform trajectory prediction on the pedestrians to obtain the future positions of the pedestrians; Visual positioning module for dynamic scenes: In the initial state, extract features (ORB feature descriptors and descriptors) from the input environmental scene image, restore the three-dimensional coordinates of the feature points using the point cloud data, and establish an initial map. In each subsequent frame, screen the extracted feature points according to the mask results output by the pedestrian detection network, and match the screened feature points with the feature points in the initial map to estimate the pose of each frame; Dynamic map update module: Receive the pose information and local map point cloud information from the visual positioning module, as well as the semantic information (i.e., mask) and pedestrian position information obtained by the pedestrian detection and prediction module. Generate a semantic map using these data and dynamically fuse the future positions of the pedestrians into the semantic map. Adopt a maximum pooling method with temporal variation to fuse the semantic map with the semantic information and obstacle information detected in each frame to obtain an updated semantic map; Autonomous hybrid path planning module: Based on the updated semantic map and the target position, ensure that the robot reaches the target position based on adaptive path planning: When the target point is far from the robot, use a heuristic search algorithm to calculate the global path; When dynamic obstacles appear near the robot or the scene changes significantly, adjust the local path planning in a timely manner through the time elastic band algorithm to prevent collisions; When the global path is blocked by new obstacles, use the heuristic search algorithm again to calculate a new global path; When the robot approaches the target object, automatically switch to the fast travel algorithm, and ensure that the robot accurately finds the target object by calculating the optimal action on the gradient field, thereby improving the overall efficiency and success rate.
[0025] In this embodiment, by constructing a diverse dynamic pedestrian simulator and designing a navigation method with self-adaptive ability to dynamic environments, the performance of the navigation system in dynamic environments is significantly improved. It mainly includes the following five parts: 1) Dynamic pedestrian simulator: By introducing pedestrian behavior data in the real world, a dynamic pedestrian model is constructed to support complex pedestrian behavior simulations such as walking, running, dancing, etc., providing a highly simulated environment for the training and testing of the navigation system. 2) Pedestrian detection and prediction module: The dual-stream network model is used to detect dynamic pedestrians and associated objects, and the Kalman filter is used to track pedestrians in real time and predict the future movement trajectories of pedestrians to avoid collisions. 3) Visual localization module for dynamic scenarios: By extracting ORB feature sub and descriptors of visual images, and combining point cloud information to restore the three-dimensional coordinates of feature points, initial map points are established, and feature points are dynamically filtered according to pedestrian detection results in subsequent frames to achieve accurate estimation of the robot pose. 4) Dynamic map update module: Receiving the pose information and local map point cloud information from the visual localization module, as well as the semantic information (i.e., mask) and pedestrian position information obtained from the pedestrian detection and prediction module, a semantic map is generated, and the future positions of pedestrians are dynamically fused into the semantic map. The semantic map is updated using the maximum pooling method with temporal changes to ensure the real-time and accuracy of the map. 5) Autonomous hybrid path planning: Based on the updated semantic map and the target position, global path planning algorithms and local path planning algorithms are adaptively selected, combined with heuristic search algorithms, time elastic band algorithms, and fast marching algorithms, to provide a hybrid strategy for global and local path planning to adapt to complex changes in dynamic environments and improve navigation efficiency and success rate.
[0026] Among them, for 1) the dynamic pedestrian simulator, in order to complete the training of the network module of the self-adaptive visual navigation method, and the testing and parameter adjustment of the overall framework of the method. Based on the existing static home simulator, by introducing dynamic digital humans and designing dynamic digital human action algorithms, this embodiment constructs a fully automatic dynamic pedestrian simulator that can be quickly deployed, achieving a simulated scene that meets the actual application requirements for training. The specific usage process is as Figure 2 shown. Set scene configuration, robot position, target position, and pedestrian pose parameters according to requirements. This dynamic pedestrian simulator can quickly generate diverse pedestrian postures according to requirements, and model pedestrian postures using a parameterized 3D human body model, dynamically adjusting pedestrian postures in the global coordinate system and local coordinate system. When collecting training data, all configurations of the simulator can be randomly set within a reasonable range, including the initial position of the robot, map information, target information, dynamic pedestrian behavior, etc., and it runs automatically. The collected data such as RGBD observations are used to complete the training of the network module of the self-adaptive visual navigation method. The overall self-adaptive visual navigation method is also first tested on the dynamic pedestrian simulator and relevant parameters are adjusted to avoid the complexity of parameter tuning in the actual scenario.
[0027] It can be seen from Figure 2 that the initial positions and behaviors of pedestrians are jointly determined by multiple factors. Specifically, the initial positions of the dynamic pedestrian models are dynamically determined based on the current position of the robot, the obstacles in the environment, and the relative positions of other pedestrians. This means that the behaviors of pedestrians in the scene are not only static but also affected in real time by the surrounding environment and other pedestrians. For example, the setting of the current position of the robot will affect the behavior patterns of nearby pedestrians, and the distance between the robot and the pedestrians will determine the avoidance behaviors of the pedestrians. At the same time, the "Real-time 2D Map" and "Behavior Update" panels in the figure show how to adjust the positions and behaviors of pedestrians in the scene in real time through these dynamic parameters. The real-time 2D map is directly obtained through the 3D projection of the virtual environment.
[0028] The long-distance movement behaviors of the dynamic pedestrian models are further controlled by the collision avoidance algorithm (ORCA, Optimal Reciprocal Collision Avoidance), which can adjust the movement trajectories of pedestrians according to real-time environmental data (such as the distances and speeds between pedestrians) to ensure that pedestrians can reasonably avoid obstacles and other pedestrians in different environments. Especially in the simulation, when multiple pedestrians exist simultaneously, the collision avoidance algorithm can calculate the optimal avoidance paths, making the movements between pedestrians more natural and avoiding unnecessary collisions. The "Pedestrian Initialization" part in the figure shows the diversity of pedestrian behaviors. Different pedestrians can choose different movement modes, such as walking, running, or waiting, and the system adaptively adjusts according to environmental conditions and task requirements.
[0029] The main advantage of this step is that the introduction of the collision avoidance algorithm enables pedestrians to adjust their paths more flexibly when facing a dynamic environment, enhancing the realism of pedestrian simulation. When the dynamic pedestrian planning model is applied in an actual scenario, it may adjust the behaviors of pedestrians according to real-time sensor data, and some simulation steps can be omitted at this time. The "Update Module" in the figure shows how to synchronize these behavior updates with the environmental changes in the scene, making the model training and testing more efficient.
[0030] In summary, combining Figure 2 the content shown above, the dynamic pedestrian simulator provides a highly flexible and highly simulated training and testing platform for the adaptive vision navigation method and system through multi-parameter scene configuration, the collision avoidance algorithm, and the real-time pedestrian behavior update mechanism. This can not only generate dynamic pedestrian behaviors but also cope with complex environmental changes, providing strong support for the development and optimization of the system.
[0031] It should be noted that the dynamic pedestrian simulator can be omitted when this dynamic pedestrian planning model is applied in an actual scenario.
[0032] Regarding 2) Pedestrian Detection and Prediction Module; In a dynamic pedestrian environment, the navigation system monitors pedestrians and their associated objects in real time through the pedestrian detection and prediction module. The robot uses an RGBD vision sensor to collect RGB images and depth images of the surrounding environment and aligns them to obtain the point cloud data of the environment. On this basis, a pre-trained two-stream branch neural network model is used to detect pedestrians and their associated objects in the scene in real time, obtaining the corresponding mask results. The mask results not only contain the position information of pedestrians but also cover the information of other objects related to pedestrians, thus providing more comprehensive environmental perception for subsequent navigation decisions.
[0033] The detected pedestrian position information Corresponds to spatial coordinates respectively, and maps this position information to the global map, so that the spatial position of pedestrians is incorporated into the layout of the entire environment. The global map is a comprehensive and complete environmental map that contains all elements in the physical space and provides basic environmental information for the robot's navigation.
[0034] To ensure the safety and efficiency of navigation, the motion trajectory of the detected pedestrians is predicted through a Kalman filter. The position and velocity state of the pedestrian Is expressed as: ; Where, Is The three-axis position information of the pedestrian at time Is The velocity information of the pedestrian at time Is The motion direction of the pedestrian at time
[0035] The future position of the pedestrian is predicted by the following formula: ; Where, Is The pedestrian prediction result at time Is The pedestrian detection result at time Is the state transition matrix, which is set in advance and corresponds to the preset time interval, Is the process noise. Through this pedestrian detection and prediction module, the motion of pedestrians can be sensed in advance and the obstacle information can be updated.
[0036] In addition, the formula for the point cloud position prediction result of pedestrians and their associated objects is as follows: ; Among them, and are respectively and the point cloud positions at moments, indicating the position transformation matrix of the point cloud under this position change. In this way, the navigation system can track the position and movement trajectory of pedestrians in real time, providing accurate dynamic environment information for path planning.
[0037] Therefore, the role of the pedestrian detection and prediction module is to use the Kalman filter for prediction, store the movement trajectories of pedestrians in the past time period, and combine with the point cloud position prediction to achieve accurate prediction of the pedestrian position at future moments, thus providing support for the safe navigation of the robot.
[0038] Regarding 3) the visual localization module for dynamic scenarios; In a dynamic pedestrian environment, the navigation system achieves accurate estimation of the robot's own position and attitude through the visual localization module for dynamic scenarios. The robot uses the RGB images and depth images obtained by the RGBD visual sensor, and after alignment processing, obtains the point cloud data of the environment, providing basic information for visual localization. On this basis, the visual localization module first extracts the ORB feature points and descriptors of the input visual image (environmental scene image) in the initial state, uses the point cloud information to restore the three-dimensional coordinates of the feature points, and then establishes the initial map points to provide a reference for subsequent localization.
[0039] In each subsequent frame, the visual localization module filters the extracted feature points according to the mask output by the pedestrian detection and prediction module, removing the feature points occluded by pedestrians or related to pedestrians to reduce the impact of pedestrian dynamic changes on the localization accuracy. Then, the filtered feature points are matched with the feature points in the map, and the pose information of the robot for each frame is solved through non-linear optimization. The specific solution formula is as follows: ; where represents the homogeneous coordinates of the 2D image after filtering through , is the function that projects 3D points onto the 2D image, represents the camera internal parameter matrix, represents the external parameter matrix of the camera, that is, the pose information of the robot for each frame to be solved, represents the homogeneous coordinates of the 3D feature points in the map.
[0040] In this way, the visual localization module for dynamic scenarios can estimate the position and pose of the robot in real time and accurately in a dynamic environment, providing reliable localization information for path planning. The key to this module lies in its ability to effectively filter and match feature points in a scene with dynamic pedestrians, and accurately solve the pose through an optimization algorithm, thus ensuring the stability and reliability of the navigation system.
[0041] Regarding 4) the dynamic map update module; One of the cores of the adaptive visual navigation method is the dynamic map update module, which is particularly important in dynamic and pedestrian-dense scenarios. Traditional navigation systems mainly rely on static environment maps, but in a dynamic environment, the movement and position changes of pedestrians will continuously affect the navigation path and safety. Therefore, it is necessary to update the map information in the environment in real time, especially in dynamic scenarios with a large number of pedestrians, to ensure the effectiveness and safety of the navigation system in a dynamic environment.
[0042] To this end, the system uses the depth information obtained by the RGB-D sensor and combines it with the semantic segmentation module to process the image data. Through semantic segmentation, the system can label objects, pedestrians, etc. in the scene by category and generate a preliminary semantic map, in which the classification information of each pixel point (such as pedestrians, walls, target objects) is clearly identified, enabling the robot to more clearly recognize the environment.
[0043] In the real-time update process, the pedestrian detection and prediction module plays a particularly important role. This module uses sensors to detect pedestrians appearing in each frame in real time and performs tracking and prediction based on these detection results. By fusing the current detection results with future prediction information, the pedestrian detection and prediction module provides key data support for dynamic map update.
[0044] During the generation of the semantic map, the robot labels the obtained objects and pedestrians according to their categories in the global map, and at the same time projects the obstacle point cloud onto a 2D plane to obtain an obstacle map, and then generates a semantic map. This process not only includes the annotation of static objects, but also dynamically integrates the position and future trajectory information of pedestrians, ensuring that the semantic map can reflect the changes in the environment in real time.
[0045] When a pedestrian moves out of the field of view or cannot be detected, the corresponding obstacle will still be retained in the semantic map and will only be updated when this part is observed again, or until it exceeds the time threshold and then removed. In addition, when the robot encounters an area that cannot be passed through, a new obstacle area will be generated in this area and gradually expand until the robot changes the planned path. This mechanism effectively avoids the fluctuations in the semantic map caused by the temporary disappearance or accidental occlusion of pedestrians, ensuring the stability and reliability of the semantic map.
[0046] The pedestrian detection and prediction results of the pedestrian detection and prediction module are dynamically fused into the semantic map to ensure that the semantic map can reflect the latest positions of pedestrians and their future trajectories, thus helping the robot better plan paths and avoid collisions. This fusion process is achieved through the max-pooling layer Specifically, the max-pooling layer combines the semantic information detected in each frame with the original semantic map to generate an updated semantic map , that is: ; This max-pooling method ensures the dynamic update of pedestrians and obstacles by retaining the maximum semantic information at each position, preventing the loss or over-update of information caused by the introduction of new data.
[0047] This update mechanism can not only respond to the dynamic environment in real time but also ensure the stability of the semantic map. To avoid frequent updates of the semantic map due to the temporary disappearance of pedestrians or occlusion, etc., the system introduces a time threshold . When a pedestrian moves out of the field of view or cannot be detected for a period of time, the corresponding obstacle information is not immediately removed from the semantic map but remains in the semantic map until the set time threshold is exceeded. In this way, the system can avoid fluctuations in the semantic map caused by the temporary disappearance or occasional occlusion of pedestrians. In this way, the semantic map can retain important information while avoiding unnecessary frequent updates, maintaining the stability and efficiency of the system.
[0048] This dynamic map update module based on depth information, semantic segmentation, pedestrian detection and prediction results enables the adaptive vision navigation system to cope with complex and dynamic pedestrian environments, ensuring the accuracy and safety of navigation. The max-pooling layer map update method for temporal changes effectively balances the update frequency and map stability, providing more reliable navigation support for the robot in the real environment.
[0049] Regarding 5) the autonomous hybrid path planning module; In the path planning stage, the system first calculates the global optimal path through a heuristic search algorithm. When the target point is far from the robot, the heuristic search algorithm is used to calculate the global path . This algorithm can efficiently plan the optimal path from the current position of the robot to the target position, ensuring the efficient navigation of the robot in a static environment.
[0050] However, in a dynamic environment, the robot needs to cope with local dynamic changes, such as the sudden appearance or movement of pedestrians. To this end, the system uses a local planning algorithm (the Time Elastic Band algorithm, TEB) to perform local path optimization within a certain range around the robot. The TEB algorithm can adjust the local path planning in a timely manner according to the predicted information of the pedestrian position and the real-time environment information provided by the dynamic map update module to prevent collisions. The specific adjustment formula is as follows: ; where, the optimized local path, is the adjustment amount based on the dynamic changes of the current environment.
[0051] When the robot approaches the target object, in order to stop more precisely at a suitable peripheral position of the target object, the system switches to the Fast Marching Method (FMM). The FMM algorithm ensures that the robot reaches the target position accurately by calculating the optimal path on the gradient field. This algorithm provides higher accuracy when the robot approaches the target, ensuring the successful completion of the navigation task.
[0052] In addition, when the global path is blocked by new obstacles, the system will use the heuristic search algorithm again to calculate a new global path to cope with sudden changes in the environment and ensure that the robot can continue to move towards the target.
[0053] The autonomous hybrid path planning module adaptively selects the most suitable path planning algorithm for the current state according to different environmental conditions and task requirements. By combining the heuristic search algorithm, the Time Elastic Band algorithm, and the Fast Marching Method, the system can achieve efficient and safe navigation in a dynamic environment. This adaptive hybrid path planning strategy not only improves the success rate of navigation but also significantly enhances the navigation efficiency and reduces the calculation time of path planning.
[0054] The adaptive visual navigation method for dynamic pedestrian environments proposed in this embodiment realizes efficient and reliable navigation of robots in complex dynamic environments by constructing a complete navigation framework. This method uses a two-stream branch network to detect dynamic pedestrians and their associated objects, and utilizes the Kalman filter for pedestrian trajectory prediction to provide real-time pedestrian information for navigation. By combining visual positioning information and pedestrian prediction results, the semantic map is dynamically updated to ensure that the map can reflect environmental changes in real time and support the efficient execution of path planning. By adaptively selecting global path planning and local path optimization algorithms, this method can flexibly adjust the navigation strategy according to environmental dynamics and target positions, significantly improving the navigation efficiency and safety of robots in dynamic environments. In addition, by constructing a dynamic pedestrian simulator with diverse pedestrian behavior models, this method can simulate more realistic pedestrian behaviors, providing a highly simulated environment for the training and testing of navigation systems and further enhancing the performance of navigation systems in complex dynamic environments. This method can be applied to multiple fields such as smart homes, service robots, and public place guidance, and has broad market prospects.
[0055] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. An adaptive visual navigation method for a dynamic pedestrian environment, characterized in that, Including: Obtain the environmental scene image through the robot to get the point cloud data of the environment, and input the environmental scene image into the visual positioning and navigation framework, which includes a pedestrian detection and prediction module, a visual positioning module for dynamic scenes, a dynamic map update module, and an autonomous hybrid path planning module; The pedestrian detection and prediction module uses a pre-trained neural network model to perform real-time detection on pedestrians and associated items in the environmental point cloud data to obtain a mask, and predicts the future positions of pedestrians based on this; The visual positioning module for dynamic scenes extracts features based on the input environmental scene image, combines the point cloud data to restore the three-dimensional coordinates of the feature points, establishes an initial map, and in each subsequent frame, filters the extracted feature points according to the mask result and matches them with the feature points in the initial map to estimate the pose information of each frame; The dynamic map update module generates a semantic map using the received pose information, local map point cloud information, semantic information, and pedestrian position information, dynamically integrates the future positions of pedestrians into the semantic map, and updates it using the maximum pooling method with temporal changes; The autonomous hybrid path planning module ensures that the robot reaches the target position based on the updated semantic map and the target position, based on adaptive path planning; 2. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, characterized in that, The pedestrian detection and prediction module realizes the detection of dynamic pedestrians and their associated objects through a two-stream branch network. The two branch networks simultaneously utilize the features of the backbone network to calculate semantic information and optical flow information. Finally, the prediction head combines these two types of information and outputs the masks of dynamic pedestrians and associated objects. .
3. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, characterized in that Solve the pose information of the robot in each frame through a non-linear optimization method, and the solution formula is as follows: ; Among them, represents the homogeneous coordinates of the 2D image after filtering through the mask is a function that projects 3D points onto the 2D image represents the camera intrinsic matrix represents the extrinsic matrix of the camera represents the homogeneous coordinates of the 3D feature points in the initial map is the homogeneous coordinates of the 3D feature points in the initial map 4. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, wherein In the pedestrian detection and prediction module, Kalman filtering is used for pedestrian tracking, prediction of motion trajectories, and prediction of dynamic obstacle point clouds; The prediction result formula for the future position of pedestrians at the next moment is as follows: ; Among them, are respectively the pedestrian prediction results at the moment, is the pedestrian detection result at the moment, is the state transition matrix, is the process noise; The prediction result formula for the point cloud of pedestrians and their associated objects is as follows: ; Among them, and are respectively and the point cloud positions at the moments, represents and the position transformation matrix of the point cloud under the change of the point cloud positions at the moments.
5. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, wherein When updating the semantic map, when a pedestrian moves out of sight or cannot be detected, the corresponding obstacle will still remain in the semantic map and will only be updated when this part is observed again, or until the time threshold is exceeded and then removed; when the robot encounters an impassable area, a new obstacle area will be generated in this area and gradually expand until the robot changes the planned path.
6. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, wherein In the autonomous hybrid path planning module, the path planning stage is specifically as follows: The robot calculates the global optimal path through a heuristic search algorithm; When static obstacles do not obstruct the global path, use the time elastic band algorithm to plan the local path to avoid collisions; When new changes occur in the scene, use the heuristic search algorithm again to update the global optimal path; When the robot is within the range where the target position is located, switch to the fast marching algorithm, and ensure that the robot accurately reaches the area near the target object by calculating the optimal action on the gradient field; 7. The adaptive visual navigation method for a dynamic pedestrian environment according to claim 1, characterized in that, A dynamic pedestrian simulator is constructed to complete the training, testing, and parameter adjustment of the network module in the visual positioning and navigation framework; the dynamic pedestrian simulator sets scene configurations, robot positions, target positions, and pedestrian pose parameters according to requirements, quickly generates diverse pedestrian poses, and models the pedestrian poses using a parameterized 3D human model, dynamically adjusting the pedestrian poses in the global coordinate system and the local coordinate system; when collecting training data, all configurations of the dynamic pedestrian simulator are randomly set within a reasonable range for the robot's initial position, map information, target information, dynamic pedestrian behaviors, etc., and it runs automatically; 8. An adaptive visual navigation system for a dynamic pedestrian environment, characterized in that, Including: Including an image acquisition module and a visual positioning and navigation framework, which includes a pedestrian detection and prediction module, a visual positioning module for dynamic scenes, a dynamic map update module, and an autonomous hybrid path planning module; The image acquisition module is used to obtain the environmental scene image through the robot to get the point cloud data of the environment, and input the environmental scene image into the visual positioning and navigation framework; The pedestrian detection and prediction module uses a pre-trained neural network model to perform real-time detection on pedestrians and associated objects in the environmental scene image to obtain a mask, and predicts the future positions of the pedestrians based on this; The visual positioning module for dynamic scenes extracts features based on the input environmental scene image, combines the point cloud data to restore the three-dimensional coordinates of the feature points, establishes an initial map, and filters the restored feature points according to the mask and matches them with the feature points in the initial map in subsequent frames to estimate the pose of each frame; The dynamic map update module generates a semantic map using the received pose information, local map point cloud information, semantic information, and pedestrian position information, dynamically integrates the future positions of the pedestrians into the semantic map, and updates it using the maximum pooling method with temporal changes; The autonomous hybrid path planning module ensures that the robot reaches the target position based on the updated semantic map and the target position, based on adaptive path planning.
9. The adaptive visual navigation system for a dynamic pedestrian environment according to claim 8, characterized in that The pedestrian detection and prediction module realizes the detection of dynamic pedestrians and their associated objects through a two-stream branch network. The two branch networks simultaneously use the features of the backbone network to calculate semantic information and optical flow information. Finally, the prediction head combines these two types of information and outputs the masks of dynamic pedestrians and associated objects. .
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the adaptive visual navigation method according to any one of claims 1-7.
Citation Information
Cited By
Target navigation method and system for active perception and rule-guided reinforcement learning
CN120831116A