Autonomous navigation method based on POI point extraction and target fusion
Through the autonomous navigation method based on the POI point extraction and target fusion, combined with Fast ViT and hierarchical SAC algorithm, the problem of inefficiency and poor adaptability in long-distance tasks in unknown environments is solved, dynamic POI extraction and global local collaborative optimization are realized, and the robustness and migratory performance of the model are improved.
Patent Information
- Application Number
- CN202510560055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing autonomous navigation technology is inefficient, reward sparsity and poor adaptability in long-distance tasks in unknown environments.
The autonomous navigation method based on the POI point extraction and target fusion is adopted, and candidate POI points are dynamically generated through the laser sensor to obtain environmental data in real time. The spatial relationship between visual features and POI sub-targets is deeply coupled with the Fast ViT perception module. The reward weight is adjusted based on the hierarchical SAC algorithm and the dangerous area transfer function, and the control instructions are output.
Dynamic POI extraction and global local collaborative optimization are realized, attention to the information related to the target in the environment is improved, reward sparsity is solved, and model robustness and migratory performance are enhanced.
Smart Images

Figure CN120088761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous navigation, and particularly relates to an autonomous navigation method based on POI point extraction and target fusion. Background Art
[0002] In the prior art, autonomous navigation systems mainly rely on traditional path planning algorithms (such as A*, Dijkstra, etc.). Such methods need to rely on global environmental prior map information, and the navigation efficiency decreases significantly in unknown environments due to the lack of map support. Although the navigation method based on deep reinforcement learning (DRL) reduces the dependence on environmental priors, it still faces many challenges: the physical state of the existing method's target is separated from the scene representation, resulting in the decision-making network being difficult to effectively fuse task objectives and environmental perception information, and the data efficiency is low; although the navigation based on lidar (LiDAR) has high accuracy, it has a large amount of data, high real-time processing computational resource consumption, and cannot distinguish approximate objects in complex environments, while visual sensors, although low-cost, are prone to falling into local optima in long-distance tasks, and their robustness and transferability are insufficient; in addition, the design of traditional DRL reward functions is sparse, resulting in slow convergence speed of agent training, low navigation success rate, and most existing methods focus on short-distance navigation scenarios, lacking a mechanism for dynamically extracting sub-goals and global-local collaboration, and it is difficult to handle complex environments with multiple regions and long distances. Existing navigation research based on vision transformers (ViT) is mostly limited to small-scale scenarios, and does not fully solve the problem of deep coupling between target states and scene representations, resulting in insufficient adaptability of the model to dynamic obstacles and complex lighting changes in long-distance tasks; at the same time, the laser-based global exploration strategy lacks real-time collaboration with visual perception, and it is difficult to balance the efficiency of local obstacle avoidance and global path planning.
[0003] In addition, existing methods often experience a significant decline in performance due to the differences between virtual and real environments in Sim2Real migration, especially in practical scenarios such as sensor noise and dynamic obstacle interference, and their robustness is insufficient. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide an autonomous navigation method based on POI point extraction and target fusion, which solves the problems of low efficiency, sparse rewards, and poor adaptability existing in the prior autonomous navigation technology in long-distance tasks in unknown environments.
[0005] Technical Solution: An autonomous navigation method based on POI point extraction and target fusion according to the present invention includes: (1) Real-time obtain environmental distance data through a laser sensor, dynamically generate candidate POI points based on two methods: continuous ranging difference exceeding a threshold and open area detection with a ranging upper limit, and construct a POI repository by combining obstacle proximity verification and access area filtering; (2) The perception module with Fast ViT as the core performs multi-scale feature fusion on the multi-frame stacked fisheye camera images, encodes the spatial relationship of POI sub-goals, i.e., relative distance and heading deviation, as a physical state vector and deeply couples it with visual features; (3) Based on the hierarchical SAC algorithm, the fused multi-modal state is input into the Critic network with a safety reward mechanism, and the reward weight is dynamically adjusted through the dangerous area transfer function, and the linear velocity and angular velocity control commands are output; (4) A multi-scenario training environment is constructed on the GAZEBO platform, and the domain randomization technology is used to perturb the parameters of sensor noise, lighting conditions and obstacle distribution to enhance the Sim2Real migration ability.
[0006] Furthermore, the dynamic POI extraction stage is as follows: when the difference between two adjacent laser rangefinder readings exceeds the preset threshold, it is determined as a passable gap and POIs are generated at the corresponding coordinates; when the continuous laser readings exceed the upper limit of the rangefinder, the center point of the open area is converted into a POI in the global coordinate system; through real-time obstacle detection and comparison with the historical path, the POIs located in the adjacent area of the obstacle or the visited area are dynamically cleared.
[0007] Furthermore, the evaluation function of the POI point is: ; where is the Euclidean distance between the position p of the vehicle at the current time step and the candidate point, the Euclidean distance between the candidate point and the global target, is the map information of the candidate point at the current time step; is used to adjust the weight of the Euclidean distance between the current vehicle position and the candidate point in the formula.
[0008] Furthermore, the core of the perception part adopts the FastViT model, including: a multi-scale feature extraction module that performs local feature extraction on the image through convolutional kernels of different sizes and performs multi-scale feature splicing; a target state encoding module that encodes the relative distance and heading deviation of the POI sub-goal as a physical state vector; and a feature coupling module that deeply couples the visual feature and the target state vector through the multi-head attention mechanism of Fast ViT to generate the fused multi-modal feature.
[0009] Furthermore, the dangerous area transfer function is defined as: The dangerous area transfer function includes: dynamic division of the safe area, based on real-time monitoring by the laser sensor, setting a safe distance threshold d safe = 0.5m, if the distance of the obstacle d obs ≥ dsafe , marked as the safe area ( ); conversely, marked as the dangerous area (F); the state transition reward mechanism includes positive rewards. When the vehicle enters the safe area ( ) from the dangerous area (F), it is negatively punished. When the vehicle enters the dangerous area from the safe area.
[0010] Furthermore, the reinforcement learning reward function includes: the reinforcement learning reward function includes: goal-oriented rewards, that is, rewards for approaching the POI point, which encourages reducing the distance from the sub-goal; heading alignment rewards, which reduce the heading deviation. Action smoothness penalties, penalties for sudden changes in linear velocity and angular velocity. Safety and task rewards, the reward value for reaching the POI point, collision penalties, and transfer rewards for the dangerous area as described above.
[0011] Furthermore, the domain randomization technique specifically includes: sensor noise injection: adding Gaussian noise to the laser ranging value; applying color jitter and motion blur to the RGB image; illumination condition perturbation: randomly adjusting the image brightness and contrast by ±30% in the HSV space; deformable obstacle parameters: randomly setting the obstacle size, material friction coefficient, and elastic coefficient.
[0012] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory. When the processor executes the program, it implements the steps of any one of the methods described above.
[0013] A computer-readable storage medium according to the present invention stores a computer program. When the program is executed by a processor, it implements the steps of any one of the methods described above.
[0014] Compared with the prior art, the present invention has the following remarkable advantages: The dynamic POI extraction method is optimized globally and locally in a coordinated manner. The optimal secondary target point is selected through the POI evaluation function, which solves the local optimum problem of long-distance navigation. FastViT performs multi-scale fusion on multi-frame stacked images, encodes the relative distance and heading deviation of the POI sub-goal into a physical state vector, and deeply couples it with visual features, enhancing the attention to target-related information in the environment. The multi-layer reward function framework solves the sparse problem of the reward function, and the dangerous area transfer function realizes the safety of the planned path. The sensor noise, illumination conditions, and obstacle distribution randomization techniques improve the robustness and transferable performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the flowchart of the present invention; Figure 2 is the algorithm framework diagram of the present invention; Figure 3 is the method diagram for extracting the point of interest (POI) of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings.
[0017] As Figure 1 shown, an autonomous navigation method based on POI point extraction and target fusion provided by an embodiment of the present invention includes the following steps: Step 1: Obtain the secondary target points of the POI points: Install a laser sensor on the vehicle to collect environmental data in real time at a sampling frequency of 40 Hz, which is used to obtain environmental distance data in real time and dynamically extract points of interest (POI). As Figure 3 shown, continuously obtain the ranging data of the laser sensor in real time, and calculate the difference between adjacent two measurement data in each direction. If the ranging difference in a certain direction exceeds a preset threshold (for example, the difference ≥ 0.5 m), it is determined that there is a passable gap at this position, and the system adds the corresponding coordinates as candidate POIs and stores them in the POI list in the memory. If the laser data in a certain direction exceeds the ranging upper limit of the sensor, it is determined as an obstacle-free area. According to the angle information in the polar coordinate system of the sensor, the coordinates of the center point of the open area are converted to the global coordinate system, and the system adds a candidate POI point to the POI storage library at this position, marked as "potentially passable area". Combining the current laser ranging data or the feedback of the vision sensor, if there are obstacles around a certain POI, delete the POI from the list. Record the historical path coordinates of the trolley. If the candidate POI is located in the visited area (for example, the Euclidean distance from the historical path point ≤ 0.2 m), prohibit it from joining the list of points of interest to avoid repeated exploration. Before each navigation decision, give priority to cleaning up invalid or redundant POIs to ensure that the list of points of interest only contains feasible candidate points in the current environment.
[0018] The evaluation formula for the POI points is: ; where is the Euclidean distance between the position p of the trolley at the current time step and the candidate point, the Euclidean distance between the candidate point and the global target, is the map information of the candidate point at the current time step; is used to adjust the weight of the Euclidean distance between the current trolley position and the candidate point in the formula.
[0019] Evaluate the POI points in the POI storage library through the above formula, and obtain the optimal POI point as the next secondary target point. When the trolley completes the navigation task from the current position to the next secondary target point, the system re-evaluates the POI storage library to select the next secondary target point, continuously divides the global target into short-distance secondary target points, selects the optimal POI point as the next target point, and provides accurate secondary target guidance for long-distance autonomous navigation.
[0020] Step 2: Use a perception module with Fast ViT as the core to perform multi-scale feature fusion on multi-frame stacked fisheye camera images, encode the spatial relationship of POI sub-goals, i.e., relative distance and heading deviation, as a physical state vector, and deeply couple it with visual features; In the embedding fusion step, the process from the current position of the vehicle to the next secondary path point is called local navigation. As the leading part of the Actor Network, a Transformer architecture based on FastVit that supports multi-modal input is used to perform local feature extraction on the input fisheye camera images and then fuse the target physical information to extract high-level abstract features in the Transformer. First, the image data obtained by the fisheye camera is input, and the FastVit architecture is used. Convolution operations are performed using different-sized convolutional kernels, and then they are concatenated to achieve multi-scale feature fusion. At the lower layer of the architecture, relevant units focus on local regions to obtain local features, which are the basis for subsequent fusion and feature extraction. By integrating key information such as the distance difference and angle deviation between the current vehicle and the target, the target value is calculated. Define to reflect the spatial relationship between the current position of the vehicle and the expected target position, which contains the key connection between the vehicle and the target in space and is used for subsequent fusion with image features. The image state input is , and the resulting state fusion features will be passed into the Vision Transformer Architecture. The attention mechanism of the hybrid structure is one of the key design elements to improve information utilization and task accuracy. This network adopts an encoder structure repeated twice, including a multi-head self-attention module, a multi-layer perceptron, a fully connected layer, and a layer normalization module. After completing the feature extraction of the perceptual data, combined with the subsequent decision-making system, it shows the target-based scene representation through visual attention flow mapping.
[0021] Step 3: Based on the hierarchical SAC algorithm, input the fused multi-modal state into the Critic network with a safety reward mechanism, dynamically adjust the reward weight through the dangerous area transfer function, and output the linear velocity and angular velocity control commands; In the decision-making step, as Figure 2 shown, the input of the SAC algorithm consists of two parts, namely the target state of the POI point and the visual state of the original RGB image. Using a raw image of 160×120, visual data is obtained from a fisheye camera with a FOV of 220 degrees, and the images of the nearest four frames are stacked. Gaussian noise is added to the stacked images to enhance the robustness and transferability of the training model. The local navigation target state obtained from the upstream task is defined in two dimensions as the relative distance and heading deviation. The first dimension of the target state is defined as the relative position between the current vehicle position and the secondary target point: ; wherein, is shown as the real-time position of the current trolley, represents the position of the target point, is the Euclidean norm operation, is an ordinary regularization factor, and its value range is [0, 1]. The second dimension of the target state is associated with the heading error between the heading of the trolley and the direction vector pointing to the target position.
[0022] ; In the formula represents the heading of the trolley, The function calculates the angle, represents the y coordinates of the target point, represents the x coordinates of the target point, is shown as the y coordinates of the real-time position of the current trolley, is shown as the x coordinates of the real-time position of the current trolley. It is input into the SAC reinforcement learning network by constructing the relationship between the current position and the target state; according to the differential drive trolley chassis structure adopted, the action instruction to be executed is , the linear velocity is limited in the range of [0, 1], and the angular velocity , and then it is sent to the mobile trolley according to the trolley operating system (ROS). In the GAZEBO training environment, the trolley is regarded as an agent, and the position of the agent and the navigation target point are randomly initialized. The optimal POI point is selected through the laser data and guided to the target point. The agent selects the optimal action to execute in the environment according to the obtained environmental and state information. The environment gives the agent corresponding feedback and reward and punishment mechanisms based on the reward function and sensor information; for one interaction completed by the agent and the environment, a quadruple data of state, reward, action, and next state will be generated, and the interaction data is stored in the experience replay pool (Replay buffer). The reward function consists of heuristic performance, action, reward for reaching the target point, collision penalty, and performance index for transferring in the dangerous area; the heuristic performance is the reward and punishment for encouraging the machine to move towards the target position, the action item is to limit the action of the trolley to prevent the trolley from getting into self-rotation and reduce the number of turns, and the dangerous area transfer function includes the distance of the trolley itself from the obstacle and the probability of possible collision, and corresponding positive and negative rewards are set based on this. The reinforcement learning algorithm selects the optimal action according to the current policy and the environment, that is, it can output the terminal instructions of the linear velocity and angular velocity that the trolley can understand.
Claims
1. An autonomous navigation method based on POI point extraction and target fusion, characterized in that: include: (1) Real-time acquisition of environmental distance data through laser sensors, dynamic generation of candidate POI points based on continuous ranging difference exceeding threshold and ranging upper limit open area detection, and construction of a POI repository by combining obstacle proximity verification and access area filtering; (2) Using Fast ViT as the core perception module, we perform multi-scale feature fusion on multi-frame stacked fisheye camera images, encode the spatial relationship of POI sub-targets, i.e., relative distance and heading deviation, into physical state vectors and deeply couple them with visual features; (3) Based on the hierarchical SAC algorithm, the fused multimodal state is input into the Critic network with a safety reward mechanism, the reward weight is dynamically adjusted through the dangerous area transfer function, and the linear velocity and angular velocity control instructions are output; (4) A multi-scenario training environment is built on the GAZEBO platform, and domain randomization technology is used to perturb the parameters of sensor noise, lighting conditions, and obstacle distribution to enhance the Sim2Real migration capability.
2. The autonomous navigation method based on POI point extraction and target fusion according to claim 1, characterized in that: The dynamic POI extraction stage is as follows: when the difference between two adjacent laser ranging measurements exceeds the preset threshold, it is determined that the gap can be passed and a POI is generated at the corresponding coordinates; when continuous laser readings exceed the upper limit of ranging, the center point of the open area is converted to a POI in the global coordinate system; through real-time obstacle detection and historical path comparison, POIs located in the area adjacent to the obstacle or in the visited area are dynamically cleared.
3. The autonomous navigation method based on POI point extraction and target fusion according to claim 2 is characterized in that: The POI evaluation function is: ; in, is the Euclidean distance between the car position p at the current time step and the candidate point, The Euclidean distance between the candidate point and the global target, is the map information of the candidate points of the current time step; Used to adjust the weight of the Euclidean distance between the current car position and the candidate point in the formula.
4. The autonomous navigation method based on POI point extraction and target fusion according to claim 2, characterized in that: The core of the perception part adopts the FastViT model, which includes: a multi-scale feature extraction module, which extracts local features of the image through convolution kernels of different sizes and performs multi-scale feature splicing; a target state encoding module, which encodes the relative distance and heading deviation of the POI sub-target into a physical state vector; a feature coupling module, which deeply couples the visual features with the target state vector through the multi-head attention mechanism of Fast ViT to generate fused multi-modal features.
5. The autonomous navigation method based on POI point extraction and target fusion according to claim 1, characterized in that: Dangerous area transfer functions include: dynamic division of safe areas, real-time monitoring based on laser sensors, and setting safety distance thresholds d safe = 0.5m, if the obstacle distance d obs ≥ d safe , marked as a safe area ; Otherwise it is marked as dangerous area F; the state transfer reward mechanism includes positive rewards, the car enters the safe area from the dangerous area F Negative penalty, the car enters the danger zone from the safe zone.
6. The autonomous navigation method based on POI point extraction and target fusion according to claim 1, characterized in that: The reinforcement learning reward function includes: goal-oriented reward, that is, reward for approaching the POI point, which encourages reducing the distance to the sub-goal; heading alignment reward, which reduces heading deviation; action smoothness penalty, linear velocity and angular velocity mutation penalty; safety and task reward, reward value for reaching the POI point, collision penalty, and reward for transferring to the dangerous area.
7. The autonomous navigation method based on POI point extraction and target fusion according to claim 1, characterized in that: The domain randomization technology specifically includes: sensor noise injection: adding Gaussian noise to the laser ranging value; applying color jitter and motion blur to the RGB image; lighting condition perturbation: randomly adjusting the image brightness and contrast by ±30% in the HSV space; obstacle deformable parameters: randomly setting the obstacle size, material friction coefficient and elastic coefficient.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Indoor navigation method based on vision and radar information fusion and reinforcement learning
CN116263335A
Cited By
Visual navigation method and system based on reinforcement learning
CN120890466A