A Target Perception and Avoidance Decision-Making Method in Visual Navigation of Intelligent Equipment
By combining obstacle perception and avoidance decision-making methods with visual sensors and radar data, a semantically enhanced risk field is generated, which solves the problems of insufficient path planning and poor adaptability to environmental changes in the visual navigation of intelligent equipment, and achieves more efficient obstacle avoidance and path planning.
Patent Information
- Application Number
- CN202511197833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing intelligent equipment visual navigation systems lack comprehensive utilization of environmental semantic information in path planning, resulting in insufficiently refined path planning, poor adaptability to environmental changes, and low accuracy in obstacle detection.
By utilizing computational vision sensors and radar data, combined with deformable convolutional kernels and spatiotemporal graph neural networks, a semantically enhanced risk field is generated. The path planning is then performed using the dynamic step size RRT* algorithm, enabling accurate obstacle perception and avoidance.
It improves the precision of path planning and adaptability to environmental changes, reduces perception errors, enhances the safety and reliability of the system in complex environments, and improves the robustness and real-time performance of path planning.
Smart Images

Figure CN120685119B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent decision-making technology, and more specifically, to a target perception and avoidance decision-making method in visual navigation of intelligent equipment. Background Technology
[0002] In visual navigation systems for intelligent equipment, path planning and obstacle avoidance have always been key technologies. With the continuous development of autonomous driving and intelligent robotics, visual navigation systems increasingly rely on sensors (such as cameras, radar, and lidar) to perceive the surrounding environment. These sensors provide detailed data on obstacles, road conditions, traffic signs, and other environmental features, thus supporting the decision-making process of intelligent equipment. By fusing this information, the system can more comprehensively understand and adapt to complex driving environments; however, traditional path planning methods often rely on data from single sensors and frequently only consider the geometric features of obstacles, lacking comprehensive utilization of environmental semantic information. Semantic information (such as lane type, traffic signs, and speed limits) is often not fully integrated into path planning, making existing systems lack sufficient flexibility and adaptability when facing complex dynamic environments.
[0003] Existing path planning technologies often rely solely on the position and velocity of obstacles for calculation, neglecting semantic information in the environment. This results in insufficiently refined path planning and poor adaptability to environmental changes. Furthermore, the accuracy of feature fusion and obstacle detection is low. Therefore, this paper proposes a target perception and avoidance decision-making method for visual navigation in intelligent equipment. Summary of the Invention
[0004] The purpose of this invention is to provide a target perception and avoidance decision-making method in visual navigation of intelligent equipment, so as to solve the problems of insufficient path planning and poor adaptability to environmental changes mentioned in the background art.
[0005] To achieve the above objectives, the present invention aims to provide a target perception and avoidance decision-making method in visual navigation of intelligent equipment, comprising the following steps:
[0006] Step S1: Use a computer vision sensor to acquire image feature maps of obstacles, use radar to collect 3D point clouds of obstacles in real time, perform feature alignment on the image feature maps in BEV space based on deformable convolution kernels, and introduce transmission delay for optimization. Then, fuse the aligned image feature maps with the radar feature maps to generate a fused feature map.
[0007] Step S2: Construct a spatiotemporal graph neural network, using the fused feature map as the input to the spatiotemporal graph neural network, introducing semantic label encoding to define the nodes and edges of the spatiotemporal graph neural network, and then predicting the trajectory of each obstacle through a regression network;
[0008] Step S3: Define environmental semantic labels and couple them with the predicted obstacle trajectories to generate a semantically enhanced risk field;
[0009] Step S4: Construct a semantic road network graph based on the semantically enhanced risk field, generate a global path point sequence, and use the current global path point as a temporary target to call the dynamic step size RRT* algorithm to avoid obstacles.
[0010] Furthermore, in S1, the specific steps for feature alignment of the image feature map in the BEV space based on deformable convolution kernels are as follows:
[0011] Step S11: Divide the obstacle into a two-dimensional grid on the horizontal plane, project the 3D point cloud onto the grid, and generate a radar feature map. The image feature map is projected onto the BEV space to generate the image BEV feature map. ;
[0012] Step S12: Transfer radar feature map and image BEV feature map The concatenation is used as input to a convolutional network, which then predicts the offset of the deformable convolutional kernel at each location. ;
[0013] Step S13: Use the predicted offset BEV feature map of image Perform deformable convolution to obtain the aligned BEV feature map of the image, as shown below:
[0014]
[0015] In the formula, Spatial location on the BEV feature map of the image; For in position BEV feature map of the image after deformable convolution alignment; ; The sampling point location for the standard convolution kernel; For the offset position BEV feature map extracted from the image; The weights are the weights of the convolution kernel.
[0016] Furthermore, in S11, the obstacle is divided into a two-dimensional grid on the horizontal plane, wherein each two-dimensional grid includes the maximum height, height distribution variance, point cloud density, reflection intensity entropy, temporal stability index, multi-frame motion vector variance, semantic confidence probability, and height gradient.
[0017] Furthermore, in step S1, a transmission delay is introduced for optimization, and the aligned image feature map and radar feature map are fused to generate a fused feature map, specifically including:
[0018] The optimized BEV feature map of the image is obtained by optimizing the transmission delay.
[0019] The optimized image BEV feature map is concatenated with the radar feature map, and the optimized offset of the deformable convolution kernel at each location is predicted based on the convolutional network.
[0020] The optimized predicted offset is used to perform deformable convolution operation on the BEV feature map of the image to obtain the aligned optimized BEV feature map of the image.
[0021] The aligned and optimized image BEV feature map is fused with the radar feature map to generate a fused feature map.
[0022] Furthermore, in S2, the nodes and edges of the spatiotemporal graph neural network are defined as follows:
[0023] For each obstacle Extract the feature vectors of the corresponding positions from the fused feature map. ;
[0024] Define the graph nodes and edges of a spatiotemporal graph neural network:
[0025] V i =[ x i , y i , v xi , v yi , S emi i , f i ]
[0026] in, For obstacles node; For obstacles The coordinates of the centroid; For obstacles velocity components; For obstacles semantic tags;
[0027] The interaction between different obstacles is represented by the following formula:
[0028]
[0029] In the formula, For obstacles and obstacles The interaction relationship; For obstacles and obstacles The relative position vector, Δ p ij =[ x j - x i , y j - y i ] ; For obstacles and obstacles The relative velocity vector, Δ v ij =[ v xj - v xi , v yj - v yi ] ; For obstacles semantic tags; For encoding functions; For obstacles eigenvectors;
[0030] Finally, update the node and generate obstacles. final state .
[0031] Furthermore, in S2, the specific steps for predicting the trajectory of each obstacle using a regression network are as follows:
[0032] Step S21: Based on the semantic labels of obstacles, generate a set of initial behavioral intent vectors according to preset rules. ;
[0033] Step S22: Obstacles output by the spatiotemporal graph neural network Final state Initial behavioral intent vector and obstacles eigenvectors The data is concatenated and input into the Transformer decoder, which then predicts the next step in turn. The trajectory at time Location ;
[0034] Step S23: For each obstacle, the decoder can output... Each trajectory contains 1 independent trajectory. The predicted location sequence at each time step is calculated, and the probability of each trajectory is also calculated.
[0035] Furthermore, in S3, the specific steps for generating the semantically enhanced risk field are as follows:
[0036] Step S31: Obtain the environmental semantic risk field corresponding to the environmental semantic label. Different categories of environmental semantic labels correspond to different risk values.
[0037] Step S32: For each obstacle, calculate the obstacle risk field based on its predicted trajectory. The obstacle risk field is the sum of the risks generated by all predicted trajectories of the obstacle.
[0038] Step S33: Generate a spatially related modulation function based on the environmental semantic tags. This is used to adjust the impact of obstacle risk; the modulation function can modify the obstacle risk field according to the semantic category of the location.
[0039] Step S34: Overlay the environmental semantic risk field and the corrected obstacle risk field to generate a semantically enhanced risk field.
[0040] Furthermore, the environmental semantic tags include static environmental tags and dynamic environmental tags, and the environmental semantic risk field is jointly determined by the static environmental tags and the dynamic environmental tags.
[0041] Furthermore, in step S4, the specific steps for avoiding obstacles are as follows:
[0042] Step S41: Based on the semantically enhanced risk field, construct a semantic road network graph where nodes represent intersection centers and lane sampling points, and edges carry road type, speed limit, and driving direction constraints. In semantic road network graph The A* algorithm is used, with a cost function that includes road semantic cost and risk field risk value. Perform a search to generate a global pathpoint sequence. ;
[0043] Step S42: The local layer uses the current global path point For the temporary target, the Dynamic Step Size (RRT) algorithm is invoked, where the dynamic step size is based on the semantically enhanced risk field at the current position. Adaptive computation, and real-time querying of the semantically enhanced risk field as nodes expand. Risk field threshold As an extended hard constraint, it is also determined whether the replanning trigger condition is met. If it is met, replanning is executed.
[0044] Step S43: When the number of consecutive replanning failures reaches the preset limit. If the current path is not found, return to the global layer and re-execute step S41, freezing the current path region in the semantic road network graph; otherwise, use the current position as the new starting point and return to the previous global path point. Continue with step S42.
[0045] Furthermore, in S42, the replanning triggering conditions include:
[0046] The cumulative risk value of a local path exceeds the preset cumulative risk threshold;
[0047] Local path costs are higher than the historical average.
[0048] Semantic conflict, i.e., semantically compatible functions Less than the semantically compatible function threshold.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. In the target perception and avoidance decision-making method of the intelligent equipment's visual navigation, a semantically enhanced risk field is used to combine environmental semantic labels with obstacle trajectories. This enables more accurate assessment of risks in different areas and adjusts risk values based on obstacle categories and locations. In path planning, the semantically enhanced risk field is combined with the dynamic step size (RRT*) algorithm, which automatically adjusts the step size according to risk changes during real-time obstacle avoidance. This results in more refined path planning in high-risk areas and faster search in low-risk areas, improving the system's safety and reliability in complex environments and reducing decision-making errors caused by environmental changes.
[0051] 2. In the target perception and avoidance decision-making method of the intelligent equipment's visual navigation, the transmission delay is introduced for optimization, which can achieve precise spatiotemporal synchronization compensation of visual and radar data, effectively eliminating feature misalignment caused by differences in sensor sampling frequency and processing delay; at the same time, by "backtracking" sampling through the product of velocity vector and delay, the position of dynamic obstacles in BEV space is accurately reproduced, thereby significantly improving the spatial consistency and semantic integrity of multimodal feature fusion, further reducing perception error, and enhancing the robustness and real-time performance of subsequent trajectory prediction and path planning. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 This is a flowchart of the overall method of the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0055] Please see Figure 1 As shown, this embodiment provides a target perception and avoidance decision-making method in visual navigation of intelligent equipment, which includes the following steps:
[0056] Step S1: Use a computer vision sensor to acquire image feature maps of obstacles, use radar to collect 3D point clouds of obstacles in real time, perform feature alignment on the image feature maps in BEV space based on deformable convolution kernels, and introduce transmission delay for optimization. Then, fuse the aligned image feature maps with the radar feature maps to generate a fused feature map.
[0057] Visual sensors (such as cameras) acquire real-time image data of the surrounding environment. These images typically contain information such as the color, shape, and texture of obstacles, representing the semantic features of objects. Radar (such as millimeter-wave radar or lidar) collects point cloud data or distance information of the environment in real time. Radar data typically provides spatial geometric features of objects, such as distance, velocity, and orientation.
[0058] In step S1, the specific steps for dynamically aligning the image feature map in the BEV space using deformable convolution kernels are as follows:
[0059] Step S11: Divide the obstacle into a two-dimensional grid on the horizontal plane. Each grid cell represents a fixed region in the physical world. Project the 3D point cloud onto the grid to generate a radar feature map. The image feature map is projected onto the BEV space to generate the image BEV feature map. Since perspective transformation causes image features to be deformed and occluded in BEV space, further alignment is required.
[0060] In step S11, the obstacle is divided into a two-dimensional grid on the horizontal plane. Each two-dimensional grid includes the maximum height, height distribution variance, point cloud density, reflection intensity entropy, temporal stability index, multi-frame motion vector variance, semantic confidence probability, and height gradient.
[0061] Among them, the height distribution variance is the variance of the height of all point clouds within the statistical grid cell, which reflects the degree of undulation of the terrain in the area and helps to distinguish between flat ground and uneven obstacles.
[0062] Reflection intensity entropy describes the uniformity of echo intensity distribution in the form of information entropy, enhancing the ability to identify multi-material, multi-directional reflective surfaces.
[0063] The temporal stability index records the rate of change of the point cloud presence ratio of the grid over the past N frames, which is used to determine whether the obstacle is static or dynamic, thereby improving the response speed to moving targets.
[0064] The multi-frame motion vector variance is calculated based on the point cloud position difference of consecutive frames, and the variance of all motion vectors within the grid is used to quantify the motion consistency of dynamic targets within the region.
[0065] Semantic confidence probability projects the semantic classification probabilities (pedestrians, vehicles, curbs, etc.) from vision onto the grid and calculates the average confidence to enhance multimodal semantic perception.
[0066] The height gradient calculation is the average height difference between the current grid and its neighboring grids, which helps to identify terrain features such as slopes and steps, and improves the accuracy of path planning.
[0067] Step S12: Transfer radar feature map (For reference) and image BEV feature map (Features to be aligned) are concatenated and used as input to a convolutional network. The convolutional network then predicts the offset of the deformable convolutional kernel at each location. This offset is applied to each sampling point of the BEV feature map of the image.
[0068] A convolutional network consists of an input layer, convolutional layers, and a fusion layer. The convolutional network predicts the offset of the deformable convolutional kernel at each location by performing convolution operations on the concatenated feature map of the input.
[0069] Due to sensor calibration errors, dynamic object motion, and time asynchrony, radar feature maps and image BEV feature maps may have local misalignment in BEV space. Deformable convolution is used to learn the offset of each sampling point, thereby adaptively aligning the features.
[0070] The BEV feature map of the image to be aligned is ;in, Indicates the height of the feature map. Indicates the width of the feature map. This represents the number of channels in the BEV feature map of the image.
[0071] The reference radar feature map is ;in, This indicates the number of channels in the radar feature map.
[0072] splicing features ; This indicates a splicing operation.
[0073] Through convolutional layers Predicted offset:
[0074] ;
[0075] in, The offset predicted for the k-th sampling point; , The number of sampling points for deformable convolution kernels ( Convolution kernel, Each sampling point has two offsets (in the x and y directions); This indicates the operation of the convolutional layer.
[0076] Step S13: Use the predicted offset BEV feature map of image Perform deformable convolution operation to obtain the aligned BEV feature map of the image;
[0077]
[0078] In the formula, Spatial location on the BEV feature map of the image; For in position BEV feature map of the image after deformable convolution alignment; ; The sampling point location for the standard convolution kernel; For the offset position BEV feature map extracted from the image; The weights are the weights of the convolution kernel.
[0079] Traditional methods typically rely on fixed calibration parameters, while this method learns dynamic offsets through deformable convolutions, enabling it to adapt to environmental changes (such as sensor deformation caused by temperature variations) and the motion of dynamic objects. The entire feature alignment module can be integrated into the neural network for end-to-end training along with subsequent tasks such as object detection and tracking, making the alignment process more accurate. Deformable convolutions can handle non-rigid transformations (such as local distortions), while traditional methods (such as affine transformations) can only handle global rigid transformations.
[0080] Due to the timing inconsistency between visual sensor and radar data, BEV feature alignment errors occur. Therefore, a transmission delay is introduced to optimize the acquisition of the BEV feature map in the image. The specific optimization formula is as follows:
[0081]
[0082] In the formula, The optimized BEV feature map of the image at position p; for The motion velocity corresponding to the image features at the location; This refers to the transmission delay. The position after time-compensated represents the position of the pixel in the BEV space several seconds ago, used for "backtracking" sampling from the BEV feature map of the image.
[0083] BEV feature maps of images optimized at all locations Constructing the optimized image BEV feature map .
[0084] The optimized image BEV feature map With radar feature map The data is stitched together, and the optimized offset of the deformable convolutional kernel at each location is predicted based on the convolutional network. ;
[0085] Optimized stitching features ;
[0086] Through convolutional layers Predicted offset: ;
[0087] in, The offset predicted for the kth sampling point after optimization;
[0088] Using the optimized predicted offset BEV feature map of image Perform deformable convolution to obtain the aligned and optimized BEV feature map of the image, as follows:
[0089]
[0090] This represents the BEV feature map of the image after alignment optimization at position p; Indicates the offset position The optimized BEV feature map of the image.
[0091] BEV feature map of the image after all positions are aligned and optimized Constructing the aligned and optimized image BEV feature map .
[0092] Alignment-optimized BEV feature map of the image With radar feature map Perform fusion to generate a fused feature map. ;
[0093]
[0094] In the formula, This is a standard convolution operation with a kernel size of 3×3;
[0095] By introducing transmission delay for optimization, precise spatiotemporal synchronization compensation of visual and radar data can be achieved, effectively eliminating feature misalignment caused by differences in sensor sampling frequency and processing delay. At the same time, by "backtracking" sampling through the product of velocity vector and delay, the position of dynamic obstacles in BEV space is accurately reproduced, thereby significantly improving the spatial consistency and semantic integrity of multimodal feature fusion, further reducing perception error, and enhancing the robustness and real-time performance of subsequent trajectory prediction and path planning.
[0096] Step S2: Construct a spatiotemporal graph neural network, using the fused feature map as the input to the spatiotemporal graph neural network, introducing semantic label encoding to define the nodes and edges of the spatiotemporal graph neural network, and then predicting the trajectory of each obstacle through a regression network.
[0097] The core of a spatiotemporal graph neural network (SPNN) is to represent obstacles as nodes in a graph. The position, velocity, and other information of each obstacle constitute the node's features, and the relationships between obstacles are represented by edges. These edges connect different obstacles, representing their interactions. SPNNs learn the spatial and temporal behavior of obstacles through this graph structure, thereby enabling trajectory prediction.
[0098] In step S2, the nodes and edges of the spatiotemporal graph neural network are defined as follows:
[0099] For each obstacle From the fused feature map Extract the feature vector at the corresponding position. ;
[0100] Define the graph nodes and edges of a spatiotemporal graph neural network:
[0101] V i =[ x i , y i , v xi , v yi , S emi i , f i ]
[0102] in, For obstacles node; For obstacles The coordinates of the centroid; For obstacles velocity components; For obstacles Semantic labels are used to distinguish the motion characteristics of different categories of obstacles; traditional methods usually only focus on position and velocity (or even only position). Obstacle category information (such as pedestrians, cars, trucks, bicycles, traffic cones, etc.) is explicitly included in the graph nodes. Different categories of obstacles have distinctly different motion characteristics and potential intentions (pedestrians can suddenly turn / stop, cars usually travel along lanes, and trucks have large turning radii). The interaction relationships between different obstacles are represented by the following formula:
[0103]
[0104] In the formula, For obstacles and obstacles The interaction relationship; For obstacles and obstacles The relative position vector, Δ p ij =[ x j - x i , y j - y i ] ; For obstacles and obstacles The relative velocity vector, Δ v ij =[ v xj - v xi , v yj - v yi ] ; For obstacles semantic tags; For obstacles semantic tags; For encoding functions; For obstacles eigenvectors; For obstacles eigenvectors.
[0105] Encoding function The specific steps for using a multilayer perceptron are as follows:
[0106] Concatenate all input features into a long vector :
[0107] f ij =[Δ p ij ,Δ v ij , S emi i , S emi j , f i , f j ]
[0108] A multilayer perceptron is used to perform a nonlinear transformation on the stitched input features to obtain the obstacle. and obstacles The interaction relationship, i.e., obstacles and obstacles The characteristics of the edges are represented as follows:
[0109] ;
[0110] The MLP consists of a series of fully connected layers, with the following structure:
[0111]
[0112] In the formula, and This is the weight matrix. and For bias terms; This represents the activation function in a spatiotemporal graph neural network.
[0113] This generates encoded edge features. It is used for message passing processes in spatiotemporal graph neural networks.
[0114] Finally, update the node and generate obstacles. final state , For obstacles go through The final node state after propagation in the layered spatiotemporal graph neural network;
[0115] In step S2, the specific steps for predicting the trajectory of each obstacle using a regression network are as follows:
[0116] The regression network predicts the location of obstacles over a future period of time based on the spatiotemporal features learned by the spatiotemporal graph neural network. In this embodiment, the regression network specifically uses the Transformer Decoder architecture.
[0117] Step S21: Based on the semantic labels of obstacles, generate a set of initial behavioral intent vectors according to preset rules. .
[0118] The preset rules determine the behavior patterns of obstacles in advance based on their categories. For vehicles, the rules might generate intentions such as "go straight," "change lanes to the left," or "change lanes to the right." For pedestrians, the rules might generate intentions such as "walk at a constant speed," "accelerate across," or "stay." These intentions are represented by low-dimensional vectors (e.g., 16-64 dimensions) to ensure that different categories of obstacles have diverse intentions that match their motion characteristics, such as preventing trees from generating the intention to "move."
[0119] Step S22: Obstacles output by the spatiotemporal graph neural network Final state Initial behavioral intent vector and obstacles eigenvectors The data is concatenated and input into the Transformer decoder, which then predicts the next step in turn. The trajectory at time Location ;
[0120] Step S23: For each obstacle, the decoder can output... Each trajectory contains 1 independent trajectory. The predicted location sequence for each time step (e.g., 50 location points in 5 seconds, one point every 0.1 seconds) is calculated, and the probability of each trajectory is also calculated.
[0121]
[0122]
[0123] In the formula, It is an obstacle The The trajectory at time The predicted location, ; Indicates obstacles The The trajectory at time The x-coordinate; Indicates obstacles The The trajectory at time The ordinate; It is an obstacle The The probability of a trajectory; ; The number of trajectories; ; The number of time steps; For obstacles The final state is mapped to the first The score of each trajectory; Indicates obstacles The set of all possible trajectories, including Each independent predicted trajectory.
[0124] Step S3: Define environmental semantic labels and couple them with the predicted obstacle trajectories to generate a semantically enhanced risk field.
[0125] In step S3, the specific steps for generating the semantically enhanced risk field are as follows:
[0126] Step S31: Obtain the environmental semantic risk field corresponding to the environmental semantic label, wherein different categories of environmental semantic labels correspond to different risk values.
[0127] Environmental semantic tags include static environmental tags and dynamic environmental tags, and the environmental semantic risk field It is determined by both static and dynamic environment tags.
[0128] Static environment labels include lane lines, curbs, and buildings. Dynamic environment labels include traffic light status and variable lane lines; the environmental semantic risk field is calculated as follows:
[0129]
[0130] In the formula, For position Environmental semantic risk field; This represents the static environmental risk value. For position The static semantic category, i.e., the static environment label; For dynamic risk weighting coefficients; This represents a dynamic environmental risk value. For position The dynamic semantic state, i.e., dynamic environment tags.
[0131] Step S32: For each obstacle, calculate the obstacle risk field based on its predicted trajectory. The obstacle risk field is the sum of the risks generated by all predicted trajectories of that obstacle. The risk value of each trajectory is calculated based on its probability and its spatial distance from the current position. The obstacle risk field is calculated as follows:
[0132]
[0133] In the formula, For position Obstacle risk field; For position To the trajectory point The Euclidean distance; For the first Trajectory in time The time-varying diffusion coefficient.
[0134] Step S33: Generate a spatially related modulation function based on the environmental semantic tags. The modulation function is used to adjust the impact of obstacle risk. It can modify the obstacle risk field according to the semantic category of the location.
[0135] Modulation function The sensitivity of a location to obstacle risk is determined by matching the obstacle category with the environmental semantic category; the sensitivity is divided into multiple levels, each corresponding to a preset modulation coefficient. The modulation function is expressed as follows:
[0136]
[0137]
[0138] In the formula, This is a semantically compatible function that maps environmental semantic labels to obstacle category information to a scalar weight. For position Environmental semantic tags, For the corrected position Obstacle risk field.
[0139] in, The two types of information, "environmental semantics" and "obstacle semantics," are jointly mapped to a scalar weight, which is used to evaluate the obstacle's position within the environment. The risk value at a certain point is amplified or attenuated. If the "pedestrian" is on the "zebra crossing," then priority is given to them. (Magnified twice); If a "car" drives onto the "pedestrian crossing," it constitutes a violation. (Attenuated to one-tenth); if the "vehicle" is in the "lane" and complies with the rules, then (Keep the original value);
[0140] Through semantic compatibility functions The system mandates that risks be amplified or suppressed in both "legal" and "illegal" scenarios to ensure that risk assessments are consistent with traffic regulations. It also considers environmental and obstacle categories to make the risk scenarios more closely resemble real driving intentions. The system suppresses risks associated with irrelevant or impossible high-risk behaviors (such as a truck suddenly crossing a pedestrian crossing) to reduce excessive braking or false alarms.
[0141] Step S34: Overlay the environmental semantic risk field and the corrected obstacle risk field to generate a semantically enhanced risk field.
[0142]
[0143] In the formula, For position The risk field of semantic enhancement; Spatial relevance weights; The total number of obstacles; ;
[0144] Step S4: Construct a semantic road network graph based on the semantically enhanced risk field, generate a global path point sequence, and use the current global path point as a temporary target to call the dynamic step size RRT* algorithm to avoid obstacles.
[0145] In step S4, the specific steps to avoid obstacles are as follows:
[0146] Step S41: Based on the semantically enhanced risk field, construct a semantic road network graph where nodes represent intersection centers and lane sampling points, and edges carry road type, speed limit, and driving direction constraints. In semantic road network graph The A* algorithm is used, with a cost function that includes road semantic cost and risk field risk value. Perform a search to generate a global pathpoint sequence. .
[0147] The semantic road network graph is constructed by extracting topological features from a high-precision map, using intersection centers as nodes, sampling lane lines in segments, connecting lane sampling points, injecting semantically enhanced risk field attributes, and finally outputting the semantic road network graph. .
[0148] In semantic road network graph superior, Nodes representing semantic road network graphs Representing the edges of a semantic road network graph, the cost function of the A* algorithm. Combining road semantic cost and semantic enhancement risk field, the cost function is calculated as follows:
[0149]
[0150] In the formula, For nodes and The Euclidean distance between them represents the path length; Distance weights; For road type cost, based on edge The type of road (such as highway, main road, auxiliary road, etc.) is mapped to its value. Weights for road types; Assign a cost to the edge based on the speed limit information, and then assign the cost to the road speed limit. Speed limiting weight; For position In the semantically enhanced risk field, the risk traversed by the path is indicated; Representing an edge The x-coordinate of the midpoint Representing an edge The ordinate of the midpoint; Risk weights.
[0151] The A* algorithm prioritizes expanding the node with the lowest total cost. If the current node is the target node, it backtracks the path to generate the optimal path from the starting point to the target.
[0152] Step S42: The local layer uses the current global path point For the temporary target, the Dynamic Step Size (RRT) algorithm is invoked, where the dynamic step size is based on the semantically enhanced risk field at the current position. Adaptive computation, and real-time querying of the semantically enhanced risk field as nodes expand. Risk field threshold As an extended hard constraint, it also determines whether the replanning trigger condition is met. If it is met, replanning is executed.
[0153] In step S42, the re-planning of triggering conditions includes:
[0154] The cumulative risk value of a local path exceeds the preset cumulative risk threshold;
[0155] Local path costs are higher than the historical average.
[0156] Semantic conflict, i.e., semantically compatible functions Less than the semantically compatible function threshold.
[0157] Dynamic step size Risk field with semantic enhancement based on current location Adaptive calculation is performed, and the calculation method is as follows:
[0158]
[0159] In the formula, For risk field threshold; This is the step size sensitivity coefficient, used to adjust the changes in the high-risk region of the step size response.
[0160] In this embodiment, the dynamic step size Risk field value with semantic enhancement of current location The exponential inverse relationship allows for smaller step sizes in high-risk areas, enabling more refined searches, while larger step sizes in low-risk areas accelerate searches and optimize obstacle avoidance.
[0161] node For an application to be accepted, the risk value of that node must not exceed the risk threshold.
[0162] That is, only when the risk value of the new node does not exceed the risk threshold. The node will only be accepted if it meets the criteria; otherwise, it will be discarded. By introducing hard risk constraints during path expansion, we ensure that each step is within an acceptable risk range, prevent the path from passing through high-risk areas, and enhance the safety of path planning.
[0163] Step S43: When the number of consecutive replanning failures reaches the preset limit. If the current path is not found, return to the global layer and re-execute step S41, freezing the current path region in the semantic road network graph; otherwise, use the current position as the new starting point and return to the previous global path point. Continue with step S42;
[0164] When a local path fails to meet the requirements, the system can trigger replanning and automatically fall back to the global layer to regenerate the path. By judging the number of consecutive replanning failures, the system avoids over-reliance on local planning, thus enhancing its fault tolerance. By freezing the current path region, the system avoids repeatedly performing path planning in known infeasible areas, improving computational efficiency and reducing redundant computational overhead.
[0165] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A target perception and avoidance decision-making method in visual navigation of intelligent equipment, characterized in that, Includes the following steps: Step S1: Use a computer vision sensor to acquire image feature maps of obstacles, use radar to collect 3D point clouds of obstacles in real time, perform feature alignment on the image feature maps in BEV space based on deformable convolution kernels, and introduce transmission delay for optimization. Then, fuse the aligned image feature maps with the radar feature maps to generate a fused feature map. Step S2: Construct a spatiotemporal graph neural network, using the fused feature map as the input to the spatiotemporal graph neural network, introducing semantic label encoding to define the nodes and edges of the spatiotemporal graph neural network, and then predicting the trajectory of each obstacle through a regression network; Step S3: Define environmental semantic labels and couple them with the predicted obstacle trajectories to generate a semantically enhanced risk field; Step S4: Construct a semantic road network map based on the semantically enhanced risk field, generate a global path point sequence, and use the current global path point as a temporary target to call the dynamic step size RRT* algorithm to avoid obstacles; In step S1, the specific steps for feature alignment of the image feature map in the BEV space based on deformable convolution kernels are as follows: Step S11: Divide the obstacle into a two-dimensional grid on the horizontal plane, project the 3D point cloud onto the two-dimensional grid, and generate a radar feature map. The image feature map is projected onto the BEV space to generate the image BEV feature map. ; Step S12: Transfer radar feature map and image BEV feature map The concatenation is used as input to a convolutional network, which then predicts the offset of the deformable convolutional kernel at each location. ; Step S13: Use the predicted offset BEV feature map of image Perform deformable convolution to obtain the aligned BEV feature map of the image, as shown below: In the formula, Spatial location on the BEV feature map of the image; For in position BEV feature map of the image after deformable convolution alignment; ; The sampling point location for the standard convolution kernel; For the offset position BEV feature map extracted from the image; These are the weights of the convolution kernel; In step S3, the specific steps for generating the semantically enhanced risk field are as follows: Step S31: Obtain the environmental semantic risk field corresponding to the environmental semantic label. Different categories of environmental semantic labels correspond to different risk values. Step S32: For each obstacle, calculate the obstacle risk field based on its predicted trajectory. The obstacle risk field is the sum of the risks generated by all predicted trajectories of the obstacle. Step S33: Generate a spatially related modulation function based on the environmental semantic tags. It is used to adjust the impact of the obstacle risk field. The modulation function can modify the obstacle risk field according to the semantic category of the location. Step S34: Overlay the environmental semantic risk field and the corrected obstacle risk field to generate a semantically enhanced risk field.
2. The target perception and avoidance decision-making method in visual navigation of intelligent equipment according to claim 1, characterized in that: In step S11, the obstacle is divided into a two-dimensional grid on the horizontal plane. Each two-dimensional grid includes the maximum height, height distribution variance, point cloud density, reflection intensity entropy, temporal stability index, multi-frame motion vector variance, semantic confidence probability, and height gradient.
3. The target perception and avoidance decision-making method in intelligent equipment visual navigation according to claim 2, characterized in that: In step S1, a transmission delay is introduced for optimization, and the aligned image feature map and radar feature map are fused to generate a fused feature map, specifically including: The optimized BEV feature map of the image is obtained by optimizing the transmission delay. The optimized BEV feature map of the image is concatenated with the radar feature map, and the optimized offset of the deformable convolution kernel at each location is predicted based on the convolutional network. The optimized offset is used to perform deformable convolution operation on the BEV feature map of the image to obtain the aligned optimized BEV feature map of the image. The aligned and optimized BEV feature map of the image is fused with the radar feature map to generate a fused feature map.
4. The target perception and avoidance decision-making method in visual navigation of intelligent equipment according to claim 1, characterized in that: In step S2, the nodes and edges of the spatiotemporal graph neural network are defined as follows: For each obstacle Extract the feature vectors of the corresponding positions from the fused feature map. ; Define the graph nodes and edges of a spatiotemporal graph neural network: in, For obstacles node; For obstacles The coordinates of the centroid; For obstacles velocity components; For obstacles semantic tags; The interaction between different obstacles is represented by the following formula: In the formula, For obstacles and obstacles The interaction relationship; For obstacles and obstacles The relative position vector, ; For obstacles and obstacles The relative velocity vector, ; For obstacles semantic tags; For encoding functions; For obstacles eigenvectors; Finally, update the node and generate obstacles. final state .
5. The target perception and avoidance decision-making method in intelligent equipment visual navigation according to claim 4, characterized in that: In step S2, the specific steps for predicting the trajectory of each obstacle using a regression network are as follows: Step S21: Based on the semantic labels of obstacles, generate a set of initial behavioral intent vectors according to preset rules. ; Step S22: Obstacles output by the spatiotemporal graph neural network Final state Initial behavioral intent vector and obstacles eigenvectors The data is concatenated and input into the decoder, which then predicts the next data step by step. The trajectory at time Location ; Step S23: For each obstacle, the decoder can output... Each trajectory contains 1 independent trajectory. The predicted location sequence at each time step is calculated, and the probability of each trajectory is also calculated.
6. The target perception and avoidance decision-making method in visual navigation of intelligent equipment according to claim 5, characterized in that: The environmental semantic label includes static environmental labels and dynamic environmental labels, and the environmental semantic risk field is jointly determined by the static environmental labels and the dynamic environmental labels.
7. The target perception and avoidance decision-making method in intelligent equipment visual navigation according to claim 6, characterized in that: In step S4, the specific steps for avoiding obstacles are as follows: Step S41: Based on the semantically enhanced risk field, construct a semantic road network graph where nodes represent intersection centers and lane sampling points, and edges carry road type, speed limit, and driving direction constraints. In semantic road network graph The A* algorithm is used, with a cost function that includes road semantic cost and risk field risk value. Perform a search to generate a global pathpoint sequence. ; Step S42: Using the current global path point For the temporary target, the Dynamic Step Size (RRT) algorithm is invoked, where the dynamic step size is based on the semantically enhanced risk field at the current position. Adaptive computation, and real-time querying of the semantically enhanced risk field as nodes expand. Risk field threshold As an extended hard constraint, it is also determined whether the replanning trigger condition is met. If it is met, replanning is executed. Step S43: When the number of consecutive replanning failures reaches the preset limit. If the current path region is frozen in the semantic road network graph, then repeat step S41; otherwise, return to the previous global path point with the current position as the new starting point. Continue with step S42.
8. The target perception and avoidance decision-making method in intelligent equipment visual navigation according to claim 7, characterized in that: In step S42, the replanning trigger conditions include: the cumulative risk value of the local path exceeds a preset cumulative risk threshold; the cost of the local path is higher than the historical average; and the semantic compatibility function... Less than the semantic compatibility function threshold.
Citation Information
Patent Citations
Dynamic obstacle environment navigation method and device based on visual semantic information
CN111367318A
Tracking method of video moving target
CN116993774A