Autonomous navigation method for local attention based on multi-channel enhancement and high efficiency
Through multi-channel enhancement and efficient local attention methods, the problems of multi-scale target perception and dynamic obstacle response lag in complex dynamic scenarios are solved, and the robustness and stability of efficient perception of dynamic obstacles and navigation path planning are achieved.
Patent Information
- Application Number
- CN202510451802.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing visual autonomous navigation methods have problems such as insufficient perception of multi-scale targets, inaccurate modeling of local spatial relationships, omission of global context information, and lagging responses of dynamic obstacles in complex dynamic scenarios.
Multi-channel enhancement and efficient local attention methods are adopted to generate significant anchors through multi-scale feature fusion module, direction-sensitive local attention module, context broadcast method and deformable convolution, and combined with reinforcement learning decision-making module, efficient perception and decision-making of dynamic obstacles are achieved.
It improves the representation ability of visual features, enhances the spatial rationality and real-time environmental adaptability of navigation path planning, and ensures the robustness and stability of the navigation system.
Smart Images

Figure CN120368976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving and robot navigation, and specifically provides an autonomous navigation method based on multi-channel enhancement and efficient local attention. Background Art
[0002] Currently, the existing vision-based autonomous navigation methods generally have problems with insufficient adaptability of the model architecture to dynamic scenarios. Although the standard Vision Transformer (ViT) model shows advantages in global feature extraction, its insufficient channel sensitivity causes the edge features of dynamic obstacles to be easily interfered by environmental noise. Especially in scenarios with sudden light changes or occlusions, the positioning accuracy of the target boundary significantly decreases. In addition, traditional vision models are limited by the fixed receptive field and single-scale feature extraction mechanism, and have insufficient perception ability for targets of different scales in dense obstacle scenarios. Super-large targets are easily submerged by local features, while super-small targets are missed due to insufficient feature resolution, seriously affecting the reliability of navigation path planning.
[0003] Existing solutions usually adopt a global attention mechanism or a conventional feature pyramid structure to make up for the multi-scale perception defect. However, the quadratic computational complexity introduced by such methods is quadratic with the length of the input sequence, making it difficult to meet the low-latency requirements of real-time navigation systems. At the same time, the context modeling method based on a fixed pooling strategy or static anchor generation has limited ability to model the spatio-temporal correlation of dynamic obstacles, resulting in a lag in the response of the path planning strategy to sudden obstacles, manifested as conservatism (such as redundant detours) or radicalism (such as sudden stops and jitters) of the obstacle avoidance path.
[0004] The pure convolutional network, fixed pooling strategy, or conventional attention mechanism relied on by traditional methods has an inherent bottleneck in the collaborative optimization of efficiency and accuracy. There is an urgent need for an autonomous navigation technology that integrates lightweight multi-scale perception, dynamic anchor association, and anti-interference decision-making to break through the limitations of the existing technology in adapting to complex scenarios. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides an autonomous navigation method based on multi-channel enhancement and efficient local attention, which solves the problems of insufficient multi-scale target perception, inaccurate local spatial relationship modeling, omission of global context information, and lag in response to dynamic obstacles existing in the existing vision navigation methods in complex dynamic scenarios.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: An autonomous navigation method based on multi-channel enhancement and efficient local attention, comprising the following steps:
[0007] S1. Input the RGB visual image and fuse the position coordinates of the unmanned vehicle, convert it into a serialized feature matrix through block embedding, input the feature matrix into the efficient multi-scale attention fusion module, and generate multi-scale features through channel dimension grouping and cross-dimensional interaction;
[0008] S2. Input the multi-scale features into the Vision Transformer, and embed a lightweight direction-sensitive local attention module before each Transformer block, generate position-encoded features along the height and width directions respectively through direction-sensitive 1D horizontal convolution and 1D vertical convolution, and capture local context relationships;
[0009] S3. Inject uniform dense attention into each layer of the Vision Transformer through the context broadcasting method to supplement the global information missed by sparse attention;
[0010] S4. Input the feature matrix into the improved context anchor attention fusion module, generate significant anchors by replacing average pooling with deformable convolution, suppress noise and retain global context;
[0011] S5. Input the significant anchors into the reinforcement learning decision module, optimize the action output based on the SAC algorithm, and output navigation action instructions.
[0012] Preferably, the implementation of the efficient multi-scale attention fusion module in step S1 includes:
[0013] Divide the input features into multiple subgroups along the channel dimension, and reshape some channels into the batch dimension to reduce the computational amount;
[0014] Execute 1×1 convolution and 3×3 convolution in parallel to capture global and local features respectively;
[0015] Fuse the global and local features through cross-dimensional matrix multiplication, and the formula is:
[0016]
[0017] Among them, represents element-wise multiplication, Y is the output feature matrix, W1 and W2 are learnable weight parameters, F GAP is the global feature, F 3×3 is the local feature.
[0018] Preferably, the implementation of the lightweight direction-sensitive local attention module in step S2 includes:
[0019] Apply a 1D horizontal convolution kernel along the height direction to generate vertical position encoding;
[0020] Apply a 1D vertical convolution kernel along the width direction to generate horizontal position encoding;
[0021] Fusing bidirectional encodings through depthwise separable convolutions, the formula is:
[0022] P cross = DWConv([P h ; P v )
[0023] where DWConv(·) represents the depthwise separable convolution operation, and P h , P v are the horizontal encoding feature and the vertical encoding feature respectively.
[0024] Preferably, the implementation of the context broadcasting method in step S3 includes:
[0025] Defining a uniform attention matrix, whose element values are covering all spatial positions; where H and W are the height and width of the input feature map respectively;
[0026] Adding the uniform attention to the sparse attention output of the Vision Transformer, the formula is:
[0027] Attention final = A uniform ·V + Attention sparse
[0028] where Attention final is the fused attention matrix, A uniform is the uniform weight matrix, V is the value matrix, and Attention sparse is the original sparse attention matrix generated by the self-attention mechanism in the Vision Transformer.
[0029] Preferably, the implementation of the deformable convolution in step S4 includes:
[0030] Generating an initial offset through a 3×3 convolution to dynamically adjust the pooling region;
[0031] The pooling formula is:
[0032] F pool = ∑ k w k ·X(p + p k + Δp k )
[0033] where p is the coordinate of the current feature position; p k are the coordinates of k preset sampling points; Δp k is the learnable offset representing the dynamic position adjustment of the k-th sampling point; w kis the weight parameter for the k-th sampling point;
[0034] The generation of significant anchor points includes:
[0035] Dynamically adjusting the anchor point distribution density based on the spatial gradient of the high-level features, with the formula:
[0036]
[0037] where H and W are the height and width of the high-level feature map respectively, is the spatial gradient of the high-level feature map, and ρ is the dynamically calculated anchor point density;
[0038] Applying a 1×1 convolution to the pooled features for channel dimension recalibration to generate a set of significant anchor points.
[0039] Preferably, the design of the reward function of the SAC algorithm in step S5 includes:
[0040] Anchor point significance reward term: Calculate the minimum Euclidean distance between the current position and the set of anchor points;
[0041] Policy entropy reward term: Maximize the policy entropy to enhance the exploration ability;
[0042] The formula of the reward function is:
[0043]
[0044] where p t is the current position coordinate of the unmanned vehicle, is the set of significant anchor points, represents the minimum Euclidean distance function, is the policy entropy, and α and β are normalization weight coefficients.
[0045] Preferably, the design of the reward function further includes:
[0046] Path tracking reward term: Calculate the distance between the current position and the target path point;
[0047] The formula of the reward function is:
[0048]
[0049] where p target is the target path point, and γ is the normalization weight coefficient.
[0050] Preferably, the constraint condition of the action output is:
[0051] The action output by the policy network follows the maximum entropy optimization objective, with the formula:
[0052]
[0053] wherein, a t represents the action sampled at time t, ~ represents sampling from a probability distribution, and π(·|s t ) represents the conditional probability distribution of the action under state s t .
[0054] The present invention also provides an autonomous navigation system based on multi-channel enhancement and efficient local attention, including:
[0055] A multi-modal sensor module for collecting RGB images and position coordinates;
[0056] An EMA fusion module for multi-scale feature enhancement and cross-dimensional interaction;
[0057] A ViT-ELA module integrating ELA-T local attention and CB dense attention injection;
[0058] A CAA module for generating significant anchor points through deformable convolution;
[0059] An SAC reinforcement learning decision module for outputting navigation actions based on anchor point saliency.
[0060] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method as described above is implemented.
[0061] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as described above is implemented.
[0062] The present invention provides an autonomous navigation method based on multi-channel enhancement and efficient local attention.
[0063] It has the following beneficial effects:
[0064] 1. Through the channel grouping and cross-dimensional interaction mechanism of the efficient multi-scale attention fusion module (EMA) of the present invention, the adaptive fusion of global context information and local detail features is realized. By parallelly extracting global statistical features and local spatial features, the perception limitations of traditional single-scale convolution for small targets or distant obstacles in complex scenarios are overcome, and the representation ability of visual features is significantly improved, providing a robust environmental perception input for subsequent navigation decisions.
[0065] 2. The direction-sensitive local attention module (ELA-T) of the present invention generates direction-specific position-encoded features by independently designing horizontal and vertical convolutional kernels, effectively capturing the local context associations in the row and column dimensions of the image. Compared with the orientation information confusion problem of traditional two-dimensional convolution, the module can accurately model the relative orientation relationship between the unmanned vehicle and the obstacle, improving the spatial rationality of the navigation path planning.
[0066] 3. The present invention injects a uniform attention matrix into the vision Transformer through the context broadcasting method (CB), forcing the model to focus on the equal-weight associations of all position pairs. This mechanism makes up for the deficiency in modeling long-distance dependencies caused by local window calculations or sampling strategies in sparse attention, ensuring the global visibility of key regions (such as the movement trajectories of dynamic obstacles) in long-range path planning.
[0067] 4. The improved context anchor attention fusion module (CAA) of the present invention dynamically adjusts the feature sampling area through deformable convolution and combines the gradient magnitude to drive the anchor density distribution, enabling the significant anchors to adaptively focus on texture-rich or moving regions. Compared with traditional fixed pooling operations, this design significantly enhances the response sensitivity to dynamic obstacles and the noise suppression ability, improving the real-time environment adaptation ability of the navigation system.
[0068] 5. The action optimization strategy based on the maximum entropy reinforcement learning framework of the present invention realizes the dynamic balance of obstacle avoidance, goal orientation, and exploration capabilities through a multi-objective function that combines anchor saliency rewards, path tracking rewards, and policy entropy rewards. At the same time, the action space constraint mechanism ensures that the output actions strictly meet the mechanical dynamics limit requirements, guaranteeing the physical feasibility and system stability of the navigation instructions at the algorithm level. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is one of the schematic diagrams of the method flow of the present invention;
[0070] Figure 2 is the second schematic diagram of the method flow of the present invention;
[0071] Figure 3 is the schematic diagram of the EMA module structure of the present invention;
[0072] Figure 4 is the schematic diagram of the 1D direction convolution and cross-direction interaction process of the ELA-T module of the present invention;
[0073] Figure 5 is the schematic diagram of the offset generation and anchor density calculation of the deformable CAA module of the present invention;
[0074] Figure 6It is a comparison chart of the reinforcement learning training effects of adding EMA, ELA-T, CB, and CAA modules in the present invention;
[0075] Figure 7 It is a comparison chart of the training error effects of the EMA, ELA-T, CB, and CAA modules in the present invention;
[0076] Figure 8 It is a schematic diagram of the system structure of the present invention;
[0077] Figure 9 It is a schematic diagram of the computer device structure of the present invention.
[0078] Among them, 100, multi-modal sensor module; 200, EMA fusion module; 300, ViT-ELA module; 400, CAA module; 500, SAC reinforcement learning decision-making module; 40, computer device; 41, processor; 42, memory; 43, storage medium. Specific embodiments
[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0080] Please refer to the attached Figure 1 - attached Figure 7 , the present invention provides an autonomous navigation method based on multi-channel enhancement and efficient local attention. This method realizes robust navigation in complex dynamic scenarios through the collaborative design of multi-scale feature fusion, direction-sensitive local attention, dense global information compensation, dynamic anchor detection, and reinforcement learning decision-making.
[0081] As Figure 1 shown, the autonomous navigation method based on multi-channel enhancement and efficient local attention may include the following steps:
[0082] S1. Input the RGB visual image and fuse the position coordinates of the unmanned vehicle, convert it into a serialized feature matrix through block embedding, input the feature matrix into the efficient multi-scale attention fusion module, and generate multi-scale features through channel dimension grouping and cross-dimensional interaction;
[0083] S2. Input the multi-scale features into the vision Transformer, and embed a lightweight direction-sensitive local attention module before each Transformer block. Generate position-encoded features along the height and width directions respectively through direction-sensitive 1D horizontal convolution and 1D vertical convolution to capture local context relationships;
[0084] S3. Inject uniform dense attention into each layer of the vision Transformer through the context broadcasting method to supplement the global information missed by the sparse attention.
[0085] S4. Input the feature matrix into the improved context anchor attention fusion module, generate significant anchors by replacing average pooling with deformable convolution, suppress noise and retain the global context.
[0086] S5. Input the significant anchors into the reinforcement learning decision module, optimize the action output based on the SAC algorithm, and output the navigation action instructions.
[0087] The following is a detailed description of each step in the method of the present invention, comprehensively elaborating on the specific implementation principles, technical details and processes for each step.
[0088] For step S1, in this embodiment, the multi-scale feature fusion process described in step S1 is implemented through the following scheme:
[0089] Input processing and position fusion:
[0090] After the input RGB visual image is preprocessed, it is spatially aligned and fused with the current position coordinates of the unmanned vehicle. Specifically, the spatial position information of each pixel point in the RGB image is jointly encoded with the global coordinates of the unmanned vehicle to generate enhanced input features with position perception. The position coordinates are mapped to the image plane through normalization processing to ensure scale consistency with the visual features.
[0091] Chunk embedding and serialization processing:
[0092] To adapt to the serialization input requirements of the vision Transformer, the chunk embedding method is used to convert the image into a serialized feature matrix. Specifically, the enhanced input features are divided into multiple non-overlapping image chunks, and each image chunk is mapped to a feature vector of a fixed dimension through a linear projection layer. Preferably, the linear projection layer is implemented by a fully connected layer or a 1×1 convolution, which is used to convert the pixel values of each image chunk into a unified feature space. The dimension of the feature matrix after chunk embedding is where N is the number of image chunks and C is the number of channels.
[0093] Efficient multi-scale attention fusion module (EMA):
[0094] The EMA module realizes multi-scale feature extraction through the channel dimension grouping and cross-dimension interaction mechanism. First, the input feature matrix X is divided into G subgroups along the channel dimension, and the number of channels in each subgroup is Preferably, the features of some subgroups are reshaped into the batch dimension to reduce the computational complexity and retain the spatial correlation.
[0095] Furthermore, global and local feature extraction operations are performed in parallel for each subgroup:
[0096] Global feature branch: Perform global average pooling (GAP) on the input features to compress the spatial dimension and capture global context information to generate a global feature vector Where B is the batch size;
[0097] Local feature branch: Use 3×3 convolution kernel to extract local spatial features and generate local feature vectors Preferably, the convolution kernel is a depthwise separable convolution to reduce the number of parameters.
[0098] Cross-dimensional interaction and feature fusion:
[0099] The global features and local features are spliced along the channel dimension and dynamically weighted fused through the learnable weight parameters W1 and W2. Specifically, the Sigmoid function is used to perform nonlinear activation on the weighted features to generate a multi-scale attention weight matrix, which is expressed as:
[0100]
[0101] Among them, W1 and W2 are learnable weight parameters, Represents an element-by-element multiplication operation. The weight matrix is used to recalibrate the input features and highlight the salient areas at multiple scales.
[0102] The fused feature matrix Y is used as a multi-scale representation output and input into the subsequent visual Transformer module. Preferably, the number of channel groupings G of the module can be dynamically adjusted according to the input resolution to balance computational efficiency and feature diversity.
[0103] Through the above design, step S1 realizes the hierarchical extraction and adaptive fusion of multi-scale features, providing a robust feature representation that takes into account both global and local information for subsequent modules.
[0104] For step S2, in this embodiment, the lightweight direction-sensitive local attention module (ELA-T module) of step S2 enhances the modeling ability of the visual transformer for local spatial relationships through direction-aware convolution operations and cross-directional feature interaction mechanisms. The specific implementation method is as follows:
[0105] Direction-sensitive encoding process:
[0106] Input multi-scale feature matrix Perform independent convolution processing along the height direction (vertical dimension) and width direction (horizontal dimension) of the image respectively. Specifically, use 1D convolution kernels to extract direction-sensitive position encoding features:
[0107] Horizontal direction encoding: Apply a 1D horizontal convolution kernel Slide along the width direction, perform convolution operations on each row of pixels, and generate horizontal position encoding features Preferably, the size k of the convolution kernel can be adapted to the input resolution to capture horizontal spatial dependencies at different scales.
[0108] Vertical direction encoding: Apply a 1D vertical convolution kernel Slide along the height direction, perform convolution operations on each column of pixels, and generate vertical position encoding features The parameters of the vertical convolution kernel are independently initialized from those of the horizontal convolution kernel to ensure direction specificity.
[0109] Cross-direction feature interaction:
[0110] Concatenate the horizontal encoding feature P h and the vertical encoding feature P v along the channel dimension to obtain Realize cross-direction feature fusion through Depthwise Separable Convolution (DWConv). Specifically, depthwise separable convolution includes two-stage operations: channel-wise convolution and pointwise convolution:
[0111] Channel-wise convolution: Perform spatial convolution on the concatenated feature map by grouping channels, and each group of channels is independently processed to preserve direction information;
[0112] Pointwise convolution: Use a 1×1 convolution kernel to adjust the number of channels and generate cross-direction interaction features Its mathematical expression is:
[0113] P cross = DWConv([P h ; P v )
[0114] Attention weight generation:
[0115] Normalize and non-linearly activate the cross-direction interaction feature P cross to generate a direction-sensitive attention weight matrix. Specifically, first apply Group Normalization (GN) to eliminate the distribution difference between channels, and then map the feature values to the interval [0, 1] through the Sigmoid function to obtain the attention weight matrix The formula is:
[0116]
[0117] Feature recalibration and output:
[0118] Element-wise multiply the attention weight matrix with the input feature matrix to achieve direction-sensitive feature enhancement. Specifically, the final output feature is generated through the following formula:
[0119] X ELA = A ELA ⊙ X
[0120] where ⊙ represents the element-wise multiplication operation, and the output feature serves as the input to the vision Transformer. Preferably, the module is embedded before each self-attention layer of the vision Transformer to inject local position prior information before the global attention calculation.
[0121] Through the above solution, step S2 provides a local feature representation with strong position awareness for the subsequent global attention calculation, thereby enhancing the robustness of the navigation system to the direction change of obstacles in complex dynamic scenarios.
[0122] For step S3, in this embodiment, the context broadcasting method (CB method) of step S3 realizes effective compensation for global context information by fusing the uniform attention matrix with the sparse attention of the vision Transformer. The specific implementation process is as follows:
[0123] Construction of the uniform attention matrix:
[0124] Based on the block sequence length N = H × W of the input feature map, construct the uniform weight matrix The value of each element in the matrix is a constant value The mathematical expression is:
[0125]
[0126] where H and W are the height and width of the input feature map respectively, and N represents the sequence length after partitioning. The physical meaning of the uniform matrix is to assign equal-weight associations to all position pairs, ensuring unbiased coverage of the global spatial relationship.
[0127] Attention fusion mechanism:
[0128] Weightedly fuse the sparse attention output of the vision Transformer with the uniform attention. Specifically, the final enhanced attention output is generated through the following formula:
[0129] Attentionfinal = A uniform ·V + Attention sparse
[0130] Among them, is the value matrix in the Vision Transformer, and d is the feature dimension. Preferably, the sparse attention is generated by the standard multi-head self-attention mechanism, and its calculation process follows the dot product form of Query, Key, and Value.
[0131] Dimension consistency constraint:
[0132] Uniform matrix A uniform and the sparse attention matrix Attention sparse should satisfy the additivity condition in dimension. Specifically, the sparse attention matrix adjusts the output dimension through a linear projection layer to ensure consistency with the dimension of A uniform ·V. Preferably, the weight matrix of the linear projection layer can be dynamically optimized during the training process to adapt to different levels of feature distributions.
[0133] Global information compensation mechanism:
[0134] The introduction of the uniform attention matrix aims to make up for the deficiency of the traditional sparse attention mechanism in modeling long-range dependencies. Since sparse attention is usually calculated based on local windows or sampling strategies, it may ignore the potential correlations between distant positions. By superimposing uniform weights, each position can obtain globally averaged context information, thereby suppressing the feature bias caused by attention sparsification.
[0135] Output feature transfer:
[0136] The fused attention matrix Attention final is used as the output of the current Transformer layer and is passed to the next network layer for further processing. Preferably, the context broadcasting method can be embedded in each layer of the Vision Transformer to form a hierarchical global-local attention complementary structure.
[0137] Through the above technical solutions, the introduction of the prior uniform distribution constraint enhances the model's robust perception ability of the spatial distribution of dynamic obstacles in complex scenarios while retaining the efficient calculation advantages of sparse attention.
[0138] For step S4, in this embodiment, the improved context anchor point attention fusion module (CAA module) of step S4 replaces the traditional average pooling operation with deformable convolution, dynamically generates significant anchor points, and suppresses noise interference. The specific implementation method is as follows:
[0139] Generation of the offset of the deformable convolution:
[0140] Based on the input high-level feature map Generate the initial offset matrix through a 3×3 convolutional kernel Where K is the preset number of sampling points. The physical meaning of the initial offset is to provide the basic offset for K sampling points at each spatial position (i, j), and each sampling point corresponds to the offset components in the horizontal and vertical directions. Preferably, the activation function of the 3×3 convolutional layer is a linear function to preserve the continuous spatial distribution characteristics of the offset.
[0141] Dynamic offset adjustment:
[0142] The initial offset Δp0 is dynamically optimized through learnable parameters to generate the final offset Specifically, use another 3×3 convolutional layer to extract features from the high-level feature map F high to obtain the residual offset correction term, which is superimposed on the initial offset:
[0143] Δp = Conv 3×3 (F high ) + Δp0
[0144] The residual correction term is optimized through gradient backpropagation, enabling the offset to adapt to the local structural features of the input content.
[0145] Deformable pooling operation:
[0146] Based on the dynamic offset Δp, perform irregular sampling and weighted pooling on the input feature map. For the current position p = (i, j), its pooled output feature F pool is calculated as follows:
[0147] F pool = ∑ k w k ·X(p + p k + Δp k )
[0148] Where:
[0149] p k = (x k , y k ) is the coordinate of the kth preset regular sampling point. For example, in a 3×3 grid, p k can cover the nine-grid positions centered on p;
[0150] Δp k = (Δx k , Δy k ) is the dynamic offset of the kth sampling point, which is parsed from the channel data at the corresponding position in Δp;
[0151] X(·) represents bilinear interpolation sampling performed on the input feature map X to ensure effective feature extraction at non-integer coordinate positions;
[0152] w k is the convolution kernel weight parameter for the k-th sampling point, which is optimized through backpropagation.
[0153] Dynamic adjustment of anchor density:
[0154] Calculate the anchor distribution density ρ based on the spatial gradient magnitude of the high-level feature map, driving the anchors to gather in the regions with significant edges. Specifically, the calculation formula for dynamically adjusting the anchor distribution density based on the spatial gradient of the high-level feature is:
[0155]
[0156] The density ρ serves as a guiding parameter for subsequent anchor generation, ensuring an increased anchor distribution density in regions with complex textures or dynamic obstacles.
[0157] Channel recalibration and anchor generation:
[0158] Perform a 1×1 convolution operation on the feature map F output by the deformable pooling pool to generate a set of significant anchors
[0159] Preferably, a Sigmoid activation function is connected after the 1×1 convolution layer to map the anchor confidence to the interval [0,1], which is used to indicate the significance degree of each position.
[0160] Through the above technical solutions, the collaborative design of deformable convolution and gradient density guidance can, while retaining the global context information, achieve accurate localization of significant regions in dynamic scenarios, providing a robust environmental perception input for the subsequent reinforcement learning decision-making module.
[0161] For step S5, in this embodiment, the reinforcement learning decision-making module based on the maximum entropy optimization framework realizes the robust output of navigation actions through the collaborative design of a multi-objective reward function and action space constraints. The specific implementation method is as follows:
[0162] State and action space modeling:
[0163] The state space S of the reinforcement learning decision-making module is defined as a vector set containing the positions of significant anchors, the current motion state of the unmanned vehicle, and the coordinates of the target path points. Specifically, the state vector s t ∈S has a dimension d s which is jointly determined by the number of anchors K, the dimension of speed information, and the dimension of target point coordinates, and the mathematical expression is d s= 2K + 4. Preferably, the anchor position information is extracted from the set of significant anchors output in step S4 to ensure end-to-end information transfer in the perception and decision-making module.
[0164] The action space A is defined as the combination of the linear velocity and angular velocity of the unmanned vehicle. Preferably, the action space is represented by a continuous value range to support fine-grained motion control.
[0165] Multi-objective reward function design:
[0166] The reward function r t is composed of the following three weighted terms:
[0167] 1. Anchor saliency reward term: Calculate the minimum Euclidean distance between the current position of the unmanned vehicle pt = (x t , y t ) and the nearest anchor:
[0168]
[0169] This reward term drives the unmanned vehicle to approach the salient area and enhances the obstacle avoidance ability for dynamic obstacles.
[0170] 2. Path tracking reward term: Calculate the Euclidean distance between the current position pt and the target path point p target :
[0171]
[0172] This reward term guides the unmanned vehicle to drive along the preset path to ensure the achievement of the mission goal.
[0173] 3. Policy entropy reward term: Calculate the entropy value of the policy network π(·|s t ):
[0174]
[0175] This reward term encourages exploration by maximizing the policy entropy to avoid premature convergence of the local optimal policy.
[0176] The comprehensive reward function is defined as the linear combination of the above three terms:
[0177]
[0178] where the weight coefficients α, β, γ satisfy the normalization constraint α + β + γ = 1.
[0179] Maximum entropy policy optimization:
[0180] The training objective of the policy network is to maximize the cumulative discounted reward with an entropy regularization term, and its optimization objective function is:
[0181]
[0182] where γ is the discount factor, ρ π represents the state-action distribution induced by policy π. Preferably, the policy network adopts an Actor-Critic architecture, and the value estimation bias is alleviated through a dual Q-network.
[0183] Action output and mechanical constraints:
[0184] The action a t ~ π(·|s t ) output by the policy network needs to satisfy the dynamic constraints of the unmanned vehicle mechanical system. Specifically, the action value range is constrained to where a min and a max are the preset physical limits of the linear velocity and angular velocity. Preferably, the constraint is implemented through the activation function (such as the Tanh function) of the output layer of the policy network to ensure that the action sampling meets the actual execution requirements.
[0185] Policy sampling and training process:
[0186] At each decision-making moment t, according to the current state s t sample the action a t from the policy distribution, and the mathematical expression is:
[0187] a t ~ π(·|s t )
[0188] where the symbol ~ represents random sampling from the probability distribution. The sampling process introduces policy randomness to support online exploration and interaction with the environment. In the training stage, the Soft Actor-Critic (SAC) algorithm is adopted to gradually improve the expected cumulative reward of the policy by alternately optimizing the policy network and the value network.
[0189] Through the above technical solutions, through the closed-loop optimization mechanism of perception-decision-making, the significant anchors generated by the front-end visual perception module are seamlessly connected to the back-end motion control strategy, and finally navigation action instructions that take into account safety, efficiency, and robustness are output.
[0190] Generally speaking, the present invention extracts global and local complementary features through a multi-scale feature fusion module, generates local position encoding by combining direction-sensitive convolution to enhance the local spatial perception ability of the vision Transformer, uses a uniform attention compensation mechanism to supplement global context information, adopts deformable convolution to dynamically generate saliency anchors driven by gradient density to suppress noise interference, and finally fuses the multi-objective reward functions of anchor saliency, path tracking, and policy exploration through a maximum entropy reinforcement learning framework, and outputs navigation action instructions under the condition of satisfying mechanical dynamics constraints, realizing the robust autonomous navigation of collaborative optimization of perception and decision-making in complex dynamic scenarios.
[0191] To verify the effectiveness of the technical solution of the present invention, a comparative experiment was carried out on a reinforcement learning decision-making system (hereinafter referred to as the "model group") including an EMA module, an ELA-T local attention module, a CB global compensation module, and a CAA anchor detection module and a benchmark method based on the original vision Transformer (ViT). During the experiment, the changing trends of four key reward indicators with the number of training rounds were recorded (as shown in the appendix Figure 6 as follows):
[0192] Crystal Reward: Reflects the dynamic fitting degree between the navigation path and the saliency anchor;
[0193] Heuristic Reward: Characterizes the comprehensive navigation effect under heuristic rules;
[0194] Action Reward: Quantifies the smoothness and physical feasibility of action output;
[0195] Freeze Reward: Evaluates the policy stability under the interference of dynamic obstacles.
[0196] The experimental results show (as shown in the appendix annotation "Reinforcement learning module using the model") that compared with the original ViT benchmark method, the method of the present invention has better characteristics in the following aspects:
[0197] Reward convergence: The fluctuation ranges of Crystal Reward and Heuristic Reward decrease significantly with the increase of the number of training rounds, indicating that the multi-module collaborative design improves the adaptability of the policy to complex scenarios;
[0198] Action stability: The increase in the mean value and the decrease in the variance of Action Reward verify the guarantee effect of the mechanical constraint module on action feasibility;
[0199] Anti-interference: The periodic fluctuation of Freeze Reward weakens, indicating that the dynamic anchor generation mechanism effectively suppresses the interference of environmental noise on decision-making.
[0200] Experimental data shows that through the collaborative design of multi-scale feature fusion, local-global attention complementarity, dynamic anchor point detection, and maximum entropy strategy optimization, the technical solution of the present invention can achieve end-to-end optimization of the perception and decision-making modules, improving the comprehensive performance of the autonomous navigation system.
[0201] The autonomous navigation system based on multi-channel enhancement and efficient local attention described below can be correspondingly referred to the autonomous navigation method based on multi-channel enhancement and efficient local attention described above.
[0202] Please refer to the attached Figure 8 , the present invention also provides an autonomous navigation system based on multi-channel enhancement and efficient local attention, including:
[0203] A multi-modal sensor module 100 for collecting RGB images and position coordinates;
[0204] An EMA fusion module 200 for multi-scale feature enhancement and cross-dimensional interaction;
[0205] A ViT-ELA module 300 integrating ELA-T local attention and CB dense attention injection;
[0206] A CAA module 400 for generating significant anchor points through deformable convolution;
[0207] An SAC reinforcement learning decision module 500 for outputting navigation actions based on anchor point saliency.
[0208] The system of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, which will not be elaborated here.
[0209] Please refer to the attached Figure 9 , the present invention also provides a computer device 40, including: a processor 41 and a memory 42, the memory 42 stores a computer program executable by the processor, and when the computer program is executed by the processor, it executes the above method.
[0210] The present invention also provides a storage medium 43, on which a computer program is stored, and when the computer program is run by the processor 41, it executes the above method.
[0211] Among them, the storage medium 43 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.
[0212] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An autonomous navigation method based on multi-channel enhancement and efficient local attention, characterized in that It includes the following steps: S1. Input an RGB visual image and fuse the position coordinates of the unmanned vehicle, convert it into a serialized feature matrix through block embedding, input the feature matrix into an efficient multi-scale attention fusion module, and generate multi-scale features through channel dimension grouping and cross-dimension interaction; S2. Input the multi-scale features into a vision Transformer, and embed a lightweight direction-sensitive local attention module before each Transformer block, generate position-encoded features along the height and width directions respectively through direction-sensitive 1D horizontal convolution and 1D vertical convolution, and capture local context relationships; S3. Inject uniform dense attention into each layer of the vision Transformer through the context broadcasting method to supplement the global information missed by sparse attention; S4. Input the feature matrix into an improved context anchor attention fusion module, generate significant anchors by replacing average pooling with deformable convolution, suppress noise and retain global context; S5. Input the significant anchors into a reinforcement learning decision module, optimize the action output based on the SAC algorithm, and output a navigation action instruction.
2. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, wherein The implementation of the efficient multi-scale attention fusion module in step S1 includes: Divide the input features into multiple subgroups along the channel dimension, and reshape some channels into the batch dimension to reduce the amount of calculation; Parallelly execute 1×1 convolution and 3×3 convolution to capture global and local features respectively; Fuse global and local features through cross-dimension matrix multiplication, and the formula is: Among them, represents element-wise multiplication, Y is the output feature matrix, W1 and W2 are learnable weight parameters, F GAP is the global feature, F 3×3 is the local feature.
3. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, wherein The implementation of the lightweight direction-sensitive local attention module in step S2 includes: Apply a 1D horizontal convolution kernel along the height direction to generate vertical position encoding; Apply a 1D vertical convolution kernel along the width direction to generate horizontal position encoding; Fuse the two-direction encoding through depthwise separable convolution, and the formula is: P cross = DWConv([P h ; P v ) Among them, DWConv(·) represents the depthwise separable convolution operation, and P h , P v are the horizontal encoded feature and the vertical encoded feature respectively.
4. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, wherein, The implementation of the context broadcasting method in step S3 includes: Define a uniform attention matrix, whose element values are covering all spatial positions; where H and W are the height and width of the input feature map respectively; Add the uniform attention to the output of the sparse attention of the vision Transformer, and the formula is: Attention final = A uniform ·V + Attention sparse Among them, Attention final is the fused attention matrix, A uniform is the uniform weight matrix, V is the value matrix, and Attention sparse is the original sparse attention matrix generated by the self-attention mechanism in the vision Transformer.
5. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, wherein The implementation of the deformable convolution in step S4 includes: Generate an initial offset through 3×3 convolution to dynamically adjust the pooling area; The pooling formula is: F pool = Σ k w k · X(p + p k + Δp k ) Among them, p is the coordinate of the current feature position; p k is the coordinate of k preset sampling points; Δp k is the learnable offset, indicating the dynamic position adjustment of the k-th sampling point; w k is the weight parameter of the k-th sampling point; The generation of significant anchors includes: Dynamically adjust the anchor distribution density based on the spatial gradient of high-level features, and the formula is: where H and W are the height and width of the high-level feature map respectively, is the spatial gradient of the high-level feature map, and ρ is the dynamically calculated anchor density; Apply 1×1 convolution to the pooled features for channel dimension recalibration to generate a set of significant anchors.
6. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, characterized in that The design of the reward function of the SAC algorithm in step S5 includes: Anchor significance reward term: Calculate the minimum Euclidean distance between the current position and the set of anchors; Policy entropy reward term: Maximize the policy entropy to improve the exploration ability; The formula of the reward function is: Among them, p t is the current position coordinate of the driverless vehicle, is the set of significant anchor points, represents the minimum Euclidean distance function, is the policy entropy, and α and β are normalization weight coefficients.
7. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 6, wherein The design of the reward function also includes: Path tracking reward term: Calculate the distance between the current position and the target path point; The formula of the reward function is: where p target is the target path point and γ is the normalized weight coefficient.
8. The autonomous navigation method based on multi-channel enhancement and efficient local attention according to claim 1, characterized in that The constraint condition of the action output is: The action output by the policy network follows the maximum entropy optimization objective, and the formula is: Among them, a t represents the action sampled at time t, ∼ represents sampling from a probability distribution, and π(·|s t ) represents the conditional probability distribution of the action under state s t .
9. An autonomous navigation system based on multi-channel enhancement and efficient local attention for implementing the method according to any one of claims 1-8, characterized in that, It includes: A multi-modal sensor module for collecting RGB images and position coordinates; An EMA fusion module for multi-scale feature enhancement and cross-dimension interaction; A ViT-ELA module integrating ELA-T local attention and CB dense attention injection; The CAA module is used to generate significant anchor points through deformable convolution; The SAC reinforcement learning decision-making module outputs navigation actions based on the saliency of the anchor points.
10. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-8.
Citation Information
Cited By
Method and device for generating driving guide line
CN120890465A
Unmanned ship robust collision avoidance decision-making method considering ship load change
CN122195014A