Method for detecting positive and negative obstacles on off-road surface of unmanned emergency rescue vehicle

By adopting YOLOv8 and multi-scale feature extraction and fusion strategies in unmanned emergency rescue vehicles, the problem of detection difficulty of positive and negative obstacles in off-road environments is solved, and accurate identification and intelligent decision-making support for off-road road obstacles are achieved.

CN119992514AActive Publication Date: 2025-05-13CHINA UNIV OF MINING & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510131343.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

In off-road environments, complex terrain and diverse positive and negative obstacles increase the difficulty of unmanned emergency rescue vehicles in perceive the surrounding environment, and existing detection methods are difficult to consider the problems of scale differences and noise interference at the same time.

Method used

Backbone of YOLOv8 is used for feature extraction, and branches are purified through multi-scale fine-grained feature, multi-scale coarse-grained feature extraction branches and multi-scale feature fusion strategies are reduced to reduce noise interference, extract multi-scale local and global information, and fuse feature map information to improve detection accuracy.

Benefits of technology

It realizes accurate identification of positive and negative obstacles on off-road roads, enhances the intelligent decision-making and control capabilities of unmanned emergency rescue vehicles, and improves the accuracy and robustness of the detection system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992514A_ABST
    Figure CN119992514A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting positive and negative obstacles on a cross-country road surface of an unmanned emergency rescue vehicle in the technical field of obstacle detection, and the method comprises the following steps: firstly, carrying out the feature extraction of an input cross-country road surface environment containing positive and negative obstacles through the Backbone of YOLOv8, thereby obtaining the information of different feature layers of the positive and negative obstacles; wherein the features of the second layer, the fourth layer, the sixth layer and the ninth layer are respectively marked as # imgabs0 #, # imgabs1 #, # imgabs2 # and # imgabs3 #, then, the four feature maps are sequentially sent to a constructed multi-scale fine-grained feature purification branch, a constructed multi-scale coarse-grained feature extraction branch and a constructed multi-scale feature fusion strategy, and finally, the feature maps passing through the multi-scale feature fusion strategy are sent to a Head part to obtain a positive and negative obstacle detection result of the cross-country road surface. According to the invention, accurate recognition of positive and negative obstacles on the cross-country road surface is realized, and key information is provided for intelligent decision and control of a subsequent unmanned emergency rescue vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of obstacle detection, and in particular to a method for detecting positive and negative obstacles on an off-road road surface of an unmanned emergency rescue vehicle. Background Art

[0002] Emergency rescue vehicles often need to perform rescue missions in off-road environments, such as earthquake rescue, forest fire rescue, mudslide rescue, etc. The complex and changeable terrain in off-road environments poses a threat to the safety of drivers, so the application of unmanned emergency rescue vehicles is particularly important. During the driving process of unmanned emergency rescue vehicles, their perception technology is like eyes, which can perceive the surrounding environment in real time, so as to respond quickly and accurately to ensure the safe driving of the vehicle. However, in off-road environments, the diverse and complex positive and negative obstacles (positive obstacles such as piles of earth, rocks, and trees on the ground, and negative obstacles such as pits and trenches) greatly increase the difficulty of unmanned vehicles to perceive the surrounding environment. Therefore, it is urgent to study the detection method of positive and negative obstacles on off-road roads for unmanned emergency rescue vehicles.

[0003] Deep learning is highly favored in the field of positive and negative obstacle detection in off-road environments due to its automatic feature extraction, excellent generalization performance, high precision and strong robustness. Detection methods based on deep learning can be divided into single-stage and two-stage. Although the existing two-stage detection method Fast-RCNN based on the improvement of Fast-RCNN achieves high-precision obstacle detection, its detection speed is slow and is not suitable for unmanned vehicles. In contrast, the single-stage method YOLO (You Only Look Once) series balances detection accuracy and speed, and therefore has received widespread attention in positive and negative obstacle detection.

[0004] There are many types of off-road obstacles, and the characteristics of positive and negative obstacles are significantly different. Among them, positive obstacles (such as mounds, bushes, rocks and trees) have large scale differences, and negative obstacles (such as pits and trenches, puddles, etc.) are easily affected by changes in illumination, water surface reflections, etc., which produce visual noise interference. In view of the characteristics of these obstacles, a variety of positive and negative obstacle detection models based on improved YOLO have emerged. In order to solve the problem of fuzzy feature information of negative obstacles caused by complex lighting conditions, the improved YOLOv5 negative obstacle detection model achieves dual feature fusion through path enhancement, and achieves good detection effect, but still does not fully consider the multi-scale characteristics of positive obstacles. In addition, the improved YOLOv7-tiny model enhances the network's perception of feature maps and realizes the simultaneous detection of positive and negative obstacles, but it mainly focuses on positive obstacles with small scale differences and negative obstacles in urban environments, and cannot be applied to positive and negative obstacle detection in off-road environments. Therefore, when designing a positive and negative obstacle detection network for off-road environments, it is necessary to deeply consider the joint effects of scale differences and noise interference to ensure the accuracy of the detection system. Summary of the invention

[0005] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0006] Therefore, the purpose of the present invention is to provide a method for identifying positive and negative obstacles on off-road surfaces for unmanned emergency rescue vehicles, so as to achieve accurate identification of positive and negative obstacles on off-road surfaces and provide key information for subsequent intelligent decision-making and control of unmanned emergency rescue vehicles.

[0007] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions: A method for detecting positive and negative obstacles on an off-road road surface of an unmanned emergency rescue vehicle, the steps are as follows: S1. First, the feature extraction of the input off-road road environment containing positive and negative obstacles is performed through the Backbone of YOLOv8, so as to obtain the different feature layer information of positive and negative obstacles, where the 2nd, 4th, 6th and 9th layer features are recorded as , , , ; S2. Secondly, the four feature maps are sequentially sent to the constructed multi-scale fine-grained feature purification branch, multi-scale coarse-grained feature extraction branch and multi-scale feature fusion strategy. The multi-scale fine-grained feature purification branch is used to reduce the interference of noise and extract multi-scale local information. The multi-scale coarse-grained feature extraction branch is used to extract multi-scale global information. The multi-scale feature fusion strategy makes the feature map sent to the Head contain more target feature information. S3. Finally, the feature map after the multi-scale feature fusion strategy is sent to the Head part to obtain the positive and negative obstacle detection results of the off-road road.

[0008] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads for unmanned emergency rescue vehicles described in the present invention, the multi-scale fine-grained feature purification branch is composed of multi-receptive field noise filtering feature learning and information interweaving.

[0009] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to the present invention, the multi-receptive field noise filter feature learning is composed of a noise filter and a multi-receptive field feature learning, and the multi-receptive field noise filter feature learning performs the following steps: Assume the input feature map is , the output feature map is , first, use the convolution kernel size of The convolutional layer pair The number of channels is adjusted, and then the transformed feature map is separated into channels. The feature map is obtained by formula (1) , , , , and use 4 convolution kernels of different sizes respectively , , , Act on 4 feature maps to extract multi-receptive field features and generate feature maps , , , ; Then, the differential attention mechanism is used on these feature maps through formula (2) to obtain the feature map , , , ; (1) (2) In the formula, Indicates channel segmentation; represents the differential attention mechanism; Indicates that the convolution kernel size is Convolution operation; Secondly, , , , Perform feature concatenation on the channel dimension, disrupt the order of feature channels through channel shuffling operation, and apply it again Convolution, through formula (3) to get the feature map ; (3) In the formula, Indicates channel splicing, Indicates channel shuffle; Finally, the processed feature map is added to the original feature map, which not only helps to alleviate the gradient vanishing problem in deep networks, but also promotes cross-layer interaction of information. The added feature map is sent to MRL to further extract multi-receptive field features. The output feature map is obtained by expression (4): ; (4) in, Represents multi-receptive field feature learning.

[0010] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to the present invention, the differential attention mechanism performs the following steps: Assume the input feature map is , the output feature map is , then the difference feature map is generated through formula (5) ; (5) In the formula, Represents a bicubic interpolation upsampling operation; represents the average pooling downsampling operation; Indicates element subtraction; After differential processing, the Sigmoid activation function is used on the differential feature map to generate a series of weight values ​​between 0 and 1. The output feature map is obtained by formula (6): .

[0011] (6) In the formula, represents the Sigmoid activation function, represents the addition of elements, Represents element-wise multiplication.

[0012] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles described in the present invention, the multi-receptive field feature learning is composed of two symmetrical branches; One of the branches adopted and Depth-wise separable convolution, used for feature extraction in horizontal and vertical directions respectively; The other branch uses different convolution kernels as and The depth of separable convolution is used to increase the size of the convolution kernel and expand the receptive field.

[0013] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles described in the present invention, the information interleaving is composed of Minish activation function, up and down sampling operation and 3×3 convolution, and the information interleaving performs the following steps: Use bilinear interpolation upsampling and average pooling downsampling operations to scale the two feature maps to the same size; Next, the two different feature maps are fused by element-wise multiplication to introduce more nonlinear factors. Then, the 3×3 convolution mines the feature information of the target without significantly increasing the number of parameters. Suppose the two input features are , The output features are , the specific expression is as follows: (7) (8) In the formula, Represents the Minish activation function; Represents a bilinear interpolation operation.

[0014] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles described in the present invention, the multi-scale coarse-grained feature extraction branch is composed of a directional attention mechanism and information interweaving, and the directional attention mechanism is composed of a vertical directional feature extraction branch, a horizontal directional feature extraction branch and a channel feature extraction branch. The multi-scale coarse-grained feature extraction branch performs the following steps: The vertical feature extraction branch is performed through the horizontal maximum pooling, average pooling and The convolution operation deeply mines the coarse-grained features in the vertical direction; The horizontal feature extraction branch uses the maximum pooling, average pooling and Convolution captures horizontal context information; The vertical feature extraction branch and the horizontal feature extraction branch use the Softmax activation function to normalize the extracted features, generate a series of weights to dynamically adjust the input feature information, and realize the adaptive adjustment of the features. Suppose the input feature map is , then the output of the horizontal feature extraction branch and the vertical feature extraction branch , Through equations (9), (10) and (11), we can get: (9) (10) (11) in, represents the average pooling in the x direction, represents the maximum pooling in the x direction, represents the average pooling in the y direction, represents the maximum pooling in the y direction, represents maximum variance pooling, Represents a transpose operation; At the same time, the channel feature extraction branch introduces maximum variance pooling, which highlights the features with the largest changes by selecting the area with the largest variance in the feature map, and performs deep modeling in the channel direction of the feature map through a multi-layer perceptron to promote the deep fusion of information between channels. The output of the channel feature extraction branch is generated through formula (12): ; (12) In the formula, represents a multi-layer perceptron, ReLU activation function. Finally, the features extracted from different directions by the three branches are fused by element-wise addition. At the same time, in order to prevent network degradation, the input feature map is connected by residual connection. Introduce and generate the output feature map as , enhancing the model’s sensitivity to spatial features.

[0015] As a preferred solution of the method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles described in the present invention, the multi-scale feature fusion strategy consists of average pooling, deconvolution and channel splicing.

[0016] Compared with the prior art, the present invention has the following beneficial effects: the core of the present invention lies in the CFFI part it constructs, which integrates a multi-scale fine-grained feature purification branch, a multi-scale coarse-grained feature extraction branch, and a multi-scale feature fusion strategy. Among them, the multi-scale fine-grained feature purification branch aims to reduce environmental noise interference and deeply explore the fine-grained features of positive and negative obstacles; while the multi-scale coarse-grained feature extraction branch is mainly used to extract the coarse-grained features of positive and negative obstacles at different scales; the multi-scale feature fusion strategy further fuses the feature map information before inputting the feature map into the detection head, so that each feature map contains more positive and negative obstacle target feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in combination with the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them: Figure 1 This is an overall framework diagram of a method for identifying positive and negative obstacles of an unmanned emergency rescue vehicle in an off-road environment according to the present invention; Figure 2 A flow chart of the multi-scale fine-grained feature purification branch provided by the present invention; Figure 3 A flow chart of information interweaving provided by the present invention; Figure 4 A flow chart of the directional attention mechanism provided by the present invention; DETAILED DESCRIPTION

[0018] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0019] A method for identifying positive and negative obstacles on off-road surfaces for unmanned emergency rescue vehicles Figure 1 As shown in Figure 1, first, Backbone is used to extract features from the input off-road road environment containing positive and negative obstacles, thereby obtaining different feature layer information of positive and negative obstacles, where the 2nd, 4th, 6th, and 9th layer features are recorded as , , , ; Secondly, these four feature maps are sent to the constructed coarse-fine granularity feature interaction (CFFI) part to improve the information loss problem caused by the long fusion path. CFFI includes a multi-scale fine-grained feature purification branch, a multi-scale coarse-grained feature extraction branch and a multi-scale feature fusion strategy (Multi-Scale Feature FusionStrategy, MFFS). The multi-scale fine-grained feature purification branch is mainly used to reduce the interference of noise and extract multi-scale local information, such as edge details; the multi-scale coarse-grained feature extraction branch is mainly used to extract multi-scale global information, such as shape structure; MFFS makes the feature map sent to the Head part contain more target feature information. Finally, the feature map after MFFS is sent to the Head part to obtain the positive and negative obstacle detection results on the off-road road.

[0020] (1) Multi-scale fine-grained feature purification branch The multi-scale fine-grained feature purification branch is mainly composed of the constructed multi-field noise filtering feature learning (Multi-Field Noise Filtering Feature Learning, MNFFL) and information interweaving (Information Interweaving, IIW).

[0021] like Figure 2 As shown in Figure 2, MNFFL consists of a noise filter and multi-receptive field feature learning (MRL). The overall process is as follows: Assume that the input feature map is , the output feature map is First, use a convolution kernel size of The convolutional layer pair The number of channels is adjusted to increase the diversity and complexity of the features. In order to ensure the computational complexity of the algorithm, the transformed feature map is then channel separated, and the feature map is obtained by formula (1): , , , , and use 4 convolution kernels of different sizes respectively , , , Act on 4 feature maps to extract multi-receptive field features and generate feature maps , , , , then use the Differential Attention Mechanism (DA) on these feature maps to get the feature maps , , , The specific expression is shown in formula (2).

[0022] (1) (2) In the formula, Indicates channel segmentation; represents the differential attention mechanism; Indicates that the convolution kernel size is The convolution operation.

[0023] Secondly, , , , Perform feature concatenation in the channel dimension, disrupt the order of feature channels through channel shuffling operations, increase the nonlinearity of model learning, avoid information redundancy, and apply it again Convolution not only restores the number of channels of the feature map, but also realizes further fusion and refinement of features. The feature map is obtained through formula (3): .

[0024] (3) In the formula, Indicates channel splicing; Indicates channel shuffle; Finally, the processed feature map is added to the original feature map, which not only helps to alleviate the gradient vanishing problem in deep networks, but also promotes cross-layer interaction of information. The added feature map is sent to MRL to further extract multi-receptive field features. The output feature map is obtained by expression (4): The specific expression is as follows: (4) in, Represents multi-receptive field feature learning.

[0025] like Figure 2 As shown in Figure 1, DA reduces noise interference through bicubic interpolation upsampling, average pooling downsampling and differential operation, thereby achieving feature enhancement. Downsampling extracts the background information of the feature map, while upsampling focuses on fine-grained information. The differential operation subtracts the feature map processed by upsampling and downsampling from the original image. Assume that the input feature map is , the output feature map is , then the difference feature map is generated through formula (5) .

[0026] (5) In the formula, Represents a bicubic interpolation upsampling operation; represents the average pooling downsampling operation; Represents element-wise subtraction.

[0027] After differential processing, the feature map can better capture fine-grained features and reduce the interference of background noise. To further enhance the expressiveness of fine-grained information, the Sigmoid activation function is used on the feature map after differentiation to generate a series of weight values ​​between 0 and 1. These weight values ​​are then dynamically applied to the original input feature map to achieve adaptive weighting of the feature map, so that the model can flexibly adjust the focus information according to the importance of the feature. According to the above, the output feature map is obtained by formula (6): .

[0028] (6) In the formula, Represents the Sigmoid activation function; It represents the addition of elements; Represents element-wise multiplication.

[0029] MRL consists of two symmetrical branches. Through its two symmetrical and complementary branch designs, it further realizes the capture and fusion of feature information and promotes the extraction of coarse-grained information in the subsequent process. and Depthwise separable convolution is used for feature extraction in the horizontal and vertical directions respectively. It effectively reduces the number of parameters while enhancing the model’s sensitivity to spatial dimension features, which helps capture fine-grained features of small targets (targets with a resolution less than 32 pixels × 32 pixels). In addition, the embedded batch normalization (BatchNorm, BN) solves the problem of numerical instability in deep neural networks by standardizing the feature layers in the network, making the distribution of features in the same batch similar and the network easier to train. GELU is selected as the activation function. Compared with traditional activation functions such as ReLU, GELU has a smooth transition in the negative region, which helps alleviate the problem of gradient vanishing, introduces more nonlinearities, and enhances the expressiveness of the model. The feature map fed into MRL is denoted as , then the branch can get the output through formula (7) .

[0030] (7) In the formula, represents batch normalization; represents the GELU activation function; Represents the convolution kernel as Depthwise separable convolution.

[0031] Similarly, the other branch uses different convolution kernels as and The depthwise separable convolution increases the size of the convolution kernel and expands the receptive field, enabling the model to capture contextual information and help extract fine-grained features of large targets (targets with a resolution greater than 32 pixels × 32 pixels). The introduction of this dual branch complements each other, that is, one branch focuses on the fine-grained features of small targets, while the other helps capture the fine-grained features of large targets. The outputs of the two branches are effectively fused by feature addition, improving the model's feature extraction capabilities for objects of different sizes. The output of this branch It is obtained by formula (8).

[0032] (8) In order to better utilize the features extracted by Backbone and enhance the representation ability of the features, IIW is proposed, such as Figure 3 As shown in the figure, the captured features are fused with the features extracted by Backbone in the form of an attention mechanism. IIW is mainly composed of the Minish activation function, up and down sampling operations, and 3×3 convolution. The introduction of the Minish activation function improves the nonlinear fitting ability of the network through its unique nonlinear mapping ability. The two feature maps are scaled to the same size using bilinear interpolation upsampling and average pooling downsampling operations; then, the two different feature maps are fused by element multiplication, introducing more nonlinear factors to further improve the expression ability of the model. In addition, the 3×3 convolution effectively mines the feature information of the target without significantly increasing the number of parameters. Assume that the two input features are , The output features are , the specific expression is as follows: (9) (10) In the formula, Represents the Minish activation function; Represents a bilinear interpolation operation.

[0033] The fine-grained features after MNFFL are separated by channels and the feature maps extracted from Backbone are extracted. , Feature maps with the same number of channels and , then, using IIW and Respectively , Interact to generate input for the coarse-grained feature extraction branch , This process not only refines the fine-grained feature information, but also provides diversified information for the subsequent coarse-grained feature extraction, enhancing the expressiveness and robustness of the model.

[0034] (2) Multi-scale coarse-grained feature extraction branch The multi-scale coarse-grained feature extraction branch is mainly composed of the designed directional attention mechanism (DRA) and IIW. The input of the coarse-grained feature extraction branch is the output of the fine-grained feature extraction branch. , Extracted from backbone branch .

[0035] In order to improve the positioning ability of the model and promote the capture of coarse-grained feature information, while effectively eliminating redundant information. The traditional coordinate attention mechanism has limitations in exploring the feature correlation in the horizontal and vertical directions. Based on this, a directional attention (DRA) mechanism is constructed to enhance the model's understanding of the spatial features of the target object, such as spatial layout and relative position relationship, by mining the feature information of the feature map in the horizontal and vertical dimensions.

[0036] The DRA flow chart is as follows Figure 4 As shown in Figure 1, DRA consists of three major branches: vertical feature extraction branch (Vertical Feature Extraction Branch, VFB), horizontal feature extraction branch (Horizontal Feature Extraction Branch, HFB) and channel feature extraction branch (Channel Feature Extraction Branch, CFB). VFB uses horizontal maximum pooling, average pooling and The convolution operation of , deeply explores the coarse-grained features in the vertical direction. Similarly, HFB uses the maximum pooling, average pooling and Convolution captures the context information in the horizontal direction. At the same time, both branches use the transpose (Reshape) operation to explore the potential correlation between different features and enrich the coarse-grained information of the feature map. To further improve the feature representation capability, the VFB and HFB branches use the Softmax activation function to normalize the extracted features, generate a series of weights to dynamically adjust the input feature information, and realize the adaptive adjustment of the features. Assume that the input feature map is , then VFB and HFB output , It can be generated by equations (11), (12) and (13).

[0037] (11) (12) (13) in, represents the average pooling in the x direction, represents the maximum pooling in the x direction, represents the average pooling in the y direction, represents the maximum pooling in the y direction, represents maximum variance pooling, Represents a transpose operation.

[0038] At the same time, CFB introduces the maximum variance pooling (Max Variance Pooling), which selects the area with the largest variance in the feature map, highlights the features with the largest changes, and effectively extracts the coarse-grained information in the channel direction. In addition, a multilayer perceptron (MLP) is used to perform deep modeling in the channel direction of the feature map to promote the deep fusion of information between channels and improve the richness of feature information. The output of CFB is generated through formula (14): .

[0039] (14) In the formula, represents a multi-layer perceptron, Represents the ReLU activation function.

[0040] Furthermore, the features extracted from different directions by the three branches are fused by element-wise addition. At the same time, in order to prevent network degradation, the input feature map is connected in a residual connection. Introduce and generate the output feature map as , enhancing the model’s sensitivity to spatial features.

[0041] (3) Multi-scale feature fusion strategy In the traditional YOLOv8 structure, the feature maps after Neck fusion are directly fed into the detection head, but these feature maps are often insufficient in balancing local details and global information, and fail to fully contain the fine-grained and coarse-grained feature information of the target. In order to improve the detection performance of large and small targets, MFFS is proposed to ensure that the feature maps contain more fine-grained and coarse-grained information of the target.

[0042] The structure diagram of MFFS is as follows: Figure 1 As shown in the last dotted box in the GFFI in , it is mainly composed of average pooling, deconvolution and channel splicing. Average pooling, as a means of downsampling, effectively captures and strengthens the overall features and coarse-grained information of the feature map by reducing the spatial size of the feature map. Next, deconvolution upsampling is used to reduce the loss of spatial resolution caused by average pooling. Deconvolution improves the spatial resolution of the feature map by learning the association between neighboring pixels, and enhances the expression ability of fine-grained features while retaining the original information. By alternating downsampling and upsampling, a step-by-step guided mechanism is formed. In this process, high-resolution feature maps (rich in fine-grained information) are continuously fused with low-resolution feature maps (rich in coarse-grained information), and each fusion is based on the further optimization of the previous fusion result. Through this step-by-step refinement method, feature maps containing rich coarse-grained and fine-grained information are generated, thereby further improving the detection performance. The overall process is as follows: (15) (16) in, The convolution kernel is Deconvolution.

[0043] Although the present invention has been described above with reference to the embodiments, various modifications may be made thereto and parts thereof may be replaced by equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the embodiments disclosed in the present invention may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles, characterized in that: Here are the steps: S1. First, the feature extraction of the input off-road road environment containing positive and negative obstacles is performed through the Backbone of YOLOv8, so as to obtain the different feature layer information of positive and negative obstacles, where the 2nd, 4th, 6th and 9th layer features are recorded as , , , ; S2. Secondly, the four feature maps are sequentially sent to the constructed multi-scale fine-grained feature purification branch, multi-scale coarse-grained feature extraction branch and multi-scale feature fusion strategy. The multi-scale fine-grained feature purification branch is used to reduce the interference of noise and extract multi-scale local information. The multi-scale coarse-grained feature extraction branch is used to extract multi-scale global information. The multi-scale feature fusion strategy makes the feature map sent to the Head contain more target feature information. S3. Finally, the feature map after the multi-scale feature fusion strategy is sent to the Head part to obtain the positive and negative obstacle detection results of the off-road road.

2. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 1, characterized in that: The multi-scale fine-grained feature purification branch consists of multi-receptive field noise filtering feature learning and information interweaving.

3. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 2, characterized in that: The multi-receptive field noise filter feature learning consists of a noise filter and a multi-receptive field feature learning. Multi-receptive field noise filter feature learning performs the following steps: Assume the input feature map is , the output feature map is , first, use the convolution kernel size of The convolutional layer pair The number of channels is adjusted, and then the transformed feature map is separated into channels. The feature map is obtained by formula (1) , , , , and use 4 convolution kernels of different sizes respectively , , , Act on 4 feature maps to extract multi-receptive field features and generate feature maps , , , ; Then, the differential attention mechanism is used on these feature maps through formula (2) to obtain the feature map , , , ; (1); (2); In the formula, Indicates channel segmentation; represents the differential attention mechanism; Indicates that the convolution kernel size is Convolution operation; Secondly, , , , Perform feature concatenation on the channel dimension, disrupt the order of feature channels through channel shuffling operation, and apply it again Convolution, through formula (3) to get the feature map ; (3); In the formula, Indicates channel splicing, Indicates channel shuffle; Finally, the processed feature map is added to the original feature map, which not only helps to alleviate the gradient vanishing problem in deep networks, but also promotes cross-layer interaction of information. The added feature map is sent to MRL to further extract multi-receptive field features. The output feature map is obtained by expression (4): ; (4); in, Represents multi-receptive field feature learning.

4. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 3, characterized in that: The differential attention mechanism performs the following steps: Assume the input feature map is , the output feature map is , then the difference feature map is generated through formula (5) ; (5); In the formula, Represents a bicubic interpolation upsampling operation; represents the average pooling downsampling operation; Indicates element subtraction; After differential processing, the Sigmoid activation function is used on the differential feature map to generate a series of weight values ​​between 0 and 1. The output feature map is obtained by formula (6): ; (6); In the formula, represents the Sigmoid activation function, represents the addition of elements, Represents element-wise multiplication.

5. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 3, characterized in that: The multi-receptive field feature learning consists of two symmetrical branches; One of the branches adopted and Depth-wise separable convolution, used for feature extraction in horizontal and vertical directions respectively; The other branch uses different convolution kernels as and The depth of separable convolution is used to increase the size of the convolution kernel and expand the receptive field.

6. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 2, characterized in that: The information interleaving is composed of the minish activation function, up and down sampling operations and 3×3 convolution. The information interleaving performs the following steps: Use bilinear interpolation upsampling and average pooling downsampling operations to scale the two feature maps to the same size; Next, the two different feature maps are fused by element-wise multiplication to introduce more nonlinear factors. Then, the 3×3 convolution mines the feature information of the target without significantly increasing the number of parameters. Suppose the two input features are , The output features are , the specific expression is as follows: (7); (8); In the formula, Represents the Minish activation function; Represents a bilinear interpolation operation.

7. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 1, characterized in that: The multi-scale coarse-grained feature extraction branch is composed of a directional attention mechanism and information interweaving. The directional attention mechanism is composed of a vertical feature extraction branch, a horizontal feature extraction branch and a channel feature extraction branch. The multi-scale coarse-grained feature extraction branch performs the following steps: The vertical feature extraction branch is performed through the horizontal maximum pooling, average pooling and The convolution operation deeply mines the coarse-grained features in the vertical direction; The horizontal feature extraction branch uses the maximum pooling, average pooling and Convolution captures horizontal context information; The vertical feature extraction branch and the horizontal feature extraction branch use the Softmax activation function to normalize the extracted features, generate a series of weights to dynamically adjust the input feature information, and realize the adaptive adjustment of the features. Suppose the input feature map is , then the output of the horizontal feature extraction branch and the vertical feature extraction branch , Through equations (9), (10) and (11), we can get: (9); (10); (11); in, represents the average pooling in the x direction, represents the maximum pooling in the x direction, represents the average pooling in the y direction, represents the maximum pooling in the y direction, represents maximum variance pooling, Represents a transpose operation; At the same time, the channel feature extraction branch introduces maximum variance pooling, which highlights the features with the largest changes by selecting the area with the largest variance in the feature map, and performs deep modeling in the channel direction of the feature map through a multi-layer perceptron to promote the deep fusion of information between channels. The output of the channel feature extraction branch is generated through formula (12): ; (12); In the formula, represents a multi-layer perceptron, ReLU activation function. Finally, the features extracted from different directions by the three branches are fused by element-wise addition. At the same time, in order to prevent network degradation, the input feature map is connected by residual connection. Introduce and generate the output feature map as , enhancing the model’s sensitivity to spatial features.

8. The method for detecting positive and negative obstacles on off-road roads of unmanned emergency rescue vehicles according to claim 1, characterized in that: The multi-scale feature fusion strategy consists of average pooling, deconvolution and channel splicing.

Citation Information

Patent Citations

  • Multi-granularity cross modal feature fusion pedestrian re-identification method and re-identification system

    CN110598654A

  • Multi-scale pedestrian re-identification method based on multi-granularity depth feature fusion

    CN112818931A

  • Satellite image recognition method based on machine learning

    CN118015488A