Robot terrain recognition method and system facing easily-confused terrain categories
Patent Information
- Application Number
- CN202610908227.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-08
AI Technical Summary
[0005]有鉴于此,本申请提出了一种面向易混淆地形类别的机器人地形识别方法及系统,用以解决现有技术中机器人对视觉特征高度相似的易混淆地形类别识别准确率低、误判率高的技术问题
(1)本申请实施例公开的方法,通过多尺度地形纹理感知调制、混淆差异引导的纹理方向注意力、综合混淆度主动配对、语义引导纹理迁移增强以及联合训练相结合的方案,解决了现有技术中机器人对视觉特征高度相似的易混淆地形类别识别准确率低、误判率高的技术问题,有效提升了模型对沥青与水泥、碎石与砂砾等易混淆地形类别的细粒度判别能力,降低了误判率,为机器人在复杂多地形环境中的自主导航和运动控制提供了可靠的地形感知支持。
Smart Images

Figure CN122714801A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot environmental perception technology, and in particular to a robot terrain recognition method and system for easily confused terrain categories. Background Technology
[0002] Terrain recognition is a key technology in autonomous navigation and motion control, and its accuracy directly affects the robot's safety, path planning efficiency, and energy consumption control in complex environments. With the widespread application of ground mobile robots in search and rescue, agricultural inspection, and logistics distribution, the requirements for terrain recognition accuracy are increasing. This is especially true when robots need to operate continuously between various similar road surfaces, such as asphalt and concrete, or gravel and sand. Simply distinguishing between passable and impassable surfaces is no longer sufficient for precise control; accurate identification of specific terrain materials is essential.
[0003] However, existing terrain recognition methods generally suffer from low accuracy and high false positive rates when faced with easily confused terrain categories that have highly similar visual features. For example, asphalt and cement are extremely similar in macroscopic color and texture, and it is difficult to capture the subtle differences between them in terms of microscopic texture granularity and surface reflectivity by relying solely on global features or conventional convolutional neural networks; gravel and sand also highly overlap in shape and distribution, and traditional classification models often classify them into the same category. This confusion severely restricts the robot's autonomous decision-making ability in scenarios with alternating terrains.
[0004] The root cause of these problems lies in the fact that existing classification models, when faced with terrain categories that have highly similar visual features, lack sufficient feature representation capabilities and the fineness of their decision boundaries to reliably distinguish such subtle differences. When the global appearance of two terrain categories highly overlaps, the model struggles to spontaneously extract sufficiently discriminative information from microscopic dimensions such as local texture and surface granularity, leading to misclassification. Currently, there is a lack of mature technical solutions in this field that can effectively address the challenge of identifying such easily confused terrain categories. Summary of the Invention
[0005] In view of this, this application proposes a robot terrain recognition method and system for easily confused terrain categories, in order to solve the technical problems of low accuracy and high misjudgment rate of robots in recognizing easily confused terrain categories with highly similar visual features in the prior art.
[0006] The technical solution of this application is implemented as follows: In a first aspect, this application provides a robot terrain recognition method for easily confused terrain categories, characterized by the following steps: S1: Acquire and preprocess the terrain image ahead of the robot's path to obtain the preprocessed terrain image; S2: Input the preprocessed terrain image into the backbone feature extraction network to extract the depth feature map, wherein the preprocessed terrain image is used as the target sample; S3: Extract the texture response map of each branch from the depth feature map by using multiple depth-separable convolutional branches of different scales; calculate the branch weights based on the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map; S4: Obtain candidate confused samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confused samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and spatial attention to the modulated feature map to obtain the enhanced feature map; S5: Pool the enhanced feature map to obtain a feature vector, calculate the overall confusion degree between the target sample and each candidate confused sample based on the feature vector, and select a high-risk confused sample for each target sample according to the overall confusion degree to form a high-risk confused sample pair; S6: Input the target sample and the high-risk confusion sample into the semantically guided texture transfer subnetwork, generate a transfer mask and extract local texture patterns from the high-risk confusion sample, determine the injection position on the non-subject region of the target sample through the transfer mask, and superimpose the local texture patterns onto the injection position to obtain a confusion-enhanced sample that maintains the semantics of the target category. S7: Input the confused enhanced samples and the original samples into the classification network for joint training to obtain a trained terrain recognition model; S8: Input the terrain image to be identified into the trained terrain recognition model and output the terrain category recognition result.
[0007] In some embodiments, step S3, calculating the branch weights based on the texture energy distribution of each texture response map, specifically includes: performing high-frequency texture energy estimation on each texture response map to obtain the texture energy value of each branch, and mapping the texture energy value to the weight of each branch using the Softmax function.
[0008] In some embodiments, step S4 specifically includes: S41: Perform global average pooling on the feature maps of the target sample and the candidate confused sample respectively to obtain their respective feature vectors, and calculate the absolute value of the difference between the two as the difference vector. S42: After concatenating the difference vector with the feature vector of the target sample, the convolution is performed using a one-dimensional convolution and the Sigmoid function to generate channel attention weights. The channel attention weights are then multiplied with the feature map of the target sample to obtain a channel-weighted feature map. S43: Perform convolution processing on the channel weighted feature map in the horizontal, vertical and diagonal directions respectively to obtain the texture response in each direction. Generate directional weights through the Softmax function based on the global average pooling value of the texture response in each direction, and perform weighted fusion on the texture response in each direction according to the directional weights to obtain the directional texture feature map. S44: Calculate the sum of the absolute values of the channel-by-channel differences between the feature maps of the target sample and the candidate confused samples, and obtain the difference region map after normalization; S45: The average pooling map and max pooling map of the directional texture feature map are concatenated with the difference region map, and a spatial attention map is generated by two-dimensional convolution and the Sigmoid function. The spatial attention map is then multiplied with the channel weighted feature map to obtain the enhanced feature map.
[0009] In some embodiments, step S5, calculating the overall confusion degree between the target sample and each candidate confused sample, specifically includes: calculating the feature distance between the target sample and the candidate confused sample, the confidence that the target sample is correctly predicted by the classifier, the probability that the candidate confused sample is misclassified as the target category, the inter-category risk weight determined according to the pre-statistical confusion matrix, and the similarity response of the target sample and the candidate confused sample in the attention difference region, as five indicators, and weighting and summing the five indicators to obtain the overall confusion degree, wherein the attention difference region is obtained by normalizing the channel-by-channel difference of the feature maps of the target sample and the candidate confused sample.
[0010] In some embodiments, in step S5, when selecting high-risk confusion samples for each target sample based on the overall confusion degree, one or more samples with the highest overall confusion degree are selected from the candidate confusion samples as high-risk confusion samples.
[0011] In some embodiments, in step S6, the semantically guided texture transfer subnetwork includes a texture localization branch, a texture transformation branch, and a semantic fusion branch; the texture localization branch generates the transfer mask based on the feature map of the target sample, the feature map of the high-risk confusion sample, and the difference between the two; the texture transformation branch extracts local texture patterns from the high-risk confusion sample and performs normalization transformation; the semantic fusion branch superimposes the transformed local texture patterns onto the non-subject region of the target sample at the position and intensity controlled by the transfer mask.
[0012] In some embodiments, when the semantic fusion branch superimposes the local texture pattern, it calculates the probability that the confusing and enhanced sample is determined to be the target category by the classifier, and constructs a semantic preservation constraint loss based on the probability to constrain the confusing and enhanced sample to preserve the semantics of the target category.
[0013] In some embodiments, the joint training in step S7 specifically includes: training using a total loss function that includes the classification loss of the original samples, the classification loss of the confusion-enhanced samples, the feature rejection loss, the attention focus loss, and the semantic preservation constraint loss; the feature rejection loss is calculated based on the difference between the feature distance and the dynamic interval between the target sample and the high-risk confusion sample, and the dynamic interval is determined based on the comprehensive confusion degree; the attention focus loss is calculated based on the degree of overlap between the spatial attention map and the difference region map; the semantic preservation constraint loss is constructed based on the probability that the confusion-enhanced sample is classified as the target category by the classifier; and, during training, the weight of the classification loss of the confusion-enhanced samples in the total loss is adaptively adjusted based on the ratio of the error rate between the target category and the confusion category on the validation set to the initial error rate.
[0014] Secondly, this application discloses a robot terrain recognition system for easily confused terrain categories, comprising: The image processing module is used to acquire and preprocess the terrain image in front of the robot's path to obtain the preprocessed terrain image. The feature extraction module is used to input the preprocessed terrain image into the backbone feature extraction network to extract depth feature maps; wherein, the preprocessed terrain image is used as the target sample; The multi-scale texture modulation module is used to extract the texture response map of each branch from the depth feature map through multiple depth-separable convolutional branches of different scales, calculate the branch weights according to the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map. The confusion difference attention module is used to acquire candidate confusion samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confusion samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and the spatial attention to the modulated feature map to obtain the enhanced feature map. The confusion calculation and pairing module is used to pool the enhanced feature map to obtain a feature vector, calculate the comprehensive confusion between the target sample and each candidate confusion sample based on the feature vector, and select a high-risk confusion sample for each target sample according to the comprehensive confusion, thus forming a high-risk confusion sample pair. The texture transfer enhancement module is used to input the target sample and the high-risk confusion sample into the semantic guidance texture transfer sub-network, generate a transfer mask and extract local texture patterns from the high-risk confusion sample, and use the transfer mask to control the injection position to superimpose the local texture patterns onto the non-subject region of the target sample to obtain a confusion enhancement sample that maintains the semantics of the target category. The joint training module is used to input the confused and enhanced samples and the original samples into the classification network for joint training to obtain a trained terrain recognition model; The recognition output module is used to input the terrain image to be recognized into the trained terrain recognition model and output the terrain category recognition result.
[0015] Thirdly, embodiments of this application also disclose an electronic device, including a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the robot terrain recognition method for easily confused terrain categories described in the first aspect.
[0016] This application has the following advantages over the prior art: (1) The method disclosed in this application solves the technical problems of low accuracy and high misjudgment rate of robots in identifying easily confused terrain categories with highly similar visual features by combining multi-scale terrain texture perception modulation, texture direction attention guided by confusion difference, active pairing of comprehensive confusion degree, semantically guided texture transfer enhancement and joint training. It effectively improves the model's fine-grained discrimination ability for easily confused terrain categories such as asphalt and cement, gravel and sand, and reduces the misjudgment rate, providing reliable terrain perception support for the autonomous navigation and motion control of robots in complex multi-terrain environments.
[0017] (2) By introducing a texture direction attention enhancement mechanism guided by confusion differences, refined attention weighting of the modulated feature map is achieved. This mechanism first uses the global feature differences between the target sample and the candidate confusion sample to generate channel attention weights, enhancing the feature channels related to confusion category discrimination; then, it extracts texture direction information through multi-directional convolution and adaptively weights and fuses it to highlight the most discriminative texture direction; then, it combines the channel-wise difference region map to locate the key confusion region in space; finally, it fuses the directional texture features with the difference region map to generate a spatial attention map, enabling the model to focus on the local position with the greatest difference between the two types of terrain in the spatial dimension. Through the combined effect of channel attention and spatial attention, the texture direction and spatial region related to the easy confusion category discrimination in the enhanced feature map are significantly strengthened, and irrelevant information is effectively suppressed, thus providing a more discriminative feature expression for subsequent sample pairing and classification, directly improving the model's fine-grained discrimination ability for highly similar terrains such as asphalt and cement.
[0018] (3) The comprehensive confusion level is calculated by integrating information from five dimensions: feature distance, classification confidence, misclassification probability, class risk weight, and attention difference response, overcoming the limitations of relying on a single metric for sample pairing. This comprehensive metric can more comprehensively assess the confusion risk between two samples, avoiding mispairing caused by accidental feature similarity or local uncertainty of the classifier. This makes the selected high-risk confusion sample pairs closer to the model's true confusion boundary, thus providing a more accurate and effective pairing basis for subsequent semantic-guided texture transfer enhancement, and significantly improving the targeting of enhanced samples to improve the model's fine-grained discrimination ability.
[0019] (4) Through the collaborative work of texture localization, texture transformation, and semantic fusion branches, precise localization, adaptive transformation, and semantically preserved injection of discriminative local textures in high-risk confusion samples are achieved. This approach differs from traditional random mixing or region replacement enhancement methods. Instead, it extracts and transfers local texture patterns that lead to misjudgment based on the current confusion state of the model, while strictly controlling the injection location and intensity to ensure that the semantics of the target category are not destroyed. The resulting confusion enhancement samples are highly targeted and realistic, effectively strengthening the model's discriminative ability near the boundaries of easily confused categories, and providing high-quality enhancement data for subsequent joint training.
[0020] (5) By constructing a total loss function containing five loss terms, the model is optimized collaboratively from multiple dimensions. The original sample classification loss maintains the overall classification performance, the confusion-enhanced sample classification loss strengthens the discrimination of easily confused boundaries, the feature rejection loss widens the distance between confused sample pairs in the feature space, the attention focus loss guides attention to the difference region, and the semantic preservation constraint loss ensures the semantic integrity of the enhanced samples. At the same time, the dynamic weight adjustment mechanism enables the model to adaptively adjust the optimization focus according to the current confusion state, forming a closed-loop iterative process of discovering confusion, generating samples, optimizing the model, and re-evaluating. This joint training scheme effectively improves the model's fine-grained discrimination ability for highly similar terrain categories such as asphalt and cement, and gravel and sand, reduces the misclassification rate, and enhances the adaptability and stability of the model in different training stages. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a robot terrain recognition method for easily confused terrain categories disclosed in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a robot terrain recognition system for easily confused terrain categories disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the device structure of the hardware operating environment of the electronic device disclosed in the embodiments of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0024] like Figure 1 As shown in the figure, this application discloses a robot terrain recognition method for easily confused terrain categories, including steps S1-S8.
[0025] Step S1: Obtain and preprocess the terrain image ahead of the robot's travel path to obtain the preprocessed terrain image.
[0026] Specifically, in this embodiment, a vision sensor installed at the front of the robot collects images of the ground ahead of the travel path, which can reflect the surface material condition that the robot is about to pass through in real time.
[0027] In this step, a robot terrain image dataset is constructed, and samples are labeled according to terrain categories. These terrain categories include, but are not limited to, asphalt, cement, sand, and snow. Let the input terrain image be... Then a single RGB image can be represented as .
[0028] in, Indicates RGB three channels, and These represent the height and width of the original image, respectively. To meet the network input requirements, the original image undergoes size normalization, for example, by scaling the image to [size value missing]. Then cut from the center ,get: .
[0029] The images are then normalized to ensure that different images maintain consistent channel mean and variance. The normalization process can be represented as follows: .
[0030] in, This represents the mean value of each color channel. This represents the standard deviation of each color channel. This represents the normalized input image. and The value of is obtained by statistically analyzing the training set images, specifically by calculating the pixel mean and standard deviation of all images in the training set across the R, G, and B channels. This normalization operation eliminates color shifts caused by differences in lighting conditions and camera parameters between different images, making the distribution of input data more stable and beneficial for the subsequent training and convergence of the network.
[0031] To improve the model's adaptability to changes in shooting angle, lighting, partial occlusion, and minor noise perturbations, data augmentation operations such as random horizontal flipping, random rotation, brightness perturbation, and contrast perturbation can be performed during the training phase. After these processing steps, the image is converted into tensor form and used as input to the subsequent backbone feature extraction network.
[0032] Step S2: Input the preprocessed terrain image into the backbone feature extraction network to extract the depth feature map, wherein the preprocessed terrain image is used as the target sample.
[0033] The backbone network can be a residual convolutional network with the final classification layer removed, or it can be replaced with other convolutional neural networks capable of image feature extraction. Taking ResNet-50 as an example, the input image is processed through multiple convolutions, batch normalization, ReLU activation function, and residual connections to obtain a high-dimensional terrain feature map.
[0034] Let the backbone network be Then its feature extraction process can be expressed as: .
[0035] in, For the first The depth feature map corresponding to the image. For a size of Given the input image, the typical feature map size output by the backbone network is: .
[0036] The feature map has 2048 channels, with a spatial height and width of 7. These 2048 channels correspond to different semantic and textural feature responses; for example, some channels may be sensitive to granular textures, while others may be sensitive to edge or linear structures. The 7×7 spatial resolution preserves the relative positional information of different regions in the image, allowing for refined operations in the spatial dimension during subsequent texture modulation and attention enhancement.
[0037] This deep feature map simultaneously contains global semantic information and local texture information of the terrain image. Global semantic information reflects the overall category attributes of the terrain, such as the overall tone and texture of asphalt pavement; local texture information depicts microscopic differences, such as the fine particle distribution of asphalt and the slab-like texture structure of cement. This feature map serves as the common foundation for subsequent multi-scale texture modulation, perplexing difference attention weighting, hard sample matching, and classification prediction. Using the preprocessed terrain image as the target sample means that each input image participates in the computation as the subject to be analyzed and enhanced in subsequent steps.
[0038] Step S3: Extract the texture response map of each branch from the depth feature map through multiple depth-separable convolutional branches of different scales. Calculate the branch weights based on the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map.
[0039] Specifically, convolutional kernels of different scales can capture texture information of different frequencies and granularities. For example, small-scale convolutional kernels are suitable for extracting fine-grained textures, while large-scale convolutional kernels are suitable for extracting plate-like or block-like textures. Depthwise separable convolution reduces computational overhead while maintaining feature extraction capabilities. By calculating the texture energy distribution of the texture response maps of each branch, it is possible to dynamically assess which scale of texture dominates in the current terrain image, and then assign corresponding weights to branches of different scales. This texture energy-driven adaptive weighted fusion method enhances texture responses related to easily confused category discrimination in the modulated feature map, while suppressing noise textures unrelated to classification, providing a more discriminative feature representation for subsequently distinguishing highly similar terrains such as asphalt and cement.
[0040] Step S4: Obtain candidate confused samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confused samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and spatial attention to the modulated feature map to obtain the enhanced feature map.
[0041] Specifically, candidate confusion samples are terrain images that belong to different categories from the target sample but are visually easily confused. By calculating the differences between the two at the feature level, key discriminative information that can distinguish between the two categories can be obtained. Channel attention generated based on this feature difference can enhance feature channels related to the boundaries of the confusion categories, while spatial attention can focus on the local regions with the greatest differences between the two types of terrain. Applying these two attention methods simultaneously to the modulated feature map allows the model to actively focus on the texture directions and locations that truly help distinguish the categories when processing easily confused terrain, rather than processing the entire image uniformly. This attention mechanism based on confusion sample pairs significantly improves the model's fine-grained discrimination ability for highly similar terrain.
[0042] Step S5: Pool the enhanced feature map to obtain feature vectors, calculate the overall confusion degree between the target sample and each candidate confusion sample based on the feature vectors, and select high-risk confusion samples for each target sample according to the overall confusion degree to form a high-risk confusion sample pair.
[0043] Specifically, pooling operations compress the spatial dimension into a one-dimensional feature vector, facilitating distance measurement and comparison between samples. The overall confusion factor comprehensively reflects information from multiple dimensions, including the proximity of the target sample and candidate confused samples in the feature space, the uncertainty of the model's classification of the target sample, and the probability of a candidate confused sample being misclassified as the target category. High-risk confused samples selected based on the overall confusion factor represent the sample pairs most easily confused by the model. Targeted enhancement of these samples directly strengthens the model's discriminative ability near the confusion boundary. This step achieves precise selection of the most valuable training samples from a large number of candidate samples, avoiding ineffective enhancements that might result from random pairing or simple nearest neighbor pairing.
[0044] Step S6: Input the target sample and the high-risk confusion sample into the semantic-guided texture transfer subnetwork to generate a transfer mask and extract local texture patterns from the high-risk confusion sample. Use the transfer mask to determine the injection position on the non-subject region of the target sample and superimpose the local texture pattern onto the injection position to obtain a confusion-enhanced sample that maintains the semantics of the target category.
[0045] Specifically, the semantically guided texture transfer subnetwork first generates a transfer mask through a texture localization branch. This transfer mask marks which local texture regions in high-risk confusion samples are most likely to cause misclassification. The texture localization branch generates the transfer mask based on the feature map of the target sample, the feature map of the high-risk confusion sample, and the difference between the two. Each pixel value of the transfer mask represents the intensity of the local texture pattern at that location being transferred. The closer the value is to 1, the more likely the texture pattern at that location is to cause confusion and should be transferred to the target sample.
[0046] The texture transformation branch extracts these key texture patterns from high-risk obfuscated samples and performs necessary normalization transformations to adapt to the brightness and scale of the target samples. The extracted local texture patterns correspond to image patches or feature blocks covered by high-response regions in the transfer mask. The normalization transformation adjusts the mean and variance of the texture patterns to match the statistical characteristics of the corresponding regions of the target samples, avoiding unnatural enhancement of samples due to differences in lighting or viewing angle.
[0047] The semantic fusion branch overlays the transformed texture patterns onto the non-subject regions of the target sample, controlling the position and intensity with a transfer mask, ensuring that the main semantics of the target sample are not compromised. Injection locations are determined in the non-subject regions of the target sample using a transfer mask. Regions with higher median values in the transfer mask correspond to texture locations in high-risk obfuscation samples that are prone to misclassification; these locations are identified as injection sites. The transformed local texture patterns are then overlaid onto these injection sites while preserving the main region, thereby generating obfuscation-enhanced samples.
[0048] The generated ambiguity-enhanced samples simulate local texture ambiguity in real-world scenarios, such as a cement repair area on an asphalt road. This allows the model to be exposed to targeted ambiguity perturbations while maintaining the semantics of the target category, thereby improving robustness to easily confused categories.
[0049] Step S7: Input the confused enhanced samples and the original samples into the classification network for joint training to obtain the trained terrain recognition model.
[0050] Specifically, the original samples are used to maintain the model's classification ability for regular terrain images, while the confused and enhanced samples are used to strengthen the model's ability to distinguish samples near the boundaries of highly similar categories. Joint training allows the model to be optimized specifically for easily confused categories without affecting its overall classification performance, thereby significantly reducing the misclassification rate between highly similar categories such as asphalt and cement, and gravel and sand, while maintaining the accuracy of recognition for other categories.
[0051] Step S8: Input the terrain image to be identified into the trained terrain recognition model and output the terrain category recognition result. This step applies the trained model to actual robot perception tasks. The output terrain category information can be directly used to guide downstream control tasks such as robot path planning, motion mode switching, and speed adjustment, realizing a complete closed loop from image acquisition to terrain recognition.
[0052] The method disclosed in this application solves the technical problems of low accuracy and high misjudgment rate in the prior art for robots to identify easily confused terrain categories with highly similar visual features by combining multi-scale terrain texture perception modulation, texture direction attention guided by confusion difference, active pairing of comprehensive confusion degree, semantically guided texture transfer enhancement, and joint training. It effectively improves the model's fine-grained discrimination ability for easily confused terrain categories such as asphalt and cement, and gravel and sand, and reduces the misjudgment rate, providing reliable terrain perception support for the robot's autonomous navigation and motion control in complex multi-terrain environments.
[0053] In some embodiments, the present application defines a specific method for calculating branch weights in step S3 based on the texture energy distribution of each texture response map, namely, performing high-frequency texture energy estimation on each texture response map to obtain the texture energy value of each branch, and mapping the texture energy value to the weight of each branch through the Softmax function.
[0054] Specifically, a multi-scale terrain texture-aware modulation module is used to enhance the texture of the depth feature map. This module contains multiple depthwise separable convolutional branches at different scales, each responsible for extracting texture responses at a specific scale. Taking three scales as an example, the scale set can be set as follows: These correspond to small, medium, and large receptive fields, respectively. Depthwise separable convolution decomposes standard convolution into channel-wise convolution and point-wise convolution, significantly reducing the number of parameters and computational overhead while maintaining feature extraction capabilities.
[0055] For each scale branch, the depth feature map is first processed by a depthwise separable convolution at that scale. Perform convolution to obtain the texture response map of this branch. Denote the scale. s The depthwise separable convolution operator is Then the texture response map of each branch can be represented as These texture response maps reflect the texture structure information of the terrain surface at different scales. For example, the 3×3 branch is sensitive to fine granular textures, while the 7×7 branch is sensitive to larger plate-like or blocky textures.
[0056] To evaluate the contribution of each scale branch to the current terrain image, high-frequency texture energy is estimated for each texture response map. High-frequency texture energy reflects the richness of local abrupt changes and detail information in the texture response map. For each scale... s The texture response map is first calculated by calculating its local average response. The local average response can be obtained by average pooling the texture response map in the spatial dimension, and is used to estimate the baseline response level at this scale. Then, the absolute value of the difference between the texture response map and the local average response is calculated to obtain the high-frequency components. .
[0057] The absolute value of this difference highlights localized abrupt changes in the texture response map that deviate from the average level, indicating where the true texture details reside. Subsequently, global average pooling is performed on this absolute value to obtain the texture energy value for this scale branch: .
[0058] GAP stands for Global Average Pooling, which compresses the spatial dimension into a scalar value. The larger the value, the richer the high-frequency texture information detected by that scale branch in the current image, and the more important that scale is for distinguishing terrain materials. In this way, each branch obtains a scalar texture energy value to characterize the texture activity at that scale.
[0059] After obtaining the texture energy values of each branch, they need to be converted into weights for each branch. The texture energy values of all branches are then combined into a vector, denoted as... Introduce a learnable mapping matrix. This matrix maps the texture energy vector to unnormalized weight scores: .
[0060] in, The network has a 3×3 dimension (assuming three branches), and its parameters are automatically learned during training via backpropagation. In this way, the network can adaptively adjust the importance of each branch based on the texture energy value, rather than using fixed preset weights. Finally, the scores are mapped to probabilistic weights using the Softmax function. .
[0061] The Softmax function ensures that the sum of the weights of all branches is 1, and that each weight is between 0 and 1. This means that branches with higher texture energy will have a larger weight, thus dominating the subsequent weighted fusion process.
[0062] After obtaining the weights of each branch, the texture response maps of each branch are weighted and fused. The fusion process can be represented as: .
[0063] in, This represents the texture modulation intensity for the current sample. It is a learnable scalar parameter that can be initially set to 0.2 and adjusted through end-to-end optimization during training. This is the modulated feature map, its size and depth feature map. To maintain consistency, this fusion method adds the original feature map to the weighted multi-scale texture response, preserving the global semantic information in the original feature map while enhancing the texture response related to easily confused category discrimination.
[0064] Through the aforementioned high-frequency texture energy estimation and Softmax mapping mechanism, the multi-scale terrain texture perception modulation module can dynamically adjust the contribution of each scale branch according to the actual texture features of the input terrain. This enables the model to automatically enhance the response to the discriminative texture scale when dealing with terrains of similar height, such as asphalt and cement, while suppressing irrelevant noise. This provides a more discriminative feature basis for subsequent attention enhancement and sample pairing.
[0065] In some embodiments, step S4 includes five sub-steps S41 to S45, each of which is described in detail below.
[0066] Step S41: Perform global average pooling on the feature map of the target sample and the feature map of the candidate confused sample respectively to obtain their respective feature vectors, and calculate the absolute value of the difference between the two as the difference vector.
[0067] Let the feature map of the target sample after multi-scale texture modulation be... The corresponding feature map of the candidate confused sample is We then perform global average pooling on both, compressing the spatial dimension into a one-dimensional vector: , .
[0068] Here, GAP stands for Global Average Pooling, which is an average value taken over the height and width dimensions of the pool. and Both are feature vectors of length equal to the number of channels (e.g., 2048), representing the average feature responses of the target sample and the candidate confusion sample globally, respectively. Then, the absolute value of the difference between the two is calculated to obtain the difference vector: .
[0069] This difference vector reflects the degree of difference in the responses of the two samples across various feature channels. Channels with large differences may be key channels for distinguishing between the two types of terrain. Through this step, difference information is obtained for subsequent channel attention generation.
[0070] Step S42: After concatenating the difference vector with the feature vector of the target sample, the convolution is performed using a one-dimensional convolution and the Sigmoid function to generate channel attention weights. The channel attention weights are then multiplied with the feature map of the target sample to obtain a channel-weighted feature map.
[0071] Specifically, the difference vector The feature vector of the target sample Concatenate the vectors to obtain the concatenated vector. The concatenated vector is input into a one-dimensional convolutional layer for feature mapping. The kernel size of the one-dimensional convolution is 1, and the number of output channels is the same as the number of channels in the target sample feature map. Then, the output values are mapped to between 0 and 1 using the sigmoid function to obtain the channel attention weights. .
[0072] in, This represents the Sigmoid function, and Conv1D represents a one-dimensional convolution operation. Its parameters are learned through backpropagation during training. Channel attention weights. It is a vector of length *number* channels, where each element represents the importance of the corresponding channel. This weight is then compared with the feature map of the target sample. Multiplying each channel sequentially yields a channel-weighted feature map: .
[0073] in, This represents element-wise multiplication, where the feature map of each channel is multiplied by its corresponding weight value. This operation enhances feature channels relevant to confusing category discrimination while suppressing irrelevant or interfering channels.
[0074] Step S43: Perform convolution processing on the channel weighted feature map in the horizontal, vertical and diagonal directions to obtain the texture response in each direction. Generate directional weights through the Softmax function based on the global average pooling value of the texture response in each direction, and perform weighted fusion on the texture response in each direction according to the directional weights to obtain the directional texture feature map.
[0075] In this embodiment, considering that terrain textures often have directionality, such as tire tracks, crack directions, and stone arrangement, four directional convolution branches are set, with the direction set as follows: For each direction Use the corresponding directional convolution operator Channel weighted feature map Perform convolution processing to obtain the directional texture response: .
[0076] Directional convolution operators can be implemented using convolution kernels of specific shapes, such as a 1×k kernel for the horizontal direction, a k×1 kernel for the vertical direction, and a sparse kernel for the diagonal direction. The parameters of these convolution kernels are learned during training.
[0077] To determine the importance of each direction, global average pooling is performed on the texture response for each direction to obtain the global statistics for that direction: .
[0078] Then, the global statistics for all directions are input into a lightweight mapping function. This function can be a fully connected layer or a convolutional layer, outputting unnormalized orientation scores. Orientation weights are generated using the Softmax function. .
[0079] Directional weights This reflects which direction of texture is most discriminative in the current terrain image. Finally, the texture responses of each direction are weighted and fused according to their direction weights to obtain the directional texture feature map. : .
[0080] This directional texture feature map integrates texture information from various directions, and the weights are adaptively determined by the data, which can highlight the texture directions that are related to confusing category discrimination.
[0081] Step S44: Calculate the sum of the absolute values of the channel-by-channel differences between the feature maps of the target sample and the candidate confused samples, and obtain the difference region map after normalization.
[0082] Feature map of target sample Feature maps of candidate confused samples The absolute value of the channel difference at each spatial location is calculated along the channel dimension, and then summed over all channels to obtain a two-dimensional difference map. .
[0083] in, c For channel indexing, Indicates the first c The spatial feature map of each channel. This difference map reflects the overall degree of difference between the two samples at each spatial location; regions with large differences may be key areas causing confusion. The difference map is then normalized to obtain a difference region map. : .
[0084] The normalization operation scales the values of the difference map to between 0 and 1, making it easier to compare and merge with the attention map later.
[0085] Step S45: Concatenate the average pooling map and max pooling map of the directional texture feature map with the difference region map, generate a spatial attention map through two-dimensional convolution and the Sigmoid function, and multiply the spatial attention map with the channel weighted feature map to obtain the enhanced feature map.
[0086] First, the directional texture feature map Performing average pooling and max pooling along the channel dimension respectively yields two two-dimensional feature maps: and The average pooling plot reflects the average response intensity at each spatial location, while the max pooling plot highlights the location of the strongest response. These two pooling plots are then compared with the difference region plot. By concatenating along the channel dimension, a three-channel concatenated feature map is obtained: .
[0087] The concatenated feature map is input into a two-dimensional convolutional layer, typically with a kernel size of 3×3 or 7×7, and outputs a single-channel spatial attention map. Then, the output value is mapped to a range of 0 to 1 using the sigmoid function to obtain the spatial attention map. : .
[0088] Here, Conv2D represents a two-dimensional convolution operation, whose parameters are learned during training. Spatial attention map. Each pixel value represents the importance of that spatial location; a larger value indicates that the location is more noteworthy. Finally, the spatial attention map is combined with the channel-weighted feature map. Element-wise multiplication yields the enhanced feature map: .
[0089] This operation allows the model to focus spatially on the region where the target sample and the candidate confused sample differ the most, while preserving the discriminative channel information enhanced by channel attention. This enhanced feature map integrates multi-scale texture modulation, channel difference attention, and spatial orientation attention, providing high-quality feature representations for subsequent sample pairing and classification.
[0090] By introducing a texture orientation attention enhancement mechanism guided by confusion differences, refined attention weighting of the modulated feature map is achieved. This mechanism first uses the global feature differences between the target sample and candidate confusion samples to generate channel attention weights, enhancing feature channels related to confusion category discrimination. Then, it extracts texture orientation information through multi-directional convolution and adaptively weights and fuses it to highlight the most discriminative texture direction. Next, it combines channel-wise difference region maps to locate key confusion regions in space. Finally, it fuses the directional texture features with the difference region map to generate a spatial attention map, enabling the model to focus on the local locations with the greatest differences between the two terrain types in the spatial dimension. Through the combined effect of channel attention and spatial attention, the texture orientation and spatial regions related to easily confused category discrimination in the enhanced feature map are significantly strengthened, while irrelevant information is effectively suppressed. This provides more discriminative feature representations for subsequent sample pairing and classification, directly improving the model's fine-grained discrimination ability for highly similar terrains such as asphalt and cement.
[0091] In some embodiments, this application provides a specific method for calculating the overall confusion degree between the target sample and each candidate confused sample, namely, obtaining five indicators and weighting and summing the five indicators to obtain the overall confusion degree.
[0092] The five metrics include: the feature distance between the target sample and the candidate confused sample, the confidence level of the target sample being correctly predicted by the classifier, the probability that the candidate confused sample is misclassified as the target category, the inter-category risk weight determined based on the pre-calculated confusion matrix, and the similarity response of the target sample and the candidate confused sample in the attention difference region. The attention difference region is obtained by normalizing the channel-by-channel difference between the feature maps of the target sample and the candidate confused sample. The following details the acquisition method of each metric and the calculation process of the overall confusion score.
[0093] The first metric is the feature distance between the target sample and the candidate confused sample. In step S5, global average pooling is performed on the enhanced feature map to obtain the feature vector, and the feature vector of the target sample is denoted as... The feature vector of the candidate confused sample is Feature distance The feature distance is used to measure how close two samples are in the feature space. Metrics such as Euclidean distance or cosine distance can be used. The smaller the feature distance, the closer the two samples are in the feature space, and the easier it is for the model to confuse them. This distance value directly reflects the feature similarity between samples and is a fundamental indicator for judging the risk of confusion.
[0094] The second metric is the confidence level that the classifier correctly predicts the target sample. The feature vector of the target sample is input into the classifier to determine if the sample belongs to its true class. confidence level This confidence level reflects the model's certainty about the classification result of the current target sample. The lower the confidence level, the less certain the model is about classifying the sample, and the more likely the sample is to be located near the class boundary and easily confused with other classes.
[0095] The third metric is the probability that a candidate confusing sample is misclassified as the target class. The feature vector of the candidate confusing sample is input into the classifier to obtain the prediction that the sample is classified as the target class. probability The higher this probability value, the more similar the candidate confusing sample is to the target category in appearance or features, and the more likely the model is to misclassify it as the target category. This metric directly measures the degree of confusion threat posed by the candidate confusing sample to the target category.
[0096] The fourth metric is the inter-category risk weight determined based on a pre-calculated confusion matrix. During training, a confusion matrix can be constructed by analyzing historical classification results. The rows of the confusion matrix represent the true categories, and the columns represent the predicted categories. Each element in the matrix represents the number or proportion of times a sample with the true category 'a' is predicted as category 'b'. Based on this confusion matrix, the risk weight between the target category 'a' and the confused category 'b' can be determined. The risk weight can be set as the misclassification rate between the two classes, or it can be assigned a higher weight according to the needs of the robot control task. For example, when asphalt is misclassified as cement, which may lead to serious motion control errors, the weight can be appropriately increased.
[0097] The fifth metric is the similarity response between the target sample and the candidate confused sample in the attention difference region. The attention difference region is obtained by normalizing the sum of the absolute values of the channel-by-channel differences between the feature maps of the target sample and the candidate confused sample; the specific calculation method has been described in step S44. Within this difference region, the similarity of the feature responses of the target sample and the candidate confused sample is calculated, for example, by calculating the cosine similarity or Pearson correlation coefficient of the feature vectors of the two samples within the difference region. This similarity response reflects whether the local features of the two samples are still close within the difference region that the model is most concerned with. If they are still highly similar within the difference region, it indicates that the two samples are highly likely to be confused.
[0098] After obtaining the above five indicators, assign corresponding weight coefficients to each. , , , and The overall confusion level is obtained by weighting and summing the five indicators: .
[0099] in, This represents similar responses in regions of attention difference. Each weight coefficient is an adjustable hyperparameter, which can be set according to the characteristics of different terrain datasets, or optimized during training through cross-validation. Overall confusion. The larger the value, the higher the risk of confusion between the target sample and the candidate confused sample, and the more worthy the sample pair is to be selected as a high-risk confused sample pair for subsequent texture transfer enhancement.
[0100] The above technical solution calculates the overall confusion level by integrating information from five dimensions: feature distance, classification confidence, misclassification probability, class risk weight, and attention difference response. This overcomes the limitations of relying on a single metric for sample pairing. This comprehensive metric can more fully assess the confusion risk between two samples, avoiding mispairing caused by accidental feature similarity or local uncertainty in the classifier. This ensures that the selected high-risk confusion pairs are closer to the model's true confusion boundary, thus providing a more accurate and effective pairing basis for subsequent semantic-guided texture transfer enhancement. This significantly improves the targeted enhancement of samples to improve the model's fine-grained discriminative ability.
[0101] In some embodiments, after calculating the overall confusion between the target sample and each candidate confused sample, each target sample corresponds to a set of overall confusion scores, and each candidate confused sample has a corresponding overall confusion value. Overall Confusion Reflects the target sample Confusion samples with candidates The higher the value, the more likely the model is to confuse the two samples, and the more worthy the sample pair is to be processed first.
[0102] For each target sample, its overall confusion score compared to all candidate confusion samples is ranked from highest to lowest. One or more candidate confusion samples with the highest overall confusion score are selected as high-risk confusion samples for that target sample. The specific number of samples selected can be set according to actual application needs; for example, the top three samples in terms of overall confusion score can be selected, or a threshold can be set to select all samples with an overall confusion score exceeding that threshold. The number of selected samples determines the number of subsequent confusion enhancement samples generated, requiring a balance between enhancement effect and computational cost.
[0103] When multiple high-risk obfuscation samples are selected, each high-risk obfuscation sample forms a high-risk obfuscation sample pair with the target sample. These sample pairs are then input into the subsequent semantically guided texture transfer sub-network to generate corresponding obfuscated enhanced samples. In this way, a single target sample can generate multiple enhanced samples with different obfuscated texture patterns, further enriching the diversity of the training data and enabling the model to be exposed to more types of obfuscation scenarios.
[0104] This selection method is directly based on the ranking results of the overall confusion level, which is simple to operate and highly interpretable. The sample pairs with the highest overall confusion level are precisely the sample pairs that the model is currently most likely to confuse. Targeted enhancement of these sample pairs can directly address the weak links in the model's classification boundary, allowing limited enhancement resources to be invested in the areas that most need improvement, thereby maximizing the efficiency of enhancement training.
[0105] In some embodiments, in step S6, the semantically guided texture transfer subnetwork includes a texture localization branch, a texture transformation branch, and a semantic fusion branch. The texture localization branch generates a transfer mask based on the feature map of the target sample, the feature map of the high-risk confusion sample, and the difference between the two; the texture transformation branch extracts local texture patterns from the high-risk confusion sample and performs normalization transformation; the semantic fusion branch superimposes the transformed local texture patterns onto the non-subject region of the target sample with the position and intensity controlled by the transfer mask. Each branch is described in detail below.
[0106] The texture localization branch generates a migration mask, which identifies which local texture regions in high-risk confusion samples are most likely to cause model misclassification. Examples include cement slab texture, fine asphalt particles, crack orientation, or localized reflective areas. Let the feature map of the target sample be... The feature map of high-risk confusion samples is The difference between the two is These three elements are concatenated along the channel dimension to obtain the concatenated feature map: .
[0107] The concatenated feature map is input into a lightweight convolutional mapping function. This function consists of several convolutional and activation layers, whose parameters are learned through backpropagation during training. The output of the convolutional mapping function is processed by the sigmoid function, mapping the value of each pixel to between 0 and 1, thus obtaining the transfer mask. .
[0108] in, Represents the Sigmoid function. (Transfer mask) Spatial dimensions and characteristics Figure 1 Each pixel value represents the intensity of the local texture region at that location being transferred. A value closer to 1 indicates that the texture pattern at that location is more likely to cause confusion and should be transferred to the target sample; a value closer to 0 indicates that the texture pattern at that location is not related to confusion and should not be transferred. In this way, the texture localization branch can automatically identify the local texture regions in high-risk confusion samples that contribute the most to misjudgment, providing precise spatial guidance for subsequent texture extraction and injection.
[0109] The texture transformation branch is used to extract local texture patterns to be transferred from high-risk confusion samples and perform normalization transformations to adapt to the brightness and scale features of the target samples. First, local texture patterns are extracted from the image space or feature space of the high-risk confusion samples. This local texture pattern corresponds to the image patch or feature patch covered by the high-response region in the transfer mask. Extraction methods can include directly cropping the pixel values of the corresponding region or decoding the texture information from the feature map.
[0110] After extracting the local texture pattern, it needs to be normalized to eliminate differences in brightness, contrast, and scale between the target sample and high-risk confounding samples. The normalization process adjusts the local statistical characteristics of the target sample; for example, adjusting the mean and variance of the extracted texture pattern to match the mean and variance of the corresponding region in the target sample, or using histogram matching to ensure the brightness distribution of the texture pattern is consistent with the target sample. The transformed texture pattern is denoted as... It retains the discriminative features of high-risk confused samples in terms of texture structure, but is compatible with the target sample in terms of brightness and scale, thereby avoiding unnatural enhancement of samples due to differences in lighting or viewing angle.
[0111] The semantic fusion branch injects the transformed local texture patterns into the target samples in a controlled manner, generating obfuscated enhanced samples. This branch receives the original image of the target samples. migration mask and the transformed texture mode As input, the injection location and intensity are first determined using a transfer mask: regions with higher median values in the transfer mask correspond to texture locations in high-risk confusion samples that are prone to misclassification; these locations will be used to replace or overlay texture patterns. The injection operation is performed in non-subject regions of the target sample to ensure that the overall semantics of the target category are not compromised.
[0112] The injection process can be represented as: .
[0113] Here, ⊙ represents element-wise multiplication. This represents a texture adaptation function used to adapt the transformed texture mode. With target sample Fusion is performed at the injection site. The texture adaptation function can be a simple direct replacement, or a more complex operation such as weighted superposition or Poisson fusion. The goal is to ensure that the injected texture pattern blends naturally with the surrounding background, avoiding obvious stitching artifacts. Through this operation, the non-subject regions of the target sample are injected with local texture patterns from high-risk obfuscated samples, while the subject regions remain unchanged, thus generating obfuscated enhanced samples that preserve the semantics of the target category. .
[0114] The generated obfuscated augmented samples simulate real-world scenarios where the main terrain is consistent but local textures are contaminated or mixed with obfuscated textures. Examples include a small cement patch on an asphalt road or a few grains of sand mixed in with gravel. These types of samples frequently occur in real-world environments but are difficult to generate using conventional data augmentation methods. By semantically guiding texture transfer, the model can be exposed to these targeted obfuscated situations during the training phase, thereby improving its robustness to easily obfuscated categories.
[0115] By employing the aforementioned technical solution, and through the collaborative work of texture localization, texture transformation, and semantic fusion branches, precise localization, adaptive transformation, and semantically preserved injection of discriminative local textures in high-risk confusion samples are achieved. Unlike traditional random mixing or region replacement enhancement methods, this approach extracts and transfers local texture patterns that lead to misjudgments based on the model's current confusion state, while strictly controlling the injection location and intensity to ensure that the target category's semantics are not compromised. The resulting enhanced confusion samples are highly targeted and realistic, effectively strengthening the model's discriminative ability near easily confused category boundaries, and providing high-quality augmentation data for subsequent joint training.
[0116] In some embodiments, after the semantic fusion branch overlays the transformed local texture pattern onto the non-subject region of the target sample, a fuzzy enhanced sample is obtained. To ensure that the augmented sample still belongs to the target class, and that the class semantics are not altered due to texture transfer, a semantic preservation constraint needs to be imposed. Specifically, the obfuscated augmented sample is input into the current classification network, and the classifier's prediction of its class affiliation is obtained. The classifier outputs a probability distribution vector, where each element represents the probability that the sample belongs to the corresponding class. The sample is then identified by the classifier as its original target class. The probability of is denoted as . The probability value ranges from 0 to 1. The closer it is to 1, the better the semantics of the enhanced sample are preserved. The closer it is to 0, the more severe the semantic shift caused by texture transfer.
[0117] Based on this probability, a semantic preservation constraint loss is constructed. In this embodiment, the approach used is to employ negative log-likelihood loss, i.e. .
[0118] when When the probability is close to 1, the loss value is close to 0; when the probability is low, the loss value is large, resulting in a large gradient in backpropagation. This prompts the network to adjust the intensity and location of texture transfer, allowing the semantics of the enhanced samples to revert to the target category. Another approach is to use cross-entropy loss, which is equivalent in form to the above. This loss term is optimized together with the classification loss, feature rejection loss, and attention focusing loss during training. Its weight coefficient λ can be set experimentally, for example, to 0.1 or 0.2, to balance the relationship between semantic preservation and other training objectives.
[0119] By introducing a semantic preservation constraint loss, the texture transfer enhancement process is no longer a blind texture overlay, but rather performed while preserving the semantics of the target category. This loss term effectively prevents the enhanced sample's category from substantially changing due to excessive texture confusion caused by excessive transfer, thus avoiding the generation of misleading training data. Simultaneously, this loss term also influences the transfer mask generated by the texture localization branch, forcing the transfer mask to preferentially select regions that do not disrupt the main semantics of the target category for texture injection, further improving the quality of the enhanced samples.
[0120] In some embodiments, this application illustrates the specific structure and dynamic adjustment mechanism of the total loss function used in the joint training in step S7. The total loss function includes the classification loss of the original samples, the classification loss of the confused and enhanced samples, the feature rejection loss, the attention focus loss, and the semantic preservation constraint loss.
[0121] Among them, the feature rejection loss is calculated based on the difference between the feature distance and the dynamic interval between the target sample and the high-risk confused sample, and the dynamic interval is determined based on the comprehensive confusion degree; the attention focus loss is calculated based on the degree of overlap between the spatial attention map and the difference region map; the semantic preservation constraint loss is constructed based on the probability that the confused enhanced sample is judged as the target category by the classifier; and during the training process, the weight of the classification loss of the confused enhanced sample in the total loss is adaptively adjusted according to the ratio of the error rate between the target category and the confused category on the validation set to the initial error rate. The following is a detailed explanation of each loss and the dynamic adjustment mechanism.
[0122] During joint training, both the original samples and the scrambled / enhanced samples are input into the classification network for training. Let the classifier be... The prediction outputs for the original sample and the scrambled / enhanced sample are as follows: , .
[0123] in, This represents the predicted probability distribution of the original sample. This represents the predicted probability distribution of the confused augmented samples. After passing through a fully connected classification layer and a softmax function, the network output yields the predicted probabilities for various terrain categories, including asphalt, cement, sand, and snow. The original category scores output by the classification head are also represented. The Softmax prediction probability can be expressed as: .
[0124] in, express The original sample was predicted as the first... The probability of terrain-like features, The output of the classifier represents the first... kind , Indicates the total number of terrain categories. This represents the class index in the summation term. For the confusion-enhanced sample, its predicted probability can be expressed as: .
[0125] in, Indicates the first The confusion-enhancing sample was predicted as the _th _th The probability of terrain-like features, Indicates the confusion enhancement sample in the 1st Class-based categorization output.
[0126] Specifically, raw samples refer to training samples that have not undergone any augmentation. The raw samples are then input into the classification network to obtain their predicted probability distribution. , with real labels Calculate the classification loss of the original samples : .
[0127] in, B Indicates batch size, c Indicates category index, For the true label of one-hot encoding, when the sample i The true category is c The value is 1 if the condition is met, otherwise it is 0. For network prediction samples i Category c The probability of this loss term. This term is used to maintain the model's ability to classify regular terrain images, ensuring that the model maintains a high accuracy rate on the overall classification task.
[0128] The obfuscation-enhanced samples are generated in step S6 by injecting high-risk obfuscation sample local texture patterns into non-subject regions of the target sample. These obfuscation-enhanced samples are then input into the classification network to obtain their predicted probability distribution. The true label of the original target sample Calculate the classification loss of the confused augmented samples : .
[0129] Because the obfuscated augmented samples preserve the semantics of the target category, their true labels remain the original target category. This loss term forces the model to correctly identify the main category when encountering samples with obfuscated textures, thereby strengthening the model's ability to discriminate near the boundaries of easily obfuscated categories.
[0130] Feature repulsion loss is used to directly increase the distance between the target sample and the high-risk confusion sample in the feature space, making their feature distributions more separated. Let the feature vector of the target sample be... The feature vector of the corresponding high-risk confusion sample is Then feature rejection loss Defined as: .
[0131] in, This represents the Euclidean distance between two eigenvectors. The interval is dynamic, based on the overall confusion level. Confirmed, specifically: .
[0132] in, The base interval is a positive constant, such as 1.0. This is the interval adjustment coefficient, which can be set to, for example, 0.5; The overall confusion level of the sample pair is calculated in detail in the preceding embodiments. The dynamic interval design requires sample pairs with higher overall confusion levels to have larger feature distances, thereby forming stronger separation boundaries in the feature space. When the feature distance between two samples is less than the dynamic interval, the loss is positive, pushing the network to increase the distance between them; when the feature distance is already greater than the dynamic interval, the loss is zero, and no additional repulsive force is applied.
[0133] The attention focus loss is used to constrain the spatial attention map to cover the difference regions between the target sample and the candidate confused samples, ensuring that the learning objective of the attention mechanism remains consistent with the difference regions of the confused sample pairs. Let the spatial attention map be... The difference region map is as follows The attention focus loss is defined as: .
[0134] in, Represents spatial location coordinates, The numerator is a small positive constant to prevent the denominator from being zero. The numerator is the sum of the element-wise products of the spatial attention map and the difference region map, measuring their spatial overlap; the denominator is the sum of all elements in the difference region map, used for normalization. When the spatial attention map completely covers the difference region map, the loss value approaches 0; when the attention map completely deviates from the difference region, the loss value approaches 1. This loss term prompts the spatial attention module to actively focus on the local region with the greatest difference between the target sample and the candidate confused sample, thus enabling attention learning to directly serve the discrimination of the confused category.
[0135] The semantic preservation constraint loss, used to constrain the obfuscated augmented samples to preserve the semantics of the target category, has been described in detail in the foregoing embodiments. The obfuscated augmented samples... Inputting the data into a classification network yields the classification result, which is then assigned to the target category. probability The semantic preservation constraint loss is: .
[0136] This loss term ensures that the texture transfer operation does not destroy the main category semantics of the target sample, so that the generated enhanced sample contains both confusing texture perturbations and class attributes.
[0137] The classification loss during training is based on the confusion boosting samples. The weights in the total loss are not fixed, but adaptively adjusted based on the classification error rate of specific confused class pairs on the validation set. Let the first... t After one round of training, the target category on the validation set a With confusion category b The error rate between The initial error rate is Then the confusion enhances the loss weights Adaptive representation is: .
[0138] in, and These represent the minimum and maximum values of the weights, for example, they can be set to 0.1 and 1.0. To prevent small constants with a denominator of zero; The function restricts the calculation result to [ Within the range. When the error rate of a certain confusion category pair is high, As the number of samples increases, the model will focus more on the confusing augmented samples for that class pair; as the error rate decreases, Automatically reduced to avoid overfitting to a small number of augmented samples.
[0139] The total loss function is the weighted sum of the losses of each component: .
[0140] in, and The weight coefficients for feature rejection loss and attention focus loss can be adjusted according to the training rounds or the level of confusion in the validation set. The weighting coefficients for semantic preservation constraint loss are typically set to small values.
[0141] By constructing a total loss function containing five loss terms, the model is collaboratively optimized from multiple dimensions. The original sample classification loss maintains the overall classification performance, the confusion-enhancing sample classification loss strengthens the discrimination of easily confused boundaries, the feature rejection loss widens the distance between confused sample pairs in the feature space, the attention focus loss guides attention to the difference regions, and the semantic preservation constraint loss ensures the semantic integrity of the enhanced samples. Simultaneously, a dynamic weight adjustment mechanism enables the model to adaptively adjust its optimization focus based on the current confusion state, forming a closed-loop iterative process of discovering confusion, generating samples, optimizing the model, and re-evaluating. This joint training scheme effectively improves the model's fine-grained discrimination ability for highly similar terrain categories such as asphalt and cement, and gravel and sand, reduces the misclassification rate, and enhances the model's adaptability and stability at different training stages.
[0142] In step S8, the terrain image to be identified is input into the trained terrain recognition model. This model includes optimized backbone feature extraction network parameters, multi-scale terrain texture perception modulation module parameters, confusion-guided texture orientation attention module parameters, and classifier parameters. During the deployment phase, the confusion sample generation branch and closed-loop update branch are only enabled during offline training or periodic incremental training; only the feature extraction and classification branches are retained during the online inference phase to meet the real-time requirements of the robot's edge device.
[0143] For a terrain image to be identified, the same preprocessing steps as in step S1 are performed, including size standardization and normalization, to ensure that the distribution of the input data is consistent with the training data. The preprocessed image is then input into the trained terrain recognition model. The model sequentially passes through a backbone feature extraction network to extract depth feature maps, a multi-scale terrain texture perception modulation module for texture enhancement, a texture direction attention module guided by confusion differences for attention weighting, and global average pooling to obtain feature vectors. Finally, a classifier and a softmax function output the predicted probabilities of each terrain category. The category with the highest probability is the model's final recognition result for the terrain image.
[0144] The output terrain category recognition results can include specific terrain material categories such as asphalt, cement, sand, and snow. This recognition result can be further provided to the robot's motion control system for tasks such as path planning, landing area selection, wheel-foot movement mode switching, speed adjustment, and drivability assessment. For example, when the recognition result is asphalt, the robot can choose a higher driving speed and a smoother movement mode; when the recognition result is sand, the robot can reduce its speed and switch to a movement mode more suitable for soft ground to avoid getting stuck or slipping.
[0145] Step S8 applies the trained terrain recognition model to the actual robot perception task, realizing a complete reasoning process from image acquisition to terrain category output. This step fully utilizes the model parameters optimized in steps S3 to S7, enabling the model to stably distinguish highly similar terrain categories such as asphalt and cement, and gravel and sand in actual operation, providing reliable terrain perception information for the robot's autonomous navigation and motion control in complex multi-terrain environments.
[0146] See attached document Figure 2 As shown in the embodiments, this application also discloses a robot terrain recognition system for easily confused terrain categories, including: The image processing module is used to acquire and preprocess the terrain image in front of the robot's path to obtain the preprocessed terrain image. The feature extraction module is used to input the preprocessed terrain image into the backbone feature extraction network to extract depth feature maps; wherein, the preprocessed terrain image is used as the target sample; The multi-scale texture modulation module is used to extract the texture response map of each branch from the depth feature map through multiple depth-separable convolutional branches of different scales, calculate the branch weights according to the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map. The confusion difference attention module is used to acquire candidate confusion samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confusion samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and the spatial attention to the modulated feature map to obtain the enhanced feature map. The confusion calculation and pairing module is used to pool the enhanced feature map to obtain a feature vector, calculate the comprehensive confusion between the target sample and each candidate confusion sample based on the feature vector, and select a high-risk confusion sample for each target sample according to the comprehensive confusion, thus forming a high-risk confusion sample pair. The texture transfer enhancement module is used to input the target sample and the high-risk confusion sample into the semantic guidance texture transfer sub-network, generate a transfer mask and extract local texture patterns from the high-risk confusion sample, and use the transfer mask to control the injection position to superimpose the local texture patterns onto the non-subject region of the target sample to obtain a confusion enhancement sample that maintains the semantics of the target category. The joint training module is used to input the confused and enhanced samples and the original samples into the classification network for joint training to obtain a trained terrain recognition model; The recognition output module is used to input the terrain image to be recognized into the trained terrain recognition model and output the terrain category recognition result.
[0147] This application also provides an electronic device. (See attached document.) Figure 3 As shown, it includes a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the above-described robot terrain recognition method for easily confused terrain categories.
[0148] Specifically, the processor may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor may also include onboard memory for caching purposes. The processor may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.
[0149] Memory can be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, memory can include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, instruments, or propagation media. Specific examples of memory include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); and also random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0150] This application also provides a computer-readable medium storing a computer program that, when executed by a processor, implements the aforementioned robot terrain recognition method for easily confused terrain categories. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0151] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0152] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A robot terrain recognition method for easily confused terrain categories, characterized in that, The steps include the following: S1: Acquire and preprocess the terrain image ahead of the robot's path to obtain the preprocessed terrain image; S2: Input the preprocessed terrain image into the backbone feature extraction network to extract the depth feature map, wherein the preprocessed terrain image is used as the target sample; S3: Extract the texture response map of each branch from the depth feature map by using multiple depth-separable convolutional branches of different scales; calculate the branch weights based on the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map; S4: Obtain candidate confused samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confused samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and spatial attention to the modulated feature map to obtain the enhanced feature map; S5: Pool the enhanced feature map to obtain a feature vector, calculate the overall confusion degree between the target sample and each candidate confused sample based on the feature vector, and select a high-risk confused sample for each target sample according to the overall confusion degree to form a high-risk confused sample pair; S6: Input the target sample and the high-risk confusion sample into the semantically guided texture transfer subnetwork, generate a transfer mask and extract local texture patterns from the high-risk confusion sample, determine the injection position on the non-subject region of the target sample through the transfer mask, and superimpose the local texture patterns onto the injection position to obtain a confusion-enhanced sample that maintains the semantics of the target category. S7: Input the confused enhanced samples and the original samples into the classification network for joint training to obtain a trained terrain recognition model; S8: Input the terrain image to be identified into the trained terrain recognition model and output the terrain category recognition result.
2. The robot terrain recognition method for easily confused terrain categories as described in claim 1, characterized in that, In step S3, the branch weights are calculated based on the texture energy distribution of each texture response map. Specifically, this includes: performing high-frequency texture energy estimation on each texture response map to obtain the texture energy value of each branch, and mapping the texture energy value to the weight of each branch using the Softmax function.
3. The robot terrain recognition method for easily confused terrain categories as described in claim 1, characterized in that, Step S4 specifically includes: S41: Perform global average pooling on the feature maps of the target sample and the candidate confused sample respectively to obtain their respective feature vectors, and calculate the absolute value of the difference between the two as the difference vector. S42: After concatenating the difference vector with the feature vector of the target sample, the convolution is performed using a one-dimensional convolution and the Sigmoid function to generate channel attention weights. The channel attention weights are then multiplied with the feature map of the target sample to obtain a channel-weighted feature map. S43: Perform convolution processing on the channel weighted feature map in the horizontal, vertical and diagonal directions respectively to obtain the texture response in each direction. Generate directional weights through the Softmax function based on the global average pooling value of the texture response in each direction, and perform weighted fusion on the texture response in each direction according to the directional weights to obtain the directional texture feature map. S44: Calculate the sum of the absolute values of the channel-by-channel differences between the feature maps of the target sample and the candidate confused samples, and obtain the difference region map after normalization; S45: The average pooling map and max pooling map of the directional texture feature map are concatenated with the difference region map, and a spatial attention map is generated by two-dimensional convolution and the Sigmoid function. The spatial attention map is then multiplied with the channel weighted feature map to obtain the enhanced feature map.
4. The robot terrain recognition method for easily confused terrain categories as described in claim 1, characterized in that, In step S5, the overall confusion degree between the target sample and each candidate confused sample is calculated. Specifically, this includes: calculating the feature distance between the target sample and the candidate confused sample, the confidence that the target sample is correctly predicted by the classifier, the probability that the candidate confused sample is misclassified as the target category, the risk weight between categories determined according to the pre-statistical confusion matrix, and the similarity response between the target sample and the candidate confused sample in the attention difference region. These five indicators are used as indicators, and the overall confusion degree is obtained by weighted summation of the five indicators. The attention difference region is obtained by normalizing the channel-by-channel difference between the feature maps of the target sample and the candidate confused sample.
5. The robot terrain recognition method for easily confused terrain categories as described in claim 4, characterized in that, In step S5, when selecting high-risk confusion samples for each target sample based on the overall confusion degree, one or more samples with the highest overall confusion degree are selected from the candidate confusion samples as high-risk confusion samples.
6. The robot terrain recognition method for easily confused terrain categories as described in claim 1, characterized in that: In step S6, the semantically guided texture transfer subnetwork includes a texture localization branch, a texture transformation branch, and a semantic fusion branch. The texture localization branch generates the transfer mask based on the feature map of the target sample, the feature map of the high-risk confusion sample, and the difference between the two. The texture transformation branch extracts local texture patterns from the high-risk confusion sample and performs normalization transformation. The semantic fusion branch superimposes the transformed local texture patterns onto the non-subject region of the target sample at the position and intensity controlled by the transfer mask.
7. The robot terrain recognition method for easily confused terrain categories as described in claim 6, characterized in that: When superimposing the local texture pattern, the semantic fusion branch calculates the probability that the confusing and enhanced sample is identified as the target category by the classifier, and constructs a semantic preservation constraint loss based on the probability to constrain the confusing and enhanced sample to maintain the semantics of the target category.
8. The robot terrain recognition method for easily confused terrain categories as described in claim 1, characterized in that, The joint training in step S7 specifically includes: training using a total loss function that includes the classification loss of the original samples, the classification loss of the confusion-enhanced samples, the feature rejection loss, the attention focus loss, and the semantic preservation constraint loss; the feature rejection loss is calculated based on the difference between the feature distance and the dynamic interval between the target sample and the high-risk confusion sample, and the dynamic interval is determined based on the comprehensive confusion degree; the attention focus loss is calculated based on the degree of overlap between the spatial attention map and the difference region map; the semantic preservation constraint loss is constructed based on the probability that the confusion-enhanced sample is classified as the target category by the classifier; and, during training, the weight of the classification loss of the confusion-enhanced samples in the total loss is adaptively adjusted based on the ratio of the error rate between the target category and the confusion category on the validation set to the initial error rate.
9. A robot terrain recognition system for easily confused terrain categories, characterized in that, include: The image processing module is used to acquire and preprocess the terrain image in front of the robot's path to obtain the preprocessed terrain image. The feature extraction module is used to input the preprocessed terrain image into the backbone feature extraction network to extract depth feature maps; wherein, the preprocessed terrain image is used as the target sample; The multi-scale texture modulation module is used to extract the texture response map of each branch from the depth feature map through multiple depth-separable convolutional branches of different scales, calculate the branch weights according to the texture energy distribution of each texture response map and perform weighted fusion to obtain the modulated feature map. The confusion difference attention module is used to acquire candidate confusion samples that are different from the target sample category, calculate the feature differences between the target sample and the candidate confusion samples, generate channel attention and spatial attention based on the feature differences, and apply the channel attention and the spatial attention to the modulated feature map to obtain the enhanced feature map. The confusion calculation and pairing module is used to pool the enhanced feature map to obtain a feature vector, calculate the comprehensive confusion between the target sample and each candidate confusion sample based on the feature vector, and select a high-risk confusion sample for each target sample according to the comprehensive confusion, thus forming a high-risk confusion sample pair. The texture transfer enhancement module is used to input the target sample and the high-risk confusion sample into the semantic guidance texture transfer sub-network, generate a transfer mask and extract local texture patterns from the high-risk confusion sample, and use the transfer mask to control the injection position to superimpose the local texture patterns onto the non-subject region of the target sample to obtain a confusion enhancement sample that maintains the semantics of the target category. The joint training module is used to input the confused and enhanced samples and the original samples into the classification network for joint training to obtain a trained terrain recognition model; The recognition output module is used to input the terrain image to be recognized into the trained terrain recognition model and output the terrain category recognition result.
10. An electronic device, characterized in that: It includes a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the robot terrain recognition method for easily confused terrain categories as described in any one of claims 1 to 8.