Humanoid robot field obstacle identification method based on transform feature enhancement

By combining camera images and lidar data, using transformer feature enhancement method, the problems of low obstacle recognition efficiency and high computational complexity in field environments are solved, and efficient and accurate obstacle recognition is achieved.

CN120299002AInactive Publication Date: 2025-07-11JIANGSU YUNMU ZHIZAO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510345445.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In unstructured environments in the wild, image-based methods are susceptible to light, while lidar-based methods have challenges in road segmentation and small-object detection. The fusion of the two modes is less studied, resulting in low barrier recognition efficiency and high computational complexity.

Method used

The humanoid robot obstacle recognition method based on transformer feature enhancement is adopted, combined with camera images and lidar point cloud data, features are extracted through ResNet50 and SECOND architecture, and feature fusion is performed using dynamic feature indexing module and adaptive fusion module to reduce projection errors and improve recognition accuracy.

Benefits of technology

It realizes robust recognition of 3D obstacles in the wild environment, reduces the time and space complexity of system identification, and improves the recognition efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299002A_ABST
    Figure CN120299002A_ABST
Patent Text Reader

Abstract

The invention discloses a humanoid robot field obstacle recognition method based on transform feature enhancement. The method comprises the following steps: S1, firstly, acquiring image data of an environment by using a camera and acquiring point cloud data of the environment by using a laser radar; s2, making the data collected in the desert environment into a three-dimensional data set for subsequent model training detection; s3, constructing a network model; and S4, for a radar point cloud branch, voxel features and voxel coordinates are obtained, a laser radar branch follows a SECOND architecture, point cloud and image data of a 3D obstacle area encountered on a patrol detection path are collected, information of a 3D obstacle target is output through analysis and calculation, a dynamic feature index (DFI) module is introduced, and the 3D obstacle target is obtained. The voxel features of the point cloud are projected to the image coordinate system by using the internal and external parameters obtained through offline calibration, the initial projection coordinates are obtained, and the method is more accurate than the method of directly projecting the voxel features to the image by using the internal and external parameters of the radar.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of humanoid robot patrol navigation, and specifically relates to a method for identifying field obstacles of a humanoid robot based on transformer feature enhancement. Background Art

[0002] In an unstructured field environment, 3D obstacle object detection can be classified into the following types according to its input method: methods based on camera images and methods based on lidar detection.

[0003] Methods based on camera images. Images encapsulate rich semantic and texture information, making them valuable resources for various applications. Taking advantage of the success of numerous algorithms, many researchers have strived to adopt image-based methods. For example, Karunasekera et al. used a binocular camera to obtain a disparity map. Subsequently, obstacles were detected by scanning pixels in the disparity map. Matthies et al. utilized the Yolov4 real-time object detection model to conduct meticulous training on 1000 lunar images from Apollo, while detecting rocks and craters. Ishida et al. learned objectness scores by calculating gradient images from a crater dataset using a linear SVM. Meanwhile, candidate regions based on their image features were used to identify the presence of craters by CNN. Tewari et al. merged optical images and DEM data and then performed crater detection through Mask RCNN. Lin et al. input an elevation map into a neural network to generate prediction results of the resulting craters. However, image-based deep learning methods are simple and are easily affected by lighting conditions and lack of precise position information.

[0004] LiDAR-based method. Compared with camera-based methods, LiDAR-based methods can limit the influence of light and provide high-precision scene position information, making them widely used in 3D object detection tasks. Shang et al. used LiDAR to collect obstacle data, strategically fused width and background information using a Bayesian network, and used an SVM classifier for final detection. Maties et al. calculated the surface normals around each point in the point cloud to determine the position of the crater, and then matched it with the input crater to determine the size. In a simulated environment, Zhou et al. used plane fitting to separate the points below the ground, used DBSCAN clustering to identify crater positions, and used RANSAC to perform circle fitting on the clustering results to determine obstacle sizes. Zhong et al. introduced a multi-LiDAR crater detection method that identifies potential negative obstacle feature point pairs by extracting radial distance jumps and then filters them according to geometric features. Goodin et al. introduced a LiDAR-based model to detect obstacles in off-road environments. The LiDAR is installed on an AGV to illustrate the movement and speed of the vehicle. Due to the large undulations and obvious curvature changes on the lunar surface, traditional LiDAR-based road segmentation methods face challenges in the road segmentation stage. In addition, the sparsity of point cloud features makes it difficult to effectively detect small targets.

[0005] In summary, in the unstructured bumpy environment of the wild, image-based methods mainly use deep learning methods, while most current LiDAR-based methods are associated with traditional methods. In addition, there is little research on the fusion of these two modalities in the wild environment. Summary of the Invention

[0006] The object of the present invention is to provide a method for identifying wild obstacles of a humanoid robot based on transformer feature enhancement, which can achieve robust identification of 3D obstacles in the detection area by the humanoid robot during the inspection and detection process in the wild environment, and reduce the time and space complexity of system identification calculation.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is: a method for identifying wild obstacles of a humanoid robot based on transformer feature enhancement, comprising the following steps:

[0008] Step S1: First, use a camera to obtain image data of the environment and a LiDAR to obtain point cloud data of the environment;

[0009] Step S2: Make the data collected in the desert environment into a three-dimensional data set for subsequent model training and detection;

[0010] Step S3: Build a network model, including two branches: an image branch and a point cloud branch. For the image branch, obtain multi-scale features. The camera branch uses ResNet50 as the backbone network and FPN as the neck structure;

[0011] Step S4: For the radar point cloud branch, obtain voxel features and voxel coordinates. The lidar branch follows the SECOND architecture, starting with dynamic voxel encoding;

[0012] Step S5: Dynamically fuse the features of the encoded point cloud with the extracted image features. The fused features are then further encoded through the backbone network and neck structure of SECOND, and the detection results are obtained through AnchorHead.

[0013] Furthermore, in the said Step S1, the lidar can obtain 3D point cloud data of the real-time scene ahead and face the field of view forward, and the camera can obtain image data of the real-time scene ahead.

[0014] Furthermore, in the said Step S2, data is collected under four different lighting conditions: front light, side light, low light, and ultra-low light. Due to the undulating characteristics of the ground, the sensor experiences different degrees of vibration at different speeds. Therefore, lidar point clouds and camera image data are collected at different speeds (10 km / h, 15 km / h, and 20 km / h).

[0015] Furthermore, in the said Step S3, for the image branch, obtain multi-scale features where N = 3, H i and W i represent the height and width of the feature map, and C i represents the number of feature channels. The camera branch uses ResNet50 as the backbone network and FPN as the neck structure.

[0016] Furthermore, in the said Step S4, for the radar point cloud branch, obtain voxel features and voxel coordinates where M represents non-empty voxels and C V represents the number of feature channels. The voxel coordinates include batch size, x, y, and z. Then, a Dynamic Feature Indexing (DFI) module is introduced, and the voxel features of the point cloud are projected into the image coordinate system using the internal and external parameters obtained through offline calibration to obtain the initial projection coordinates.

[0017] Furthermore, in the said Step S5, a simple and effective dynamic fusion module is adopted. This module adaptively extracts features from the two modalities, replacing the method of directly weighting the features. After the low-level image features have been successfully extracted in the point cloud branch, the cross-attention module is used to further extract image features.

[0018] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:

[0019] 1. The present invention collects point cloud and image data of 3D obstacle areas encountered on the inspection and detection path, outputs information on 3D obstacle targets through parsing and calculation, and by introducing a Dynamic Feature Indexing (DFI) module, projects the voxel features of the point cloud into the image coordinate system using the internal and external parameters obtained through offline calibration to obtain the initial projection coordinates, which is more accurate than directly projecting the voxel features onto the image using the internal and external parameters of the radar.

[0020] 2. This patent adopts an adaptive fusion module that can adaptively extract features from two modalities, replacing the method of directly weighting features, which is more effective than directly weighting features. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings.

[0022] Figure 1 It is an example diagram of dataset annotation provided by the present invention.

[0023] Figure 2 It is a schematic diagram of the network structure of the pre-fusion feature scheme provided by the present invention.

[0024] Figure 3 It is a schematic diagram of the dynamic feature extraction module provided by the present invention.

[0025] Figure 4 It is a schematic diagram of the adaptive fusion module provided by the present invention.

[0026] Figure 5 It is a schematic diagram of the Transformer-based feature enhancement module provided by the present invention.

[0027] Figure 6 It is a schematic diagram of dot product attention (upper) and multi-head attention (lower) provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts in the embodiments of the present invention belong to the scope of protection of the present invention.

[0029] Embodiment 1

[0030] The present invention provides a method for identifying field obstacles of a humanoid robot based on transformer feature enhancement, which includes the following steps:

[0031] First, a three-dimensional dataset is created. The validation scenario in the U2Ground dataset belongs to a typical desert environment, which is characterized by undulating ground and the presence of a large number of rocks and potholes.

[0032] The data acquisition vehicle is equipped with a forward-facing camera located 149 cm above the ground, operating in the FLIR41c6c mode with a resolution of 2448×2048. In addition, it is integrated with an 180-line LS180S2 lidar with a field of view (FOV) of 120°(H) x 25°(V), located 79 cm above the ground.

[0033] To enhance the diversity of the dataset, data is collected under four different lighting conditions: frontal light, side light, low light, and ultra-low light. Due to the undulating nature of the ground, the sensor experiences different degrees of vibration at different speeds. Therefore, this patent collects lidar point clouds and camera image data at different speeds (10 km / h, 15 km / h, and 20 km / h) to cover different degrees of sensor vibration.

[0034] This patent selects 1463 frames from the collected data as the most representative scenarios, which require laborious 3D annotation. An annotation example is Figure 1 as shown.

[0035] The training set consists of 1317 radar-camera frames, including 5020 "rock" annotations and 2662 "pothole" annotations. The validation set consists of 146 radar-camera frames, including 514 "rocks" and 264 "potholes". The dataset of this patent shows a consistent distribution under ambient light conditions: frontal, side, low light, and ultra-low light. Specifically, the training set and validation set include 357, 328, 334, 298 frames and 39, 38, 33, 36 frames for each lighting condition.

[0036] A model is constructed. The model takes point clouds and front view images as inputs, denoted as and Here, N represents the number of points in the point cloud, (x i , y i , z i ) represents the spatial coordinates of the i-th point in the point cloud, f i represents the reflection intensity, H and W represent the height and width of the image. The formal representation of the model output is a 3D detection target set, denoted as where M is the number of detected targets. Each target is represented by its 3D spatial coordinates (x i , y i , z i ), dimensions (l i , wi , h i ), and class information c i to describe.

[0037] The proposed DyFTNet is as Figure 2 shown, including two branches: the image branch and the point cloud branch. For the image branch, multi-scale features are obtained where N = 3, H i and W i represent the height and width of the feature map, and C i represents the number of feature channels. For the radar point cloud branch, voxel features and voxel coordinates are obtained, where M represents non-empty voxels, and C V represents the number of feature channels. The voxel coordinates include the batch size, x, y, and z.

[0038] The details of the camera branch and the lidar branch are as follows. The camera branch uses ResNet50 as the backbone network and FPN as the neck structure. The lidar branch follows the SECOND architecture, starting with dynamic voxel encoding. The encoded features are dynamically fused with the extracted image features. The fused features are then further encoded through the backbone network and neck structure of SECOND, and the detection results are obtained through AnchorHead.

[0039] Dynamic Feature Extraction

[0040] After the camera branch, the image feature F I is obtained, and then it is indexed using the voxel features of the point cloud. Usually, the voxel features are projected onto the image using the internal and external parameters of the lidar. Based on the position of the projected points in the image coordinate system, the corresponding image features are extracted. However, this process is usually prone to inaccuracies because the internal and external parameter matrices from the lidar to the image often contain errors. These errors mainly originate from the roughness of the surface in the field environment, which causes sensor vibrations during vehicle movement. Therefore, this leads to inaccuracies in the initial calibration parameters.

[0041] To solve this problem, this patent introduces a Dynamic Feature Indexing (DFI) module, as Figure 3 shown. First, this patent projects the voxel features of the point cloud into the image coordinate system using the internal and external parameters obtained from offline calibration to obtain the initial projection coordinates:

[0042] P uv = KTP w ;

[0043] where K represents the internal parameter matrix, T is the external parameter matrix, represents the coordinates of the point cloud in the world coordinate system, Denote the coordinates projected onto the image coordinate system. Then, extract the corresponding image feature blocks based on the projected coordinates in the image coordinate system, and perform pooling on these image feature blocks to obtain the corresponding image features:

[0044] F indexed = f Pooling (f Indexing (P uv , F I ));

[0045] Next, by learning the dynamic offset F' related to the index feature uv , this patent dynamically corrects the initial projected coordinates:

[0046] P' uv = σ(Conv(F indexed ));

[0047]

[0048] Finally, by using the corrected coordinates and repeating Equation (3-38), the indexed image features are obtained Through these operations, the mitigation of the projection error is achieved.

[0049] Adaptive Fusion

[0050] After DFI, this patent has obtained the image features corresponding to the voxel features of the point cloud. Some methods directly combine the features of the two modalities to obtain the fused features for downstream module processing. However, this is not reasonable because the features of the two modalities have different effectiveness for different detection targets.

[0051] A simple and effective dynamic fusion module is adopted, as Figure 4 shown. This module adaptively extracts features from the two modalities, replacing the method of directly weighting the features, and undergoes the following steps:

[0052] First, connect the features of the two modalities along the feature dimension, that is and F V . Subsequently, this patent uses a set of convolutional layers to obtain the feature U. Then, inspired by the Squeeze-and-Excitation concept

[47] , the feature U undergoes a squeezing operation to aggregate the feature map across the spatial dimensions H×W to generate a channel descriptor. Next is an excitation operation, which learns the sample-specific activation of each channel through a self-gating mechanism based on channel dependence to control the excitation of each channel. The resulting feature is denoted as U se . Finally, the output U sePerform element-wise multiplication with feature U to obtain the fusion result F fuse 。

[0053] This process can be formalized as follows:

[0054]

[0055] where f conv represents a set of convolution operations, and f se is a Squeeze-and-Excitation module defined as follows:

[0056] f se (U) = σ(W favg (U)) · U;

[0057] The ablation study shows that the adaptive fusion module of this patent is more effective than directly weighting the features.

[0058] Feature Enhancement Module Based on Transformer

[0059] As Figure 5 shown, after the above steps, the point cloud branch has successfully extracted low-level image features. To enhance the fusion effect of the final features, this patent uses a cross-attention module to further extract the image feature F I 。

[0060] After adaptive fusion, this patent obtains voxel-level fusion features. These features are processed by a classical point cloud backbone and neck, similar to the SECOND architecture, to obtain the feature F bn , and then it is flattened into a query sequence.

[0061] Subsequently, this patent compresses the image feature map along the height direction to generate a key-value sequence. This compression strategy along the height axis utilizes camera geometry and can easily establish an association between spatial positions and image columns. Generally, each image column contains at most one target. Therefore, folding along the height axis can significantly reduce the computational cost without losing important information. Then, the query and key-value sequences are input into the Transformer decoder layer for fusion enhancement. The decoder layer follows the design principle of DETR. The Transformer network structure is as Figure 6 shown.

[0062] Training Loss

[0063] In DyFTNet, the weights of the image branch are frozen, and only the lidar branch is trained. The overall training loss L includes the classification loss L cls , the bounding box regression loss L bbox , and the orientation loss L dir, where λ1, λ2, and λ3 are hyperparameters used to balance various losses and are set to 1.0, 2.0, and 0.2 respectively in the experiment.

[0064] L = λ1 × L cls + λ2 × L bbox + λ3 × L dir ;

[0065] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation manners. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A method for identifying field obstacles of a humanoid robot based on transformer feature enhancement, characterized in that, It includes the following steps: Step S1: First, use a camera to obtain image data of the environment and a lidar to obtain point cloud data of the environment; Step S2: Make the data collected in the desert environment into a three-dimensional data set for subsequent model training and detection; Step S3: Build a network model, including two branches: an image branch and a point cloud branch. For the image branch, obtain multi-scale features. The camera branch uses ResNet50 as the backbone network and FPN as the neck structure; Step S4: For the radar point cloud branch, obtain voxel features and voxel coordinates. The lidar branch follows the SECOND architecture, starting with dynamic voxel encoding; Step S5: Dynamically fuse the features of the encoded point cloud and the extracted image features. The fused features are then further encoded through the backbone network and neck structure of SECOND, and the detection results are obtained through AnchorHead.

2. The method for identifying field obstacles of a humanoid robot based on transformer feature enhancement according to claim 1, wherein: In the step S1, the lidar can obtain 3D point cloud data of the real-time scene ahead and face the field of view forward, and the camera can obtain image data of the real-time scene ahead.

3. A method for identifying field obstacles of a humanoid robot based on transformer feature enhancement according to claim 1, characterized in that: In the step S2, data is collected under four different lighting conditions: front light, side light, low light, and ultra-low light. Due to the undulating characteristics of the ground, the sensor experiences different degrees of vibration at different speeds. Therefore, radar point clouds and camera image data are collected at different speeds (10 km / h, 15 km / h, and 20 km / h).

4. A method for identifying field obstacles of a humanoid robot based on transformer feature enhancement according to claim 1, characterized in that: In the step S3, for the image branch, multi-scale features are obtained. where N = 3, H i and W i represent the height and width of the feature map, and C i represents the number of feature channels. The camera branch uses ResNet50 as the backbone network and FPN as the neck structure.

5. A method for identifying field obstacles of a humanoid robot based on transformer feature enhancement according to claim 1, characterized in that: In the step S4, for the radar point cloud branch, voxel features are obtained and voxel coordinates where M represents non-empty voxels, and C V represents the feature channel. The voxel coordinates include the batch size, x, y, and z. Then, a Dynamic Feature Indexing (DFI) module is introduced, and the voxel features of the point cloud are projected onto the image coordinate system using the internal and external parameters obtained through offline calibration to obtain the initial projection coordinates.

6. A method for identifying field obstacles of a humanoid robot based on transformer feature enhancement according to claim 1, characterized in that: In the step S5, a simple and effective dynamic fusion module is adopted. This module adaptively extracts features from the two modalities, replacing the method of directly weighting the features. After the low-level image features have been successfully extracted in the point cloud branch, the cross-attention module is used to further extract image features.

Citation Information

Patent Citations

  • Transform feature enhancement-based lunar surface obstacle identification method

    CN118587686A