Three-dimensional terrain recognition method and apparatus, device, and computer-readable storage medium

By combining inertial sensor data and image data for dynamic compensation to generate point cloud data, the problem of inaccurate robot terrain recognition is solved, achieving higher precision terrain feature extraction and safer robot control.

WO2026092096A1PCT designated stage Publication Date: 2026-05-07HYPERSHELL CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HYPERSHELL CO LTD
Filing Date
2025-10-11
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and adapt to various terrains, resulting in inaccurate robot control and unsafe movement.

Method used

By combining inertial sensor data, depth images, and field-of-view images, point cloud data is generated through dynamic compensation, and terrain features, including slope and obstacle locations, are extracted.

Benefits of technology

It improves the accuracy and precision of terrain feature recognition, and enhances the precision and safety of robot control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127083_07052026_PF_FP_ABST
    Figure CN2025127083_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A three-dimensional terrain recognition method and apparatus, a device, and a computer-readable storage medium, relating to the technical field of computers. The method comprises: acquiring a first depth image of a target region at a first moment collected by means of a first image collection module, a first field-of-view image of the target region at the first moment collected by means of a second image collection module, and first inertial data collected by means of an inertial sensor; on the basis of the first inertial data, performing dynamic compensation on the first depth image and the first field-of-view image to obtain a compensated second depth image and a compensated second field-of-view image; generating point cloud data of the target region on the basis of the second depth image and the second field-of-view image; and on the basis of the point cloud data, extracting terrain features of the target region, the terrain features being used for representing the terrain of the target region. The terrain features of the target region extracted by the method have higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional terrain recognition methods, devices, equipment and computer-readable storage media

[0001] This disclosure claims priority to Chinese Patent Application No. 202411560674.7, filed on November 4, 2024, entitled “Three-dimensional terrain recognition method, apparatus, device and computer-readable storage medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a three-dimensional terrain recognition method, apparatus, device, and computer-readable storage medium. Background Technology

[0003] With the continuous development of computer technology, the robotics industry is also developing rapidly, and the application fields of robots are becoming increasingly wide. For example, robots can be attached to the outside of users to enhance their strength, agility, and endurance. The widespread application of robots in outdoor scenarios such as rehabilitation medicine, industrial safety, and tourism has made the ability of robots to adapt to various terrains one of the key technologies of robotics.

[0004] Therefore, a three-dimensional terrain recognition method is needed to identify terrain features in order to better control the robot based on these features, enabling the robot to adapt to various terrains. Summary of the Invention

[0005] This application provides a three-dimensional terrain recognition method, apparatus, device, and computer-readable storage medium, which can be used to determine the terrain features of a target area. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a three-dimensional terrain recognition method, the method comprising:

[0007] The system acquires a first depth image of the target region acquired by a first image acquisition module at a first moment, a first field-of-view image of the target region acquired by a second image acquisition module at the first moment, and first inertial data acquired by an inertial sensor. The first depth image is used to indicate the depth information of each pixel in the target region at the first moment, the first field-of-view image is used to indicate the color information of each pixel in the target region at the first moment, and the first inertial data is used to indicate the motion information of the first image acquisition module and the second image acquisition module at the first moment.

[0008] Based on the first inertial data, the first depth image and the first field of view image are dynamically compensated to obtain the compensated second depth image and the second field of view image.

[0009] Point cloud data of the target region is generated based on the second depth image and the second field-of-view image;

[0010] Based on the point cloud data, the terrain features of the target area are extracted, and the terrain features are used to characterize the terrain of the target area.

[0011] In one possible implementation, the method further includes:

[0012] The first field-of-view image is detected to obtain the illumination intensity of the first field-of-view image;

[0013] The step of dynamically compensating the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image includes:

[0014] Since the illumination intensity of the first field-of-view image does not meet the intensity requirements, dynamic compensation is performed on the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image.

[0015] In one possible implementation, the first inertial data includes attitude data and acceleration;

[0016] The step of dynamically compensating the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image includes:

[0017] Based on the pose data, the first depth image, and the first field of view image, a third depth image and a third field of view image are obtained, wherein the third depth image and the third field of view image are located in a reference coordinate system;

[0018] Based on the acceleration, the third depth image is height-compensated to obtain the second depth image, and the third field-of-view image is height-compensated to obtain the second field-of-view image.

[0019] In one possible implementation, obtaining the third depth image and the third field-of-view image based on the pose data, the first depth image, and the first field-of-view image includes:

[0020] Process the first depth image to obtain a reference depth image after denoising the first depth image;

[0021] The first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image;

[0022] Based on the attitude data, a rotation matrix is ​​determined, which is used to transform the reference depth image and the reference field of view image to the reference coordinate system;

[0023] The reference depth image is transformed according to the rotation matrix to obtain the third depth image;

[0024] The reference field-of-view image is transformed according to the rotation matrix to obtain the third field-of-view image.

[0025] In one possible implementation, transforming the reference depth image according to the rotation matrix to obtain the third depth image includes:

[0026] Obtain the position and depth information of each first pixel in the reference depth image;

[0027] Based on the rotation matrix, update the position and depth information of each first pixel to obtain the updated position and depth information of each first pixel.

[0028] The third depth image is generated based on the updated position and depth information of each first pixel.

[0029] In one possible implementation, transforming the reference field-of-view image according to the rotation matrix to obtain the third field-of-view image includes:

[0030] Obtain the position and color information of each second pixel in the reference field image;

[0031] Based on the rotation matrix, update the position information and color information of each second pixel to obtain the updated position information and color information of each second pixel.

[0032] The third field-view image is generated based on the updated position and color information of each second pixel.

[0033] In one possible implementation, the step of performing height compensation on the third depth image based on the acceleration to obtain the second depth image includes:

[0034] Based on the acceleration, determine the height change value of the inertial sensor in the reference direction;

[0035] Based on the height change value, height compensation is performed on each pixel in the third depth image to obtain the second depth image.

[0036] In one possible implementation, determining the height change value of the inertial sensor in the reference direction based on the acceleration includes:

[0037] Integrating the acceleration yields the velocity;

[0038] Based on the speed and the target duration, the height change value of the inertial sensor in the reference direction is determined, and the target duration is obtained based on the sampling frequency of the inertial sensor.

[0039] In one possible implementation, the step of performing height compensation on each pixel in the third depth image based on the height change value to obtain the second depth image includes:

[0040] The height change value is added to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image;

[0041] The second depth image is generated based on the depth information of each pixel in the third depth image after height compensation.

[0042] In one possible implementation, the method further includes:

[0043] Based on the fact that the illumination intensity of the first field-of-view image meets the intensity requirement, the first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image.

[0044] Generate reference point cloud data corresponding to the reference field-of-view image;

[0045] Based on the reference point cloud data, the terrain features of the target area at the first time point are extracted.

[0046] In one possible implementation, generating point cloud data of the target region at the first time moment based on the second depth image and the second field-of-view image includes:

[0047] Based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module, the three-dimensional coordinates of each pixel in the second depth image are determined.

[0048] Obtain the color information of each pixel in the second field-of-view image;

[0049] Point cloud data of the target region is generated based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field of view image.

[0050] In one possible implementation, after generating point cloud data of the target region based on the second depth image and the second field-of-view image, the method further includes:

[0051] Based on the point cloud data of the target area, a three-dimensional terrain model of the target area is generated. The three-dimensional terrain model of the target area is used by the robot's control device to control the robot.

[0052] In one possible implementation, the terrain feature includes terrain slope;

[0053] The step of extracting terrain features of the target area based on the point cloud data includes:

[0054] The point cloud data is fitted to a plane to obtain the plane equation of the target region;

[0055] Determine the normal vector of the target region based on the plane equation;

[0056] The angle between the normal vector and the horizontal plane of the target area is determined as the terrain slope of the target area.

[0057] In one possible implementation, the terrain features include the location information of obstacles;

[0058] The step of extracting terrain features of the target area based on the point cloud data includes:

[0059] The point cloud data is segmented to obtain ground points and non-ground points;

[0060] Cluster analysis is performed on the non-ground points to obtain at least one set of non-ground points;

[0061] For any set of non-ground points, if the volume of any set of non-ground points is greater than a volume threshold, the target location information is determined based on the location information of any set of non-ground points.

[0062] The target location information is determined to be the location information of obstacles included in the target area.

[0063] In one possible implementation, the method further includes:

[0064] Based on the first field-of-view image, the terrain material of the target area is determined.

[0065] On the other hand, embodiments of this application provide a three-dimensional terrain recognition device, the device comprising:

[0066] The acquisition module is used to acquire a first depth image of the target area acquired by the first image acquisition module at a first moment, a first field-of-view image of the target area acquired by the second image acquisition module at the first moment, and first inertial data acquired by the inertial sensor. The first depth image is used to indicate the depth information of each pixel in the target area at the first moment, the first field-of-view image is used to indicate the color information of each pixel in the target area at the first moment, and the first inertial data is used to indicate the motion information of the first image acquisition module and the second image acquisition module at the first moment.

[0067] The dynamic compensation module is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data to obtain a compensated second depth image and a second field of view image.

[0068] The generation module is used to generate point cloud data of the target region based on the second depth image and the second field-of-view image;

[0069] The identification module is used to extract the terrain features of the target area based on the point cloud data, and the terrain features are used to characterize the terrain of the target area.

[0070] In one possible implementation, the device further includes:

[0071] The detection module is used to detect the first field-of-view image and obtain the illumination intensity of the first field-of-view image;

[0072] The dynamic compensation module is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data, since the illumination intensity of the first field of view image does not meet the intensity requirements, to obtain a compensated second depth image and a second field of view image.

[0073] In one possible implementation, the first inertial data includes attitude data and acceleration;

[0074] The dynamic compensation module is used to obtain a third depth image and a third field of view image based on the attitude data, the first depth image, and the first field of view image, wherein the third depth image and the third field of view image are located in a reference coordinate system; and to perform height compensation on the third depth image based on the acceleration to obtain a second depth image, and to perform height compensation on the third field of view image to obtain a second field of view image.

[0075] In one possible implementation, the dynamic compensation module is used to process the first depth image to obtain a denoised reference depth image; process the first field-of-view image to obtain a denoised and feature-extracted reference field-of-view image; determine a rotation matrix based on the pose data, the rotation matrix being used to transform the reference depth image and the reference field-of-view image to the reference coordinate system; transform the reference depth image based on the rotation matrix to obtain the third depth image; and transform the reference field-of-view image based on the rotation matrix to obtain the third field-of-view image.

[0076] In one possible implementation, the dynamic compensation module is used to acquire the position information and depth information of each first pixel in the reference depth image; update the position information and depth information of each first pixel based on the rotation matrix to obtain the updated position information and depth information of each first pixel; and generate the third depth image based on the updated position information and depth information of each first pixel.

[0077] In one possible implementation, the dynamic compensation module is used to acquire the position and color information of each second pixel in the reference field-of-view image; update the position and color information of each second pixel based on the rotation matrix to obtain the updated position and color information of each second pixel; and generate the third field-of-view image based on the updated position and color information of each second pixel.

[0078] In one possible implementation, the dynamic compensation module is used to determine the height change value of the inertial sensor in the reference direction based on the acceleration; and to perform height compensation on each pixel in the third depth image based on the height change value to obtain the second depth image.

[0079] In one possible implementation, the dynamic compensation module is used to integrate the acceleration to obtain a velocity; and to determine the height change value of the inertial sensor in the reference direction based on the velocity and the target duration, wherein the target duration is obtained based on the sampling frequency of the inertial sensor.

[0080] In one possible implementation, the dynamic compensation module is used to add the height change value to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image; and to generate the second depth image based on the height-compensated depth information of each pixel in the third depth image.

[0081] In one possible implementation, the generation module is further configured to process the first field-of-view image based on the illumination intensity of the first field-of-view image meeting the intensity requirement, to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image; and generate reference point cloud data corresponding to the reference field-of-view image.

[0082] The identification module is further configured to extract the terrain features of the target area at the first moment based on the reference point cloud data.

[0083] In one possible implementation, the generation module is configured to determine the three-dimensional coordinates of each pixel in the second depth image based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module; acquire the color information of each pixel in the second field-of-view image; and generate point cloud data of the target region based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field-of-view image.

[0084] In one possible implementation, the generation module is further configured to generate a three-dimensional terrain model of the target area based on the point cloud data of the target area, and the three-dimensional terrain model of the target area is used by the robot's control device to control the robot.

[0085] In one possible implementation, the terrain feature includes terrain slope;

[0086] The recognition module is used to perform plane fitting on the point cloud data to obtain the plane equation of the target region; determine the normal vector of the target region based on the plane equation; and determine the angle between the normal vector and the horizontal plane of the target region as the terrain slope of the target region.

[0087] In one possible implementation, the terrain features include the location information of obstacles;

[0088] The identification module is used to segment the point cloud data to obtain ground points and non-ground points; perform cluster analysis on the non-ground points to obtain at least one group of non-ground points; for any group of non-ground points, if the volume of any group of non-ground points is greater than a volume threshold, determine the target location information based on the location information of the any group of non-ground points; and determine the target location information as the location information of obstacles included in the target area.

[0089] In one possible implementation, the device further includes:

[0090] The determination module is used to determine the terrain material of the target area based on the first field-of-view image.

[0091] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor, so that the computer device implements any of the three-dimensional terrain recognition methods described above.

[0092] On the other hand, a computer-readable storage medium is also provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the at least one piece of program code being loaded and executed by a processor to enable a computer to implement any of the three-dimensional terrain recognition methods described above.

[0093] On the other hand, a computer program or computer program product is also provided, wherein the computer program or computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement any of the above-mentioned three-dimensional terrain recognition methods.

[0094] The technical solution provided in this application has at least the following beneficial effects:

[0095] The technical solution provided in this application, when determining the terrain features of a target area, considers not only the first inertial data from the inertial sensor at the first moment, but also the first field-of-view image and the first depth image of the target area at the first moment. Based on the first inertial data, the first depth image and the first field-of-view image are processed to determine the terrain features of the target area, resulting in higher precision and accuracy of the determined terrain features, which better reflect the actual situation of the target area. Since terrain features can characterize the terrain of the target area, more precise and accurate terrain features can better characterize the terrain of the target area. When controlling the robot based on the terrain of the target area, the control precision of the robot is improved, as well as the walking safety of the robot and the walking safety of the robot user. Attached Figure Description

[0096] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0097] Figure 1 is a schematic diagram of the implementation environment of a three-dimensional terrain recognition method provided in an embodiment of this application;

[0098] Figure 2 is a flowchart of a three-dimensional terrain recognition method provided in an embodiment of this application;

[0099] Figure 3 is a flowchart of a three-dimensional terrain recognition method provided in an embodiment of this application;

[0100] Figure 4 is an architecture diagram of a three-dimensional terrain recognition system provided in an embodiment of this application;

[0101] Figure 5 is a schematic diagram of the structure of a three-dimensional terrain recognition device provided in an embodiment of this application;

[0102] Figure 6 is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0103] Figure 7 is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0104] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0105] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0106] First, a brief introduction to the terms used in the embodiments of this application will be given.

[0107] Depth cameras, also known as 3D (three-dimensional) cameras, are devices that capture distance information of objects in a scene. They use various technologies, such as structured light, time-of-flight (ToF), and binocular vision, to acquire the three-dimensional coordinates of objects, thereby achieving depth perception. Depth cameras have wide applications in many fields, including but not limited to facial recognition technology, 3D reconstruction, robotics and automation, and augmented reality (AR).

[0108] ToF depth camera: It directly generates depth information by measuring the time it takes for light to travel from the camera lens to the object and back.

[0109] A color (Red, Green, Blue, RGB) camera is an imaging device capable of capturing the three primary colors of light: red (R), green (G), and blue (B). It receives light from a scene through a lens and uses millions of tiny photosensitive elements on its internal image sensor (e.g., a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS)) to convert the light into electrical signals. Each photosensitive element corresponds to a pixel and is sensitive to only one of the three RGB colors. A color filter array (such as a Bayer color filter array) decomposes the received light signal into the three RGB color channels. During imaging, the RGB camera combines the RGB values ​​of each pixel to form a full-color image. This image can display rich colors and details because the human eye is most sensitive to these three colors; the RGB camera can simulate the human eye's perception of color. Finally, these electrical signals are converted from analog to digital and processed by image processing, then stored or transmitted in real-time as a digital image, widely used in photography, video conferencing, surveillance systems, and many other fields. RGB cameras typically produce 24-bit true color images, capable of displaying up to 16.77 million colors, resulting in vibrant and richly detailed images.

[0110] An Inertial Measurement Unit (IMU) is a device used to measure and report specific forces, angular velocities, and in some cases, magnetic fields around an object. The principle of an IMU is based on Newton's laws of motion: an accelerometer measures the object's acceleration, and a gyroscope measures its angular velocity. By integrating these measurements, the object's velocity and position can be calculated, thus determining its attitude and trajectory in space. An IMU typically consists of an accelerometer, a gyroscope, and a magnetometer. Depending on its onboard measuring elements, an IMU can output data such as acceleration, angular velocity, magnetic field strength, and attitude angles. IMUs are widely used in aerospace, robotics, automotive, and mobile phone industries for navigation, positioning, and attitude control.

[0111] Figure 1 is a schematic diagram of the implementation environment of a three-dimensional terrain recognition method provided in an embodiment of this application. As shown in Figure 1, the implementation environment includes a computer device 101, which can be a terminal device or a server; this embodiment of the application does not limit the specific type of device. The computer device 101 is used to execute the three-dimensional terrain recognition method provided in this embodiment of the application.

[0112] Optionally, computer device 101 is a terminal device. A terminal device can be any electronic device that allows human-computer interaction with a user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablets, smart car systems, smart TVs, smart speakers, and smartwatches.

[0113] A terminal device can refer to one of multiple terminal devices; this embodiment uses only one terminal device as an example. Those skilled in the art will understand that the number of terminal devices can be more or less. For example, there may be only one terminal device, or there may be dozens or hundreds, or even more. This application embodiment does not limit the number or type of terminal devices.

[0114] When computer device 101 is a server, the server can be a single server, a server cluster consisting of multiple servers, or any of the following: a cloud computing platform or a virtualization center. This application embodiment does not limit this. The server and terminal devices communicate via a wired or wireless network. The server has data receiving, data processing, and data sending functions. Of course, the server may also have other functions, which this application embodiment does not limit.

[0115] Those skilled in the art should understand that the above-described terminal devices and servers are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0116] This application provides a three-dimensional terrain recognition method, which can be applied to the implementation environment shown in Figure 1. Taking the flowchart of a three-dimensional terrain recognition method provided by this application embodiment shown in Figure 2 as an example, the method can be executed by the computer device 101 in Figure 1. As shown in Figure 2, the method includes the following steps 201 to 204.

[0117] In step 201, the first depth image of the target area at the first moment is acquired by the first image acquisition module, the first field-of-view image of the target area at the first moment is acquired by the second image acquisition module, and the first inertial data is acquired by the inertial sensor.

[0118] The first depth image indicates the depth information of each pixel in the target area at the first moment, and the first field-of-view image indicates the color information of each pixel in the target area at the first moment. The first inertial data is the inertial data of the inertial sensor at the first moment, and it indicates the motion information of the first and second image acquisition modules at the first moment. The first inertial data includes acceleration and attitude data. The acceleration is the acceleration of the first and second image acquisition modules at the first moment, and the attitude data is the attitude data of the first and second image acquisition modules at the first moment. The attitude data includes pitch angle, roll angle, and yaw angle.

[0119] In one possible implementation, the computer device is either mounted on the robot or capable of remotely controlling the robot; this application embodiment does not limit this. The robot is also equipped with a first image acquisition module, a second image acquisition module, and an inertial sensor. The first image acquisition module is a depth camera; optionally, the depth camera can be any one of a binocular vision depth camera, a structured light depth camera, a ToF depth camera, or a light field camera. The second image acquisition module can be an RGB camera. The first image acquisition module, the second image acquisition module, and the inertial sensor are all connected to the computer device via a wired or wireless network. The robot can be an exoskeleton robot or other types of robots; this application embodiment does not limit this.

[0120] In some embodiments, the first image acquisition module and the second image acquisition module described above can be implemented as the same image acquisition module. For example, an RGB-D camera is a device that can simultaneously capture color images (RGB) and depth information (Depth), that is, it acquires depth images and field-of-view images (RGB images) through an RGB-D camera.

[0121] In some embodiments, the first image acquisition module acquires the depth information of each pixel in the target area at a first moment by transmitting and receiving light signals to each pixel in the target area at a first moment. The depth information of any pixel in the target area at the first moment is used to indicate the distance between any pixel in the target area and the first image acquisition module at the first moment. After acquiring the depth information of each pixel in the target area at the first moment, the first image acquisition module sends the depth information of each pixel in the target area at the first moment to a computer device, so that the computer device can acquire the depth information of each pixel in the target area at the first moment, and then generate a first depth image of the target area at the first moment based on the depth information of each pixel in the target area at the first moment.

[0122] Alternatively, after acquiring the depth information of each pixel in the target area at the first moment, the first image acquisition module generates a first depth image of the target area at the first moment based on the depth information of each pixel in the target area at the first moment, and then sends the first depth image of the target area at the first moment to the computer device so that the computer device can acquire the first depth image of the target area at the first moment.

[0123] In one possible implementation, the sampling rate of the first image acquisition module is 30 fps (frames per second), ensuring real-time acquisition of depth information in dynamic environments. The resolution of the first image acquisition module is 100×100, balancing image accuracy and processing speed. Of course, the sampling rate and resolution of the first image acquisition module can also be other, and this application embodiment does not limit them.

[0124] The second image acquisition module is used to acquire the first field-of-view image (i.e., RGB image) of the target area at the first moment, and then send the first field-of-view image of the target area at the first moment to the computer device so that the computer device can obtain the first field-of-view image of the target area at the first moment. The first field-of-view image is used to generate sparse 3D point clouds. The first field-of-view image has a good effect on the generation of 3D terrain models in well-lit environments, and consumes less resources.

[0125] The second image acquisition module has a sampling rate of 30fps, the same as the first image acquisition module, and both modules start at the same time, allowing them to acquire depth and field-of-view images simultaneously. The second image acquisition module has a resolution of 1280×720, enabling the acquisition of clear images for subsequent terrain feature extraction and texture identification. Of course, the sampling rate and resolution of the second image acquisition module can also be other than those specified in this embodiment.

[0126] An inertial sensor is used to acquire inertial data in real time, and this inertial data is used to monitor the motion state of the image acquisition module. The inertial data can be used to compensate for motion deviations of the image acquisition module during movement, ensuring the stability of depth and field-of-view images. The sampling rate of the inertial sensor can be the same as that of the first and second image acquisition modules, and the startup time of the inertial sensor is the same as that of the first and second image acquisition modules, so that the first and second image acquisition modules and the inertial sensor can acquire depth images, field-of-view images, and inertial data at the same time.

[0127] The sampling rate of the inertial sensor can also differ from the sampling rates of the first image acquisition module and the second image acquisition module. For example, the sampling rate of the inertial sensor may be greater than or less than the sampling rates of the first and second image acquisition modules. For instance, the sampling rate of the inertial sensor is 500 fps. When the sampling rate of the inertial sensor differs from the sampling rates of the first and second image acquisition modules, the inertial sensor fails to acquire the first inertial data at the first moment. Therefore, it is necessary to determine the first inertial data of the inertial sensor at the first moment. This application does not limit the process of determining the first inertial data of the inertial sensor at the first moment. Optionally, the first inertial data of the inertial sensor at the first moment can be determined by a computer device or by the inertial sensor itself; this application does not limit this. The process of the inertial sensor determining its first inertial data at the first moment is similar to the process of a computer device determining its first inertial data at the first moment. This application only uses the process of a computer device determining the first inertial data of the inertial sensor at the first moment as an example for illustration.

[0128] Optionally, the process by which the computer device determines the first inertial data of the inertial sensor at the first moment includes: acquiring multiple candidate inertial data collected by the inertial sensor during a reference time period and the acquisition time corresponding to each candidate inertial data, wherein the reference time period includes the first moment; and determining the first inertial data of the inertial sensor at the first moment based on the multiple candidate inertial data, the acquisition time corresponding to each candidate inertial data, and the first moment.

[0129] Optionally, the process of determining the first inertial data of the inertial sensor at the first moment based on multiple candidate inertial data, the acquisition time corresponding to each candidate inertial data, and the first moment includes: determining the first reference inertial data and the second reference inertial data among multiple candidate inertial data; and determining the first inertial data of the inertial sensor at the first moment based on the first reference inertial data, the acquisition time corresponding to the first reference inertial data, the second reference inertial data, the acquisition time corresponding to the second reference inertial data, and the first moment.

[0130] Specifically, any two candidate inertial data from multiple candidate inertial data are used as the first reference inertial data and the second reference inertial data, respectively.

[0131] In one possible implementation, the process of determining the first inertial data of the inertial sensor at the first moment based on the first reference inertial data, the acquisition time corresponding to the first reference inertial data, the second reference inertial data, the acquisition time corresponding to the second reference inertial data, and the first moment includes: determining the first inertial data of the inertial sensor at the first moment by downsampling or linear interpolation based on the first reference inertial data, the acquisition time corresponding to the first reference inertial data, the second reference inertial data, the acquisition time corresponding to the second reference inertial data, and the first moment.

[0132] Optionally, based on the first reference inertial data, the acquisition time corresponding to the first reference inertial data, the second reference inertial data, the acquisition time corresponding to the second reference inertial data, and the first moment, the first inertial data of the inertial sensor at the first moment is determined by linear interpolation according to the following formula (1).

[0133] In the above formula (1), This is the first inertial data from the inertial sensor at the first moment. As the primary reference inertial data, The second reference inertial data is t1, the acquisition time corresponding to the first reference inertial data is t2, the acquisition time corresponding to the second reference inertial data is t, and the first moment is t.

[0134] In one possible implementation, the first inertial data of the inertial sensor at a first moment is obtained in the manner described above, so that the time corresponding to the first inertial data, the first depth image and the first field of view image are all at the first moment, thereby aligning the first inertial data, the first depth image and the first field of view image in time.

[0135] It should be noted that when the first inertial data is determined by the inertial sensor, after the inertial sensor determines the first inertial data at the first moment, it sends the first inertial data at the first moment to the computer device so that the computer device can obtain the first inertial data at the first moment.

[0136] In step 202, based on the first inertial data, the first depth image and the first field of view image are dynamically compensated to obtain the compensated second depth image and the second field of view image.

[0137] In one possible implementation, after acquiring the first field-of-view image, the first field-of-view image is detected to obtain the illumination intensity of the first field-of-view image; based on the illumination intensity of the first field-of-view image meeting the intensity requirements, the first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction; reference point cloud data corresponding to the reference field-of-view image is generated; and based on the reference point cloud data, the terrain features of the target area at the first time step are extracted.

[0138] Optionally, the illumination intensity of the first field-of-view image meeting the intensity requirement means that the illumination intensity of the first field-of-view image is greater than the intensity threshold. The intensity threshold is set based on experience or adjusted according to the implementation environment; this embodiment does not limit this. The process of extracting the terrain features of the target area at the first moment based on the reference point cloud data is similar to the process described below for extracting the terrain features of the target area at the first moment based on point cloud data, as detailed below. This embodiment will not repeat the details here.

[0139] The process of detecting the first field-of-view image and obtaining its illumination intensity in this embodiment is not limited. Optionally, the illumination intensity of the first field-of-view image can be obtained by detecting the first field-of-view image using an illumination intensity determination model. The illumination intensity model can be any model capable of determining illumination intensity, and this embodiment is not limited to it. The process of processing the first field-of-view image to obtain a reference field-of-view image after denoising and feature extraction is described below and will not be repeated here.

[0140] In this implementation, since the terrain features of the target area at the first moment can be determined solely from the field-of-view image of the target area at the first moment, the algorithm for determining the terrain features is simple, thereby saving computing resources and extending the usage time of computer equipment.

[0141] In one possible implementation, when the illumination intensity of the first field-of-view image does not meet the intensity requirement, that is, when the illumination intensity of the first field-of-view image is not greater than the intensity threshold, the terrain features of the target area at the first moment are determined through steps 202 to 204. In step 202, the first inertial data includes attitude data and acceleration. The process of dynamically compensating the first depth image and the first field-of-view image based on the first inertial data to obtain the compensated second depth image and second field-of-view image includes: acquiring a third depth image and a third field-of-view image based on the attitude data, the first depth image, and the first field-of-view image, wherein the third depth image and the third field-of-view image are located in a reference coordinate system; performing height compensation on the third depth image based on the acceleration to obtain the second depth image, and performing height compensation on the third field-of-view image to obtain the second field-of-view image. The reference coordinate system is a horizontal coordinate system.

[0142] Optionally, the process of obtaining a third depth image and a third field of view image based on pose data, a first depth image, and a first field of view image includes: processing the first depth image to obtain a reference depth image after denoising the first depth image; processing the first field of view image to obtain a reference field of view image after denoising and feature extraction of the first field of view image; determining a rotation matrix based on pose data, the rotation matrix being used to transform the reference depth image and the reference field of view image to a reference coordinate system; transforming the reference depth image based on the rotation matrix to obtain the third depth image; and transforming the reference field of view image based on the rotation matrix to obtain the third field of view image.

[0143] When acquiring depth information, the first image acquisition module may be affected by ambient lighting, surface reflectivity, and its own noise, resulting in irregular noise in the first depth image. Therefore, it is necessary to perform denoising processing on the first depth image to obtain a reference depth image, which is smoother and more reliable.

[0144] Optionally, Gaussian filtering can be used to denoise the first depth image to obtain a reference depth image. Gaussian filtering is a linear filtering method that reduces the influence of random noise by smoothing the pixel values ​​of each pixel and its neighborhood in the first depth image.

[0145] Optionally, the process of using Gaussian filtering to denoise the first depth image to obtain a reference depth image includes: adjusting the depth information of each pixel in the first depth image using Gaussian filtering to obtain the adjusted depth information of each pixel in the first depth image; and generating a reference depth image based on the adjusted depth information of each pixel in the first depth image.

[0146] In this process, Gaussian filtering is used to adjust the depth information of each pixel in the first depth image according to the following formula (2) to obtain the adjusted depth information of each pixel in the first depth image.

[0147] In the above formula (2), D ′ (x,y) represents the adjusted depth information of pixel (x,y) in the first depth image; These are the normalization coefficients of the Gaussian filter function, used to ensure that the filter energy remains constant; σ is the Gaussian kernel function, which determines the weight distribution of the filter; σ is the standard deviation, used to control the width of the Gaussian kernel; D(x,y) is the depth information of pixel (x,y) in the first depth image; * is the convolution operator; x is the x-coordinate of pixel (x,y) in the first depth image; y is the y-coordinate of pixel (x,y) in the first depth image.

[0148] The process of processing the first field-of-view image includes denoising and feature extraction. This process involves: denoising the first field-of-view image to obtain a denoised first field-of-view image; and then performing feature extraction on the denoised first field-of-view image to obtain a reference field-of-view image.

[0149] The process of denoising the first field-of-view image to obtain a denoised first field-of-view image includes: denoising the first field-of-view image using bilateral filtering to obtain a denoised first field-of-view image. Bilateral filtering can not only smooth the image noise of the first field-of-view image, but also preserve the image edge information of the first field-of-view image, thereby providing a higher quality image for subsequent feature extraction.

[0150] Optionally, the process of denoising the first field-of-view image by bilateral filtering to obtain the denoised first field-of-view image includes: adjusting the pixel values ​​of each pixel in the first field-of-view image by bilateral filtering to obtain the adjusted pixel values ​​of each pixel in the first field-of-view image; and generating the denoised first field-of-view image based on the adjusted pixel values ​​of each pixel in the first field-of-view image.

[0151] Optionally, the pixel values ​​of each pixel in the first field of view image are adjusted by dual-variable filtering according to the following formula (3) to obtain the adjusted pixel values ​​of each pixel in the first field of view image.

[0152] In the above formula (3), I′(x,y) is the adjusted pixel value of pixel (x,y) in the first field of view image; W(x,y) is the normalization factor used to ensure that the sum of weights is 1; I(i,j) is the pixel value of pixel (i,j) in the first field of view image. σ represents the spatial domain weight, reflecting spatial proximity; d The standard deviation is the spatial domain. The intensity domain weights reflect the similarity of pixel values; σ r denoted as the standard deviation of the intensity domain; x is the abscissa of pixel (x,y) in the first field of view image, y is the ordinate of pixel (x,y) in the first field of view image, and pixel (i,j) is the neighboring pixel of pixel (x,y).

[0153] In one possible implementation, after obtaining the first field-view image after denoising in the above steps, feature extraction is also required on the first field-view image after denoising to obtain a reference field-view image.

[0154] Before performing feature extraction on the first field-of-view image after denoising, edge detection needs to be performed on the first field-of-view image after denoising to detect the edges of the first field-of-view image after denoising; then the edges of the first field-of-view image after denoising are removed to obtain the reference field-of-view image.

[0155] The Canny edge detection algorithm (an edge detection algorithm) is used to detect edges in the first field of view image after denoising. The Canny edge detection algorithm is a widely used algorithm in the field of image processing, designed to accurately detect edges in an image. The Canny edge detection algorithm includes steps 1 to 4 as described below.

[0156] Step 1: Use a Gaussian filter to smooth the first field of view image after denoising, and obtain the smoothed first field of view image.

[0157] Step 1 reduces noise in the first field-of-view image after denoising, preventing noise interference from causing false edge detection. The process involves applying a Gaussian filter and weighting the pixels surrounding each pixel in the denoised first field-of-view image using a Gaussian kernel. This smooths out minor fluctuations in the denoised first field-of-view image, thus reducing the impact of noise.

[0158] Step 2: Calculate the gradient intensity and gradient direction of each pixel in the smoothed first field-of-view image to generate a gradient image.

[0159] Step 2 is used to identify the gradient direction and gradient intensity of each pixel in the smoothed first field-of-view image. The horizontal and vertical gradients of each pixel in the smoothed first field-of-view image are calculated using the Sobel operator (a widely used edge detection operator in image processing). Based on the horizontal and vertical gradients of each pixel, the gradient intensity and gradient direction of each pixel are calculated.

[0160] Optionally, the gradient intensity of each pixel is calculated according to the following formula (4) based on the horizontal and vertical gradients of each pixel.

[0161] In the above formula (4), G is the gradient intensity of any pixel. x Let G be the gradient in the horizontal direction of any pixel. y is the gradient in the vertical direction of any pixel.

[0162] The gradient direction of each pixel is calculated according to the following formula (5) based on the horizontal and vertical gradients of each pixel.

[0163] In the above formula (5), θ is the gradient direction of any pixel, and G x Let G be the gradient in the horizontal direction of any pixel. y is the gradient in the vertical direction of any pixel.

[0164] Step 2 generates a gradient image containing the gradient intensity and direction of each pixel. Regions with higher gradient intensity in the gradient image are more likely to be edges.

[0165] Step 3: Traverse the gradient image using non-maximum suppression and retain the pixels with maximum values ​​in the outermost direction.

[0166] In step 3, each pixel is examined along the gradient direction of the gradient image, and pixels that are not local maxima are set to 0. This ensures the fineness of the edges.

[0167] For example, if the gradient direction is 0°, then the adjacent left and right pixels are compared. If the gradient strength of the current pixel is not the maximum, it is set to 0. This way, only the local maxima along the gradient direction are retained, and these points will be used as candidate edges.

[0168] Step 4: Use the double threshold method to detect and connect edges to ensure edge continuity.

[0169] Step 4 ensures edge coherence, eliminates false edges, and preserves true edges. The image after non-maximum suppression is processed using two thresholds (a high threshold and a low threshold). Pixels with gradient strength greater than the high threshold are considered strong edges, while pixels with gradient strength between the low and high thresholds are retained as weak edges if adjacent to a strong edge. Edge concatenation connects strong and weak edges to ensure coherence in the final edge design; weak edges not connected to strong edges are discarded. Step 4 yields a streamlined, coherent edge image.

[0170] In one possible implementation, a user wears a robot. During the user's movement, the robot's image acquisition modules undergo various dynamic changes along with the exoskeleton's motion, such as tilting, rotating, and displacing. These dynamic changes can cause discrepancies between the depth image acquired by the first image acquisition module and the field-of-view image acquired by the second image acquisition module, thus affecting the accuracy of the determined terrain features. To ensure the accuracy and stability of the determined terrain features, dynamic compensation is required for the depth and field-of-view images.

[0171] Based on the attitude data, the determined rotation matrix can eliminate the deviation caused by the tilt of the first and second image acquisition modules, so that the obtained third depth image and third field of view image are located in the horizontal coordinate system, thereby making the determined terrain features consistent with the real ground.

[0172] Optionally, the attitude data includes three-axis rotation angles: pitch, roll, and yaw. The process of determining the rotation matrix based on the attitude data includes: determining the rotation matrix based on the pitch, roll, and yaw angles.

[0173] For example, the rotation matrix is ​​determined according to the pitch angle, roll angle, and yaw angle using the following formula (6): R = R y (θ y )R x (θ x )R z (θ z ) Formula (6)

[0174] In the above formula (6), R is the rotation matrix, θ y Let θ be the heading angle. x Let θ be the pitch angle. z R is the roll angle. y (θ y Let R be a rotation matrix about the y-axis. x (θ x Let R be a rotation matrix about the x-axis. z (θ z ) is a matrix that rotates about the z-axis.

[0175] After determining the rotation matrix, the process of transforming the reference depth image to obtain the third depth image based on the rotation matrix includes: obtaining the position and depth information of each first pixel in the reference depth image; updating the position and depth information of each first pixel based on the rotation matrix to obtain the updated position and depth information of each first pixel; and generating the third depth image based on the updated position and depth information of each first pixel.

[0176] After determining the rotation matrix, the process of transforming the reference field-of-view image to obtain the third field-of-view image based on the rotation matrix includes: obtaining the position and color information of each second pixel in the reference field-of-view image; updating the position and color information of each second pixel based on the rotation matrix to obtain the updated position and color information of each second pixel; and generating the third field-of-view image based on the updated position and color information of each second pixel.

[0177] In one possible implementation, the acceleration included in the attitude data is used to dynamically compensate for the displacement deviation of the image acquisition module during movement. In this embodiment, not only the depth image is dynamically compensated, but also the field of view image is dynamically compensated to make the compensated image data more accurate.

[0178] Optionally, the process of performing height compensation on the third depth image based on acceleration to obtain the second depth image includes: determining the height change value of the inertial sensor in the reference direction based on the acceleration; and performing height compensation on each pixel in the third depth image based on the height change value to obtain the second depth image; wherein the reference direction is the vertical direction.

[0179] The process of performing height compensation on the third field of view image to obtain the second field of view image includes: performing height compensation on each pixel in the third field of view image based on the height change value to obtain the second field of view image.

[0180] The process of determining the height change of the inertial sensor in the reference direction based on acceleration includes: integrating the acceleration to obtain the velocity; and determining the height change of the inertial sensor in the reference direction based on the velocity and the target duration, where the target duration is obtained based on the sampling frequency of the inertial sensor. For example, if the sampling rate of the inertial sensor is 500 fps, then the target duration is 1 / 500 = 0.002 seconds.

[0181] The process of determining the altitude change of the inertial sensor in the reference direction based on the velocity and the target duration includes: using the product of the velocity and the target duration as the altitude change of the inertial sensor in the reference direction.

[0182] In one possible implementation, the process of obtaining a second depth image by height compensation of each pixel in the third depth image based on the height change value includes: adding a height change value to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image; and generating the second depth image based on the height-compensated depth information of each pixel in the third depth image.

[0183] The process of obtaining a second field of view image by performing height compensation on each pixel in the third field of view image based on the height change value includes: adding height change values ​​to the depth information of each pixel in the third field of view image to obtain the height-compensated depth information of each pixel in the third field of view image; and generating the second field of view image based on the height-compensated depth information of each pixel in the third field of view image.

[0184] It should be noted that by performing height compensation on depth and field images, the height error caused by slight vertical jitter during travel can be eliminated, thereby making the subsequent determination of terrain features more accurate.

[0185] In step 203, point cloud data of the target region is generated based on the second depth image and the second field-of-view image.

[0186] In one possible implementation, the process of generating point cloud data of the target region based on the second depth image and the second field-of-view image includes: determining the three-dimensional coordinates of each pixel in the second depth image based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module; acquiring the color information of each pixel in the second field-of-view image; and generating point cloud data of the target region based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field-of-view image, wherein the point cloud data is the point cloud data of the target region at the first moment.

[0187] The process of generating point cloud data of the target region based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field of view image includes: for any pixel in the second depth image, determining the three-dimensional coordinates of any pixel in the three-dimensional coordinates of each pixel in the second depth image, determining the color information of any pixel in the color information of each pixel in the second field of view image, combining the three-dimensional coordinates and color information of any pixel into a data set, traversing each pixel in the second depth image to obtain multiple data sets, and combining the multiple data sets into point cloud data of the target region.

[0188] In one possible implementation, after generating point cloud data of the target area, a three-dimensional terrain model of the target area can be generated based on the point cloud data. After obtaining the three-dimensional terrain model of the target area, it can be sent to the robot's control device so that the control device can control the robot based on the three-dimensional terrain model. The robot's control device can be a device installed on the robot or a device capable of remotely controlling the robot; this embodiment does not limit the specific device used.

[0189] In step 204, the terrain features of the target area are extracted based on the point cloud data. The terrain features are used to characterize the terrain of the target area.

[0190] The terrain features include terrain slope and obstacle location information. Obstacles refer to any obstacles included in the target area.

[0191] In one possible implementation, after obtaining point cloud data based on terrain features, including terrain slope, the process of extracting terrain features of the target area based on the point cloud data includes: performing plane fitting on the point cloud data to obtain the plane equation of the target area; determining the normal vector of the target area based on the plane equation; and determining the angle between the normal vector and the horizontal plane of the target area as the terrain slope of the target area.

[0192] Optionally, a plane fitting algorithm is used to fit the point cloud data to obtain the plane equation of the target region. The fitting algorithm can be the Random Sample And Consensus (RANSAC) algorithm, or other fitting algorithms, which are not limited in this embodiment.

[0193] In one possible implementation, after acquiring point cloud data based on terrain features, including obstacle location information, the process of extracting terrain features of the target area from the point cloud data includes: segmenting the point cloud data to obtain ground points and non-ground points; performing cluster analysis on the non-ground points to obtain at least one set of non-ground points; for any set of non-ground points, if the volume of any set of non-ground points is greater than a volume threshold, determining the target location information based on the location information of any set of non-ground points; and determining the target location information as the location information of obstacles included in the target area. The volume threshold is set based on experience or adjusted according to the implementation environment; this embodiment does not limit its application.

[0194] Optionally, the point cloud data can be segmented using a planar segmentation algorithm to obtain ground points and non-ground points. The non-ground points can then be clustered using either the DBSCAN clustering algorithm (a density-based clustering algorithm) or the k-means clustering algorithm (a distance-based clustering algorithm).

[0195] In one possible implementation, terrain features may further include obstacle types. This application does not limit the method for determining obstacle types. Optionally, the type corresponding to the volume of an obstacle is determined as the obstacle type. Alternatively, based on the obstacle's location information, the image region corresponding to the obstacle is determined in the field-of-view image, and the obstacle type is obtained by identifying the image region corresponding to the obstacle. Obstacle types include, but are not limited to, stones, wood, walls, fences, etc.

[0196] The process of identifying the type of obstacle by recognizing the image region corresponding to the obstacle includes: inputting the image region corresponding to the obstacle into an obstacle type determination model to obtain the type of obstacle. The obstacle type determination model can be any model capable of determining the type of obstacle, and this application embodiment does not limit this type.

[0197] In some embodiments, after determining the terrain features of the target area, the terrain features are sent to the robot's control device, which then controls the robot based on the terrain features. Optionally, the control device determines a gait control strategy corresponding to the terrain features and controls the robot according to the gait control strategy corresponding to the terrain features.

[0198] In one possible implementation, the terrain features of the target area include terrain slope. The control device can also determine the terrain type based on the terrain slope; determine the gait control strategy corresponding to the terrain features and terrain type; and control the robot according to the gait control strategy corresponding to the terrain features and terrain type. The terrain type includes uphill, downhill, and flat land. The process of determining the terrain type based on the terrain slope includes: if the terrain slope is greater than 0 degrees, the terrain type is uphill; if the terrain slope is less than 0 degrees, the terrain type is downhill; and if the terrain slope is 0 degrees, the terrain type is flat land.

[0199] In one possible implementation, the first view-of-view image of the target area acquired by the second image acquisition module at the first moment can also determine the terrain material of the target area, which includes, but is not limited to, grassland, snow, and sand. After acquiring the first view-of-view image of the target area at the first moment, the terrain material of the target area can also be determined based on the first view-of-view image.

[0200] The process of determining the terrain material of the target area based on the first field-of-view image includes: denoising and feature extraction of the first field-of-view image to obtain a reference field-of-view image, and determining the image features of the reference field-of-view image; then, determining the terrain material of the target area based on the image features. Specifically, the process of determining the image features of the reference field-of-view image includes: determining the image features of the reference field-of-view image using convolutional neural networks (CNNs); or, determining the image features of the reference field-of-view image using image processing methods such as color histograms and texture features.

[0201] After obtaining image features, the process of determining the terrain material of the target area based on these features includes: inputting the image features into a terrain material classifier, and outputting the terrain material classifier as the terrain material of the target area. The terrain material classifier is a deep learning model.

[0202] Optionally, after determining the terrain material of the target area, a material label corresponding to the terrain material of the target area can also be determined. The material label is used to uniquely indicate the terrain material of the target area. Optionally, the material label can be the name of the terrain material of the target area. The material label can also be sent to the control device so that the control device can determine the gait control strategy corresponding to the terrain features and the material label, and control the robot according to the gait control strategy corresponding to the terrain features and the material label.

[0203] For example, if the material label is snow, the robot will be controlled according to the gait control strategy corresponding to snow and terrain features to make the robot adapt to a lower coefficient of friction. As another example, if the material label is grass, the robot will be controlled according to the gait control strategy corresponding to grass and terrain features to make the robot's gait smoother.

[0204] In one possible implementation, after determining the terrain material of the target area, a compensation value corresponding to the terrain material of the target area can also be determined; the 3D terrain model of the target area is compensated according to the compensation value to obtain a compensated 3D terrain model, which is more in line with the actual situation of the target area.

[0205] For example, when the terrain material of the target area is grassland, a compensation value corresponding to the grassland is determined, and the rate of change of the depth value of the 3D terrain model of the target area is reduced according to the compensation value corresponding to the grassland to obtain a compensated 3D terrain model, so that the compensated 3D terrain model reflects a relatively smooth ground. As another example, when the terrain material of the target area is snow, a compensation value corresponding to snow is determined, and the surface roughness of the 3D terrain model is adjusted according to the compensation value corresponding to snow to obtain a compensated 3D terrain model, so that the compensated 3D terrain model reflects a rough ground.

[0206] The system can also fine-tune the 3D terrain model of the target area based on its terrain material. Optionally, for soft or irregular terrain materials (such as grass or sand), the hard edges in the 3D terrain model of the target area can be softened to make it more supple. For harder terrain materials (such as stone), the surface details of the 3D terrain model of the target area can be increased to more accurately reflect irregular shapes.

[0207] In one possible implementation, after obtaining the compensated 3D terrain model, it can be sent to the robot's control device, which then controls the robot based on this model. The robot's control device dynamically adjusts its gait control strategy according to the compensated terrain model to adapt the robot to the friction coefficient and hardness of different terrain materials. For example, if the compensated terrain model detects that the target area's terrain material is slippery snow, the control device increases the robot's gait stability or reduces its walking speed to prevent slipping, thereby improving the user's safety.

[0208] The method described above, when determining the terrain features of a target area, considers not only the first inertial data from the inertial sensor at the first moment, but also the first field-of-view image and the first depth image of the target area at the first moment. Based on the first inertial data, the first depth image and the first field-of-view image are processed to determine the terrain features of the target area. This results in higher precision and accuracy of the determined terrain features, making them more consistent with the actual situation of the target area. Since terrain features can characterize the terrain of the target area, more precise and accurate terrain features can better characterize the terrain of the target area. When controlling the robot based on the terrain of the target area, the control precision of the robot is improved, as well as the walking safety of the robot and the walking safety of the user.

[0209] Furthermore, the method provided in this application, by taking into account the depth image of the target area at the first moment, can determine the terrain features of the target area under various lighting conditions, overcoming the problems of being unable to determine terrain features or having low accuracy in determining terrain features under low light conditions. By employing a scheme that fuses field-of-view and depth images, accurate terrain features can be determined even under complex terrain conditions, compensating for the shortcomings of traditional robots that rely solely on inertial sensors and cannot acquire environmental images, thus improving the accuracy of the determined terrain features and the adaptability of the terrain recognition method.

[0210] In addition, terrain features can be fed back to the robot's control equipment to help the control equipment dynamically adjust the robot's gait control strategy, thereby improving the robot's stability and further enhancing the robot's safety as well as the walking safety of the robot's users.

[0211] Figure 3 is a flowchart of a three-dimensional terrain recognition method provided in an embodiment of this application. As shown in Figure 3, the method includes the following steps:

[0212] 1. Acquire the first depth image of the target area at the first moment acquired by the first image acquisition module, the first field-of-view image of the target area at the first moment acquired by the second image acquisition module, and the first inertial data acquired by the inertial sensor.

[0213] In one possible implementation, the acquisition of the first depth image of the target area acquired by the first image acquisition module at the first moment, the first field-of-view image of the target area acquired by the second image acquisition module at the first moment, and the first inertial data acquired by the inertial sensor have been described in step 201 above, and will not be repeated here.

[0214] 2. Detect the first field of view image to obtain the illumination intensity of the first field of view image.

[0215] In one possible implementation, the process of detecting the first field-of-view image and obtaining the illumination intensity of the first field-of-view image has been described in step 202 above, and will not be repeated here.

[0216] 3. If the illumination intensity of the first field of view image meets the intensity requirements, process the first field of view image to obtain a reference field of view image after denoising and feature extraction of the first field of view image.

[0217] In one possible implementation, the process of processing the first field-of-view image to obtain a reference field-of-view image after denoising and feature extraction, provided that the illumination intensity of the first field-of-view image meets the intensity requirements, has been described in step 202 above and will not be repeated here.

[0218] 4. Generate reference point cloud data corresponding to the reference field-of-view image.

[0219] In one possible implementation, the process of generating reference point cloud data corresponding to the reference view image has been described in step 202 above, and will not be repeated here.

[0220] 5. Extract terrain features of the target area based on the reference point cloud data.

[0221] In one possible implementation, the process of extracting the terrain features of the target area at the first moment based on the reference point cloud data has been described in step 202 above, and will not be repeated here.

[0222] 6. If the illumination intensity of the first field-of-view image does not meet the intensity requirements, dynamic compensation is performed on the first depth image and the first field-of-view image based on the first inertial data to obtain the compensated second depth image and the second field-of-view image.

[0223] In one possible implementation, the process of dynamically compensating the first depth image and the first field of view image based on the first inertial data to obtain the compensated second depth image and the second field of view image has been described in step 202 above, and will not be repeated here.

[0224] 7. Generate point cloud data of the target area based on the height of the second depth image and the second field of view image.

[0225] In one possible implementation, the process of generating point cloud data of the target region based on the height of the second depth image and the second field of view image has been described in step 203 above, and will not be repeated here.

[0226] 8. Extract terrain features of the target area based on point cloud data.

[0227] In one possible implementation, the process of extracting the terrain features of the target area based on point cloud data has been described in step 204 above, and will not be repeated here.

[0228] This application embodiment also provides a three-dimensional terrain recognition system, as shown in Figure 4. The system includes a first image acquisition module 401, a second image acquisition module 402, an inertial sensor 403, and a computer device 404.

[0229] The first image acquisition module 401 is used to acquire a first depth image of the target area at a first moment and send the first depth image to the computer device 404.

[0230] The second image acquisition module 402 is used to acquire a first field-of-view image of the target area at a first moment and send the first field-of-view image to the computer device 404.

[0231] Inertial sensor 403 is used to acquire first inertial data at a first moment and send the first inertial data to computer device 404;

[0232] Computer device 404 is used to determine the terrain features of the target area at a first moment based on a first field-of-view image, a first depth image, and first inertial data.

[0233] In one possible implementation, the process of the first image acquisition module 401 acquiring the first depth image has been described in step 201 above, and will not be repeated here.

[0234] In one possible implementation, the process of the second image acquisition module 402 acquiring the first field-of-view image has been described in step 201 above, and will not be repeated here.

[0235] In one possible implementation, the process of the inertial sensor 403 acquiring the first inertial data has been described in step 201 above, and will not be repeated here.

[0236] In one possible implementation, the process by which computer device 404 determines the terrain features of the target area based on the first field-of-view image, the first depth image, and the first inertial data has been described in steps 202 to 204 above, and will not be repeated here.

[0237] Figure 5 shows a schematic diagram of a three-dimensional terrain recognition device provided in an embodiment of this application. As shown in Figure 5, the device includes:

[0238] The acquisition module 501 is used to acquire a first depth image of the target area at a first moment, a first field-of-view image of the target area at a first moment, acquired by the first image acquisition module, acquired by the second image acquisition module, and first inertial data acquired by the inertial sensor. The first depth image is used to indicate the depth information of each pixel in the target area at a first moment, the first field-of-view image is used to indicate the color information of each pixel in the target area at a first moment, and the first inertial data is used to indicate the motion information of the first image acquisition module and the second image acquisition module at a first moment.

[0239] The dynamic compensation module 502 is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data to obtain the compensated second depth image and the second field of view image.

[0240] The generation module 503 is used to generate point cloud data of the target region based on the second depth image and the second field-of-view image;

[0241] The recognition module 504 is used to extract the terrain features of the target area based on the point cloud data. The terrain features are used to characterize the terrain of the target area.

[0242] In one possible implementation, the device further includes:

[0243] The detection module is used to detect the first field-of-view image and obtain the illumination intensity of the first field-of-view image;

[0244] The dynamic compensation module 502 is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data when the illumination intensity of the first field of view image does not meet the intensity requirements, so as to obtain the compensated second depth image and second field of view image.

[0245] In one possible implementation, the first inertial data includes attitude data and acceleration;

[0246] The dynamic compensation module 502 is used to obtain a third depth image and a third field of view image based on the attitude data, the first depth image, and the first field of view image. The third depth image and the third field of view image are located in the reference coordinate system. Based on the acceleration, the third depth image is height compensated to obtain a second depth image, and the third field of view image is height compensated to obtain a second field of view image.

[0247] In one possible implementation, the dynamic compensation module 502 is used to process the first depth image to obtain a reference depth image after denoising the first depth image; process the first field-of-view image to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image; determine a rotation matrix based on the pose data, the rotation matrix being used to transform the reference depth image and the reference field-of-view image to a reference coordinate system; transform the reference depth image based on the rotation matrix to obtain a third depth image; and transform the reference field-of-view image based on the rotation matrix to obtain a third field-of-view image.

[0248] In one possible implementation, the dynamic compensation module 502 is used to acquire the position information and depth information of each first pixel in the reference depth image; update the position information and depth information of each first pixel based on the rotation matrix to obtain the updated position information and depth information of each first pixel; and generate a third depth image based on the updated position information and depth information of each first pixel.

[0249] In one possible implementation, the dynamic compensation module 502 is used to acquire the position information and color information of each second pixel in the reference field image; update the position information and color information of each second pixel based on the rotation matrix to obtain the updated position information and color information of each second pixel; and generate a third field image based on the updated position information and color information of each second pixel.

[0250] In one possible implementation, the dynamic compensation module 502 is used to determine the height change value of the inertial sensor in the reference direction based on the acceleration; and to perform height compensation on each pixel in the third depth image based on the height change value to obtain the second depth image.

[0251] In one possible implementation, the dynamic compensation module 502 is used to integrate the acceleration to obtain the velocity; and to determine the height change value of the inertial sensor in the reference direction based on the velocity and the target duration, wherein the target duration is obtained based on the sampling frequency of the inertial sensor.

[0252] In one possible implementation, the dynamic compensation module 502 is used to add height change values ​​to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image; and to generate a second depth image based on the height-compensated depth information of each pixel in the third depth image.

[0253] In one possible implementation, the generation module 503 is further configured to process the first field-of-view image based on the illumination intensity of the first field-of-view image meeting the intensity requirements, to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image; and to generate reference point cloud data corresponding to the reference field-of-view image.

[0254] The identification module 504 is also used to extract the terrain features of the target area at the first moment based on the reference point cloud data.

[0255] In one possible implementation, the generation module 503 is used to determine the three-dimensional coordinates of each pixel in the second depth image based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module; acquire the color information of each pixel in the second field of view image; and generate point cloud data of the target area based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field of view image.

[0256] In one possible implementation, the generation module 503 is further configured to generate a three-dimensional terrain model of the target area based on the point cloud data of the target area, and the three-dimensional terrain model of the target area is used by the robot's control device to control the robot.

[0257] In one possible implementation, the terrain feature includes terrain slope;

[0258] The recognition module 504 is used to perform plane fitting on the point cloud data to obtain the plane equation of the target area; based on the plane equation, the normal vector of the target area is determined; and the angle between the normal vector and the horizontal plane of the target area is determined as the terrain slope of the target area.

[0259] In one possible implementation, the terrain features include the location information of obstacles;

[0260] The identification module 504 is used to segment the point cloud data to obtain ground points and non-ground points; perform cluster analysis on the non-ground points to obtain at least one set of non-ground points; for any set of non-ground points, if the volume of any set of non-ground points is greater than a volume threshold, determine the target location information based on the location information of any set of non-ground points; the target location information is determined to be the location information of obstacles included in the target area.

[0261] In one possible implementation, the device further includes:

[0262] The determination module is used to determine the terrain material of the target area based on the first field-of-view image.

[0263] When determining the terrain features of a target area, the aforementioned device considers not only the first inertial data from the inertial sensor at the first moment, but also the first field-of-view image and the first depth image of the target area at the first moment. Based on the first inertial data, the first depth image and the first field-of-view image are processed to determine the terrain features of the target area. This results in higher precision and accuracy of the determined terrain features, making them more consistent with the actual situation of the target area. Since terrain features can characterize the terrain of the target area, more precise and accurate terrain features can better characterize the terrain of the target area. When controlling the robot based on the terrain of the target area, the control precision of the robot is improved, as well as the walking safety of the robot and the walking safety of the robot user are also enhanced.

[0264] It should be understood that the above-described apparatus is only illustrated by the division of the functional modules described above when implementing its functions. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0265] Figure 6 shows a structural block diagram of a terminal device 600 provided in an exemplary embodiment of this application. The terminal device 600 can be any electronic device product capable of human-computer interaction with a user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablet computers, smart car systems, smart TVs, smart speakers, and smartwatches.

[0266] Typically, terminal device 600 includes a processor 601 and a memory 602.

[0267] Processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0268] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 are used to store at least one instruction, which is executed by the processor 601 to implement the three-dimensional terrain recognition method provided in the method embodiments of this application.

[0269] In some embodiments, the terminal device 600 may also optionally include a peripheral device interface 603 and at least one peripheral device. The processor 601, memory 602, and peripheral device interface 603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 604, a display screen 605, a camera assembly 606, an audio circuit 607, and a power supply 608.

[0270] Peripheral interface 603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 601 and memory 602. In some embodiments, processor 601, memory 602 and peripheral interface 603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 601, memory 602 and peripheral interface 603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0271] The radio frequency (RF) circuit 604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 604 can communicate with other terminal devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 604 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0272] Display screen 605 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 601 for processing. In this case, display screen 605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 605, disposed on the front panel of terminal device 600; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal device 600 or in a folded design; in still other embodiments, display screen 605 may be a flexible display screen, disposed on a curved or folded surface of terminal device 600. Furthermore, display screen 605 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 605 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0273] The camera assembly 606 is used to acquire images or videos. Optionally, the camera assembly 606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal device 600, and the rear-facing camera is located on the back of the terminal device 600. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0274] The audio circuit 607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 601 for processing, or input to the radio frequency circuit 604 to achieve voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 600. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 601 or the radio frequency circuit 604 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 607 may also include a headphone jack.

[0275] Power supply 608 is used to supply power to the various components in terminal device 600. Power supply 608 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 608 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0276] In some embodiments, the terminal device 600 further includes one or more sensors 609. The one or more sensors 609 include, but are not limited to, an accelerometer 610, a gyroscope 611, a pressure sensor 612, an optical sensor 613, and a proximity sensor 614.

[0277] Accelerometer 610 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 600. For example, accelerometer 610 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 601 can control display screen 605 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 610. Accelerometer 610 can also be used for games or for acquiring user motion data.

[0278] The gyroscope sensor 611 can detect the orientation and rotation angle of the terminal device 600. The gyroscope sensor 611 can work in conjunction with the accelerometer sensor 610 to collect the user's 3D movements on the terminal device 600. Based on the data collected by the gyroscope sensor 611, the processor 601 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0279] The pressure sensor 612 can be disposed on the side bezel of the terminal device 600 and / or on the lower layer of the display screen 605. When the pressure sensor 612 is disposed on the side bezel of the terminal device 600, it can detect the user's grip signal on the terminal device 600, and the processor 601 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 612. When the pressure sensor 612 is disposed on the lower layer of the display screen 605, the processor 601 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0280] An optical sensor 613 is used to collect ambient light intensity. In one embodiment, the processor 601 can control the display brightness of the display screen 605 based on the ambient light intensity collected by the optical sensor 613. Specifically, when the ambient light intensity is high, the display brightness of the display screen 605 is increased; when the ambient light intensity is low, the display brightness of the display screen 605 is decreased. In another embodiment, the processor 601 can also dynamically adjust the shooting parameters of the camera assembly 606 based on the ambient light intensity collected by the optical sensor 613.

[0281] The proximity sensor 614, also known as a distance sensor, is typically mounted on the front panel of the terminal device 600. The proximity sensor 614 is used to detect the distance between the user and the front of the terminal device 600. In one embodiment, when the proximity sensor 614 detects that the distance between the user and the front of the terminal device 600 is gradually decreasing, the processor 601 controls the display screen 605 to switch from a screen-on state to a screen-off state; when the proximity sensor 614 detects that the distance between the user and the front of the terminal device 600 is gradually increasing, the processor 601 controls the display screen 605 to switch from a screen-off state to a screen-on state.

[0282] Those skilled in the art will understand that the structure shown in FIG6 does not constitute a limitation on the terminal device 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0283] Figure 7 is a schematic diagram of the server structure provided in the embodiments of this application. The server 700 can vary considerably due to different configurations or performance. It may include one or more processors (Central Processing Units, CPUs) 701 and one or more memories 702. The one or more memories 702 store at least one piece of program code, which is loaded and executed by the one or more processors 701 to implement the three-dimensional terrain recognition method provided in the above-described method embodiments. Of course, the server 700 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 700 may also include other components for implementing device functions, which will not be elaborated here.

[0284] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one piece of program code that is loaded and executed by a processor to enable a computer to implement any of the three-dimensional terrain recognition methods described above.

[0285] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0286] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement any of the three-dimensional terrain recognition methods described above.

[0287] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0288] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0289] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A three-dimensional terrain recognition method, wherein, The method includes: The system acquires a first depth image of the target region acquired by a first image acquisition module at a first moment, a first field-of-view image of the target region acquired by a second image acquisition module at the first moment, and first inertial data acquired by an inertial sensor. The first depth image is used to indicate the depth information of each pixel in the target region at the first moment, the first field-of-view image is used to indicate the color information of each pixel in the target region at the first moment, and the first inertial data is used to indicate the motion information of the first image acquisition module and the second image acquisition module at the first moment. Based on the first inertial data, the first depth image and the first field of view image are dynamically compensated to obtain the compensated second depth image and the second field of view image. Point cloud data of the target region is generated based on the second depth image and the second field-of-view image; Based on the point cloud data, the terrain features of the target area are extracted, and the terrain features are used to characterize the terrain of the target area.

2. The method according to claim 1, wherein, The method further includes: The first field-of-view image is detected to obtain the illumination intensity of the first field-of-view image; The step of dynamically compensating the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image includes: Since the illumination intensity of the first field-of-view image does not meet the intensity requirements, dynamic compensation is performed on the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image.

3. The method according to claim 2, wherein, The first inertial data includes attitude data and acceleration; The step of dynamically compensating the first depth image and the first field-of-view image based on the first inertial data to obtain a compensated second depth image and a second field-of-view image includes: Based on the pose data, the first depth image, and the first field of view image, a third depth image and a third field of view image are obtained, wherein the third depth image and the third field of view image are located in a reference coordinate system; Based on the acceleration, the third depth image is height-compensated to obtain the second depth image, and the third field-of-view image is height-compensated to obtain the second field-of-view image.

4. The method according to claim 3, wherein, The step of obtaining a third depth image and a third field of view image based on the pose data, the first depth image, and the first field of view image includes: Process the first depth image to obtain a reference depth image after denoising the first depth image; The first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image; Based on the attitude data, a rotation matrix is ​​determined, which is used to transform the reference depth image and the reference field of view image to the reference coordinate system; The reference depth image is transformed according to the rotation matrix to obtain the third depth image; The reference field-of-view image is transformed according to the rotation matrix to obtain the third field-of-view image.

5. The method according to claim 4, wherein, The step of transforming the reference depth image according to the rotation matrix to obtain the third depth image includes: Obtain the position and depth information of each first pixel in the reference depth image; Based on the rotation matrix, update the position and depth information of each first pixel to obtain the updated position and depth information of each first pixel. The third depth image is generated based on the updated position and depth information of each first pixel.

6. The method according to claim 4, wherein, The step of transforming the reference field-of-view image according to the rotation matrix to obtain the third field-of-view image includes: Obtain the position and color information of each second pixel in the reference field image; Based on the rotation matrix, update the position information and color information of each second pixel to obtain the updated position information and color information of each second pixel. The third field-view image is generated based on the updated position and color information of each second pixel.

7. The method according to claim 3, wherein, The step of performing height compensation on the third depth image based on the acceleration to obtain the second depth image includes: Based on the acceleration, determine the height change value of the inertial sensor in the reference direction; Based on the height change value, height compensation is performed on each pixel in the third depth image to obtain the second depth image.

8. The method according to claim 7, wherein, Determining the height change value of the inertial sensor in the reference direction based on the acceleration includes: Integrating the acceleration yields the velocity; Based on the speed and the target duration, the height change value of the inertial sensor in the reference direction is determined, and the target duration is obtained based on the sampling frequency of the inertial sensor.

9. The method according to claim 7, wherein, The step of performing height compensation on each pixel in the third depth image based on the height change value to obtain the second depth image includes: The height change value is added to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image; The second depth image is generated based on the depth information of each pixel in the third depth image after height compensation.

10. The method according to claim 2, wherein, The method further includes: Based on the fact that the illumination intensity of the first field-of-view image meets the intensity requirement, the first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image. Generate reference point cloud data corresponding to the reference field-of-view image; Based on the reference point cloud data, the terrain features of the target area at the first time point are extracted.

11. The method according to any one of claims 1 to 10, wherein, The step of generating point cloud data of the target region based on the second depth image and the second field-of-view image includes: Based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module, the three-dimensional coordinates of each pixel in the second depth image are determined. Obtain the color information of each pixel in the second field-of-view image; Point cloud data of the target region is generated based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field of view image.

12. The method according to any one of claims 1 to 10, wherein, After generating point cloud data of the target region based on the second depth image and the second field-of-view image, the method further includes: Based on the point cloud data of the target area, a three-dimensional terrain model of the target area is generated. The three-dimensional terrain model of the target area is used by the robot's control device to control the robot.

13. The method according to any one of claims 1 to 10, wherein, The terrain features include terrain slope; The step of extracting terrain features of the target area based on the point cloud data includes: The point cloud data is fitted to a plane to obtain the plane equation of the target region; Determine the normal vector of the target region based on the plane equation; The angle between the normal vector and the horizontal plane of the target area is determined as the terrain slope of the target area.

14. The method according to any one of claims 1 to 10, wherein, The terrain features include the location information of obstacles; The step of extracting terrain features of the target area based on the point cloud data includes: The point cloud data is segmented to obtain ground points and non-ground points; Cluster analysis is performed on the non-ground points to obtain at least one set of non-ground points; For any set of non-ground points, if the volume of any set of non-ground points is greater than a volume threshold, the target location information is determined based on the location information of any set of non-ground points. The target location information is determined to be the location information of obstacles included in the target area.

15. The method according to any one of claims 1 to 10, wherein, The method further includes: Based on the first field-of-view image, the terrain material of the target area is determined.

16. A three-dimensional terrain recognition device, wherein, The device includes: The acquisition module is used to acquire a first depth image of the target area acquired by the first image acquisition module at a first moment, a first field-of-view image of the target area acquired by the second image acquisition module at the first moment, and first inertial data acquired by the inertial sensor. The first depth image is used to indicate the depth information of each pixel in the target area at the first moment, the first field-of-view image is used to indicate the color information of each pixel in the target area at the first moment, and the first inertial data is used to indicate the motion information of the first image acquisition module and the second image acquisition module at the first moment. The dynamic compensation module is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data to obtain a compensated second depth image and a second field of view image. The generation module is used to generate point cloud data of the target region based on the second depth image and the second field-of-view image; The identification module is used to extract the terrain features of the target area based on the point cloud data, and the terrain features are used to characterize the terrain of the target area.

17. The apparatus according to claim 16, wherein, The device further includes: The detection module is used to detect the first field-of-view image and obtain the illumination intensity of the first field-of-view image; The dynamic compensation module is used to dynamically compensate the first depth image and the first field of view image based on the first inertial data, since the illumination intensity of the first field of view image does not meet the intensity requirements, to obtain a compensated second depth image and a second field of view image.

18. The apparatus according to claim 17, wherein, The first inertial data includes attitude data and acceleration; The dynamic compensation module is used to obtain a third depth image and a third field of view image based on the attitude data, the first depth image, and the first field of view image, wherein the third depth image and the third field of view image are located in a reference coordinate system; Based on the acceleration, the third depth image is height-compensated to obtain the second depth image, and the third field-of-view image is height-compensated to obtain the second field-of-view image.

19. The apparatus according to claim 18, wherein, The dynamic compensation module is used to process the first depth image to obtain a reference depth image after denoising the first depth image. The first field-of-view image is processed to obtain a reference field-of-view image after denoising and feature extraction of the first field-of-view image; Based on the attitude data, a rotation matrix is ​​determined, which is used to transform the reference depth image and the reference field of view image to the reference coordinate system; The reference depth image is transformed according to the rotation matrix to obtain the third depth image; The reference field-of-view image is transformed according to the rotation matrix to obtain the third field-of-view image.

20. The apparatus according to claim 19, wherein, The dynamic compensation module is used to obtain the position and depth information of each first pixel in the reference depth image; Based on the rotation matrix, update the position and depth information of each first pixel to obtain the updated position and depth information of each first pixel. The third depth image is generated based on the updated position and depth information of each first pixel.

21. The apparatus according to claim 19, wherein, The dynamic compensation module is used to obtain the position information and color information of each second pixel in the reference field image; Based on the rotation matrix, update the position information and color information of each second pixel to obtain the updated position information and color information of each second pixel. The third field-view image is generated based on the updated position and color information of each second pixel.

22. The apparatus according to claim 18, wherein, The dynamic compensation module is used to determine the height change value of the inertial sensor in the reference direction based on the acceleration. Based on the height change value, height compensation is performed on each pixel in the third depth image to obtain the second depth image.

23. The apparatus according to claim 22, wherein, The dynamic compensation module is used to integrate the acceleration to obtain the velocity; Based on the speed and the target duration, the height change value of the inertial sensor in the reference direction is determined, and the target duration is obtained based on the sampling frequency of the inertial sensor.

24. The apparatus according to claim 22, wherein, The dynamic compensation module is used to add the height change value to the depth information of each pixel in the third depth image to obtain the height-compensated depth information of each pixel in the third depth image. The second depth image is generated based on the depth information of each pixel in the third depth image after height compensation.

25. The apparatus according to claim 17, wherein, The generation module is further configured to process the first field-view image based on the illumination intensity of the first field-view image meeting the intensity requirement, and obtain a reference field-view image after denoising and feature extraction of the first field-view image. Generate reference point cloud data corresponding to the reference field-of-view image; The identification module is further configured to extract the terrain features of the target area at the first moment based on the reference point cloud data.

26. The apparatus according to any one of claims 16 to 25, wherein, The generation module is used to determine the three-dimensional coordinates of each pixel in the second depth image based on the depth information of each pixel in the second depth image and the intrinsic parameters of the first image acquisition module. Obtain the color information of each pixel in the second field-of-view image; Point cloud data of the target region is generated based on the three-dimensional coordinates of each pixel in the second depth image and the color information of each pixel in the second field of view image.

27. The apparatus according to any one of claims 16 to 25, wherein, The generation module is further configured to generate a three-dimensional terrain model of the target area based on the point cloud data of the target area, and the three-dimensional terrain model of the target area is used by the robot's control device to control the robot.

28. The apparatus according to any one of claims 16 to 25, wherein, The terrain features include terrain slope; The recognition module is used to perform plane fitting on the point cloud data to obtain the plane equation of the target region; Determine the normal vector of the target region based on the plane equation; The angle between the normal vector and the horizontal plane of the target area is determined as the terrain slope of the target area.

29. The apparatus according to any one of claims 16 to 25, wherein, The terrain features include the location information of obstacles; The identification module is used to segment the point cloud data to obtain ground points and non-ground points; Cluster analysis is performed on the non-ground points to obtain at least one set of non-ground points; For any set of non-ground points, if the volume of any set of non-ground points is greater than a volume threshold, the target location information is determined based on the location information of any set of non-ground points. The target location information is determined to be the location information of obstacles included in the target area.

30. The apparatus according to any one of claims 16 to 25, wherein, The device further includes: The determination module is used to determine the terrain material of the target area based on the first field-of-view image.

31. A computer device, wherein, The computer device includes a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to enable the computer device to implement the three-dimensional terrain recognition method as described in any one of claims 1 to 15.

32. A computer-readable storage medium, wherein, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to enable the computer to implement the three-dimensional terrain recognition method as described in any one of claims 1 to 15.

33. A computer program product, wherein, The computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement the three-dimensional terrain recognition method as described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method for discretizing complex road condition into footholds of multi-foot robot based on three-dimensional imaging

    CN107644441A

  • Robot control method and system based on terrain trafficability

    CN115185266A

  • Terrain and force fused quadruped robot reachability map construction method and system

    CN116147642A

  • Self-adaptive environment reconstruction and obstacle detection method for autonomous fire-fighting robot

    CN116360423A

  • Modeling method and apparatus using three-dimensional (3D) point cloud

    US20180211399A1