Ground contact behavior identification method and apparatus, device, medium, and product

By combining depth maps and field-of-view images, the distance difference between the object and the ground is identified, solving the problems of inconvenient equipment and limited application scenarios in traditional ground contact detection methods, and realizing non-contact, high-accuracy ground contact detection.

WO2026092134A1PCT designated stage Publication Date: 2026-05-07HYPERSHELL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HYPERSHELL CO LTD
Filing Date
2025-10-14
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Traditional ground contact detection methods rely on physical sensors, which are inconvenient to wear and have limited application scenarios, making it difficult to meet the needs of non-contact, long-distance detection.

Method used

By acquiring depth maps and field-of-view images, and utilizing the distance information from the depth maps and the texture information from the field-of-view images, the distance difference between the object and the ground is identified, thus achieving ground contact detection.

Benefits of technology

It achieves non-contact ground contact detection, allowing the device to be configured in any location, ensuring the accuracy and stability of the detection results, and meeting real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127638_07052026_PF_FP_ABST
    Figure CN2025127638_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A ground contact behavior identification method and apparatus, a device, a medium, and a product. The method comprises: obtaining a depth map acquired by a first image acquisition module at a first moment, and a field-of-view image acquired by a second image acquisition module at the first moment (210); on the basis of distance information represented by the depth map, and texture information represented by the field-of-view image, identifying first distance information corresponding to a first object among scene objects, and second distance information corresponding to a ground object among the scene objects (220); and determining a ground contact detection result of the first object at the first moment on the basis of a distance difference between the first distance information and the second distance information (230).
Need to check novelty before this filing date? Find Prior Art

Description

Ground contact behavior recognition methods, devices, equipment, media and products

[0001] This application claims priority to Chinese Patent Application No. 202411560677.0, filed on November 4, 2024, entitled “Method, Apparatus, Device, Medium and Product for Recognizing Ground Touch Behavior”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a method, apparatus, device, medium and product for recognizing ground contact behavior. Background Technology

[0003] Real-time prediction and recognition of foot contact behavior is of great significance in fields such as gait analysis, motion monitoring, human behavior recognition, and robot control.

[0004] Traditional ground contact detection methods mainly rely on physical sensors such as pressure sensors and accelerometers. That is, sensors placed on the foot are used to detect pressure and velocity data in real time, and then the foot is in contact with the ground or off the ground is determined based on the pressure and velocity data.

[0005] However, the sensors in the above methods need to directly contact the human body or the sole of the shoe, which has problems such as inconvenience in wearing the device and limited application scenarios, making it difficult to meet the needs of non-contact, long-distance detection. Summary of the Invention

[0006] This application provides a method, apparatus, device, medium, and product for recognizing ground contact behavior. The technical solution is as follows:

[0007] On the one hand, a method for recognizing ground contact behavior is provided, the method comprising:

[0008] The depth map acquired by the first image acquisition module at a first moment and the field-view image acquired by the second image acquisition module at the first moment are obtained. The depth map includes distance information between the scene object and the first image acquisition module, and the field-view image includes texture information of the scene.

[0009] Based on the distance information represented by the depth map and the texture information represented by the field of view image, a first distance information corresponding to a first object in the scene object and a second distance information corresponding to a ground object in the scene object are identified; wherein, there is a ground contact relationship between the first object and the ground object, the first distance information is used to express the first distance between the identified first object and the first image acquisition module, and the second distance information is used to express the second distance between the identified ground object and the first image acquisition module;

[0010] Based on the distance difference between the first distance information and the second distance information, the ground contact detection result of the first object at the first moment is determined.

[0011] On the other hand, a ground-touching behavior recognition device is provided, the device comprising:

[0012] The acquisition module is used to acquire a depth map acquired by the first image acquisition module at a first moment, and a field-of-view image acquired by the second image acquisition module at the first moment. The depth map includes distance information between the scene object and the first image acquisition module, and the field-of-view image includes texture information of the scene.

[0013] The recognition module is used to recognize, based on the distance information represented by the depth map and the texture information represented by the field-of-view image, a first distance information corresponding to a first object in the scene objects and a second distance information corresponding to a ground object in the scene objects; wherein, there is a ground contact relationship between the first object and the ground object, the first distance information is used to express the first distance between the recognized first object and the first image acquisition module, and the second distance information is used to express the second distance between the recognized ground object and the first image acquisition module;

[0014] The judgment module is used to determine the ground contact detection result of the first object at the first moment based on the distance difference between the first distance information and the second distance information.

[0015] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the ground contact behavior recognition method as described in any of the embodiments of this application above.

[0016] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the ground contact behavior recognition method as described in any of the embodiments of this application above.

[0017] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the ground contact behavior recognition methods described in the above embodiments.

[0018] The technical solution provided in this application includes at least the following beneficial effects:

[0019] When performing ground contact detection on the first object, the distance information represented by the depth map and the texture information represented by the field-of-view image, acquired simultaneously, is used to identify the distance information between the first object to be detected and the ground in the scene. The difference between these two distances is used to determine whether the first object is in contact with the ground, thus obtaining the ground contact detection result. In other words, since ground contact detection can be achieved using only the depth map and field-of-view image provided by the image acquisition module, the geodesic detection recognition device can be configured at any location in the scene where the first object can be captured. This satisfies the need for non-contact ground contact detection scenarios. Furthermore, the reliable depth information provided by the depth map and the rich visual information provided by the field-of-view image ensure stable output of ground contact detection results, guaranteeing accuracy even in real-time detection scenarios. Attached Figure Description

[0020] Figure 1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;

[0021] Figure 2 is a flowchart of a ground-touching behavior recognition method provided in an exemplary embodiment of this application;

[0022] Figure 3 is a flowchart of a ground-touching behavior recognition method provided in an exemplary embodiment of this application;

[0023] Figure 4 is a flowchart of a ground-touching behavior recognition method provided in an exemplary embodiment of this application;

[0024] Figure 5 is a schematic diagram of the structure of a ground contact detection functional module provided in an exemplary embodiment of this application;

[0025] Figure 6 is a structural block diagram of a ground-touching behavior recognition device provided in an exemplary embodiment of this application;

[0026] Figure 7 is a structural block diagram of a terminal device provided in an exemplary embodiment of this application. Detailed Implementation

[0027] First, a brief introduction to the terms used in the embodiments of this application will be given.

[0028] Depth cameras, also known as 3D cameras, are devices that capture distance information of objects in a scene. They acquire the three-dimensional coordinates of objects using various technologies, such as structured light, time-of-flight (ToF), and binocular vision, thereby achieving depth perception. Depth cameras have wide applications in many fields, including but not limited to facial recognition technology, 3D reconstruction, robotics and automation, and augmented reality (AR).

[0029] ToF depth camera: It directly generates depth information by measuring the time it takes for light to travel from the camera lens to the object and back.

[0030] A color (Red, Green, Blue, RGB) camera is an imaging device capable of capturing the three primary colors of light: red (R), green (G), and blue (B). It receives light from a scene through a lens and uses millions of tiny photosensitive elements on its internal image sensor (e.g., a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS)) to convert the light into electrical signals. Each photosensitive element corresponds to a pixel and is sensitive to only one of the three RGB colors. A color filter array (such as a Bayer color filter array) decomposes the received light signal into the three RGB color channels. During imaging, the RGB camera combines the RGB values ​​of each pixel to form a full-color image. This image can display rich colors and details because the human eye is most sensitive to these three colors; the RGB camera can simulate the human eye's perception of color. Finally, these electrical signals are converted from analog to digital and processed by image processing, then stored or transmitted in real-time as a digital image, widely used in photography, video conferencing, surveillance systems, and many other fields. RGB cameras typically produce 24-bit true color images, capable of displaying up to 16.77 million colors, resulting in vibrant and richly detailed images.

[0031] An Inertial Measurement Unit (IMU) is a device used to measure and report specific forces, angular velocities, and, in some cases, magnetic fields around an object. The IMU operates on the principles of Newton's laws of motion: an accelerometer measures the object's acceleration, and a gyroscope measures its angular velocity. By integrating these measurements, the object's velocity and position can be calculated, thus determining its attitude and trajectory in space. An IMU typically consists of an accelerometer, a gyroscope, and a magnetometer. Depending on its onboard measuring elements, the IMU can output data such as acceleration, angular velocity, magnetic field strength, and attitude angles. IMUs are widely used in aerospace, robotics, automotive, and mobile phone industries for navigation, positioning, and attitude control.

[0032] Figure 1 illustrates a schematic diagram of an implementation environment provided by an exemplary embodiment of this application. The implementation environment 100 includes: an identification device 110.

[0033] The identification device 110 is a device that provides ground contact behavior recognition function. Schematic, the identification device 110 has an application that supports ground contact behavior recognition function installed and running, which will call the ground contact behavior recognition function when ground contact detection is required.

[0034] The device type of identification device 110 includes at least one of the following: game console, desktop computer, smartphone, tablet computer, laptop computer, virtual reality (VR) head-mounted display, robot, rehabilitation training instrument, medical instrument, and exoskeleton device.

[0035] In some optional embodiments, the recognition device 110 is equipped with a first image acquisition module 111 and a second image acquisition module 112. The first image acquisition module 111 is an image acquisition module composed of a depth camera, and the second image acquisition module 112 is an image acquisition module composed of an RGB camera or a grayscale camera. Alternatively, the first image acquisition module 111 and the second image acquisition module 112 can be integrated into a single image acquisition module, such as an RGB-D camera; this application does not impose any limitations.

[0036] In one example, the ground contact behavior recognition method provided in this application embodiment is executed independently by the recognition device 110. Illustratively, the recognition device 110 acquires a depth map through a first image acquisition module 111 and a field-of-view image through a second image acquisition module 112. Based on the distance information represented by the depth map and the texture information represented by the field-of-view image, the recognition device 110 identifies first distance information corresponding to the foot of the detected object in the scene and second distance information corresponding to the ground in the scene. Based on the distance difference between the first and second distance information, it determines whether the foot of the detected object is in contact with the ground at a first moment, thereby achieving ground contact detection.

[0037] In some alternative embodiments, the first image acquisition module 111 and the second image acquisition module 112 are independently configured acquisition devices, and the identification device 110 is connected to the acquisition device to acquire depth maps and field-of-view maps. Optionally, the identification device 110 and the acquisition device are connected via at least one of Bluetooth, ZigBee Technology (ZigBee), Wireless Fidelity (WiFi), and a data cable.

[0038] In some alternative embodiments, the implementation environment further includes a server 120. The ground contact behavior recognition method provided in this application embodiment is executed collaboratively by the recognition device 110 and the server 120. Illustratively, the recognition device 110 acquires depth maps and field-of-view images by calling the first image acquisition module 111 and the second image acquisition module 112, and uploads the acquired depth maps and field-of-view images to the server 120. The server 120, by calling the ground contact behavior recognition service, identifies the first distance information corresponding to the foot of the detected object in the scene and the second distance information corresponding to the ground in the scene based on the distance information represented by the depth map and the texture information represented by the field-of-view image. Based on the distance difference between the first distance information and the second distance information, it determines whether the foot of the detected object is in contact with the ground at the first moment, obtains the ground contact detection result, and sends the ground contact detection result to the recognition device 110, or sends the ground contact detection result to a remote terminal for data analysis.

[0039] Optionally, server 120 may include at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, server 120 undertakes the primary computing task, and identification device 110 undertakes the secondary computing task; or, server 120 undertakes the secondary computing task, and identification device 110 undertakes the primary computing task; or, server 120 and identification device 110 collaborate on computing using a distributed computing architecture.

[0040] It is worth noting that the aforementioned server 120 can be implemented as a physical server or as a cloud server in the cloud. In some embodiments, the aforementioned server 120 can also be implemented as a node in a blockchain system.

[0041] Those skilled in the art will understand that the number of the aforementioned devices can be more or less. For example, there may be only one device, or there may be dozens or hundreds of devices, or even more. This application does not limit the number or type of devices.

[0042] Based on the above-mentioned terminology and implementation environment, the ground contact behavior recognition method provided in this application will be described. Taking the method being executed by a recognition device as an example, as shown in Figure 2, the method includes the following steps 210 to 230.

[0043] Step 210: Obtain the depth map acquired by the first image acquisition module at the first moment, and the field-of-view image acquired by the second image acquisition module at the first moment.

[0044] A depth map is a two-dimensional image used to represent the distance between objects in a scene and the first image acquisition module; that is, the depth map includes distance information between scene objects and the first image acquisition module. In this embodiment, the depth map is acquired by the first image acquisition module, which is implemented as a depth camera. Optionally, the depth camera includes at least one of a binocular vision depth camera, a structured light depth camera, a ToF depth camera, and a light field camera.

[0045] The field-of-view image includes texture information of the captured scene. Texture information describes the spatial arrangement of colors or intensities in the image, reflecting the visual characteristics of homogeneous phenomena and embodying the structural arrangement attributes of object surfaces. In this embodiment, the field-of-view image is acquired by a second image acquisition module, which is implemented as an RGB camera or a grayscale camera. Optionally, when the second image acquisition module is implemented as an RGB camera, the field-of-view image includes color and texture information of the scene; that is, the field-of-view image is implemented as an RGB image. Optionally, when the second image acquisition module is implemented as a grayscale camera, the field-of-view image includes brightness and texture information of the scene; that is, the field-of-view image is implemented as a grayscale image.

[0046] In some embodiments, the first image acquisition module and the second image acquisition module described above can be implemented as the same image acquisition module. For example, an RGB-D camera is a device that can simultaneously capture color images (RGB) and depth information (Depth), that is, it acquires depth maps and field-of-view images (RGB images) through an RGB-D camera.

[0047] In this embodiment, the aforementioned scene objects refer to objects within the field of view of the scene captured by the first and second image acquisition modules during image acquisition. Optionally, the aforementioned scene objects include items, ground, buildings, people, animals, plants, etc., in the scene.

[0048] In some embodiments, the first image acquisition module and the second image acquisition module are set to the same image acquisition frequency and the same startup time, so that the first image acquisition module and the second image acquisition module acquire depth map and field of view image at the same time.

[0049] In other embodiments, a first image acquisition module continuously records a first scene video, a second image acquisition module continuously records a second scene video, a depth map at a first moment is extracted from the first scene video, and a field-of-view image at a first moment is extracted from the second scene video.

[0050] In some embodiments, the acquired field-view image and depth map are preprocessed. Optionally, the data preprocessing includes at least one of depth map filtering and field-view image enhancement. Illustratively, depth map filtering is implemented by applying a noisy Gaussian filter to the depth map to remove environmental noise and improve the stability of the depth data; field-view image enhancement is implemented by adjusting the brightness and contrast of the field-view image to ensure visual clarity under different lighting conditions.

[0051] Step 220: Based on the distance information represented by the depth map and the texture information represented by the field of view image, identify the first distance information corresponding to the first object in the scene and the second distance information corresponding to the ground object in the scene.

[0052] In this embodiment, the first object is the object to be detected for ground contact in the scene. In one example, when the recognition device is implemented as a rehabilitation training instrument or a VR headset, the first object is the foot of the user wearing the rehabilitation training instrument or VR headset; in another example, when the recognition device is implemented as a robot, the first object is the robot's mechanical foot; in yet another example, when the recognition device is implemented as an animal behavior detection device, the first object is the foot of the animal wearing the animal behavior detection device.

[0053] The ground object is the accessible area of ​​the first object in the scene. There is a ground contact relationship between the first object and the ground object. The ground contact relationship is used to indicate the contact status between the first object and the ground object. The ground contact relationship includes at least two of the following: direct contact, dummy contact, and no contact. Among them, the direct contact relationship indicates that the center of gravity of the detected object is on the first object and the first object is in contact with the ground object. The dummy contact relationship indicates that the first object is in contact with the ground object, but the center of gravity of the detected object is not on the first object (e.g., the foot is not firmly planted). The no contact state indicates that the first object and the ground object are not in contact.

[0054] Optionally, the first image acquisition module and the second image acquisition module can be configured on the main body of the first object. For example, a VR headset worn by a user can realize ground-touching behavior recognition. The first and second image acquisition modules are mounted on the VR headset, and the viewing angles of the first and second image acquisition modules are downward. Optionally, the first and second image acquisition modules can be configured in a scene. Schematic, a recognition device is set at a first position in the scene, and the viewing angles of the first and second image acquisition modules on the recognition device are directed towards a second position where the main body of the first object is located; that is, the second position is within the field of view of the first and second image acquisition modules.

[0055] In this embodiment, the first distance information is used to express the first distance between the identified first object and the first image acquisition module, and the second distance information is used to express the second distance between the identified ground object and the first image acquisition module.

[0056] In some embodiments, a first object region where a first object is located in a field-of-view image and a first ground region where a ground object is located are identified. Based on the depth map and the field-of-view image, first distance information corresponding to the first object region and second distance information corresponding to the first ground region are determined.

[0057] Optionally, the identification method for the first object region and the first ground region can be implemented as at least one of the following:

[0058] The first method identifies the regions where the first object and the ground object are located by using the texture information represented by the field-of-view image.

[0059] In a schematic way, based on the texture information represented by the field of view image, the first object region corresponding to the first object in the field of view image and the first ground region corresponding to the ground object are identified, wherein the first object region is the image region where the first object is located in the field of view image and the first ground region is the image region where the ground object is located in the field of view image.

[0060] In other words, by using the texture information represented by the field-of-view image, the first object region where the first object is located and the first ground region where the ground object is located are identified in the field-of-view image. Texture information contains rich details, and different objects have different texture information. For example, the instep of a human object (the first object) may contain the texture of shoe material, while the texture of a ground object may show the graininess of the road surface, cracks, etc. Therefore, by quickly identifying the first object region where the first object is located and the first ground region where the ground object is located through the texture information represented by the field-of-view image, it is possible to quickly identify the regions where different objects are located with a single image, thereby improving the detection efficiency of ground contact detection.

[0061] Optionally, the size and shape of the identified first ground region are determined based on the ground distribution in the field of view image. That is, any area in the field of view image that belongs to a ground object is classified as the first ground region. Optionally, the size and shape of the identified first ground region are preset. In some embodiments, the preset size and shape are the shape and size that cover the first object region. That is, only the ground objects within a specified range that the first object touches are segmented and identified. The first ground region is a region larger than the first object region.

[0062] In some embodiments, a pre-trained region recognition model is used to identify and segment the field-of-view image, thereby determining the first object region where the first object is located and the first ground region where the ground object is located. Illustratively, the field-of-view image is input into the region recognition model, which identifies the first object region where the first object is located and the first ground region where the ground object is located based on the texture information represented by the field-of-view image. The region recognition model is a pre-trained machine learning model used to analyze the spatial location information expressed by the depth map and the texture information expressed by the field-of-view image to identify the first object and the ground object from scene objects.

[0063] Optionally, the aforementioned region recognition model can be implemented using neural network models such as Convolutional Neural Networks (CNN), Feedforward Neural Network (FNN), Residual Network (ResNet), Transformer, and U-Net, without specific limitations.

[0064] In one example, the region recognition model can be implemented as a lightweight U-Net model with depthwise separable convolutions. The lightweight U-Net model with depthwise separable convolutions is a lightweight neural network model that combines depthwise separable convolutions and the U-Net network structure. It aims to reduce model parameters and computational cost while maintaining high performance. Depthwise separable convolutions are a type of convolution operation that decomposes standard convolution into depthwise convolution (convolution over the input channels) and pointwise convolution (1x1 convolution over the result of the depthwise convolution to change the number of channels). The lightweight U-Net model is typically implemented by replacing the ordinary convolutions in U-Net with depthwise separable convolutions.

[0065] In some embodiments, the region recognition model includes a first recognition model and a second recognition model, wherein the first recognition model is used to identify the region where a first object is located in the field of view image, and the second recognition model is used to identify the region where a ground object is located. Illustratively, the field of view image is input to the first recognition model, and the first recognition model identifies the region where the first object is located in the field of view image based on the texture information represented by the field of view image; the field of view image is input to the second recognition model, and the second recognition model identifies the region where a ground object is located in the field of view image based on the texture information represented by the field of view image.

[0066] The second method combines texture information represented by the field-of-view image with distance information represented by the depth map to identify the regions where the first object and the ground object are located respectively.

[0067] Schematic, based on the distance information represented by the depth map and the texture information represented by the field-of-view image, a first object region corresponding to a first object in the field-of-view image and a first ground region corresponding to a ground object are identified, wherein the first object region is the image region where the first object is located in the field-of-view image and the first ground region is the image region where the ground object is located in the field-of-view image; based on the depth map and the field-of-view image, first distance information corresponding to the first object region and second distance information corresponding to the first ground region are determined.

[0068] That is, by using the texture information represented by the field-of-view image and the distance information represented by the depth map, the first object region where the first object is located and the first ground region where the ground object is located are identified in the field-of-view image.

[0069] In some embodiments, a pre-trained region recognition model is used to identify and segment the field-of-view image, thereby determining the first object region where the first object is located, and the first ground region where the ground object is located. Illustratively, the depth map and the field-of-view image are input into the region recognition model, which identifies the first object and the ground object, determining the first object region and the first ground region in the field-of-view image. The region recognition model is used to identify the first object and the ground object from scene objects based on the spatial location information expressed by the depth map and the texture information expressed by the field-of-view image.

[0070] By combining spatial location information represented by depth maps and texture information represented by view images with a region recognition model, the first object and ground objects are identified from the scene. The first object and ground objects are segmented from the view images to obtain the first object region and the first ground region. The recognition efficiency of the first object region and the first ground region is improved by using a machine learning model, while the recognition accuracy of the first object region and the first ground region is improved by combining spatial location information and texture information, thereby improving the detection accuracy of ground contact detection.

[0071] In some embodiments, since there may be differences in the field of view between the first image acquisition module and the second image acquisition module, it is necessary to combine the pixel mapping relationship between the field of view image and the depth map when combining the texture information represented by the field of view image and the distance information represented by the depth map. Illustratively, a pixel mapping relationship is established between the first pixel in the field of view image and the second pixel in the depth map; the depth map, the field of view image, and the pixel mapping relationship are input into a region recognition model, which identifies the first object and the ground object, determining the first object region and the first ground region in the field of view image. That is, by establishing the pixel mapping relationship between the first pixel in the field of view image and the second pixel in the depth map, the region recognition model can combine the pixel mapping relationship between different images when recognizing the first object region and the first ground region, ensuring the accuracy of combining spatial location information and texture information, thereby improving the recognition accuracy of the first object region and the first ground region.

[0072] In some embodiments, a pixel mapping relationship between the field-of-view image and the depth map is established by pixel synchronization. That is, the pixels in the depth map are synchronized with the first coordinate system corresponding to the pixel image as a reference to obtain the pixel-synchronized depth map. The field-of-view image and the pixel-synchronized depth map are then input into the region recognition model to identify the first object region and the first ground region.

[0073] In illustrative terms, the pixel synchronization process includes: determining a first mapping relationship between the first coordinate system corresponding to the field-view image and the world coordinate system; determining a second mapping relationship between the second coordinate system corresponding to the depth map and the world coordinate system; and aligning the second pixel in the depth map with the first pixel in the field-view image based on the first and second mapping relationships to obtain the pixel-synchronized depth map.

[0074] In some embodiments, the region recognition model includes a first recognition model for recognizing a first object and a second recognition model for recognizing ground objects. Illustratively, a field-of-view image and a depth map (pixel-synchronized depth map) are input to the first recognition model, which identifies the first object region in the field-of-view image based on texture information represented by the field-of-view image and distance information represented by the depth map. The field-of-view image and the depth map (pixel-synchronized depth map) are then input to the second recognition model, which identifies the first ground region in the field-of-view image based on texture information represented by the field-of-view image and distance information represented by the depth map.

[0075] By designing different recognition models to identify different objects, the recognition models can learn the features of different types of objects in a targeted manner during the training phase, thereby improving the recognition accuracy of the first object region and the first ground region.

[0076] Step 230: Determine the ground contact detection result of the first object at the first moment based on the distance difference between the first distance information and the second distance information.

[0077] In some embodiments, the ground contact detection result includes a ground contact state and a ground-free state, wherein the ground contact state indicates that the first object is in contact with the ground object, and the ground-free state indicates that the first object is not in contact with the ground object. In some embodiments, when the ground contact relationship includes three relationships: direct contact, quasi-contact, and no contact, the ground contact state includes a direct contact relationship, that is, the state when the center of gravity of the detected object is on the first object and the first object is in contact with the ground object is the ground contact state; the ground-free state includes quasi-contact and no contact relationships, that is, the state when the center of gravity of the detected object is not on the first object but the first object is in contact with the ground object, and the state when the first object and the ground object are not in contact is the ground-free state.

[0078] In some embodiments, a preset difference threshold is used to determine the ground contact detection result of the first object at a first moment. Illustratively, the preset difference threshold is obtained; if the distance difference between the first distance information and the second distance information is less than the preset difference threshold, the first object is determined to be in a ground contact state; if the distance difference between the first distance information and the second distance information is greater than or equal to the preset difference threshold, the first object is determined to be in a ground-off state. The preset difference threshold provides a clear quantitative standard for determining ground contact and ground-off states, making the detection results more accurate and reliable.

[0079] Optionally, the aforementioned preset difference threshold can be system-preset or determined based on the terrain type of the ground object. In some embodiments, the value of the preset difference threshold is negatively correlated with the terrain material hardness of the terrain type.

[0080] In some embodiments, the identification device is further equipped with an inertial measurement unit (IMU). Illustratively, the IMU acquires the attitude information of the first and second image acquisition modules; the attitude information is used to perform viewpoint compensation on the depth map and the field-of-view image to obtain the compensated depth map and the compensated field-of-view image. That is, by acquiring the attitude information of the image acquisition modules through the IMU, viewpoint compensation can be performed on the depth map and the field-of-view image. This means that even if the image acquisition modules are in motion or tilted when capturing images, the attitude information can be used to adjust the images so that they are essentially captured from a fixed, standard viewpoint, thereby improving the accuracy of downstream information recognition and ultimately improving the accuracy of ground contact detection.

[0081] In some embodiments, the inertial measurement unit (IMU) can also output motion information of the first and second image acquisition modules. During the image acquisition process of the first and second image acquisition modules, the acquisition angles of the first and second image acquisition modules can be dynamically adjusted using the aforementioned motion information. Illustratively, the inertial measurement unit acquires motion information of the first and second image acquisition modules; based on the motion information, the first angle of the first image acquisition module and / or the second angle of the second image acquisition module are adjusted. That is, during image acquisition, the motion information acquired by the IMU can adjust the angle of the image acquisition modules in real time, meaning that even in dynamic environments, the image acquisition modules can always maintain the optimal acquisition angle, thereby improving the accuracy of downstream information recognition and ultimately improving the accuracy of ground contact detection.

[0082] In some embodiments, the first image acquisition module and the second image acquisition module are image acquisition modules configured within a dual-light module. That is, the first image acquisition module and the second image acquisition module are mounted on the same rigid hardware entity, and the attitude information of the first image acquisition module and the second image acquisition module is the same. Therefore, an inertial measurement unit is installed in the aforementioned dual-light module. Schematic, the attitude information of the dual-light module is acquired through the inertial measurement unit, and the depth map and field-of-view image are compensated using the attitude information to obtain compensated depth maps and compensated field-of-view images. In other words, the ground contact detection result of the first object is identified using the compensated depth map and compensated field-of-view image.

[0083] Optionally, the inertial measurement unit can also dynamically compensate the first and second image acquisition modules. For example, the inertial measurement unit acquires the motion and attitude information of the dual-light module, and uses this motion and attitude information to correct the viewing angle and position of the first and second image acquisition modules in real time, ensuring the accuracy of ground contact detection. Optionally, the aforementioned motion information includes acceleration information and / or angular velocity information.

[0084] Optionally, the inertial measurement unit can also dynamically adjust the distance information corresponding to the depth map. For example, when the inertial measurement unit detects that the first image acquisition module moves forward or backward and / or jitters, it dynamically adjusts the positional relationship between the first object area and the first ground area in the second coordinate system corresponding to the first image acquisition module through acceleration and angular velocity data, so as to ensure that the calculation of the first distance information and the second distance information is not affected by the camera movement.

[0085] In summary, when performing ground contact detection on the first object, the distance information represented by the depth map and the texture information represented by the field-of-view image, acquired simultaneously, are used to identify the distance information between the first object to be detected and the ground in the scene. The difference between these two distances is used to determine whether the first object is in contact with the ground, thus obtaining the ground contact detection result. That is, since ground contact detection can be achieved using only the depth map and field-of-view image provided by the image acquisition module, the geodesic detection recognition device can be configured at any location in the scene where the first object can be captured. This satisfies the need for non-contact ground contact detection scenarios. Furthermore, the reliable depth information provided by the depth map and the rich visual information provided by the field-of-view image ensure stable output of ground contact detection results, guaranteeing accuracy even in real-time detection scenarios.

[0086] In some optional embodiments, a first object region where a first object is located in the field of view image is identified by a first recognition model, and a first ground region where a ground object is located in the field of view image is identified by a second recognition model. Based on the identified regions, distance information of the identified regions is extracted from the depth map. Please refer to Figure 3, which shows a flowchart of a ground contact behavior recognition method provided in an exemplary embodiment of this application. The method includes steps 221 to 223, where steps 221 to 223 are subordinate steps to step 220.

[0087] Step 221: Input the depth map and field image into the first recognition model, identify the first object through the first recognition model, and determine the first object region in the field image.

[0088] In this embodiment of the application, the first object region in the field of view image is identified by the distance information represented by the depth map and the texture information represented by the field of view image. That is, the depth map and the field of view image are used as multimodal input data of the first recognition model, so as to identify the first object region in the field of view image by the first recognition model.

[0089] In some embodiments, since there may be a difference in the field of view between the first image acquisition module and the second image acquisition module, it is necessary to combine the pixel mapping relationship between the field of view image and the depth map when combining the texture information represented by the field of view image and the distance information represented by the depth map. Illustratively, a pixel mapping relationship is established between a first pixel in the field of view image and a second pixel in the depth map; the depth map, the field of view image, and the pixel mapping relationship are input into a first recognition model, and the first object is identified through the first recognition model to determine the first object region in the field of view image.

[0090] In some embodiments, a pixel mapping relationship between the field-view image and the depth map is established by pixel synchronization. That is, the pixels in the depth map are synchronized with the first coordinate system corresponding to the pixel image as a reference to obtain the pixel-synchronized depth map. The field-view image and the pixel-synchronized depth map are then input into the first recognition model to recognize the first object region.

[0091] Optionally, the first recognition model mentioned above can be implemented using neural network models such as CNN, FNN, ResNet, Transformer, and U-Net, without specific limitations.

[0092] In one example, the first recognition model is implemented as a U-Net model, which performs well in semantic segmentation tasks and can preserve detailed information in images. The U-Net model takes a view image and a depth map as input and outputs a binary mask of the first object region, where the pixel value of the first object region is 1, and the pixel values ​​of other regions are 0.

[0093] In some embodiments, the training process of the first recognition model includes: S1, data annotation: collecting sample depth map-sample field-of-view image pairs containing different scenes, terrains and lighting conditions, and accurately annotating the first object region to generate a binary mask for the first object region; S2, loss function: using a combination of cross-entropy loss function and Dice coefficient loss to improve the segmentation accuracy of the edge of the first object region and ensure that the boundary of the first object region is clear; S3, training details: expanding the training dataset through image enhancement techniques (rotation, scaling, translation, etc.) to improve the generalization ability and adaptability of the model and ensure segmentation accuracy in various terrains and environments.

[0094] In some embodiments, morphological operations (e.g., dilation and erosion) are applied to the binary mask of the first object region output by the first recognition model to remove noise in the segmentation result and ensure that the boundary of the first object region is smooth and coherent.

[0095] Step 222: Input the depth map and field-of-view image into the second recognition model, identify ground objects through the second recognition model, and determine the first ground region in the field-of-view image.

[0096] In this embodiment of the application, the first ground region where the ground object is located in the field of view image is identified by the distance information represented by the depth map and the texture information represented by the field of view image. That is, the depth map and the field of view image are used as multimodal input data of the second recognition model, so as to identify the first ground region in the field of view image by the second recognition model.

[0097] In some embodiments, since there may be differences in the field of view between the first image acquisition module and the second image acquisition module, it is necessary to combine the pixel mapping relationship between the field of view image and the depth map when combining the texture information represented by the field of view image and the distance information represented by the depth map. Illustratively, a pixel mapping relationship is established between a first pixel in the field of view image and a second pixel in the depth map; the depth map, the field of view image, and the pixel mapping relationship are input into a second recognition model, and the second recognition model identifies ground objects to determine a first ground region in the field of view image.

[0098] In some embodiments, a pixel mapping relationship between the field-of-view image and the depth map is established by pixel synchronization. That is, the pixels in the depth map are synchronized with the first coordinate system corresponding to the pixel image as a reference to obtain the pixel-synchronized depth map. The field-of-view image and the pixel-synchronized depth map are then input into the second recognition model to identify the first ground region.

[0099] Optionally, the second recognition model described above can be implemented using neural network models such as CNN, FNN, ResNet, Transformer, and U-Net, without any specific limitations.

[0100] In one example, the second recognition model is implemented as the U-Net model, which performs well in semantic segmentation tasks and can preserve detailed information in the image. The U-Net model takes the view image and depth map as input and outputs a binary mask of the first ground region, where the pixel value of the first ground region is 1 and the pixel value of other regions is 0.

[0101] In some embodiments, the training process of the second recognition model includes: S1, data annotation: collecting sample depth map-sample field-of-view image pairs containing different scenes, terrains and lighting conditions, and accurately annotating the first ground region to generate a binary mask for the first ground region; S2, loss function: using a combination of cross-entropy loss function and Dice coefficient loss to improve the segmentation accuracy of the edge of the first ground region and ensure that the boundary of the first ground region is clear; S3, training details: expanding the training dataset through image enhancement techniques (rotation, scaling, translation, etc.) to improve the generalization ability and adaptability of the model and ensure segmentation accuracy in various terrains and environments.

[0102] In some embodiments, morphological operations (e.g., dilation and erosion) are applied to the binary mask of the first ground region output by the second recognition model to remove noise in the segmentation result and ensure that the boundary of the first ground region is smooth and coherent.

[0103] Step 223: Determine the first distance information corresponding to the first object region and the second distance information corresponding to the first ground region based on the depth map and the field of view image.

[0104] In some embodiments, since there may be differences in the field of view between the first image acquisition module and the second image acquisition module, when determining the distance information corresponding to the segmented region in the field of view image using the depth map, it is necessary to combine the pixel mapping relationship between the field of view image and the depth map. Illustratively, a pixel mapping relationship is established between a first pixel in the field of view image and a second pixel in the depth map; based on the pixel mapping relationship, a second object region corresponding to the first object region is determined in the depth map; first distance information is obtained based on the depth value indicated by the second pixel in the second object region of the depth map; a second ground region corresponding to the first ground region is determined in the depth map based on the pixel mapping relationship; and second distance information is obtained based on the depth value indicated by the second pixel in the second ground region of the depth map. That is, by establishing a pixel mapping relationship between the first pixel in the field of view image and the second pixel in the depth map, and combining the pixel mapping relationship to determine the second object region associated with the first object and the second ground region associated with the ground object in the depth map, the recognition accuracy of the second object region and the second ground region is improved.

[0105] In some embodiments, the average depth value is calculated for the depth values ​​indicated by the second pixel values ​​in the second object region of the depth map, and the resulting average depth value is used as the first distance information; the average depth value is calculated for the depth values ​​indicated by the second pixel values ​​in the second ground region of the depth map, and the resulting average depth value is used as the second distance information.

[0106] In some embodiments, the plane containing the ground object is detected on the depth map to further ensure the depth accuracy of ground object recognition. A Random Sample Consensus (RANSAC) algorithm is used to perform plane fitting on the regional depth data within the second ground region to determine the average depth of the ground object as the second distance information. That is, outliers in the regional depth data are filtered out, making the ground object a more stable and reliable reference plane. Illustratively, regional depth data corresponding to the second ground region in the depth map is obtained, where the regional depth data includes the depth values ​​corresponding to each second pixel within the second ground region. The RANSAC algorithm is used to perform plane fitting on the regional depth data to determine the average depth value corresponding to the second ground region, and this average depth value is used as the second distance information.

[0107] In some optional embodiments, before performing plane fitting using the random sampling consensus algorithm, the second pixel in the depth map is horizontally corrected using the shooting angle information of the first image acquisition module at the first moment. This ensures that the corrected depth map indicates the true horizontal plane where the ground is located, thereby improving the recognition accuracy of the second distance information. Illustratively, the shooting angle information of the first image acquisition module at the first moment is acquired; the second pixel in the depth map is horizontally corrected based on the shooting angle information to obtain the corrected depth map; the regional depth data corresponding to the second ground region in the corrected depth map is acquired; the regional depth data is plane-fitted using the random sampling consensus algorithm to determine the average depth value corresponding to the second ground region, and the average depth value is used as the second distance information.

[0108] Optionally, the aforementioned shooting angle information can be data detected by an inertial measurement unit. For example, the shooting angle of the depth camera when acquiring the depth map is obtained by the inertial measurement unit, and the horizontal plane is corrected for each pixel in the depth map according to the shooting angle so that the true horizontal plane can be obtained when detecting the ground.

[0109] In summary, when performing ground contact detection on the first object, the distance information represented by the depth map and the texture information represented by the field-of-view image, acquired simultaneously, are used to identify the distance information between the first object to be detected and the ground in the scene. The difference between these two distances is used to determine whether the first object is in contact with the ground, thus obtaining the ground contact detection result. That is, since ground contact detection can be achieved using only the depth map and field-of-view image provided by the image acquisition module, the geodesic detection recognition device can be configured at any location in the scene where the first object can be captured. This satisfies the need for non-contact ground contact detection scenarios. Furthermore, the reliable depth information provided by the depth map and the rich visual information provided by the field-of-view image ensure stable output of ground contact detection results, guaranteeing accuracy even in real-time detection scenarios.

[0110] In some optional embodiments, when determining the ground contact detection result based on the distance difference between the first distance information and the second distance information, the accuracy of the ground contact detection result is improved by combining the terrain type recognition result of the ground object. Please refer to Figure 4, which shows a flowchart of a ground contact behavior recognition method provided in an exemplary embodiment of this application. The method includes steps 231 to 234, wherein steps 231 to 234 are subordinate steps of step 230.

[0111] Step 231: Detect the terrain type of the ground object to obtain the terrain information of the ground object.

[0112] In this embodiment of the application, the detected terrain information is used to indicate the terrain type of the ground object contacted by the first object. Optionally, the terrain type includes at least one of the following: grassland type, concrete type, hard floor type, carpet type, sand type, etc.

[0113] In some embodiments, after identifying a first ground region in the field-of-view image using a second recognition model, the terrain type is confirmed using a pre-trained terrain recognition model. Illustratively, a first ground region in the field-of-view image is determined; this first ground region is the image region where the ground object is located. The field-of-view image with annotation information of the first ground region is input into the terrain recognition model, which identifies the terrain type of the ground object to obtain terrain information. The terrain recognition model is used to classify the terrain type of the ground object based on the texture of the first ground region. That is, using a machine learning model to identify the terrain type of the first ground region in the field-of-view image improves the efficiency and accuracy of terrain type recognition.

[0114] Optionally, the terrain recognition model described above can be implemented using neural network models such as CNN, FNN, ResNet, Transformer, and U-Net, without any specific limitations.

[0115] In other embodiments, the identification of terrain information and the segmentation of the first ground region are achieved through the same model. Schematic, the depth map and the field-of-view image are used as multimodal input data of the second identification model, thereby identifying the first ground region in the field-of-view image and the terrain information corresponding to the first ground region through the second identification model.

[0116] Step 232: Obtain a preset difference threshold that matches the terrain information.

[0117] In this embodiment of the application, in order to adapt to the ground contact detection requirements of different terrains, this solution dynamically adjusts the ground contact judgment threshold according to the terrain recognition results to ensure accurate ground contact judgment on both soft ground (grass, sand) and hard ground (floor, concrete). To illustrate, the value of the preset difference threshold is negatively correlated with the hardness of the terrain material of the terrain type.

[0118] In some embodiments, the dynamic threshold can be preset according to the characteristics of different terrains, or obtained through experimental optimization. Based on the terrain recognition results, the system can select an appropriate preset difference threshold in real time. A mapping table between terrain types and preset difference thresholds is obtained, and the corresponding preset difference threshold is retrieved from the mapping table according to the terrain type indicated by the recognized terrain information.

[0119] In one example, on a hard surface, since the first object (sole) does not significantly embed itself in the ground, a smaller fixed threshold (e.g., 0.05 meters) is used. In another example, on a soft surface such as grass or sand, the first object (sole) may sink slightly, resulting in a slightly larger depth difference between the first object area and the first ground area. To avoid misjudgment, a larger dynamic threshold (e.g., 0.08 meters or higher) is used in this case.

[0120] Step 233: If the distance difference between the first distance information and the second distance information is less than a preset difference threshold, determine that the first object is in a ground-touching state.

[0121] Step 234: If the distance difference between the first distance information and the second distance information is greater than or equal to a preset difference threshold, determine that the first object is in an off-ground state.

[0122] In some embodiments, the ground contact detection result includes a ground contact state and a ground lift state, wherein the ground contact state indicates that the first object is in contact with a ground object, and the ground lift state indicates that the first object is not in contact with a ground object.

[0123] In some optional embodiments, when performing ground contact detection, in addition to determining ground contact based on the depth difference between the first object and the ground object, ground contact can also be determined based on the relative velocity of the first object. Illustratively, the object velocity information of the first object at a first moment is obtained, as well as a preset velocity threshold. If the distance difference between the first distance information and the second distance information is less than the preset difference threshold, and the object velocity information is less than the preset velocity threshold, the first object is determined to be in the ground contact state. If the distance difference between the first distance information and the second distance information is greater than or equal to the preset difference threshold, and / or the object velocity information is greater than or equal to the preset velocity threshold, the first object is determined to be in the off-ground state. That is, when both the distance difference and the object velocity information are less than the preset velocity threshold are satisfied simultaneously, the obtained ground contact detection result is that the first object is in the ground contact state. Introducing the relative velocity of the first object during ground contact determination improves the accuracy of ground contact determination.

[0124] In some embodiments, the relative velocity of the first object is calculated using depth maps. Illustratively, multiple consecutive depth maps continuously acquired by a first image acquisition module within a first time period are obtained, wherein the first time period includes a first moment; at least one key point corresponding to the first object is determined in each of the multiple depth maps; the inter-frame pixel displacement of the key point between the multiple depth maps is determined; the inter-frame pixel displacement is converted into spatial displacement; and based on the spatial displacement and the inter-frame time difference between the multiple depth maps, the object velocity information of the first object at the first moment is determined. The first time period can be a time period of preset duration ending at the first moment; or, the first time period can be a time period of preset duration centered at the first moment.

[0125] In other words, in the process of calculating the relative velocity of the first object by collecting depth maps, the depth map can represent the distance information between the first object and the ground. The relative velocity of the first object can be determined by multiple depth maps. Since the depth map needs to be collected when identifying the first distance information and the second distance information, using the depth map to determine the relative velocity of the first object can reduce the detection cost of the relative velocity of the first object. At the same time, the relative velocity of the first object is introduced when determining the ground contact to improve the accuracy of ground contact detection.

[0126] Specifically, the process of calculating the relative velocity of the first object includes the following steps:

[0127] S1, Pixel Position Tracking

[0128] In a depth map spanning multiple frames, a keypoint (or a set of keypoints) of a first object is selected. For example, if the first object is a foot, the keypoint could be the center point of the sole or a point near the heel and toes. An optical flow algorithm (e.g., Lucas-Kanade optical flow) is used to track the position of the determined keypoint in the depth map across multiple frames, obtaining the pixel displacements between them.

[0129] S2, pixel displacement converted to actual distance

[0130] The depth information captured by the depth camera is converted into actual distance. Assuming each pixel value in the depth map can be represented as depth Z, the actual 3D coordinates of the feet in each frame are calculated as shown in Formula 1. The keypoint positions in each frame are then transformed to obtain the 3D coordinates of the first object.

[0131] Where (x,y) are pixel coordinates, (c x ,c y (f) is the principal point of the depth camera. x ,f y () is the focal length, obtained through camera calibration.

[0132] S3, calculate inter-frame relative velocity

[0133] Velocity is calculated using the 3D coordinate changes of keypoints in the depth map across multiple frames. Assuming the inter-frame time is Δt, the relative velocity... As shown in Formula 2.

[0134] in, and These represent the three-dimensional coordinates of the first object in two adjacent frames.

[0135] S4, eliminates camera motion effects

[0136] IMU data is used to compensate for errors caused by depth camera motion. The IMU can provide the acceleration and velocity of the depth camera in each frame. Through pose estimation of the IMU, the rotation and translation changes of the depth camera itself can be calculated. By subtracting the influence of the depth camera's motion from the relative position changes in the depth map, the calculated relative velocity ensures that only the actual movement of the feet relative to the ground is reflected.

[0137] In a schematic way, the first moment is defined as the moment extracted at a preset extraction frequency within the second time period, and the ground contact detection result of the first object is continuously detected within the second time period. In some embodiments, when the first object is detected to change from an off-ground state to a ground contact state, the start timestamp of the ground contact is recorded, and when the first object is detected to change from a ground contact state to an off-ground state, the end timestamp of the ground contact is recorded. The duration of the ground contact is determined based on the start timestamp and the end timestamp.

[0138] In some optional embodiments, the ground contact detection results can be used for gait analysis. Schematic, the ground contact detection results for the first object at each detection moment within a second time period are obtained. Based on the ground contact detection results within the second time period, gait analysis is performed on the first object to obtain gait analysis results, which indicate whether there are any abnormalities in the first object's gait during the second time period. That is, the ground contact state of the first object at each detection moment can be accurately recorded through the ground contact detection results. This accurate ground contact detection provides a high-quality data foundation for gait analysis, thereby improving the accuracy and reliability of the gait analysis.

[0139] Optionally, the gait cycle of the first object can be identified through the ground contact detection results. For example, the system tracks the alternation of ground contact state and ground lift state in multiple consecutive frames to identify the gait cycle and calculate the gait cycle step length and step frequency.

[0140] Optionally, the stride of the first object can be identified through the ground contact detection results. For example, the stride of the first object can be obtained by detecting the change in distance of the first object's position between two gait cycles (taking a gait cycle as an example, which includes a continuous ground contact state and a continuous ground lift state).

[0141] Schematic illustration: In gait analysis, gait anomaly detection can be achieved by at least one of the following:

[0142] 1. Determine whether the ground contact time is within the specified duration range (including the shortest duration threshold and the longest duration threshold). For example, when the ground contact time of the first object in a single gait cycle is detected to be lower than the shortest duration threshold or higher than the longest duration threshold, it is determined that there is a gait abnormality.

[0143] 2. Determine whether the stride variation between multiple consecutive gait cycles is less than the first variation threshold. For example, when the stride difference between two adjacent gait cycles of the first object is detected to be greater than or equal to the first variation threshold, it is determined that there is a gait abnormality.

[0144] 3. Determine whether the change in gait frequency within multiple time intervals is less than the second change threshold. For example, when the gait frequency difference of the first object in adjacent time intervals is detected to be greater than or equal to the second change threshold, it is determined that there is a gait abnormality.

[0145] In summary, when performing ground contact detection on the first object, the distance information represented by the depth map and the texture information represented by the field-of-view image, acquired simultaneously, are used to identify the distance information between the first object to be detected and the ground in the scene. The difference between these two distances is used to determine whether the first object is in contact with the ground, thus obtaining the ground contact detection result. That is, since ground contact detection can be achieved using only the depth map and field-of-view image provided by the image acquisition module, the geodesic detection recognition device can be configured at any location in the scene where the first object can be captured. This satisfies the need for non-contact ground contact detection scenarios. Furthermore, the reliable depth information provided by the depth map and the rich visual information provided by the field-of-view image ensure stable output of ground contact detection results, guaranteeing accuracy even in real-time detection scenarios.

[0146] In one example, please refer to FIG5, which shows a schematic diagram of the structure of a ground contact detection functional module 500 provided in an exemplary embodiment of this application. The functional module 500 includes a depth camera 510, an RGB camera 520, an IMU 530, and a data processing unit 540.

[0147] During ground contact detection, the depth camera 510 sends the acquired depth map to the data processing unit 540, the RGB camera 520 sends the acquired RGB image to the data processing unit 540, and the IMU 530 sends the acquired motion and attitude information to the data processing unit 540.

[0148] The data processing unit 540 adjusts the depth map and RGB image based on motion and posture information, inputs the adjusted depth map and RGB image into the first recognition model, identifies the foot region in the RGB image, and obtains the corresponding binary mask of the foot region. The adjusted depth map and RGB image are then input into the second recognition model, which identifies the ground region in the RGB image and obtains the corresponding binary mask of the ground region. The mean depth value of the binary mask corresponding to the foot region in the corresponding region of the depth map is determined as the first distance information corresponding to the foot region, and the mean depth value of the binary mask corresponding to the ground region in the corresponding region of the depth map is determined as the second distance information corresponding to the ground region. The difference between the first distance information and the second distance information and a preset difference threshold are used to determine whether the foot is in contact with the ground.

[0149] The data processing unit 540 also adjusts the viewing angle of the depth camera 510 and the RGB camera 520 based on the motion and pose information.

[0150] It is worth noting that the ground contact behavior recognition method provided in this application embodiment has at least the following beneficial effects:

[0151] 1. Non-contact ground contact detection improves convenience and applicability: By combining a depth camera, RGB camera, and IMU, non-contact foot ground contact detection is achieved, avoiding the limitations of traditional physical sensors that rely on direct contact with the human body or shoe sole. This method makes device installation more flexible and easier to use, especially suitable for applications requiring long-distance and non-contact detection in gait analysis, rehabilitation training, and other scenarios.

[0152] 2. Multimodal fusion improves detection accuracy: By fusing color and texture information from an RGB camera, precise depth data from a depth camera, and dynamic compensation from an IMU, the segmentation of the shoe sole from the ground and depth detection become more accurate. Especially in complex environments, the high-resolution images provided by the RGB camera can clearly identify the boundary between the shoe sole and the ground, while the depth data provided by the camera ensures the accuracy of spatial positioning, solving the problem of insufficient accuracy of single-vision detection in poor lighting or complex environments.

[0153] 3. Dynamic Threshold Ground Contact Judgment Enhances Adaptability: By introducing terrain recognition, the system can dynamically adjust the ground contact judgment threshold according to different ground types (hard floors, grass, sand, etc.), ensuring accurate identification of the sole's ground contact status on both soft and hard surfaces. This approach effectively solves the problem of misjudgment easily occurring on different terrains using traditional fixed threshold methods, greatly improving the system's environmental adaptability.

[0154] 4. IMU Dynamic Compensation Enhances System Robustness: The acceleration, angular velocity, and attitude information provided by the IMU are used to correct camera motion deviations in real time, ensuring that the segmentation of the shoe sole and the ground, as well as depth detection, remain accurate even when the camera is moving. This dynamic compensation mechanism greatly improves the system's stability in dynamic environments, avoiding detection errors caused by camera shake or tilt, and enabling the system to operate stably in various mobile scenarios.

[0155] 5. Gait cycle analysis supports rehabilitation training and behavior monitoring: The system identifies and analyzes gait cycles through continuous multi-frame ground contact state analysis. Combining ground contact duration with data such as stride length and cadence, the system provides accurate gait assessment results. This function is particularly suitable for rehabilitation training and behavior monitoring scenarios, helping to monitor users' walking stability and gait abnormalities, and preventing fall risks.

[0156] 6. Adaptability to Complex Terrain and Diverse Lighting Conditions: The multimodal fusion design of this solution enables the system to maintain stable ground contact detection capabilities under various complex terrains and lighting conditions. The RGB camera provides rich visual details under lit conditions, while the camera can still generate accurate depth data in low-light or no-light environments. The IMU's attitude correction function further ensures that the detection of the shoe sole and the ground is unaffected during camera movement. The system can stably handle the detection needs of different environments, including indoors, outdoors, and alternating light and dark conditions.

[0157] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0158] Please refer to Figure 6, which shows a structural block diagram of a ground contact behavior recognition device provided in an exemplary embodiment of this application. The device includes the following modules:

[0159] The acquisition module 610 is used to acquire a depth map acquired by the first image acquisition module at a first moment, and a field-of-view image acquired by the second image acquisition module at the first moment. The depth map includes distance information between the scene object and the first image acquisition module, and the field-of-view image includes texture information of the scene.

[0160] The recognition module 620 is used to recognize first distance information corresponding to a first object in the scene object and second distance information corresponding to a ground object in the scene object based on the distance information represented by the depth map and the texture information represented by the field of view image; wherein, there is a ground contact relationship between the first object and the ground object, the first distance information is used to express the first distance between the recognized first object and the first image acquisition module, and the second distance information is used to express the second distance between the recognized ground object and the first image acquisition module;

[0161] The judgment module 630 is used to determine the ground contact detection result of the first object at the first moment based on the distance difference between the first distance information and the second distance information.

[0162] In some optional embodiments, the recognition module 620 is further configured to identify a first object region corresponding to the first object in the field of view image and a first ground region corresponding to the ground object, based on the distance information represented by the depth map and the texture information represented by the field of view image; or, based on the texture information represented by the field of view image, identify the first object region corresponding to the first object in the field of view image and the first ground region corresponding to the ground object; wherein, the first object region is the image region where the first object is located in the field of view image, and the first ground region is the image region where the ground object is located in the field of view image;

[0163] The recognition module 620 is further configured to determine, based on the depth map and the field-of-view image, the first distance information corresponding to the first object region and the second distance information corresponding to the first ground region.

[0164] In some optional embodiments, the recognition module 620 is further configured to input the depth map and the field-of-view image into a region recognition model, identify the first object and the ground object through the region recognition model, and determine the first object region and the first ground region in the field-of-view image. The region recognition model is a pre-trained machine learning model, which is used to analyze the spatial location information expressed by the depth map and the texture information expressed by the field-of-view image to identify the first object and the ground object from the scene objects.

[0165] In some optional embodiments, the recognition module 620 is further configured to establish a pixel mapping relationship between a first pixel in the field of view image and a second pixel in the depth map;

[0166] The recognition module 620 is further configured to input the depth map, the field-of-view image, and the pixel mapping relationship into the region recognition model, and identify the first object and the ground object through the region recognition model, thereby determining the first object region and the first ground region in the field-of-view image.

[0167] In some optional embodiments, the region recognition model includes a first recognition model and a second recognition model, wherein the first recognition model is used to identify the region where the first object is located in the field of view image, and the second recognition model is used to identify the region where the ground object is located;

[0168] The recognition module 620 is further configured to input the depth map and the field of view image into the first recognition model, recognize the first object through the first recognition model, and determine the first object region in the field of view image;

[0169] The recognition module 620 is further configured to input the depth map and the field-of-view image into the second recognition model, and identify the ground object through the second recognition model to determine the first ground region in the field-of-view image.

[0170] In some optional embodiments, the recognition module 620 is further configured to establish a pixel mapping relationship between a first pixel in the field of view image and a second pixel in the depth map;

[0171] The recognition module 620 is further configured to determine a second object region corresponding to the first object region in the depth map based on the pixel mapping relationship;

[0172] The recognition module 620 is further configured to obtain the first distance information based on the depth value indicated by the second pixel in the second object region of the depth map;

[0173] The identification module 620 is further configured to determine a second ground region corresponding to the first ground region in the depth map based on the pixel mapping relationship;

[0174] The identification module 620 is further configured to obtain the second distance information based on the depth value indicated by the second pixel in the second ground region of the depth map.

[0175] In some optional embodiments, the acquisition module 610 is further configured to acquire the region depth data corresponding to the second ground region in the depth map, the region depth data including the depth value corresponding to each second pixel in the second ground region;

[0176] The identification module 620 is further configured to perform planar fitting on the regional depth data using a random sampling consistency algorithm to determine the average depth value corresponding to the second ground region, and use the average depth value as the second distance information.

[0177] In some optional embodiments, the acquisition module 610 is further configured to acquire the shooting angle information of the first image acquisition module at the first moment;

[0178] The recognition module 620 is further configured to perform pixel horizontal plane correction on the second pixel in the depth map based on the shooting angle information to obtain a corrected depth map;

[0179] The acquisition module 610 is further configured to acquire the regional depth data corresponding to the second ground region in the corrected depth map.

[0180] In some optional embodiments, the acquisition module 610 is further configured to acquire attitude information of the first image acquisition module and the second image acquisition module through an inertial measurement unit;

[0181] The acquisition module 610 is further configured to perform viewpoint compensation on the depth map and the field of view image using the attitude information to obtain a compensated depth map and a compensated field of view image.

[0182] In some optional embodiments, the acquisition module 610 is further configured to acquire motion information of the first image acquisition module and the second image acquisition module through the inertial measurement unit;

[0183] The device further includes:

[0184] An adjustment module (not shown in the figure) is used to adjust the first viewpoint of the first image acquisition module and / or the second viewpoint of the second image acquisition module based on the motion information.

[0185] In some optional embodiments, the ground contact detection result includes ground contact state and ground lift state;

[0186] The acquisition module 610 is also used to acquire a preset difference threshold.

[0187] The judgment module 630 is further configured to determine that the first object is in the ground-touching state when the distance difference between the first distance information and the second distance information is less than the preset difference threshold.

[0188] The judgment module 630 is further configured to determine that the first object is in the off-ground state when the distance difference between the first distance information and the second distance information is greater than or equal to the preset difference threshold.

[0189] In some optional embodiments, the identification module 620 is further configured to detect the terrain type of the ground object and obtain terrain information of the ground object, the terrain information being used to indicate the terrain type of the ground object contacted by the first object;

[0190] The acquisition module 610 is further configured to acquire the preset difference threshold that matches the terrain information, wherein the value of the preset difference threshold is negatively correlated with the hardness of the terrain material of the terrain type.

[0191] In some optional embodiments, the recognition module 620 is further configured to determine the first ground region in the field of view image, wherein the first ground region is the image region in the field of view image where the ground object is located;

[0192] The recognition module 620 is further configured to input the view image with the annotation information of the first ground area into the terrain recognition model, identify the terrain type of the ground object through the terrain recognition model, and obtain the terrain information. The terrain recognition model is used to classify the terrain type of the ground object according to the texture of the first ground area.

[0193] In some optional embodiments, the acquisition module 610 is further configured to acquire the object speed information of the first object at the first moment, and to acquire a preset speed threshold.

[0194] The identification module 620 is further configured to determine that the first object is in the ground-contact state when the distance difference between the first distance information and the second distance information is less than the preset difference threshold and the object speed information is less than the preset speed threshold.

[0195] In some optional embodiments, the acquisition module 610 is further configured to acquire multiple consecutive depth maps continuously acquired by the first image acquisition module within a first time period, the first time period including the first moment.

[0196] At least one key point corresponding to the first object is determined in each of the multiple depth maps;

[0197] The identification module 620 is also used to determine the inter-frame pixel displacement of the key point between the multiple depth maps;

[0198] The recognition module 620 is also used to convert the inter-frame pixel displacement into spatial displacement;

[0199] The identification module 620 is further configured to determine the object velocity information of the first object at the first moment based on the spatial displacement and the inter-frame time difference between the plurality of depth maps.

[0200] In some optional embodiments, the acquisition module 610 is further configured to acquire the ground contact detection results corresponding to each detection time of the first object in the second time period;

[0201] The identification module 620 is further configured to perform gait analysis on the first object based on the ground contact detection results during the second time period, and obtain gait analysis results, which are used to indicate whether there is any abnormality in the gait of the first object during the second time period.

[0202] It should be noted that the ground contact behavior recognition device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the ground contact behavior recognition device and the ground contact behavior recognition method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0203] Figure 7 shows a structural block diagram of a terminal device 700 provided in an exemplary embodiment of this application. The terminal device 700 may be a smartphone, tablet computer, Moving Picture Experts Group Audio Layer III (MP3) player, Moving Picture Experts Group Audio Layer IV (MP4) player, laptop computer, or desktop computer. The terminal device 700 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.

[0204] Typically, terminal device 700 includes a processor 701 and a memory 702.

[0205] Processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 701 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 701 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.

[0206] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one instruction, which is executed by the processor 701 to implement the grounding behavior recognition method provided in the method embodiments of this application.

[0207] Indicatively, the terminal device 700 also includes other components 703. Those skilled in the art will understand that the structure shown in FIG7 does not constitute a limitation on the terminal device 700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0208] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may be a computer-readable storage medium included in the memory described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into the terminal. The computer-readable storage medium stores at least one instruction, at least one program segment, a code set, or an instruction set. The at least one instruction, the at least one program segment, the code set, or the instruction set is loaded and executed by the processor to implement any of the ground-touching behavior recognition methods described in the above embodiments.

[0209] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0210] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

Claims

1. A method for recognizing ground contact behavior, the method being executed by a computer device, the method comprising: The depth map acquired by the first image acquisition module at a first moment and the field-view image acquired by the second image acquisition module at the first moment are obtained. The depth map includes distance information between the scene object and the first image acquisition module, and the field-view image includes texture information of the scene. Based on the distance information represented by the depth map and the texture information represented by the field of view image, a first distance information corresponding to a first object in the scene object and a second distance information corresponding to a ground object in the scene object are identified; wherein, there is a ground contact relationship between the first object and the ground object, the first distance information is used to express the first distance between the identified first object and the first image acquisition module, and the second distance information is used to express the second distance between the identified ground object and the first image acquisition module; Based on the distance difference between the first distance information and the second distance information, the ground contact detection result of the first object at the first moment is determined.

2. The method according to claim 1, wherein, The step of identifying first distance information corresponding to a first object in the scene object and second distance information corresponding to a ground object in the scene object based on the distance information represented by the depth map and the texture information represented by the view image includes: Based on the distance information represented by the depth map and the texture information represented by the field-of-view image, a first object region corresponding to the first object and a first ground region corresponding to the ground object are identified in the field-of-view image; or, based on the texture information represented by the field-of-view image, the first object region corresponding to the first object and the first ground region corresponding to the ground object are identified in the field-of-view image; wherein, the first object region is the image region where the first object is located in the field-of-view image, and the first ground region is the image region where the ground object is located in the field-of-view image; Based on the depth map and the field of view image, determine the first distance information corresponding to the first object region and the second distance information corresponding to the first ground region.

3. The method according to claim 1 or 2, wherein, The step of identifying a first object region corresponding to the first object in the field of view image and a first ground region corresponding to the ground object, based on the distance information represented by the depth map and the texture information represented by the field of view image, includes: The depth map and the field-of-view image are input into a region recognition model. The region recognition model identifies the first object and the ground object, and determines the first object region and the first ground region in the field-of-view image. The region recognition model is a pre-trained machine learning model. The region recognition model is used to analyze the spatial location information expressed by the depth map and the texture information expressed by the field-of-view image to identify the first object and the ground object from the scene objects.

4. The method according to any one of claims 1 to 3, wherein, The step of inputting the depth map and the field-of-view image into a pre-trained region recognition model, and using the region recognition model to identify the first object and the ground object within the field of view, and determining the first object region and the first ground region in the field-of-view image, includes: Establish a pixel mapping relationship between the first pixel in the field-of-view image and the second pixel in the depth map; The depth map, the field-of-view image, and the pixel mapping relationship are input into the region recognition model. The region recognition model identifies the first object and the ground object, and determines the first object region and the first ground region in the field-of-view image.

5. The method according to any one of claims 1 to 4, wherein, The region recognition model includes a first recognition model and a second recognition model. The first recognition model is used to recognize the region where the first object is located in the field of view image, and the second recognition model is used to recognize the region where the ground object is located. The step of inputting the depth map and the field-of-view image into a region recognition model, identifying the first object and the ground object through the region recognition model, and determining the first object region and the first ground region in the field-of-view image includes: The depth map and the field of view image are input into the first recognition model, and the first object is identified by the first recognition model to determine the first object region in the field of view image. The depth map and the field-of-view image are input into the second recognition model, and the ground object is identified by the second recognition model to determine the first ground region in the field-of-view image.

6. The method according to any one of claims 1 to 5, wherein, The step of determining the first distance information corresponding to the first object region and the second distance information corresponding to the first ground region based on the depth map and the field of view image includes: Establish a pixel mapping relationship between the first pixel in the field-of-view image and the second pixel in the depth map; Based on the pixel mapping relationship, a second object region corresponding to the first object region is determined in the depth map; The first distance information is obtained based on the depth value indicated by the second pixel in the second object region of the depth map; Based on the pixel mapping relationship, a second ground region corresponding to the first ground region is determined in the depth map; The second distance information is obtained based on the depth value indicated by the second pixel in the second ground region of the depth map.

7. The method according to any one of claims 1 to 6, wherein, The step of obtaining the second distance information based on the depth value indicated by the second pixel in the second ground region of the depth map includes: Obtain the region depth data corresponding to the second ground region in the depth map, wherein the region depth data includes the depth value corresponding to each second pixel in the second ground region; The average depth value corresponding to the second ground area is determined by performing planar fitting on the depth data of the region using a random sampling consistency algorithm, and the average depth value is used as the second distance information.

8. The method according to any one of claims 1 to 7, wherein, The step of obtaining the regional depth data corresponding to the second ground region in the depth map includes: Obtain the shooting angle information of the first image acquisition module at the first moment; Based on the shooting angle information, the second pixel in the depth map is corrected for the horizontal plane to obtain the corrected depth map; Obtain the depth data of the region corresponding to the second ground region in the corrected depth map.

9. The method according to any one of claims 1 to 8, wherein, The method further includes: The attitude information of the first image acquisition module and the second image acquisition module is obtained through an inertial measurement unit; The pose information is used to perform viewpoint compensation on the depth map and the field of view image to obtain the compensated depth map and the compensated field of view image.

10. The method according to any one of claims 1 to 9, wherein, The method further includes: Motion information of the first image acquisition module and the second image acquisition module is obtained through an inertial measurement unit; Adjust the first viewpoint of the first image acquisition module and / or the second viewpoint of the second image acquisition module based on the motion information.

11. The method according to any one of claims 1 to 10, wherein, The ground contact detection results include ground contact state and ground lift state; Determining the ground contact detection result of the first object at the first moment based on the distance difference between the first distance information and the second distance information includes: Get the preset difference threshold; If the distance difference between the first distance information and the second distance information is less than the preset difference threshold, it is determined that the first object is in the ground-touching state; If the distance difference between the first distance information and the second distance information is greater than or equal to the preset difference threshold, the first object is determined to be in the off-ground state.

12. The method according to any one of claims 1 to 11, wherein, The step of obtaining the preset difference threshold includes: The terrain type of the ground object is detected to obtain the terrain information of the ground object, and the terrain information is used to indicate the terrain type of the ground object that the first object comes into contact with. Obtain the preset difference threshold that matches the terrain information. The value of the preset difference threshold is negatively correlated with the hardness of the terrain material of the terrain type.

13. The method according to any one of claims 1 to 12, wherein, The step of detecting the terrain type of the ground object to obtain the terrain information of the ground object includes: Determine the first ground region in the field of view image, wherein the first ground region is the image region in the field of view image where the ground object is located; The view image with the annotation information of the first ground area is input into the terrain recognition model. The terrain recognition model identifies the terrain type of the ground object to obtain the terrain information. The terrain recognition model is used to classify the terrain type of the ground object according to the texture of the first ground area.

14. The method according to any one of claims 1 to 13, wherein, The step of determining that the first object is in the ground-touching state when the distance difference between the first distance information and the second distance information is less than the preset difference threshold includes: Obtain the object velocity information of the first object at the first moment, and obtain the preset velocity threshold; If the distance difference between the first distance information and the second distance information is less than the preset difference threshold, and the object speed information is less than the preset speed threshold, then the first object is determined to be in the ground-contact state.

15. The method according to any one of claims 1 to 14, wherein, The step of obtaining the object velocity information of the first object at the first moment includes: The first image acquisition module continuously acquires multiple depth maps during a first time period, which includes the first moment. At least one key point corresponding to the first object is determined in each of the multiple depth maps; Determine the inter-frame pixel displacement of the key points between the multiple depth maps; Convert the inter-frame pixel displacement into spatial displacement; Based on the spatial displacement and the inter-frame time difference between the multiple depth maps, the object velocity information of the first object at the first moment is determined.

16. The method according to any one of claims 1 to 15, wherein, The method further includes: Obtain the ground contact detection results of the first object at each detection time within the second time period; Based on the ground contact detection results during the second time period, gait analysis is performed on the first object to obtain gait analysis results, which are used to indicate whether there are any abnormalities in the gait of the first object during the second time period.

17. A ground-touching behavior recognition device, the device comprising: The acquisition module is used to acquire a depth map acquired by the first image acquisition module at a first moment, and a field-of-view image acquired by the second image acquisition module at the first moment. The depth map includes distance information between the scene object and the first image acquisition module, and the field-of-view image includes texture information of the scene. The recognition module is used to recognize, based on the distance information represented by the depth map and the texture information represented by the field-of-view image, a first distance information corresponding to a first object in the scene objects and a second distance information corresponding to a ground object in the scene objects; wherein, there is a ground contact relationship between the first object and the ground object, the first distance information is used to express the first distance between the recognized first object and the first image acquisition module, and the second distance information is used to express the second distance between the recognized ground object and the first image acquisition module; The judgment module is used to determine the ground contact detection result of the first object at the first moment based on the distance difference between the first distance information and the second distance information.

18. A computer device comprising a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the ground contact behavior recognition method as claimed in any one of claims 1 to 16.

19. A computer-readable storage medium storing at least one piece of program code, the program code being loaded and executed by a processor to implement the ground contact behavior recognition method as described in any one of claims 1 to 16.

20. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the ground contact behavior recognition method as described in any one of claims 1 to 16.

Citation Information

Patent Citations

  • Method and device for synchronously acquiring depth information and color information

    CN103796001A

  • Image shooting method and device, image processing method and device, electronic equipment and storage medium

    CN111726531A

  • Depth image completion method and device and computer readable storage medium

    CN112446909A