Carpet detection method, motion control method, and mobile machine using these methods
Detecting carpet and carpet curls through RGB-D cameras and deep learning models, the problems of low carpet detection efficiency and poor safety in the prior art are solved, and efficient navigation of mobile robots in carpet scenes are achieved.
Patent Information
- Application Number
- CN202280003746.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-31
- Filing Date
- 2022-03-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-03-02
AI Technical Summary
The existing carpet detection methods are difficult to effectively detect carpets and carpet curls, resulting in the safety risks of mobile robots when navigating in carpet scenes and are cost-effective and inefficient.
Using an RGB-D camera combined with a deep learning model, carpet curled areas in RGB-D images are detected, 2D bounding boxes are generated and depth image pixels are matched, and carpet curled points are identified to achieve carpet detection and obstacle avoidance.
It improves the navigation stability and safety of mobile robots in carpet scenes, reduces inspection costs, and achieves real-time and efficient carpet and carpet curl detection.
Smart Images

Figure CN115668293B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to carpet detection technology, and in particular to a carpet detection method, a motion control method, and a mobile machine using these methods. Background Art
[0002] Mobile machines (such as mobile robots, vehicles) are usually given the ability of autonomous navigation in order to work in a more autonomous way. Especially for mobile robots of types such as domestic robots and companion robots, since they will be used in scenarios with carpets in daily life, which will affect their movement and even lead to dangerous situations such as falling, it is necessary to detect carpets during the movement of the mobile robot.
[0003] Due to the characteristics of low height of carpets and the fact that they are usually laid on the floor, they are difficult to be detected by most sensors. Therefore, existing carpet detection methods usually adopt a combination of multiple sensors, resulting in low detection efficiency and high cost. In addition, for mobile robots that move using wheels or have a small height, carpet curling is likely to cause dangerous situations, and existing carpet detection methods are not particularly optimized for the detection of carpet curling. Therefore, existing carpet detection methods are difficult to be applied to the navigation of mobile robots in scenarios with carpets.
[0004] Therefore, a carpet detection method capable of detecting carpets and carpet curling is needed, and this method can be applied to navigate mobile robots in scenarios with carpets. Summary of the Invention
[0005] The present application provides a carpet detection method, a motion control method, and a mobile machine using these methods to detect carpets and carpet curling and solve the problems existing in the existing carpet detection methods in the foregoing prior art.
[0006] An embodiment of the present application provides a carpet detection method, including:
[0007] Obtaining one or more pairs of RGB-D images through an RGB-D camera, where each pair of RGB-D images includes an RGB image and a depth image;
[0008] Detecting one or more carpet regions in the RGB image of the one or more pairs of RGB-D images, and generating a first 2D bounding box based on a first data set including carpet classes and using a first deep learning model to mark each of the carpet regions;
[0009] Detecting one or more carpet curling regions in the RGB image of the one or more pairs of RGB-D images, and generating a second 2D bounding box based on a second data set including carpet curling classes and using a second deep learning model to mark each of the carpet curling regions;
[0010] A set of carpet points corresponding to each carpet region in the RGB image within the first 2D bounding box corresponding to each such carpet region is generated by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair; and
[0011] A set of carpet curl points corresponding to each carpet curl region in the RGB image within the second 2D bounding box corresponding to each such carpet curl region is generated by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair.
[0012] Embodiments of the present application also provide a method for controlling a mobile machine to move, where the mobile machine has an RGB-D camera, and the method includes:
[0013] On one or more processors of the mobile machine:
[0014] Obtain one or more RGB-D image pairs through the RGB-D camera, where each such RGB-D image pair includes an RGB image and a depth image;
[0015] Detect one or more carpet regions in the RGB image of the one or more RGB-D image pairs, and generate a first 2D bounding box to mark each such carpet region based on a first data set including carpet classes and using a first deep learning model;
[0016] Detect one or more carpet curl regions in the RGB image of the one or more RGB-D image pairs, and generate a second 2D bounding box to mark each such carpet curl region based on a second data set including carpet curl classes and using a second deep learning model;
[0017] A set of carpet points corresponding to each carpet region in the RGB image within the first 2D bounding box corresponding to each such carpet region is generated by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair;
[0018] A set of carpet curl points corresponding to each carpet curl region in the RGB image within the second 2D bounding box corresponding to each such carpet curl region is generated by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair;
[0019] Take the generated set of carpet points corresponding to each carpet region in the RGB image corresponding to each such RGB-D image pair as a carpet;
[0020] Obtain one or more carpet curls based on the generated set of carpet curl points corresponding to each carpet curl region in the RGB image corresponding to each RGB-D image pair, and regard each carpet curl as an obstacle, where each carpet curl corresponds to each carpet curl region; and
[0021] Control the mobile machine to move with reference to the carpet in each carpet region in the RGB image corresponding to each RGB-D image pair to avoid the obstacle.
[0022] Embodiments of the present application further provide a mobile machine, including:
[0023] An RGB-D camera;
[0024] One or more processors; and
[0025] One or more memories storing one or more computer programs, which are executed by the one or more processors, where the one or more computer programs include multiple instructions for:
[0026] Obtain one or more RGB-D image pairs through the RGB-D camera, where each RGB-D image pair includes an RGB image and a depth image;
[0027] Detect one or more carpet regions in the RGB image of the one or more RGB-D image pairs, and generate first 2D bounding boxes to mark each carpet region based on a first data set including carpet classes and using a first deep learning model;
[0028] Detect one or more carpet curl regions in the RGB image of the one or more RGB-D image pairs, and generate second 2D bounding boxes to mark each carpet curl region based on a second data set including carpet curl classes and using a second deep learning model;
[0029] Generate a set of carpet points corresponding to each carpet region in the RGB image corresponding to each RGB-D image pair by matching each pixel of the RGB image within the first 2D bounding box corresponding to each carpet region with each pixel of the depth image of each RGB-D image pair; and
[0030] Generate a set of carpet curl points corresponding to each carpet curl region in the RGB image corresponding to each RGB-D image pair by matching each pixel of the RGB image within the second 2D bounding box corresponding to each carpet curl region with each pixel of the depth image of each RGB-D image pair.
[0031] As can be seen from the embodiments of the present application above, the carpet detection method provided by the present application uses an RGB-D camera and detects carpets and carpet curls based on a deep learning model, thereby solving problems such as difficulty in navigating a mobile robot in a scenario with a carpet in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. In the following drawings, the same reference numerals represent corresponding parts throughout the drawings. It should be understood that the drawings in the following description are only examples of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative labor.
[0033] Figure 1A It is a schematic diagram of the navigation scenario of a mobile machine in some embodiments of the present application.
[0034] Figure 1B It is in Figure 1A A schematic diagram of detecting a carpet in the scenario.
[0035] Figure 2A It is a perspective view of a mobile machine in some embodiments of the present application.
[0036] Figure 2B It is to illustrate Figure 2A A schematic block diagram of the mobile machine.
[0037] Figure 3A It is Figure 2A A schematic block diagram of an example of a mobile machine performing carpet detection.
[0038] Figure 3B It is Figure 1A A schematic diagram of an RGB image of the scenario.
[0039] Figure 4A It is Figure 2A A schematic block diagram of an example of the movement control of the mobile machine.
[0040] Figure 4B It is in Figure 4A A schematic block diagram of an example of determining a carpet and a carpet curl in an example of the movement control of the mobile machine.
[0041] Figure 4C It is in Figure 4A A schematic block diagram of an example of the movement control in an example of the movement control of the mobile machine.
[0042] Figure 5 It is a flowchart of a navigation method in some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, features, and advantages of this application more obvious and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without any creative work are within the scope protected by this application.
[0044] It should be understood that when used in this application and the appended claims, the terms "comprising", "including", "having" and their variants mean the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0045] It should also be understood that the terms used in the description of this application are only for the purpose of describing specific embodiments and do not limit the scope of this application. As used in the description of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0046] It should be further understood that the term "and / or" used in the description of this application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0047] The terms "first", "second", and "third" in this application are only for the purpose of description and should not be construed as indicating or implying relative importance or the number of the indicated technical features. Thus, the features defined by "first", "second", and "third" may explicitly or implicitly include at least one of the technical features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise clearly defined.
[0048] The statement "in an embodiment" or "in some embodiments" or the like described in the description of this application means that specific features, structures, or characteristics related to the description of the embodiment may be included in one or more embodiments of this application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different places in the description are not meant that the described embodiments should be cited by all other embodiments, but by "one or more but not all other embodiments", unless otherwise specifically emphasized.
[0049] This application relates to the navigation of mobile machines. As used herein, the term "mobile machine" refers to a machine capable of moving around in its environment, such as a mobile robot or vehicle. The term "carpet" (e.g., rug, blanket, and mat) refers to a textile type of floor covering, typically consisting of an upper pile attached to its backing. The term "carpet curl" refers to the curling of the edge of a carpet. The term "trajectory planning" refers to finding a series of time-parametrized valid configurations that move a mobile machine from a source to a destination, where a "trajectory" represents a sequence of poses with timestamps (for reference, a "path" represents a series of poses or positions without timestamps). The term "pose" refers to a position (e.g., x and y coordinates on the x and y axes) and an orientation (e.g., yaw angle along the z axis). The term "navigation" refers to the process of monitoring and controlling a mobile robot to move from one place to another. The term "collision avoidance" refers to preventing or reducing the severity of a collision. The term "sensor" refers to a device, module, machine, or subsystem (e.g., a camera) intended to detect events or changes in its environment and send relevant information to other electronic devices (e.g., a processor).
[0050] Figure 1A is a schematic diagram of the navigation scenario of the mobile machine 100 in some embodiments of this application. The mobile machine 100 is navigated in its environment (e.g., a living room), while dangerous situations such as collisions and unsafe states (e.g., falling, extreme temperatures, radiation, and exposure) can be prevented. In this indoor navigation, the mobile machine 100 is navigated from a starting point (e.g., the position where the mobile machine 100 is initially located) to a destination (e.g., by a user U (not shown in the figure) or the navigation / operating system of the mobile machine 100), while the carpet C c can be traversed (i.e., the mobile machine 100 can move on the carpet C c during navigation) and obstacles (e.g., the carpet curl C c of the carpet C r , walls, furniture, people, pets, and debris) can be avoided to prevent the above-mentioned dangerous situations. It is necessary to plan a trajectory (e.g., trajectories T1 and T2) for the mobile machine 100 to move from the starting point to the destination, so as to move the mobile machine 100 according to the trajectory. Each trajectory includes a series of poses (e.g., poses S1 - S9 of trajectory T2). It should be noted that the starting point and the ending point only represent the positions of the mobile machine 100 in the scenario shown in the figure, rather than the true start and end of the trajectory (the true start and end of the trajectory should be a pose respectively, e.g., Figure 1A the initial pose S i1 , S i2 and the desired pose S d1 , S d2)。In some embodiments, to implement the navigation of the mobile machine 100, it is necessary to construct an environmental map, determine the current position of the mobile machine 100 in the environment, and then plan a trajectory based on the constructed map and the determined position of the mobile machine 100.
[0051] The first initial pose S i1 is the starting point of the trajectory T1, and the second initial pose S i2 is the starting point of the trajectory T2. The first desired pose S d1 is the last one in the pose sequence S of the trajectory T1, that is, the end of the trajectory T1; the second desired pose S d2 is the last one in the pose sequence S of the trajectory T2, that is, the end of the trajectory T2. The trajectory T1 is planned according to, for example, the shortest path to the user U in the constructed map. The trajectory T2 is planned corresponding to the carpet detection performed by the mobile machine 100 (see Figure 3A ). In addition, when planning, it is necessary to consider avoiding collisions with obstacles (such as walls and furniture) in the constructed map or obstacles detected in real time (such as people and pets) in order to navigate the mobile machine 100 more accurately and safely.
[0052] In some embodiments, the navigation of the mobile machine 100 can be initiated by providing a navigation request for the mobile machine 100 by the mobile machine 100 itself (such as a control interface on the mobile machine 100) or by an electronic device such as a remote control, a smartphone, a tablet computer, a laptop computer, a desktop computer, or other electronic devices. The mobile machine 100 and the control device 200 can communicate through a network, which can include, for example, the Internet, an intranet, an extranet, a local area network (LAN), a wide area network (WAN), a wired network, a wireless network (such as a Wi-Fi network, a Bluetooth network, and a mobile network), or other suitable networks, or any combination of two or more such networks.
[0053] Figure 1B is Figure 1A a schematic diagram of detecting a carpet in the Figure 1B scene. As c shown, the field of view V of the RGB-D camera C covers the carpet C c and the floor F. Through carpet detection, the carpet C c and the carpet curl C r of the carpet C can be detected. Figure 2A is a perspective view of the mobile machine 100 in some embodiments of the present application. In some embodiments, the mobile machine 100 can be a mobile robot (such as a mobile assistance robot), which can include a walking frame F, a grasping component H, wheels E, and include an RGB-D camera C, a LiDAR (light detection and ranging) S1, and an IR (infrared sensor) S iSensors including it. The RGB-D camera C is installed on the front side (upper part) of the mobile machine 100, facing the direction in which the mobile machine 100 moves forward, so that the field of view V can cover the place where the mobile machine 100 is going to move, in order to detect, for example, the carpet C when the mobile machine 100 moves forward. c and / or obstacles (such as the carpet C c of the carpet curl C r ). The height of the RGB-D camera C on the mobile machine 100 can be changed according to actual needs (for example, the greater the height, the larger the field of view V; the smaller the height, the smaller the field of view V), and the elevation angle of the RGB-D camera C relative to the floor F can also be changed according to actual needs (for example, the larger the pitch angle, the closer the field of view V; the smaller the pitch angle, the farther the field of view V). The LiDAR S1 is installed in the middle of the front side, and the IR S i is installed in the lower part of the front side (for example, two or more IR S i are installed at a certain interval on the front side) to detect obstacles (such as walls, people, the carpet C c of the curl C r ), and can also detect the carpet C c . The grasping part H is installed on the upper edge of the walking frame F for the user U to grasp; and the wheels E are installed on the bottom (such as the chassis) of the walking frame F to move the walking frame F, so that the user U can stand and move supported by the mobile machine 100 with the assistance of the mobile machine 100. The height of the walking frame F can be adjusted manually or automatically, for example, through a telescopic mechanism such as a telescopic rod in the walking frame F, so that the grasping part H reaches a height convenient for the user U to grasp. The grasping part H may include a pair of parallel handles for the user U to grasp with both hands, and a brake lever installed on the handles for the user U to brake the mobile machine 100 with both hands, and may also include related components such as Bowden cables. It should be noted that the mobile machine 100 is just an example of a mobile machine, which may have different, more / fewer components from those shown above or below (for example, having legs instead of the wheels E), or may have different component configurations or arrangements (for example, having a single grasping component in the form of a grasping rod). In other embodiments, the mobile machine 100 may be another mobile machine such as a vehicle.
[0054] Figure 2B is to illustrate Figure 2ASchematic block diagram of mobile machine 100. Mobile machine 100 may include a processing unit 110, a storage unit 120, and a control unit 130 that communicate via one or more communication buses or signal lines L. It should be noted that mobile machine 100 is just an example of a mobile machine. Mobile machine 100 may have more or fewer components (such as units, subunits, and modules) than shown above or below, two or more components may be combined, or it may have different component configurations or arrangements. The processing unit 110 executes various (groups of) instructions stored in the storage unit 120, which may be in the form of software programs, to perform various functions of mobile machine 100 and process relevant data, and it may include one or more processors (such as a central processing unit (CPU)). The storage unit 120 may include one or more memories (such as high-speed random access memory (RAM) and non-transitory memory), one or more memory controllers, and one or more non-transitory computer-readable storage media (such as a solid-state drive (SSD) or a hard disk). The control unit 130 may include various controllers (such as a camera controller, a display controller, and a physical button controller) and a peripheral interface for coupling the input / output peripherals of mobile machine 100 to the processing unit 110 and the storage unit 120, such as external ports (such as USB), wireless communication circuits (such as RF communication circuits), audio circuits (such as speaker circuits), sensors (such as an inertial measurement unit (IMU)). In some embodiments, the storage unit 120 may include a navigation module 121 for implementing navigation functions (such as map building and trajectory planning) related to the navigation (and trajectory planning) of mobile machine 100, which may be stored in one or more memories (and one or more non-transitory computer-readable storage media).
[0055] The navigation module 121 in the storage unit 120 of mobile machine 100 may be a software module (of the operating system of mobile machine 100) that has instructions I n (such as instructions for driving the motor 1321 of mobile machine 100 to move mobile machine 100) to implement the navigation, map builder 1211, and trajectory planner 1212 of mobile machine 100. The map builder 1211 may be a software module that has instructions I b for building a map for mobile machine 100, and the trajectory planner 1212 may be a software module that has instructions I p for planning a trajectory for mobile machine 100. The trajectory planner 1212 may include a global trajectory planner for planning a global trajectory (such as trajectory T1 and trajectory T2) for mobile machine 100, and for planning a local trajectory of mobile machine 100 (such as including Figure 1AA local trajectory planner for a part of the trajectory T2 of the poses S1 - S4 (in...). The global trajectory planner can be, for example, a trajectory planner based on the Dijkstra algorithm, which plans the global trajectory based on the map constructed by the map builder 1211 through methods such as simultaneous localization and mapping (SLAM). The local trajectory planner can be a trajectory planner based on the timed elastic band (TEB) algorithm, which plans the local trajectory based on the global trajectory P g and other data collected by the mobile machine 100. For example, images can be collected by the RGB - D camera C of the mobile machine 100, and the collected images can be analyzed to identify obstacles (such as Figure 1B the obstacle O in...) c ), so that the local trajectory can be planned with reference to the identified obstacles, and the mobile machine 100 can be moved according to the planned local trajectory to avoid obstacles.
[0056] The map builder 1211 and the trajectory planner 1212 can be sub - modules separated from the instructions I for implementing the navigation of the mobile machine 100 n , or other sub - modules of the navigation module 121, or a part of the instructions I n . The trajectory planner 1212 can also have data related to the trajectory planning of the mobile machine 100 (such as input / output data and temporary data), which can be stored in one or more memories and accessed by the processing unit 110. In some embodiments, each trajectory planner 1212 can be a module separated from the navigation module 121 in the storage unit 120.
[0057] In some embodiments, the instructions I n can include instructions for implementing collision avoidance of the mobile machine 100 (such as obstacle detection and trajectory replanning). In addition, the global trajectory planner can replan the global trajectory (i.e., plan a new global trajectory) in response to, for example, the original global trajectory being blocked (such as blocked by one or more unexpected obstacles) or being insufficient to avoid collisions (such as being unable to avoid the detected obstacles when adopted). In other embodiments, the navigation module 121 can be a navigation unit that communicates with the processing unit 110, the storage unit 120, and the control unit 130 through one or more communication buses or signal lines L, and can also include one or more memories (such as high - speed random access memory (RAM) and non - volatile memory) for storing the instructions I n , the map builder 1211, and the trajectory planner 1212; and one or more processors (such as MPU and MCU) for executing the stored instructions I n 、I b and Ip , to achieve the navigation of the mobile machine 100.
[0058] The mobile machine 100 may further include a communication subunit 131 and an actuation subunit 132. The communication subunit 131 and the actuation subunit 132 communicate with the control unit 130 through one or more identical communication buses or signal lines, and the one or more communication buses or signal lines may be the same as or at least partially different from the above-mentioned one or more communication buses or signal lines L. The communication subunit 131 is coupled to the communication interface of the mobile machine 10, such as a network interface 1311 for the mobile machine 100 to communicate with the control device 200 through a network, and an I / O interface 1312 (such as a physical button). The actuation subunit 132 is coupled to the components / devices for realizing the movement of the mobile machine 100 to drive the motors 1321 of the wheels E and / or joints of the mobile machine 100. The communication subunit 131 may include a controller for the above-mentioned communication interfaces of the mobile machine 100, and the actuation subunit 132 may include a controller for the above-mentioned components / devices for realizing the movement of the mobile machine 100. In other embodiments, the communication subunit 131 and / or the actuation subunit 132 may be just abstract components used to represent the logical relationship between the components of the mobile machine 100.
[0059] The mobile machine 100 may further include a sensor subunit 133, which may include a set of sensors and related controllers, such as an RGB-D camera C, a LiDAR S1, an IR S1, and an IMU 1331 (or an accelerometer and a gyroscope), for detecting its surrounding environment to achieve its navigation. The sensor subunit 133 communicates with the control unit 130 through one or more communication buses or signal lines, and the one or more communication buses or signal lines may be the same as or at least partially different from the above-mentioned one or more communication buses or signal lines L. In other embodiments, when the navigation module 121 is the above-mentioned navigation unit, the sensor subunit 133 may communicate with the navigation unit through one or more communication buses or signal lines, and the communication bus or signal line may be the same as or at least partially different from the above-mentioned one or more communication buses or signal lines L. In addition, the sensor subunit 133 may be just an abstract component used to represent the logical relationship between the components of the mobile machine 100.
[0060] In addition to detecting the carpet C c In addition, the RGB-D camera C may be useful for detecting small objects with rich features or obvious contours (such as a mobile phone and a pool of water on the carpet C c ), as well as suspended objects with a larger upper part than the lower part. For example, for a table with four legs, the RGB-D camera C can identify the surface of the table, while the LiDAR S1 can only detect the four legs and does not know that there is a surface on it. The IR S iIt may be beneficial for detecting small obstacles with a height exceeding 5 cm. In other embodiments, the number of the above sensors in the sensor S can be changed according to actual needs. The sensor S can include a part of the above types of sensors (such as the RGB-D camera C and the LiDAR S1), and can also include other types of sensors (such as sonar).
[0061] In some embodiments, the map builder 1211, the trajectory planner 1212, the sensor subunit 133, and the motor 1321 (as well as the wheels E and / or joints of the mobile machine 100 connected to the motor 1321) together form a (navigation) system to implement map building, (global and local) trajectory planning, and motor drive to achieve the navigation of the mobile machine 100. In addition, Figure 2B the various components shown can be implemented in hardware, software, or a combination of hardware and software. Two or more of the processing unit 110, the storage unit 120, the control unit 130, the navigation module 121, and other units / subunits / modules can be implemented on a single chip or circuit. In other embodiments, at least a part of them can be implemented on separate chips or circuits.
[0062] Figure 3A is Figure 2A a schematic block diagram of an example of carpet detection for a mobile machine. In some embodiments, for example, the instruction(s) I corresponding to the carpet detection method for the mobile machine 100 to detect a carpet n is stored as the navigation module 121 in the storage unit 120 and the stored instruction I is executed by the processing unit 110 n to implement the carpet detection method in the mobile machine 100. Then, the mobile machine 100 can use the RGB-D camera C to detect the carpet. The carpet detection method can be executed in response to a request for the detection of a carpet and / or carpet curling from, for example, the mobile machine 100 itself or the control device 200 (navigation / operating system). The carpet detection method can also be re-executed after each change in the direction of the mobile machine 100 (such as after the movement of the pose S5 in the trajectory T2 Figure 1A ). According to the carpet detection method, the processing unit 110 can obtain the RGB-D image pair G( Figure 3A the frame 310), where each RGB-D image pair G includes the RGB image G r and the depth image G r corresponding to the RGB image G d . One or more RGB-D image pairs G can be obtained. For example, one RGB-D image pair G (such as a qualified RGB-D image pair G that meets certain quality requirements) can be selected for use. Figure 3B is Figure 1A the RGB image G of the scener Schematic diagram. RGB image G r Includes a red channel, a green channel, and a blue channel for representing the RGB image G r of the scene objects (such as carpet C c and floor F) in the depth image G d includes data for representing the distances of the scene objects in the depth image G d .
[0063] In this carpet detection method, the processing unit 110 can also use the first deep learning model M1 to detect the RGB image G of the RGB-D image pair G r of the carpet area A c , and generate a first 2D (two-dimensional) bounding box (Bbox) B1 (which can have a label such as "carpet") to mark each carpet area A c ( Figure 3A box 320). The first deep learning model M1 is a computer model based on, for example, the YOLO algorithm, which is trained by using a large amount of labeled data - that is, the first data set D1 containing the carpet class and a neural network architecture with multiple layers, so as to directly learn from the input RGB G r image to perform a classification task to detect the carpet area A r in the RGB image G c . The first data set D1 is, for example, more than 10,000 carpet images in different scenes. The processing unit 110 can further use the second deep learning model M2 to detect the carpet curling area A r in the RGB image G r , and generate a second 2D bounding box B2 (which can have a label such as "carpet curling") to mark each carpet curling area A r ( Figure 3A block 330). The second deep learning model M2 is a computer model based on, for example, the YOLO algorithm, which is trained by using a large amount of labeled data - that is, the second data set D2 containing the carpet curling class and a neural network architecture with multiple layers, so as to directly learn from the input RGB G r image to perform a classification task to detect the carpet curling area A r in the RGB image G r . The second data set D2 is, for example, more than 10,000 carpet curling images in different scenes.
[0064] Note that in the dataset used to train the deep learning model, carpets and carpet curls are regarded as two independent classes. The first dataset D1 may include images from the COCO (Common Objects in Context) dataset, which contains carpets, blankets, and mats. In addition, the first dataset D1 can also include a dataset containing carpet classes collected from the office. The second dataset D2 can include a dataset containing carpet curl classes collected from the office. The number of annotation instances of carpets and carpet curls can also be balanced by applying data engineering techniques such as adding noise, flipping, grayscale, and shrinking to the original images of the dataset. The first deep learning model M1 and the second deep learning model M2 can be trained using the above datasets based on the YOLO v4 tiny algorithm, and respectively output the first 2D bounding box B1 of the carpet area A c and the second 2D bounding box B2 of the carpet curl area A r .
[0065] In this carpet detection method, the processing unit 110 can also generate a carpet point group P c corresponding to each carpet area A r within the first 2D bounding box B1 of the RGB image G d by matching each pixel of the RGB image G r within the first 2D bounding box B1 with each pixel of the depth image G c of the RGB-D image pair G c ( Figure 3A block 340). Since all pixel points of the RGB image G r within the first 2D bounding box B1 are regarded as the detected carpet area A c , the carpet point group P d representing all or part of the detected carpet C c can be obtained by matching the pixels within the first 2D bounding box B1 with the corresponding pixels in the depth image G c . The processing unit 110 can further generate a carpet curl point group P r corresponding to each carpet curl area A d within the RGB image G of the second 2D bounding box B1 by matching each pixel of the RGB image G with each pixel of the depth image G r of each RGB-D image pair G r ( Figure 3A block 350). Since all pixels of the RGB image G r within the second 2D bounding box B2 are regarded as the detected carpet curl area Ar , so the carpet curl points group P representing all or part of the detected carpet curl C can be obtained by matching the pixels within the second 2D bounding box B2 with the corresponding pixels in the depth image G d . r . r
[0066] Carpet area detection (i.e., Figure 3A block 320) and carpet points group generation (i.e., Figure 3A box 340) can be executed simultaneously with carpet curl area detection (i.e., Figure 3A box 330) and carpet curl points group generation (i.e., Figure 3A box 350), for example, can be executed simultaneously by different threads executed by different processors in the processing unit 110, or other parallel processing mechanisms. In one embodiment, for each RGB-D image pair G, the RGB image G r and the first 2D box B1 generated corresponding to each carpet area A r in the RGB image G c (which can have a label such as "carpet") can be displayed on the display (such as a screen) of the mobile machine 100 to mark the carpet area A c , and the second 2D box B2 generated corresponding to each carpet curl area A r in the RGB image G r (which can have a label such as "carpet curl") can also be displayed on the display to mark the carpet curl area Ar, thereby displaying the detected carpet C c and the detected carpet curl C r to the user U.
[0067] Figure 4A is Figure 2A a schematic block diagram of an example of the movement control of the mobile machine 100. In some embodiments, for example, the instruction(s) I Figure 3A corresponding to the movement control method based on the above carpet detection method (such as n ) is stored as the navigation module 121 in the storage unit 120 and the stored instruction I is executed by the processing unit 110 n to implement the movement control method in the mobile machine 100 to navigate the mobile machine 100, and then the mobile machine 100 can be navigated. The movement control method can be executed in response to a request from, for example, the mobile machine 100 itself or the control device 200 (navigation / operating system), and the obstacles detected by the RGB-D camera C of the mobile machine 100 (such as the carpet curl C c of the carpet C r, walls, furniture, people, pets, and garbage), and then the movement control method can also be re - executed after every change in direction of, for example, the mobile machine 100 (e.g., after the movement of the pose S5 in the trajectory T2 of Figure 1A ).
[0068] According to the movement control method, in addition to performing RGB - D image acquisition as in the above - mentioned carpet detection method ( Figure 4A box 410 of Figure 3A box 310 of Figure 4A box 420 of Figure 3A box 320 of Figure 4A box 430 of Figure 3A box 330 of Figure 3A box 340 of Figure 4A box 450 of Figure 3A box 350 of r each carpet area A in the RGB image G corresponding to each RGB - D image pair G c the generated carpet point group P c as a carpet C c ( Figure 4A box 460 of c the generated carpet point group P corresponding to one carpet area A c as one carpet C c , and the respective carpet point groups P corresponding to different carpet areas A c as different carpets C c . The processing unit 110 can further obtain, based on the generated carpet curl point group P corresponding to each carpet curl area A c , each carpet curl C of each carpet curl area A in the RGB image G corresponding to each RGB - D image pair G r and regard each carpet curl C r as an (obstacle of carpet curl) obstacle O r ( r box 470 of r ). The carpet curl C corresponding to one carpet curl area A r is regarded as one obstacle Or, and the carpet curls C corresponding to different carpet curl areas A r ( Figure 4A box 470 of r are regarded as different obstacles O r ). r are regarded as different obstacles O r . r . Figure 4B is inFigure 4A Schematic block diagram of an example for determining a carpet and carpet curl in an example of movement control of a mobile machine. In one embodiment, to perform carpet confirmation (i.e., Figure 4A block 460 of Figure 4B ), the processing unit 110 may combine all generated carpet point groups Pc into a carpet Cc ( Figure 4A block 461 of r ), and all generated carpet point groups Pc are regarded as a single carpet. To perform carpet curl obstacle confirmation (i.e., r block 470 of r ), the processing unit 110 may generate a 3D (three - dimensional) bounding box B3 to mark the RGB image G of each RGB - D image pair G Figure 4B in each carpet curl region A r corresponding to the generated carpet curl points P Figure 4B group ( r block 471 of r ), and the processing unit 110 may further regard the generated 3D bounding box B3 as an obstacle O r ( r block 472 of r ). By regarding the 3D bounding box B3 itself as the obstacle O r ), the position (and orientation) of the obstacle O r in the 3D coordinate system can be easily obtained, thus achieving collision avoidance. Since each carpet curl can be regarded as an independent obstacle, the generated carpet curl point groups P
[0069] belonging to different carpet curls are not combined together. A carpet curl point group P r corresponding to each carpet region A c in the RGB image G of each RGB - D image pair G c is marked by a 3D bounding box B3, and the carpet curl point groups P Figure 4A corresponding to different carpet curl regions A Figure 4C are marked by different 3D bounding boxes B3. Figure 4A In this carpet detection method, the processing unit 110 may also control the mobile machine 100 to move with reference to each carpet C Figure 4A corresponding to each carpet region A Figure 4C in the RGB image G of each RGB - D image pair G Figure 4A to avoid obstacles Or ( Figure 4A block 480 of Figure 4C ). Figure 4A is a schematic block diagram of an example of movement control in an example of movement control of a mobile machine in Figure 4A . In some embodiments, to perform movement control (i.e., Figure 4C block 480 of Figure 4A ), the processing unit 110 may plan a trajectory T that can avoid obstacles ( Figure 4C block 481 of Figure 4A ). The obstacles to be avoided may be those determined in carpet curl obstacle determination (i.e., Figure 4Athe obstacle O identified in the frame 470) r and the obstacle O detected in the obstacle detection (see Figure 4C the frame 482) o . The trajectory T includes a series of poses (e.g., the poses S1 - S9 of the trajectory T2), and each pose includes the position (e.g., coordinates in a coordinate system) and orientation (e.g., Euler angles in a coordinate system) for moving the machine 100. Based on the height and pitch angle of the RGB - D camera C, the relative position of the carpet C c (or the obstacle to be avoided) relative to the mobile machine 100 can be obtained, and the corresponding pose in the trajectory T can be obtained based on the relative position. The trajectory T can be planned by the above - mentioned global trajectory planner based on the map constructed by the map builder 1211. For example, when identifying the obstacle O r (i.e., Figure 4A the frame 470) or detecting the obstacle O o (see Figure 4C the frame 482), the trajectory T can be replanned. The processing unit 110 can also detect the obstacle O o ( Figure 4C the frame 482) through the sensors in the sensor sub - unit 133. Obstacle detection (i.e., Figure 4C the frame 481) can be performed after, before, or simultaneously with the trajectory planning. The obstacle O can be detected by one or more sensors (e.g., the RGB - D camera C, the LiDAR S1, and the IR S i ) in the sensor sub - unit 133 o . For example, sensor data can be collected by the sensors in the sensor sub - unit 133, and the collected data can be analyzed to identify, for example, obstacles O that suddenly appear or are suddenly detected when approaching o . The sensor data can be collected by different sensors (e.g., the LiDAR S1 and the IR S i ), and a fusion algorithm such as a Kalman filter can be used to fuse the sensor data received from the sensors in order to reduce the uncertainty caused by noisy data, fuse data with different frequencies, and merge data of the same object, etc. For example, an image (e.g., obtaining the RGB image G in the RGB - D image pair G) can be acquired by the RGB - D camera C in the sensor sub - unit 133 r , and other data can be collected by other sensors such as the lidar S1 and the IRS i in the sensor sub - unit 133, and then the collected data can be fused and analyzed to identify the obstacle O o .
[0070] The processing unit 110 can further plan a local trajectory based on the planned global trajectory, sensor data collected by the sensor sub-unit 133 of the mobile machine 100, and the current pose of the mobile machine 100 (i.e., the pose of the mobile machine 100 in the trajectory T at present), for example, including Figure 1A a part of the trajectory T2 including poses S1 - S4 in i . For example, the sensor data can be collected by LiDAR S1 and IR S o , and the collected sensor data can be analyzed to identify obstacles (such as obstacle O Figure 4C ), so that the local path can be planned with reference to the identified obstacles, and the mobile machine 100 can be moved along the planned local path to avoid obstacles. In some embodiments, the above trajectory planner can be used to plan a local trajectory by generating a local trajectory based on the planned global trajectory while considering the identified obstacles (such as avoiding the identified obstacles). In addition, obstacle detection may be re-executed, for example, at each predetermined time interval (such as 1 second), until all motion controls according to the planned trajectory T are completed (see
[0071] the block 483 of Figure 4A ). Figure 4C To perform motion control ( c the block 480 of Figure 4C ), the processing unit 110 can further drive the motor 1321 of the mobile machine 100 according to the trajectory T to control the movement of the mobile machine 100 along the trajectory T ( c the block 483 of c ). For example, the mobile machine 100 can rotate the motor 1321 (and turn the wheels E) according to the trajectory T so that the mobile machine 100 satisfies the position and pose at each pose of the trajectory T. The processing unit 110 can further maintain the current speed of the mobile machine 100 by adjusting the parameters of the motor 1321 in response to the mobile machine 100 being on the carpet C c ( c the block 484 of Figure 4C ). For example, the processing unit 110 first determines whether the mobile machine 100 is already on the carpet C based on the relative position of the carpet C d2 relative to the mobile machine 100 and the current position of the mobile machine 100 (such as the position obtained by the IMU1331), and then changes the power of the driving motor 1321 to offset the friction of the carpet C Figure 4Cof the frame 483), until the movement of the last pose in the trajectory T is completed.
[0072] Figure 5 is a flowchart of a navigation method in some embodiments of the present application. In some embodiments, the navigation method can be implemented in the mobile machine 100 to navigate the mobile machine 100. Therefore, in step 510, an RGB-D image pair G can be obtained by the RGB-D camera C (i.e., Figure 3A of the frame 310). In step 520, the first deep learning model M1 is used to detect the RGB image G of the RGB-D image pair G r in the carpet area A c and generate a first 2D bounding box B1 to mark each carpet area A c (i.e., Figure 3A of the frame 320), and the second deep learning model M2 is used to detect the RGB image G of the RGB-D image pair G r in the carpet curl area A r and generate a second 2D bounding box B2 to mark each carpet curl area A r (i.e., Figure 3A of the frame 330). In step 530, it is determined whether a carpet area A c or a carpet curl area A r has been detected.
[0073] If it is determined that a carpet area A c or a carpet curl area A r has been detected, step 540 is executed; otherwise, step 550 is executed. In step 540, by matching each pixel of the RGB image G within the first 2D bounding box B1 corresponding to the carpet area A c with each pixel of the depth image G of the RGB-D image pair G r , a carpet point group P d corresponding to each carpet area A r in the RGB image G of each RGB-D image pair G is generated c (i.e., c of the frame 340), and by matching each pixel of the RGB image G within the second 2D bounding box B1 corresponding to the carpet curl area A Figure 3A with each pixel of the depth image G of the RGB-D image pair G r , a carpet curl point group P corresponding to each carpet curl area A d in the depth image G of the RGB-D image pair G is generated for each carpet curl area A r in the RGB image G of each RGB-D image pair G r (i.e., Figure 3A of block 350), based on the generated corresponding to each carpet curl area Ar The carpet curl point set P r to obtain the RGB image G of each RGB-D image pair G r each carpet curl region A in r each corresponding carpet curl C r , and each blanket curl C r is taken as an obstacle O r (i.e., Figure 4A the frame 470). In step 550, a trajectory T that can avoid obstacles is planned (i.e., Figure 4C the frame 481), and the motor 1321 of the mobile machine 100 is actuated according to the trajectory T to control the mobile machine 100 to move along the planned trajectory T (i.e., Figure 4C the frame 483), and the parameters of the motor 1321 can be adjusted in response to the mobile machine 100 being on the carpet C c to further maintain the current speed of the mobile machine 100 ( Figure 4C the frame 484). The obstacles to be avoided in the trajectory T may include the obstacle O determined in step 540 r , and the obstacles O detected by the sensors in the sensor sub-unit 133 o (i.e., Figure 4C the frame 482). After step 550, if the mobile device 100 has completed the movement of the last pose in the trajectory T (e.g., the second desired pose S in the trajectory T2 d2 ), the method ends and the navigation of the mobile device 100 ends; otherwise, the method will be executed again (i.e., step 510 and its subsequent steps). For example, the method can be executed again after each direction change of the mobile machine 100 (e.g., after the movement of the pose S5 in the Figure 1A trajectory T2). It should be noted that by executing the method again after each direction change of the mobile machine 100, different angles of the carpet C c and the carpet curl C r can be obtained through the newly acquired RGB-D image pair G, so the relative position with respect to the mobile machine 100 can be updated accordingly.
[0074] As can be seen from the above navigation method, the carpet C c and the carpet curl C r will be detected before and during the movement of the mobile device 100, so that when the mobile machine 100 moves on the carpet C c , the speed of the mobile machine 100 can be maintained while avoiding the carpet curl C r serving as an obstacle O rThis method is simple because it only includes several steps, does not require complex calculations, and only needs an economical RGB-D camera C to detect the carpet C c and carpet curling C r , so the detection of the carpet C c and carpet curling C r can be carried out in real time and requires a small amount of cost. Compared with the existing navigation methods, the stability and safety of the mobile machine 100 moving in a carpeted scenario are effectively improved.
[0075] Those skilled in the art can understand that all or part of the methods in the above embodiments can be implemented by using one or more computer programs to instruct relevant hardware. In addition, one or more programs can be stored in a non-transitory computer-readable storage medium. When one or more programs are executed, all or part of the corresponding methods in the above embodiments are executed. Any reference to storage, memory, database, or other media may include non-transitory and / or transitory memory. Non-transitory memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, solid-state drive (SSD), etc. Volatile memory may include random access memory (RAM), external cache memory, etc.
[0076] The processing unit 110 (and the above-mentioned processor) may include a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate, transistor logic devices, and discrete hardware components. The general-purpose processor may be a microprocessor, and may be any conventional processor. The storage unit 120 (and the above-mentioned memory) may include internal storage units such as hard disks and internal memories. The storage unit 120 may also include external storage devices, such as plug-in hard disks, smart media cards (SMCs), secure digital (SD) cards, and flash memory cards.
[0077] The exemplary units / modules and methods / steps described in the embodiments can be implemented by software, hardware, or a combination of software and hardware. Whether these functions are implemented by software or hardware depends on the specific application and design constraints of the technical solution. The above carpet detection method and the mobile machine 100 can be implemented in other ways. For example, the division of units / modules is only a logical function division. In actual implementation, other division methods can also be adopted, that is, multiple units / modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the above mutual coupling / connection can be a direct coupling / connection or a communication connection, or an indirect coupling / connection or a communication connection through some interfaces / devices, and can also be in electrical, mechanical or other forms.
[0078] The above embodiments are only used to illustrate the technical solutions of the present invention and are not used to limit the technical solutions of the present invention. Although the present invention has been described in detail in combination with the above embodiments, the technical solutions in the above respective embodiments can still be modified, or some technical features can be equivalently replaced, so that these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the respective embodiments of the present invention, and all should be included within the scope of protection of the present invention.
Claims
1. A carpet detection method, comprising: Obtaining one or more RGB-D image pairs through an RGB-D camera, wherein each RGB-D image pair includes an RGB image and a depth image; Detecting one or more carpet regions in the RGB image of the one or more RGB-D image pairs, and generating first 2D bounding boxes based on a first data set including carpet classes and using a first deep learning model to label each of the carpet regions; Detecting one or more carpet curling regions in the RGB image of the one or more RGB-D image pairs, and generating second 2D bounding boxes based on a second data set including carpet curling classes and using a second deep learning model to label each of the carpet curling regions; Generating a set of carpet points corresponding to each of the carpet regions in the RGB image of each of the RGB-D image pairs by matching each pixel of the RGB image within the first 2D bounding box corresponding to each of the carpet regions with each pixel of the depth image in each of the RGB-D image pairs; And Generating a set of carpet curling points corresponding to each of the carpet curling regions in the RGB image of each of the RGB-D image pairs by matching each pixel of the RGB image within the second 2D bounding box corresponding to each of the carpet curling regions with each pixel of the depth image in each of the RGB-D image pairs.
2. The method according to claim 1, further comprising: Combining all the generated sets of carpet points into a carpet.
3. The method according to claim 1, further comprising: Generating 3D bounding boxes to label the generated set of carpet curling points, wherein the set of carpet curling points corresponds to each of the carpet curling regions in the RGB image of each of the RGB-D image pairs.
4. The method according to claim 1, wherein the first deep learning model and the second deep learning model are based on the YOLO algorithm.
5. The method according to claim 4, wherein the first deep learning model is trained using the first data set including the COCO data set.
6. A method for controlling a mobile machine to move, wherein the mobile machine has an RGB-D camera, The method comprises: On one or more processors of the mobile machine: Obtaining one or more RGB-D image pairs through an RGB-D camera, wherein each RGB-D image pair includes an RGB image and a depth image; Detecting one or more carpet regions in the RGB image of the one or more RGB-D image pairs, and generating first 2D bounding boxes based on a first data set including carpet classes and using a first deep learning model to label each of the carpet regions; Detecting one or more carpet curling regions in the RGB image of the one or more RGB-D image pairs, and generating second 2D bounding boxes based on a second data set including carpet curling classes and using a second deep learning model to label each of the carpet curling regions; Generating a set of carpet points corresponding to each carpet area in the RGB image within the first 2D bounding box corresponding to each such carpet area by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair; Generating a set of carpet curl points corresponding to each carpet curl area in the RGB image within the second 2D bounding box corresponding to each such carpet curl area by matching each pixel of the RGB image with each pixel in the depth image of each such RGB-D image pair; Taking the generated set of carpet points corresponding to each carpet area in the RGB image corresponding to each such RGB-D image pair as the carpet; Obtaining one or more carpet curls based on the generated set of carpet curl points corresponding to each carpet curl area in the RGB image corresponding to each such RGB-D image pair, and taking each such carpet curl as an obstacle, where each such carpet curl corresponds to each such carpet curl area; and Controlling the mobile machine to move with reference to the carpet corresponding to each carpet area in the RGB image corresponding to each such RGB-D image pair to avoid the obstacle.
7. The method according to claim 6, wherein said taking the generated set of carpet points corresponding to each carpet area in the RGB image corresponding to each such RGB-D image pair as the carpet comprises: Combining all the generated sets of carpet points into the carpet; and Controlling the mobile machine to move with reference to the carpet corresponding to each carpet area in the RGB image corresponding to each such RGB-D image pair to avoid the obstacle, comprises: Controlling the mobile machine to move with reference to the combined carpet to avoid the obstacle.
8. The method according to claim 6, wherein said obtaining the one or more carpet curls based on the generated set of carpet curl points corresponding to each carpet curl area in the RGB image corresponding to each such RGB-D image pair, and taking each such carpet curl as the obstacle comprises: Generating a 3D bounding box to mark the generated set of carpet curl points corresponding to each carpet curl area in the RGB image corresponding to each such RGB-D image pair; and Taking the generated 3D bounding box as the obstacle.
9. A mobile machine, comprising: An RGB-D camera; One or more processors; and One or more memories storing one or more computer programs, the one or more computer programs being executed by the one or more processors, wherein the one or more computer programs include a plurality of instructions for: Obtaining one or more RGB-D image pairs by an RGB-D camera, where each such RGB-D image pair includes an RGB image and a depth image; Detecting one or more carpet areas in the RGB image of the one or more RGB-D image pairs, and generating a first 2D bounding box to mark each such carpet area based on a first data set including carpet classes and using a first deep learning model; Detect one or more carpet curl regions in the RGB image of the one or more RGB-D image pairs, and generate second 2D bounding boxes to label each of the carpet curl regions based on a second dataset containing carpet curl classes, using a second deep learning model; Generate a set of carpet points corresponding to each carpet region in the RGB image of each RGB-D image pair by matching each pixel of the RGB image within the first 2D bounding box corresponding to each carpet region with each pixel of the depth image of each RGB-D image pair; and Generate a set of carpet curl points corresponding to each carpet curl region in the RGB image of each RGB-D image pair by matching each pixel of the RGB image within the second 2D bounding box corresponding to each carpet curl region with each pixel of the depth image of each RGB-D image pair.
10. The mobile machine according to claim 9, wherein the one or more computer programs further comprise a plurality of instructions for: Taking the generated set of carpet points corresponding to each carpet region in the RGB image of each RGB-D image pair as a carpet; Obtaining one or more carpet curls based on the generated set of carpet curl points corresponding to each carpet curl region in the RGB image of each RGB-D image pair, and taking each carpet curl as an obstacle, wherein each carpet curl corresponds to each carpet curl region; and Controlling the mobile machine to move with reference to the carpet corresponding to each carpet region in the RGB image of each RGB-D image pair to avoid the obstacle.
Citation Information
Patent Citations
Ground obstacle map marking method, mobile robot and storage medium
CN111899299A
Target detection method and device
CN111950543A