Detection Method, Device, Mobile Robot and Storage Medium
By using monocular cameras and depth prediction models on mobile robots, efficient and accurate ranging is achieved, and the problems of high ranging costs, strict environmental requirements and complex algorithms in the prior art are solved.
Patent Information
- Application Number
- CN202011529623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-12-22
AI Technical Summary
When measuring distances in existing mobile robots, laser rangefinders are costly and have strict environmental requirements. The algorithms of binocular stereoscopic vision methods are complex and the computing resources are large.
The monocular camera is used to collect images, predict the depth image through the depth prediction model, and determine the distance between the target object and the mobile robot.
Reduces ranging costs, improves ranging accuracy, simplifies algorithms, and reduces requirements for computing and deployment resources.
Smart Images

Figure CN114723799B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computer technology, and in particular, to a detection method, device, mobile robot, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, mobile robots are more and more widely used, and mobile robots usually have ranging requirements. Taking a sweeping robot as an example, when the sweeping robot performs a cleaning task in an operating environment, it needs to measure the distance to obstacles in the forward direction, and avoid obstacles based on the ranging result.
[0003] Currently, mobile robots usually use a laser rangefinder for ranging. However, the hardware production of the laser rangefinder is difficult, the ranging cost is high, and the laser rangefinder has relatively strict requirements on the environment and is easily affected by the environment, resulting in insufficient ranging accuracy; the binocular stereo vision measurement method has high efficiency and appropriate accuracy, but the algorithm is relatively complex and has high requirements on computing and deployment resources. Summary of the Invention
[0004] Embodiments of the present application provide a detection method, device, mobile robot, and storage medium, which are used to ensure good ranging accuracy, reduce the ranging cost, have a relatively simple algorithm, and have less requirements on computing and deployment resources.
[0005] In a first aspect, a detection method provided in an embodiment of the present application includes: obtaining a first monocular image collected by a monocular camera, and identifying a target object in the first monocular image; predicting a first depth image corresponding to the first monocular image by using a depth prediction model; and determining the distance between the target object and the mobile robot based on the first depth image.
[0006] In a second aspect, a detection method provided in an embodiment of the present application includes: collecting an image of a calibration object located at a calibration distance from the mobile robot by using a monocular camera to obtain a second monocular image; where the calibration distance is known; predicting a second depth image corresponding to the second monocular image by using a depth prediction model; determining the relative distance between the calibration object and the mobile robot based on the second depth image; and obtaining a scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibration distance.
[0007] In a third aspect, a detection device provided in an embodiment of the present application includes: a first acquisition module, configured to obtain a first monocular image collected by a monocular camera, and identify a target object in the first monocular image; a first processing module, configured to predict a first depth image corresponding to the first monocular image by using a depth prediction model; and the first processing module is further configured to determine the distance between the target object and the mobile robot based on the first depth image.
[0008] Fourth aspect, an embodiment of the present application provides a detection device, including: a second acquisition module, configured to collect an image of a calibration object located at a calibrated distance from a mobile robot by using a monocular camera, so as to obtain a second monocular image; wherein, the calibrated distance is known; a second processing module, configured to predict a second depth image corresponding to the second monocular image by using a depth prediction model; the second processing module is further configured to determine a relative distance between the calibration object and the mobile robot based on the second depth image; the second processing module is further configured to obtain a scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibrated distance.
[0009] Fifth aspect, an embodiment of the present application provides a detection device, including: a storage component and a processing component;
[0010] The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component;
[0011] The processing component is configured to: acquire a first monocular image collected by a monocular camera, and identify a target object in the first monocular image; predict a first depth image corresponding to the first monocular image by using a depth prediction model; determine a distance between the target object and the mobile robot based on the first depth image.
[0012] Sixth aspect, an embodiment of the present application provides a mobile robot, including an acquisition component, a storage component and a processing component;
[0013] Wherein, the storage component stores one or more computer instructions; the one or more computer instructions are called and executed by the processing component;
[0014] The acquisition component is configured to acquire a first monocular image;
[0015] The processing component is configured to:
[0016] acquire the first monocular image collected by the acquisition component, and identify a target object in the first monocular image; predict a first depth image corresponding to the first monocular image by using a depth prediction model; determine a distance between the target object and the mobile robot based on the first depth image
[0017] Seventh aspect, an embodiment of the present application provides a mobile robot, including an acquisition component, a storage component and a processing component;
[0018] Wherein, the storage component stores one or more computer instructions; the one or more computer instructions are called and executed by the processing component;
[0019] The acquisition component is used to collect an image of a calibration object located at a calibrated distance from the mobile robot to obtain a second monocular image; wherein, the calibrated distance is known.
[0020] The processing component is used for:
[0021] Obtain the second monocular image collected by the acquisition component; wherein, the calibrated distance is known; use a depth prediction model to predict the second depth image corresponding to the second monocular image; based on the second depth image, determine the relative distance between the calibration object and the mobile robot; based on the relative distance between the calibration object and the mobile robot and the calibrated distance, obtain a scale ratio relationship.
[0022] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, the above detection method is implemented.
[0023] In the embodiments of the present application, by using the first monocular image containing the target object collected by the monocular camera, the distance between the target object and the mobile robot can be measured. It can be achieved as long as there is a monocular camera in the mobile robot. The ranging cost is low, and the ranging result is relatively less affected by the environment, ensuring better ranging accuracy. At the same time, the ranging algorithm belongs to the category of ranging using monocular vision technology, with a relatively low algorithm complexity, not very high requirements for computing and deployment resources, relatively little occupancy of computing resources, and high ranging efficiency.
[0024] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 Shows a flowchart of an embodiment of a detection method provided by an embodiment of the present application;
[0027] Figure 2 Shows a schematic structural diagram of an embodiment of a detection system provided by an embodiment of the present application;
[0028] Figure 3a Shows a schematic diagram of marked point annotation in an actual application of an embodiment of the present application;
[0029] Figure 3bShows the schematic diagram of marker point annotation in another practical application of the embodiment of the present application;
[0030] Figure 4 Shows the principle diagram of camera pinhole imaging;
[0031] Figure 5 Shows the flowchart of another embodiment of a detection method provided by the embodiment of the present application;
[0032] Figure 6 Shows the schematic diagram of scale calibration in a practical application of the embodiment of the present application;
[0033] Figure 7 Shows the structural schematic diagram of an embodiment of a detection device provided by the embodiment of the present application;
[0034] Figure 8 Shows the structural schematic diagram of an embodiment of a detection device provided by the embodiment of the present application;
[0035] Figure 9 Shows the structural schematic diagram of an embodiment of a mobile robot provided by the embodiment of the present application;
[0036] Figure 10 Shows the structural schematic diagram of another embodiment of a detection device provided by the embodiment of the present application. Detailed implementation manners
[0037] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.
[0038] In some processes described in the specification, claims and the above-mentioned drawings of the present application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0039] The technical solution of the embodiment of the present application can be applied to various mobile robots that need to sense the distance from an object in the working environment, such as floor-sweeping robots, intelligent dialogue robots, and other robots that can move autonomously in the working environment for logistics distribution, environmental purification, intelligent customer service, etc. The moving mode can be, for example, wheeled, tracked, etc., and the present application does not limit this. Among them, the object described herein can refer to a person, an animal, or any other static or dynamic object in the working environment.
[0040] As described in the background art, mobile robots usually have a ranging requirement, that is, they need to determine the distance information between the mobile robot and certain objects in the working environment, so as to perform corresponding processing based on the distance information, such as obstacle avoidance, navigation, or mapping. However, currently, using a laser rangefinder not only has a high cost, but also has a high risk that the ranging result is greatly affected by the environment.
[0041] The inventors found in the process of implementing the present invention that although the binocular stereo vision ranging method can be used for ranging, however, the algorithm complexity of the binocular stereo vision ranging method is relatively high and will occupy a large amount of computing resources. And mobile robots need to process a large amount of data during the moving process. If ranging occupies a large amount of computing resources, it will increase the system pressure.
[0042] In order to ensure the ranging result, reduce the ranging cost, and reduce the ranging complexity, the inventors further studied and proposed the technical solution of the present application. The embodiment of the present application can achieve ranging by using the monocular image collected by a monocular camera. Compared with the laser ranging method, it can reduce the difficulty of hardware production and the ranging cost, and the ranging result is relatively less affected by the environment, ensuring better ranging accuracy; compared with the binocular stereo vision ranging method, it can reduce the algorithm complexity and reduce the occupation of computing resources.
[0043] Next, the technical solution in the embodiment of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiment of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0044] Figure 1The flowchart of an embodiment of a detection method provided by an embodiment of the present application is shown. The detection method provided by this embodiment can be executed by a mobile robot. In addition, in practical applications, the mobile robot can also establish a connection with a user terminal and can also establish a connection with a server to receive corresponding instructions from the user terminal or the server and execute corresponding operations. Therefore, the detection method provided by this embodiment can also be executed by a user terminal or a server connected to the mobile robot. The present application does not make specific restrictions on this.
[0045] Among them, the user terminal can be configured in user equipment such as mobile phones, tablets, etc., and the server can be a cloud server, etc. Refer to Figure 2 , which shows a detection system to which the technical solution of the present application may be applicable. The mobile robot 201 can establish connections with the user terminal 202 and the server 203 respectively (for the sake of easy understanding and explanation, Figure 2 the mobile robot in is a sweeping robot, the user terminal is represented by the user equipment it is configured with, and the server is represented by the server image), for example, the connection can be specifically realized through Bluetooth, wifi (Wireless Fidelity), etc. The connection establishment method is the same as that in the prior art and will not be elaborated here.
[0046] The user terminal 202 can directly or through the server 203 send corresponding control instructions to the mobile robot based on the corresponding user requests; of course, the mobile robot 201 can also set a corresponding control panel for the user to trigger corresponding control instructions, or recognize the corresponding control instructions issued by the user's voice through voice recognition, etc. The present application does not make specific restrictions on this. Among them, the mobile robot 201 can execute corresponding operations based on the control instructions.
[0047] It should be noted that Figure 2 is only an example of a system to which the technical solution of the present application can be applicable. In practical applications, the mobile robot can also be only connected to the user terminal, or only connected to the server, or an independently working device.
[0048] Optionally, since the mobile robot is usually a comprehensive system integrating multiple functions such as environmental perception, dynamic decision-making and planning, and behavior control and execution, it can automatically execute work. In order to improve the processing efficiency, etc., the technical solution of this embodiment can be executed by the mobile robot.
[0049] Refer to Figure 1 , the method may include the following steps:
[0050] 101: Obtain a first monocular image collected by a monocular camera and identify the target object in the first monocular image.
[0051] Specifically, the monocular camera can be a camera device consisting of one camera in a mobile robot, or a monocular camera in a multi-camera in a mobile robot. Among them, the multi-camera can be regarded as a camera device composed of multiple cameras, that is, the multi-camera includes multiple monocular cameras. For example, Figure 2 When the shown floor cleaning robot is provided with a monocular camera 204, image acquisition of the working environment is performed through the monocular camera. Another example is that if the floor cleaning robot is provided with a binocular camera, any one of the monocular cameras in the binocular camera is selected to perform image acquisition of the working environment.
[0052] Optionally, the first monocular image can be any image frame in the monocular video stream collected by the mobile robot using the monocular camera. Therefore, step 101 can specifically be taking any image frame in the monocular video stream collected by the monocular camera as the first monocular image. Each image frame in the monocular video stream can be processed according to the technical solution of this embodiment.
[0053] In addition, since the mobile robot does not always have a ranging requirement, in order to save resources, the operation of step 101 can be performed when the mobile robot is in a specific working mode. For the floor cleaning robot, it can be when the floor cleaning robot is in the cleaning working mode, taking any image frame in the monocular video stream collected by the monocular camera as the first monocular image. In the cleaning working mode, the floor cleaning robot performs a cleaning task, moves in its working environment and may perform path planning, etc. Therefore, obstacle avoidance is required, that is, it is necessary to identify the target object and perform corresponding obstacle avoidance processing based on the distance from the target object.
[0054] Among them, the switching of the specific working mode of the mobile robot can be realized in response to a switching instruction generated by the user terminal based on the user's switching operation. The switching instruction can be directly sent to the mobile robot, or sent to the mobile robot through the server.
[0055] When the technical solution of this embodiment is executed by the user terminal or the server connected to the mobile robot, the mobile robot can upload the first monocular image it collects to the user terminal or the server in real time.
[0056] Among them, the recognition of the first monocular image can be realized by using image recognition technology. For example, an image recognition model is used to recognize the target object in the first monocular image. The image recognition model can be pre-trained based on sample images marked with target objects. The target object can be determined according to the actual application scenario of ranging. It can refer to various static or dynamic targets in the working environment of the mobile robot, such as various objects such as people, animals, and furniture.
[0057] It can be understood that before performing image recognition on the first monocular image, some necessary preprocessing operations may also be performed on the first monocular image, such as denoising processing, image enhancement, grayscale conversion, binarization, and / or edge detection. Specifically, object detection is performed on the preprocessed first monocular image to identify the target object in the first monocular image. These preprocessing operations are the same as those in the prior art and will not be elaborated here.
[0058] Among them, the working environment referred to in this article may refer to the workplace of the mobile robot.
[0059] 102: Use the depth prediction model to predict the first depth image corresponding to the first monocular image.
[0060] Specifically, the first monocular image is an RGB color image. Among them, RGB represents the colors of the three channels of red (R, red), green (G, green), and blue (B, blue). In this embodiment, the depth prediction model can convert the RGB color image into a depth image (English: depth image). The depth prediction model can predict the depth value of each pixel point in the color image. Among them, the depth image is also called a range image (English: range image), which refers to an image in which the depth values from the camera to each point in the real three-dimensional space are used as pixel values. The depth value is also the distance from the camera to each point in the real three-dimensional space.
[0061] Among them, the depth prediction model can be obtained by pre-training, and its training process will be introduced accordingly below.
[0062] 103: Based on the first depth image, determine the distance between the target object and the mobile robot.
[0063] Specifically, each object in the first monocular image has its corresponding pixel point in the first depth image. Since the target object has a certain size, there are more than one pixel points corresponding to the target object in the first depth image.
[0064] As a possible implementation, when obtaining the depth value of the target object from the first depth image, the depth values of all its corresponding pixel points can be accumulated and averaged, and the average value is used as the depth value of the target object.
[0065] As another possible implementation, when obtaining the depth value of the target object from the first depth image, one pixel point can be randomly selected from all the corresponding pixel points, and the depth value of the selected pixel point is used as the depth value of the target object. Among them, the randomly selected pixel point is, for example, the central pixel point at the center position of all pixel points; the randomly selected pixel point is also, for example, any pixel point on the lower edge of the target object. Among them, the lower edge of the target object is the edge where the target object is close to or in contact with the ground, and the lower edge of the target object can be detected from the edge of the target object in the first monocular image.
[0066] As can be seen from the above description, the depth values in the depth image are the distances from the camera to each point in the real three-dimensional space. Therefore, the depth value of the target object can be understood as the distance between the target object and the monocular camera.
[0067] In an application scenario, for a mobile robot equipped with a monocular camera, the distance between the target object and the monocular camera is often regarded as the distance between the target object and the mobile robot. Therefore, the depth value of the target object is both the distance between the target object and the monocular camera and the distance between the target object and the mobile robot.
[0068] In another application scenario, due to the limited installation position of the monocular camera in the mobile robot, there is a distance deviation between "the distance between the target object and the mobile robot" and "the distance between the target object and the monocular camera". At this time, it is necessary to correct "the distance between the target object and the monocular camera" to obtain "the distance between the target object and the mobile robot". At this time, "the distance between the target object and the monocular camera" is the depth value of the target object, and "the distance between the target object and the mobile robot" is the sum of the depth value of the target object and the distance deviation. The distance deviation is set according to actual business requirements, and the distance deviation is, for example, 1 millimeter.
[0069] Therefore, in this application scenario, it can be to sum the depth value of the target object obtained based on the first depth image and the distance deviation to obtain the distance between the target object and the mobile robot.
[0070] In this embodiment, by using the first monocular image collected by the monocular camera containing the target object, the distance between the target object and the mobile robot can be measured. It only needs to be equipped with a monocular camera in the mobile robot to achieve this. The ranging cost is low, and the ranging result is relatively less affected by the environment, ensuring better ranging accuracy. At the same time, the ranging algorithm belongs to the category of ranging using monocular vision technology. The algorithm complexity is low, the requirements for computing and deployment resources are not very high, the occupation of computing resources is relatively small, and the ranging efficiency is high.
[0071] Since the first depth image is obtained by predicting the first monocular image using a depth prediction model, and the first monocular image is a single image collected by a monocular camera, the true size of the target object cannot be determined. Therefore, the depth information contained in the first depth image predicted based on the first monocular image is not the actual distance from the mobile robot, but a relative distance.
[0072] In some embodiments, in order to obtain the actual distance of the target object from the mobile robot and improve the ranging accuracy, the algorithm complexity of the traditional epipolar geometry method is relatively high. The inventor thought that as long as the scale ratio relationship between the relative distance and the actual distance is determined, through simple scale conversion, the actual distance can be calculated based on the relative distance.
[0073] Therefore, as an optional implementation manner, step 103 specifically is: determining the relative distance between the target object and the mobile robot based on the first depth image; calculating the distance between the target object and the mobile robot based on the relative distance and the scale ratio relationship.
[0074] Among them, the scale ratio relationship can be set before the mobile robot leaves the factory and configured in the mobile robot. Of course, it can also be detected by prompting the user to perform operations such as deploying and calibrating the object before the mobile robot is used for the first time.
[0075] In an optional implementation manner, when determining the scale ratio relationship, the monocular camera can be used to collect an image of the calibration object located at the calibration distance of the mobile robot to obtain a second monocular image; where the calibration distance is known; using the depth prediction model to predict the second depth image corresponding to the second monocular image; determining the relative distance between the calibration object and the mobile robot based on the second depth image; obtaining the scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibration distance.
[0076] Among them, the calibration distance can be understood as the actual distance between the calibration object and the mobile robot. Since the scale ratio relationship describes the ratio relationship between the "actual distance between the calibration object and the mobile robot" and the "relative distance between the calibration object and the mobile robot", therefore, in actual application, the "relative distance between the target object and the mobile robot" can be scaled based on the scale ratio relationship to obtain the "actual distance between the target object and the mobile robot".
[0077] For the convenience of understanding, it can be assumed that the calibration distance is Z2, and the relative distance between the calibration object and the mobile robot is D2. Then Z2 / D2 can be used as the scale ratio relationship Scale. At the same time, it is assumed that the relative distance between the target object and the mobile robot determined based on the first depth image is D1, and the distance between the target object and the mobile robot is Z1. Then Z1 = D1 × Scale.
[0078] Among them, the calibration object can be an object with a specific shape, specific size or specific color that is convenient for image recognition. For example, it can be a rectangular baffle with a certain height and width, etc. Of course, it can also be any other object that is easy to identify. This application does not impose specific restrictions on this.
[0079] In actual applications, some business scenarios may require obtaining more information about the target object. For example, in the obstacle avoidance scenario of a sweeping robot, in order to improve the obstacle avoidance effect, taking the target object as a sofa, there is usually a certain space between the lower edge of the sofa and the ground. If the height from the ground is greater than the height of the sweeping robot, the sweeping robot does not need to avoid obstacles on the sofa, and can clean the ground covered by the sofa to ensure the cleaning effect. In addition, if the sweeping machine knows the width information of the sofa, it can also plan the cleaning path of the area around the sofa. Therefore, when the sweeping robot performs obstacle avoidance processing, in addition to being based on the distance information between it and the target object, it can also be combined with the height information of the lower edge of the target object from the ground. In addition, it can also be combined with the width information of the target object to perform obstacle avoidance processing.
[0080] In addition, in actual applications, some target objects have suspended parts relative to the ground, such as hollow furniture, while some target objects have no suspended parts relative to the ground, such as ordinary furniture. Among them, the target object has a suspended part relative to the ground can be understood as the lower edge of the target object is at a certain height from the ground, and the target object has no suspended part relative to the ground can be understood as the lower edge of the target object is close to the ground.
[0081] Therefore, in some embodiments, before obtaining more information about the target object, it can be determined whether the target object has a suspended portion relative to the ground, and different strategies can be used for target objects in different situations to obtain more information about the target object.
[0082] In practical applications, when a target object has a suspended portion relative to the ground, there is often a need to obtain information such as the width of the target object and the height from the ground.
[0083] Therefore, as an optional implementation method, if the target object has a suspended part relative to the ground, the first boundary of the target object close to the ground and not in contact with the ground in the first monocular image can be detected; the second boundary mapped to the ground by the first boundary is identified; the two vertex pixel points on any diagonal line in the area formed by the first boundary and the second boundary are used as marking points; based on the position information of the two marking points in the first monocular image, the width information of the target object and the height information from the ground are determined.
[0084] Specifically, when detecting the first boundary, edge detection can be performed on the target object in the first monocular image, and a bounding box surrounding the target object can be determined based on the edge detection result of the target object. The boundary in the bounding box that is close to the ground and does not touch the ground is used as the first boundary.
[0085] The first boundary can refer to the lower edge of the target object close to the ground. In practical applications, there may be a situation where the lower edge of the target object is an irregular curve. Therefore, the horizontal line where the pixel point closest to the ground in the lower edge can be selected as the first boundary.
[0086] Therefore, as an alternative implementation, detecting the first boundary of the target object in the first monocular image that is close to the ground and does not touch the ground specifically includes: performing edge detection on the target object in the first monocular image; determining the lower edge of the target object close to the ground according to the edge result; using the horizontal line where the pixel point closest to the ground in the lower edge as the first boundary.
[0087] It can be understood that when using the horizontal line where the pixel point closest to the ground in the lower edge as the first boundary, the first boundary is parallel to the ground, and the area formed by the first boundary and the second boundary is a rectangular area.
[0088] It should be noted that if the irregularly shaped lower edge is used as the first boundary, the area formed by the first boundary and the second boundary at this time is an irregular area, not a rectangular area.
[0089] For ease of understanding, refer to Figure 3a , which shows a schematic diagram of marker point annotation in an actual application of an embodiment of the present application. Figure 3a The crib in [reference] can be considered as a hollowed-out piece of furniture. P1P4 is the lower edge of the crib, and this lower edge is at a certain height from the ground.
[0090] In Figure 3a , the target object is a crib, and there is a certain height between its lower edge and the ground. The line segment P1P4 is the first boundary obtained by recognition, and P2P3 is the second boundary (the second boundary is on the ground). The two vertex pixel points P1 and P3 on the diagonal can be used as marker points, or alternatively, the two vertex pixel points P4 and P2 on the diagonal can also be used as marker points.
[0091] Figure 4 shows the principle of camera pinhole imaging. Among them, p(x, y) is the coordinate of the pixel point in the image plane, and P(X c , Y c , Z c ) is an arbitrary point of the target object in the real three-dimensional space, and f is the focal length of the camera lens. According to the principle of pinhole imaging, it can be known that: x = X c / Zc ×f and y = Y c / Z c ×f. Since the scale ratio Scale = Z c / D c , given Scale, D c , f, x, and y, Z can be calculated c , and then X can be calculated c , Y c . Among them, D c is the relative distance between point P obtained from the first depth image and the mobile robot, and Z c is the target distance, and Z c is the actual distance between point P and the mobile robot.
[0092] Assume that the image position information of p1 in the image plane is p1(x1, y1). According to the above calculation method, it can be known that its corresponding point in the real three-dimensional space is P1(X1, Y1, Z1); assume that the image position information of p3 in the image plane is p3(x3, y3), and it can be known that its corresponding point in the real three-dimensional space is P3(X3, Y3, Z3). Then Figure 3a the width of the middle bed Width = |X1 - X3|, Figure 3a the height of the middle bed Height = |Y1 - Y3|. The target distance between the bed and the mobile robot can be determined based on any pixel point on the first boundary, such as p1, and the target distance between the bed and the mobile robot is Z1.
[0093] In practical applications, since the mobile robot can be a sweeping robot and the sweeping robot needs to support an obstacle avoidance function, after obtaining the distance, width information, and height information between the target object and the mobile robot, in some embodiments, the method may further include: performing obstacle avoidance processing on the mobile robot based on the distance, width information, and height information between the target object and the mobile robot. The specific obstacle avoidance processing process is similar to the prior art and will not be elaborated here.
[0094] In practical applications, for the case where the target object has no suspended part relative to the ground, there is often a need to obtain information such as the width information of the target object.
[0095] Therefore, as an alternative implementation, if the target object has no suspended part relative to the ground, the first boundary where the target object contacts the ground in the first monocular image can be detected; the two end-point pixels of the first boundary are used as marker points; based on the position information of the two marker points in the first monocular image, the width information of the target object is determined.
[0096] The first boundary may refer to the lower edge where the target object contacts the ground. In practical applications, there may be cases where the lower edge of the target object is an irregular curve. Therefore, the horizontal line where the pixel point closest to the ground in the lower edge can be selected as the first boundary.
[0097] Therefore, as an alternative implementation, specifically detecting the first boundary where the target object contacts the ground in the first monocular image is as follows: performing edge detection on the target object in the first monocular image; determining the lower edge where the target object contacts the ground according to the edge result; and taking the horizontal line where the pixel point closest to the ground in the lower edge as the first boundary.
[0098] It can be understood that when taking the horizontal line where the pixel point closest to the ground in the lower edge as the first boundary, the first boundary is parallel to the ground.
[0099] For ease of understanding, refer to Figure 3b , which shows a schematic diagram of marker point annotation in another practical application of the embodiment of the present application. Figure 3b The crib in can be considered as ordinary furniture, and its lower edge is close to the ground. The line segment p5p6 is the first boundary obtained by recognition, and the two end point pixels p5 and p6 are used as marker points. Assume that the image position information of p5 in the image plane is p5(x5,y5). According to the above calculation method, it can be known that the corresponding point in the real three-dimensional space is P5(X5,Y5,Z5); assume that the image position information of p6 in the image plane is p6(x6,y6), and it can be known that the corresponding point in the real three-dimensional space is P6(X6,Y6,Z6). Then Figure 3b the width Width of the bed in = |X5 - X6|.
[0100] Since in practical applications, the mobile robot can be a sweeping robot, and the sweeping robot needs to support an obstacle avoidance function. Therefore, after obtaining the distance and width information between the target object and the mobile robot, in some embodiments, the method may further include: performing obstacle avoidance processing on the mobile robot based on the distance and width information between the target object and the mobile robot. The specific obstacle avoidance processing process is similar to the prior art and will not be elaborated here.
[0101] In some embodiments of the present application, the depth prediction model can be trained using a supervised learning method or an unsupervised learning method. The depth prediction model can, for example, but is not limited to: CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory artificial neural network).
[0102] As an alternative implementation, the depth prediction model can be pre-trained as follows:
[0103] Obtain a sample image and the depth image corresponding to the sample image; use the sample image as the model input and the depth image corresponding to the sample image as the training label to train the depth prediction model.
[0104] As another alternative implementation, the depth prediction model can be trained using an unsupervised learning method and can be pre-trained as follows:
[0105] Obtain a first image and a second image collected by a binocular camera; use the first image as the model input and the second image as the model constraint to train the depth prediction model.
[0106] Optionally, the first image can be used as the model input, and the first depth map to be optimized is obtained by inputting it into the depth prediction model; based on the first depth map to be optimized, the first image, and the second image, first error information is obtained; the depth prediction model is optimized using the first error information.
[0107] Among them, first, a difference operation can be performed on the first depth map to be optimized and the first image to obtain a reconstructed image, and then a difference operation is performed on the reconstructed image and the second image to obtain the first error information, and the depth prediction model is optimized using the first error information.
[0108] Among them, the first image can refer to the left view collected by the binocular camera, then the second image is the right view, or the first image can be the right view, then the second image is the left view.
[0109] As yet another alternative implementation, the depth prediction model can be trained using an unsupervised learning method and can be pre-trained as follows:
[0110] Obtain adjacent first image frames and second image frames in the video stream collected by a monocular camera; use the first image frame as the model input and the second image frame as the model constraint to train the depth prediction model.
[0111] This training method is to use two adjacent image frames in the video stream collected by the monocular camera to train the model.
[0112] Specifically, the first image frame can be used as the model input, and the second depth map to be optimized is obtained by inputting it into the depth prediction model; based on the second depth map to be optimized, the first image frame, and the second image frame, second error information is obtained; the depth prediction model is optimized using the second error information.
[0113] Optionally, the pose information of the second depth map to be optimized can be obtained first. Based on the difference operation result between the second depth map to be optimized and the first image frame, and combined with the pose information, a second reconstructed image is obtained. Then, the difference operation is performed between the second reconstructed image and the second image frame to obtain second error information, and the depth prediction model is optimized using the second error information.
[0114] As described above, the scale ratio relationship can be pre-calculated, such as Figure 5 shown in the flowchart of another embodiment of a detection method provided by an embodiment of the present application. This embodiment describes the technical solution of the present application from the perspective of calculating the scale ratio relationship. The technical solution of this embodiment can be executed by a mobile robot, a client, or a server. The method may include the following steps:
[0115] 501: Use a monocular camera to collect an image of a calibration object located at the calibration distance of the mobile robot to obtain a second monocular image; wherein, the calibration distance is known.
[0116] Refer to Figure 6 , which shows a schematic diagram of scale calibration in an actual application of an embodiment of the present application. A monocular camera 602 can be set in the sweeping robot 601. It should be noted that Figure 6 only the sweeping robot 601 is taken as an example of the mobile robot for illustration, and it does not mean that the mobile robot is limited to the sweeping robot 601.
[0117] The calibration object 603 is set at the calibration distance from the sweeping robot 601, and the calibration distance is denoted as Z2. Among them, the calibration object 603 can adopt an object with a specific shape, specific size, or specific color that is convenient for image recognition. For example, it can be a baffle with a rectangular shape having a certain height and width, and of course it can also be any other object that is convenient for recognition. The present application does not make specific limitations on this.
[0118] When the sweeping robot 601 collects an image, it remains stationary at a distance of Z2 from the calibration object 603, and collects a second monocular image including the calibration object through the monocular camera 602 provided thereon.
[0119] After the second monocular image is acquired, it is necessary to perform image recognition on the second monocular image. Among them, the recognition of the second monocular image can be achieved by using image recognition technology. For example, an image recognition model is used to recognize the calibration object in the second monocular image. The image recognition model is trained using sample images labeled with calibration objects during the training process and can be used to recognize calibration objects. Of course, the image recognition model is also trained using sample images labeled with target objects during the training process. The target object can be determined according to the actual application scenario of distance measurement, and it can refer to various static or dynamic objects in the working environment of the mobile robot, such as various objects like people, animals, and furniture.
[0120] Before performing image recognition on the second monocular image, some necessary preprocessing operations such as denoising, image enhancement, grayscale conversion, binarization, and / or edge detection may also be performed on the second monocular image. Specifically, target detection is performed on the preprocessed second monocular image to recognize the calibration object in the second monocular image. These preprocessing operations are the same as those in the prior art and will not be elaborated here.
[0121] After the calibration object is recognized in the second monocular image, the bounding box of the calibration object can be determined in combination with the edge detection result of the calibration object. Among them, the boundary close to the ground in the bounding box is used as the lower boundary. The lower boundary can be the lower edge of the calibration object close to the ground. In actual applications, there may be a situation where the lower edge of the calibration object is an irregular curve. Therefore, the horizontal line where the pixel point closest to the ground in the lower edge is located can be selected as the lower boundary.
[0122] 502: Use a depth prediction model to predict the second depth image corresponding to the second monocular image.
[0123] Among them, the functions and training methods of the depth prediction model can refer to the introduction in the above embodiments and will not be elaborated here.
[0124] 503: Based on the second depth image, determine the relative distance between the calibration object and the mobile robot.
[0125] In actual applications, since the calibration object has a certain size, there is more than one pixel point corresponding to the calibration object in the second depth image. The depth value corresponding to any one of the pixel points in the second depth image can be selected as the depth value of the calibration object. Any one of these pixel points can be the central pixel point of the calibration object. Of course, it can also be other pixel points of the calibration object. Of course, the depth values of all pixel points of the calibration object can also be accumulated and averaged, and the average value can be used as the depth value of the calibration object.
[0126] It can be understood that the positioning accuracy of the image position of the calibration object in the second monocular image directly affects the accuracy of determining the pixel points corresponding to the calibration object in the second depth image. Therefore, in order to improve the accuracy of determining the pixel points corresponding to the calibration object in the second depth image, the bounding box of the calibration object can be determined by combining the edge detection result of the calibration object. Optionally, the bounding box is a rectangular box. The pixel points corresponding to the image area surrounded by the bounding box in the second depth image are used as the pixel points corresponding to the calibration object in the second depth image.
[0127] In an actual scenario, for a mobile robot equipped with a monocular camera, if the distance between the object and the monocular camera is regarded as the distance between the object and the mobile robot, then the depth value of the calibration object is used as the relative distance between the calibration object and the mobile robot.
[0128] In an actual scenario, for a mobile robot equipped with a monocular camera, if there is a distance deviation between "the distance between the object and the monocular camera" and "the distance between the object and the mobile robot", then the sum of the depth value of the calibration object and the distance deviation is used as the relative distance between the calibration object and the mobile robot. For the distance deviation, reference can be made to the introduction in the above embodiments.
[0129] 504: Obtain the scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibration distance.
[0130] Among them, the scale ratio relationship is used to perform scale conversion on the distance information in the first depth image obtained by converting the first monocular image collected by the mobile robot. For the specific application of the scale ratio relationship, details can be found in Figure 1 the embodiments shown, and will not be elaborated here.
[0131] Among them, when the scale ratio relationship is obtained by the user side or the server side, the scale ratio relationship can be configured in the mobile robot.
[0132] Among them, after knowing the calibration distance Z2 and predicting the relative distance D2 between the calibration object and the mobile robot, the ratio Z2 / D2 of Z2 to D2 can be calculated, and Z2 / D2 is used as the scale ratio relationship Scale. In this embodiment, by setting a calibration object at the known calibration distance of the mobile robot, and determining the relative distance between the calibration object and the mobile robot based on the second monocular image including the calibration object and the depth prediction model, and obtaining the scale ratio relationship based on the calibration distance and the relative distance, the scale ratio relationship is beneficial to improving the accuracy of the ranging result.
[0133] The technical solution of the embodiment of the present application can be applied to different types of mobile robots with ranging requirements. The distance between the detected target object and the mobile robot can be processed accordingly in different application scenarios. The following are several possible application scenarios to introduce the technical solution of the present application:
[0134] Application Scenario 1:
[0135] In the obstacle avoidance scenario of a floor cleaning robot, the target object refers to an obstacle, such as furniture in the working environment, etc. A monocular camera is set on the floor cleaning robot. When the floor cleaning robot is in the cleaning working mode, it can use the monocular camera to collect images of its working environment to obtain a monocular video stream. The image frames in the monocular video stream are used as the first monocular images to be processed. Obstacles can be recognized from them, and the first depth image corresponding to the first monocular image predicted by the depth prediction model can be obtained; based on the first depth image, the distance between the obstacle and the floor cleaning robot can be determined. The floor cleaning robot can perform obstacle avoidance processing based on this distance, for example, by performing path planning to bypass the obstacle, etc. In addition, the height and width information of the obstacle can be calculated in combination with camera parameters, etc., for participating in obstacle avoidance processing.
[0136] Application Scenario 2:
[0137] In the scenario where a mobile robot maps an indoor space to construct an electronic map, taking the mapping scenario of an air purification robot as an example, the target object can refer to walls, partitions, furniture such as sofas, beds, tables, chairs, and electrical appliances such as refrigerators and washing machines in the working environment. Before the first full-house purification, a home map needs to be created. A monocular camera is set on the air purification robot. In the mapping mode, the air purification robot traverses the home environment along the planned path. During the process of traversing the home environment, it collects images through the monocular camera to obtain a monocular video stream. The image frames in the monocular video stream are used as the first monocular images to be processed. The target object can be recognized from them, and the first depth image corresponding to the first monocular image predicted by the depth prediction model can be obtained; based on the first depth image, the distance between the target object and the air purification robot can be determined. It should be noted that when the air purification robot collects images, it also records its own position information. In this way, after determining the distance between the target object and the air purification robot, combined with the position information of the air purification robot itself when collecting the image frame including the target object, the position information of the target object in the world coordinates can be calculated. A home map is created based on the position information of each target object in the home environment in the world coordinates.
[0138] Figure 7 FIG. shows a schematic structural diagram of an embodiment of a detection device provided by an embodiment of the present application. The device may include:
[0139] The first acquisition module 701 is configured to acquire a first monocular image collected by a monocular camera and identify a target object in the first monocular image;
[0140] The first processing module 702 is configured to predict a first depth image corresponding to the first monocular image by using a depth prediction model;
[0141] The first processing module 702 is further configured to determine the distance between the target object and the mobile robot based on the first depth image.
[0142] In some embodiments, the first processing module 702 determines the distance between the target object and the mobile robot based on the first depth image specifically as follows:
[0143] Determine the relative distance between the target object and the mobile robot based on the first depth image;
[0144] Calculate the distance between the target object and the mobile robot based on the relative distance and the scale ratio relationship.
[0145] In some embodiments, the first processing module 702 is further configured to determine whether there is a suspended part of the target object relative to the ground.
[0146] In some embodiments, the first processing module 702 is further configured to: if there is a suspended part of the target object relative to the ground, detect a first boundary of the target object that is close to the ground and does not contact the ground in the first monocular image; identify a second boundary of the first boundary mapped to the ground; use two vertex pixel points on any diagonal line in the region formed by the first boundary and the second boundary as marker points; determine the width information of the target object and the height information from the ground based on the position information of the two marker points in the first monocular image.
[0147] In some embodiments, the first processing module 702 detects the first boundary of the target object that is close to the ground and does not contact the ground in the first monocular image specifically as follows:
[0148] Perform edge detection on the target object in the first monocular image;
[0149] Determine the lower edge of the target object close to the ground according to the edge result;
[0150] Use the horizontal line where the pixel point closest to the ground in the lower edge is located as the first boundary.
[0151] In some embodiments, the first processing module 702 is further configured to perform obstacle avoidance processing on the mobile robot based on the distance, width information, and height information between the target object and the mobile robot.
[0152] In some embodiments, the first processing module 702 is further configured to: if there is no suspended part of the target object relative to the ground, detect the first boundary where the target object contacts the ground in the first monocular image; use the two endpoint pixels of the first boundary as marking points; and determine the width information of the target object based on the position information of the two marking points in the first monocular image. In some embodiments, the first processing module 702 is further configured to:
[0153] Perform obstacle avoidance processing on the mobile robot based on the distance and width information between the target object and the mobile robot.
[0154] In some embodiments, the mobile robot is a sweeping robot; the specific manner in which the first processing module 702 obtains the first monocular image collected by the monocular camera is:
[0155] When the sweeping robot is in the cleaning working mode, any image frame in the monocular video stream collected by the monocular camera is used as the first monocular image.
[0156] In some embodiments, the method for determining the scale ratio relationship is:
[0157] Use the monocular camera to collect an image of a calibration object located at the calibration distance of the mobile robot to obtain a second monocular image; wherein, the calibration distance is known;
[0158] Use a depth prediction model to predict the second depth image corresponding to the second monocular image;
[0159] Based on the second depth image, determine the relative distance between the calibration object and the mobile robot;
[0160] Based on the relative distance between the calibration object and the mobile robot and the calibration distance, obtain the scale ratio relationship.
[0161] Figure 7 The detection device can execute Figure 1 The detection method of the illustrated embodiment, and its implementation principle and technical effects will not be elaborated. For the detection device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0162] In a possible design, Figure 7 The detection device of the illustrated embodiment can be implemented as a detection device, as Figure 8 shown, the detection device may include a storage component 801 and a processing component 802;
[0163] The storage component 801 stores one or more computer instructions, and one or more computer instructions are called and executed by the processing component.
[0164] The processing component 802 is used for:
[0165] Obtain the first monocular image collected by the monocular camera and identify the target object in the first monocular image;
[0166] Use the depth prediction model to predict the first depth image corresponding to the first monocular image;
[0167] Based on the first depth image, determine the distance between the target object and the mobile robot.
[0168] Among them, the processing component 802 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0169] The storage component 801 is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0170] Of course, the detection device may also necessarily include other components, such as input / output interfaces, communication components, etc.
[0171] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, an acquisition component, etc.
[0172] The communication component is configured to facilitate communication between the detection device and other devices in a wired or wireless manner, etc.
[0173] In a possible design, Figure 7 The detection device in the illustrated embodiment may be implemented as a mobile robot, such as Figure 9 shown, and the mobile robot may include an acquisition component 903, a storage component 901, and a processing component 902;
[0174] Among them, the storage component 901 stores one or more computer instructions, and one or more computer instructions are called and executed by the processing component.
[0175] The acquisition component 903 is used for acquiring the first monocular image;
[0176] The processing component 902 is configured to:
[0177] Obtain the first monocular image collected by the acquisition component 903 and identify the target object in the first monocular image;
[0178] Predict the first depth image corresponding to the first monocular image by using a depth prediction model;
[0179] Determine the distance between the target object and the mobile robot based on the first depth image.
[0180] Wherein, the acquisition component 903 may be a monocular camera.
[0181] Wherein, the processing component 902 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0182] The storage component 901 is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0183] Of course, the mobile robot may also necessarily include other components, such as an input / output interface, a communication component, etc.
[0184] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, an acquisition component, etc.
[0185] The communication component is configured to facilitate wired or wireless communication between the mobile robot and other devices, etc.
[0186] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the detection method of the above Figure 1 shown embodiment.
[0187] Figure 10The figure shows a schematic structural diagram of another embodiment of a detection device provided by an embodiment of the present application. The device may include:
[0188] A second acquisition module 1001, configured to collect an image of a calibration object located at a calibration distance from a mobile robot by using a monocular camera, so as to obtain a second monocular image; wherein, the calibration distance is known;
[0189] A second processing module 1002, configured to predict a second depth image corresponding to the second monocular image by using a depth prediction model;
[0190] The second processing module 1002 is further configured to determine a relative distance between the calibration object and the mobile robot based on the second depth image;
[0191] The second processing module 1002 is further configured to obtain a scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibration distance.
[0192] Figure 10 The detection device can execute Figure 5 the detection method of the embodiment shown, and its implementation principle and technical effects will not be elaborated. For the detection device in the above embodiment, the specific manners in which each module and unit perform operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0193] In a possible design, Figure 10 the detection device of the embodiment shown can be implemented as a detection device, and the structure of this detection device is Figure 8 the same as the structure of the detection device shown.
[0194] The detection device includes a storage component and a processing component;
[0195] The storage component stores one or more computer instructions, and one or more computer instructions are called by the processing component;
[0196] Wherein, the processing component is configured to:
[0197] Collect an image of a calibration object located at a calibration distance from a mobile robot by using a monocular camera, so as to obtain a second monocular image; wherein, the calibration distance is known;
[0198] Predict a second depth image corresponding to the second monocular image by using a depth prediction model;
[0199] Determine a relative distance between the calibration object and the mobile robot based on the second depth image;
[0200] Obtain a scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibration distance.
[0201] Among them, the processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0202] The storage component is configured to store various types of data to support the operation of the terminal. The storage component may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0203] Of course, the detection device may also necessarily include other components, such as input / output interfaces, communication components, etc.
[0204] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, an acquisition component, etc.
[0205] The communication component is configured to facilitate communication between the detection device and other devices in a wired or wireless manner, etc.
[0206] In a possible design, Figure 10 The detection device of the illustrated embodiment may be implemented as a mobile robot, and the structure of this mobile robot is Figure 9 the same as the structure of the illustrated mobile robot.
[0207] Specifically, the mobile robot may include an acquisition component, a storage component, and a processing component;
[0208] Among them, the storage component stores one or more computer instructions, and one or more computer instructions are called and executed by the processing component.
[0209] The acquisition component is used to acquire an image of a calibration object at a calibrated distance from the mobile robot to obtain a second monocular image; wherein, the calibrated distance is known;
[0210] The processing component is used for:
[0211] Obtain the second monocular image acquired by the acquisition component; wherein, the calibrated distance is known;
[0212] Use the depth prediction model to predict the second depth image corresponding to the second monocular image;
[0213] Based on the second depth image, determine the relative distance between the calibration object and the mobile robot;
[0214] Based on the relative distance between the calibration object and the mobile robot and the calibration distance, obtain the scale ratio relationship.
[0215] Among them, the acquisition component can be a monocular camera.
[0216] The processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.
[0217] The storage component is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0218] Of course, the mobile robot may also necessarily include other components, such as input / output interfaces, communication components, etc.
[0219] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, an acquisition component, etc.
[0220] The communication component is configured to facilitate wired or wireless communication between the mobile robot and other devices, etc.
[0221] This application embodiment also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the detection method of the above Figure 5 shown embodiment.
[0222] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0223] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0224] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0225] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A detection method, characterized in that, Including: Obtain a first monocular image collected by a monocular camera, and identify a target object in the first monocular image; Use a depth prediction model to predict a first depth image corresponding to the first monocular image; Based on the first depth image, determine the distance between the target object and the mobile robot; Judge whether there is a suspended part of the target object relative to the ground; If there is a suspended part of the target object relative to the ground, detect a first boundary of the target object that is close to the ground and does not contact the ground in the first monocular image; identify a second boundary of the first boundary mapped to the ground; use two vertex pixel points on any diagonal line in the area formed by the first boundary and the second boundary as marker points; based on the position information of the two marker points in the first monocular image, determine the width information of the target object and the height information from the ground.
2. The method according to claim 1, characterized in that, The determining the distance between the target object and the mobile robot based on the first depth image includes: Determine the relative distance between the target object and the mobile robot based on the first depth image; Based on the relative distance and the scale ratio relationship, calculate the distance between the target object and the mobile robot.
3. The method according to claim 1, characterized in that, Detecting the first boundary of the target object that is close to the ground and does not contact the ground in the first monocular image includes: Perform edge detection on the target object in the first monocular image; Determine the lower edge of the target object close to the ground according to the edge result; Use the horizontal line where the pixel point closest to the ground in the lower edge is located as the first boundary.
4. The method according to claim 1, characterized in that, Also including: Based on the distance between the target object and the mobile robot, the width information and the height information, perform obstacle avoidance processing on the mobile robot.
5. The method according to claim 1, characterized in that, Also including: If there is no suspended part of the target object relative to the ground, detect a first boundary where the target object contacts the ground in the first monocular image; Use the two endpoint pixel points of the first boundary as marker points; Based on the position information of the two marker points in the first monocular image, determine the width information of the target object.
6. The method according to claim 5, characterized in that, Also including: Based on the distance between the target object and the mobile robot and the width information, perform obstacle avoidance processing on the mobile robot.
7. The method according to claim 1, characterized in that, The mobile robot is a sweeping robot; The obtaining the first monocular image collected by the monocular camera includes: When the sweeping robot is in the cleaning working mode, use any image frame in the monocular video stream collected by the monocular camera as the first monocular image.
8. The method according to claim 2, characterized in that, The determination method of the scale ratio relationship is: Use the monocular camera to collect an image of a calibration object at the calibration distance of the mobile robot to obtain a second monocular image; wherein, the calibration distance is known; Use the depth prediction model to predict a second depth image corresponding to the second monocular image; Based on the second depth image, determine the relative distance between the calibration object and the mobile robot; Based on the relative distance between the calibration object and the mobile robot and the calibration distance, obtain the scale ratio relationship.
9. A detection method, characterized in that, Including: An image of a calibration object located at a calibrated distance from a mobile robot is captured using a monocular camera to obtain a second monocular image; wherein, the calibrated distance is known; A depth prediction model is used to predict a second depth image corresponding to the second monocular image; Based on the second depth image, the relative distance between the calibration object and the mobile robot is determined; Based on the relative distance between the calibration object and the mobile robot and the calibrated distance, a scale ratio relationship is obtained; It is determined whether the calibration object has a suspended part relative to the ground; if the calibration object has a suspended part relative to the ground, a first boundary of the calibration object that is close to the ground and does not contact the ground in the second monocular image is detected; a second boundary mapped from the first boundary to the ground is identified; two vertex pixel points on any diagonal of the region formed by the first boundary and the second boundary are used as marker points; based on the position information of the two marker points in the second monocular image, the width information of the calibration object and the height information from the ground are determined.
10. A detection device, characterized in that, Including: A first acquisition module, configured to acquire a first monocular image captured by a monocular camera and identify a target object in the first monocular image; A first processing module, configured to use a depth prediction model to predict a first depth image corresponding to the first monocular image; The first processing module is further configured to determine the distance between the target object and the mobile robot based on the first depth image; determine whether the target object has a suspended part relative to the ground; If the target object has a suspended part relative to the ground, a first boundary of the target object that is close to the ground and does not contact the ground in the first monocular image is detected; a second boundary mapped from the first boundary to the ground is identified; two vertex pixel points on any diagonal of the region formed by the first boundary and the second boundary are used as marker points; based on the position information of the two marker points in the first monocular image, the width information of the target object and the height information from the ground are determined.
11. A detection device, characterized in that, Including: A second acquisition module, configured to use a monocular camera to capture an image of a calibration object located at a calibrated distance from a mobile robot to obtain a second monocular image; wherein, the calibrated distance is known; A second processing module, configured to use a depth prediction model to predict a second depth image corresponding to the second monocular image; The second processing module is further configured to determine the relative distance between the calibration object and the mobile robot based on the second depth image; The second processing module is further configured to obtain a scale ratio relationship based on the relative distance between the calibration object and the mobile robot and the calibrated distance; The second processing module is further configured to determine whether there is a suspended part of the calibration object relative to the ground; if there is a suspended part of the calibration object relative to the ground, detect a first boundary of the calibration object in the second monocular image that is close to the ground and does not contact the ground; identify a second boundary of the first boundary mapped to the ground; use two vertex pixel points on any diagonal line in the region formed by the first boundary and the second boundary as marking points; and determine the width information of the calibration object and the height information from the ground based on the position information of the two marking points in the second monocular image.
12. A mobile robot, characterized in that, It includes a collection component, a storage component, and a processing component; wherein, the storage component stores one or more computer instructions; the one or more computer instructions are called and executed by the processing component; The collection component is used to collect a first monocular image; the processing component is used to: Obtain the first monocular image collected by the collection component and identify the target object in the first monocular image; Use a depth prediction model to predict a first depth image corresponding to the first monocular image; Based on the first depth image, determine the distance between the target object and the mobile robot; Determine whether there is a suspended part of the target object relative to the ground; if there is a suspended part of the target object relative to the ground, detect a first boundary of the target object in the first monocular image that is close to the ground and does not contact the ground; identify a second boundary of the first boundary mapped to the ground; use two vertex pixel points on any diagonal line in the region formed by the first boundary and the second boundary as marking points; and determine the width information of the target object and the height information from the ground based on the position information of the two marking points in the first monocular image.
13. A mobile robot, characterized in that, It includes a collection component, a storage component, and a processing component; wherein, the storage component stores one or more computer instructions; the one or more computer instructions are called and executed by the processing component; The collection component is used to collect an image of a calibration object at a calibration distance of the mobile robot to obtain a second monocular image; wherein, the calibration distance is known; The processing component is used to: Obtain the second monocular image collected by the collection component; wherein, the calibration distance is known; Use a depth prediction model to predict a second depth image corresponding to the second monocular image; Based on the second depth image, determine the relative distance between the calibration object and the mobile robot; Based on the relative distance between the calibration object and the mobile robot and the calibration distance, obtain a scale ratio relationship; Determine whether there is a suspended part of the calibration object relative to the ground; if there is a suspended part of the calibration object relative to the ground, detect a first boundary of the calibration object in the second monocular image that is close to the ground and does not contact the ground; identify a second boundary of the first boundary mapped to the ground; use two vertex pixel points on any diagonal line in the region formed by the first boundary and the second boundary as marking points; and determine the width information of the calibration object and the height information from the ground based on the position information of the two marking points in the second monocular image.
14. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a computer, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Floor sweeping robot, floor sweeping robot system and working method thereof
CN109213137A
Inspection method, device and equipment, and computer readable storage medium
CN109664301A
Three-dimensional scene fusion method and device based on monocular estimation
CN111340864A