Range information determination method and apparatus, and readable storage medium and robot

By combining a monocular camera and lidar, the robot can accurately detect the distance to the human body at a low cost, improve the detection frame rate and real-time performance, solve the problems of high cost and insufficient real-time performance in existing technologies, and improve the robot's following effect.

WO2025218425A1PCT designated stage Publication Date: 2025-10-23MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/083302
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2025-03-19
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

In the existing technology, the robot's estimation of the distance to the human body is costly and lacks real-time performance, especially when the frame rate is insufficient during movement, resulting in poor following effect.

Method used

By combining a monocular camera and a lidar, location information is determined through image detection, and the lidar is controlled to perform location detection within a small area, thus narrowing the detection range and improving real-time performance.

Benefits of technology

It achieves low-cost and high-real-time distance detection, and improves the accuracy of the robot's distance to the human body and the following effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025083302_23102025_PF_FP_ABST
    Figure CN2025083302_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A range information determination method, comprising: performing image detection on a first image frame, so as to obtain detection information corresponding to a target object, wherein the first image frame is collected by an image collection apparatus, and image content of the first image frame comprises the target object; determining position information on the basis of the detection information; determining a target detection region by using a first position point indicated by the position information as the center and using a preset range as the radius; and controlling a ranging apparatus to detect, in the target detection area, range information between the target object and a robot. Further provided are a range information determination apparatus, a readable storage medium, and a robot.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for determining distance information, readable storage medium and robot

[0001] The present application claims priority from the Chinese patent application No. 202410480955.5 filed on April 19, 2024, and entitled "Method and device for determining distance information, readable storage medium and robot", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of robot control, in particular to a method and device for determining distance information, a readable storage medium and a robot. BACKGROUND

[0003] In the related art, robot devices such as sweeping robots, service robots or welcome robots can realize real-time communication with users and follow the users. At this time, the robot needs to rely on the estimation result of the distance between the robot and the human body to follow the human body.

[0004] In order to realize accurate distance estimation, it is necessary to rely on a binocular vision system, that is, two or more vision sensors are installed on the robot, so the cost is high, and when the robot moves, the frame rate of the binocular vision sensor is insufficient, resulting in poor real-time performance of distance estimation. SUMMARY

[0005] The present application aims to at least solve one of the problems in the prior art or related art.

[0006] To this end, a first aspect of the present application provides a method for determining distance information.

[0007] A second aspect of the present application provides a device for determining distance information.

[0008] A third aspect of the present application provides a device for determining distance information.

[0009] A fourth aspect of the present application provides a readable storage medium.

[0010] A fifth aspect of the present application provides a robot.

[0011] Therefore, the first aspect of the present application provides a distance information determination method, which is executed by a robot, the robot comprising an image acquisition device and a distance detection device, and the distance information determination method comprising: performing image detection on a first image frame to obtain detection information corresponding to a target object; wherein the first image frame is acquired by the image acquisition device, and the image content of the first image frame comprises the target object; determining position information according to the detection information; determining a target detection region with a first position point indicated by the position information as the center and a preset distance as the radius; and controlling the distance detection device to detect distance information between the target object and the robot within the target detection region.

[0012] The technical scheme of the present application can realize accurate detection of the distance of the user by combining the monocular camera and the laser radar, and the requirement for computing power when performing image detection on the image frame acquired by the monocular camera is smaller, which can improve the detection frame rate, control the laser radar to perform position detection in a small range based on the image detection result, reduce the detection range of the laser radar, and thus significantly shorten the time consumption of position detection and improve the real-time performance of position detection.

[0013] The second aspect of the present application provides a distance information determination device, which is applied to a robot, the robot comprising an image acquisition device and a distance detection device, and the distance information determination device comprising:

[0014] a detection module, configured to perform image detection on a first image frame to obtain detection information corresponding to a target object; wherein the first image frame is acquired by the image acquisition device, and the image content of the first image frame comprises the target object; a determination module, configured to determine position information according to the detection information; and determine a target detection region with a first position point indicated by the position information as the center and a preset distance as the radius; and a control module, configured to control the distance detection device to detect distance information between the target object and the robot within the target detection region.

[0015] The technical scheme of the present application can realize accurate detection of the distance of the user by combining the monocular camera and the laser radar, and the requirement for computing power when performing image detection on the image frame acquired by the monocular camera is smaller, which can improve the detection frame rate, control the laser radar to perform position detection in a small range based on the image detection result, reduce the detection range of the laser radar, and thus significantly shorten the time consumption of position detection and improve the real-time performance of position detection.

[0016] The third aspect of the present application provides a distance information determination device, which comprises: a memory, configured to store programs or instructions; and a processor, configured to execute the programs or instructions to realize the steps of the distance information determination method provided in any of the above technical schemes, and thus has all the beneficial effects of the distance information determination method provided in any of the above technical schemes. To avoid repetition, the details are not described here.

[0017] The fourth aspect of the present application provides a readable storage medium, which stores programs or instructions, and the programs or instructions are executed by a processor to realize the steps of the distance information determination method provided in any of the above technical solutions, and thus have all the beneficial effects of the distance information determination method provided in any of the above technical solutions. To avoid repetition, details are not described here.

[0018] The fifth aspect of the present application provides a robot, which comprises the distance information determination device provided in any of the above technical solutions; and / or the readable storage medium provided in any of the above technical solutions, and thus has all the beneficial effects of the distance information determination device provided in any of the above technical solutions and / or the readable storage medium provided in any of the above technical solutions. To avoid repetition, details are not described here. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which:

[0020] FIG. 1 shows a structural schematic diagram of a robot of some embodiments of the present application;

[0021] FIG. 2 shows a flowchart of a distance information determination method of some embodiments of the present application;

[0022] FIG. 3 shows a schematic diagram of the positional relationship between a robot and a target object of some embodiments of the present application;

[0023] FIG. 4 shows a logic diagram of robot distance estimation of some embodiments of the present application;

[0024] FIG. 5 shows a structural block diagram of a distance information determination device of some embodiments of the present application;

[0025] FIG. 6 shows a structural block diagram of a distance information determination device of some embodiments of the present application.

[0026] Reference signs: 100 robot, 102 image acquisition device, 104 distance detection device. DETAILED DESCRIPTION

[0027] In order to enable a clearer understanding of the above-mentioned purposes, features and advantages of the present application, the present application is further described in detail below in conjunction with the drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0028] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the application, but the application can also be implemented in other ways different from those described herein, and therefore the scope of protection of the application is not limited by the specific embodiments disclosed below.

[0029] The method and device for determining distance information, the readable storage medium and the robot according to some embodiments of the application are described below with reference to FIGS. 1-6.

[0030] In some embodiments of the application, a method for determining distance information is provided, and the method is performed by a robot. FIG. 1 shows a structural schematic diagram of a robot according to some embodiments of the application. As shown in FIG. 1, the robot 100 includes an image acquisition device 102 and a distance detection device 104.

[0031] FIG. 2 shows a flowchart of a method for determining distance information according to some embodiments of the application. As shown in FIG. 2, the method includes:

[0032] In step 202, image detection is performed on the first image frame to obtain detection information corresponding to the target object; wherein the first image frame is acquired by the image acquisition device, and the image content of the first image frame includes the target object.

[0033] In step 204, the position information is determined according to the detection information.

[0034] In step 206, the target detection area is determined with the first position point indicated by the position information as the center and with a preset distance as the radius.

[0035] In step 208, the distance detection device is controlled to detect the distance information between the target object and the robot within the target detection area.

[0036] In this embodiment, the robot can be a sweeping robot, a service robot or a welcome robot, and the robot can respond to the interactive operation of the user and realize following the user. The interactive operation of the user can be a voice interactive operation, a remote control interactive operation or a direct control operation on the human-computer interaction interface provided on the robot.

[0037] The image acquisition device and the distance detection device are provided on the robot. Exemplarily, the image acquisition device can be a camera, such as a monocular camera, and the distance detection device can be a radar device, such as a laser radar. The image acquisition device can acquire the environment image near the robot, and when there is a user nearby, the image acquisition device can realize tracking and shooting of the user, thereby ensuring that the user is always within the image field of view of the robot. The distance detection device can be installed on the chassis of the robot, thereby detecting the position (or the distance from the robot) of a specific object.

[0038] The target object can be a user, or other mobile objects such as pets, other robots, etc., and can also be a static object such as a household appliance, furniture, wall, etc. The robot can track, follow or avoid the target object by estimating the distance information corresponding to the target object. In the case of the robot following the target object, whether the distance information estimated by the robot is accurate or not will directly affect the following effect.

[0039] In the related art, in order to achieve accurate distance measurement, it is generally necessary to use methods such as combining binocular cameras or using TOF (Time Of Flight) sensors to obtain relevant information, which has a relatively high cost, and in the case of movement of the robot and the target object, the frame rate of these detection methods is often not enough, resulting in that the distance estimation result is not timely and the robot following effect is poor.

[0040] To solve the above problems, embodiments of the present application combine image recognition and chassis radar detection to improve the real-time and accuracy of robot distance estimation while keeping the cost low. Specifically, the first image frame is an image frame acquired by the robot through an image acquisition device at a current time point, and the first image frame includes an image of a target object followed by the robot.

[0041] For example, the target object can be the entire human body of a user, or a part of the human body of a user, such as a head, a face, etc.

[0042] After the first image frame is acquired, the robot performs image detection on the first image frame, thereby determining detection information of the target object according to the result of the image detection, and determining position information corresponding to the target object through the detection information. The position information is an estimated position obtained based on image recognition.

[0043] After obtaining the position information, since the position information is an estimated position, in order to obtain the accurate position of the target object, the robot further controls the distance detection device, such as a chassis radar, to emit a detection signal to a certain range corresponding to the position information, thereby obtaining the accurate position of the target object and obtaining accurate distance information between the target object and the robot.

[0044] Specifically, FIG. 3 shows a schematic diagram of the position relationship between the robot and the target object according to some embodiments of the present application. As shown in FIG. 3, during the movement of the robot, the robot always controls and adjusts the direction of the robot so that the target object followed is always within the image field of view of the robot, and then the robot controls the distance detection device to search in a target detection region with the position information (RegX, RegZ) estimated by vision as the center and a preset distance r as the radius, thereby acquiring the detection coordinates of the target object, and determining the accurate distance information between the target object and the robot according to the detected detection coordinates of the target object.

[0045] The embodiment of the present application can realize accurate detection of the distance of the user by combining the monocular camera and the laser radar, the calculation power requirement for image detection of the image frame collected by the monocular camera is smaller, the frame rate of detection can be improved, the laser radar is controlled to detect the position in a small range based on the image detection result, the detection range of the laser radar can be reduced, the time consumption of position detection can be significantly shortened, and the real-time performance of position detection can be improved.

[0046] In some embodiments of the present application, the step of determining the position information according to the detection information specifically includes: determining a target rectification regression model according to the detection information; and applying the target rectification regression model to determine the position information according to the detection information and the image information of the first image frame.

[0047] In the embodiment of the present application, the detection information of the target object can reflect whether the target object is occluded in the first image frame, and the specific occlusion degree of the target object in the case of being occluded in the first image frame. For different occlusion degrees of the target object, a suitable rectification regression model is selected, which is beneficial to reduce the influence of image occlusion on position estimation.

[0048] Specifically, when the detection information reflects that the target object is not occluded or the occlusion of the target object is relatively slight, a common rectification regression model can be selected for position estimation, thereby reducing the calculation amount. When the detection information reflects that the occlusion of the target object is relatively serious, a rectification regression model for the occlusion can be selected, thereby reducing the influence of the occlusion and improving the accuracy of position estimation.

[0049] After determining the suitable target rectification regression model, the target rectification regression model is applied to estimate the position of the target object in combination with the detection information, and the position information corresponding to the target object is obtained.

[0050] The embodiment of the present application determines whether the target object is occluded through the detection information, selects a corresponding target rectification regression model according to the judgment result, can reduce the error of distance estimation when the target is occluded, and improve the accuracy of distance estimation.

[0051] In some embodiments of the present application, the detection information includes a first detection width and a first detection height of the target object; and the step of determining the target rectification regression model according to the detection information specifically includes:

[0052] A ratio of the first detection width and the first detection height is determined to obtain a detection width-height ratio. A width-height ratio fluctuation value is determined according to a difference between the detection width-height ratio and the average width-height ratio. A target correction regression model is determined based on a comparison result of the width-height ratio fluctuation value and a first threshold value. Position information is determined according to the first detection width, the first detection height, and the target correction regression model.

[0053] In the embodiments of the present application, the detection information specifically includes width detection information and height detection information. Specifically, the width detection information is the first detection width corresponding to the target object, and the height detection information is the first detection height corresponding to the target object.

[0054] For example, taking a user's head as the target object, when detecting the user's head, the head detection is performed on the input video image, i.e., the first image frame. In some embodiments, the head detection algorithm determines the head position through the detection frame. At this time, the first detection width is specifically the edge frame width of the head detection frame, and the first detection height is specifically the edge frame height of the head detection frame.

[0055] For example, taking a user's body as the target object, at this time, the user's body includes the head, the torso, and the limbs. When detecting the user's body, image recognition is performed on the input video image to detect the user's head, torso, or limbs, and a body detection frame is determined according to the detected body part. At this time, the first detection width is specifically the edge frame width of the body detection frame, and the first detection height is specifically the edge frame height of the body detection frame.

[0056] Let the first detection width be head_w, and let the first detection height be head_h. Then, the detection width-height ratio head_r = head_w / head_h is calculated. Further, the width-height ratio fluctuation value is determined. For example, let the width-height ratio fluctuation value be Δhead_r. Then, Δhead_r = |head_r-aHead_r|, where aHead_r is the pre-determined average width-height ratio.

[0057] Specifically, since the target object and the robot can both be in a motion state, as the positional relationship between the two changes, the target object can be blocked in the image field of view of the robot. Through the width-height ratio fluctuation value, it can be reflected whether the target object is blocked and the degree of blocking when the blocking exists.

[0058] Therefore, according to the comparison result of the width-height ratio fluctuation value and the pre-set first threshold value, it is determined whether there is blocking. According to the comparison result, the corresponding target correction regression model is determined. The target correction regression model is used to segmentally smooth the distance estimation, thereby reducing the influence of the blocking.

[0059] Exemplarily, as shown in FIG. 3, the first threshold value is ahead r thr, the target correction regression model includes a first correction regression model, specifically RegZ = RegZ1, RegX = RegX1, and a second correction regression model, specifically RegZ = RegZ2, RegX = RegX2.

[0060] Specifically, the following conditions are met:

[0061] RegZ1 = k1 x image w ÷ head w + b1;

[0062] RegX1 = k3 x (human head w ÷ head w) x (head center w offset);

[0063] RegZ2 = k2 x image h ÷ head h + b2;

[0064] RegX2 = k4 x (human head h ÷ head h) x (head center h offset);

[0065] wherein RegZ and RegX are position information, RegZ1 and RegX1 are the first correction regression model, RegZ2 and RegX2 are the second correction regression model, image w is the image width of the first image frame, image h is the image height of the first image frame, head w is the first detection width, head h is the first detection height, k1, k2, k3 and k4 are proportional coefficients of distance estimation, b1 and b2 are offset coefficients of distance estimation, human head w is the average width of the actual human head, human head h is the average height of the actual human head, head center w offset is the pixel offset value of the center point of the human head detection frame in the image in the width direction offset from the image center, and head center h offset is the pixel offset value of the center point of the human head detection frame in the image in the height direction offset from the image center.

[0066] Exemplarily, the proportional coefficients and offset coefficients can be obtained by a linear regression method or solved by a least square method.

[0067] If the comparison result of the aspect ratio fluctuation value and the first threshold value is Δhead_r≥ahead_r_thr, the first correction regression model is used for estimation, i.e., RegZ=RegZ1 and RegX=RegX1 are used for distance estimation, and if the comparison result of the aspect ratio fluctuation value and the first threshold value is Δhead_r<ahead_r_thr, the second correction regression model is used for distance estimation, i.e., RegZ=RegZ2 and RegX=RegX2 are used for distance estimation. In this way, the problem of distance estimation of the target due to left-right or front-back occlusion can be effectively reduced.

[0068] The embodiments of the present application can determine whether the target object is occluded by comparing the detection aspect ratio with the threshold value, and determine the corresponding target correction regression model according to the comparison result, so as to reduce the distance estimation error when the target is occluded and improve the accuracy of distance estimation.

[0069] In some embodiments of the present application, optionally, before the step of determining the position information according to the detection information, the determination method further comprises:

[0070] The average detection width and the average detection height of the target object are determined, wherein the average detection width is determined according to the first detection width of the target object in the continuous multiple frames of the second image frames, the average detection height is determined according to the first detection height of the target object in the continuous multiple frames of the second image frames, the second image frames are obtained by the image acquisition device, and the average aspect ratio is determined according to the average detection width and the average detection height.

[0071] In the embodiments of the present application, for the scene of the robot following the user, the image acquisition device of the robot will continuously acquire the image data of the followed user, i.e., the target object, so as to obtain continuous multiple image frames, i.e., the above-mentioned second image frames, wherein the image content of the multiple frames of the second image frames all includes the target object.

[0072] The first detection width and the first detection height of the target object in each frame of the second image frames are detected respectively, and the average values thereof are calculated to obtain the average detection width of the target object and the average detection height of the target object.

[0073] Exemplarily, assuming that there are N frames of the second image frames, there are N first detection widths corresponding to the N frames of the second image frames, which are respectively head_w1, head_w2, …, head_wN, and N first detection heights, which are respectively head_h1, head_h2, …, head_hN, the average detection width ahead_w=(head_w1+head_w2+…+head_wN) / N and the average detection height aHead_h=(head_h1+head_h2+…+head_hN) / N can be calculated.

[0074] After the average detection width ahead w and the average detection height ahead h are obtained, the aspect ratio average value aHead r = ahead w / ahead h is calculated according to the ratio of the two.

[0075] The embodiments of the present application can accurately identify whether the target object is in a state of being occluded or not by calculating the aspect ratio average value of the target object according to the continuous multiple second image frames and using the aspect ratio average value as a judgment standard for whether the target object is occluded or not, and can reduce the influence on distance estimation when the target object is occluded by selecting a corresponding correction regression model according to the identification result.

[0076] In some embodiments of the present application, optionally, before the step of determining the ratio of the first detection width and the first detection height to obtain the detection aspect ratio, the determination method further comprises:

[0077] determining a width change value according to the difference between the first detection width and the second detection width, and determining a height change value according to the difference between the first detection height and the second detection height; wherein the second detection width and the second detection height are obtained by image detection on a third image frame, and the third image frame is a previous image frame of the first image frame;

[0078] In a case where at least one of the width change value is greater than the second threshold value and the height change value is greater than the third threshold value is met, the position information is determined according to the second detection width and the second detection height; or in a case where the width change value is less than or equal to the second threshold value and the height change value is less than or equal to the third threshold value, the step of determining the ratio of the first detection width and the first detection height to obtain the detection aspect ratio is performed.

[0079] In some scenarios, the target object followed by the robot may be seriously occluded in the field of view of the robot, such as when the target object followed by the robot enters or exits a door, or turns a corner, the door frame and the corner of the wall and other objects may cause a large occlusion to the target object, which can be in the width direction or in the height direction, and is specifically reflected in that the detection height or the detection width of the target object in the image frame of the target object obtained by the robot through the image acquisition device changes greatly. Therefore, when a relatively serious occlusion occurs, it is considered that the current image frame is not suitable for accurate distance detection.

[0080] Specifically, whether the occlusion occurs is determined according to the detection result of the current first image frame and the change value of the detection result of the previous third image frame. Assuming that the first detection width and the first detection height are detected in the current first image frame, and the second detection width and the second detection height are detected in the third image frame before the first image frame, then the difference between the first detection width and the second detection width is calculated to obtain the width change value, and the difference between the first detection height and the second detection height is calculated to obtain the height change value.

[0081] When the width change value is greater than the second threshold value, or the height change value is greater than the third threshold value, or both are satisfied, it is determined that the target object is seriously occluded, and the current first image frame is not suitable as the basis for distance detection and position detection, so the first image frame is discarded at this time, and the distance estimation of the last frame is kept unchanged, that is, the position information of the target object is determined by the second detection width and the second detection height of the third image frame.

[0082] If the width change value is less than or equal to the second threshold value, and at the same time the height change value is less than or equal to the third threshold value, it can be determined that the target object is not seriously occluded, and the current first image frame can be used as the basis for distance detection and position detection, and then the step of determining the position information according to the first detection height and the first detection width corresponding to the first image frame is normally executed.

[0083] The embodiments of the present application determine whether the target object is seriously occluded by the width change value and the height change value detected by two consecutive image frames, thereby discarding the image frame with serious occlusion, which can effectively reduce the influence of image occlusion on position detection and improve the accuracy of position detection and distance estimation.

[0084] In some embodiments of the present application, optionally, the second threshold value is determined according to the product of the second detection width and a proportionality coefficient, and the third threshold value is determined according to the product of the second detection height and the proportionality coefficient, and the proportionality coefficient has a value range of greater than or equal to 0.3 and less than or equal to 0.8.

[0085] In the embodiments of the present application, whether the target object is seriously occluded in the image field of view of the robot is determined according to the comparison result of the width change value, the height change value and the corresponding threshold value. When the width change value is greater than the second threshold value, it indicates that the target object is seriously occluded in the width direction, and the second threshold value is determined according to the detection width of the previous image frame of the current image frame. Taking the current first image frame and the previous third image frame as an example, the first detection width corresponding to the first image frame is head_w1, and the second detection width corresponding to the third image frame is head_w2, and the corresponding second threshold value is the product of the proportionality coefficient k and the second detection width, that is, the second threshold value is k×head_w2.

[0086] Similarly, the first detection height corresponding to the first image frame is head_h1, and the second detection height corresponding to the third image frame is head_h2, and the corresponding second threshold is the product of the proportional coefficient k and the second detection height, that is, the second threshold is k x head_h2.

[0087] Exemplarily, the proportional coefficient ranges from 0.3 to 0.8.

[0088] Exemplarily, the proportional coefficient is 0.5.

[0089] The application determines the comparison threshold for judging whether the target object in the current image frame is occluded based on the detection width and the detection height of the previous image, which can accurately identify whether the target object is occluded, thereby improving the accuracy of position detection and distance estimation.

[0090] In some embodiments of the application, optionally, the position information includes a first coordinate and a second coordinate, the target rectification regression model includes a horizontal rectification regression model and a vertical rectification regression model, the horizontal rectification regression model is used to determine the first coordinate, and the vertical rectification regression model is used to determine the second coordinate; after the step of determining the position information according to the detection information, the determination method further includes:

[0091] obtaining distortion information of the image acquisition device; in the case of horizontal distortion of the distortion information, adjusting the horizontal rectification regression model based on the second coordinate; or, in the case of vertical distortion of the distortion information, adjusting the vertical rectification regression model based on the first coordinate.

[0092] In this embodiment, the position information, i.e., the position information of the target object determined by image detection, is exemplarily in the format of (RegX, RegZ), where RegX is the first coordinate described above, and RegZ is the second coordinate described above. Specifically, the robot's own coordinate system is an (x, y, z) coordinate system. Since the robot and the following target are generally at the same horizontal height when the robot performs the following task, the height coordinate y can be ignored when determining the position of the target object, and only the x-axis coordinate and the z-axis coordinate in the horizontal coordinate system, i.e., the first coordinate and the second coordinate, are retained.

[0093] The image acquisition device of the robot is generally a camera, which generally includes an image sensor and a lens assembly. Due to the hardware limitations of the lens assembly, in certain cases, the image frames captured by the camera of the robot may have certain edge distortion. This edge distortion is generally divided into horizontal distortion in the horizontal direction and vertical distortion in the vertical direction.

[0094] When the camera has distortion, the image content in the image frame collected by the robot can be deformed, which can cause deviation between the position information obtained by image detection to determine the position of the target object and the actual position of the target object, resulting in inaccurate position detection results, and ultimately affecting the distance estimation of the robot.

[0095] To solve the above problems, different correction regression models are used for different lens distortions in the embodiments of the present application. Specifically, the target correction regression model specifically includes a horizontal correction regression model and a vertical correction regression model, which are respectively used to correct the first coordinate and the second coordinate.

[0096] For example, the target correction regression model includes RegZ=RegZ1 and RegX=RegX1, where RegX=RegX1 is the horizontal correction regression model described above, used to correct the first coordinate RegX, and RegZ=RegZ1 is the vertical correction regression model described above, used to correct the second coordinate RegZ.

[0097] When the lens of the image acquisition device of the robot has distortion in the vertical direction, the vertical correction regression model is adjusted according to the first coordinate in the horizontal direction, i.e. RegX. For example, the adjustment method is: RegZ'=RegZ+alpha×|RegX|, where RegZ' is the adjusted vertical correction regression model, RegZ is the vertical correction regression model before adjustment, alpha is the correction coefficient, and alpha can be learned by linear regression, and RexX is the first coordinate.

[0098] Similarly, when the lens of the image acquisition device of the robot has distortion in the horizontal direction, the horizontal correction regression model is adjusted according to the second coordinate in the vertical direction, i.e. RegZ. For example, the adjustment method is: RegX'=RegX+alpha×|RegZ|, where RegX' is the adjusted horizontal correction regression model, RegX is the horizontal correction regression model before adjustment, alpha is the correction coefficient, and alpha can be learned by linear regression, and RexZ is the second coordinate.

[0099] It can be understood that the distortion information of the image acquisition device can be determined by the hardware parameters of the image acquisition device. When producing the robot, the distortion information of the image acquisition device is determined according to the selected hardware model of the image acquisition device, or by experimentally measuring the distortion information of the image acquisition device, so as to determine whether the image acquisition device has lens distortion, and whether the lens distortion is in the horizontal direction or in the vertical direction.

[0100] The embodiment of the present application adjusts the correction regression model in the direction with distortion according to the distortion information of the image acquisition device, thereby effectively reducing the influence of lens distortion on position detection, and improving the accuracy of position detection and distance estimation.

[0101] In some embodiments of the present application, optionally, the step of controlling the distance detection device to detect the distance information between the target object and the robot in the target detection area comprises:

[0102] The distance detection device detects the target object in the target detection area to obtain the detection coordinates of the target object; and the distance information is determined based on the detection coordinates.

[0103] In this embodiment, during the movement of the robot, the robot always adjusts the direction by control, so that the followed target object is always in the image field of view of the robot, and then the distance detection device controlled by the robot searches in the target detection area centered on the position information (RegX, RegZ) estimated by vision, as shown in FIG. 3, with a preset distance r as the radius, so as to collect the detection coordinates of the target object, denoted as (lidarX, lidarZ). The detection coordinates are obtained by the position detection device of the robot, such as the chassis radar, and therefore have high accuracy.

[0104] After obtaining the detection coordinates, the robot can calculate the distance information between the followed target object and the robot itself in combination with the detection coordinates and the coordinates of the robot itself.

[0105] The embodiment of the present application estimates the position of the target object by combining the vision of the robot and the radar sensor, and controls the radar sensor to perform position detection in a targeted manner according to the position estimation result, so as to effectively reduce the detection range of the radar sensor, thereby improving the detection efficiency of position detection and improving the real-time performance of position detection.

[0106] In some embodiments of the present application, optionally, the step of controlling the distance detection device to detect the distance information between the target object and the robot in the target detection area comprises:

[0107] The distance detection device detects the target object in the target detection area; and in a case where the distance detection device does not detect the target object in the target detection area, the distance information is determined based on the width change value, the height change value and the second detection coordinates, wherein the second detection coordinates are the coordinates obtained when the distance detection device last detected the target object.

[0108] In this embodiment, the robot always adjusts the direction of the robot during the movement by control, so that the target object followed is always in the image field of view of the robot, and then the robot controls the distance detection device to search in a target detection region centered on the position information (RegX, RegZ) estimated by vision, as shown in FIG. 3, with a preset distance r as the radius, so as to collect the detection coordinates of the target object.

[0109] If no target object is detected in the target detection region, it means that the detection signal of the position detection device may be blocked, such as there being an obstacle between the chassis radar of the robot and the target object, or the posture of the target object changes during movement, such as the position changes caused by the user stepping, at this time, the position coordinates of the target object can be estimated by interpolation according to the width change value and the height change value obtained by vision detection, and the position coordinates lidarX_t-1 and lidarZ_t-1 detected by the position detection device in the last frame.

[0110] For example, assuming that the width change value is △RegX, the height change value is △RegZ, and the position detection result of the last frame is (lidarX_t-1, lidarZ_t-1), the detection coordinates of the current frame obtained by interpolation are (lidarX, lidarZ), wherein idarX = lidarX_t-1 + beta × △RegX, lidarZ = lidarZ_t-1 + beta × △RegZ, and beta is a preset coefficient.

[0111] When the coordinates of the target object are not successfully detected, the embodiment of the present application estimates the coordinates of the current target object by interpolation combining the coordinate detection result of the last frame and the position estimation based on image detection, so as to ensure that the estimated distance of the target is obtained when the moving target is followed, and the real-time and continuity of distance estimation are ensured.

[0112] In some embodiments of the present application, FIG. 4 shows a logical diagram of robot distance estimation of some embodiments of the present application. As shown in FIG. 4, the input video stream, that is, the image information collected by the robot through the image collection device, is composed of a plurality of continuous image frames. The head of the person, that is, the target object, and the head detection, that is, the process of identifying the target object through image detection. The process of segmented smooth estimation is the process of determining the target correction regression model according to the width-height ratio fluctuation value of the head detection frame. The linear regression of distortion correction is the process of adjusting the target correction regression model according to the lens distortion. After the above steps are completed, the position offset information of the human body from the robot can be estimated, and the chassis radar distance estimation is performed based on the position offset information, that is, the chassis radar collects the coordinates of the user, and the interpolation compensation estimation is the process of estimating the distance information by interpolation based on the detection coordinates of the last frame and the width change value and the height change value when the target object is not detected by the chassis radar.

[0113] Specifically, step 1, for the input video image, human head detection is performed. The average width, average height and average aspect ratio of the human head target of multiple frames are calculated and cached, denoted as average width aHead_w, average height aHead_h and average aspect ratio aHead_r.

[0114] Step 2, the target is estimated to be offset from the robot in the motion process, and the offset distance RegZ and RegX of the human body from the robot are estimated, wherein RegX estimates the offset distance in the horizontal direction, and RegZ estimates the offset distance in the vertical direction. The segmented smoothing method is used in distance estimation.

[0115] Step 2.1, if the head frame width head_w or head_h of the moving target is not within the effective fluctuation range, it is considered that the target may be mutated or be too severely occluded up, down, left and right, and the distance estimation of the last frame is maintained; otherwise, step 2.2 is performed.

[0116] Step 2.2, the aspect ratio head_r of the current human head target is calculated, and the fluctuation value of the aspect ratio is obtained, and the fluctuation value of the aspect ratio is obtained. If the fluctuation value of the aspect ratio of the human head target is greater than or equal to ahead_r_thr, the distortion correction regression model (1) is used for estimation, that is, RegZ=RegZ1 and RegX=RegX1 are used for distance estimation, and if the fluctuation value of the aspect ratio of the human head target is less than ahead_r_thr, the distortion correction regression model (2) is used for estimation, that is, RegZ=RegZ2 and RegX=RegX2 are used for distance estimation. In this way, the problem of distance estimation of the target due to left and right occlusion or front and back occlusion can be effectively reduced. Here, k1, k2, k3, k4 are estimated proportion coefficients, and b1, b2 are estimated offset coefficients. The linear regression method is used to obtain, and the least square method or other methods can be used. human_head_w is the average width of the actual human head, human_head_h is the average height of the actual human head. head_center_w_offset is the pixel offset value of the center point of the human head detection frame in the w direction from the image center, and head_center_h_offset is the pixel offset value of the center point of the human head detection frame in the h direction from the image center.

[0117] wherein RegZ1=k1×image_w / head_w+b1;

[0118] RegX1=k3×(human_head_w / head_w)×(head_center_w_offset);

[0119] RegZ2=k2×image_h / head_h+b2;

[0120] RegX2=k4×(human_head_h / head_h)×(head_center_h_offset).

[0121] When there is vertical distortion in the camera, the regression model RegZ is corrected according to the left and right offset distance RegX, that is, RegZ=RegZ+alpha×|RegX|, when there is horizontal distortion in the camera, the regression model RegX is corrected according to the front and back offset distance RegZ, that is, RegX=RegX+alpha×|RegZ|. Here, alpha is obtained by linear regression learning according to the actual collected data.

[0122] During the movement of the robot, the direction of the robot is always adjusted by control, so that the human target is in the middle of the image field of view, and then the radar searches in the target area with r as the radius according to the visual estimation RegX and RegZ, if the target is searched, the target measurement (lidarX, lidarZ) is obtained as the final human position estimation, if the target is not searched, the lidarX_t-1 and lidarZ_t-1 measured in the previous frame of the radar are interpolated with the change △RegX and △RegZ of the visual estimation position information RegX and RegZ, to obtain the current position estimation lidarX=lidarX_t-1+beta×△RegX, lidarZ=lidarZ_t-1+beta×△RegZ. Finally, the obtained lidarX and lidarZ are used as the position estimation of the robot following the human body, so as to control the robot to follow and approach the human target. Because the distance output is the fusion of visual estimation and radar measurement, the estimated distance of the human body in the movement process can be obtained in real time.

[0123] The linear regression method with distortion correction proposed in the embodiments of the present application can reduce the influence of image distortion at the edge, and at the same time introduces a segmented smoothing mechanism to reduce the influence of occlusion. Moreover, the embodiments of the present application are based on monocular vision, combined with the auxiliary role of radar, to realize the distance estimation of the human body in the movement process, and the frame rate is completely real-time.

[0124] In some embodiments of the present application, a distance information determination apparatus is provided, which is applied to a robot, the robot comprising an image acquisition apparatus and a distance detection apparatus, and a structural block diagram of the distance information determination apparatus is shown in FIG. 5. As shown in FIG. 5, the distance information determination apparatus 500 comprises:

[0125] a detection module 502, configured to perform image detection on the first image frame to obtain detection information corresponding to the target object; wherein the first image frame is acquired by the image acquisition apparatus, and the image content of the first image frame comprises the target object;

[0126] a determination module 504, configured to determine the position information according to the detection information; and determine a target detection region with a first position point indicated by the position information as a center and a preset distance as a radius;

[0127] a control module 506, configured to control the distance detection apparatus to detect the distance information between the target object and the robot within the target detection region.

[0128] In this embodiment, the robot can be a sweeping robot, a service robot or a welcome robot, and the robot can respond to the interactive operation of the user and realize following of the user. The interactive operation of the user can be a voice interactive operation, a remote control interactive operation or a direct control operation on a human-computer interaction interface provided on the robot.

[0129] The image acquisition apparatus provided on the robot can be a camera, such as a monocular camera, and the distance detection apparatus can be a radar device, such as a laser radar. The image acquisition apparatus can acquire the environment image near the robot, and when there is a user nearby, the image acquisition apparatus can realize tracking and shooting of the user, so as to ensure that the user is always within the image field of view of the robot. The distance detection apparatus can be installed on the chassis of the robot, so as to detect the position (or the distance from the robot) of a specific object.

[0130] The target object can be a user, or other moving objects, such as pets, other robots, etc., and the target object can also be a static object, such as a household appliance, furniture, a wall, etc. The robot can track, follow or avoid the target object by estimating the distance information corresponding to the target object. Taking the scenario that the robot follows the target object as an example, whether the estimation of the distance information by the robot is accurate or not will directly affect the following effect.

[0131] In the related art, in order to realize accurate distance measurement, it is generally necessary to use methods such as combining binocular cameras or using TOF (Time Of Flight) sensors to obtain relevant information, which has relatively high cost, and in the scene where the robot and the target object move, the frame rate of these detection methods is often not enough, resulting in that the distance estimation result is not timely and the robot following effect is poor.

[0132] In view of the above problems, the embodiment of the present application combines image recognition and chassis radar detection to improve the real-time performance and accuracy of robot distance estimation while keeping the cost low. Specifically, the first image frame is an image frame acquired by the robot through an image acquisition device at a current time point, and the first image frame includes an image of a target object followed by the robot.

[0133] Exemplarily, the target object can be the whole human body of a user, or a part of the human body of the user, such as a head, a face, etc.

[0134] After the first image frame is acquired, the robot performs image detection on the first image frame, so as to determine detection information of the target object according to the result of the image detection, and determine position information corresponding to the target object through the detection information, which is an estimated position obtained based on image recognition.

[0135] After obtaining the position information, since the position information is an estimated position, in order to obtain the accurate position of the target object, the robot further controls the distance detection device, such as the chassis radar, to emit a detection signal to a certain range corresponding to the position information, so as to obtain the accurate position of the target object and obtain accurate distance information between the target object and the robot.

[0136] Specifically, as shown in FIG. 3, during the movement of the robot, the robot always controls the direction of the robot to make the target object followed always in the image field of view of the robot, and then the robot controls the distance detection device to search in a target detection area with a preset distance r as a radius and with the position information as a center according to the visually estimated position information (RegX, RegZ) as shown in FIG. 3, so as to acquire detection coordinates of the target object, and determine accurate distance information between the target object and the robot according to the detected detection coordinates of the target object.

[0137] The embodiment of the present application can realize accurate detection of the distance of the user by combining the monocular camera and the laser radar, and the computing power demand for image detection on the image frame acquired by the monocular camera is smaller, which can improve the detection frame rate, control the laser radar to perform position detection in a small range based on the image detection result, and can reduce the detection range of the laser radar, thereby significantly shortening the time consumption of position detection and improving the real-time performance of position detection.

[0138] In some embodiments of the present application, the determining module is further configured to determine a target rectification regression model according to the detection information; and apply the target rectification regression model to determine the position information according to the detection information and the image information of the first image frame.

[0139] In the embodiments of the present application, the detection information of the target object can reflect whether the target object is occluded in the first image frame, and the degree of occlusion of the target object in the case that the target object is occluded in the first image frame. For different degrees of occlusion of the target object, a suitable rectification regression model is selected, which is beneficial to reduce the influence of image occlusion on position estimation.

[0140] Specifically, when the detection information reflects that the target object is not occluded or the occlusion of the target object is relatively slight, a common rectification regression model can be selected for position estimation, thereby reducing the amount of calculation. When the detection information reflects that the occlusion of the target object is relatively serious, a rectification regression model for the occlusion can be selected, thereby reducing the influence of the occlusion and improving the accuracy of position estimation.

[0141] After determining the suitable target rectification regression model, the target rectification regression model is applied to estimate the position of the target object in combination with the detection information, and the position information corresponding to the target object is obtained.

[0142] The embodiments of the present application can reduce the error of distance estimation when the target is occluded and improve the accuracy of distance estimation by determining whether the target object is occluded through the detection information, and selecting a corresponding target rectification regression model according to the determination result.

[0143] In some embodiments of the present application, the detection information includes a first detection width and a first detection height of the target object; and the determining module is specifically configured to:

[0144] determine a ratio of the first detection width and the first detection height to obtain a detection width-height ratio; determine a width-height ratio fluctuation value according to a difference between the detection width-height ratio and the average value of the width-height ratio; determine the target rectification regression model based on a comparison result of the width-height ratio fluctuation value and a first threshold value; and determine the position information according to the first detection width, the first detection height, and the target rectification regression model.

[0145] In the embodiments of the present application, the detection information specifically includes width detection information and height detection information. Specifically, the width detection information is a first detection width corresponding to the target object, and the height detection information is a first detection height corresponding to the target object.

[0146] Exemplarily, taking the user's head as an example, when detecting the user's head, the head detection algorithm determines the head position through the detection frame, at this time, the first detection width is specifically the edge frame width of the head detection frame, and the first detection height is specifically the edge frame height of the head detection frame.

[0147] Exemplarily, taking the user's body as an example, at this time, the user's body includes the head, the torso and the limbs, when detecting the user's body, image recognition is performed on the input video image, the user's head, torso or limbs are detected therefrom, and a body detection frame is determined according to the detected body part, at this time, the first detection width is specifically the edge frame width of the body detection frame, and the first detection height is specifically the edge frame height of the body detection frame.

[0148] Supposing the first detection width is head_w, and the first detection height is head_h, then the detection width-height ratio head_r=head_w / head_h is calculated. Further, the width-height ratio fluctuation value is determined, exemplarily, supposing the width-height ratio fluctuation value is △head_r, then △head_r=|head_r-aHead_r|, wherein aHead_r is the pre-determined average width-height ratio.

[0149] Specifically, since the target object and the robot can both be in a motion state, as the positional relationship between the two changes, the target object can be blocked in the image field of view of the robot, and through the width-height ratio fluctuation value, it can be reflected whether the target object is blocked and the degree of blocking when the blocking exists.

[0150] Therefore, according to the comparison result of the width-height ratio fluctuation value and the pre-set first threshold value, it is determined whether there is blocking, and according to the comparison result, the corresponding target correction regression model is determined, wherein the target correction regression model is used to segmentally smooth the distance estimation, thereby reducing the influence of the blocking.

[0151] Exemplarily, supposing the first threshold value is ahead_r_thr, and the target correction regression model includes a first correction regression model, specifically including RegZ=RegZ1, RegX=RegX1, and a second correction regression model, specifically including RegZ=RegZ2, RegX=RegX2.

[0152] Specifically, the following conditions are met:

[0153] RegZ1=k1×image_w÷head_w+b1;

[0154] RegX1=k3×(human_head_w÷head_w)×(head_center_w_offset);

[0155] RegZ2=k2*image_h / head_h+b2;

[0156] RegX2=k4*(human_head_h / head_h)*(head_center_h_offset);

[0157] wherein RegZ and RegX are position information, RegZ1 and RegX1 are the first correction regression model, RegZ2 and RegX2 are the second correction regression model, image_w is the image width of the first image frame, image_h is the image height of the first image frame, head_w is the first detection width, head_h is the first detection height, k1, k2, k3 and k4 are proportional coefficients of distance estimation, b1 and b2 are offset coefficients of distance estimation, human_head_w is the average width of the actual human head, human_head_h is the average height of the actual human head, head_center_w_offset is the pixel offset value of the center point of the human head detection frame in the image in the width direction offset from the image center, and head_center_h_offset is the pixel offset value of the center point of the human head detection frame in the image in the height direction offset from the image center.

[0158] Exemplarily, the proportional coefficients and the offset coefficients can be obtained by a linear regression method or solved by a least square method.

[0159] If the comparison result of the aspect ratio fluctuation value and the first threshold value is Δhead_r>ahead_r_thr, the first correction regression model is used for estimation, that is, RegZ=RegZ1 and RegX=RegX1 are used for distance estimation, and if the comparison result of the aspect ratio fluctuation value and the first threshold value is Δhead_r<ahead_r_thr, the second correction regression model is used for distance estimation, that is, RegZ=RegZ2 and RegX=RegX2 are used for distance estimation. In this way, the problem of distance estimation of the target due to left-right or front-back occlusion can be effectively reduced.

[0160] The embodiments of the present application can reduce the error of distance estimation when the target is occluded and improve the accuracy of distance estimation by comparing the detection aspect ratio with the threshold to determine whether the target object is occluded and determining the corresponding target correction regression model according to the comparison result.

[0161] In some embodiments of the present application, optionally, the determining module is further configured to

[0162] determine an average detection width and an average detection height of the target object, wherein the average detection width is determined according to first detection widths of the target object in the continuous multiple frames of second image frames, the average detection height is determined according to first detection heights of the target object in the continuous multiple frames of second image frames, the second image frames are obtained by the image acquisition device, and the aspect ratio average value is determined according to the average detection width and the average detection height.

[0163] In the embodiments of the present application, for the scene that the robot follows the user, the image acquisition device of the robot continuously collects image data of the followed user, that is, the target object, so as to obtain continuous multiple image frames, that is, the second image frames, wherein the image content of the multiple frames of second image frames all includes the target object.

[0164] The first detection width and the first detection height of the target object in each frame of second image frame are detected respectively, and the average values thereof are calculated to obtain the average detection width of the target object and the average detection height of the target object.

[0165] Exemplarily, assuming that there are N frames of second image frames, there are N first detection widths corresponding to the N frames of second image frames, which are respectively head_w1, head_w2, …, head_wN, and N first detection heights, which are respectively head_h1, head_h2, …, head_hN, then the average detection width ahead_w=(head_w1+head_w2+…+head_wN) / N and the average detection height aHead_h=(head_h1+head_h2+…+head_hN) / N can be calculated.

[0166] After the average detection width ahead_w and the average detection height ahead_h are obtained, the aspect ratio average value aHead_r=ahead_w / ahead_h is calculated according to the ratio of the two.

[0167] The embodiments of the present application calculate the aspect ratio average value of the target object according to the continuous multiple second image frames, and use the aspect ratio average value as the judgment standard for whether the target object is blocked, so that the state of whether the target object is blocked can be accurately recognized, and the corresponding correction regression model is selected according to the recognition result, so that the influence of the target object being blocked on distance estimation can be reduced.

[0168] In some embodiments of the present application, optionally, the determining module is further used for:

[0169] determine a width change value according to a difference between the first detection width and a second detection width, and determine a height change value according to a difference between the first detection height and a second detection height, wherein the second detection width and the second detection height are obtained by image detection on a third image frame, and the third image frame is a previous image frame of the first image frame;

[0170] In a case where at least one of the width change value is greater than a second threshold value and the height change value is greater than a third threshold value is met, the position information is determined according to the second detection width and the second detection height; or in a case where the width change value is less than or equal to the second threshold value and the height change value is less than or equal to the third threshold value, a step of determining a ratio of the first detection width and the first detection height to obtain a detection width-height ratio is performed.

[0171] In the embodiments of the present application, in some scenarios, the target object followed by the robot can be seriously occluded in the field of view of the robot, such as when the target object followed by the robot enters or exits a door, or turns a corner, the door frame and the corner of the wall and other objects can cause a large occlusion to the target object, which can be in the width direction or in the height direction, and is specifically reflected in that the detection height or the detection width of the target object in the image frame of the target object obtained by the robot through the image acquisition device changes greatly. Therefore, when a relatively serious occlusion occurs, it is considered that the current image frame is not suitable for accurate distance detection.

[0172] Specifically, whether an occlusion occurs is determined according to the change value of the detection result of the current first image frame and the detection result of the previous third image frame. It is assumed that the first detection width and the first detection height are detected from the current first image frame, and the second detection width and the second detection height are detected from the third image frame before the first image frame, then the difference between the first detection width and the second detection width is calculated to obtain a width change value, and the difference between the first detection height and the second detection height is calculated to obtain a height change value.

[0173] When the width change value is greater than the second threshold value, or the height change value is greater than the third threshold value, or both are met, it is determined that the target object is seriously occluded, and the current first image frame is not suitable as a basis for distance detection and position detection, so the first image frame is abandoned at this time, and the distance estimation of the previous frame is kept unchanged, that is, the position information of the target object is determined by the second detection width and the second detection height of the third image frame.

[0174] If the width change value is less than or equal to the second threshold value, and at the same time the height change value is less than or equal to the third threshold value, it can be determined that the target object is not seriously occluded, and the current first image frame can be used as a basis for distance detection and position detection, and then the step of determining the position information according to the first detection height and the first detection width corresponding to the first image frame is normally performed.

[0175] The embodiment of the present application determines whether the target object is seriously occluded by the width change value and the height change value detected from two continuous image frames, thereby discarding the image frame with serious occlusion, which can effectively reduce the influence of image occlusion on position detection and improve the accuracy of position detection and distance estimation.

[0176] In some embodiments of the present application, optionally, the second threshold is determined according to the product of the second detection width and a proportionality coefficient, and the third threshold is determined according to the product of the second detection height and the proportionality coefficient, and the proportionality coefficient has a value range of greater than or equal to 0.3 and less than or equal to 0.8.

[0177] In the embodiment of the present application, whether the target object is seriously occluded in the image field of view of the robot is determined according to the comparison result of the width change value, the height change value and the corresponding threshold. When the width change value is greater than the second threshold, it indicates that the target object is seriously occluded in the width direction, and the second threshold is determined according to the detection width of the previous image frame of the current image frame. Taking the current as the first image frame and the previous image frame as the third image frame as an example, the first detection width corresponding to the first image frame is head_w1, and the second detection width corresponding to the third image frame is head_w2, and the corresponding second threshold is the product of the proportionality coefficient k and the second detection width, that is, the second threshold is k x head_w2.

[0178] Similarly, the first detection height corresponding to the first image frame is head_h1, and the second detection height corresponding to the third image frame is head_h2, and the corresponding second threshold is the product of the proportionality coefficient k and the second detection height, that is, the second threshold is k x head_h2.

[0179] Exemplarily, the value range of the proportionality coefficient is 0.3 to 0.8.

[0180] Exemplarily, the proportionality coefficient is 0.5.

[0181] The present application determines the comparison threshold for judging whether the target object in the current image frame is occluded based on the detection width and the detection height of the previous image frame, which can accurately identify whether the target object is occluded, thereby improving the accuracy of position detection and distance estimation.

[0182] In some embodiments of the present application, optionally, the position information includes a first coordinate and a second coordinate, the target rectification regression model includes a horizontal rectification regression model and a vertical rectification regression model, the horizontal rectification regression model is used to determine the first coordinate, and the vertical rectification regression model is used to determine the second coordinate; the determining device further includes:

[0183] The acquisition module is configured to acquire distortion information of the image acquisition device; the adjustment module is configured to, in a case where the distortion information is horizontal distortion, adjust a horizontal correction regression model based on the second coordinate; or, in a case where the distortion information is vertical distortion, adjust a vertical correction regression model based on the first coordinate.

[0184] In this embodiment, the position information is the position information of the target object determined through image detection. For example, the position information is in the format of (RegX, RegZ), where RegX is the first coordinate, and RegZ is the second coordinate. Specifically, the robot self-coordinate system is an (x, y, z) coordinate system. Since the robot and the target to be followed are generally at the same horizontal height when the robot performs a following task, the height coordinate y can be ignored when determining the position of the target object, and the x-axis coordinate and the z-axis coordinate in the horizontal coordinate system, i.e., the first coordinate and the second coordinate, are retained.

[0185] The image acquisition device of the robot is generally a camera, which generally includes an image sensor and a lens assembly. Due to the hardware limitations of the lens assembly, in certain cases, the image frame captured by the camera of the robot can have certain edge distortion. This edge distortion is generally divided into horizontal distortion in the horizontal direction and vertical distortion in the vertical direction.

[0186] When the camera has distortion, the image content in the image frame captured by the robot can be deformed, which can cause a deviation between the position information obtained by image detection to determine the position of the target object and the actual position of the target object, resulting in inaccurate position detection results and ultimately affecting the distance estimation of the robot.

[0187] To address the above problems, different correction regression models are used for different lens distortions in the embodiments of the present application. Specifically, the target correction regression model includes a horizontal correction regression model and a vertical correction regression model, which are used to correct the first coordinate and the second coordinate, respectively.

[0188] For example, the target correction regression model includes RegZ=RegZ1 and RegX=RegX1. RegX=RegX1 is the horizontal correction regression model used to correct the first coordinate RegX, and RegZ=RegZ1 is the vertical correction regression model used to correct the second coordinate RegZ.

[0189] When the lens of the image acquisition device of the robot has distortion in the vertical direction, the vertical correction regression model is adjusted according to the first coordinate in the horizontal direction, that is, RegX, and the adjustment manner is exemplarily RegZ' = RegZ + alpha x |RegX|, where RegZ' is the adjusted vertical correction regression model, RegZ is the vertical correction regression model before adjustment, alpha is a correction coefficient, and alpha can be obtained by linear regression learning, and RexX is the first coordinate.

[0190] Similarly, when the lens of the image acquisition device of the robot has distortion in the horizontal direction, the horizontal correction regression model is adjusted according to the second coordinate in the vertical direction, that is, RegZ, and the adjustment manner is exemplarily RegX' = RegX + alpha x |RegZ|, where RegX' is the adjusted horizontal correction regression model, RegX is the horizontal correction regression model before adjustment, alpha is a correction coefficient, and alpha can be obtained by linear regression learning, and RexZ is the second coordinate.

[0191] It can be understood that the distortion information of the image acquisition device can be determined by the hardware parameters of the image acquisition device. When the robot is produced, the distortion information of the image acquisition device is determined according to the selected hardware model of the image acquisition device, or by experimentally measuring the distortion information of the image acquisition device, so as to determine whether the lens of the image acquisition device has distortion, and whether the lens has distortion in the horizontal direction or the vertical direction.

[0192] The embodiment of the present application adjusts the correction regression model in the direction with distortion according to the distortion information of the image acquisition device, so as to effectively reduce the influence of the lens distortion on the position detection, and improve the accuracy of the position detection and distance estimation.

[0193] In some embodiments of the present application, the control module is further configured to control the distance detection device to detect a target object in a target detection region to obtain a detection coordinate of the target object, and determine distance information based on the detection coordinate.

[0194] In this embodiment, during the movement of the robot, the robot direction is always adjusted by the control, so that the followed target object is always in the image field of view of the robot, and then the robot controls the distance detection device to search in a target detection region with the position information (RegX, RegZ) as the center and a preset distance r as the radius, so as to collect the detection coordinate of the target object, which is recorded as (lidarX, lidarZ). The detection coordinate is obtained by the position detection device of the robot, such as the chassis radar, and therefore has high accuracy.

[0195] After obtaining the detection coordinates, the robot combines the detection coordinates and the self coordinates, and thus the distance information between the followed target object and the robot itself can be calculated.

[0196] The embodiment of the application can effectively reduce the detection range of the radar sensor, and thus improve the detection efficiency and real-time performance of the position detection, by combining the robot vision and the radar sensor, estimating the position of the target object through image recognition, and controlling the radar sensor to perform position detection in a targeted manner according to the position estimation result.

[0197] In some embodiments of the application, optionally, the control module is further configured to control the distance detection device to detect the target object in the target detection area; and the determination module is further configured to determine the distance information based on the width change value, the height change value and the second detection coordinates in a case where the distance detection device does not detect the target object in the target detection area, wherein the second detection coordinates are coordinates obtained when the distance detection device last detected the target object.

[0198] In this embodiment, the robot always adjusts the direction of the robot during movement, so that the followed target object is always in the image field of view of the robot, and then the robot controls the distance detection device to search in a target detection area with the position information (RegX, RegZ) estimated by vision as the center and a preset distance r as the radius, so as to collect the detection coordinates of the target object.

[0199] If the target object is not detected in the target detection area, it indicates that the detection signal of the position detection device can be blocked, for example, there is an obstacle between the chassis radar of the robot and the target object, or the posture of the target object changes during movement, for example, the user steps, causing the position to change. At this time, the position coordinates of the target object can be estimated by interpolation according to the width change value and the height change value obtained by vision detection, and the position coordinates (lidarX_t-1, lidarZ_t-1) detected by the position detection device in the last frame.

[0200] For example, the width change value is △RegX, the height change value is △RegZ, and the position detection result of the last frame is (lidarX_t-1, lidarZ_t-1). The detection coordinates of the current frame obtained by interpolation are (lidarX, lidarZ), wherein idarX = lidarX_t-1 + beta x △RegX, lidarZ = lidarZ_t-1 + beta x △RegZ, and beta is a preset coefficient.

[0201] The embodiment of the present application can obtain the estimated distance of the target when following the moving target by interpolating the coordinate of the current target object through combining the coordinate detection result of the last frame and the position estimation based on image detection when the coordinate of the target object is not successfully detected, thereby guaranteeing the real-time and continuity of distance estimation.

[0202] In some embodiments of the present application, a distance information determination apparatus is provided. FIG. 6 shows a structural block diagram of the distance information determination apparatus according to some embodiments of the present application. As shown in FIG. 6, the distance information determination apparatus 600 comprises a memory 602 for storing programs or instructions, and a processor 604 for executing the programs or instructions to implement the steps of the distance information determination method according to any of the above embodiments, thus having all the advantages of the distance information determination method according to any of the above embodiments. To avoid repetition, no further elaboration is made herein.

[0203] In some embodiments of the present application, a readable storage medium is provided, which stores programs or instructions. The programs or instructions are executed by a processor to implement the steps of the distance information determination method according to any of the above embodiments, thus having all the advantages of the distance information determination method according to any of the above embodiments. To avoid repetition, no further elaboration is made herein.

[0204] In some embodiments of the present application, a robot is provided, which comprises the distance information determination apparatus according to any of the above embodiments and / or the readable storage medium according to any of the above embodiments, thus having all the advantages of the distance information determination apparatus according to any of the above embodiments and / or the readable storage medium according to any of the above embodiments. To avoid repetition, no further elaboration is made herein.

[0205] The method can be implemented in various ways according to specific features and / or example applications. For example, the method can be implemented by a combination of hardware, firmware and / or software. For example, in a hardware implementation, the processor can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the above functions and / or combinations thereof.

[0206] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, and any suitable combination of the foregoing. Computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media, or electrical signals through a wire, digital or analog communication links, wireless communications links, and the like.

[0207] In the description of the present application, the term "a plurality of" means two or more, unless otherwise explicitly defined, and the terms "upper", "lower", and the like, indicate the orientation or positional relationship as shown in the drawings and are merely used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application; the terms "connection", "installation", "fixation" and the like should be understood broadly, for example, "connection" can be fixed connection, can also be detachable connection, or integral connection; can be direct connection, or indirect connection through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0208] In the description of the present application, the terms "one embodiment", "some embodiments", "a specific embodiment", and the like, mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present application, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0209] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method of determining distance information, wherein, The determining method is executed by a robot, the robot comprising an image acquisition device and a distance detection device, and the determining method comprises: performing image detection on a first image frame to obtain detection information corresponding to a target object; wherein the first image frame is acquired by the image acquisition device, and image content of the first image frame comprises the target object; determining position information according to the detection information; determining a target detection region with a first position point indicated by the position information as a center and a preset distance as a radius; controlling the distance detection device to detect distance information between the target object and the robot within the target detection region.

2. The determination method according to claim 1, wherein The step of determining position information according to the detection information specifically comprises: determining a target rectification regression model according to the detection information; applying the target rectification regression model to determine the position information according to the detection information and image information of the first image frame.

3. The determination method according to claim 2, wherein The detection information comprises a first detection width and a first detection height of the target object. The step of determining a target rectification regression model according to the detection information specifically comprises: determining a ratio of the first detection width and the first detection height to obtain a detection width-height ratio; determining a width-height ratio fluctuation value according to a difference between the detection width-height ratio and an average width-height ratio; determining a target rectification regression model based on a comparison result of the width-height ratio fluctuation value and a first threshold value.

4. The determination method according to claim 3, wherein Before the step of determining position information according to the detection information, the determining method further comprises: determining an average detection width and an average detection height of the target object; wherein the average detection width is determined according to the first detection width of the target object in continuous multiple second image frames, and the average detection height is determined according to the first detection height of the target object in the continuous multiple second image frames, the second image frames being acquired by the image acquisition device; determining the average width-height ratio according to the average detection width and the average detection height.

5. The determination method according to claim 3, wherein Before the step of determining a ratio of the first detection width and the first detection height to obtain a detection width-height ratio, the determining method further comprises: determining a width change value according to a difference between the first detection width and a second detection width, and determining a height change value according to a difference between the first detection height and a second detection height; wherein the second detection width and the second detection height are obtained by performing image detection on a third image frame, the third image frame being a previous image frame of the first image frame; in a case where at least one of the width change value being greater than a second threshold value and the height change value being greater than a third threshold value is met, determining the position information according to the second detection width and the second detection height; or in a case where the width change value is less than or equal to the second threshold value and the height change value is less than or equal to the third threshold value, performing the step of determining a ratio of the first detection width and the first detection height to obtain a detection width-height ratio.

6. The determination method according to claim 5, wherein The second threshold is determined according to a product of the second detection width and a proportionality coefficient, the third threshold is determined according to a product of the second detection height and the proportionality coefficient, and the proportionality coefficient ranges from greater than or equal to 0.3 to less than or equal to 0.

8.

7. The determination method according to any one of claims 3 to 5, wherein The position information includes a first coordinate and a second coordinate, the target correction regression model includes a horizontal correction regression model and a vertical correction regression model, the horizontal correction regression model is used to determine the first coordinate, and the vertical correction regression model is used to determine the second coordinate. After the step of determining the position information according to the detection information, the determination method further includes: obtaining distortion information of the image acquisition device; in a case where the distortion information is horizontal distortion, adjusting the horizontal correction regression model based on the second coordinate, or in a case where the distortion information is vertical distortion, adjusting the vertical correction regression model based on the first coordinate.

8. The determination method according to any one of claims 1 to 6, wherein, The step of controlling the distance detection device to detect the distance information between the target object and the robot in the target detection area includes: controlling the distance detection device to detect the target object in the target detection area to obtain a detection coordinate of the target object; and determining the distance information based on the detection coordinate.

9. The determination method according to claim 5 or 6, wherein The step of controlling the distance detection device to detect the distance information between the target object and the robot in the target detection area includes: controlling the distance detection device to detect the target object in the target detection area; in a case where the distance detection device does not detect the target object in the target detection area, determining the distance information based on a second detection coordinate and the width change value and the height change value, wherein the second detection coordinate is a coordinate obtained when the distance detection device last detected the target object.

10. A distance information determining apparatus, wherein, The determination device is applied to a robot, and the robot includes an image acquisition device and a distance detection device. The determination device includes: a detection module, configured to perform image detection on a first image frame to obtain detection information corresponding to a target object, wherein the first image frame is obtained by the image acquisition device, and image content of the first image frame includes the target object; a determination module, configured to determine position information according to the detection information; and a control module, configured to control the distance detection device to detect distance information between the target object and the robot in a target detection area, wherein the target detection area is determined with a first position point indicated by the position information as a center and with a preset distance as a radius.

11. A distance information determining apparatus, wherein, The determination device includes: a memory, configured to store programs or instructions; a processor, configured to execute the programs or instructions to implement steps of the determination method in any one of claims 1 to 9.

12. A readable storage medium, on which a program or instructions are stored, wherein, The programs or instructions are executed by the processor to implement steps of the determination method in any one of claims 1 to 9.

13. A robot, wherein, includes: the distance information determination device in claim 10 or 11; and / or the readable storage medium in claim 12.

Citation Information

Patent Citations

  • Target obstacle detection method and device and robot

    CN111337022A

  • Sweeping robot control method and sweeping robot

    CN111759227A

  • Obstacle avoidance method and device for cleaning robot, electronic equipment and medium

    CN114391777A

  • Cleaning robot and cleaning robot control method

    CN115500740A

  • Image acquisition method and device, computer equipment and storage medium

    CN115550552A