Smart home robot and task execution method and storage medium thereof

By combining the image data of the main camera of the smart home robot and the external camera, the position of the target object in the three-dimensional map is calculated and weighted calculation is performed, the task execution problem caused by the main camera being blocked is solved, and the success rate and accuracy of the task are improved.

CN119839883BActive Publication Date: 2025-05-13WOCAO TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510315482.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-13
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The ontological camera of smart home robots may be blocked in some special environments, resulting in the inability to obtain task feedback, affecting the accuracy and success rate of task execution.

Method used

By obtaining image data of the main camera of the smart home robot and the external cameras deployed in the environment, the position of the same target object collected by different cameras in the three-dimensional map is calculated, and the position weighting calculation is performed based on the position weights under each camera to determine the final position of the target object to control the robot to complete the target task.

Benefits of technology

When the main camera is blocked or the target is moved out of the field of view, the task is continued with an external camera, which improves the success rate and accuracy of task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119839883B_ABST
    Figure CN119839883B_ABST
Patent Text Reader

Abstract

The present application relates to the field of robots, and discloses a smart home robot and its task execution method and storage medium, the method comprising: responding to task instructions to obtain multiple images collected by various cameras at the same time; for images with target objects, obtaining depth information of the target objects in the images to calculate the initial pose of the target objects in the associated camera coordinate system; obtaining multiple transformed poses of the target objects in the three-dimensional map according to the initial pose of the target objects and the pose of the associated cameras in the three-dimensional map; for each transformed pose, determining the corresponding pose weight according to the depth information and the initial pose; performing pose weighted average calculation according to each transformed pose and the associated pose weight to obtain the adjusted pose of the target object; and controlling the smart home robot to perform the target task according to the adjusted pose. The present application uses data from multiple external cameras to identify and locate objects, so that the robot can perform tasks more flexibly and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robotics, and in particular to an intelligent home robot and a task execution method and storage medium thereof. Background Art

[0002] When performing tasks through the main camera set on the smart home robot, some special environments may be encountered, such as picking up items from behind a box, or a person suddenly turns and walks behind a door when performing a following task. The main camera of the smart home robot will be blocked, resulting in the main camera being unable to obtain task feedback (for example, when performing a grasping task, the camera set at the eye of the smart home robot is blocked by the manipulator itself, or the camera set at the wrist of the smart home robot enters the blind spot and becomes ineffective, etc.), resulting in the smart home robot being unable to perform the task or being unable to accurately determine whether the task is successful. Summary of the invention

[0003] In view of this, an embodiment of the present application provides a smart home robot, a task execution method thereof, and a storage medium.

[0004] In a first aspect, the present application provides a method for executing a task of a smart home robot, comprising:

[0005] Responding to the task instruction to obtain multiple images collected by each camera at the same time; the camera includes a main camera arranged on the smart home robot and multiple external cameras deployed at different positions in the home environment;

[0006] In the case where a target object is identified in a plurality of the images, for each image in which the target object exists, obtaining depth information of the target object in the image to calculate an initial position and posture of the target object in an associated camera coordinate system;

[0007] According to each of the initial poses of the target object and the pose of the associated camera in the three-dimensional map of the home environment, obtaining a plurality of transformed poses of the target object in the three-dimensional map;

[0008] For each of the transformed postures, determining a corresponding posture weight according to the depth information of the associated target object and the initial posture;

[0009] Performing posture weighted average calculation according to each of the transformed postures and the associated posture weights to obtain an adjusted posture of the target object;

[0010] The smart home robot is controlled to perform a target task according to the adjusted posture.

[0011] In an optional implementation, determining a corresponding posture weight according to the depth information of the associated target object and the initial posture includes:

[0012] Determining a depth estimation value of the target object according to the associated depth information; wherein the depth estimation value is used to characterize the distance between the target object and the camera;

[0013] Determine a first weight according to the depth estimation value of the target object; wherein the larger the depth estimation value is, the smaller the first weight is;

[0014] Determining a second weight according to the initial posture of the associated target object; wherein the closer the target object is to the edge of the image, the smaller the second weight is;

[0015] According to the first weight and the second weight, a corresponding posture weight is calculated.

[0016] In an optional implementation manner, the calculation formula of the first weight is: ;

[0017] The calculation formula of the second weight is: ;

[0018] The calculation formula of the posture weight is: ;

[0019] In the formula, is the depth estimation value; is the first weight; is the second weight; is the angle at which the target object deviates from the center line of the camera's field of view; is the posture weight.

[0020] In an optional implementation, after controlling the smart home robot to perform a target task according to the adjusted posture, the method further includes:

[0021] If it is identified that the target object does not exist in at least one of the images, it is confirmed that the corresponding associated camera has lost the target object, and the motion information of the target object is calculated according to a preset number of historical transformation postures associated before the target object is lost; wherein the motion information includes motion speed, motion direction and motion acceleration;

[0022] Predict the transformed posture associated with the target object in the three-dimensional map according to the motion information, and return to execute the step of performing posture weighted average calculation according to each transformed posture and the associated posture weight to obtain the adjusted posture of the target object.

[0023] In an optional implementation, after predicting the transformed posture associated with the target object in the three-dimensional map according to the motion information, the method further includes:

[0024] According to the loss time of the target object under the associated camera, the confidence of the predicted transformed posture is adjusted, and when the confidence drops below a preset value, the transformed posture of the target object under the associated camera in the three-dimensional map is discarded.

[0025] In an optional implementation, after discarding the transformed position of the target object under the associated camera in the three-dimensional map, the method further includes:

[0026] If it is recognized that the transformed poses of the target object under each camera in the three-dimensional map are discarded, it is determined that the target object is lost, and the smart home robot is controlled to stop executing the target task.

[0027] In an optional implementation, the calculating the initial pose of the target object in the associated camera coordinate system includes:

[0028] Constructing an object coordinate system corresponding to the target object based on the geometric features of the target object;

[0029] Determining the origin of the object coordinate system;

[0030] According to the depth point cloud data of the target object, align the object coordinate system with the camera coordinate system of the associated camera to obtain a rotation matrix from the object coordinate system to the camera coordinate system; wherein the depth point cloud data is generated based on the depth information of the target object;

[0031] Obtaining a translation vector according to the position from the origin of the object coordinate system to the camera coordinate system;

[0032] The initial position of the target object in the camera coordinate system is determined according to the rotation matrix and the translation vector.

[0033] In an optional embodiment, the three-dimensional map includes environmental image features; after acquiring the multiple images respectively captured by each camera in real time, and before acquiring the multiple transformed positions of the target object in the three-dimensional map, the method further includes:

[0034] For each of the images, feature matching is performed between the image and the environmental image features in the three-dimensional map to obtain the position and posture of the associated camera in the three-dimensional map.

[0035] In a second aspect, the present application provides a smart home robot, comprising:

[0036] An actuator is used to execute corresponding actions;

[0037] a memory storing a computer program;

[0038] A processor is used to execute the computer program to implement the smart home robot task execution method described in the above embodiment.

[0039] In a third aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed on a processor, implements the smart home robot task execution method described in the aforementioned embodiment.

[0040] The embodiments of the present application have the following beneficial effects:

[0041] This application enables the smart home robot to obtain image data from multiple external cameras deployed in the environment space in real time in addition to the image data from the main camera, and then uses coordinate transformation to calculate the position and posture of the same target object captured by different cameras in the three-dimensional map, and then combines the position and posture weights of each camera to perform weighted position and posture calculation on the target object to comprehensively determine the final position and posture of the target object, thereby controlling the robot to complete the target task. The method of this application enables the robot to use multiple cameras in the external environment to jointly identify and locate the target object, so that when the main camera is blocked, it can use other external cameras to continue to perform the task, which can improve the success rate of task completion. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A schematic diagram of the structure of the smart home robot according to an embodiment of the present application is shown;

[0044] Figure 2 A first flow chart of a method for executing a task of a smart home robot according to an embodiment of the present application is shown;

[0045] Figure 3 A schematic diagram of a scene in which multiple external cameras are deployed in a home environment space according to an embodiment of the present application is shown;

[0046] Figure 4 A second flow chart of the method for executing a task of a smart home robot according to an embodiment of the present application is shown;

[0047] Figure 5 A third flow chart of the smart home robot task execution method according to an embodiment of the present application is shown;

[0048] Figure 6 A schematic diagram showing the positional relationship between the target object and the camera in a two-dimensional plane in an embodiment of the present application is shown;

[0049] Figure 7 A structural schematic diagram of a smart home robot task execution device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments.

[0051] The components of the embodiments of the present application generally described and shown in the drawings herein may be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0052] Hereinafter, the terms "including", "having" and their cognates that can be used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or a combination of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or a combination of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or a combination of the foregoing items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and cannot be understood as indicating or implying relative importance.

[0053] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meanings as those generally understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meanings as the contextual meanings in the relevant technical field and will not be interpreted as having idealized meanings or overly formal meanings unless clearly defined in the various embodiments of the present application.

[0054] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0055] When some tasks need to be performed through the camera on the smart home robot body, the camera may be blocked, which may affect the normal execution of the task or the confirmation of the success rate. For this reason, the present application proposes to enable the smart home robot to use the camera that can be freely deployed in the external environment space to combine with the camera on the smart home robot body to jointly identify and locate the target object, so as to assist the smart home robot in performing tasks and thereby improve the success rate of task completion.

[0056] Among them, the smart home robot in this application is a robot that can realize intelligent behavior through real-time perception, interaction and action with the real environment. Among them, the smart home robot has mechanisms such as mechanical arms, mobile chassis, sensors, etc., and can act in the physical world like humans or animals (such as walking, grasping, avoiding obstacles, etc.). In addition, it can also perceive the environment in real time through sensors such as vision, touch, hearing, force, etc., and adjust behavior according to feedback. Smart home robots can include cleaning robots (such as sweeping robots, mopping robots, sweeping and mopping robots, etc.), mechanical arms, humanoid robots, grasping robots, and companion robots, etc.

[0057] Figure 1 A structural schematic diagram of the smart home robot 10 according to an embodiment of the present application is shown.

[0058] Exemplarily, the smart home robot 10 includes a processor 11, a memory 12, a sensing unit 13 and an actuator 14, wherein the sensing unit 13 is used to detect environmental information of the smart home robot 10; the actuator 14 is used to perform corresponding actions; the memory 12 stores a computer program, and the processor 11 runs the computer program, so that the smart home robot 10 performs the target task according to the smart home robot task execution method of the following embodiment.

[0059] Among them, the processor 11 can be an integrated circuit chip with signal processing capabilities. The processor 11 can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or at least one of other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application.

[0060] The memory 12 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc. The memory 12 is used to store a computer program, and the processor 11 may execute the computer program accordingly after receiving an execution instruction.

[0061] Among them, the perception unit 13 may include detection sensors, visual sensors, etc. arranged on the body of the smart home robot 10, wherein the detection sensor is mainly used to detect targets or obstacles on the path of the smart home robot 10 when performing tasks. For example, for a cleaning robot, its detection sensor may include a laser radar, and the environment space point cloud information can be obtained by scanning the environment around the path of travel by the laser radar, which can then be used for tasks such as building a three-dimensional map or avoiding obstacles. In some optional embodiments, the detection sensor may also include an infrared sensor, an ultrasonic sensor, etc. Among them, the visual sensor may include a camera (referred to as the body camera S0), such as a conventional RGB (red, green, and blue) camera, a depth camera, or a stereo camera, etc., which is mainly used to collect image data in the current environment space, and then can be used for the identification and positioning of the target object.

[0062] It is worth noting that the location and number of the main body camera on the smart home robot 10 are not limited. For example, for a grasping robot, the main body camera can be set on the grasping mechanism arm, etc.; for a cleaning robot, the main body camera can be set on the top of the main body of the smart home robot 10 to obtain a wider field of view.

[0063] Among them, the actuator 14 is mainly used to make the smart home robot 10 perform corresponding actions. For example, for a grasping robot, its actuator 14 includes a walking mechanism and a grasping mechanism, such as a walking mechanism may include but is not limited to side wheels, track wheels, tracks, rollers, etc., and a grasping mechanism includes but is not limited to a mechanical arm and a clamp connected to the mechanical arm, etc. For another example, for a cleaning robot, its actuator includes a walking mechanism and a cleaning mechanism, wherein the cleaning mechanism includes but is not limited to any one or more combinations of cleaning parts, fans, and pump bodies. For example, cleaning parts may include but are not limited to side brushes, roller brushes, and rollers, etc. Different types of cleaning parts can perform different cleaning operations; fans can be used for vacuuming, and pump bodies can be used for pumping water to rollers, etc. It can be understood that the specific structure of the actuator 14 can be determined according to the actual needs of the smart home robot 10, and no specific limitation is made here.

[0064] Figure 2 A flowchart of a method for executing a task of a smart home robot according to an embodiment of the present application is shown. Exemplarily, the method for executing a task of a smart home robot includes steps S110 to S160:

[0065] Step S110, responding to the task instruction to obtain multiple images collected by each camera at the same time; the camera includes a main camera provided on the smart home robot and multiple external cameras deployed at different positions in the home environment.

[0066] Among them, the external camera may include any one of a depth camera, a monocular camera or a binocular camera, etc. For the external camera deployed in the home environment, it can be set at different positions with a certain height in the environment space, for example, it can be deployed on the ceiling, on the wall, etc. Figure 3 Three external cameras (respectively denoted as S1 to S3) are set on the wall to cover different areas. In addition, there is no limit on the number of external cameras. If conditions permit, it is best if the field of view of these cameras can cover the entire environment space, and the external cameras and the smart home robot can communicate wirelessly (such as WiFi or Bluetooth) to achieve information exchange.

[0067] Among them, the task instructions can be sent to the smart home robot 10 by an external object (such as a user) through a terminal device (such as a mobile phone, tablet, smart speaker, etc.) or a button on the smart home robot 10, or can be sent to the smart home robot 10 in the form of voice commands.

[0068] Different smart home robots 10 often receive different task instructions. For example, for a grasping robot, its task instruction may be, but is not limited to, "please pass me the cup on the table", "grab the red block and put it in the box", "recycle the tools in the third drawer on the right", "please pass me the slippers", etc. For a following robot, its task instruction may be, but is not limited to, "follow a person (such as me, Xiaobao), etc."

[0069] After receiving the task instruction, the smart home robot 10 will first analyze the task instruction and determine the target object in the task instruction. It can be understood that the target object can be static or dynamic; and the number of target objects can be one or more. For example, for a grasping robot, if the instruction issued is "pass me the water cup on the table", it will use the water cup on the table as the target object (need to identify and locate the table and water cup); if the task instruction is "please pass me the slippers", it will use the slippers as the target object.

[0070] After determining the target object, the smart home robot 10 will collect images through the main camera installed on the smart home robot 10, and will also obtain multiple images collected by various external cameras deployed in the home environment, such as Figure 3 As shown, at the same time, the main camera S0 captures image 0, the external camera S1 captures image 1, the external camera S2 captures image 2, and the external camera S3 captures image 3, and the smart home robot 10 simultaneously captures image 0, image 1, image 2, and image 3. In this step, by acquiring multiple images captured by each camera at the same time, if the target object exists in multiple images, it can be ensured that the state of the target object captured by each camera is the same at the current moment, so that the target object can be more accurately positioned in combination with each camera in the subsequent steps.

[0071] It can be understood that since the positions of the external cameras are fixed, that is, the areas captured by the external cameras are fixed, for dynamically moving targets, as time goes by, sometimes the targets can be captured, and sometimes they cannot be captured, or for targets that are not in the area captured by a certain external camera, the external camera cannot capture the targets. Therefore, in this application, for multiple images captured by each camera, the smart home robot 10 will first identify whether there is a target in each image.

[0072] Step S120, when it is identified that there are target objects in multiple images, for each image in which the target object exists, depth information of the target object in the image is obtained to calculate the initial position and posture of the target object in the associated camera coordinate system.

[0073] Exemplarily, after all acquired images are identified, it can be determined whether each image contains a target object. If an image contains a target object, the depth information of the target object is further obtained based on the image. Then, the initial position of the target object in the associated camera coordinate system is calculated based on the depth information of the target object, where the associated camera coordinate system refers to the camera coordinate system of the camera corresponding to the image, for example Figure 3 As shown, if the image 1 captured by the external camera S1 contains a target object, the depth information of the target object is obtained according to the image 1, and the initial pose of the target object in the camera coordinate system of the corresponding external camera S1 is calculated by the depth information; if an image does not contain a target object, the image is not further processed.

[0074] It should be noted that, when the target object is identified in multiple images, multiple initial poses can be obtained from each image. Figure 3 As shown, if the image 1 captured by the external camera S1 contains the target object, and the image 2 captured by the external camera S2 also contains the target object, the initial position 1 of the target object in the camera coordinate system of the external camera S1 can be obtained based on the image 1, and the initial position 2 of the target object in the camera coordinate system of the external camera S2 can be obtained based on the image 2.

[0075] Among them, the depth information is mainly reflected by the distance between the target object and the camera. For example, the depth information can be obtained by directly or indirectly calculating the collected images of each camera (depth camera, monocular camera or dual-phase camera, etc.), and further obtaining the depth point cloud data. For example, if the camera is a depth camera, the depth information can be directly output to form a depth map; if the camera is a monocular camera, it is necessary to combine additional depth estimation algorithms (such as structured light, time-of-flight method, etc.) to estimate the depth information to form a depth map; if the camera is a binocular camera, the depth information can be calculated by comparing the parallax of the same point in the images taken by the two cameras to form a depth map. Then, each pixel in the depth map is converted into a three-dimensional coordinate to form point cloud data. The point cloud data is segmented according to the target recognition box (usually the bounding box of the target object identified by the image processing algorithm (such as Mask R-CNN, YOLO, etc.)), and finally the depth point cloud data of the target object is obtained.

[0076] It can be understood that since different images are captured by cameras at different viewing angles, the above-mentioned association means that there is a one-to-one correspondence between a certain camera and the corresponding image it captures. In addition, it should be understood that the initial posture, depth information and transformed posture mentioned later of the corresponding image also have an association relationship with the camera.

[0077] Regarding the acquisition of the above initial posture. In one embodiment, for example, Figure 4 As shown, it includes steps S210 to S250:

[0078] Step S210: constructing an object coordinate system corresponding to the target object based on the geometric features of the target object.

[0079] Step S220, determining the origin of the object coordinate system.

[0080] The geometric features include two-dimensional features and three-dimensional features of the target object, such as the vertices, faces, lines, etc. of the target object. Exemplarily, by extracting the geometric features of the target object from the image, and then constructing an object coordinate system on the target object based on these geometric features and the structural features of the target object, including the coordinate origin and three axes (i.e., XYZ axes) of the object coordinate system.

[0081] For example, assuming that the target object is a water bottle, when constructing the object coordinate system corresponding to the water bottle, the center of the bottom of the identified water bottle can be used as the origin, the direction from the center of the bottom of the bottle to the center of the bottle cap is the Z-axis direction, and the direction parallel to the bottom of the bottle is defined as the XY plane.

[0082] It can be understood that for the target objects captured by different cameras, their own object coordinate systems can be constructed based on their recognition results. The construction of the object coordinate system is mainly used to facilitate the calculation of the initial pose of the target object in the current image in the corresponding camera coordinate system, and then to facilitate the conversion of the pose of the target object under each shooting angle to the same reference object (such as a three-dimensional map) for fusion.

[0083] Step S230 , aligning the object coordinate system and the camera coordinate system of the associated camera according to the depth point cloud data of the target object, and obtaining a rotation matrix from the object coordinate system to the camera coordinate system.

[0084] The depth point cloud data can be further generated based on the depth information of the target object. Exemplarily, when obtaining the rotation matrix, some key points on the target object can be determined first. These key points usually have significant geometric features, such as corner points, edge points or plane areas, etc., and then the coordinates of these key points in the object coordinate system are determined, and the positions of these key points in the camera coordinate system are determined by the depth point cloud data. By comparing the positions of the key points in the object coordinate system and the camera coordinate system, the rotation matrix from the object coordinate system to the camera coordinate system can be calculated.

[0085] Step S240, obtaining a translation vector according to the position from the origin of the object coordinate system to the camera coordinate system.

[0086] Exemplarily, based on the above-mentioned depth point cloud data, the actual position of the point of the target object selected as the coordinate origin of the object coordinate system can be determined, and then the translation vector can be obtained.

[0087] Step S250, determining the initial position and posture of the target object in the camera coordinate system according to the rotation matrix and the translation vector.

[0088] Exemplarily, when the rotation matrix from the object coordinate system to the camera coordinate system and the translation vector of the object coordinate system relative to the camera coordinate system are known, the initial position of the target object in the corresponding camera coordinate system can be calculated using the coordinate system transformation principle.

[0089] It can be understood that by obtaining the initial pose of the target object in the corresponding camera coordinate system from the images taken at different viewing angles, and then knowing the poses of these cameras in the three-dimensional map, the poses of the target object taken from different viewing angles in the three-dimensional map can be calculated through further coordinate transformation, which provides a basis for the subsequent fusion processing of these poses.

[0090] Step S130, obtaining multiple transformed postures of the target object in the three-dimensional map according to each initial posture of the target object and the posture of the associated camera in the three-dimensional map of the home environment.

[0091] Among them, regarding the position and posture of each camera in the three-dimensional map of the environment, optionally, for the main body camera, since the main body camera is fixed on the smart home robot 10, when the position and posture of the smart home robot 10 in the three-dimensional map is known, according to the positioning function of the smart home robot itself, the position and posture of the main body camera in the three-dimensional map can be obtained. For example, for the convenience of calculation, the position and posture of the smart home robot 10 in the three-dimensional map can be directly used as the position and posture of the main body camera in the three-dimensional map. Optionally, in one embodiment, the three-dimensional map includes environmental image features. For the main body camera and the external camera, for each image, the image is feature matched with the environmental image features in the three-dimensional map to obtain the position and posture of the associated camera in the three-dimensional map, for example Figure 3As shown, the image 0 of the main camera S0, the image 1 of the external camera S1, the image 2 of the external camera S2, and the image 3 of the external camera S3 are matched with the environmental image features in the three-dimensional map respectively, so as to obtain the position and posture of the main camera S0 through image 0, the position and posture of the main camera S1 through image 1, the position and posture of the main camera S2 through image 2, and the position and posture of the main camera S3 through image 3 according to the features in each image that are the same or similar to the environmental image features in the three-dimensional map. As an optional solution, the position and posture of the external cameras in the three-dimensional map can also be obtained by pre-manual marking. For example, after the three-dimensional map is constructed, the deployment positions of each external camera are manually marked in the three-dimensional map so that the smart home robot can read the position and posture of each external camera in the three-dimensional space.

[0092] The above three-dimensional map is a three-dimensional map of the home environment space where the smart home robot 10 is currently walking. It is pre-constructed by the smart home robot 10 before step S110 and can be obtained by controlling the smart home robot 10 to walk in the environment space and perform real-time positioning and mapping. It is worth noting that the three-dimensional map constructed by the present application also carries environmental image features, and more environmental feature information can be obtained by combining environmental image features. In one embodiment, Figure 5 As shown, the construction of the above three-dimensional map includes steps S310 to S320:

[0093] Step S310, controlling the smart home robot to move in the environment, so as to collect environmental spatial data in real time through detection sensors and to collect multiple environmental images in real time through the body camera.

[0094] Step S320, performing feature fusion based on the environmental space data and the environmental image to construct a three-dimensional map.

[0095] Exemplarily, during the movement, a map construction algorithm (such as SLAM, Simultaneous Localization and Mapping) can be used to initially construct a three-dimensional map, wherein the environmental space data mainly refers to the physical environment information collected by the smart home robot 10 through detection sensors (such as laser radar, ultrasonic sensor, etc.), such as but not limited to point cloud data, distance information and location information of objects, etc. Then, the environmental image data captured by the detection sensor and the body camera are fused based on the features of the object, so that a three-dimensional environmental map carrying the image feature data of each object can be constructed.

[0096] Step S140: for each transformed posture, determine the corresponding posture weight according to the depth information and the initial posture of the associated target object.

[0097] Taking into account that although multiple cameras have captured images of the target object, the distance between each camera and the target object is different, and the angles at which the target object is photographed are different, resulting in different pose accuracy of the target object in the corresponding image. For this reason, the present application proposes to set the pose weights of the target object captured by each camera according to the depth information and initial pose of the target object in the image, and finally fuse these pose information to obtain a more accurate pose of the target object in the three-dimensional map.

[0098] Since the accuracy of the horizontal pose estimation of the camera is usually high, while the accuracy of the depth estimation value confirmed by the depth information is low, and the accuracy is lower as the distance is farther, based on this, in one embodiment, the weight of each camera's transformed pose can be set according to the initial pose and depth estimation value of the target object in each camera image. Among them, the depth estimation value associated with the target object can be calculated according to the associated depth information and through the depth estimation algorithm; wherein the depth estimation value is used to characterize the distance between the target object and the camera.

[0099] Exemplarily, in some embodiments, a first weight is determined based on the depth estimation value of the target object; wherein, the larger the depth estimation value, the smaller the first weight; conversely, the smaller the depth estimation value, the larger the first weight. At the same time, a second weight is also determined based on the initial posture of the associated target object; wherein, the closer the target object is to the edge of the image, the smaller the second weight. Finally, based on the first weight and the second weight, the total weight corresponding to the initial posture (i.e., the posture weight mentioned above) is calculated.

[0100] For ease of understanding, Figure 6 It shows the positional relationship between the target object and a single camera in a two-dimensional plane. Assume that the circle refers to the camera, the rectangle refers to the observed target object, and the dotted line is the approximate field of view of the camera. Here, the center of the camera is taken as the origin O of the camera coordinate system, and the direction of its field of view midline is defined as the camera coordinate system Direction, and The direction is perpendicular to the camera coordinate system Direction, such as Figure 6 As shown, the distance between the target object and the camera can be obtained through the depth estimation algorithm (denoted as ), which is used as the depth estimation value; through the initial position of the target in the camera coordinate system, the deviation of the target object can be calculated according to the pixels in the image The angle of direction (denoted as ).

[0101] Since the depth estimate is estimated with low certainty and increases with the depth estimate As the depth increases, the certainty will be further reduced. In one embodiment, according to the depth estimation value To calculate a weight (ie the first weight): ; For this formula, when the depth estimate When the distance between the target object and the camera approaches 0, the first weight is 1. When it approaches infinity, that is, when the distance between the target object and the camera approaches infinity, the first weight is 0.

[0102] In addition, due to the lens distortion of the camera, there will be some errors at the edge of the lens. Generally, the closer to the edge, the more serious the distortion, the lower its accuracy, and the smaller the corresponding weight. Based on this, in one embodiment, another weight (i.e., the second weight) can be calculated according to the initial posture: , for this formula, when the angle When it is equal to 0, the second weight is 1, which means that the target object is in the center line direction of the camera's field of view (the camera coordinate system). direction), the initial position at this time is relatively accurate, but as the angle of deviation Increases, that is, the angle of the target object deviating from the camera's field of view increases, and the accuracy of the initial pose becomes lower and lower. Correspondingly, the second weight Gradually reduce.

[0103] Finally, a pose weight can be determined by obtaining the first weight and the second weight :

[0104] .

[0105] Step S150, performing posture weighted average calculation according to each transformed posture and the associated posture weight to obtain the adjusted posture of the target object.

[0106] Exemplarily, the pose weights of the target object under different cameras are obtained through the above steps. Therefore, for multiple transformed poses of the same target object in the three-dimensional map, the pose weighted average can be performed, that is, each transformed pose is multiplied by its corresponding pose weight p, all products are added together, and then divided by the sum of all pose weights to calculate the comprehensive pose of the target object (i.e., the adjusted pose mentioned above).

[0107] Step S160, controlling the smart home robot to perform the target task according to the adjusted posture.

[0108] Exemplarily, after the adjusted posture of the target object is finally determined, the smart home robot can control its own movement or corresponding action execution according to the adjusted posture information.

[0109] As an optional solution, after step S160, the method further includes:

[0110] If it is recognized that there is no target object in at least one image, it means that the corresponding associated camera has not captured the target object at the current moment. At this time, it can be confirmed that the corresponding associated camera has lost the target object. Optionally, the image of the camera can be directly filtered out. Alternatively, if the associated camera was able to capture the target object before, but the target object did not appear in its field of view after a certain moment, for such a suddenly lost target object, the previous historical image data of the camera can also be used to predict the movement of the target object in the future.

[0111] For example, in some embodiments, the motion information of the target object is calculated based on a preset number of historical transformation positions associated with the target object before it is lost; and then, the transformation position of the target object in the three-dimensional map is predicted based on the motion information. The above-mentioned motion information includes but is not limited to motion speed, motion direction, and motion acceleration. The preset number can be set as needed, such as 8, 10, 13, 15, etc., and is not limited to a single limit here.

[0112] It can be understood that after the camera loses the target object, especially in a short period of time after the loss, the motion state of the target object usually does not change much. The historical transformation posture obtained by multiple frames of images associated with the camera before the target object was lost can be used to estimate the posture of the target object in a short period of time (i.e., the new transformation posture), so that the transformation posture calculated by the camera can continue to be used.

[0113] After predicting the transformed posture associated with the target object in the three-dimensional map based on the motion information, return to the step of performing posture weighted average calculation based on each transformed posture and the associated posture weight to obtain the adjusted posture of the target object.

[0114] However, as the target object is lost for a longer time, the accuracy of the predicted transformed posture will decrease. To this end, the present application also proposes to evaluate the confidence of the predicted transformed posture to further ensure the accuracy of the posture data.

[0115] For example, in some embodiments, after predicting the transformed position associated with the target object in the three-dimensional map according to the motion information, the method further includes:

[0116] The confidence of the predicted transformed position and posture is adjusted according to the loss time of the target object under the associated camera. Then, when the confidence drops below a preset value, the transformed position and posture data of the target object under the associated camera in the three-dimensional map is discarded.

[0117] In this embodiment, the confidence level is adjusted by the loss time, wherein the longer the loss time is. The smaller the confidence level is, for example, the initial confidence level is 100%, and there is no loss; when the loss time is 3 seconds, the confidence level will drop to 80%, when the loss time is 5 seconds, the confidence level will drop to 10%, and so on. It can be understood that the correspondence between the confidence level and the loss time can be set by itself, or determined according to a preset formula, which is not limited here. In addition, the preset value of the confidence level is mainly used to determine whether to use the data of the predicted transformation posture, such as 80%, 70%, etc., which is only an example here. Generally, when the confidence level is less than the preset value, the predicted transformation posture is discarded; otherwise, it is retained.

[0118] As an optional solution, after discarding the transformed position of the target object under the associated camera in the three-dimensional map, the method further includes:

[0119] If it is recognized that the transformed positions of the target objects under each camera in the three-dimensional map are all discarded, it can be determined that the target objects are lost. At this time, the smart home robot 10 should be controlled in time to stop executing the target task.

[0120] The present application uses multiple external cameras and the main camera set on the smart home robot 10 to collect environmental images in real time during the task execution process to identify and locate the target object. This can avoid the situation where the main camera of the smart home robot 10 is blocked or the target object moves out of its field of view, and the task can still be continued with the help of other external cameras.

[0121] For example, when the task is to follow a person, if the person being followed suddenly turns behind a door, the main camera of the smart home robot 10 will not be able to recognize the target object. At this time, the image data of other external cameras in the room can be used to locate the person's position, and then the robot can be controlled to move behind the door.

[0122] For another example, when the task is to grab a bottle on a table, the smart home robot 10 can identify the positions of the table and the bottle in the three-dimensional map through the image data of each camera (here refers to the adjusted position calculated after weighted processing), and then control it to move to the table; then, in the process of grabbing the bottle, assuming that the main camera at the head of the robot is blocked by the robotic arm used for grabbing, the image data of other external cameras can be used to obtain the relative position between the end of the robotic arm (claw) and the bottle, so as to control the moving direction and position of the end of the robotic arm, and then complete the task of grabbing the bottle.

[0123] Figure 7 A schematic diagram of the structure of a smart home robot task execution device according to an embodiment of the present application is shown. Exemplarily, the smart home robot task execution device 100 includes:

[0124] The image acquisition module 110 is used to respond to the task instruction to acquire multiple images collected by each camera at the same time; the camera includes a main camera provided on the smart home robot 10 and multiple external cameras deployed at different positions in the home environment;

[0125] A posture rotation module 120 is used to obtain depth information of the target object in each image in which the target object exists, so as to calculate the initial posture of the target object in the associated camera coordinate system when the target object exists in the multiple images;

[0126] A posture transformation module 130, for obtaining a plurality of transformed postures of the target object in the three-dimensional map according to each initial posture of the target object and the posture of the associated camera in the three-dimensional map of the home environment;

[0127] A weight determination module 140 is used to determine, for each transformed posture, a corresponding posture weight according to the depth information of the associated target object and the initial posture;

[0128] A posture adjustment module 150 is used to perform posture weighted average calculation according to each transformed posture and the associated posture weight to obtain an adjusted posture of the target object;

[0129] The task execution module 160 is used to control the smart home robot 10 to execute the target task according to the adjusted posture.

[0130] It can be understood that the device of this embodiment corresponds to the smart home robot task execution method of the above embodiment, and the optional items in the above embodiment are also applicable to this embodiment, so they will not be repeated here.

[0131] The present application also provides a computer-readable storage medium for storing the computer program used in the above-mentioned smart home robot. For example, the computer-readable storage medium may include but is not limited to: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0132] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or the flow diagram, and the combination of boxes in the structure diagram and / or the flow diagram, can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0133] In addition, the functional modules or units in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0134] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0135] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A method for executing tasks of an intelligent home robot, characterized in that: include: Responding to the task instruction to obtain multiple images collected by each camera at the same time; the camera includes a main camera arranged on the smart home robot and multiple external cameras deployed at different positions in the home environment; In the case where a target object is identified in a plurality of the images, for each image in which the target object exists, obtaining depth information of the target object in the image to calculate an initial position and posture of the target object in an associated camera coordinate system; According to each of the initial poses of the target object and the pose of the associated camera in the three-dimensional map of the home environment, obtaining a plurality of transformed poses of the target object in the three-dimensional map; For each of the transformed postures, determining a corresponding posture weight according to the depth information of the associated target object and the initial posture; Performing posture weighted average calculation according to each of the transformed postures and the associated posture weights to obtain an adjusted posture of the target object; The smart home robot is controlled to perform a target task according to the adjusted posture.

2. The smart home robot task execution method according to claim 1, characterized in that: The determining a corresponding posture weight according to the depth information of the associated target object and the initial posture comprises: Determining a depth estimation value of the target object according to the associated depth information; wherein the depth estimation value is used to characterize the distance between the target object and the camera; Determine a first weight according to the depth estimation value of the target object; wherein the larger the depth estimation value is, the smaller the first weight is; Determining a second weight according to the initial posture of the associated target object; wherein the closer the target object is to the edge of the image, the smaller the second weight is; According to the first weight and the second weight, a corresponding posture weight is calculated.

3. The smart home robot task execution method according to claim 2, characterized in that: The calculation formula of the first weight is: ; The calculation formula of the second weight is: ; The calculation formula of the posture weight is: ; In the formula, is the depth estimation value; is the first weight; is the second weight; is the angle at which the target object deviates from the center line of the camera's field of view; is the pose weight.

4. The smart home robot task execution method according to claim 1, characterized in that: After controlling the smart home robot to perform the target task according to the adjusted posture, the method further includes: If it is identified that the target object does not exist in at least one of the images, it is confirmed that the corresponding associated camera has lost the target object, and the motion information of the target object is calculated according to a preset number of historical transformation postures associated before the target object is lost; wherein the motion information includes motion speed, motion direction and motion acceleration; Predict the transformed posture associated with the target object in the three-dimensional map according to the motion information, and return to execute the step of performing posture weighted average calculation according to each transformed posture and the associated posture weight to obtain the adjusted posture of the target object.

5. The smart home robot task execution method according to claim 4, characterized in that: After predicting the transformed posture associated with the target object in the three-dimensional map according to the motion information, the method further includes: According to the loss time of the target object under the associated camera, the confidence of the predicted transformed posture is adjusted, and when the confidence drops below a preset value, the transformed posture of the target object under the associated camera in the three-dimensional map is discarded.

6. The smart home robot task execution method according to claim 5, characterized in that: After discarding the transformed position of the target object under the associated camera in the three-dimensional map, the method further includes: If it is recognized that the transformed poses of the target object under each camera in the three-dimensional map are discarded, it is determined that the target object is lost, and the smart home robot is controlled to stop executing the target task.

7. The smart home robot task execution method according to any one of claims 1 to 6, characterized in that: The calculating the initial position and posture of the target object in the associated camera coordinate system includes: Constructing an object coordinate system corresponding to the target object based on the geometric features of the target object; Determining the origin of the object coordinate system; According to the depth point cloud data of the target object, align the object coordinate system with the camera coordinate system of the associated camera to obtain a rotation matrix from the object coordinate system to the camera coordinate system; wherein the depth point cloud data is generated based on the depth information of the target object; Obtaining a translation vector according to the position from the origin of the object coordinate system to the camera coordinate system; The initial position of the target object in the camera coordinate system is determined according to the rotation matrix and the translation vector.

8. The method for executing a task of an intelligent home robot according to any one of claims 1 to 6, characterized in that: The three-dimensional map includes environmental image features; after acquiring the multiple images respectively collected by the cameras in real time, and before acquiring the multiple transformed positions of the target object in the three-dimensional map, the method further includes: For each of the images, feature matching is performed between the image and the environmental image features in the three-dimensional map to obtain the position and posture of the associated camera in the three-dimensional map.

9. A smart home robot, characterized in that: include: An actuator is used to execute corresponding actions; a memory storing a computer program; A processor, wherein the processor is used to execute the computer program to implement the smart home robot task execution method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed on a processor, implements the smart home robot task execution method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Target following method and device for robot

    CN108673501A

  • Smart home robot based on embedded WEB and application thereof

    CN109739097A