Method and system for controlling a body-equipped robot to return to a base station, and body-equipped robot

By acquiring the projection outline and orientation of the base station, the embodied robot is controlled to return to the base station, solving the docking failure problem caused by environmental factors and achieving stable docking in complex environments.

CN121028651BActive Publication Date: 2026-02-27WOCAO TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511554831.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-02-27
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

In existing technologies, when a robot returns to a base station when its battery is low, docking may fail due to environmental factors such as base station movement or infrared reflection.

Method used

By obtaining the projection contour of the base station in a two-dimensional plane, the front projection line segment and orientation of the base station are determined, and the embodied robot is controlled to return to the base station.

Benefits of technology

In complex environments, embodied robots can autonomously identify the orientation of base stations, achieve stable docking, and improve the accuracy of returning to base stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121028651B_ABST
    Figure CN121028651B_ABST
Patent Text Reader

Abstract

The application relates to the field of embodied robots, and discloses a control method and system for returning an embodied robot to a base station and the embodied robot. The method comprises: acquiring a base station projection contour of the base station in a two-dimensional plane; determining a front projection line segment of the base station according to the base station projection contour, and determining an orientation of the base station based on the front projection line segment; and controlling the embodied robot to return to the base station based on the orientation. The application identifies the base station contour through the projection of the base station, and then identifies the orientation of the base station based on the base station contour, so that the embodied robot can return to the base station according to the orientation of the base station, and the embodied robot can autonomously identify the orientation of the base station in a complex environment, thereby realizing stable docking of the embodied robot and the base station, and improving the robustness and accuracy of the embodied robot returning to the base station.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of embodied robots, and in particular to a control method and system for an embodied robot to return to a base station, and an embodied robot. BACKGROUND

[0002] In the prior art, a robot automatically returns to a base station for charging when the power is low. This requires the robot to locate the base station to achieve docking between the robot and the base station. In the prior art, the base station position is marked based on a positioning marker or manually marked when the robot is detached from the base station. After the robot returns to the vicinity of the marked position of the base station, the base station emits an infrared signal or a light beam, and the robot receives the signal through an infrared sensor and adjusts its direction and position, and then guides the docking of the base station. Specifically, the robot records the base station position based on a positioning marker (such as an infrared lamp) or a manual marker (such as a visual marker) when it is detached from the base station. When the power is low, the robot will go to this marked position and dock through the infrared signal emitted by the base station in the vicinity of the marked position.

[0003] However, the method of marking the base station position based on a positioning marker or manually in the prior art is prone to cause the robot to fail to correctly find the base station or to fail to dock due to abnormal movement of the base station or environmental factors such as infrared reflection or infrared range limitation. SUMMARY

[0004] Therefore, in view of the technical problem that an embodied robot and a base station cannot dock due to environmental factors, the present application provides a control method and system for an embodied robot to return to a base station, and an embodied robot.

[0005] In a first aspect, the present application provides a control method for an embodied robot to return to a base station, comprising:

[0006] obtaining a base station projection contour of the base station in a two-dimensional plane;

[0007] determining a front projection line segment of the base station according to the base station projection contour, and determining an orientation of the base station based on the front projection line segment;

[0008] controlling the embodied robot to return to the base station based on the orientation.

[0009] In a second aspect, the present application provides a control system for an embodied robot to return to a base station, comprising:

[0010] an obtaining module configured to obtain a base station projection contour of the base station in a two-dimensional plane;

[0011] an orientation determining module configured to determine a front projection line segment of the base station according to the base station projection contour, and determine an orientation of the base station based on the front projection line segment.

[0012] A control module is used to control the android to return to the base station based on the orientation.

[0013] Thirdly, this application provides a cloaked robot, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned control method for the cloaked robot to return to a base station.

[0014] The embodiments of this application have the following beneficial effects:

[0015] This application provides a control method for a robot to return to a base station. The method identifies the outline of the base station by projecting it onto the ground station, and then identifies its orientation based on the outline. This allows the robot to return to the base station according to its orientation, enabling it to autonomously identify the orientation of the base station even in complex environments. Furthermore, when entering an unfamiliar environment for the first time or after the base station has been moved, the robot can still autonomously explore and locate the base station without relying on human markings, initial positioning information, or infrared signals. This achieves stable docking between the robot and the base station and improves the accuracy of the robot's return to the base station. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and therefore should not be considered as a limitation on the scope of protection of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A schematic diagram of the control system for the embodied robot returning to the base station in an embodiment of this application is shown;

[0018] Figure 2 This paper illustrates the first flowchart of a control method for a robot returning to a base station according to an embodiment of this application.

[0019] Figure 3 The second flowchart of the control method for the embodied robot to return to the base station in an embodiment of this application is shown;

[0020] Figure 4 The third flowchart of the control method for the embodied robot to return to the base station in an embodiment of this application is shown;

[0021] Figure 5 This invention illustrates a vector diagram between two adjacent points on the projected outline of a base station in an embodiment of this application.

[0022] Figure 6A vertical line diagram of a front projection line segment of a base station in the embodiment of the present application is shown;

[0023] Figure 7 A fourth flow diagram of a control method for a humanoid robot returning to a base station in the embodiment of the present application is shown;

[0024] Figure 8 Another structural diagram of a control system for a humanoid robot returning to a base station in the embodiment of the present application is shown. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application.

[0026] The components of the embodiments of the present application generally described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0027] Hereinafter, the terms "include", "have", and their conjugates in various embodiments of the present application merely mean to indicate that there are certain features, numbers, steps, operations, elements, components, or combinations thereof, and should not be construed as excluding the presence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations thereof.

[0028] In addition, the terms "first", "second", "third", and the like are used only to distinguish descriptions, and should not be understood as indicating or implying relative importance.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present application belong. The terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning that is the same as in the context of the relevant art and should not be interpreted in an idealized or overly formal sense, unless clearly defined in various embodiments of the present application.

[0030] Some embodiments of the present application will be described below in detail with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.

[0031] Embodied Robot in this application refers to a robot with a physical body (entity) and realizes intelligent behavior through real-time perception, interaction and action with the real environment. Its core idea comes from the theory of Embodied Intelligence, that is, intelligence not only depends on algorithms and data processing, but also needs to learn and evolve through dynamic interaction between body and environment. Among them, the physical entity refers to the fact that the embodied robot has a real body (such as a mechanical arm, a mobile chassis, a sensor, etc.), which can act in the physical world like humans or animals (such as walking, grabbing, obstacle avoidance, etc.). In addition, it can also perceive the environment in real time through visual, tactile, auditory, force, etc. sensors, and adjust behavior according to feedback.

[0032] Figure 1 A structural schematic diagram of an embodied robot 10 of an embodiment of the present application is shown.

[0033] Exemplarily, the embodied robot 10 includes a processor 11 and a memory 12; the memory 12 stores a computer program, and the processor 11 executes the computer program, so that the embodied robot 10 executes the control method for the embodied robot to return to the base station in the following embodiments.

[0034] The processor 11 can be an integrated circuit chip with a signal processing capability. The processor 11 can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), and a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or at least one of them. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can implement or execute computer programs to implement the disclosed methods, steps and logic block diagrams in the embodiments of the present application.

[0035] The memory 12 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory 12 is configured to store a computer program. After receiving an execution instruction, the processor 11 can execute the computer program accordingly.

[0036] Further, the embodied robot 10 can further include a perception unit 13 and an execution mechanism 14. The perception unit 13 is configured to detect environmental information (such as target objects, obstacles, etc.) of the embodied robot and motion state information (such as motion pose, turning, acceleration, etc.) of the embodied robot. The execution mechanism 14 is configured to perform corresponding actions.

[0037] For example, for obtaining the environmental information, the perception unit 13 can include detection sensors arranged on the body of the embodied robot 10, such as a vision sensor, a laser radar, etc. The vision sensor can include a camera, such as a conventional RGB (red, green, and blue) camera, a depth camera, or a stereo camera, etc. For example, for a cleaning robot, the vision sensor can be used to collect image data in the current environment space, so as to obtain information of the front obstacles in the environment space, and further be used for target object recognition, positioning, etc. The laser radar can be used to obtain point cloud data of the cleaning robot when moving in the current environment space. In some optional embodiments, the detection sensor can further include an infrared sensor, an ultrasonic sensor, etc. For obtaining the motion state information, the perception unit 13 can further include an inertial measurement unit (IMU), a position encoder, etc. The IMU is configured to collect IMU data (acceleration and angular velocity) of the embodied robot during motion.

[0038] It should be understood that the embodied robot 10 described above can include, but is not limited to, a sweeping robot (also referred to as a cleaning robot), a mop robot driven by a sweeping robot (i.e., a sweeping and mopping integrated robot), a food delivery robot, an autonomous driving object carrying robot, a companion robot, a service robot, etc. It should be understood that the existence forms of the above-mentioned various robots are not limited, for example, the robot can be a wheeled robot, or can be a humanoid robot with two legs or a multi-legged robot, etc.

[0039] Based on the above embodiment of the somatotropic robot, the embodiment of the present application further provides a control method for returning the somatotropic robot to the base station.

[0040] Referring to Figure 2 , the control method for returning the somatotropic robot to the base station comprises the following steps S210-S230:

[0041] S210, obtaining a base station projection contour of the base station in a two-dimensional plane.

[0042] The base station projection contour refers to the outermost contour line of the base station projected in the two-dimensional plane.

[0043] As an optional embodiment, the two-dimensional plane can be a two-dimensional environment image, and the somatotropic robot is provided with a body camera, and in the process of returning to the base station, the two-dimensional environment image containing the base station can be collected through the body camera.

[0044] Demonstratively, in the process of returning the somatotropic robot to the base station, in the case that the somatotropic robot has been in a preset range in which the base station can be detected, or in the case that the base station is detected through the body camera of the somatotropic robot, the base station projection contour can be obtained from the two-dimensional environment image containing the base station.

[0045] It can be understood that the above-mentioned preset range can be the area range of the initial position of the base station or the pre-marked position observed before returning to the base station, for example, the circumferential area obtained with the initial position as the center and with a first preset radius, wherein the value of the first preset radius can be set according to actual needs, which is not limited here.

[0046] As an optional embodiment, the somatotropic robot is provided with a laser radar, and in the process of returning to the base station, the point cloud data of the base station can be collected through the laser radar, and the point cloud data is projected to the two-dimensional plane, and the base station projection contour of the base station in the two-dimensional plane is obtained.

[0047] As an optional embodiment, in the process of returning the somatotropic robot to the base station, as Figure 3 shown, the process of obtaining the base station projection contour of the base station in the two-dimensional plane specifically comprises steps S211-S215:

[0048] S211, collecting perception data through the perception sensor on the somatotropic robot, and detecting the base station in the perception data to obtain a detection frame of the base station.

[0049] The perception sensor can be at least one of a camera, a laser radar and an inertial measurement unit, and correspondingly, the perception data can be at least one of image data, point cloud data and IMU data.

[0050] In this embodiment, the perception sensor includes a camera, and the perception data includes a first environment image of the base station, i.e., the base station is in the first environment image.

[0051] Optionally, the collected environment image can be detected by a target detection model (e.g., YOLO or SSD or EfficientDet), and if the base station is detected in the collected environment image, the environment image is taken as the first environment image, and the detection box of the base station in the first environment image is further obtained by the target detection model.

[0052] S212, extracting depth information in the detection box.

[0053] In an example, the depth estimation can be performed only in combination with the first environment image to obtain the depth information. Specifically, the first environment image containing the base station is first subjected to two-dimensional target detection to obtain a detection box of the base station, and then the depth estimation is performed on the detection box of the base station to detect the depth information of the base station. In the process of depth estimation, a corresponding target detection algorithm or target detection network (such as a 2D target detection network), a depth estimation algorithm or a depth estimation network (such as a monocular depth estimation network) can be used according to actual needs, which is not limited in this embodiment.

[0054] S213, converting the depth information into initial three-dimensional point cloud data.

[0055] Exemplarily, after obtaining the depth information, the depth information is converted to a three-dimensional space to obtain initial three-dimensional point cloud data. For example, according to the depth information, each pixel point in the first environment image is correspondingly mapped to a point in the three-dimensional space in the camera coordinate system, i.e., the initial three-dimensional point cloud data is obtained.

[0056] In a preferred embodiment, after obtaining the depth information, the depth information is converted to a three-dimensional space to obtain the candidate three-dimensional point cloud data, and the candidate three-dimensional point cloud data is subjected to noise reduction processing to reduce data redundancy, remove noise points obviously deviating from the main area, and improve data quality to obtain the initial three-dimensional point cloud data. For example, the outlier detection algorithm is used to perform noise reduction processing on the candidate three-dimensional point cloud data to obtain the initial three-dimensional point cloud data. The outlier detection algorithm used in this embodiment includes, but is not limited to, one or more combinations of Gaussian filtering algorithm and voxel network filtering algorithm.

[0057] Since the base station is usually placed on the ground, after the depth information is converted into the initial three-dimensional point cloud data, the initial three-dimensional point cloud data contains three-dimensional point cloud data belonging to the ground because the detection box contains the base station and the ground. In order to improve the accuracy of obtaining the projection contour of the base station, the target three-dimensional point cloud data for determining the projection contour of the base station can be further screened from the initial three-dimensional point cloud data.

[0058] S214: Identify and remove planar point cloud data belonging to the ground from the initial 3D point cloud data to obtain the target 3D point cloud data.

[0059] Optionally, ground point cloud and non-ground point cloud can be identified in the initial 3D point cloud data, and then planar point cloud data belonging to the ground (i.e., ground point data) can be removed. For example, if the normal vector of the plane formed by the extracted planar point cloud data is close to the vertical direction, it is considered that the plane is close to the horizontal, and the planar point cloud data can be confirmed to belong to the ground and removed from the initial 3D point cloud data.

[0060] In one implementation, the Random Sample Consensus Algorithm (RANSAC) can be used to identify planar point cloud data belonging to the same plane in the initial 3D point cloud data and calculate the normal vector of the plane; if the Z-axis component of the normal vector falls within the target range, the plane is determined to be the ground, and the planar point cloud data corresponding to the ground is removed from the initial 3D point cloud data.

[0061] The target range is a range defined based on the characteristics of the Z-axis component of the ground; for example, the target range may be that the Z-component of the normal vector is between 0.95 and 1.05, or the Z-component is between -0.95 and -1.05.

[0062] Specifically, based on the random sample consensus algorithm, three non-collinear points are randomly selected from the initial 3D point cloud data each time, and these points are fitted into a plane to identify planar point cloud data belonging to the same plane in the initial 3D point cloud data. The normal vector of this plane is calculated. If the Z-axis component of the normal vector of this plane falls within the target range, the plane is determined to be the ground. For example, if the Z-axis component of the normal vector of this plane is close to 1 or -1, it means that the normal vector is close to the vertical direction, and the plane can be determined to be the ground. Then, the planar point cloud data corresponding to the ground is removed from the initial 3D point cloud data. The above process is repeated, iteratively updating the three non-collinear points to traverse multiple planes and their normal vectors in the initial 3D point cloud data, in order to remove all ground points in the initial 3D point cloud data and obtain the target 3D point cloud data.

[0063] S215, Obtain the base station projection contour in a two-dimensional plane based on the target three-dimensional point cloud data.

[0064] Specifically, the base station projection outline can be obtained by projecting the target's three-dimensional point cloud data onto a two-dimensional plane.

[0065] In one implementation, the target three-dimensional point cloud data can be projected onto a two-dimensional plane to obtain the projection area of ​​the base station in the two-dimensional plane; convex hull detection is performed on the two-dimensional point cloud data within the projection area to obtain the contour of the base station projection.

[0066] Convex hull detection refers to analyzing the set of contour points in an image using convex hull algorithms to determine the smallest convex polygon containing all points. Examples of such algorithms include, but are not limited to, the Graham Scan algorithm, the Jarvis March algorithm, the QuickHull algorithm, and the Andrew monotonic chain algorithm.

[0067] In the above steps, the target 3D point cloud data can be projected onto a 2D plane by vertical projection or arbitrary planar projection. Then, the convex hull algorithm is used to perform convex hull detection on the 2D point cloud in the projection area of ​​the 2D plane. From the original 2D point cloud data, the points located on the outermost edge and forming the smallest convex polygon (i.e., convex hull vertices) are selected. Then, the convex polygon obtained by connecting these convex hull vertices is used as the base station projection contour.

[0068] As an optional implementation, this embodiment can also simplify the convex polygon obtained in the above steps, so as to effectively remove redundant points in the convex polygon while retaining the key geometric features of the base station's outer contour, thereby reducing the complexity of subsequent processing of the base station's projected contour.

[0069] For example, a trajectory thinning algorithm can be used to simplify the convex polygon, and the simplified convex polygon can be used as the projection contour of the base station. This trajectory thinning algorithm includes, but is not limited to, any one or more of the following algorithms: Douglas-Peucker algorithm, Top-Down Time-Ratio algorithm, Bellman algorithm, etc.

[0070] S220, determine the front projection line segment of the base station based on the base station projection outline, and determine the orientation of the base station based on the front projection line segment.

[0071] The front of the base station is the side where the base station docks with the embodied robot; the front projection line segment of the base station is the outline line segment of the front of the base station in the base station projection outline.

[0072] It should be noted that, since base stations vary in length and width and are usually placed against walls or other objects, the environmental images collected by the embodied robot can only detect the side or front of the base station. Therefore, in the base station projection outline obtained based on the environmental image, the longest line segment is the front projection line segment of the base station.

[0073] The orientation of a base station refers to the orientation of its front. In this article, the orientation of a base station refers to the orientation of its front in the camera coordinate system.

[0074] In some examples, such as Figure 4 As shown, the implementation process of "determining the front projection line segment of the base station based on the base station projection outline" in step S220 may include the following steps S221~S223:

[0075] S221, calculate the vector formed by every two adjacent points on the base station projection contour, and calculate the vector angle between every two adjacent vectors.

[0076] In this context, each pair of adjacent points on the base station projection profile can be each pair of adjacent convex hull vertices on the base station projection profile.

[0077] For example, in one embodiment, if two adjacent vectors are considered as a group of adjacent vectors, the angle between the vectors in the group can be obtained by calculating the inverse cosine of the two adjacent vectors (denoted as α). The formula for calculating the angle between the vectors is as follows:

[0078] ;

[0079] In the formula, the coordinates of the first vector among two adjacent vectors are: The coordinates of the second vector among two adjacent vectors are: ,in,( )and( ) represents the coordinates of two adjacent points used to form the first vector, ( )and( ) represents the coordinates of two adjacent points used to form the second vector.

[0080] S222, the intersection point between two vectors whose included angle is greater than a preset threshold is taken as the inflection point.

[0081] As an example, for the included angle of a set of adjacent vectors, if the included angle is greater than a preset threshold, the intersection point between the two adjacent vectors is taken as the inflection point. The value of the preset threshold is not limited here; for example, the preset threshold could be 80 degrees, 90 degrees, 110 degrees, etc.

[0082] It should be noted that the intersection point is the common point between each pair of adjacent points corresponding to two adjacent vectors, and this intersection point is also located on the projection outline of the base station; for example, points 1 and 2 are two adjacent points, points 2 and 3 are two adjacent points, calculate vector 1 between points 1 and 2, and vector 2 between points 2 and 3. Vector 1 and vector 2 are two adjacent vectors. If the angle between vector 1 and vector 2 is greater than a preset threshold, then point 2 (i.e., the intersection point) is taken as the inflection point.

[0083] S223, the base station projection outline is divided into several line segments based on multiple inflection points, and the longest line segment among the several line segments is taken as the front projection line segment of the base station.

[0084] The base station's projected outline is segmented based on multiple inflection points, and the longest line segment is selected as the frontal projection line segment of the base station. For example, as shown...Figure 5 As shown, the inflection point 1 and the inflection point 2 divide the base station projection profile into the line segment 1, the line segment 2 and the line segment 3, wherein the line segment 3 is the longest one, i.e. the line segment 3 can be taken as the front projection line segment of the base station.

[0085] As an optional solution, considering that the longest line segment (i.e. the front projection line segment) in the base station projection profile found in the above steps is not necessarily a straight line, in order to realize accurate identification of the base station front projection profile and further improve the accuracy of determining the base station orientation based on the base station front projection profile, the embodiment can further perform fitting processing on the obtained front projection line segment, so that the front projection line segment gradually approximates the base station front projection profile, and then determine the base station orientation based on the fitted straight line.

[0086] It can be understood that, since the embodied robot aligns to the center line position of the base station (i.e. the midpoint of the base station front projection profile) when returning to the base station, and then in the process of determining the base station orientation based on the fitted straight line, the midpoint of the straight line can be taken as the origin of the base station, and the direction of the perpendicular line corresponding to the midpoint of the straight line can be taken as the orientation of the base station.

[0087] Therefore, after determining the front projection line segment of the base station, the orientation of the base station is determined according to the front projection line segment. In some examples, the implementation process of "determining the orientation of the base station based on the front projection line segment" in step S220 includes:

[0088] The least square method is used to perform straight line fitting on the front projection line segment to obtain a straight line corresponding to the front projection line segment; according to the straight line and the midpoint of the straight line, a perpendicular line of the straight line is determined, and the direction of the perpendicular line is determined as the orientation of the base station.

[0089] Exemplarily, assuming that the expression of the straight line equation corresponding to the front projection line segment of the base station obtained by fitting is set as: y=kx+b; in the equation, (x, y) is the coordinate of any point on the straight line, k is the slope, and b is the constant. The perpendicular line of the straight line passes through the midpoint of the straight line. According to the expression of the straight line equation, the perpendicular line equation of the midpoint of the straight line equation can be obtained as:

[0090] ;

[0091] In the equation, is the midpoint of the front projection line segment.

[0092] It can be understood that the result of the perpendicular line of the straight line is shown by the arrow in Figure 6 , and the direction of the perpendicular line passing through the midpoint of the straight line is the orientation of the base station.

[0093] S230, based on the orientation, controlling the embodied robot to return to the base station.

[0094] After the orientation of the base station is determined, the embodied robot is controlled to move towards the base station in the orientation direction until the embodied robot returns to the base station. It should be noted that since the base station orientation is the direction of the front of the base station in the camera coordinate system, the base station orientation needs to be converted into the direction of the base station on the map when guiding the embodied robot to move by the base station orientation.

[0095] In an embodiment, the embodied robot can be controlled to return to the base station in combination with the map and the base station orientation.

[0096] Exemplarily, the direction of the perpendicular line in the two-dimensional plane rectangular coordinate system is first determined, and then the direction of the perpendicular line in the plane rectangular coordinate system is converted into the direction of the base station in the camera coordinate system according to the intrinsic parameters of the camera, to obtain the direction of the base station in the camera coordinate system. The direction of the base station in the camera coordinate system is converted into the direction of the base station in the embodied robot coordinate system by using the homogeneous coordinate transformation matrix, and then converted into the direction of the base station in the map coordinate system according to the extrinsic parameters of the camera. The movement path of the embodied robot is planned in combination with the map and the direction of the base station on the map to guide the embodied robot to return to the base station. The map is a map constructed by the embodied robot in advance.

[0097] In another embodiment, the distance information between the base station and the embodied robot can also be calculated according to the depth information of the base station, and the path of the embodied robot is planned in combination with the map, the orientation of the base station on the map and the distance information to guide the embodied robot to return to the base station.

[0098] It can be understood that in the process of returning the embodied robot to the base station, the base station orientation is determined by recognizing the projection outline of the base station in the environment image containing the base station to control the embodied robot to return to the base station according to the orientation, so that the adaptive base station orientation recognition is realized without the base station position marker, and the efficiency and accuracy of the robot returning to the base station are improved.

[0099] For the specific process of extracting the depth information in the detection frame in step S212, the application further provides another way of acquiring depth information.

[0100] In another embodiment, if an inertial measurement unit (IMU) is also provided on the embodied robot, the IMU data acquired by the inertial measurement unit during the process of returning the embodied robot to the base station can also be fused with the first environment image to acquire the depth information in the first environment image. It can be understood that by fusing the IMU data and the two-dimensional environment image, the depth information of the base station is calculated by using the visual features in the image and the motion information in the IMU data, which can further improve the estimation accuracy of the depth information and reduce the scale ambiguity in the depth estimation.

[0101] Specifically, the sensing sensor also includes an inertial measurement unit (IMU), and the sensing data includes IMU data. In step S211, the sensing data is acquired through the sensing sensor, including: simultaneously acquiring IMU data through the inertial measurement unit on the android and acquiring a first environmental image containing the base station through the camera on the android; in step S212, the depth information within the detection box is extracted, including: inputting the IMU data and the first environmental image into a depth information estimation network model to output depth information. The depth information estimation network model can be pre-trained.

[0102] In the first example, an IMU photometric loss function and a cross-sensor photometric consistency loss function can be constructed, and the depth information estimation network model can be obtained by jointly training based on these two functions.

[0103] Exemplary, based on sample IMU data and sample environmental images including base stations, an IMU photometric loss function and a cross-sensor photometric consistency loss function are constructed; based on the IMU photometric loss function and the cross-sensor photometric consistency loss function, a scale-aware depth network and a self-motion network are jointly trained to obtain a depth information estimation network model.

[0104] It should be noted that the IMU photometric loss function uses motion information (such as rotation and translation) provided by the sample IMU data to predict changes between environmental images (i.e., predict self-motion changes), and uses photometric error to measure the difference between this change and the actual sample environmental image.

[0105] Specifically, first use a deep network (denoted as...) Predicted depth value Combining the rotation estimation from the current frame to the adjacent frame in the sample environment image (denoted as...) Translation and shift estimation (denoted as) ), to obtain the pixels in the current frame In adjacent frames (which could be the previous frame) δ =-1) or the next frame ( δ =1) Pixel intensity after reprojection (denoted as ) Then calculate the pixels in the current frame. pixel intensity (denoted as) The pixel intensity of the two frames before and after (i.e.) The photometric loss of all pixels is calculated by summing the minimum photometric errors of all pixels and taking the average value, resulting in the final IMU photometric loss function (denoted as ). ).

[0106] The mathematical expression for the IMU photometric loss function is as follows:

[0107]

[0108] In the formula:

[0109] ;

[0110] in, Indicates photometric error; These are the weights of the Structural Similarity Index (SSIM). Represents the structural similarity index; Indicates the source from the sample environment image arrive The self-motion estimation results include rotation and translation estimation results, which are obtained by fusing IMU pre-integration and a deep network in the camera coordinate system using a Kalman filter (EKF) framework. Obtained from predicted self-movement; The pixels in the current frame pixel intensity, yes The pixel intensity after reprojection in adjacent frames; This means for each pixel Calculate the time from the current frame to the previous frame respectively. δ =-1) and the current frame to the next frame ( δ =1) of the photometric error, and take the minimum of the two;

[0111] This represents the summation of the minimum photometric errors of all pixels, followed by the average value, to obtain the final IMU photometric loss function; K and N are the camera intrinsic parameters and the number of pixels used, respectively. and These are sample environment images. The pixel coordinates in the deep network and their composition Predicted depth value.

[0112] It is understood that this embodiment utilizes the self-motion estimation results (i.e., the rotation estimation results and translation estimation results mentioned above) obtained by fusing IMU pre-integration and self-motion network prediction from the Kalman filter (i.e., EKF) framework to construct a scale-aware IMU photometric loss function. In this way, the photometric loss provides dense supervision signals for the scale-aware deep network and self-motion network, so as to jointly train the deep network and the self-motion estimation network.

[0113] After constructing the IMU photometric loss function, this embodiment also constructs a cross-sensor photometric consistency loss function to align the IMU pre-integration and the self-motion estimation results predicted by the self-motion network in the aforementioned IMU photometric loss function. Specifically, this cross-sensor photometric consistency loss function utilizes the difference between the self-motion estimation obtained from IMU pre-integration and the self-motion estimation predicted by the self-motion network, measuring this difference through photometric error.

[0114] Exemplary, the process of constructing this cross-sensor photometric consistency loss function includes:

[0115] Rotation matrix predicted by its own motion network Translation vector And the rotation estimate obtained by IMU pre-integration Translation estimation Calculate the pixels in adjacent frames (which can be the previous frame) respectively. δ =-1) or the next frame ( δ =1) reprojection and .

[0116] Then, the photometric error is calculated for two cases for adjacent frames, and the minimum error is selected as the final loss value. Finally, the minimum photometric error of all pixels is averaged to obtain the final cross-sensor photometric consistency loss (denoted as ). ).

[0117] The mathematical expression for the cross-sensor photometric consistency loss function is as follows:

[0118] In the formula: , These are the reprojected pixel intensities predicted by their own motion network and pre-integrated by the IMU in adjacent frames, respectively.

[0119] Since adjacent frames can be the previous frame ( δ =-1) or the next frame ( δ =1),

[0120] This means that the photometric error is calculated separately for two cases between adjacent frames, and the minimum error is selected as the final loss value.

[0121] This represents the average of the minimum photometric errors across all pixels, resulting in the final cross-sensor photometric consistency loss function.

[0122] It can be understood that the IMU photometric loss function and the cross-sensor consistency loss function constructed based on the above steps can be used to jointly train the scale-aware depth network and the ego-motion network to obtain a depth information estimation network model; and subsequently, the depth information estimation network model can be used to fuse the IMU data and the first environment image to obtain the depth information in the first environment image.

[0123] In the second example, on the basis of the joint training of the IMU photometric loss function and the cross-sensor consistency loss function, a visual photometric loss function (denoted as ), a disparity smoothness loss function (denoted as ), and a weak L2 norm loss function based on velocity and gravity prediction (denoted as ) can be further combined for joint training to obtain the depth information estimation network model. For example, in an embodiment, the mathematical expressions of the joint training of the above loss functions are as follows:

[0124] ;

[0125] In the formula, { , , , } are loss weights of the loss functions.

[0126] After obtaining the depth information estimation network model by training the above loss functions, the depth information can be obtained by using the depth information estimation network model to combine the IMU data for depth estimation of the first environment image containing the base station.

[0127] It can be understood that by fusing the IMU data in the first environment image, the IMU data can be used to solve the problem of visual scale uncertainty in image depth estimation to restore the real scale, thereby providing more accurate depth information estimation. In addition, in a weak texture, dynamic scene, or environment with severe light changes, vision is easy to fail, and by fusing the IMU data, the influence of motion blur on visual matching can be reduced to ensure the accuracy of depth information estimation, so that subsequent base station pose estimation based on the depth information can be implemented with higher precision, and thus in the case that the infrared signal is unreliable or missing, the stable docking of the embodied robot and the base station can also be achieved according to the base station pose estimation result. Obviously, the embodiment not only improves the robustness and accuracy of the embodied robot returning to the base station, but also provides an important reference for the autonomous navigation of the embodied robot in the future dynamic environment.

[0128] Further, as an optional implementation, before the embodied robot returns to the base station, the embodiment can continuously update the base station position by the embodied robot long-term acquiring environment images containing the base station in the process of performing the task, and pre-mark the base station position on the map of the current environment space, so that when the instruction of returning to the base station is received, the embodied robot is guided to move towards the latest pre-marked position. When the embodied robot has been in a preset range containing the pre-marked position, the embodied robot can acquire the first environment image containing the base station in the preset range through the camera, and then guide the embodied robot to return to the base station according to the above steps S210-S230.

[0129] By pre-marking the base station position before the embodied robot returns, when the base station position moves, the embodied robot can also guide it to return to the base station based on the long-term observed base station position, instead of only relying on the initial position calibration result of the base station to cause the embodied robot to blindly search for the base station in the return process.

[0130] Exemplarily, as shown in Figure 7 The above pre-marking of the base station position so that when the embodied robot returns to the base station, the process of guiding the embodied robot to move towards the pre-marked position includes the following steps S710-S730:

[0131] S710, acquiring a plurality of second environment images containing the base station through the camera in the process of the embodied robot performing the task.

[0132] Among them, the plurality of second environment images can be acquired by the embodied robot at different time points or positions. Exemplarily, if the embodied robot is a sweeping robot, the task performed by the embodied robot can be a whole house cleaning task.

[0133] S720, determining the target position of the base station in the pre-constructed map according to the plurality of second environment images.

[0134] Among them, the pre-constructed map is constructed based on the environment data acquired at different positions in the current environment space before the embodied robot performs the task in the process of moving in the current environment space; the map construction method is not limited here. Optionally, the environment data includes but is not limited to one or more of environment images, laser point cloud data and IMU data. Optionally, for each second environment image, the position of the base station in the map can be determined according to the second environment image, a plurality of positions are obtained, and the plurality of positions are averaged to obtain the target position.

[0135] In an embodiment, for each second environment image, a preset detection model is adopted to detect the base station in the second environment image to obtain a plurality of detection positions of the base station in a plurality of second environment images and a confidence degree corresponding to each detection position, wherein the detection position in the second environment image is a position coordinate of the base station in the environment image. The preset detection model includes but is not limited to YOLO, Faster, R-CNN or SSD model, and the confidence degree reflects the probability that the detection position is the position coordinate of the base station.

[0136] The clustering algorithm is used to cluster the detection positions in all the environment images obtained in the above steps to classify the detection positions close in space into the same class, thereby obtaining a plurality of clustering results, wherein each clustering result includes at least one detection position. The clustering algorithm includes but is not limited to any one or more of K-Means, DBSCAN, Mean Shift and the like, and the present embodiment is not limited thereto.

[0137] For each clustering result, the number of detection positions in the clustering result is counted, wherein the number of each detection position reflects the number of times the detection position is detected. The detection position with the highest confidence degree is selected from the clustering result with the largest number as a credible detection position, which is the position coordinate of the most likely real base station.

[0138] It can be understood that a plurality of detection positions of the base station can be detected in a plurality of environment images, wherein some of the detection positions can be false detection or noise. Since the position of the real base station appears repeatedly in a plurality of environment images, the false detection or noise position usually appears in only a few environment images. Therefore, by clustering a plurality of environment images, the position coordinate of the real base station can be determined from the environment images to eliminate the noise in the environment images and avoid false detection, thereby improving the stability and accuracy of the credible detection position, so that the target position of the base station can be accurately determined from the map based on the credible detection position in the environment image.

[0139] In an embodiment, the credible detection position in the environment image can be converted into the target position on the map by twice coordinate system conversion. Specifically, the credible detection position is first converted to the camera coordinate system according to the intrinsic parameters of the camera, and then converted from the camera coordinate system to the map coordinate system according to the extrinsic parameters of the camera to obtain the target position corresponding to the credible detection position in the map. The target position is recorded in the map as a pre-marked position.

[0140] S730, moving to the target position according to the map.

[0141] Exemplarily, after obtaining the pre-marked position of the base station, a path planning is generated in combination with the map and the target position on the map to guide the robot to return to the base station.

[0142] It can be understood that the base station position is pre-marked before returning to the base station, and then the embodied robot is guided to move according to the pre-marked position to quickly approach the approximate area of the base station, thereby avoiding the low efficiency caused by blind search of the base station position in the process of starting to return to the base station.

[0143] In addition, as shown in Figure 8 The embodiment of the present application also provides a control system for returning an embodied robot to a base station, which comprises:

[0144] The acquisition module 810 is configured to acquire a base station projection contour of the base station in a two-dimensional plane.

[0145] The orientation determination module 820 is configured to determine a front projection line segment of the base station according to the base station projection contour, and determine an orientation of the base station based on the front projection line segment.

[0146] The control module 830 is configured to control the embodied robot to return to the base station based on the orientation.

[0147] It can be understood that the control system for returning an embodied robot to a base station in the embodiment corresponds to the control method for returning an embodied robot to a base station in the above-mentioned embodiment, and the optional items in the above-mentioned embodiment are also applicable to the embodiment, so the description is not repeated here.

[0148] The present application also provides a computer storage medium for storing the computer program used in the above-mentioned embodied robot. The computer storage medium can be a readable storage medium, a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium can include but is not limited to a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, in alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0150] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0151] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a robot (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0152] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A control method for a body-equipped robot to return to a base station, characterized by, The method comprises the following steps: obtaining a base station projection contour of the base station in a two-dimensional plane based on a first environment image containing the base station; the base station projection contour is the outermost contour line of the base station projected in the two-dimensional plane; calculating a vector formed by every two adjacent convex hull vertices on the base station projection contour, calculating the vector included angle between every two adjacent vectors; taking the intersection point between the two vectors with the vector included angle greater than a preset threshold as a inflection point; dividing the base station projection contour into several line segments based on a plurality of the inflection points; taking the longest line segment in the several line segments as the front projection line segment of the base station, and determining the orientation of the base station based on the front projection line segment; controlling the embodied robot to return to the base station based on the orientation.

2. The somatic robot base-return control method according to claim 1, wherein The method for determining the orientation of the base station based on the front projection line segment comprises the following steps: performing linear fitting on the front projection line segment by using the least square method to obtain a straight line corresponding to the front projection line segment; determining a perpendicular line of the straight line according to the straight line and the midpoint of the straight line; wherein the perpendicular line passes through the midpoint of the straight line; determining the direction of the perpendicular line as the orientation of the base station.

3. The somatic robot base return control method according to claim 1, wherein The method for obtaining the base station projection contour of the base station in a two-dimensional plane based on a first environment image containing the base station comprises the following steps: collecting perception data through a perception sensor on the embodied robot, and detecting the base station in the perception data to obtain a detection box of the base station; the perception sensor comprises a camera, and the perception data comprises the first environment image containing the base station; extracting depth information in the detection box; converting the depth information into initial three-dimensional point cloud data; identifying and removing planar point cloud data belonging to the ground in the initial three-dimensional point cloud data to obtain target three-dimensional point cloud data; obtaining the base station projection contour of the base station in a two-dimensional plane according to the target three-dimensional point cloud data.

4. The somatic robot base-return control method according to claim 3, wherein The method for identifying and removing planar point cloud data belonging to the ground in the initial three-dimensional point cloud data comprises the following steps: identifying planar point cloud data belonging to the same plane in the initial three-dimensional point cloud data based on a random sample consensus algorithm, and calculating the normal vector of the plane; if the Z-axis component of the normal vector falls within a target range, it is determined that the plane is the ground, and the planar point cloud data corresponding to the ground is removed from the initial three-dimensional point cloud data.

5. The somatic robot base-return control method according to claim 3, wherein The method for obtaining the base station projection contour of the base station in a two-dimensional plane according to the target three-dimensional point cloud data comprises the following steps: projecting the target three-dimensional point cloud data to the two-dimensional plane to obtain a projection area of the base station in the two-dimensional plane; performing convex hull detection on the two-dimensional point cloud data in the projection area to obtain the base station projection contour.

6. The somatic robot base return control method according to claim 3, wherein Before the step of collecting perception data through the perception sensor on the embodied robot, the method further comprises the following steps: collecting a plurality of second environment images containing the base station through the camera during the process of the embodied robot performing a task; determining a target position of the base station in a pre-constructed map according to a plurality of the second environment images; moving to the target position according to the map.

7. The somatic robot base return control method according to claim 6, wherein The target position of the base station in a pre-constructed map is determined according to a plurality of the second environment images, and the method comprises the following steps: For each of the second environment images, a preset detection model is used to detect the detection position of the base station in the second environment image in the second environment image, thereby obtaining a plurality of detection positions corresponding to a plurality of the second environment images; Each of the detection positions is clustered to obtain a plurality of clustering results, and at least one of the detection positions is included in the clustering result; For each of the clustering results, the number of the detection positions in the clustering result is counted; The detection position with the highest confidence in the clustering result with the largest number is selected as a credible detection position; The target position is determined according to the credible detection position. 8.The control method of the embodied robot returning to the base station according to claim 3, wherein the perception sensor further comprises an inertial measurement unit, and the perception data further comprises IMU data. The perception data is collected through the perception sensor on the embodied robot, and the method comprises the following steps: At the same time, the IMU data is obtained through the inertial measurement unit on the embodied robot, and the first environment image containing the base station is collected through the camera on the embodied robot; The depth information in the detection frame is extracted, and the method comprises the following steps: The IMU data and the environment image are input into a depth information estimation network model to output the depth information; The depth information estimation network model is obtained by the following process: Based on the sample IMU data and the sample environment image containing the base station, an IMU photometric loss function and a cross-sensor photometric consistency loss function are constructed; Based on the IMU photometric loss function and the cross-sensor photometric consistency loss function, the scale perception depth network and the self-motion network are jointly trained to obtain the depth information estimation network model.

9. A control system for a body-attached robot returning to a base station, characterized by, The system comprises: An acquisition module is configured to acquire a base station projection contour of the base station in a two-dimensional plane based on a first environment image containing the base station; the base station projection contour is the outermost contour line of the base station projected in the two-dimensional plane; A direction determination module is configured to calculate a vector formed by each adjacent two convex vertices of the base station projection contour, calculate a vector included angle between each adjacent two vectors, take an intersection point between two vectors with a vector included angle greater than a preset threshold as a turning point, divide the base station projection contour into a plurality of line segments based on a plurality of turning points, take a longest line segment in the plurality of line segments as a front projection line segment of the base station, and determine the direction of the base station based on the front projection line segment; A control module is configured to control the embodied robot to return to the base station based on the direction.

10. A humanoid robot, characterized by The embodied robot comprises a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the control method for the embodied robot returning to the base station according to any one of claims 1-8.

Citation Information

Patent Citations

  • Sweeper base station centerline detection method, device and equipment and storage medium

    CN117911466A

  • Self-moving device and recharging method thereof

    CN120715921A