Mobile Robot Control Method, Computer-Implemented Storage Medium, and Mobile Robot

The image-based mobile robot control method uses a single camera to control robot movement by extracting features from images of desired and current poses, addressing the cost and complexity issues of existing methods by eliminating the need for expensive sensors and pose measurement.

JP7692676B2Active Publication Date: 2025-06-16UBKANG (QINGDAO) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023539797
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-31
Filing Date
2021-12-28
Publication Date
2025-06-16
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

Existing mobile robot control methods rely on expensive and bulky sensors like lidar, RGB-D, and stereo cameras for pose measurement, which is costly and complex, especially when the target model is unknown in advance.

Method used

An image-based mobile robot control method that uses a single camera to capture images of the robot's desired and current poses, extracts matching feature points, projects them onto a virtual unit sphere, and calculates image invariant features and rotation vector features to control the robot's movement without requiring pose measurement.

Benefits of technology

This method allows for cost-effective and efficient control of mobile robots using only a camera, eliminating the need for expensive sensors and reducing computational complexity by not requiring knowledge of the target model or pose estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007692676000019
    Figure 0007692676000019
  • Figure 0007692676000020
    Figure 0007692676000020
  • Figure 0007692676000021
    Figure 0007692676000021
Patent Text Reader

Abstract

The present invention relates to a mobile robot control method, a computer-implemented storage medium, and a mobile robot. The method includes the steps of: acquiring a first image taken by a camera on the robot when the robot is in a desired pose; acquiring a second image taken by a camera on the robot when the robot is in a current pose; extracting a plurality of pairs of matching feature points from the first and second images, projecting the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, the center of which overlaps with the optical center of the camera coordinate system; acquiring image invariant features and rotation vector features based on the plurality of projected feature points, and controlling the robot to move according to the image invariant features and rotation vector features until the robot reaches the desired pose. Since the desired pose of the robot is determined using the image invariant features and rotation vector features, there is no need to know a target model in advance or estimate the pose of the robot relative to the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to mobile robots, and more particularly to an image-based mobile robot control method and a mobile robot that do not perform pose measurement.

Background Art

[0002] Ground mobile robots are actively developing in fields such as rear service, operation assistance, and monitoring in order to perform various repetitive and dangerous activities. The adjustment control of many ground mobile robots is mainly carried out in the Cartesian coordinate system. That is, the output of the system is the metric coordinates (along the x-axis and y-axis), and the angle measured in degrees or radians (around the z-axis). For this purpose, it is necessary to use the assumed information collected by the sensor for control calculation. However, commonly used distance sensors such as lidar, RGB-D, and stereo cameras are expensive and large in volume.

[0003] The relative pose of a robot can be reconstructed with a monocular camera. However, for pose reconstruction based on a monocular camera, it is necessary to know the target model and the recovery of pose estimation in advance, but it is not always possible to know the target model in advance. Furthermore, the calculations required for pose estimation are very complex.

[0004] Therefore, in order to overcome the above problems, it is necessary to provide an image-based mobile robot control method that does not perform pose measurement.

Summary of the Invention

[0005] An object of the present invention is to provide a robot assistant in order to solve the above existing problems.

[0006] The present invention is realized as follows. A method realized by a computer for controlling a mobile robot, which is executed by one or more processors, the method comprising: obtaining a first image captured by a camera on the robot when the robot is in a desired pose; obtaining a second image captured by a camera on the robot when the robot is in a current pose; extracting a plurality of pairs of matching feature points from the first image and the second image, and projecting the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, wherein the center of the virtual unit sphere coincides with the optical center of the camera coordinates; obtaining an image invariant feature and a rotation vector feature based on the plurality of projected feature points, and controlling the robot to move until the robot reaches the desired pose according to the image invariant feature and the rotation vector feature.

[0007] Furthermore, the step of controlling the robot to move includes: controlling the robot to perform a translational movement according to the image invariant feature; and controlling the robot to rotate according to the rotation vector feature.

[0008] Furthermore, the step of controlling the robot to perform a translational movement according to the image invariant feature includes: calculating a first angular velocity and a first linear velocity according to the image invariant feature and a preset control model; controlling the robot to perform a translational movement according to the first angular velocity and the first linear velocity; determining whether a translational error is less than a first preset threshold; returning to the step of obtaining the second image when the translational error is greater than or equal to the first preset threshold; and controlling the robot to rotate according to the rotation vector feature when the translational error is less than the first preset threshold.

[0009] Furthermore, according to the rotation vector feature, the step of controlling the robot to rotate includes: calculating a second angular velocity and a second linear velocity based on the rotation vector feature and a control model; controlling the robot to rotate according to the second angular velocity and the second linear velocity; determining whether a direction error after the rotation of the robot is smaller than a second preset threshold; when the direction error of the robot is greater than or equal to the second preset threshold, acquiring a third image by the camera; extracting a plurality of pairs of matching feature points from the first image and the third image, projecting the extracted feature points onto the unit sphere to obtain a plurality of projected feature points; obtaining the rotation vector feature based on the plurality of projected feature points extracted from the first image and the third image, and then returning to the step of calculating the second angular velocity and the second linear velocity based on the rotation vector feature and the control model.

[0010] Furthermore, the image invariant feature includes one or more of the reciprocal of the distance between two of the projected feature points, image moments, and area.

[0011] Furthermore, the step of obtaining the image invariant feature based on the plurality of projected feature points includes: obtaining at least two of the distance between two of the projected feature points, image moments, and area; calculating an average value of at least two of the distance between two of the projected feature points, image moments, and area, and using this average value as the image invariant feature.

[0012] Furthermore, the step of obtaining the rotation vector feature based on the plurality of projected feature points includes: determining the acceleration direction of the robot based on the plurality of projected feature points; using the included angle between the acceleration direction and the x-axis of the robot coordinate system as the rotation vector feature.

[0013] Furthermore, the step of extracting a plurality of pairs of matching feature points from the first image and the second image includes: using scale-invariant feature transform to extract a first number of original feature points from each of the first image and the second image; and comparing and matching the extracted original feature points to obtain a second number of pairs of matching feature points.

[0014] Furthermore, the method further includes: controlling the robot to stop at a desired position in a preset pose; and using an image of the environment in front of the robot captured by the camera as the first image, wherein the environment in front of the robot includes at least three feature points.

[0015] The present invention also provides a non-transitory computer-readable storage medium storing one or more programs to be executed by a mobile robot. When the one or more programs are executed by one or more processors of the robot, the processing to be executed by the robot includes: obtaining a first image captured by a camera on the robot when the robot is in a desired pose; obtaining a second image captured by the camera on the robot when the robot is in a current pose; extracting a plurality of pairs of matching feature points from the first image and the second image, and projecting the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, wherein the center of the virtual unit sphere coincides with the optical center of the camera coordinates; obtaining an image invariant feature and a rotation vector feature based on the plurality of projected feature points; and controlling the robot to move until the robot reaches the desired pose according to the image invariant feature and the rotation vector feature.

[0016] Furthermore, the image invariant feature includes one or more of the reciprocal of the distance between two of the projection feature points, image moments, and area. The step of obtaining the rotation vector feature based on the plurality of projection feature points includes determining the acceleration direction of the robot based on the plurality of projection feature points, and using the included angle between the acceleration direction and the x-axis of the robot coordinate system as the rotation vector feature.

[0017] The present invention further includes a mobile robot. This mobile robot includes one or more processors, a memory, and one or more programs stored in the memory and arranged to be executed by the one or more processors. The one or more programs include instructions to obtain a first image captured by a camera on the robot when the robot is in a desired pose, instructions to obtain a second image captured by a camera on the robot when the robot is in the current pose, instructions to extract a plurality of pairs of matching feature points from the first image and the second image, project the extracted feature points onto a virtual unit sphere to obtain a plurality of projection feature points, where the center of the virtual unit sphere coincides with the optical center of the camera coordinates, instructions to obtain an image invariant feature and a rotation vector feature based on the plurality of projection feature points, and instructions to control the robot to move until the robot reaches the desired pose according to the image invariant feature and the rotation vector feature.

[0018] Furthermore, controlling the robot to move includes controlling the robot to perform translational movement according to the image invariant feature, and controlling the robot to rotate according to the rotation vector feature.

[0019] Furthermore, the step of controlling the robot to perform translational movement according to the image invariant features includes: calculating a first angular velocity and a first linear velocity according to the image invariant features and a preset control model; controlling the robot to perform translational movement according to the first angular velocity and the first linear velocity; determining whether a translational movement error is smaller than a first preset threshold; if the translational movement error is greater than or equal to the first preset threshold, returning to the step of acquiring the second image; and if the translational movement error is smaller than the first preset threshold, controlling the robot to rotate according to the rotation vector features. Further, the step of controlling the robot to rotate according to the rotation vector features includes: calculating a second angular velocity and a second linear velocity based on the rotation vector features and a control model; controlling the robot to rotate according to the second angular velocity and the second linear velocity; determining whether a direction error after rotation of the robot is smaller than a second preset threshold; if the direction error of the robot is greater than or equal to the second preset threshold, acquiring a third image by the camera; extracting a plurality of pairs of matching feature points from the first image and the third image, projecting the extracted feature points onto the unit sphere to obtain a plurality of projected feature points; acquiring the rotation vector features based on the plurality of projected feature points extracted from the first image and the third image, and then returning to the step of calculating the second angular velocity and the second linear velocity based on the rotation vector features and the control model.

[0020] Furthermore, the image invariant features include one or more of the reciprocal of the distance between two of the projected feature points, image moments, and area.

[0021] Furthermore, obtaining the image invariant features based on the plurality of projected feature points includes: obtaining at least two of the distance between two of the projected feature points, the image moment, and the area; and calculating an average value of at least two of the distance between two of the projected feature points, the image moment, and the area, and using this average value as the image invariant feature.

[0022] Furthermore, obtaining the rotation vector features based on the plurality of projected feature points includes: determining the acceleration direction of the robot based on the plurality of projected feature points; and using the included angle between the acceleration direction and the x-axis of the robot coordinate system as the rotation vector feature.

[0023] Furthermore, extracting a plurality of pairs of matching feature points from the first image and the second image includes: using scale-invariant feature transform to extract a first number of original feature points from each of the first image and the second image; and comparing and matching the extracted original feature points to obtain a second number of pairs of matching feature points.

[0024] Furthermore, it further includes: controlling the robot to stop at a desired position in a preset pose; and using an image of the environment in front of the robot captured by the camera as the first image, wherein the environment in front of the robot includes at least three feature points.

[0025] The technical effect of the present invention compared with the prior art is that the mobile robot control method of the present invention does not depend on expensive and bulky sensors in the prior art, but only requires one camera. On the other hand, in order to determine the desired pose of the robot using the image invariant features and the rotation vector features, it is not necessary to know the target model in advance or estimate the pose of the robot with respect to the target.

Brief Description of the Drawings

[0026] To further clarify the technical means of the embodiments of the present invention, the following briefly describes the drawings necessary for the description of the embodiments or the prior art of the present invention. Naturally, the drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative labor.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Modes for Carrying Out the Invention

[0027] Hereinafter, embodiments of the present invention will be described in detail, and exemplary embodiments are shown in the accompanying drawings. The same or similar reference numerals denote the same or similar elements, or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are for the purpose of explaining the present invention, and should not be construed as limiting the present invention.

[0028] In the description of the present invention, the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is only for the convenience of description and simplification of the description of the present invention based on the orientation or positional relationship shown in the drawings, and does not indicate or imply that the referred device or component must have a specific orientation and be structured and operated in a specific orientation. Therefore, it should be understood that it does not limit the present invention.

[0029] In addition, technical terms such as "first" and "second" are used only for the purpose of description and cannot be understood as indicating or suggesting relative importance or implicitly indicating the number of the shown technical features. Thus, the features limited by "first" and "second" can include one or more of the features explicitly or implicitly. In the description of the present invention, "a plurality" means two or more unless specifically and clearly limited.

[0030] In the present invention, unless specifically specified and limited, terms such as "mounting", "connecting", "connecting", "fixing", etc. need to be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integrated case. It may be a mechanical connection or an electrical connection. It may be a direct connection, an indirect connection via an intermediate medium, an internal communication between two elements, or an interaction relationship between two elements. For those skilled in the art, the specific meaning of the above terms in the present invention can be understood according to specific situations.

[0031] To further clarify the objectives, technical means, and advantages of the present invention, the following will describe the present invention in more detail with reference to the drawings and embodiments.

[0032] FIG. 1 is a schematic block diagram of a robot 11 according to an embodiment. The robot 11 may be a mobile robot. The robot 11 may include a processor 110, a memory 111, and one or more computer programs 112 stored in the memory 111 and executable by the processor 110. When the processor 110 executes the computer program 112, steps in an embodiment of a method for controlling the robot 11, for example, steps S41 to S44 in FIG. 5, steps S51 to S56 in FIG. 6, steps S551 to S56 in FIG. 10, and steps S561 to S566 in FIG. 11 are implemented.

[0033] Exemplarily, the one or more computer programs 112 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 111 and executed by the processor 110. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the one or more computer programs 112 in the robot 11.

[0034] The processor 110 may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, discrete gates, transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0035] The memory 111 may be an internal storage unit of the robot 11, such as a hard disk or a memory. The memory 111 may also be an external storage device of the robot 11, such as a plug-in hard disk, a smart memory card (SMC) and a secure digital (SD) card, or any suitable flash memory card. Further, the memory 111 may include an internal storage unit and an external storage device. The memory 111 is used to store computer programs, other programs, and data required by the robot. The memory 111 can also be used to temporarily store the output or output data.

[0036] Note that FIG. 1 is only an example of the robot 11, and it should be noted that it does not limit the robot 11. The robot 11 may include a different number of members from those shown, or may include some other different members. For example, in the embodiments shown in FIGS. 2 and 3, the robot 11 may further include an actuator 113, a moving mechanism 114, a camera 115, and a communication interface module 106. The robot 11 may include input / output devices, network access devices, a bus, etc.

[0037] In one embodiment, the actuator 113 may include one or more motors and / or steering gears. The moving mechanism 114 may include one or more wheels and / or crawlers. The actuator 113 is electrically coupled to the moving mechanism 114 and the processor 110 and can drive the movement of the moving mechanism 114 according to the instructions of the processor 110. The camera 115 may be, for example, a camera mounted on the robot 11. The camera 115 is electrically connected to the processor 110 and is configured to transmit the captured image to the processor 110. The communication interface module 116 may include a wireless transmitter, a wireless receiver, and a computer program executable by the processor 110. The communication interface module 116 is electrically connected to the processor 103 and is used for communication between the processor 110 and external devices. In one embodiment, the processor 110, the memory 111, the actuator 113, the moving mechanism 114, the camera 115, and the communication interface module 116 can be connected to each other via a bus.

[0038] In one embodiment, the robot 11 can observe a wall socket for automatic charging, observe items on a shelf for picking, and observe indoor furniture for home services. FIG. 4 shows an exemplary scenario in which the robot 11 can automatically dock to a charging station to execute an automatic charging process. Before controlling the robot 11 to perform self-charging, first, stop the robot 11 in front of the charging station and face it towards the charging station. Use the pose of the robot 11 at this time as the desired pose and capture an image of the environment in front of the robot 11 at this time as the desired image. Then, if the robot 11 needs to be charged, regardless of what pose the robot 11 is in, use the current pose as the initial pose and capture an image of the environment in front of the robot 11 as the current image. Next, by executing the steps of the control method provided in the following embodiments, the robot 11 is controlled to move and rotate to the desired pose based on the feature points in the desired image and the current image.

[0039] FIG. 5 shows a flowchart of a mobile robot control method according to an embodiment. This method can be used to control the movement of the robot 11 in FIGS. 1 to 3. This method can be executed by one or more processors of the robot or one or more processors of other control devices electrically coupled to the robot. The control device may include, but is not limited to, a desktop computer, a tablet computer, a notebook computer, a multimedia player, a server, a smart mobile device (such as a smartphone, a mobile phone, etc.), and a smart wearable device (such as a smartwatch, smart glasses, a smart camera, a smart bracelet, etc.), and other computing devices with computing and control functions. In one embodiment, this method may include steps S41 to S44.

[0040] In step S41, when the robot is in a desired pose, a first image captured by a camera on the robot is acquired.

[0041] In step S42, when the robot is in the current pose, a second image captured by a camera on the robot is acquired.

[0042] In step S43, a plurality of pairs of matching feature points are extracted from the first image and the second image, and the extracted feature points are projected onto a virtual unit sphere to obtain a plurality of projected feature points.

[0043] The virtual unit sphere is an integrated spherical model that simulates a central imaging system using a combination of a virtual sphere and perspective projection. The center of the virtual unit sphere coincides with the optical center of the camera coordinates. When a plurality of pairs of matching feature points extracted from the first image and the second image are projected onto the virtual unit sphere, the corresponding projection points on the unit sphere become the projected feature points.

[0044] In step S44, based on the plurality of projected feature points, an image invariant feature and a rotation vector feature are obtained, and according to the image invariant feature and the rotation vector feature, the robot is controlled to move until it reaches the desired pose.

[0045] In one embodiment, the image invariant feature may include one or more of the reciprocal of the distance between two of the projected feature points, the image moment, and the area. The rotation vector feature may be an angle vector or a direction vector. The control of the robot's movement includes the control of the robot's translational movement and the control of the robot's rotation.

[0046] FIG. 6 shows a flowchart of a mobile robot control method according to another embodiment. This method can be used to control the movement of the robot in FIGS. 1 to 3. This method can be executed by one or more processors of the robot or one or more processors of other control devices electrically coupled to the robot. The control device may include, but is not limited to, a desktop computer, a tablet computer, a notebook computer, a multimedia player, a server, a smart mobile device (such as a smartphone, a mobile phone, etc.), and a smart wearable device (such as a smartwatch, smart glasses, a smart camera, a smart bracelet, etc.), and other computing devices with computing and control functions. In one embodiment, this method may include steps S51 to S56.

[0047] In step S51, when the robot is in the desired pose, a first image captured by a camera on the robot is obtained.

[0048] In one embodiment, the desired pose may include the desired position and the desired orientation of the robot. The robot is expected to move to the desired position in the desired direction. The robot can be controlled (e.g., by the user) to stop at the desired position and in the desired direction, whereby an image of the environment in front of the robot (e.g., a wall with an outlet) can be captured by a camera mounted on the robot as the first image (i.e., the desired image). When the robot stops at the desired position and in the desired direction, there must be at least three feature points in the environment within the robot's field of view. The camera of the robot may be a fish-eye pinhole camera or a non-pinhole camera such as a refractive-reflective camera.

[0049] In step S52, when the robot is in the current pose, a second image captured by the camera on the robot is acquired.

[0050] When it is necessary to control the robot to move to the desired position and be in the desired direction, the current image of the environment in front of the robot is captured by the camera mounted on the robot as the second image (i.e., the current image). When capturing the second image, there must be at least three feature points in the environment within the robot's field of view.

[0051] In step S53, a plurality of pairs of matching feature points are extracted from the first image and the second image, and the extracted feature points are projected onto a virtual unit sphere to obtain a plurality of projected feature points. The center of this virtual unit sphere coincides with the optical center of the coordinates of this camera.

[0052] In one embodiment, extracting a plurality of pairs of matching feature points from the first image and the second image can be realized as follows.

[0053] First, using a Scale-invariant feature transform (SIFT) descriptor, a first number of original feature points are extracted from each of the first image and the second image. Next, the extracted original feature points are compared and matched to obtain a second pair of matching feature points.

[0054] For example, using the SIFT descriptor, 200 SIFT feature points from any position in the first image and 200 SIFT feature points from any position in the second image can be extracted. Next, using the closest Euclidean distance algorithm, the 400 (200 pairs) of extracted SIFT feature points are compared and matched to obtain at least 3 pairs of matching feature points. Optionally, in the process of comparison and matching, a balanced binary search tree such as a KDTree can be used to speed up the search process.

[0055] Note that the positions of the above-mentioned feature points are not limited. That is, these feature points may be on the same plane or on non-same planes, and the plane on which these feature points are located may be perpendicular or not perpendicular to the motion plane.

[0056] One goal of the present disclosure is to adjust the pose of a robot based on the visual feedback of at least three non-collinear static feature points in a first image and make it independent of the pose requirements of the robot with respect to an inertial system or a visual target. To achieve this goal, it is necessary to project a plurality of pairs of matching feature points extracted from the first image and the second image onto a unit sphere.

[0057] Referring to FIG. 7, in one embodiment, the camera mounted on the robot is a calibrated onboard camera having a coordinate system F in which the coordinates of the robot and the coordinates of the camera are the same. The robot moves on the plane shown in FIG. 7. The x-axis and y-axis in the coordinate system F define the moving plane. The positive direction of the x-axis is the traveling direction of the robot, and the y-axis coincides with the axis about which the wheels of the robot rotate. Point P in the coordinate system F i is projected onto a virtual unit sphere and represented by h i . The virtual unit sphere is an integrated spherical model that simulates a central imaging system using a combination of a virtual sphere and a perspective projection.

[0058] In step S54, based on a plurality of projected feature points, image invariant features and rotation vector features are obtained.

[0059] Image invariant features are a class of image features that are invariant to the rotation of the camera. Image invariant feature S ∈ R 2 represents the translational motion of the robot as a system output, and R 2 represents a two-dimensional real coordinate space. The dynamic model of the image invariant feature is JPEG0007692676000001.jpg6170Here, J ∈ R 2x2 is the interaction matrix, and υ = [υ x , υ y T represents the linear velocity of the robot in the coordinate system F and is not restricted by the inconsistency constraint. The acceleration of the robot coordinate system satisfies α ∈ R 2 . JPEG0007692676000002.jpg12170

[0060] In one embodiment, the invariant image features may include one or more of the reciprocal of the distance between two of the projected feature points, image moments, and area.

[0061] In an example as shown in FIG. 7, the distance d i between the projected feature points h j and h ij ​Its reciprocal is used as an image invariant feature because it has the best linearization characteristics between the task space and the image space. Point h i and h j represent, for example, the projections of feature points extracted from the same image, such as the feature points P i and P j in FIG. 7, onto the unit sphere. That is, points h i and h j can represent the projections of feature points extracted from the second image (i.e., the current image) onto the unit sphere. Points h i and h j can represent the projections of feature points extracted from the first image (i.e., the desired image) onto the unit sphere. Note that it should be noted that two projected feature points can be connected to each other to form a small line segment. As the robot approaches the desired position, the line segment becomes longer. The purpose of the translational control is to make the selected line segment equal to the desired value. For example, assuming there are three projected feature points, i = 0, 1, 2, j = 0, 1, 2, and i ≠ j, there are three combinations (0, 1)(0, 2)(1, 2), and two of them can be selected.

[0062] JPEG0007692676000003.jpg represents the homogeneous coordinates of the feature points extracted from the 7170 image plane, and i represents a positive integer greater than 0. A virtual image plane called a retina is constructed, which is associated with a virtual perspective projection camera. The corresponding coordinates in the retina are JPEG0007692676000004.jpg8170 where A represents a generalized camera projection matrix related to the mirror and the camera internal parameters. The projected feature points projected onto the unit sphere are JPEG0007692676000005.jpg31170

[0063] Based on the above principle, at least three pairs of feature points extracted from the real environment represented by the desired image and the current image are projected onto a virtual unit sphere with h i points. The corresponding image invariant features are JPEG0007692676000006.jpg51170 In this embodiment, they can be replaced with constant values close to the desired values. According to experiments, it is found that the closed-loop system is stable under this approximation.

[0064] In one embodiment, when image moments are used as image invariant features, the image invariant features are JPEG0007692676000007.jpg20170

[0065] In one embodiment, when the area is used as an image invariant feature, the image invariant feature is JPEG0007692676000008.jpg10170 represents the three-dimensional coordinates of the i-th projected feature point on the unit sphere.

[0066] JPEG0007692676000009.jpg43170 represents the desired value.

[0067] JPEG0007692676000010.jpg36170

[0068] By calculating the rotation vector feature using all the projected feature points obtained by projection, the robustness can be improved. The calculated rotation vector feature can provide better characteristics because it provides a direct mapping between the rotation speed and the rotation vector speed.

[0069] Note that regardless of whether the image invariant feature is distance, image moment, or area, the rotation vector feature JPEG0007692676000011.jpg42170

[0070] In step S55, the robot is controlled to move translationally according to the image invariant feature.

[0071] In step S56, the robot is controlled to rotate according to the rotation vector feature. Note that the movement of the robot is a rigid body movement and can be decomposed into rotational movement and translational movement. Referring to FIG. 9, in this embodiment, when adjusting the pose of the robot, it is necessary to perform switching control between two steps. The first step is to control the robot to translate using image invariant features. The goal is to eliminate the parallel error (also called the invariant feature error). That is, the goal is to make the translational error equal to 0. The second step is to control the robot to rotate after removing the translational error using the rotation vector feature. The goal is to eliminate the direction error, that is, to make the direction error equal to 0.

[0072] In the above first step, in order to represent the translational state using image invariant features, even if the desired pose cannot be obtained, the robot can be controlled to translate. By using image invariant features, the control of direction and position is separated. JPEG0007692676000012.jpg7170

[0073] JPEG0007692676000013.jpg14170 It is controlled to rotate to align the mass center to a desired value.

[0074] Referring to FIG. 10, in one embodiment, the first step (i.e., step S55) may include the following steps.

[0075] In step S551, according to the image invariant features and the preset control model, the first angular velocity and the first linear velocity are calculated.

[0076] In step S552, the robot is controlled to translate according to the first angular velocity and the first linear velocity.

[0077] In step S553, it is determined whether the translational error is smaller than the first preset threshold.

[0078] When the translational error is less than the first preset threshold, the process proceeds to step S56 that controls the robot to rotate according to the rotational vector feature.

[0079] When the translational error is greater than or equal to the first preset threshold, it means that the translational error has not been eliminated, the robot has not reached the desired position, and the robot needs to translate towards the required position. Then, the process returns to step S52 to acquire the second image. Steps S551, S552, S553, and S52 are repeated until the translational error becomes less than the first preset threshold, and then the second step of controlling the direction of the robot is started.

[0080] JPEG0007692676000014.jpg52170 is one of the sufficient conditions for.

[0081] JPEG0007692676000015.jpg14170 can be used to determine whether the translational error is zero.

[0082] According to the above formula, it can be seen that during the position control process of the first step, the trajectory of the robot is not a straight line but a curve. That is, in the process of controlling the translational movement of the robot in the first step, not only the magnitude of the robot's movement needs to be adjusted, but also accordingly, the direction of the robot needs to be adjusted in order to move the robot to the desired position.

[0083] JPEG0007692676000016.jpg6170 The direction is flipped over by 180 degrees. Such a design is for enabling the robot to keep the visual target in front when the robot moves backward and the target is behind. When the FOV is limited, it helps to keep the visual target within the field of view (FOV) of the camera.

[0084] Referring to FIG. 11, in one embodiment, the second step (i.e., step S56) may include the following steps.

[0085] In step S561, based on the rotation vector feature and the control model, a second angular velocity and a second linear velocity are calculated.

[0086] In step S562, the robot is controlled to rotate according to the second angular velocity and the second linear velocity.

[0087] In step S563, it is determined whether the direction error after the rotation of the robot is smaller than a second preset threshold.

[0088] In step S564, if the direction error of the robot is greater than or equal to the second preset threshold, it means that the direction error has not been eliminated and the robot is not in the desired direction. In this case, a third image is acquired by the camera of the robot.

[0089] In step S565, a plurality of pairs of matching feature points are extracted from the first image and the third image, and the extracted feature points are projected onto the unit sphere to obtain a plurality of projected feature points.

[0090] In step S566, based on the plurality of projected feature points extracted from the first image and the third image, the rotation vector feature is obtained. Next, the process returns to step S561 where the second angular velocity and the second linear velocity are calculated based on the rotation vector feature and the control model. Steps S551, S552, S553, and S56 are repeated until the direction error is smaller than the second preset threshold.

[0091] In step S567, if the direction error of the robot is smaller than the second preset threshold, it means that the direction error has been eliminated and the robot is in the desired direction, and the process ends. In one embodiment, steps S561 to S567 are based on the rotation vector feature and the control model JPEG0007692676000017.jpg represents the 16170 quantity center.

[0092] JPEG0007692676000018.jpg22170

[0093] In one embodiment, image invariant features can be determined based on two or all combinations of distance, image moment, and area. Specifically, the average value of at least two of the distances between two of the projection feature points, image moments, and areas is used as the image invariant feature. Next, the image invariant feature is substituted into the above formula to realize the position control of the robot. In this way, by combining multiple parameters to determine the invariant characteristics of the final application, the accuracy of the control result is improved.

[0094] On the one hand, the image features in the image coordinate system are used to represent the kinematics / dynamics of the mobile robot instead of the distance and angle defined in the Cartesian coordinate system. Instead of relying on conventional expensive and bulky sensors (such as lidar, sonar, wheel odometers, etc.) for the movement control of the robot, only one camera is required, which can reduce the manufacturing cost of the robot and reduce the size of the robot.

[0095] On the other hand, in order to determine the desired pose of the robot using image invariant features and rotation vector features, it is not necessary to know the target model in advance or estimate the pose of the robot with respect to the target. That is, since it is not necessary to calculate and decompose the homography matrix or the fundamental matrix, the computational complexity can be reduced and the computational speed can be improved.

[0096] On the one hand, since the desired image and the feature points in the current image can be arbitrarily selected, there is no requirement for the physical position of the points in the environment. For example, the feature points in the environment may be on the same plane or on non - same planes. Therefore, this method can be applied to more scenarios, improving the versatility of this method.

[0097] In one embodiment, a mobile robot control device can be constructed in the same manner as the robot of FIG. 1. That is, the mobile robot control device may include a processor and a memory. The memory is electrically coupled to the processor and stores one or more computer programs executable by the processor. When the processor executes the computer program, steps in an example of a method for controlling the robot 11 are performed, such as steps S41 to S44 in FIG. 5, steps S551 to S56 in FIG. 6, steps S561 to S56 in FIG. 10, and steps S561 to S566 in FIG. 11. The processor, memory, and computer program of the mobile robot control device may be the same as the above-described processor 110, memory 111, and computer program 112, and will not be repeated here. The robot movement control device may be various computer system devices having a data interaction function, including, but not limited to, mobile phones, smartphones, other wireless communication devices, personal digital assistants, audio players, other media players, music recorders, video recorders, cameras, other media recorders, radios, vehicle transportation devices, programmable remote controls, laptop computers, desktop computers, printers, netbook computers, portable game devices, portable Internet devices, data storage devices, smart wearable devices (e.g., head-mounted devices (HMDs), such as smart glasses, smart wearables, smart bracelets, smart necklaces, or smart watches), and combinations thereof.

[0098] Exemplarily, one or more computer programs can be divided into one or more modules / units, the one or more modules / units can be stored in a memory, and executed by a processor. The one or more modules / units may be a series of computer program instruction segments capable of executing a specific function. The instruction segments are used to describe the execution process of one or more computer programs. For example, one or more computer programs can be divided into a first acquisition module, a second acquisition module, a feature extraction module, a calculation module, and a control module.

[0099] The first acquisition module is arranged to acquire a first image captured by a camera on the robot when the robot is in a desired pose.

[0100] The second acquisition module is arranged to acquire a second image captured by a camera on the robot when the robot is in the current pose.

[0101] The feature extraction module is arranged to extract a plurality of pairs of matching feature points from the first image and the second image, and project the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points. The center of the virtual unit sphere overlaps with the optical center of the camera coordinates.

[0102] The calculation module is arranged to acquire an image invariant feature and a rotation vector feature based on the plurality of projected feature points.

[0103] The control module is arranged to control the robot to move until it reaches the desired pose according to the image invariant feature and the rotation vector feature.

[0104] It should be noted that the mobile robot control device may include a different number of components from those described above, or may be combined with some other different components. For example, the mobile robot control device may include an input / output device, a network access device, a bus, etc.

[0105] As can be understood by those skilled in the art, for the convenience of description and to simplify, the assignment of each of the above functional units and modules has been described by way of example. However, in practice, depending on requirements, the above functions may be realized by different functional units and modules with different function assignments. That is, in order to realize all or part of the functions described above, the internal structure of the device may be assigned to different functional units or modules. Each functional unit in each embodiment may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated units may be realized in the form of hardware or in the form of software functional units. Furthermore, the specific names of each functional unit and module are only for the convenience of distinguishing from each other and are not used to limit the protection scope of the present disclosure. For the specific working processes of the units and modules in the above system, reference may be made to the corresponding processes of the method embodiments described above, and details will not be described again in this specification.

[0106] In one embodiment, the non-transitory computer-readable storage medium can be disposed in the robot 11 or the mobile robot control device as described above. The non-transitory computer-readable storage medium may be a storage unit disposed in the main control chip and the data collection chip in the above-described embodiments. One or more computer programs are stored in the non-transitory computer-readable storage medium, and when the computer programs are executed by one or more processors, the robot control method described in the above embodiments is implemented.

[0107] In the foregoing embodiments, the description of each embodiment focuses on itself. However, for parts that are not detailed or described in one embodiment, reference may be made to the related descriptions of other embodiments.

[0108] Those skilled in the art will understand that the various example units and algorithm steps described in connection with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software depends on the specific application of the technical solution and design constraints. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered as exceeding the scope of this disclosure.

[0109] In the embodiments provided by this disclosure, it should be understood that the disclosed apparatus / terminal device and method may be implemented in other ways. For example, the embodiments of the apparatus / terminal device as described above are merely illustrative. For example, the division of the above-mentioned module or unit is only a logical function division, and there may be other division methods when actually implemented. For example, a plurality of units or assemblies may be combined or integrated into another system, or some features may be ignored or not executed. Also, the shown or discussed couplings or direct couplings or communication connections between each other may be indirect couplings or communication connections through some interfaces, devices or units, and may be in electrical, mechanical or other forms.

[0110] The units described as separate members may or may not be physically separated, and the members shown as units may or may not be physical units, may be located in the same place, or may be distributed on a plurality of network units. According to actual requirements, some or all of the units can be selected to achieve the purpose of this embodiment.

[0111] In each embodiment of the present disclosure, each functional unit may be integrated into one processing unit, each unit may physically exist alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0112] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the process in the method for implementing the above embodiments of the present disclosure can also be realized by instructing related hardware with a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be realized. The computer program includes computer program code, and the computer program code may be in source code form, object code form, executable file, or some intermediate form. The computer-readable recording medium may include any primitive or device, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, communication signal, and software distribution medium that can carry the computer program code. It should be noted that the content included in the computer-readable recording medium can be appropriately increased or decreased according to the laws of the jurisdiction and the requirements for patent enforcement. For example, in some jurisdictions, according to laws and patent enforcement, computer-readable media do not include electrical carrier signals and electrical communication signals.

[0113] The foregoing description has been presented for purposes of illustration and description, and has been made with reference to specific embodiments. However, the foregoing illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application, to enable others skilled in the art to best utilize the disclosure and various embodiments with various modifications as are suited to the particular use contemplated.

Claims

1. A method implemented by a computer for controlling a mobile robot, which is executed by one or more processors, comprising: obtaining a first image captured by a camera on the robot when the robot is in a desired pose; obtaining a second image captured by a camera on the robot when the robot is in a current pose; extracting a plurality of pairs of matching feature points from the first image and the second image using scale-invariant feature transform descriptors and the nearest neighbor Euclidean distance algorithm, and projecting the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, wherein the center of the virtual unit sphere coincides with the optical center of the camera coordinates; obtaining image invariant features, which are a type of image features invariant to the rotation of the camera, and rotation vector features representing the angle between the x-axis direction, which is the traveling direction of the robot, and the acceleration direction of the robot in a coordinate system where the coordinates of the robot and the camera are the same, based on the plurality of projected feature points, and controlling the robot to move until it reaches the desired pose according to the image invariant features and the rotation vector features.

2. The step of controlling the robot to move comprises: controlling the robot to perform translational movement according to the image invariant features; and controlling the robot to rotate according to the rotation vector features. The method according to claim 1, characterized by comprising the above steps.

3. The step of controlling the robot to perform translational movement according to the image invariant features comprises: calculating a first angular velocity and a first linear velocity according to the image invariant features and a preset control model; and controlling the robot to perform translational movement according to the first angular velocity and the first linear velocity. A step of determining whether the translational error is less than a first preset threshold value, When the translational error is greater than or equal to the first preset threshold value, a step of returning to the step of acquiring the second image, When the translational error is less than the first preset threshold value, a step of controlling the robot to rotate according to the rotation vector feature, characterized in that the method according to claim 2 includes the steps.

4. The step of controlling the robot to rotate according to the rotation vector feature includes: Calculating a second angular velocity and a second linear velocity based on the rotation vector feature and a control model; Controlling the robot to rotate according to the second angular velocity and the second linear velocity; Determining whether a direction error after rotation of the robot is less than a second preset threshold value; When the direction error of the robot is greater than or equal to the second preset threshold value, acquiring a third image by the camera; Extracting a plurality of pairs of matching feature points from the first image and the third image, projecting the extracted feature points onto the unit sphere to obtain a plurality of projected feature points; Based on the plurality of projected feature points extracted from the first image and the third image, acquiring the rotation vector feature, and then returning to the step of calculating the second angular velocity and the second linear velocity based on the rotation vector feature and the control model, characterized in that the method according to claim 2 includes the steps.

5. The image invariant feature includes one or more of the reciprocal of the distance between two of the projected feature points, image moments, and area, characterized in that the method according to claim 1 includes the steps.

6. The step of acquiring the image invariant feature based on the plurality of projected feature points includes: Acquiring at least two of the distance between two of the projected feature points, image moments, and area. calculating at least two average values among the distances between two of the projection feature points, image moments, and areas, and using the average values as image invariant features;

7. The step of obtaining a rotation vector feature based on the plurality of projection feature points includes: determining the acceleration direction of the robot based on the plurality of projection feature points; and using the included angle between the acceleration direction and the x-axis direction of the robot's coordinate system as the rotation vector feature. The method according to claim 1 is characterized by this.

8. The step of extracting a plurality of pairs of matching feature points from the first image and the second image includes: using the scale-invariant feature transform descriptor to extract a first number of original feature points from each of the first image and the second image; and comparing and matching the extracted original feature points to obtain a second number of pairs of matching feature points. The method according to claim 1 is characterized by this.

9. controlling the robot to stop at a desired position in a preset pose; and using an image of the environment in front of the robot taken by the camera as the first image, where the environment in front of the robot includes at least three feature points. The method according to claim 1 is further characterized by this.

10. A non-transitory computer-readable storage medium storing one or more programs to be executed by a mobile robot, where when the one or more programs are executed by one or more processors of the robot, the processing to be executed by the robot includes: obtaining a first image taken by a camera on the robot when the robot is in a desired pose; When the robot is in the current pose, obtaining a second image captured by a camera on the robot; Extracting a plurality of pairs of matching feature points from the first and second images using scale-invariant feature transform descriptors and the nearest Euclidean distance algorithm, and projecting the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, wherein the center of the virtual unit sphere coincides with the optical center of the camera coordinates; Based on the plurality of projected feature points, obtaining an image invariant feature which is a type of image feature invariant to the rotation of the camera, and a rotation vector feature representing the included angle between the x-axis direction which is the traveling direction of the robot and the acceleration direction of the robot in a coordinate system where the coordinates of the robot and the camera are the same, and controlling the robot to move until it reaches the desired pose according to the image invariant feature and the rotation vector feature. A non-transitory computer-readable storage medium characterized by comprising:

11. The image invariant feature includes one or more of the reciprocal of the distance between two of the projected feature points, image moments, and area; The step of obtaining the rotation vector feature based on the plurality of projected feature points includes: Determining the acceleration direction of the robot based on the plurality of projected feature points; Taking the included angle between the acceleration direction and the x-axis direction of the robot's coordinate system as the rotation vector feature. The non-transitory computer-readable storage medium according to claim 10, characterized by comprising:

12. A mobile robot, One or more processors; A memory; One or more programs stored in the memory and arranged to be executed by the one or more processors, The one or more programs are, A command to obtain a first image captured by a camera on the robot when the robot is in a desired pose, A command to obtain a second image captured by a camera on the robot when the robot is in the current pose, A command to extract a plurality of pairs of matching feature points from the first image and the second image using scale-invariant feature transform descriptors and the nearest neighbor Euclidean distance algorithm, and project the extracted feature points onto a virtual unit sphere to obtain a plurality of projected feature points, wherein the center of the virtual unit sphere overlaps with the optical center of the camera coordinates, Based on the plurality of projected feature points, an image invariant feature, which is a type of image feature invariant to the rotation of the camera, and a rotation vector feature representing the included angle between the x-axis direction, which is the traveling direction of the robot, and the acceleration direction of the robot in a coordinate system where the coordinates of the robot and the camera are the same are obtained, and according to the image invariant feature and the rotation vector feature, the robot is controlled to move until it reaches the desired pose. A robot characterized by including:

13. Controlling the robot to move includes: Controlling the robot to perform translational movement according to the image invariant feature; Controlling the robot to rotate according to the rotation vector feature. The robot according to claim 12, characterized by including:

14. The step of controlling the robot to perform translational movement according to the image invariant feature includes: Calculating a first angular velocity and a first linear velocity according to the image invariant feature and a preset control model; Controlling the robot to perform translational movement according to the first angular velocity and the first linear velocity; Determining whether the translational movement error is smaller than a first preset threshold; When the translational movement error is greater than or equal to the first preset threshold, returning to the step of obtaining the second image. When the translational error is smaller than the first preset threshold, the step of controlling the robot to rotate according to the rotation vector feature, and the robot according to claim 13, characterized in that it includes this.

15. The step of controlling the robot to rotate according to the rotation vector feature is Based on the rotation vector feature and the control model, the step of calculating a second angular velocity and a second linear velocity, and According to the second angular velocity and the second linear velocity, the step of controlling the robot to rotate, and The step of determining whether the direction error after the rotation of the robot is smaller than a second preset threshold, and When the direction error of the robot is greater than or equal to the second preset threshold, the step of acquiring a third image by the camera, and Extracting a plurality of pairs of matching feature points from the first image and the third image, projecting the extracted feature points onto the unit sphere to obtain a plurality of projected feature points, and Based on the plurality of projected feature points extracted from the first image and the third image, obtaining the rotation vector feature, and then returning to the step of calculating the second angular velocity and the second linear velocity based on the rotation vector feature and the control model, and the robot according to claim 13, characterized in that it includes this.

16. The image invariant feature includes one or more of the reciprocal of the distance between two of the projected feature points, the image moment, and the area, and the robot according to claim 12, characterized in that it is like this.

17. Obtaining the image invariant feature based on the plurality of projected feature points is The step of obtaining at least two of the distance between two of the projected feature points, the image moment, and the area, and Calculating at least two average values among the distances between two of the projected feature points, image moments, and area, and using this average value as an image invariant feature; and the robot according to claim 12, characterized in that it includes this step.

18. Based on the plurality of projected feature points, obtaining the rotation vector feature includes: Based on the plurality of projected feature points, determining the acceleration direction of the robot; and Taking the included angle between the acceleration direction and the x-axis direction of the robot's coordinate system as the rotation vector feature; and the robot according to claim 12, characterized in that it includes these steps.

19. Extracting a plurality of pairs of matching feature points from the first image and the second image includes: Using the scale-invariant feature transform descriptor to extract a first number of original feature points from each of the first image and the second image; and Comparing and matching the extracted original feature points to obtain a second number of pairs of matching feature points; and the robot according to claim 12, characterized in that it includes these steps.

20. Controlling the robot to stop at a desired position in a preset pose; and Taking the image of the environment in front of the robot captured by the camera as the first image, wherein the environment in front of the robot includes at least three feature points; and the robot according to claim 12, further characterized in that it includes this step.

Citation Information

Patent Citations

  • Image processing apparatus, robot control system, robot, image processing method, and image processing program

    JP2014238687A

  • Autonomous mobile device, memory defragmentation method and program

    JP2019159354A

  • Feature value extraction device and location inference device

    WO2014073204A1