Self-moving device system and method for determining pose of charging station thereof
Through camera component parameter configuration and target following technology, combined with vision and laser sensors, the accuracy and efficiency problems of self-mobile devices when positioning the boundaries of the working area and the positioning of the charging pile are solved, and efficient and accurate self-mobile device operation is achieved.
Patent Information
- Application Number
- PCT/CN2024/131456
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-10
AI Technical Summary
The existing self-mobile devices have time-consuming and labor-intensive and inaccurate positioning when positioning the boundaries of the working area and the charging pile positioning. In particular, the parameter deviation occurs after the image acquisition device is aging, which affects the working accuracy and recharge efficiency.
The camera component parameter configuration system is used to calibrate three-dimensional feature points through preset feature information, combine binocular cameras and infrared lasers to achieve fast and accurate calibration of camera components, and use target follow-up technology to identify the boundaries of the working area, and combine visual and laser sensors to determine the precise position of the charging pile.
It improves the working accuracy of the camera components, reduces the need for manual deployment of physical encirclement lines, improves the identification efficiency of the work area boundaries of the self-mobile device and the accuracy of the charging pile position, and ensures efficient charging of the self-mobile device.
Smart Images

Figure CN2024131456_10072025_PF_FP_ABST
Abstract
Description
Position determination method for self-propelled equipment system and charging pile position
[0001] This application claims priority to the Chinese patent application with application number 202410021420.1 filed with the China Patent Office on January 5, 2024, the Chinese patent application with application number 202410098903.1 filed with the China Patent Office on January 23, 2024, and the Chinese patent application with application number 202410030288.0 filed with the China Patent Office on January 9, 2024. The entire contents of the above applications are incorporated by reference into this application. Technical Field
[0002] The present application relates to the technical field of electric tools, for example, to a self-propelled equipment system. Background Art
[0003] With the rapid development of power tool technology, autonomous devices are gaining increasing attention and are being widely used in various work scenarios, such as autonomous lawn mowers. Autonomous devices typically incorporate image acquisition devices (such as cameras) to enable positioning and movement.
[0004] Before using a self-propelled lawn mower to mow, it is necessary to first determine the boundaries of the mower's working area so that the scope of the mower's working area can be determined based on the boundaries of the working area. In related art, physical boundaries are manually installed on the boundaries of the mower's working area. When the mower is mowing, the working area boundaries are located by identifying the physical boundaries, and then the scope of the working area determined by the working area boundaries is located. However, this solution requires relying on manually installed physical boundaries, which is time-consuming, labor-intensive, and requires a certain amount of financial investment.
[0005] After mowing, the autonomous device needs to determine the exact position of the charging station during its self-charging process. Related technologies use vision technology or laser sensor technology to locate the charging station, but due to the limitations of these technologies, the positioning results are inaccurate, which in turn affects the recharging efficiency of the autonomous device.
[0006] When an image acquisition device in a mobile device operates for a long time, structural issues such as lens or circuit board aging may occur. This can cause the device's parameters to deviate from reality, leading to inaccurate parameters and thus affecting the device's operating accuracy. Therefore, to ensure the device's operating accuracy, accurate calibration of the device's parameters is necessary.
[0007] This section provides background information related to the present application which is not necessarily prior art.
[0008] Summary of the Invention
[0009] An object of the present application is to solve or at least alleviate some or all of the above problems.
[0010] One object of the present application is to provide a self-supporting equipment system for configuring camera component parameters.
[0011] In order to achieve the above objectives, this application adopts the following technical solutions:
[0012] A self-moving device system for configuring camera component parameters includes a self-moving device and a charging pile, wherein the charging pile is configured to charge the self-moving device; the self-moving device includes: a body; a walking wheel assembly, configured to support the body; a camera assembly, mounted to the body, configured to capture images from the periphery of the self-moving device; a controller, communicatively connected to the camera assembly, the controller pre-stores three-dimensional feature points on the charging pile; wherein the three-dimensional feature points are pre-calibrated based on preset feature information; the controller is configured to: obtain a target image corresponding to the charging pile captured by the camera assembly, and extract feature points from the target image based on the preset feature information to obtain two-dimensional feature points; and calibrate the camera assembly parameters based on the correspondence between the two-dimensional feature points and the three-dimensional feature points.
[0013] In some embodiments, the camera assembly includes a binocular camera including a left camera and a right camera.
[0014] In some embodiments, the target image includes a first image and a second image, the first image is an image captured by the left camera, and the second image is an image captured by the right camera.
[0015] In some embodiments, feature point extraction is performed on a target image based on preset feature information to obtain two-dimensional feature points, including: extracting feature points on a first image based on the preset feature information to obtain two-dimensional feature points of a first image corresponding to the first image; and extracting feature points on a second image based on the preset feature information to obtain two-dimensional feature points of a second image corresponding to the second image.
[0016] In some embodiments, the camera assembly is calibrated based on the correspondence between the two-dimensional feature points and the three-dimensional feature points, including: determining the first intrinsic parameter of the left camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image; and determining the second intrinsic parameter of the right camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image.
[0017] In some embodiments, the internal parameters include focal length and optical center.
[0018] In some embodiments, based on the correspondence between two-dimensional feature points and three-dimensional feature points, the camera component is calibrated, including: determining a first external parameter of the left camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image and the first internal parameter; wherein the first external parameter is used to characterize the relative posture information between the left camera and the charging pile; determining a second external parameter of the right camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image and the second internal parameter; wherein the second external parameter is used to characterize the relative posture information between the right camera and the charging pile; determining a third external parameter of the binocular camera based on the first external parameter and the second external parameter; wherein the third external parameter is used to characterize the relative posture information between the left camera and the right camera.
[0019] In some embodiments, the preset feature information includes structure information and additional mark information.
[0020] In some embodiments, the structural information includes at least one of the following: edge information, corner information, and texture information.
[0021] In some embodiments, the additional marking information includes preset pattern information, and the preset pattern information includes QR code information.
[0022] In some embodiments, the self-propelled device system further includes an infrared laser, and the camera assembly further includes a low-light camera.
[0023] In some embodiments, the controller is further configured to: determine a working scene of the self-mobile device, where the working scene is daytime or nighttime; and determine a working mode of a camera component in the self-mobile device according to the working scene.
[0024] In some embodiments, determining the working scene of the mobile device includes determining the working scene of the mobile device based on the working time of the mobile device or the brightness of the environment in which the mobile device is located.
[0025] In some embodiments, the working mode of the camera component in the mobile device is determined according to the working scene, including: if the working scene is daytime, determining that the working mode of the camera component is a low-light camera and a binocular camera in a first state, or only a binocular camera in the first state; wherein the first state is that the infrared laser is in an off state; if the working scene is night, determining that the working mode of the camera component is a low-light camera and a binocular camera in a second state; wherein the second state is that the infrared laser is in an on state.
[0026] The benefits of this application are:
[0027] The present application provides a self-moving device system for configuring camera component parameters, including a self-moving device and a charging pile, wherein the charging pile is configured to charge the self-moving device; the self-moving device includes: a body; a walking wheel assembly, configured to support the body; a camera assembly, mounted on the body, configured to capture images from the periphery of the self-moving device; a controller, communicatively connected to the camera assembly, the controller pre-storing three-dimensional feature points on the charging pile; wherein the three-dimensional feature points are pre-calibrated based on preset feature information; the controller is configured to: obtain a target image corresponding to the charging pile captured by the camera assembly, and extract feature points from the target image based on the preset feature information to obtain two-dimensional feature points; and calibrate the camera assembly parameters based on the correspondence between the two-dimensional feature points and the three-dimensional feature points. The self-moving device system provided in the present application can achieve rapid and accurate calibration of camera assembly parameters based on the correspondence between the three-dimensional feature points of the charging pile and the two-dimensional feature points of the charging pile image under the same feature information, thereby solving the problem of inaccurate calibration of camera assembly parameters caused by structural factors such as aging of the lens or circuit board of the camera assembly, and helping to improve the working accuracy of the camera assembly.
[0028] An object of the present application is to provide a self-moving lawn mower based on target following.
[0029] In order to achieve the above objectives, this application adopts the following technical solutions:
[0030] A self-propelled lawn mower based on target following comprises: a body; a walking wheel assembly configured to support the body; a camera assembly mounted to the body and configured to capture images of the self-propelled lawn mower's periphery; and a controller communicatively connected to the camera assembly, the controller being configured to: perform target detection of preset categories on a current image captured by the camera assembly to obtain candidate detection bounding boxes; determine one of the candidate detection bounding boxes as a target detection bounding box; and control the movement of the walking wheel assembly based on the target detection bounding box.
[0031] In some embodiments, the camera assembly includes a binocular camera including a left camera and a right camera.
[0032] In some embodiments, the current image is an RGB image or a grayscale image captured by the left camera or the right camera.
[0033] In some embodiments, the candidate detection bounding boxes are multiple bounding boxes; determining one from the candidate detection bounding boxes as the target detection bounding box includes: determining a current depth map corresponding to the current image based on the image captured by the binocular camera; determining the candidate depth information corresponding to each candidate detection bounding box according to the current depth map; performing information statistics on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box; and using the candidate detection bounding box corresponding to the minimum value in the depth statistical information as the target detection bounding box.
[0034] In some embodiments, determining a current depth map corresponding to a current image based on images captured by a binocular camera includes: using an RGB image or a grayscale image captured by a left camera of the binocular camera as a first image, and using an RGB image or a grayscale image captured by a right camera of the binocular camera as a second image; determining a current disparity map between the first image and the second image; and determining a current depth map corresponding to the current image based on the current disparity map.
[0035] In some embodiments, determining the candidate depth information corresponding to each candidate detection bounding box according to the current depth map includes: pixel-aligning the current depth map with each candidate detection bounding box; and determining the candidate depth information corresponding to each candidate detection bounding box based on the current depth map after pixel alignment.
[0036] In some embodiments, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: determining the depth average of the candidate depth information corresponding to each candidate detection bounding box; and using the depth average as the depth statistical information corresponding to the candidate detection bounding box.
[0037] In some embodiments, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: determining the depth median of the candidate depth information corresponding to each candidate detection bounding box; and using the depth median as the depth statistical information corresponding to the candidate detection bounding box.
[0038] In some embodiments, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: for each candidate detection bounding box, the depth information at the center pixel position of the candidate detection bounding box is obtained as the depth statistical information corresponding to the candidate detection bounding box.
[0039] In some embodiments, the candidate detection bounding boxes are multiple bounding boxes; determining one from the candidate detection bounding boxes as the target detection bounding box includes: obtaining a reference image, the reference image including a reference object under a preset category; performing target detection of the preset category on the reference image to obtain a reference detection bounding box; respectively determining the similarity between the reference detection bounding box and each candidate detection bounding box; and selecting the candidate detection bounding box with the greatest similarity as the target detection bounding box.
[0040] In some embodiments, controlling the movement of the walking wheel assembly based on the target detection bounding box includes: determining target depth information corresponding to the target detection bounding box; converting the target detection bounding box into three-dimensional space according to the target depth information to obtain target point cloud information corresponding to the target detection bounding box; and controlling the movement of the walking wheel assembly based on the target point cloud information.
[0041] In some embodiments, converting the target detection bounding box into a three-dimensional space according to the target depth information includes: obtaining camera internal parameters of a binocular camera; and converting the target detection bounding box into a three-dimensional space according to the target depth information and the camera internal parameters.
[0042] In some embodiments, controlling the movement of the walking wheel assembly based on the target point cloud information includes: determining the target center of mass coordinates corresponding to the target point cloud information; and controlling the movement of the walking wheel assembly according to the target center of mass coordinates.
[0043] In some embodiments, the default category is person.
[0044] In some embodiments, performing target detection of a preset category on the current image captured by the camera component includes: performing target detection of a preset category on the current image captured by the camera component based on a preset deep learning model; wherein the preset deep learning model is one of CNN, R-CNN, SSD and YOLO.
[0045] The benefits of this application are:
[0046] The present application provides a self-propelled lawn mower based on target following, comprising: a body; a walking wheel assembly configured to support the body; a camera assembly mounted to the body and configured to capture images of the surrounding area of the self-propelled lawn mower; a controller in communication with the camera assembly, the controller being configured to: perform target detection of a preset category on the current image captured by the camera assembly to obtain candidate detection bounding boxes; determine one of the candidate detection bounding boxes as a target detection bounding box; and control the movement of the walking wheel assembly based on the target detection bounding box. The present application utilizes the camera assembly of the self-propelled lawn mower to perform target detection of a preset category, and determines the following object of the self-propelled lawn mower based on the target detection result, so as to determine the boundary of the working area of the self-propelled lawn mower according to the target following result. This eliminates the need for manually deploying physical bounding lines, effectively saves manpower and material resources, and improves the efficiency of the self-propelled lawn mower in recognizing the boundary of the working area.
[0047] One object of the present application is to provide a method for determining the posture of a charging pile of a self-mobile equipment system.
[0048] In order to achieve the above objectives, this application adopts the following technical solutions:
[0049] A method for determining the position and posture of a charging pile of an autonomous mobile device system, wherein the autonomous mobile device system includes the autonomous mobile device, the autonomous mobile device includes a camera assembly and an auxiliary device, the auxiliary device includes a laser sensor, or the auxiliary device includes an external positioning unit and a memory, the memory stores a preset position and a three-dimensional model of the charging pile, and the position and posture determination method includes:
[0050] Collect images through the camera component;
[0051] Determine the rough position and auxiliary information of the charging pile relative to the mobile device based on the image and auxiliary equipment;
[0052] Based on the rough pose and auxiliary information, the precise pose of the charging pile relative to the mobile device is calculated.
[0053] In some embodiments, if the auxiliary device includes a laser sensor;
[0054] Based on the image and auxiliary equipment, the approximate position and auxiliary information of the charging pile relative to the mobile device are determined, including:
[0055] Recognize the landmark information of charging piles from images;
[0056] Calculate the rough position of the charging pile relative to the mobile device based on the landmark information;
[0057] Emitting laser light through a laser sensor;
[0058] Calculate the yaw angle between the laser and the charging pile and use the yaw angle as auxiliary information.
[0059] In some embodiments, if the auxiliary device includes an external positioning unit and a memory;
[0060] Based on the image and auxiliary equipment, the approximate position and auxiliary information of the charging pile relative to the mobile device are determined, including:
[0061] Obtain the current location of the mobile device through an external positioning unit;
[0062] Use deep learning programs to detect whether there are charging stations in the image;
[0063] When the confidence level of the charging pile detected from the image is greater than or equal to a preset threshold, a rough position of the charging pile relative to the mobile device is calculated based on the preset position of the charging pile and the current position of the mobile device;
[0064] Feature points are calculated based on the image and the three-dimensional model of the charging pile, and the feature points are used as auxiliary information.
[0065] In some embodiments, calculating the yaw angle between the laser and the charging pile includes:
[0066] According to the landmark information, the target laser point cloud data corresponding to the charging pile is filtered out;
[0067] The yaw angle is calculated using the target laser point cloud data.
[0068] In some embodiments, the identifying information is an identification code.
[0069] In some embodiments, the landmark information is a QR code.
[0070] In some embodiments, the external positioning unit includes a satellite positioning device;
[0071] The current location of the mobile device is obtained through an external positioning unit, including:
[0072] The current location is determined by receiving satellite signals through a satellite positioning device.
[0073] In some embodiments, calculating feature points based on the image and the three-dimensional model of the charging station includes:
[0074] Obtaining a rendering image of the charging pile according to the three-dimensional model and image of the charging pile;
[0075] Perform feature matching between the image and the charging pile rendering image to calculate feature points.
[0076] In some embodiments, performing feature matching between the image and the charging station rendering image includes:
[0077] Use SIFT or SURF to perform feature matching between the image and the charging station rendering image.
[0078] In some embodiments, calculating the precise position of the charging station relative to the mobile device based on the rough position and auxiliary information includes:
[0079] The scale information of the charging pile is fixed by using the rough pose and combined with the feature points to calculate the precise pose of the charging pile relative to the mobile device in the image.
[0080] The benefit of the present application lies in that: by combining the camera assembly and auxiliary equipment, the rough posture and auxiliary information of the charging pile relative to the self-mobile device are obtained, and then the precise posture of the charging pile relative to the self-mobile device is obtained based on the rough posture and auxiliary information. The auxiliary information obtained by the auxiliary equipment is used to further accurately optimize the rough posture obtained according to the camera assembly, thereby improving the accuracy of the charging pile posture determination, and thus ensuring the charging effectiveness of the self-mobile device. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] FIG1 is a schematic diagram of a self-propelled equipment system provided by the present application;
[0082] FIG2 is a schematic diagram of a self-propelled lawn mower provided by the present application;
[0083] FIG3 is a schematic diagram of a charging pile provided in this application;
[0084] FIG4 is a flow chart of a camera component parameter configuration (parameter calibration) provided by the present application;
[0085] FIG5 is a schematic diagram of a pinhole imaging model provided by the present application;
[0086] FIG6 is a schematic diagram of a simplified model of FIG5 provided by this application;
[0087] FIG7 is a flow chart of another camera component parameter configuration (working mode determination) provided by the present application;
[0088] FIG8 is a module schematic diagram of a controller provided by the present application;
[0089] FIG9 is a module schematic diagram of another controller provided by the present application;
[0090] FIG10 is a flow chart of a target following method provided by the present application;
[0091] FIG11 is a flowchart of another target following method provided by the present application;
[0092] FIG12 is a module schematic diagram of another controller provided by the present application;
[0093] FIG13 is a schematic diagram of a method for determining the posture of a charging pile of a mobile equipment system according to an embodiment of the present application;
[0094] FIG14 is a schematic diagram of a method for determining the posture of a charging pile of a self-mobile equipment system according to an embodiment of the present application;
[0095] FIG15 is a schematic diagram of a method for determining the posture of a charging pile of a self-mobile equipment system according to an embodiment of the present application;
[0096] FIG16 is a module diagram of a device for determining the posture of a charging pile of a self-mobile equipment system provided in an embodiment of the present application.
[0097] Reference numerals:
[0098] 10. Self-equipment system; 100. Self-equipment; 200. Charging station;
[0099] 110. Body; 120. Travel wheel assembly; 130. Camera assembly; 140. Controller; 150. LiDAR; 160. External positioning unit; 170. Communication module; 180. Working assembly. DETAILED DESCRIPTION
[0100] Before any embodiments of the present application are explained in detail, it is to be understood that the application is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the foregoing drawings.
[0101] In this application, the terms "comprises," "includes," "has," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0102] In this application, the term "and / or" describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " in this application generally indicates that the related objects are in an "and / or" relationship.
[0103] In this application, the terms "connect," "combine," "couple," and "install" may refer to direct connection, combination, coupling, or installation, or indirect connection, combination, coupling, or installation. For example, a direct connection refers to two parts or components being connected together without an intermediary, and an indirect connection refers to two parts or components being connected to at least one intermediary, with the two parts or components being connected via the intermediary. Furthermore, "connect" and "couple" are not limited to physical or mechanical connections or couplings and may include electrical connections or couplings.
[0104] In this application, it will be understood by those skilled in the art that relative terms (e.g., "about," "approximately," "substantially," etc.) used in conjunction with quantities or conditions include the values and have the meaning indicated by the context. For example, the relative terms include at least the degree of error associated with the measurement of a specific value, the tolerance caused by manufacturing, assembly, use, etc. associated with a specific value. Such terms should also be considered to disclose a range defined by the absolute values of the two endpoints. Relative terms may refer to plus or minus a certain percentage (e.g., 1%, 5%, 10% or more) of the indicated value. Numerical values that do not use relative terms should also be disclosed as specific values with tolerances. In addition, "substantially" may refer to plus or minus a certain degree (e.g., 1 degree, 5 degrees, 10 degrees or more) on the basis of the indicated angle when expressing a relative angular position relationship (e.g., substantially parallel, substantially perpendicular).
[0105] In this application, it will be understood by those skilled in the art that the function performed by an assembly can be performed by one assembly, multiple assemblies, one part, or multiple parts. Similarly, the function performed by a part can also be performed by one part, one assembly, or a combination of multiple parts.
[0106] In the present application, the terms "upper", "lower", "left", "right", "front", "back" and other directional words are described based on the orientation and positional relationship shown in the accompanying drawings, and should not be understood as limiting the embodiments of the present application. In addition, in the context, it is also necessary to understand that when it is mentioned that an element is connected to another element "upper" or "lower", it can not only be directly connected to the other element "upper" or "lower", but also be indirectly connected to the other element "upper" or "lower" through an intermediate element. It should also be understood that directional words such as upper side, lower side, left side, right side, front side, back side, etc. not only represent the positive orientation, but can also be understood as the lateral orientation. For example, below can include directly below, lower left, lower right, lower front and lower back, etc.
[0107] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0108] In this application, the terms "controller," "processor," "central processing unit," "CPU," and "MCU" are used interchangeably. Where a unit "controller," "processor," "central processing unit," "CPU," or "MCU" is used to perform a particular function, unless otherwise specified, the function may be performed by a single unit or multiple units.
[0109] In this application, the terms "device", "module" or "unit" can be implemented in the form of hardware or software to achieve specific functions.
[0110] In this application, the terms "calculate", "judge", "control", "determine", "identify", etc. refer to the operations and processes of a computer system or similar electronic computing device (e.g., controller, processor, etc.).
[0111] The technical solution proposed in this application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0112] Referring to Figure 1, the self-moving device system 10 in this embodiment includes a self-moving device 100 and a charging pile 200. Among them, the self-moving device 100 may refer to an intelligent device that can realize automatic positioning and movement, and the following is an illustration of a self-moving lawn mower. Referring to Figure 2, the self-moving lawn mower includes a body 110, a walking wheel assembly 120, a camera assembly 130 and a controller 140. Among them, the walking wheel assembly 120 can be configured to support the body 110; the camera assembly 130 is installed to the body 110 and can be configured to collect images around the self-moving lawn mower; the controller 140 is located inside the self-moving lawn mower and is communicatively connected to the camera assembly 130 (such as an electrical connection). Exemplarily, the camera assembly can be a camera or a camera. Referring to Figure 3, the charging pile 200 can be understood as an external charging device with a fixed size, which can be configured to charge the self-moving device 100 (such as a self-moving lawn mower).
[0113] It should be noted that the three-dimensional feature points on the charging pile are pre-stored in the controller of the self-mobile device, for example, stored in a register or memory. The three-dimensional feature points are pre-calibrated based on preset feature information. The preset feature information may refer to feature information pre-set according to actual application requirements. Optionally, the preset feature information includes structural information and additional mark information. Optionally, the structural information includes at least one of the following: edge information, corner information and texture information. Optionally, the additional mark information includes preset pattern information, wherein the preset pattern information includes QR code information, barcode information, etc. Exemplarily, the three-dimensional feature points can be represented by world coordinate information. Specifically, a world coordinate system can be established in advance, and the charging pile with pre-calibrated three-dimensional feature points based on the preset feature information can be projected into the world coordinate system, thereby obtaining the spatial coordinate information corresponding to each three-dimensional feature point, and then storing the spatial coordinate information in the controller of the self-mobile device.
[0114] 4 , the controller of the mobile device is configured to perform the following steps A1-A2:
[0115] A1, obtain the target image corresponding to the charging pile captured by the camera component, and extract feature points of the target image based on preset feature information to obtain two-dimensional feature points.
[0116] Specifically, when a mobile device walks near a charging pile, the camera component in the mobile device can be used to capture an image of the charging pile to obtain a target image corresponding to the charging pile. Feature point detection and extraction are then performed on the target image based on the preset feature information to obtain two-dimensional feature points in the target image. The two-dimensional feature points have the same number as the three-dimensional feature points pre-stored in the controller, and the two-dimensional feature points can be represented by two-dimensional coordinate information. Specifically, a pixel coordinate system can be established as a reference using the target image as a reference. For example, a pixel coordinate system can be established with the upper left corner of the target image as the center and the horizontal pixel direction (i.e., the width) and vertical pixel direction (i.e., the height) of the target image as the horizontal and vertical axes, respectively. This allows the two-dimensional coordinate information of each two-dimensional feature point in the pixel coordinate system to be obtained.
[0117] It should be noted that the feature point detection method in this embodiment is not specifically limited and can be flexibly selected according to actual application requirements. For example, for edge information detection, a first-order detection algorithm (such as Roberts operator, Prewitt operator, Sobel operator, Canny operator, etc.) can be used, or a second-order detection algorithm (such as Laplacian operator, etc.) can be used; for corner detection, SIFT (Scale Invariant Feature Transform), Harris, SURF or FAST (features from accelerated segment test) and other algorithms can be used; for texture detection, GLCM (Gray-level co-occurrence matrix) or LBP (Local Binary Pattern) and other methods can be used.
[0118] A2, based on the correspondence between two-dimensional feature points and three-dimensional feature points, the camera component parameters are calibrated.
[0119] In this embodiment, since both 2D and 3D feature points are obtained based on preset feature information, the only difference is the feature point dimensions. This difference in feature point dimensions is caused by the camera component capturing the image, and the parameters of the camera component play a key role. Therefore, the camera component can be calibrated based on the correspondence between the 2D and 3D feature points under the same preset feature information. The parameters may include internal parameters, which may optionally include focal length and optical center.
[0120] Referring to Figure 5, if the camera (camera assembly) is considered as a pinhole, the correspondence between two-dimensional feature points and three-dimensional feature points can be described by the pinhole imaging model. Specifically, a three-dimensional feature point P in the world coordinate system is projected onto the physical imaging plane through the optical center O of the camera to become the corresponding two-dimensional feature point P'. Here, Oxyz is the camera coordinate system, with the z-axis pointing directly in front of the camera, the x-axis pointing to the right, and the y-axis pointing downward. The distance between the physical imaging plane and the optical center (i.e., the OO' length) is the focal length f of the camera.
[0121] Assuming that the coordinates of point P in the world coordinate system are (X, Y, Z), and the coordinates of point P' in the O'-x'-y' coordinate system are (X', Y'), the pinhole imaging model in Figure 5 can be simplified to the similar triangles in Figure 6. Based on Figure 6, the following relationship can be obtained: Then we can get the coordinate relationship between P and P': From this we can get the coordinates of point P' in the O'-x'-y' coordinate system:
[0122] Furthermore, it is necessary to transform the point P' on the physical imaging plane to the pixel plane, that is, transform the point P' from the O'-x'-y' coordinate system to the pixel coordinate system (o'-uv), where the origin o' is located in the upper left corner of the image, the u axis is parallel to the x axis to the right, and the v axis is parallel to the y axis downward. Assume that the coordinates of the corresponding point P' in the o'-uv coordinate system are (U, V), and the coordinates of the point P' are scaled by α times on the u axis and β times on the v axis, and the origin is translated by [c x ,c y ], then the coordinate relationship between P′ and P″ can be obtained as Substituting the coordinate relationship between P and P' into the above formula, we can get The corresponding matrix form is Among them, f x and f y The unit is pixel, and K is the intrinsic parameter matrix of the camera. Therefore, the intrinsic parameters of the camera component can be calibrated based on the correspondence between the two-dimensional feature points and the three-dimensional feature points under the same preset feature information.
[0123] In some embodiments, the camera assembly optionally includes a binocular camera, wherein the binocular camera includes a left camera and a right camera. Accordingly, the target image includes a first image and a second image, wherein the first image is an image of the charging pile captured by the left camera, and the second image is an image of the charging pile captured by the right camera. In this case, feature point extraction of the target image based on preset feature information to obtain two-dimensional feature points may specifically include the following steps B1-B2:
[0124] B1. Extract feature points of the first image based on preset feature information to obtain first image two-dimensional feature points corresponding to the first image.
[0125] The first image two-dimensional feature points may refer to the two-dimensional feature points corresponding to the first image obtained after feature point extraction based on the preset feature information. The feature point extraction method may refer to the relevant description in the above step A1 and will not be repeated here.
[0126] B2. Extract feature points of the second image based on the preset feature information to obtain two-dimensional feature points of the second image corresponding to the second image.
[0127] The two-dimensional feature points of the second image may refer to the two-dimensional feature points corresponding to the second image obtained after feature point extraction of the second image based on the preset feature information. The feature point extraction method may refer to the relevant description in the above step A1 and will not be repeated here.
[0128] It should be noted that the binocular camera can use grayscale mode or color mode (such as RGB). When the binocular camera operates in grayscale mode, the target image (including the first image and the second image) it captures is a grayscale image; when the binocular camera operates in RGB mode, the target image (including the first image and the second image) it captures is an RGB image. Furthermore, to improve the feature point extraction effect of the target image, if the target image (including the first image and the second image) captured by the binocular camera is an RGB image, it is necessary to first convert the RGB image into a grayscale image, and then perform feature point extraction on the grayscale image based on the preset feature information.
[0129] In some embodiments, optionally, step A2 may specifically include the following steps C1-C2:
[0130] C1: Determine a first intrinsic parameter of the left camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image.
[0131] Among them, the first internal parameter can refer to the internal parameter corresponding to the left camera, and the method for determining the first internal parameter can refer to the relevant description in the above step A2, which will not be repeated here.
[0132] C2: Determine a second intrinsic parameter of the right camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image.
[0133] Among them, the second internal parameter can refer to the internal parameter corresponding to the right camera, and the method for determining the second internal parameter can refer to the relevant description in the above step A2, which will not be repeated here.
[0134] In some embodiments, optionally, step A2 may further include the following steps D1-D3:
[0135] D1, based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image and the first intrinsic parameter, determine the first extrinsic parameter of the left camera; wherein the first extrinsic parameter is used to represent the relative posture information between the left camera and the charging pile.
[0136] For example, the first extrinsic parameter of the left camera can be determined based on the least squares method or the PnP (Perspective-n-Point) algorithm. The PnP algorithm represents the extrinsic parameter based on the rotation matrix and the translation vector. Taking the PnP algorithm as an example, the first extrinsic parameter of the left camera can be determined by the following formula:
[0137] Among them, p represents the coordinates of the point in the pixel coordinate system, P C Represents the coordinates of the point in the camera coordinate system, P WRepresents the coordinates of the point in the world coordinate system, ω represents the depth of the point, and K represents the intrinsic parameter matrix of the left camera. CW and It is used to characterize the pose transformation from the world coordinate system to the camera coordinate system (i.e., the first external parameter), where R CW Represents the rotation matrix from the world coordinate system to the camera coordinate system (converting the same vector in the world coordinate system to the camera coordinate system), Represents the corresponding translation vector (i.e., the vector from the origin of the camera coordinate system to the origin of the world coordinate system in the camera coordinate system).
[0138] Specifically, the first external parameter solution process of the left camera can be described as the following problem: The coordinates of n three-dimensional feature points in the world coordinate system are known to be And the two-dimensional feature points corresponding to these n three-dimensional feature points in the pixel coordinate system are p1, p2, ..., p n , and the intrinsic parameter matrix of the left camera is known to be K, solve the position R of the camera coordinate system relative to the world coordinate system CW and Exemplarily, the problem can be solved based on methods such as DLT (direct linear transformation), P3P, EPnP or BA (bundle adjustment). For details, please refer to the relevant implementation process in the relevant technology, which will not be described here.
[0139] D2, based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image and the second intrinsic parameter, determine the second extrinsic parameter of the right camera; wherein the second extrinsic parameter is used to represent the relative position information between the right camera and the charging pile.
[0140] The method for determining the second external parameter may refer to the relevant description in the above step D1, which will not be repeated here.
[0141] D3, determining a third extrinsic parameter of the binocular camera based on the first extrinsic parameter and the second extrinsic parameter; wherein the third extrinsic parameter is used to represent the relative pose information between the left camera and the right camera.
[0142] It can be understood that since the first extrinsic parameter and the second extrinsic parameter respectively represent the relative posture information between the left camera and the right camera and the charging pile, when the first extrinsic parameter and the second extrinsic parameter are known, the relative posture information between the left camera and the right camera (that is, the third extrinsic parameter of the binocular camera) can be quickly determined.
[0143] In some embodiments, optionally, the self-mobile device system further includes an infrared laser, and the camera assembly further includes a low-light camera.
[0144] Exemplarily, the low-light camera can be an RGB camera. Furthermore, the low-light camera can be set as an ultra-sensitive camera, and the minimum illumination is less than or equal to 0.01Lux, which helps to improve the quality of images taken at night when the ambient brightness is low. Specifically, the low-light camera can work during the day and at night. When the low-light camera works during the day, it can be used for AI semantic perception and positioning; because the color effect at night is not as good as during the day, when the low-light camera works at night, it cannot be used for AI semantic perception, and can only be used for positioning, which can specifically include charging pile visual positioning (such as positioning the target or the relative position of the target) and visual SLAM (such as positioning itself or the absolute position), etc. Furthermore, the infrared laser can also be a near-infrared laser projector.
[0145] 7 , the controller is further configured to perform the following steps E1-E2:
[0146] E1, determine the working scene of the mobile device, which is daytime or nighttime.
[0147] In some embodiments, optionally, step E1 may specifically include the following steps: determining a working scene of the mobile device based on the working time of the mobile device or the brightness of the environment in which the mobile device is located.
[0148] Specifically, a first mapping relationship between the working hours of the self-mobile device and the working scene can be pre-set according to the actual application scenario, and a second mapping relationship between the ambient brightness of the self-mobile device and the working scene can be set. For example, the first mapping relationship can be set as: [7:00-19:00] is daytime; [19:00-7:00] is nighttime. The second mapping relationship can be set as: ambient brightness greater than a preset brightness threshold is daytime; ambient brightness less than or equal to the preset brightness threshold is nighttime. The preset brightness threshold can refer to an ambient brightness reference value pre-set according to actual application requirements.
[0149] Based on the pre-set first mapping relationship, the current working time of the self-mobile device can be determined first, and the first mapping relationship can be searched according to the current working time, from which the working scene that matches the current working time is determined as the current working scene of the self-mobile device. Based on the pre-set second mapping relationship, the current ambient brightness of the self-mobile device can be determined first, and the second mapping relationship can be searched according to the current ambient brightness, from which the working scene that matches the current ambient brightness is determined as the current working scene of the self-mobile device. Based on the pre-set first mapping relationship and the second mapping relationship, one of the methods can be selected to determine the current working scene of the self-mobile device.
[0150] E2, determines the working mode of the camera component in the mobile device according to the working scenario.
[0151] After determining the working scenario of the self-moving device, the working mode of the camera component in the self-moving device can be determined according to the working scenario, so that the special needs of the self-moving device for positioning (such as visual SLAM) and obstacle avoidance (visual obstacle detection) functions can be simultaneously realized in different working scenarios.
[0152] In some embodiments, optionally, step E2 may specifically include the following steps F1-F2:
[0153] F1. If the working scene is daytime, determine that the working mode of the camera assembly is a low-light camera and a binocular camera in a first state, or only a binocular camera in the first state; wherein the first state is that the infrared laser is in an off state.
[0154] Specifically, if the working scene is during the day, passive binocular can be achieved by turning off the infrared laser. At this time, the binocular camera can be used to calculate the visual SLAM and depth map of the charging pile. Therefore, the positioning and obstacle avoidance functions of the self-moving device during the day can be achieved by using only the binocular camera. Among them, the depth map is mainly used for visual obstacle detection and can also assist in visual SLAM. Furthermore, on the basis of passive binocular, a low-light camera can also be used in combination to combine positioning information in different dimensions, which helps to improve the positioning accuracy of the self-moving device.
[0155] F2: If the working scene is at night, the working mode of the camera assembly is determined to be a low-light camera and a binocular camera in a second state; wherein the second state is that the infrared laser is turned on.
[0156] Specifically, if the working scene is at night, active binocular can be achieved by turning on the infrared laser. In this case, the binocular camera images are speckle images and cannot be used to calculate the visual SLAM of the charging pile. It can only be used to calculate the depth map, which is not able to meet the positioning function required by the self-moving device at night. If you want to achieve the positioning and obstacle avoidance functions of the self-moving device at night, you need to use a low-light camera in addition to the active binocular, and use the low-light camera to provide the visual SLAM information of the charging pile.
[0157] Referring to Figure 8 , the controller 140 in the mobile device may specifically include a 2D feature point extraction module 1401 and a camera component parameter calibration module 1402. Specifically, the 2D feature point extraction module 1401 is configured to obtain a target image corresponding to a charging station captured by the camera component and extract feature points from the target image based on preset feature information to obtain 2D feature points. The camera component parameter calibration module 1402 is configured to calibrate the camera component parameters based on the correspondence between the 2D feature points and the 3D feature points.
[0158] The three-dimensional feature points are pre-calibrated feature points on the charging pile based on preset feature information stored in the controller. Optionally, the preset feature information includes structural information and additional marking information. The structural information includes at least one of the following: edge information, corner point information, and texture information. The additional marking information includes preset pattern information, which includes QR code information, barcode information, and the like.
[0159] In some embodiments, the two-dimensional feature point extraction module 1401 is optionally configured to: extract feature points from the first image based on preset feature information to obtain two-dimensional feature points of the first image corresponding to the first image; and extract feature points from the second image based on the preset feature information to obtain two-dimensional feature points of the second image corresponding to the second image. The first image is an image of the charging pile captured by the left camera in the binocular camera (camera assembly), and the second image is an image of the charging pile captured by the right camera in the binocular camera (camera assembly).
[0160] In some embodiments, the camera assembly parameter calibration module 1402 is optionally configured to determine a first intrinsic parameter of the left camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image; and to determine a second intrinsic parameter of the right camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image. The intrinsic parameters include focal length and optical center.
[0161] In some embodiments, optionally, the camera component parameter calibration module 1402 is further configured to: determine a first external parameter of the left camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the first image and the first internal parameter; wherein the first external parameter is used to characterize the relative posture information between the left camera and the charging pile; determine a second external parameter of the right camera based on the correspondence between the two-dimensional feature points and the three-dimensional feature points of the second image and the second internal parameter; wherein the second external parameter is used to characterize the relative posture information between the right camera and the charging pile; determine a third external parameter of the binocular camera based on the first external parameter and the second external parameter; wherein the third external parameter is used to characterize the relative posture information between the left camera and the right camera.
[0162] 9 , the controller 140 in the mobile device may further include a mobile device working scene determination module 1403 and a camera component working mode determination module 1404. The mobile device working scene determination module 1403 is configured to determine the working scene of the mobile device based on the working time of the mobile device or the brightness of the environment in which the mobile device is located. The camera component working mode determination module 1404 is configured to determine, if the working scene is daytime, the working mode of the camera component as a low-light camera and a binocular camera in a first state, or only a binocular camera in the first state; wherein the first state is that the infrared laser is in an off state; and if the working scene is nighttime, the working mode of the camera component as a low-light camera and a binocular camera in a second state; wherein the second state is that the infrared laser is in an on state.
[0163] Another application scenario of the present application is described as follows: a target object of a preset category moves along the boundary of the working area of the self-moving lawn mower, and the camera component of the self-moving lawn mower collects an image of the target object at a certain detection frequency (other objects may exist in the image), and the controller 140 of the self-moving lawn mower performs target detection on the collected image, and the object that the self-moving lawn mower needs to follow (i.e., the target object) is determined based on the target detection result. The boundary of the working area of the self-moving lawn mower can be effectively identified through the target tracking result of the self-moving lawn mower for the follow-up object.
[0164] Among them, the working area boundary, detection frequency and preset category can be pre-set according to actual application needs, and this application does not make specific restrictions on this. Exemplarily, the preset category can be a person or other object. Specifically, if the preset category is a person, then the person can be directly moved along the working area boundary of the self-moving lawn mower, and the self-moving lawn mower can perform target detection and target tracking on the person, so that the self-moving lawn mower can quickly and conveniently identify the working area boundary. If the preset category is other items, such as self-moving devices or static markers (such as hats, armbands, etc.) that can be controlled by the outside, the self-moving device can be controlled by the person to move along the working area boundary of the self-moving lawn mower, or a person wearing a static marker can be moved along the working area boundary of the self-moving lawn mower, and the self-moving lawn mower can perform target detection and target tracking on the self-moving device or static marker, so that the self-moving lawn mower can quickly and conveniently identify the working area boundary. It can be understood that setting the preset category to a person is the simplest and most efficient case.
[0165] 10 , the controller 140 of the self-propelled lawn mower is configured to perform the following steps G1-G3:
[0166] G1, performs target detection of preset categories on the current image captured by the camera component to obtain candidate detection bounding boxes.
[0167] The current image may refer to an image of the surrounding area of the self-propelled lawn mower captured by the camera assembly at the current moment, and the current image includes a target object of a preset category. Optionally, the preset category is a person. A candidate detection bounding box can be understood as an image region of interest containing the target object, obtained by performing target detection of the preset category on the current image, and can be used to represent the location of the target object of the preset category in the current image. It should be noted that the current image may contain one or more target objects of the preset category, and the number of target objects is consistent with the number of candidate detection bounding boxes. That is, if the current image contains only one target object of the preset category, then target detection of the preset category will result in one candidate detection bounding box; if the current image contains multiple target objects of the preset category, then target detection of the preset category will result in multiple candidate detection bounding boxes. The size and shape of the candidate detection bounding box can be predetermined based on actual application requirements and are not specifically limited in this embodiment. For example, the shape of the candidate detection bounding box can be set to be rectangular, square, circular, or elliptical.
[0168] In some embodiments, optionally, the camera assembly includes a binocular camera, wherein the binocular camera includes a left camera and a right camera. Accordingly, the current image can be an RGB image or a grayscale image captured by the left camera or the right camera.
[0169] In addition, the camera assembly may also include other cameras in addition to the binocular camera. For example, the other camera may be a low-light camera. Accordingly, the current image may be an RGB image or grayscale image captured by the low-light camera, or an RGB image or grayscale image captured by the left camera or the right camera of the binocular camera.
[0170] In some embodiments, optionally, performing target detection of a preset category on the current image captured by the camera component includes: performing target detection of a preset category on the current image captured by the camera component based on a preset deep learning model; wherein the preset deep learning model is one of CNN, R-CNN, SSD and YOLO.
[0171] Among them, the preset deep learning model may refer to a pre-trained deep learning model that can be used to detect targets of preset categories in images. For example, taking the preset category of pedestrians as an example, the transfer learning technology can be used to use one's own pedestrian image data set on a target detection basic model to obtain a preset deep learning model that meets the prediction accuracy and speed through model training, or the trained pedestrian target detection model can be directly obtained as the preset deep learning model. Among them, the input of the preset deep learning model is an image containing pedestrians, and the output is the detection bounding box corresponding to the pedestrian in the image. Therefore, when using the preset deep learning model to perform target detection on pedestrians in the image, it is only necessary to input the current image captured by the camera component into the preset deep learning model, and the output of the model is the candidate detection bounding box corresponding to each pedestrian in the current image.
[0172] Specifically, the preset deep learning model can be one of CNN, R-CNN, SSD or YOLO. Among them, CNN (Convolutional Neural Networks) is a convolutional neural network, which specifically includes a convolution layer (mainly used for feature extraction), a pooling layer (mainly used for downsampling) and a fully connected layer (mainly used for classification). R-CNN is a regional convolutional neural network, which applies a regional recommendation strategy on the basis of a convolutional neural network to form a bottom-up target positioning model. SSD (Single Shot MultiBox Detector) is a multi-scale target detection algorithm that uses a convolutional neural network to extract features and selects different feature layers for target detection. YOLO (You Only Look Once) is another target detection algorithm. Its core idea is to convert the target detection task into a regression problem, and simultaneously locate and classify the target through a single neural network, thereby achieving real-time and efficient target detection.
[0173] G2, determine one of the candidate detection bounding boxes as the target detection bounding box.
[0174] In this embodiment, after obtaining candidate detection bounding boxes corresponding to a preset category in the current image, one of the candidate detection bounding boxes needs to be determined as the target detection bounding box. Specifically, if there is only one candidate detection bounding box in the current image, that candidate detection bounding box can be directly determined as the target detection bounding box. If there are multiple candidate detection bounding boxes in the current image, one of the multiple candidate detection bounding boxes needs to be selected as the target detection bounding box. The target object of the preset category corresponding to the target detection bounding box is the target object that the self-propelled lawn mower needs to follow.
[0175] In some embodiments, optionally, the candidate detection bounding boxes are multiple bounding boxes; determining one from the candidate detection bounding boxes as the target detection bounding box includes: obtaining a reference image, the reference image including a reference object under a preset category; performing target detection of a preset category on the reference image to obtain a reference detection bounding box; respectively determining the similarity between the reference detection bounding box and each candidate detection bounding box; and selecting the candidate detection bounding box with the greatest similarity as the target detection bounding box.
[0176] In this embodiment, a user-friendly interactive interface can be used to allow the user to pre-specify a follow-up object that meets a preset category as a reference object, and an image of the user-specified reference object can be taken as a reference image. When the controller 140 determines the target detection bounding box from the candidate detection bounding boxes, it is necessary to first obtain a reference image and perform target detection of a preset category on the reference image using a preset deep learning model to obtain a reference detection bounding box. The similarity between the reference detection bounding box and each candidate detection bounding box is then calculated, and the similarities are sorted in descending order, and the candidate detection bounding box with the greatest similarity is selected as the target detection bounding box. It can be understood that the greater the similarity between the reference detection bounding box and the candidate detection bounding box, the higher the degree of similarity between the target object corresponding to the candidate detection bounding box and the reference object. Therefore, it can be considered that the target object corresponding to the candidate detection bounding box with the greatest similarity is the reference object. At this time, the candidate detection bounding box with the greatest similarity can be used as the target detection bounding box, thereby ensuring that the self-propelled lawn mower can follow the reference object specified by the user.
[0177] In some embodiments, optionally, the candidate detection bounding boxes are multiple bounding boxes; determining one from the candidate detection bounding boxes as the target detection bounding box includes: determining a current depth map corresponding to the current image based on the image captured by the binocular camera; determining the candidate depth information corresponding to each candidate detection bounding box according to the current depth map; performing information statistics on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box; and using the candidate detection bounding box corresponding to the minimum value in the depth statistical information as the target detection bounding box.
[0178] In this embodiment, based on the depth map of the current image, a target object of a preset category that is closest to the self-moving lawn mower can be selected as the actual follow-up object. Each pixel value in the depth map represents the distance between the pixel and the camera component (i.e., depth information), and the depth information can provide distance information between the target object of the preset category and the camera component, which helps to better characterize the position and posture of the target object of the preset category. It should be noted that the smaller the pixel value in the depth map, the closer the distance between the pixel and the camera component, that is, it can be reflected that the distance between the target object of the preset category corresponding to the pixel and the camera component is closer.
[0179] Specifically, it is first necessary to determine a current depth map corresponding to the current image based on the images captured by the binocular camera. Optionally, determining the current depth map corresponding to the current image based on the images captured by the binocular camera includes: using an RGB image or grayscale image captured by the left camera of the binocular camera as a first image, and using an RGB image or grayscale image captured by the right camera of the binocular camera as a second image; determining a current disparity map between the first image and the second image; and determining a current depth map corresponding to the current image based on the current disparity map.
[0180] Specifically, the intrinsic and extrinsic parameters of the binocular camera need to be pre-calibrated. Intrinsic parameters can include focal length, optical center, and other parameters, while extrinsic parameters are used to determine the position and orientation of the binocular camera in space (i.e., pose information). First, the RGB image or grayscale image captured by the left camera of the binocular camera is determined as the first image, and the RGB image or grayscale image captured by the right camera of the binocular camera is determined as the second image. The first and second images are RGB images or grayscale images captured at different viewing angles for the same detection area. Image preprocessing operations such as dedistortion and grayscale conversion can then be performed on the first and second images, respectively. After preprocessing, corresponding pixel points in the first and second images can be matched. By matching corresponding pixels in the two images, the distance between each corresponding pixel can be obtained, thereby obtaining a current disparity map between the first and second images. Based on the current disparity map, the depth information at each pixel position can be calculated in combination with the intrinsic and extrinsic parameters of the binocular camera. Finally, the depth information at each pixel position is mapped to the current image to generate a current depth map corresponding to the current image.
[0181] After determining a current depth map corresponding to the current image, candidate depth information corresponding to each candidate detection bounding box can be determined based on the current depth map. Optionally, determining the candidate depth information corresponding to each candidate detection bounding box based on the current depth map includes: pixel-aligning the current depth map with each candidate detection bounding box; and determining the candidate depth information corresponding to each candidate detection bounding box based on the current depth map after pixel alignment.
[0182] It should be noted that in order to currently match the depth map with the candidate detection bounding box in the current image, it is necessary to ensure that the depth map and the current image have been spatially positioned and calibrated. Specifically, it is first necessary to align the current depth map with each candidate detection bounding box separately. Exemplarily, the left camera of the binocular camera can be used as a reference, and the intrinsic and extrinsic parameters of the binocular camera are first calibrated. Then, based on the spatial transformation of the internal and external parameters of the binocular camera, each candidate detection bounding box is mapped to the current depth map, thereby achieving pixel alignment. Then, the depth information corresponding to each candidate detection bounding box can be obtained from the current depth map after pixel alignment as candidate depth information. It can be understood that since the candidate detection bounding box includes multiple pixels, the candidate depth information corresponding to each candidate detection bounding box also includes multiple depth information.
[0183] After determining the candidate depth information corresponding to each candidate detection bounding box, each candidate depth information can be further statistically analyzed to obtain the depth statistical information corresponding to each candidate detection bounding box. It should be noted that the depth statistical information is a statistical value that can objectively reflect the depth information of each candidate detection bounding box as a whole. For example, the parameter corresponding to the depth statistical information can be an average value, a median, or a representative depth value, etc., which can be flexibly set according to actual application requirements, and this embodiment does not limit this.
[0184] In some embodiments, optionally, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: for each candidate depth information corresponding to the candidate detection bounding box, determining the depth average of the candidate depth information; and using the depth average as the depth statistical information corresponding to the candidate detection bounding box.
[0185] Specifically, for each candidate depth information corresponding to the candidate detection bounding box, the average value of each candidate depth information is calculated as the depth average value, and the depth average value is used as the depth statistical information corresponding to each candidate detection bounding box.
[0186] In some embodiments, optionally, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: determining the depth median of the candidate depth information corresponding to each candidate detection bounding box; and using the depth median as the depth statistical information corresponding to the candidate detection bounding box.
[0187] Specifically, for each candidate depth information corresponding to the candidate detection bounding box, the median of each candidate depth information is calculated as the depth median, and the depth median is used as the depth statistical information corresponding to each candidate detection bounding box.
[0188] In some embodiments, optionally, information statistics are performed on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: for each candidate detection bounding box, obtaining the depth information at the center pixel position of the candidate detection bounding box as the depth statistical information corresponding to the candidate detection bounding box.
[0189] Specifically, the center pixel position of each candidate detection bounding box is first located, and then the depth information at the center pixel position is respectively obtained as the depth statistical information corresponding to each candidate detection bounding box.
[0190] After obtaining the depth statistics corresponding to each candidate detection bounding box, the depth statistics can be sorted in order of arrival from the smallest to the smallest, and the smallest depth statistics can be determined. The candidate detection bounding box corresponding to the smallest depth statistics is then used as the target detection bounding box. In this way, the detection bounding box closest to the camera assembly (i.e., the target detection bounding box) can be quickly and accurately located. The target object of the preset category corresponding to the target detection bounding box is the actual object being followed by the self-propelled lawn mower.
[0191] G3 controls the movement of the walking wheel assembly based on the target detection bounding box.
[0192] In this embodiment, after determining the target detection bounding box, the walking wheel assembly can be controlled to move based on the target detection bounding box at different detection moments, so that the self-propelled lawn mower can identify the boundary of the working area by performing target detection and target tracking on the target object of the preset category corresponding to the target detection bounding box.
[0193] In some embodiments, optionally, controlling the movement of the walking wheel assembly based on the target detection bounding box includes: determining target depth information corresponding to the target detection bounding box; converting the target detection bounding box into three-dimensional space according to the target depth information to obtain target point cloud information corresponding to the target detection bounding box; and controlling the movement of the walking wheel assembly based on the target point cloud information.
[0194] Specifically, based on the determination of the current depth map corresponding to the current image based on the image captured by the binocular camera, the current depth map can be pixel-aligned with the target detection bounding box, and the target depth information corresponding to the target detection bounding box can be determined based on the pixel-aligned current depth map. The target depth information can be used to represent the depth information corresponding to each pixel in the target detection bounding box.
[0195] After determining the target depth information, the target detection bounding box can be converted into a three-dimensional space according to the target depth information to obtain the target point cloud information corresponding to the target detection bounding box. Among them, the target point cloud information can be used to reflect the point cloud representation of the target object of the preset category corresponding to the target detection bounding box in the three-dimensional space. It can be understood that since the target detection bounding box includes multiple pixels, the corresponding target point cloud information includes multiple point clouds. Exemplarily, the target point cloud information can be represented based on the spatial coordinates under the world coordinate system (established in advance according to actual application requirements). In this case, one target object will correspond to multiple spatial coordinates.
[0196] In some embodiments, optionally, converting the target detection bounding box into a three-dimensional space according to the target depth information includes: obtaining camera internal parameters of a binocular camera; and converting the target detection bounding box into a three-dimensional space according to the target depth information and the camera internal parameters.
[0197] Specifically, we first need to obtain the camera internal parameters (i.e., camera intrinsic parameters) of the binocular camera, such as focal length and optical center. For example, the camera internal parameters can be described in the form of an intrinsic parameter matrix, where the intrinsic parameter matrix can be expressed as After obtaining the intrinsic parameter matrix of the binocular camera, the formula Convert the target detection bounding box into three-dimensional space. Where p represents the coordinates of the point in the pixel coordinate system, P W represents the coordinates of the point in the world coordinate system, ω represents the depth of the point (determined by the target depth information), K represents the internal parameter matrix, R CW Represents the rotation matrix from the world coordinate system to the camera coordinate system (converting the same vector in the world coordinate system to the camera coordinate system), Represents the corresponding translation vector (i.e., the vector from the origin of the camera coordinate system to the origin of the world coordinate system in the camera coordinate system). Among them, the pixel coordinate system (o′-uv) can be established with the upper left corner of the image as the origin o′, and the horizontal pixel direction and vertical pixel direction as the u axis and v axis respectively. The camera coordinate system (O′-xyz) can be established with the optical center of the camera as the origin O and the front of the camera as the z axis. Based on the above formula, by solving P W The target point cloud information corresponding to the target detection bounding box can be obtained.
[0198] After determining the target point cloud information, the walking wheel assembly can be controlled to move based on the target point cloud information. Optionally, controlling the walking wheel assembly to move based on the target point cloud information includes: determining the target center of mass coordinates corresponding to the target point cloud information; and controlling the walking wheel assembly to move according to the target center of mass coordinates.
[0199] For example, assume that the target point cloud information is represented by spatial coordinates in the world coordinate system. For example, the target point cloud information includes the spatial coordinates corresponding to four point clouds, which are represented as (x1, y1, z1), (x2, y2, z2), (x3, y3, z3), and (x4, y4, z4) in the world coordinate system (OXYZ). In this case, the average spatial coordinates corresponding to the four point clouds on each coordinate axis are calculated, and the target center of mass coordinates are determined based on the average spatial coordinates on each coordinate axis. Specifically, the average value of the coordinates on the X-axis can be expressed as (x1+x2+x3+x4) / 4, the average value of the coordinates on the Y-axis can be expressed as (y1+y2+y3+y4) / 4, and the average value of the coordinates on the Z-axis can be expressed as (z1+z2+z3+z4) / 4, so that the target center of mass coordinates are ((x1+x2+x3+x4) / 4, (y1+y2+y3+y4) / 4, (z1+z2+z3+z4) / 4). According to the above process, the target center of mass coordinates corresponding to each detection moment are calculated respectively, and then the controller 140 can be used to adjust the movement of the self-moving lawn mower according to the target center of mass coordinates, so that the self-moving lawn mower can quickly and accurately locate the boundary of the working area by following the target center of mass coordinates.
[0200] In some embodiments, referring to FIG. 11 , the controller 140 of the self-propelled lawn mower is configured to perform the following steps J1-J8:
[0201] J1, performs target detection of preset categories on the current image captured by the camera component to obtain candidate detection bounding boxes; wherein the camera component includes a binocular camera.
[0202] J2, determine one of the candidate detection bounding boxes as the target detection bounding box.
[0203] J3, determine a current depth map corresponding to the current image based on the image captured by the binocular camera.
[0204] J4, determine the target depth information corresponding to the target detection bounding box according to the current depth map.
[0205] J5, obtains the camera internal parameters of the binocular camera.
[0206] J6, converts the target detection bounding box into three-dimensional space according to the target depth information and the internal parameters of the camera, and obtains the target point cloud information corresponding to the target detection bounding box.
[0207] J7, determine the target center of mass coordinates corresponding to the target point cloud information.
[0208] J8 controls the movement of the walking wheel assembly according to the target center of mass coordinates.
[0209] Referring to Figure 12 , the controller 140 of the self-propelled lawn mower may specifically include a candidate detection bounding box determination module 141, a target detection bounding box determination module 142, and a wheel assembly movement control module 143. Specifically, the candidate detection bounding box determination module 141 is configured to perform target detection of a preset category on the current image captured by the camera assembly to obtain a candidate detection bounding box. The target detection bounding box determination module 142 is configured to determine one of the candidate detection bounding boxes as the target detection bounding box. The wheel assembly movement control module 143 is configured to control the movement of the wheel assembly based on the target detection bounding box.
[0210] In some embodiments, optionally, the target detection bounding box determination module 142 includes: a current depth map determination unit, which is configured to determine a current depth map corresponding to the current image based on the image captured by the binocular camera; a candidate depth information determination unit, which is configured to determine the candidate depth information corresponding to each candidate detection bounding box according to the current depth map; a depth statistical information determination unit, which is configured to perform information statistics on each candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box; and a target detection bounding box determination unit, which is configured to use the candidate detection bounding box corresponding to the minimum value in the depth statistical information as the target detection bounding box.
[0211] In some embodiments, optionally, the current depth map determination unit is specifically configured to: use the RGB image or grayscale image captured by the left camera of the binocular camera as the first image, and use the RGB image or grayscale image captured by the right camera of the binocular camera as the second image; determine a current disparity map between the first image and the second image; and determine a current depth map corresponding to the current image based on the current disparity map.
[0212] In some embodiments, optionally, the candidate depth information determining unit is specifically configured to: perform pixel alignment on the current depth map and each candidate detection bounding box respectively; and determine the candidate depth information corresponding to each candidate detection bounding box based on the pixel-aligned current depth map.
[0213] In some embodiments, optionally, the depth statistics information determining unit is configured to: determine the average depth value of the candidate depth information corresponding to each candidate detection bounding box; and use the average depth value as the depth statistics information corresponding to the candidate detection bounding box.
[0214] In some embodiments, optionally, the depth statistics information determination unit is further configured to: determine the depth median of the candidate depth information corresponding to each candidate detection bounding box; and use the depth median as the depth statistics information corresponding to the candidate detection bounding box.
[0215] In some embodiments, optionally, the depth statistics information determining unit is further configured to: for each candidate detection bounding box, respectively obtain depth information at a center pixel position of the candidate detection bounding box as the depth statistics information corresponding to the candidate detection bounding box.
[0216] In some embodiments, optionally, the target detection bounding box determination module 142 is further configured to: obtain a reference image, wherein the reference image includes a reference object under a preset category; perform target detection of a preset category on the reference image to obtain a reference detection bounding box; respectively determine the similarity between the reference detection bounding box and each candidate detection bounding box; and use the candidate detection bounding box with the greatest similarity as the target detection bounding box.
[0217] In some embodiments, optionally, the walking wheel assembly movement control module 143 includes: a target depth information determination unit, configured to determine the target depth information corresponding to the target detection bounding box; a target point cloud information determination unit, configured to convert the target detection bounding box into a three-dimensional space according to the target depth information to obtain the target point cloud information corresponding to the target detection bounding box; and a walking wheel assembly movement control unit, configured to control the movement of the walking wheel assembly based on the target point cloud information.
[0218] In some embodiments, optionally, the target point cloud information determining unit is specifically configured to: obtain camera internal parameters of the binocular camera; and convert the target detection bounding box into a three-dimensional space according to the target depth information and the camera internal parameters.
[0219] In some embodiments, optionally, the walking wheel assembly movement control unit is specifically configured to: determine the target center of mass coordinates corresponding to the target point cloud information; and control the walking wheel assembly to move according to the target center of mass coordinates.
[0220] In some embodiments, optionally, the candidate detection bounding box determination module is specifically configured to: perform target detection of a preset category on the current image captured by the camera component based on a preset deep learning model; wherein the preset deep learning model is one of CNN, R-CNN, SSD and YOLO.
[0221] Figure 13 is a schematic diagram of a method for determining the position of a charging station in a self-mobile device system according to an embodiment of the present application. This embodiment is applicable to situations where the method for determining the position of a charging station is optimized. The method can be performed by a device for determining the position of a charging station in the self-mobile device system. This device can be implemented in software and / or hardware and integrated into an electronic device. The electronic device involved in this embodiment can be a device with computing capabilities, such as a server. In one embodiment, the device for determining the position of a charging station in the self-mobile device system is the controller 140 of the self-mobile device.
[0222] The self-propelled device system includes a self-propelled device, which includes a visual sensor and an auxiliary device. The visual sensor is the aforementioned camera assembly 130, and the auxiliary device includes a laser sensor, or an external positioning unit and a memory that stores a preset position and a three-dimensional model of the charging station.
[0223] Among them, the self-moving device is a device that can move autonomously to perform corresponding tasks, for example, the self-moving device can be an unmanned vehicle or an automatic lawn mower. The visual sensor is an instrument that uses optical elements and imaging devices to obtain image information of the external environment. For example, the visual sensor can be a camera or a video camera set on the self-moving device. The auxiliary device refers to a device used to assist in positioning the position information of the self-moving device, which may include a laser sensor, or an external positioning unit and a memory, wherein the laser sensor can obtain laser point cloud data by emitting laser to obtain auxiliary information related to the position information of the self-moving device; the external positioning unit refers to a device that obtains the position information of the self-moving device through an external positioning device, and obtains auxiliary information related to the position information of the self-moving device through the preset position and three-dimensional model of the charging pile stored in the external positioning unit and the memory.
[0224] Specifically, the self-mobile device system 10 includes at least a self-mobile device 100 and a charging station 200. The self-mobile device is a smart lawn mower. FIG2 shows a schematic diagram of the structure of the smart lawn mower. In FIG2, reference numeral 150 represents a laser sensor, which may be a lidar. Reference numeral 130 represents a camera assembly, such as a camera module, which may include a binocular camera. Reference numeral 170 represents a communication module, such as a radio, through which the lawn mower communicates with other devices in the system, such as an RTK base station. Reference numeral 160 represents an external positioning unit, such as a satellite positioning device. Reference numeral 120 represents a wheel assembly. Reference numeral 180 represents a working assembly, which is used to support the main working functions of the smart lawn mower, such as mowing. FIG3 shows a schematic diagram of the structure of the charging station 200.
[0225] Specifically, referring to FIG13 , the posture determination method specifically includes the following steps:
[0226] S110 , collecting images through the camera assembly 130 .
[0227] After the mobile device is started, the camera assembly 130 on the mobile device is turned on, and image information around the mobile device 100 is obtained through the camera assembly 130. For example, after the mobile lawn mower is started, the camera on the lawn mower is turned on at the same time. During the movement of the mobile lawn mower, the camera captures images or videos around the lawn mower as images.
[0228] S120: Determine the rough position and auxiliary information of the charging pile relative to the mobile device based on the image and the auxiliary device.
[0229] Among them, the charging pile refers to the charging device associated with the self-mobile device. The charging pile can be pre-set at a certain location. When the power of the self-mobile device is less than the preset power threshold, it can move to the charging pile location for charging. Therefore, in order to ensure that the self-mobile device accurately reaches the charging pile location, it is necessary to determine the posture information of the charging pile relative to the self-mobile device, so that the self-mobile device can move to the charging pile location according to the posture information. The posture includes position information and direction information.
[0230] Specifically, positioning is performed through the visual dimension, and the rough posture of the charging pile relative to the self-mobile device is determined based on the image, that is, the charging pile is roughly positioned through visual information; auxiliary positioning is performed through other dimensions, and auxiliary information of the charging pile relative to the self-mobile device is determined based on other information obtained by the auxiliary device. The auxiliary information describes the posture information of the charging pile relative to the self-mobile device from dimensions other than vision.
[0231] For example, if the auxiliary device is a laser sensor, the laser is emitted by the laser sensor, and the yaw angle between the laser and the charging pile is calculated as auxiliary information; if the auxiliary device is an external positioning unit and a memory, the feature points are calculated as auxiliary information based on the image and the three-dimensional model of the charging pile.
[0232] S130: Calculate the precise position of the charging pile relative to the mobile device based on the rough position and auxiliary information.
[0233] Since the rough pose is mainly used to determine the pose of the charging pile relative to the mobile device from a visual perspective, the auxiliary information is used to determine the pose of the charging pile relative to the mobile device by adding a non-visual perspective. In addition, since direct positioning based solely on the visual perspective will lead to certain errors in the positioning results, the rough pose and auxiliary information are combined to calculate the precise pose of the charging pile relative to the mobile device, which is equivalent to combining two dimensions of information to accurately locate the charging pile, thereby improving the accuracy of the charging pile pose determination.
[0234] The benefit of the present application lies in that: by combining the camera assembly and auxiliary equipment, the rough posture and auxiliary information of the charging pile relative to the self-mobile device are obtained, and then the precise posture of the charging pile relative to the self-mobile device is obtained based on the rough posture and auxiliary information. The auxiliary information obtained by the auxiliary equipment is used to further accurately optimize the rough posture obtained according to the camera assembly, thereby improving the accuracy of the charging pile posture determination, and thus ensuring the charging effectiveness of the self-mobile device.
[0235] Figure 14 is a schematic diagram of a method for determining the posture of a charging station of a self-mobile device system according to an embodiment of the present application. This embodiment is a further refinement of the above technical solution. The auxiliary equipment includes a laser sensor. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. As shown in Figure 14, posture determination includes the following:
[0236] S210: Capture images through a camera assembly.
[0237] S220: Identify the landmark information of the charging pile from the image.
[0238] Among them, the landmark information refers to the marker information pre-deployed on the charging pile. For example, the landmark information can be a barcode or some pre-set special marks. In a feasible embodiment, the landmark information is an identification code. The identification code refers to an identification code that includes location information. For example, the coordinate information of the four corner points of the identification code can be obtained by scanning and identifying the identification code. In a feasible embodiment, the landmark information is a QR code. That is, by scanning the QR code, the world coordinate system coordinate information of the positions of the four corner points of the QR code can be obtained. The landmark information can be deployed in the front panel of the charging pile represented by the dotted box as shown in Figure 3.
[0239] Specifically, an image is acquired by a camera component of a mobile device, and computer vision technology and image processing algorithms are used to perform feature matching on the image according to the feature information of the landmark information, to determine the image including the landmark information, and to obtain the identification code coordinate information included in the landmark information by identifying the landmark information in the image.
[0240] S230: Calculate a rough position of the charging pile relative to the mobile device based on the landmark information.
[0241] Using a pre-set visual technology algorithm or deep learning model, the rough pose of the charging pile relative to the robot is calculated based on the identified landmark information, where the rough pose provides the approximate position and direction of the charging pile in the robot coordinate system.
[0242] For example, the landmark information includes the coordinates of the four corner points of the identification code in the world coordinate system. This coordinate information is converted through PNP (Perspective-n-Point) to obtain the coordinates of the four corner points of the identification code in the robot coordinate system. Combined with the pre-deployed position information of the identification code on the charging pile, the approximate position and orientation of the charging pile in the robot coordinate system is obtained as the rough pose.
[0243] S240, emitting laser light through the laser sensor.
[0244] After the mobile device is turned on, the laser sensor deployed on the mobile device is turned on, and the laser sensor emits laser to collect laser point cloud data from the environment around the mobile device.
[0245] S250: Calculate the yaw angle between the laser and the charging pile, and use the yaw angle as auxiliary information.
[0246] Because the laser sensor's laser beam is emitted in a straight line and is located in front of the mobile device, the angle between the laser beam and the charging station is the yaw angle, or the yaw angle between the mobile device and the charging station. This yaw angle indicates the direction between the mobile device and the charging station, and is therefore used as auxiliary information.
[0247] Exemplarily, at the moment of capturing the image in which the landmark information of the charging pile is recognized, the yaw angle between the laser and the charging pile is calculated.
[0248] In a feasible embodiment, calculating the yaw angle between the laser and the charging pile includes:
[0249] According to the landmark information, the target laser point cloud data corresponding to the charging pile is filtered out;
[0250] The yaw angle is calculated using the target laser point cloud data.
[0251] When calculating the yaw angle between the laser and the charging station, the landmark information is used as a reference. The laser point cloud data near the landmark information is filtered from the large amount of laser point cloud data to obtain the target laser point cloud data corresponding to the charging station. The yaw angle is calculated based on the target laser point cloud data. The yaw angle provides information about the charging station's orientation in the robot coordinate system.
[0252] Exemplarily, the landmark information includes the coordinate information of the four corner points of the identification code in the world coordinate system. The coordinate information of the identification code is converted through PNP (Perspective-n-Point) to obtain the coordinate information of the four corner points of the identification code in the laser coordinate system. Based on the coordinate information of the four corner points of the identification code in the laser coordinate system, the identification code laser point position data corresponding to the identification code is determined from all laser point cloud data. Combined with the pre-deployed position information of the identification code on the charging pile, the laser point cloud data within a certain range around the identification code is determined as the target laser point cloud data corresponding to the charging pile based on the identification code laser point cloud data. For example, the laser point cloud data within a preset distance range extended on both sides of the identification code laser point cloud data in the horizontal direction is determined as the target laser point cloud data corresponding to the charging pile. The preset distance needs to be determined based on the pre-deployed position information of the identification code on the charging pile. After obtaining the target laser point cloud data, if the laser sensor is a line laser sensor, a straight line fitting is performed on the target laser point cloud data, and the normal direction of the fitted straight line is the yaw angle; if the laser sensor is a surface laser sensor, a surface fitting is performed on the target laser point cloud data, and the normal vector direction of the fitted surface is the yaw angle.
[0253] S260: Calculate the precise position of the charging pile relative to the mobile device based on the rough position and yaw angle.
[0254] Since the rough pose is obtained through visual positioning, and the yaw angle in the direction information obtained through vision is inaccurate due to defects in the camera component, the yaw angle is determined by laser point cloud data and used to replace the yaw angle in the rough pose to obtain the precise pose.
[0255] For example, the rough pose and yaw angle are used as the initial pose for pose optimization, and an optimization algorithm (such as least squares or gradient descent) is used to iteratively optimize the initial pose to obtain a more accurate pose of the charging pile relative to the lawn mower. The goal of pose optimization is to minimize the difference between the precise pose and the rough pose and yaw angle. The precise pose of the charging pile relative to the lawn mower is calculated and updated to the lawn mower's control system or navigation system to guide the lawn mower's actions.
[0256] The benefit of the present application lies in that: by combining the camera assembly and the laser sensor, the rough position and yaw angle of the charging pile relative to the self-moving device are obtained, and then the precise position of the charging pile relative to the self-moving device is obtained based on the rough position and yaw angle. The yaw angle obtained by the laser sensor is used to further accurately optimize the rough position obtained according to the camera assembly, thereby improving the accuracy of the charging pile position determination and ensuring the charging effectiveness of the self-moving device.
[0257] Figure 15 is a schematic diagram of a method for determining the posture of a charging station of a self-mobile device system according to an embodiment of the present application. This embodiment is a further refinement of the above technical solution. The auxiliary device includes an external positioning unit and a memory. The technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. As shown in Figure 15, posture determination includes the following:
[0258] S310: Capture images through a camera assembly.
[0259] S320: Obtain the current location of the mobile device through an external positioning unit.
[0260] The external positioning unit refers to a device that obtains the location information of the mobile device through an external positioning device. Specifically, the current location of the mobile device is obtained by positioning the external positioning unit on the mobile device. The external positioning unit can be a positioning device integrated with a positioning system.
[0261] In one possible embodiment, the external positioning unit includes a satellite positioning device;
[0262] The current location of the mobile device is obtained through an external positioning unit, including:
[0263] The current location is determined by receiving satellite signals through a satellite positioning device.
[0264] The term "satellite positioning device" refers to a device that uses a satellite positioning system for positioning. When installed in a mobile device, the device can receive satellite signals and determine the current real-time location of the mobile device. Specifically, the satellite positioning device can be a differential positioning system (RTK). Specifically, the current position is the position of the mobile device in a world coordinate system, which is the approximate position of the mobile device.
[0265] S330: Use a deep learning program to detect whether there is a charging pile in the image.
[0266] During the recharging process, the camera component continuously captures images and uses 2D object detection technology to process the images in real time to detect whether a charging station is present in the image. For example, 2D object detection technology is used to perform feature matching on the charging station's features, and the presence of a charging station is determined based on the matching results.
[0267] Specifically, deep learning training is performed in advance based on a training image set of charging piles to obtain a deep learning charging pile detection model. The deep learning charging pile detection model is used to detect images collected in real time to determine the detection information of the charging piles in the images. For example, the detection information of the charging piles includes the detection confidence of the charging piles in the image.
[0268] S340: When the confidence level of the charging pile detected from the image is greater than or equal to a preset threshold, calculate a rough position of the charging pile relative to the mobile device according to the preset position of the charging pile and the current position of the mobile device.
[0269] The confidence of the charging pile can represent the accuracy of the charging pile in the image, so the confidence of the charging pile in the real-time image is determined. If it is determined that the confidence of the charging pile is greater than or equal to the preset threshold, the frame image is determined to be the first frame target image, and the current position of the self-mobile device obtained by the external positioning unit is determined according to the acquisition time of the first frame target image. The preset position of the charging pile stored in the memory can be used to determine the position information of the pre-set charging pile in the world coordinate system. Combined with the current position of the self-mobile device, that is, the position information of the self-mobile device in the world coordinate system, the rough posture of the charging pile relative to the self-mobile device can be obtained, that is, the rough position information and rough direction information of the charging pile relative to the self-mobile device are obtained.
[0270] S350: Calculate feature points based on the image and the three-dimensional model of the charging pile, and use the feature points as auxiliary information.
[0271] A three-dimensional model of the charging pile is pre-stored in the memory, and the feature points of the charging pile in the image are determined through feature matching results between the three-dimensional model and the first frame target image.
[0272] In a feasible embodiment, calculating feature points based on the image and the three-dimensional model of the charging pile includes:
[0273] Obtaining a rendering image of the charging pile according to the three-dimensional model and image of the charging pile;
[0274] Perform feature matching between the image and the charging pile rendering image to calculate feature points.
[0275] The three-dimensional model of the charging pile includes reference charging pile images from multiple perspectives. The reference charging pile images from multiple perspectives are matched with the first frame target image, and the reference charging pile image that is closest to the charging pile acquisition perspective in the first frame target image is determined as the charging pile rendering image.
[0276] The first frame of the target image is used to match the charging pile rendering image to obtain the charging pile feature points. The charging pile feature points are corner points on the charging pile where the gradient change is greater than a preset gradient threshold.
[0277] In a feasible embodiment, performing feature matching between the image and the charging pile rendering image includes:
[0278] Use SIFT or SURF to perform feature matching between the image and the charging station rendering image.
[0279] The feature extraction algorithm used for feature matching between the first frame target image and the charging pile rendering image can be SIFT or SURF. Among them, SIFT (Scale-invariant feature transform) is a machine vision algorithm used to detect and describe local features in images. It finds extreme points in the spatial scale and extracts their position, scale, and rotation invariants. SURF feature extraction technology is a feature extraction method based on two-dimensional gradients. It is based on the Haar wavelet transform and uses the translation-invariant Haar wavelet to extract local feature points of the image. In the image processing process, SURF mainly uses two methods: direction of angle (DOA) and contrast measurement (CMT). DOA uses the direction of the gradient to describe the local features of the image, while CMT uses the degree of grayscale change of the image to describe the local features of the image.
[0280] The SURF detector uses the Direction of Gradients (DOA) and Contrast Metrics (CMT) to divide an image into multiple parts and perform specific processing on each part to detect key points in the image. Key points are typically distributed at the edges and corners of the image. The coordinates of the key points are used to distinguish similar points in the image and extract local features of the image.
[0281] S360: Calculate the precise position of the charging pile relative to the mobile device based on the rough position and feature points.
[0282] Based on the perspective information of the rendered image of the charging pile, the direction information of the charging pile relative to the mobile device is determined according to the position information of the feature points in the image. Combined with the position information in the rough pose, the precise pose of the charging pile relative to the mobile device is obtained.
[0283] In a feasible embodiment, calculating the precise position of the charging pile relative to the mobile device based on the rough position and auxiliary information includes:
[0284] The scale information of the charging pile is fixed by using the rough pose and combined with the feature points to calculate the precise pose of the charging pile relative to the mobile device in the image.
[0285] Since the scale information of the charging pile is not fixed when the direction information is determined by feature points, in order to ensure the accuracy of the direction information determination, it is necessary to first use the position information in the rough pose to fix the scale information of the charging pile. That is, the scale information of the charging pile in the charging pile rendering image and the first frame target image is fixed by the position information in the rough pose. Then, the direction information of the charging pile relative to the self-moving device is determined based on the feature points obtained by feature matching. Then, the direction information is used to replace the rough direction information in the rough pose to obtain the precise pose of the charging pile relative to the self-moving device.
[0286] Exemplarily, the scale information of the charging pile in the rendered image of the charging pile and the first frame target image is fixed using the position information in the rough pose, and then the direction information of the charging pile relative to the mobile device is determined based on the feature points obtained by feature matching. The difference between the direction information and the rough direction information in the rough pose is determined. If the difference is greater than a preset threshold, the next frame target image with a confidence level of the charging pile greater than or equal to the preset threshold is re-determined, and the precise pose is determined based on the next frame target image.
[0287] The benefit of the present application lies in that: by combining the camera assembly, the external positioning unit and the memory, the rough posture and feature points of the charging pile relative to the self-mobile device are obtained, and then the precise posture of the charging pile relative to the self-mobile device is obtained based on the rough posture and feature points. The feature points obtained by the external positioning unit and the memory are used to further accurately optimize the rough posture obtained according to the camera assembly, thereby improving the accuracy of the charging pile posture determination and ensuring the charging effectiveness of the self-mobile device.
[0288] Figure 16 is a schematic diagram of a module of a device for determining the position of a charging pile in a self-propelled device system provided by an embodiment of the present application. The self-propelled device system includes a self-propelled device, the self-propelled device includes a camera assembly and an auxiliary device, the auxiliary device includes a laser sensor, or the auxiliary device includes an external positioning unit and a memory, the memory stores the preset position and three-dimensional model of the charging pile. As shown in Figure 16, the device includes:
[0289] An image acquisition module 410 is configured to acquire images through the camera assembly;
[0290] a rough pose and auxiliary information determination module 420, configured to determine a rough pose and auxiliary information of the charging pile relative to the mobile device based on the image and the auxiliary device;
[0291] The precise posture determination module 430 is configured to calculate the precise posture of the charging pile relative to the mobile device based on the rough posture and the auxiliary information.
[0292] In some embodiments, if the auxiliary device includes a laser sensor;
[0293] Rough pose and auxiliary information determination module, including:
[0294] a landmark information recognition unit, configured to recognize landmark information of the charging pile from the image;
[0295] a first rough pose calculation unit, configured to calculate a rough pose of the charging pile relative to the mobile device according to the landmark information;
[0296] a laser emitting unit, configured to emit laser light through the laser sensor;
[0297] The first auxiliary information determination unit is configured to calculate a yaw angle between the laser and the charging pile, and use the yaw angle as the auxiliary information.
[0298] In some embodiments, if the auxiliary device includes an external positioning unit and a memory;
[0299] Rough pose and auxiliary information determination module, including:
[0300] A self-mobile device position determination unit, configured to obtain the current position of the self-mobile device through the external positioning unit;
[0301] a charging pile detection unit, configured to detect whether there is a charging pile in the image using a deep learning program;
[0302] a second rough pose calculation unit, configured to calculate a rough pose of the charging pile relative to the self-mobile device according to a preset position of the charging pile and a current position of the self-mobile device when a confidence level of the charging pile detected from the image is greater than or equal to a preset threshold;
[0303] The second auxiliary information determining unit is configured to calculate feature points based on the image and the three-dimensional model of the charging pile, and use the feature points as the auxiliary information.
[0304] In some embodiments, the first auxiliary information determining unit is specifically configured to:
[0305] Filter and obtain target laser point cloud data corresponding to the charging pile according to the landmark information;
[0306] The yaw angle is calculated using the target laser point cloud data.
[0307] In some embodiments, the identifying information is an identification code.
[0308] In some embodiments, the landmark information is a QR code.
[0309] In some embodiments, the external positioning unit includes a satellite positioning device;
[0310] The unit for determining the location of the mobile device is specifically configured as follows:
[0311] The current position is determined by receiving satellite signals through the satellite positioning device.
[0312] In some embodiments, the second auxiliary information determining unit is specifically configured to:
[0313] Obtaining a rendering image of the charging pile based on the three-dimensional model of the charging pile and the image;
[0314] Feature matching is performed between the image and the charging pile rendering image to calculate feature points.
[0315] In some embodiments, the second auxiliary information determination unit includes a feature matching subunit, which is specifically configured to:
[0316] SIFT or SURF is used to perform feature matching between the image and the charging pile rendering image.
[0317] In some embodiments, the precise pose determination module is specifically configured to:
[0318] The rough posture fixed charging pile scale information is used in combination with the feature points to calculate the precise posture of the charging pile relative to the mobile device in the image.
[0319] The device for determining the posture of a charging pile of a self-equipped device system provided in an embodiment of the present application can execute the method for determining the posture of a charging pile of a self-equipped device system provided in any embodiment of the present application, and has functional modules and beneficial effects corresponding to the execution method.
[0320] The acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.
[0321] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0322] The above shows and describes the basic principles, main features and advantages of this application. Those skilled in the art should understand that the above embodiments do not limit this application in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of this application.
Claims
1. A self - moving device system for camera component parameter configuration, comprising a self - moving device and a charging pile, where the charging pile is configured to charge the self - moving device; the self - moving device includes: Body; Traveling wheel assembly, configured to support the body; Camera assembly, mounted to the body, configured to collect images around the self - moving device; Controller, communicatively connected to the camera assembly, wherein the controller pre - stores three - dimensional feature points on the charging pile; wherein, the three - dimensional feature points are pre - calibrated based on preset feature information; The controller is configured to: Obtain a target image corresponding to the charging pile collected by the camera assembly, and perform feature point extraction on the target image based on the preset feature information to obtain two - dimensional feature points; Perform parameter calibration on the camera assembly based on the correspondence between the two - dimensional feature points and the three - dimensional feature points.
2. The self - moving device system according to claim 1, wherein, The camera assembly includes a binocular camera, and the binocular camera includes a left camera and a right camera.
3. The self-moving device system according to claim 2, wherein, The target image includes a first image and a second image, the first image is the image collected by the left camera, and the second image is the image collected by the right camera.
4. The self - moving device system according to claim 3, wherein, Performing feature point extraction on the target image based on the preset feature information to obtain two - dimensional feature points includes: Performing feature point extraction on the first image based on the preset feature information to obtain first - image two - dimensional feature points corresponding to the first image; and, Performing feature point extraction on the second image based on the preset feature information to obtain second - image two - dimensional feature points corresponding to the second image.
5. The self - moving device system according to claim 4, wherein, Performing parameter calibration on the camera assembly based on the correspondence between the two - dimensional feature points and the three - dimensional feature points includes: Determining a first internal parameter of the left camera based on the correspondence between the first - image two - dimensional feature points and the three - dimensional feature points; Determining a second internal parameter of the right camera based on the correspondence between the second - image two - dimensional feature points and the three - dimensional feature points.
6. The self - moving device system according to claim 5, wherein, The internal parameter includes focal length and optical center.
7. The self - moving device system according to claim 5, wherein, Performing parameter calibration on the camera assembly based on the correspondence between the two - dimensional feature points and the three - dimensional feature points includes: Determining a first external parameter of the left camera based on the correspondence between the first - image two - dimensional feature points and the three - dimensional feature points and the first internal parameter; wherein, the first external parameter is used to represent the relative pose information between the left camera and the charging pile; Determining a second external parameter of the right camera based on the correspondence between the second - image two - dimensional feature points and the three - dimensional feature points and the second internal parameter; wherein, the second external parameter is used to represent the relative pose information between the right camera and the charging pile; Determining a third external parameter of the binocular camera according to the first external parameter and the second external parameter; wherein, the third external parameter is used to represent the relative pose information between the left camera and the right camera.
8. The self-moving device system according to any one of claims 1-7, wherein, The preset feature information includes structural information and additional flag information.
9. The self - moving device system according to claim 8, wherein, The structural information includes at least one of the following: edge information, corner point information, and texture information.
10. The self - moving device system according to claim 8, wherein, The additional flag information includes preset pattern information, and the preset pattern information includes two - dimensional code information.
11. The self-moving device system according to claim 2, wherein, The self - moving device system further includes an infrared laser, and the camera assembly further includes a low - light camera.
12. The self-moving device system according to claim 11, wherein, The controller is further configured to: Determine the working scenario of the self - moving device, where the working scenario is day or night; Determine the working mode of the camera component in the self - moving device according to the working scenario.
13. The self - moving device system according to claim 12, wherein, Determine the working scenario of the self - moving device, including: Based on the working time of the self - moving device or the ambient brightness where the self - moving device is located, determine the working scenario of the self - moving device.
14. The self-moving device system according to claim 12 or 13, wherein, Determine the working mode of the camera component in the self - moving device according to the working scenario, including: If the working scenario is day, determine that the working mode of the camera component is a low - light camera and a binocular camera in the first state, or only the binocular camera in the first state; where the first state is that the infrared laser is in the off state; If the working scenario is night, determine that the working mode of the camera component is a low - light camera and a binocular camera in the second state; where the second state is that the infrared laser is in the on state.
15. A self - propelled lawn mower based on target following, comprising: Body; A walking wheel assembly, configured to support the body; A camera component, installed on the body, configured to collect images around the self - moving lawn mower; A controller, communicatively connected to the camera component, and the controller is configured to: Perform object detection of a preset category on the current image collected by the camera component to obtain candidate detection bounding boxes; Determine one of the candidate detection bounding boxes as the target detection bounding box; Control the walking wheel assembly to move based on the target detection bounding box.
16. The self-propelled lawn mower according to claim 15, wherein, The camera component includes a binocular camera, and the binocular camera includes a left camera and a right camera.
17. The self - propelled lawn mower according to claim 16, wherein, The current image is an RGB image or a grayscale image collected by the left camera or the right camera.
18. The self-propelled lawn mower according to claim 16, wherein, The candidate detection bounding boxes are multiple bounding boxes; Determine one of the candidate detection bounding boxes as the target detection bounding box, including: Determine a current depth map corresponding to the current image based on the images collected by the binocular camera; Determine candidate depth information corresponding to each candidate detection bounding box according to the current depth map; Perform information statistics on each of the candidate depth information respectively to obtain depth statistical information corresponding to each candidate detection bounding box; Take the candidate detection bounding box corresponding to the minimum value in the depth statistical information as the target detection bounding box.
19. The self-propelled lawn mower according to claim 18, wherein, Determine a current depth map corresponding to the current image based on the images collected by the binocular camera, including: Take the RGB image or grayscale image collected by the left camera of the binocular camera as the first image, and take the RGB image or grayscale image collected by the right camera of the binocular camera as the second image; Determine the current disparity map between the first image and the second image; Determine the current depth map corresponding to the current image according to the current disparity map.
20. The self-propelled lawn mower according to claim 19, wherein, Determine candidate depth information corresponding to each candidate detection bounding box according to the current depth map, including: Align the current depth map with each of the candidate detection bounding boxes pixel - by - pixel; Determine the candidate depth information corresponding to each candidate detection bounding box based on the current depth map after pixel alignment.
21. The self-propelled lawn mower according to claim 18, wherein, Perform information statistics on each of the candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: For the candidate depth information corresponding to each candidate detection bounding box, respectively determine the average depth of the candidate depth information; Use the average depth as the depth statistical information corresponding to the candidate detection bounding box.
22. The self-propelled lawn mower according to claim 18, wherein, Perform information statistics on each of the candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: For the candidate depth information corresponding to each candidate detection bounding box, respectively determine the median depth of the candidate depth information; Use the median depth as the depth statistical information corresponding to the candidate detection bounding box.
23. The self-propelled lawn mower according to claim 18, wherein, Perform information statistics on each of the candidate depth information to obtain depth statistical information corresponding to each candidate detection bounding box, including: For each candidate detection bounding box, respectively obtain the depth information at the center pixel position of the candidate detection bounding box as the depth statistical information corresponding to the candidate detection bounding box.
24. The self-propelled lawn mower according to claim 15, wherein, The candidate detection bounding boxes are multiple bounding boxes; Determine one of the candidate detection bounding boxes as the target detection bounding box, including: Obtain a reference image, which includes a reference object under the preset category; Perform target detection of the preset category on the reference image to obtain a reference detection bounding box; Respectively determine the similarity between the reference detection bounding box and each candidate detection bounding box; Use the candidate detection bounding box with the maximum similarity as the target detection bounding box.
25. The self-propelled lawn mower according to claim 18, wherein, Control the walking wheel assembly to move based on the target detection bounding box, including: Determine the target depth information corresponding to the target detection bounding box; Convert the target detection bounding box into three-dimensional space according to the target depth information to obtain the target point cloud information corresponding to the target detection bounding box; Control the walking wheel assembly to move based on the target point cloud information.
26. The self-propelled lawn mower according to claim 25, wherein, Convert the target detection bounding box into three-dimensional space according to the target depth information, including: Obtain the internal parameters of the camera of the binocular camera; Convert the target detection bounding box into three-dimensional space according to the target depth information and the internal parameters of the camera.
27. The self-propelled lawn mower according to claim 25, wherein, Control the walking wheel assembly to move based on the target point cloud information, including: Determine the target centroid coordinates corresponding to the target point cloud information; Control the walking wheel assembly to move according to the target centroid coordinates.
28. The self-propelled lawn mower according to any one of claims 15-27, wherein, The preset category is a person.
29. The self-propelled lawn mower according to any one of claims 15-27, wherein, Perform target detection of the preset category on the current image collected by the camera assembly, including: Perform target detection of the preset category on the current image collected by the camera assembly based on a preset deep learning model; wherein, the preset deep learning model is one of CNN, R-CNN, SSD, and YOLO.
30. A method for determining the pose of a charging pile of a self - moving device system, the self - moving device system includes a self - moving device, the self - moving device includes a camera assembly and an auxiliary device, the auxiliary device includes a laser sensor, or the auxiliary device includes an external positioning unit and a memory, and the memory stores the preset position and three - dimensional model of the charging pile. The pose determination method includes: Collect an image through the camera assembly; Determine the rough pose and auxiliary information of the charging pile relative to the self - moving device according to the image and the auxiliary device; Calculate the precise pose of the charging pile relative to the self - moving device according to the rough pose and the auxiliary information.
31. The pose determination method according to claim 30, wherein, If the auxiliary device includes a laser sensor; Determine the rough pose and auxiliary information of the charging pile relative to the self - moving device according to the image and the auxiliary device, including: Identify the landmark information of the charging pile from the image; Calculate the rough pose of the charging pile relative to the self - moving device according to the landmark information; Emit laser through the laser sensor; Calculate the yaw angle between the laser and the charging pile, and use the yaw angle as the auxiliary information.
32. The pose determination method according to claim 30, wherein, If the auxiliary device includes an external positioning unit and a memory; Determine the rough pose and auxiliary information of the charging pile relative to the self - moving device according to the image and the auxiliary device, including: Obtain the current position of the self - moving device through the external positioning unit; Use a deep - learning program to detect whether there is a charging pile in the image; When the confidence level of detecting the charging pile from the image is greater than or equal to a preset threshold, calculate the rough pose of the charging pile relative to the self - moving device according to the preset position of the charging pile and the current position of the self - moving device; Calculate feature points according to the image and the three - dimensional model of the charging pile, and use the feature points as the auxiliary information.
33. The pose determination method according to claim 31, wherein, Calculate the yaw angle between the laser and the charging pile, including: Filter out the target laser point cloud data corresponding to the charging pile according to the landmark information; Calculate the yaw angle using the target laser point cloud data.
34. The pose determination method according to claim 31, wherein, The landmark information is an identification code.
35. The pose determination method according to claim 34, wherein, The landmark information is a two - dimensional code.
36. The pose determination method according to claim 32, wherein, The external positioning unit includes a satellite positioning device; Obtain the current position of the self - moving device through the external positioning unit, including: Receive satellite signals through the satellite positioning device to determine the current position.
37. The pose determination method according to claim 32, wherein, Calculate feature points according to the image and the three - dimensional model of the charging pile, including: Obtain a rendered image of the charging pile according to the three - dimensional model of the charging pile and the image; Perform feature matching in the image and the rendered image of the charging pile to calculate feature points.
38. The pose determination method according to claim 37, wherein, Perform feature matching in the image and the rendered image of the charging pile, including: Use SIFT or SURF to perform feature matching in the image and the rendered image of the charging pile.
39. The pose determination method according to claim 32, wherein, Calculate the precise pose of the charging pile relative to the self - moving device according to the rough pose and the auxiliary information, including: Using the rough pose to fix the scale information of the charging pile, and combining with the feature points, calculate the accurate pose of the charging pile relative to the self-moving device in the image.
Citation Information
Patent Citations
Method and system for automatic butt joint of robot and charging pile
CN111679671A
Camera calibration method of sweeping robot, terminal and computer readable storage medium
CN112634377A
Self-moving device and control method thereof
CN116430838A
Charging base positioning method and system, self-moving device and storage medium
CN117115476A
Self-moving equipment, control method, control device and storage medium thereof
US20220007913A1