Method for generating texture map
By combining LiDAR and cameras to generate textured maps, the problem of marking ground types on maps for mobile devices has been solved, enabling more intuitive map interaction and automatic region segmentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2024-10-10
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to effectively mark different ground types, such as carpet and tile boundaries, on maps for mobile devices, making it difficult for users to define accessible or restricted areas.
A method combining LiDAR sensors and cameras is used to generate texture maps. By synchronizing depth information and images, texture maps of the ground surface are extracted and generated. Multi-view compositing and machine learning techniques are used to identify and segment ground types.
The generated textured maps are easier to understand and interact with. Users can define areas more intuitively, and the automatic segmentation method can learn passable or prohibited areas from user input, improving the accuracy and consistency of the map.
Smart Images

Figure CN122055680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for generating a textured map of a ground surface in an environment in which a mobile device (particularly a transport vehicle or robot that is at least partially automated) is moving or about to move, and the present invention relates to a data processing system, mobile device, and computer program for performing the method. Background Technology
[0002] Mobile devices, such as at least partially automated vehicles or robots, typically move within an environment, particularly in environments or work areas, such as residences, gardens, factory floors, streets, or in the air or water. One of the fundamental problems with such or other mobile devices is localization, that is, understanding how the environment looks, especially where obstacles or other objects are located, and their own (absolute) position. To this end, mobile devices can be equipped with various sensors, such as cameras, lidar sensors, or inertial sensors, which detect the motion of the environment and the mobile device in two or three dimensions, for example. This enables the mobile device to move locally, identify obstacles in a timely manner, and navigate around them. Summary of the Invention
[0003] According to the present invention, a method for generating textured maps, having the features of the independent patent claims, is proposed, along with a system, mobile device, and computer program for performing the method for data processing. Advantageous designs are the subject matter of the dependent claims and the description thereof below.
[0004] The present invention generally relates to mobile devices that move or are to be moved in an environment or in a work area therein, and particularly to generating texture maps of ground surfaces in such environments. A texture map is to be understood herein as a map in which the texture or type of a ground surface (e.g., carpet, flat ground, etc.) can be identified or seen, i.e., in the form of a photograph or image, for example.
[0005] Examples of such mobile devices (or mobile work devices) are: robots and / or drones and / or partially or (fully) automated (on land, in water, or in the air) vehicles. For example, robots could be considered as household robots such as vacuuming and / or mopping robots, floor or street cleaning devices, or lawnmower robots, or other so-called service robots; vehicles that are at least partially automated could be considered as passenger or freight transport vehicles (also known as industrial transport vehicles, for example, in warehouses), as well as air transport vehicles such as so-called drones, or water transport vehicles.
[0006] Such mobile devices typically include control and / or regulation units and drive units for moving the device, enabling it to move within an environment, such as along a path or trajectory. Furthermore, the mobile device may have one or more sensors that can detect the environment or information within it.
[0007] The invention will be illustrated below with a cleaning robot as a mobile device as an example, although the principle can also be applied to other types of mobile devices.
[0008] Mobile devices can create and / or use maps within their navigation within their environment. For example, household robots, such as cleaning robots, can interact with users using 2D map views. Here, for example, the structure of the environment can be seen on an artificially colored map, displaying obstacles and walls as boundaries or other objects. Maps can also be enriched with pictographs, displaying, for example, the location of docking stations, the robot's current position, the path traveled, etc. Users can also interactively place object placeholders, such as sofa or chair models, to improve the map.
[0009] However, it has been shown that one problem for users is the difficulty in marking different floor types (such as carpet, tiled boundaries between open kitchens and living rooms) on such maps. However, this kind of information about the floor surface can be important when users want to define specific "Go" or "No-Go" zones (i.e., specific areas in the environment where robots are allowed or not allowed to move).
[0010] Against this backdrop, a method for generating texture maps of ground surfaces in an environment, particularly indoors, is proposed. The mobile device should move within this environment. For this purpose, device and / or environmental information detected from the mobile device and / or the environment by at least one sensor of the mobile device is provided. The at least one sensor herein includes (at least) a depth sensor. As a depth sensor, a lidar sensor (or laser scanner) is particularly considered, and more specifically, a 2D lidar sensor, i.e., a lidar sensor that scans the environment on only one plane. This plane is typically parallel to the ground surface, at least when the mobile device is on a flat surface.
[0011] In principle, other types of depth sensors, such as time-of-flight cameras, can also be considered. Relatedly, mobile devices can also include one or more sensors other than depth sensors, such as inertial measurement units (IMUs), cameras, or odometry. These other sensors can be used, in particular, for the localization and / or navigation of mobile devices, for example, within the so-called SLAM (Simultaneous Localization and Mapping) framework.
[0012] In SLAM, there are various methods for representing maps and locations. Traditional methods for SLAM are typically based on geometric information, such as nodes and edges. Nodes and edges are usually components of the SLAM graph. Nodes and edges in a SLAM graph can have different designs; traditionally, nodes correspond, for example, to the pose (position and orientation) of a mobile device or a specific environmental feature at a specific point in time, while edges represent relative measurements between the mobile device and the environmental feature, or relative measurements between different poses of the mobile device at different points in time. The map of the environment in which the mobile device moves can then be determined or can be determined based on such a SLAM graph. The map (or SLAM graph) can be expanded or updated using each new dataset containing environmental information obtained from or based on the mobile device's sensors. The aforementioned map or 2D map for the user can also be generated here based on the SLAM map or SLAM graph.
[0013] In addition to the device and / or environmental information, images detected from the environment by a camera on the mobile device are also provided. The mobile device therefore has at least a camera in addition to a depth sensor; this is independent of any other sensors that may be present. The camera, or the images it detects, can also be used for the localization and / or navigation of the mobile device, for example, within the scope of SLAM. Visual SLAM is also discussed in this context.
[0014] Then, for one or more selected images provided, the following steps are performed: The image is synchronized spatially and / or temporally with at least a portion of device and / or environmental information (i.e., depth information) detected by the depth sensor to obtain synchronized depth information.
[0015] This part of the device and / or environmental information is therefore, for example, data from a lidar scan. A lidar scan typically corresponds to a set of points, known as a lidar point cloud. Here, each point corresponds to the location where the lidar sensor's laser beam is reflected from an object or obstacle in the environment. Therefore, each point also indicates the distance from the obstacle to the mobile device or lidar sensor. The point cloud as a whole also provides a depth map. Time-of-flight cameras provide similar type of depth information.
[0016] Synchronization of image and depth information specifically involves temporal synchronization (so-called pairing), and subsequent possible data transformations to place them in a common reference coordinate system for further processing. This is because the camera image and depth sensor data are typically not, or not always, detected at the same point in time or in a temporally synchronized manner. Since the mobile device may move between these two points in time, spatial synchronization may be required. Synchronized depth information therefore specifically refers to depth information that corresponds to the camera image in both time and the pose of the mobile device.
[0017] Then, based on the synchronized depth information, a ground image of the environment is extracted from the image. It's important to note that the camera on a mobile device is typically pointed forward (or possibly backward or to the side), which is often necessary for navigation and obstacle detection. From this point onward, the ground image should be understood as at least approximately showing (only) the ground surface, excluding areas of the environment significantly above ground level.
[0018] Then, based on the one or more ground images, a texture map of the ground surface of the environment is generated, and subsequently provided to the user, for example.
[0019] Compared to textureless representations, textured maps with photorealistic views, for example, are easier for users to understand and interact with. This facilitates manual spatial segmentation and region labeling (e.g., No-Go areas). Furthermore, the generated textured maps provide a foundation for automatic segmentation methods that can utilize machine learning techniques, for example, to learn Go-To or No-Go areas from previous user input.
[0020] Compared to other techniques for generating texture maps, the use of depth sensors, such as LiDAR scanning, offers an effective and efficient way to determine which points or areas are assigned to the ground surface or plane. This allows for the efficient creation of image boundaries and avoids mapping distorted objects in the texture map. This combination of sensors enables the creation of more consistent images of the ground surface or plane than when using only bird's-eye view (top-down view) camera images. Furthermore, this texture map can be created online (i.e., while a mobile device is running, even in real time; alternatively, it is conceivable to perform the method at least partially, for example, in the cloud or on the internet). Additionally, unlike a real-time view from above, a consistent texture map of the ground surface can be created using known locations from the SLAM system and stored, for example, for subsequent reference.
[0021] In one implementation, extracting a ground image includes the following steps: Determining a free region, and more specifically, determining the free region based on synchronized depth information. This free region here covers at least a portion of the ground surface of the environment. This could be, for example, a predetermined angular range in front of a mobile device. The depth information, or synchronized depth information, here allows for the specific selection of only one area free of obstacles or other objects (such as walls). Therefore, this free region here is particularly located in a plane, for example, a plane parallel to the ground surface, especially at a vertical distance corresponding to the height of the depth sensor above the ground.
[0022] In the specific case of a lidar sensor or a 2D lidar sensor, the free region can be defined, for example, as a set of points on straight lines corresponding to the detection direction of the lidar sensor, i.e., the lines on which the lidar sensor emits a laser beam. These points should not be confused with points in a lidar point cloud; rather, the points on the straight lines can be multiple points on each line, thus generally at least substantially uniformly distributed across the free region and representing that region.
[0023] The free region is then projected onto the (camera's) image, and image information from the image is assigned to at least a portion of the free region. The free region with the assigned image information is then used as the ground image for generating the texture map. In this way, depth information can be utilized particularly efficiently to generate or determine the ground image for the texture map.
[0024] Alternatively, extracting the ground image may also include the following steps. First, a perspective transformation is performed on the (camera's) image to obtain a transformed image. Furthermore, based on synchronized depth information, a free region is determined that covers at least a portion of the ground surface of the environment. In principle, this can be done as described above. Then, the portion of the transformed image corresponding to this free region is used as the ground image.
[0025] In one implementation, generating a texture map includes determining one or more image regions for which image information from multiple ground images exists. Such image regions may be, for example, individual image points and / or groups of image points (so-called pixels). Then, for one or more, preferably all, of the multiple image regions, a value for at least a portion of the image information is determined based on the multiple ground images. This value is then used in the texture map.
[0026] This can be referred to as multi-view ensemble. One goal of multi-view ensemble is specifically to create a globally consistent image of the ground surface from multiple individual images (multiple ground images). Each ground image is assigned a pose of the mobile device within a global map. To combine the individual ground images, averaging can be performed, for example, over a global pixel grid. For each pixel in this grid, all pixels of the ground image that fall within that grid can be collected or provided. Subsequently, for example, the color values of all these pixels can be averaged, where gamma correction can be considered to obtain a natural blend of the observed colors. Here, individual pixels can be weighted differently based on their distance from the camera of the mobile device. This is because pixel or image information mapped to areas closer to the camera generally has less distortion caused by inverse perspective mapping and is also less affected by pose estimation noise. In this way, a (final) value (e.g., a color value), as an average or otherwise derived value, can be found for use in the texture map.
[0027] Another method for determining such values (e.g., by averaging different colors) is to average these colors in, for example, a so-called LAB color space. Alternatively, after defining and applying quality criteria (e.g., camera distance or moderate brightness at the time of observation) to select one value from multiple pixels as the final value, only the "best" color in each pixel may be retained or used.
[0028] In one implementation, one or more selected images are selected from the provided images according to at least one of the criteria. Therefore, it is not necessary to use all images; instead, images can be selectively chosen or omitted to make the method faster and more efficient, or to improve the results themselves. Here, consider the following criteria for the respective images: The position of the mobile device relative to the last detected image has changed by less than a pre-given distance when detecting the image. Or, the orientation of the mobile device relative to the last detected image has changed by less than a pre-given angle when detecting the image. Here, position and orientation can also be referred to as pose. Therefore, for example, only images in which the mobile device has performed a limited number of translational and angular movements are included. In this way, errors due to insufficient external calibration and motion blur can be limited. It is also conceivable that the position or orientation has changed by more than one (and thus another) pre-given distance and / or pre-given angle.
[0029] It is also possible to consider whether the orientation of the mobile device is within a pre-defined range when detecting images, or whether the (potentially global) pose of the mobile device has not appeared in a predetermined number of last detected images. Therefore, for example, only images in which the mobile device or its camera is pointed in a specific direction can be considered to limit the influence of specular reflections from the light source on the final result. In practice, this is particularly applicable to mobile devices that perform coverage travel to complete a task (e.g., cleaning, vacuuming, and lawnmowing robots). This also allows determining when a pose has been visited and deciding whether a new image should be used at that location or the last image should be retained.
[0030] It is also possible to consider whether one or more image parameters meet pre-defined requirements. Therefore, image quality can be considered. For example, some fundamental parameters of the image itself can be evaluated to exclude images that are likely severely underexposed or overexposed and thus fail to provide useful information within the area of interest.
[0031] It can also be considered whether the distance to the obstacle in the direction detected by the camera is greater than a predetermined value, where this distance is determined specifically based on device and / or environmental information detected by the depth sensor. Therefore, for example, the decision to use an image can be made based on how much free space is in front of a moving device as can be identified using the depth sensor. In other words, images where the moving device is close to the object and therefore has little visible ground should not be used.
[0032] In one implementation, a segmented map of the environmental floor is also generated based on the texture map. For example, the floor texture can then be post-processed using a convolutional neural network (CNN) or other machine learning algorithm to deduce different types of floor coverings. This can, for example, help with the correct segmentation of rooms, since different rooms typically have different floor coverings; furthermore, it can also, for example, lead to recommendations on cleaning or cleaning methods based on the floor type (e.g., do not wet mop carpets).
[0033] In one implementation, active detection can also be considered. Here, the mobile device can be directly instructed to go to a specific location or position to detect new information or images for texture mapping.
[0034] As mentioned earlier, in addition to lidar sensors, other depth sensors can also be used, such as stereo cameras, depth cameras (based on time-of-flight or structured light), and ultrasonic sensors.
[0035] In one implementation, so-called stitching can be used. For example, the multi-view combination described above can be improved by using a stitching method for stitching panoramic images. This can, for example, result in a more convincing and consistent bottom or ground surface image.
[0036] In one implementation, reflection modeling can be performed. For example, the multi-view combination described above can also be designed to detect reflections on the ground. One possibility is to create, for example, multiple global pixel grids, where each pixel grid is filled with views taken from different global directions or orientations of the mobile device. For example, one grid would be filled with only a single image view of the robot pointing between 0° and 30°, another grid would be filled with views between 30° and 60°, and so on. This would result in multiple ground mapping images where the reflections in each mapping image are consistent. These mapping images can then be used in the user interface and display different maps depending on the direction the user rotates the map in the application, resulting in a more realistic and immersive user experience. Another possible implementation is to train a so-called NeRF-based model (NeRF stands for Neural Radiation Field and is a method for synthesizing new views of complex scenes) based on a single scene, limited to 2D application cases, thereby simplifying it.
[0037] In one implementation, panoramic mapping can be performed. Besides mapping color data directly from the image, semantic or panoramic segmentation can be performed on the color image, for example, and the results of these methods can be assigned. In this context, panoramic segmentation describes assigning each image point to a unique object category, and possibly to the distinction between different instances of that object category. This results in a classification assignment of the ground, which is also easier for the user to interpret by simply showing them which areas, for example, contain carpet, tile, wood flooring, etc. To achieve this, averaging methods different from those discussed above can be used, where the confidence level predicted by the neural network used for segmentation may need to be considered.
[0038] As previously mentioned, it is preferable to also determine positioning and / or navigation information for the mobile device, and more specifically, to determine this information based on the device and / or environmental information and SLAM. The generation of the texture map can then be further based on the positioning and / or navigation information.
[0039] The system for data processing according to the invention includes apparatus for performing the method or method steps according to the invention. The system may be a computer or server, such as a computer or server in a so-called cloud or cloud environment. There, images and environmental and / or device information can be acquired from a mobile device, and texture maps can be transmitted back to the mobile device (e.g., via a wireless data connection). However, it is also contemplated that such a system for data processing is a computer or control device within such a mobile device. Similarly, it is contemplated that some steps of the method are performed on a computer or server in the sense of a cloud environment, while other steps are performed on a computer or control device within the mobile device.
[0040] The present invention also relates to a mobile device having a depth sensor and a camera, as well as the aforementioned system for data processing (in the sense of a computer or control unit in the mobile device). Preferably, the mobile device further includes a drive unit for moving the mobile device and a control and / or adjustment unit.
[0041] The mobile device is preferably constructed as a transport vehicle that is at least partially automated in its movement, particularly a passenger transport vehicle or a freight transport vehicle, and / or a robot, particularly a household robot, such as a vacuuming and / or mopping robot, a floor or street cleaning device or a lawn mowing robot, and / or a drone, as previously described.
[0042] It is also advantageous to implement the method according to the invention in the form of a computer program or computer program product having program code for performing all method steps, because this results in particularly low costs, especially when the control device for execution is also used for other tasks and therefore already exists. Finally, a machine-readable storage medium is provided on which the computer program as described above is stored. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical, and electrical memories, such as hard disks, flash memory, EEPROM, DVDs, etc. The program can also be downloaded via a computer network (Internet, intranet, etc.). Such downloading can be performed via a wired / cable connection or wirelessly (e.g., via a WLAN network, 3G, 4G, 5G, or 6G connection, etc.).
[0043] Other advantages and embodiments of the invention will become apparent from the description and drawings. Attached Figure Description
[0044] The present invention is illustrated schematically based on the embodiments shown in the accompanying drawings, and is described below with reference to the drawings.
[0045] Figure 1 A mobile device in an environment illustrating the invention is shown schematically in a preferred embodiment.
[0046] Figure 2 The flowchart of the method according to the invention in a preferred embodiment is illustrated schematically.
[0047] Figure 3 A movable device for illustrating the present invention is shown schematically. Detailed Implementation
[0048] exist Figure 1The diagram schematically illustrates a mobile device 100 in an environment 120 illustrating the invention, according to a preferred embodiment. The mobile device 100 is exemplarily a cleaning or vacuuming robot having a control or adjustment unit 102 and a drive unit 104 (with wheels) for moving the vacuuming robot 100 within the environment 120 (e.g., a residence or room). Here, the vacuuming robot 100 moves on a generally at least substantially flat ground surface 122 of the environment 120.
[0049] Furthermore, the vacuum robot 100 exemplarily includes a sensor 106 configured as a 2D LiDAR sensor, which has a detection field. For better illustration, the detection field is chosen to be relatively small here; however, in practice, the field of view can also reach 360° (e.g., but at least 180° or at least 270°). With the aid of the 2D LiDAR sensor 106, the distance between the mobile device 100 (or sensor 106) and objects in the environment can be detected or determined. The 2D LiDAR sensor 106 is specifically arranged here such that it measures the distance to these objects in a plane parallel to the ground surface 122.
[0050] Therefore, the 2D LiDAR sensor 106 is an example of a depth sensor through which depth information can be detected, for example as part of device and / or environmental information.
[0051] Furthermore, the vacuum robot 100 has a camera 110, through which it can detect images of the environment 120. Here, the camera is specifically pointed forward, i.e., in the normal driving direction, and thus can detect images of the environment. The ground surface 122 is distortedly included in the camera's image only at the edges.
[0052] Furthermore, the vacuum robot 100 has a system 108 for data processing, such as a control unit, through which it can exchange data with a higher-level data processing system 150 via an indicated wireless connection. In system 150 (e.g., a server, which may also represent a so-called cloud), navigation information can be determined, for example, and then transmitted to system 108 within the vacuum robot 100, according to which the vacuum robot should be operated. However, it can also be specified that the navigation information is determined within system 108 itself or otherwise obtained. In addition to navigation information, system 108 can also, for example, obtain control information determined based on the navigation information, according to which control or adjustment unit 102 can move the vacuum robot 100 via drive unit 104. Here, the vacuum robot 100 can move according to a movement path or trajectory 130.
[0053] In addition, two obstacles 124 and 128, such as furniture, are exemplarily shown in environment 120. A carpet 126 is also exemplarily shown, on which the vacuum robot 100 can move, for example. The properties or texture of the surface of carpet 126 (which is also typically a floor surface) differ from other areas of floor surface 122, such as areas where wooden floors exist.
[0054] Furthermore, a mobile input device 140, such as a smartphone, is shown, which can, for example, connect to the vacuum robot 100 or its system 108 and computing system 150 via data transmission, for example, wireless communication. An application (so-called an App) can be set on the mobile input device 140, for example, to control the vacuum robot 100, such as starting and stopping its movement; and the movement of the vacuum robot 100 itself is typically automated or autonomous, and specifically, particularly using SLAM as described above.
[0055] Furthermore, the mobile input device 140 or the application may also provide the possibility of displaying a map of the environment 120, on which the user can define specific areas as Go-To or No-Go areas, or define specific areas in other ways.
[0056] Within the scope of this invention, in particular, a possibility for generating texture maps as such maps is proposed. Texture map 142 in Figure 1 The diagram is schematically shown. In particular, the different properties or textures of the ground surface in the environment can be seen here.
[0057] Figure 2 A method for generating a texture map according to the present invention, in a preferred embodiment, is schematically illustrated in the form of a flowchart or diagram. For example, as shown in... Figure 1 The vacuum cleaner robot shown is an example of a mobile device.
[0058] In step 200, device and / or environmental information is provided during this process. This device and / or environmental information is detected from the mobile device and / or environment by at least one sensor of the mobile device. Here, the device and / or environmental information includes, for example, information 202 from the IMU and information 204 from the odometer, which are sensors of the mobile device. Furthermore, the device and / or environmental information also includes depth information 206, which is, for example, from a lidar sensor of the mobile device (such as...). Figure 1 The depth information detected (as shown) can exist, for example, as a point cloud.
[0059] In step 208, a camera, such as a mobile device, is provided. Figure 1The image 210 shown is detected by the camera from the environment.
[0060] The device and / or environmental information and images are typically detected and provided repeatedly, but not usually in a time- or space-synchronized manner.
[0061] In step 212, the selected image 210 is then synchronized with the depth information 206 in spatial and / or temporal aspects to obtain synchronized depth information 214. Temporal synchronization is particularly important here, and the data may subsequently be transformed to lie in a common reference coordinate system for further processing.
[0062] In step 216, ground images 218 of the environment are then extracted from the image based on synchronized depth information 214.
[0063] These steps 212, 216 can be performed on all images 210, or only on selected images among these images 210. As explained in more detail above, various criteria 220 are considered to select those images that can be used for this purpose. The resulting ground image 218 can then be temporarily stored, for example, in step 222.
[0064] In addition to the steps of synchronizing images and extracting ground images, positioning and / or navigation information 224 can also be very routinely generated for the operation of the mobile device, for example, within the scope of SLAM described above. For this purpose, device and / or environmental information is used, particularly information 202 from the IMU, information 204 from the odometry, and depth information 206. In step 226, this information 202, 204, and 206 is processed within the scope of, for example, continuous or quasi-continuous trajectory planning. Based on this, pose optimization may be performed using the SLAM map in step 228, possibly through intermediate steps.
[0065] Step 226 can also be performed on all received device and / or environmental information, while the optimization in step 228 is performed, for example, only on so-called keyframes. Based on the trajectory planning in step 226, criterion 220 can also be determined in particular.
[0066] Based on the ground image 218, which may be temporarily stored, a texture map 232 is then generated in step 230. For this purpose, in particular, the multi-view combination mentioned in step 234 is used. Here, in particular, image regions in the texture map that contain image information from multiple ground images can be identified; furthermore, values can then be determined for these image regions based on these ground images.
[0067] As previously stated, the goal of multi-view compositing is specifically to create a globally consistent image of the ground surface from ground images. Each ground image is assigned the pose of the mobile device in the global map. For this purpose, for example, the aforementioned averaging on a global pixel grid can be performed, but other types of multi-view compositing can also be performed, as previously explained in detail. For this purpose, in particular, the results of the optimization in step 228 can be utilized.
[0068] Texture map 232 (e.g., corresponding to) Figure 1 The texture map 142 shown includes, or is based on, a 2D map 236 already obtained within the scope of SLAM (localization and / or navigation information 224 and possible mobile device trajectories), and ground images extracted from camera images and then connected in an appropriate manner. In step 240, the texture map may then be provided, for example, for further use.
[0069] It should be noted that such texture maps can be created from environments of varying sizes and scales, not just from a single room. Thus, for example, a texture map can be created from a residence with many rooms, each containing multiple different types of floor surfaces. Similarly, texture maps can be created from environments such as factory workshops, and in principle, from outdoor environments. Furthermore, different types of floor surfaces do not necessarily have to have different roughnesses, such as distinguishing between carpet and wood flooring; for example, different colors of the floor can also be used to differentiate between different types of floor surfaces.
[0070] exist Figure 3 The diagram schematically illustrates a mobile device 300 used to illustrate the present invention, and specifically shows the mobile device 300 used to illustrate the present invention with respect to the aspect of extracting ground images. The mobile device 300 exemplarily includes a 2D lidar sensor 306 and a camera 310. This could be... Figure 1 The mobile device 100 shown is a vacuum cleaner robot.
[0071] In this context, 360 represents a lidar beam or laser beam emitted by a lidar sensor. When the environment is detected by a lidar sensor, such laser beams 360 are typically emitted at angular intervals Δθ between each other. The angular interval between any two laser beams 360 can be, but is not necessarily, the same.
[0072] The laser beam 360 is then reflected off an obstacle (if the obstacle is close enough) and the reflected beam is detected. In this way, a point, denoted here 362, can be obtained along each of these laser beams, where the point lies on the obstacle surface and thus indicates the distance to the obstacle. The set of points 362 can also be referred to here as a lidar point cloud.
[0073] As previously described, extracting ground images may include: determining a free region, and more specifically, determining the free region based on synchronized depth information, such as a lidar point cloud synchronized with the image. This free region covers at least a portion of the ground surface of the environment. In particular, the free region may be defined as a set of points on straight lines corresponding to the detection directions of the lidar sensors.
[0074] Such points are Figure 3 The coordinates are represented by 364. These points 364 can lie on a straight line defined by the laser beam. For example, in the coordinate system of a lidar sensor (represented here by x, y, z), at the minimum radius r... min and maximum radius r max Within the region between them, a set of such points 364 can be assigned to each of the laser beams. These points 364 can be along the laser beams or straight lines, for example, with a specific interval Δr (sampling interval).
[0075] In general, this involves determining "how far" points can or should be assigned to the free region along an imaginary line. These points 364 are not necessarily mandatory to be sampled on a straight line with an interval of Δθ, but can also be sampled more densely (e.g., 1 / 2 Δθ). Therefore, there is not a single point 362 defining the boundary of each bundle. In this case, the free region can be defined, for example, by linear interpolation between two points 362, or (e.g., for caution) by choosing points with smaller radii.
[0076] Here, the maximum radius r max It is a pre-defined constant that describes why distances exceeding r should not be considered under any circumstances. max Points (e.g., because it is assumed that the image will be overly distorted or because the depth sensor is noisy when measuring at a distance, etc.).
[0077] With r max Similarly, r min These are also pre-set parameters that define the sampling range, for example, to skip areas in the image where parts of the robot can be seen or where the depth measurement of the LiDAR is inaccurate (LiDAR typically has a minimum distance from which it provides a reliable depth value).
[0078] In this way, the free region 366 can be represented by a set of points 364, specifically in a coordinate system. This free region 366—at least at the height of the lidar sensor—is unobstructed. Therefore, a portion of the ground surface lies within the free region 366 and is also visible in the relevant images from the camera.
[0079] The set of free regions 366 or points 364 can then be projected into the image, for which, for example, external and / or internal camera parameters can be used. Thus, image information from the image, i.e., colors or sets of pixels with specific colors, can be assigned to at least a portion of the free regions, specifically each point 364.
[0080] Alternatively, as previously mentioned, a perspective transformation can also be performed on the image.
Claims
1. A method for generating a texture map (142, 232) of a ground surface (122) in an environment (120), in which a mobile device (100), particularly a transport vehicle or robot with at least partially automated movement, especially a cleaning robot, is moving or about to move, the method comprising: Provide (200) device and / or environment information detected from the mobile device and / or the environment by at least one sensor of the mobile device (202, 204, 206), wherein the at least one sensor includes a depth sensor (106, 306). Provide (208) an image (210) detected from the environment by the camera (110, 310) of the mobile device; For one selected image or each of the provided images: The image is synchronized spatially and / or temporally with at least a portion of the device and / or environmental information detected by the depth sensor (212) to obtain synchronized depth information (214), and Based on the synchronized depth information, extract (216) the ground image of the environment (218) from the image. Based on the one or more ground images, generate (230) a texture map (232) of the ground surface of the environment; and Provide the texture map described in (240).
2. The method according to claim 1, wherein extracting the ground image comprises: A free region (366) is determined based on synchronized depth information, the free region covering at least a portion of the ground surface of the environment. Projecting the free region onto the image, and Assigning image information from the image to at least a portion of the free region, and Use the free regions with assigned image information as ground images.
3. The method according to claim 2, wherein the free region is defined as a set of points (364) on a straight line corresponding to the detection direction of a depth sensor configured as a lidar sensor.
4. The method according to claim 1, wherein extracting the ground image comprises: Perform a perspective transformation on the image to obtain the transformed image. Based on synchronized depth information, a free region is determined, the free region covering at least a portion of the ground surface of the environment, and The portion of the transformed image corresponding to the free region is used as the ground image.
5. The method according to any one of claims 2 to 4, wherein the free region lies in a plane.
6. The method according to any one of the preceding claims, wherein generating the texture map comprises: Determine one or more image regions of the texture map (234), for which image information from multiple ground images exists; For one or more of the plurality of image regions, determine (234) a value for at least a portion of the image information based on the plurality of ground images; as well as Use the value in the texture map.
7. The method according to any one of the preceding claims, the method further comprising: Based on the device and / or environmental information and based on SLAM, determine (226) positioning and / or navigation information for the mobile device (224). The texture map is further generated based on the positioning and / or navigation information (224).
8. The method according to any one of the preceding claims, wherein the one or more selected images are selected from the provided images according to at least one criterion (220), wherein the at least one criterion includes at least one of the following criteria for the respective image: When detecting the image, the position of the mobile device has changed less than a predetermined distance relative to the last detected image. When detecting the image, the orientation of the mobile device has changed less than a pre-given angle relative to the last detected image. When detecting the image, the orientation of the mobile device is within a predetermined range. The pose of the mobile device did not appear in any of the predetermined number of last detected images during the image detection process. One or more image parameters meet the pre-defined requirements. The distance to the obstacle in the detection direction of the camera is greater than a predetermined value, wherein the distance is determined in particular based on device and / or environmental information detected by the depth sensor.
9. The method according to any one of the preceding claims, the method further comprising: Based on the texture map, a segmented map of the ground of the environment is generated.
10. The method according to any one of the preceding claims, wherein the depth sensor is configured as a lidar sensor, particularly a 2D lidar sensor.
11. A system (108) for data processing, the system comprising means for performing the method according to any one of the preceding claims.
12. A mobile device (100, 300) having a depth sensor (106, 306), a camera (110, 310), and a system (108) according to claim 11, wherein the mobile device preferably has a control or adjustment unit (102) and a drive unit (104) for moving the mobile device according to navigation information.
13. The mobile device (100, 300) according to claim 12, configured as a transport vehicle, particularly a passenger transport vehicle or a freight transport vehicle, and / or a robot, particularly a household robot, such as a cleaning robot, a ground or street cleaning device or a lawnmower robot, and / or a drone.
14. A computer program comprising instructions that, when executed by a computer, cause the computer to perform method steps of the method according to any one of claims 1 to 10 when the computer program is executed on the computer.
15. A computer-readable storage medium having a computer program stored thereon according to claim 14.