Method and device for mapping a use environment for at least one mobile unit and locating at least one mobile unit in the use environment, and positioning system for the use environment
By employing random feature detection and binary descriptors in base plate texture localization, the problem of high computational workload is solved, achieving efficient and accurate base plate texture localization that adapts to different environmental conditions.
Patent Information
- Application Number
- CN202180071140.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-19
- Filing Date
- 2021-10-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2041-10-15
AI Technical Summary
Existing localization methods based on substrate texture involve large computational workloads and are difficult to perform efficient feature detection and correspondence lookup, especially in substrate textures with strong repetitive patterns where it is difficult to distinguish different features.
Feature detection is performed by randomly or pseudo-randomly selecting feature image regions, using binary descriptors such as BRIEF, BRISK, LATCH, or AKAZE to reduce computational workload, and determining feature locations through random processes or predefined distribution schemes, combined with data processing devices for parallel processing.
It effectively reduces the workload of feature extraction and localization calculations, improves the localization accuracy and robustness under strongly repeating pattern substrate textures, reduces computational complexity, and adapts to different environmental conditions.
Smart Images

Figure CN116324886B_ABST
Abstract
Description
Technical Field
[0001] This invention is based on the apparatus or method of the type claimed in the independent claims. Computer programs are also the subject of this invention. Background Technology
[0002] Various methods are known for feature-based localization based on substrate texture, or absolute or map-based localization based on substrate texture features. For example, random feature image regions can be used for feature extraction, but so far this has been particularly only applicable to applications that are not looking for correspondences between image pairs and therefore not specifically related to localization tasks, but rather to applications where the goal is to understand the content of the image. Examples of such applications are image classification or object recognition. Such applications particularly involve methods that operate entirely within a single system (e.g., autonomous vehicles or robots).
[0003] DE 10 2017 220 291 A1 discloses a method for automatically guiding vehicles along a virtual track system. This is a method that enables an autonomous system with a downward-oriented camera to follow a previously learned virtual track. Summary of the Invention
[0004] Against this backdrop, methods according to the independent claims and apparatus for using these methods are proposed using the solutions presented herein, and a corresponding computer program is provided. Advantageous extensions and modifications of the methods described in the independent claims are possible through the measures listed in the dependent claims.
[0005] According to the implementation, in particular, it enables efficient feature detection for texture-based drawing and / or localization in the context of at least one mobile unit. More precisely, for this purpose, a simplified scheme for feature detection for texture-based drawing and / or localization can be used. In other words, an efficient process for feature detection for localization using images from a downward-oriented camera can be provided. In this case, feature detection can be performed, in particular, independently of the actual image content, or by defining feature locations through a random process or by a fixed pattern. This can be done, for example, substantially independently of subsequent localization steps (such as feature description, correspondence lookup, and pose determination). In particular, the implementation can be based on the fact that, rather than actual feature detection, any image region can be used for the drawing and / or localization process. That is, computational time can be saved by determining the image region for feature extraction randomly, pseudo-randomly, or based on a static pattern. This type of feature detection is an efficient scheme for feature-based localization based on texture, particularly for the following reasons.
[0006] First, the probability of randomly selecting similar image regions can be categorized as relatively high in this case. This is because the camera pose can be well approximated using only three parameters: the x and y coordinates in the substrate plane and the orientation angle. In particular, the distance to the substrate is known, so the feature size of the image region can remain constant. For example, if the current pose estimate already exists during localization, the complexity can be further reduced. Especially when the orientation is estimated with sufficient accuracy, the parameters to be determined for the feature image region can be reduced to the image coordinates of the feature image region. If a feature descriptor with some translation robustness is used, such that slightly displaced image regions can also be evaluated as similar descriptors, then even though the feature image regions used are randomly selected, there is, for example, a high probability of finding the correct correspondence between overlapping image pairs. Second, it can be assumed that the substrate texture has a lot of information content. Correspondences can be found through appropriate feature descriptors, especially without special feature image regions, i.e., feature image regions with particularly high information content. Instead, typical substrate textures, such as concrete, asphalt, or carpet, can have sufficient characteristic properties anywhere to make it possible to find correspondences, provided that sufficiently overlapping feature image regions are used in the localization and reference images.
[0007] According to the implementation method, the computational workload for high-precision rendering and / or localization can be reduced, particularly based on the substrate texture features. This avoids the use of conventional, typically computationally intensive feature detection methods, such as SIFT (Lowe, 2004), which determine image regions suitable for finding correspondences. In this conventional approach, image regions with the most prominent specific attributes are identified, also known as global optimization. An example of such attributes is contrast with the local environment. Compared to this approach, the implementation method reduces computational workload, for example, by abandoning optimization, where the use of randomly or uniformly distributed feature image regions proposed herein results in reduced computational workload without compromising localization capabilities.
[0008] Compared to existing technologies, the following advantages are particularly achieved. One advantage is the reduction in computational workload for feature extraction, as mentioned above. Although a larger number of features can be used due to the lower probability of finding corresponding features, the overall computational workload for localization can be significantly reduced by using processes that are efficient in describing features, finding correspondences, and eliminating incorrect correspondences or selecting correct ones. Efficient methods suitable for describing features include, for example, binary descriptors such as BRIEF (Calonder et al., 2010), BRISK (Leutenegger et al., 2011), LATCH (Levi and Hassner, 2016), or AKAZE (Alcantarilla et al., 2013). In addition to the lower computational workload, the use of random feature image regions also offers advantages for specific substrate texture types. This is due to a phenomenon that may occur in substrate textures with strong repetitive patterns. In such textures, using classical feature detectors may result in the same location of a pattern being consistently identified as a feature image region. In this case, localization methods may no longer be able to, or may find it difficult to, distinguish patterns of different features. This can be prevented by using random or pseudo-random feature image regions.
[0009] A method is proposed for providing mapping data for a map of the usage environment of at least one mobile unit, wherein the method comprises the following steps:
[0010] Reference image data is read from the interface with the image capture device of the mobile unit, wherein the reference image data represents a plurality of reference images, the reference images being captured by means of the image capture device of a sub-segment specific to each reference image of the substrate of the use environment, wherein adjacent sub-segments partially overlap.
[0011] Multiple reference image features are extracted for each reference image using the reference image data, wherein the position of the reference image features in each reference image is determined by means of a random process and additionally or alternatively according to a predefined distribution scheme;
[0012] The drawing data is generated, wherein a reference feature descriptor is determined at the location of each reference image feature using the reference image data, wherein the drawing data has the reference image data, the location of the reference image feature, and the reference feature descriptor.
[0013] The operating environment can be the interior of one or more buildings that can be traversed by at least one mobile unit, and additionally or alternatively, the exterior surfaces of one or more buildings. The operating environment can have predefined boundaries. The at least one mobile unit can be designed as a vehicle, robot, etc., for highly automated driving. The image capturing device can have at least one camera of the mobile unit. The image capturing device can be arranged relative to the mobile unit in a predetermined orientation. The image capturing device can be a camera.
[0014] A method for creating a map of the usage environment for at least one mobile unit is also proposed, wherein the method comprises the following steps:
[0015] Receive drawing data from the communication interface with the at least one mobile unit, wherein the drawing data is provided according to an embodiment of the method provided above;
[0016] Using the rendering data and based on the correspondence between reference image features of overlapping reference images determined using the reference feature descriptor, the reference pose of the image capturing device relative to the reference coordinate system for each reference image is determined; and
[0017] The reference image is combined based on the reference pose, the position of the reference image features, the reference feature descriptor, and the reference pose to create a map of the usage environment.
[0018] The method for creating the map can be executed, for example, on a data processing device or by using a data processing device. In this case, the data processing device can be arranged separately from at least one mobile unit inside or outside the usage environment.
[0019] According to one embodiment, in the determination step, the reference pose can be determined based on the correspondence between reference image features, for which reference feature descriptors satisfying similarity criteria have been determined for the reference image features in the overlapping reference images. This embodiment offers the advantage that using such reproducible conditions improves robustness to image transformations such as translation and rotation of the image capturing device, as well as photometric transformations, which in turn can have a beneficial effect on localization because a more accurate correspondence can be found.
[0020] Furthermore, a method for determining location data for the positioning of at least one mobile unit in a usage environment is proposed, wherein the method comprises the following steps:
[0021] Image data is read from the interface with the image capture device of the mobile unit, wherein the image data represents at least one image, the at least one image being captured by means of the image capture device of a sub-segment of the base plate of the usage environment;
[0022] Multiple image features are extracted from the image data, wherein the location of the image features in the image is determined by means of a random process and additionally or alternatively according to a predefined distribution scheme;
[0023] The image data is used to generate a feature descriptor at the location of each image feature to determine the positioning data, wherein the positioning data has the location of the image feature and the feature descriptor.
[0024] At least one moving unit may be at least one moving unit from one of the methods described above, or at least one other moving unit corresponding to or similar to at least one moving unit from one of the methods described above. At least some steps of the method may be performed repeatedly or periodically for each image. Images, and therefore adjacent sub-segments represented by said images, may overlap.
[0025] Optionally, the reference feature descriptor and / or feature descriptor can be a binary descriptor. Using binary descriptors can be advantageous because they can typically be computed faster than non-binary descriptors or floating-point descriptors, and because binary descriptors enable particularly efficient mapping.
[0026] According to one embodiment, the determination method may include the step of outputting the positioning data to an interface with a data processing device. In this case, the positioning data can be output in multiple data packets. Each data packet may include at least one location of an image feature and at least one feature descriptor. Once at least one feature descriptor has been generated, the data packet can be output. This embodiment offers the advantage of efficiently implementing a centralized positioning method based on substrate texture features, where image processing and relative pose determination or visual ranging are performed on the mobile unit or robot, while absolute pose determination is outsourced to a central server based on previously captured maps. In this case, in particular, the three steps of the positioning method, namely image processing, communication, and positioning, can be performed in parallel or partially overlapping. Using random or predefined feature regions allows one feature to be calculated after another and then directly sent to the data processing device or server. Thus, the server obtains a constant stream of extracted image features from the mobile unit, allowing image processing and communication to be performed in parallel or partially overlapping. Subsequent positioning based on image features obtained from the positioning image and a map of the application area performed on the server can also be performed in parallel or partially overlapping. It is also possible to systematically and / or entirely search images based on feature criteria, and whenever an image region that meets the criteria is found, a descriptor can be calculated, and then the feature information can be sent to the server. In terms of feature detectors, the search for globally optimal image regions can be abandoned in order to form features.
[0027] The determination method may further include the step of finding a correspondence between image features of the positioning data and reference image features of the previous image using feature descriptors of the positioning data and reference feature descriptors of the previous image. Further, the determination method may also include the step of determining the pose of the image capturing device relative to the reference coordinate system based on the correspondence found in the finding step to perform positioning. Such an implementation provides the advantage that incremental or relative positioning can be performed, wherein a relative camera pose can be determined relative to a previous camera pose, even if, for example, the data connection with the data processing device should be temporarily interrupted.
[0028] According to an embodiment of the method provided above and / or the determination method described above, a random process and, additionally or alternatively, a predefined distribution scheme may be used in the extraction step, wherein a list of all possible image locations having reference image features or image features is generated, and the list is pseudo-randomly mixed or locations are pseudo-randomly selected from the list, and, additionally or alternatively, a fixed location pattern or one of multiple pseudo-randomly created location patterns is used. The localization method based on substrate texture uses arbitrary or unrelated feature image regions (whether random or predefined) to form correspondences. Its advantage lies in reducing the computational workload of image processing because, compared to using a traditional feature detector, it is not necessary to fully process the entire image to identify the optimal feature image region. Instead, features can be computed at any location in the reference image or image. This method also has the advantage that image processing can still proceed while the information is used in the next processing step.
[0029] According to another embodiment of the method provided above and / or the determination method described above, a random process and, additionally or alternatively, a predefined distribution scheme may be used in the extraction step, wherein a variable or predetermined number of positions are used, and additionally or alternatively, different position distribution densities are set for different sub-regions of the reference image or the image. Such an implementation offers the advantage of improved adaptation to actual conditions in the usage environment, more precisely, adaptation to the substrate.
[0030] A method for locating at least one mobile unit in a usage environment is also proposed, wherein the method comprises the following steps:
[0031] Location data is received from the communication interface with the at least one mobile unit, wherein the location data is determined according to an embodiment of the determination method described above;
[0032] Using the feature descriptors of the location data and the reference feature descriptors of the map, the correspondence between the image features of the location data and the reference image features of the map created according to the implementation method described above is determined.
[0033] Based on the correspondence determined in the determining step and using the reference pose of the map, the pose of the image capturing device relative to the reference coordinate system is determined to generate pose information representing the pose; and
[0034] The attitude information is output to the communication interface with the at least one mobile unit to perform the positioning.
[0035] This positioning method can be performed, for example, on a data processing device or by using a data processing device. In this case, the data processing device can be arranged separately from at least one mobile unit inside or outside the operating environment.
[0036] According to one implementation, in the attitude determination step, a weighted value and, additionally or alternatively, a confidence value can be applied to the correspondence determined in the correspondence determination step to generate an evaluated correspondence. In this case, the attitude can be determined based on the evaluated correspondence. Such an implementation provides the advantage of further improving the reliability, robustness, and accuracy of positioning, particularly because incorrect or less accurate correspondences can be excluded.
[0037] Each of the methods mentioned above can be implemented, for example, in software or hardware or a combination of both, such as in a control device or apparatus.
[0038] The proposed solution also creates a device for a mobile unit, wherein the device is configured to perform, manipulate, or implement, within a corresponding apparatus, the steps of the proposed providing method and / or variations of the proposed determining method. This embodiment of the invention, in the form of a device for a mobile unit, can also quickly and efficiently solve the task on which the invention is based. The proposed solution also creates a device for a data processing apparatus, wherein the device is configured to perform, manipulate, or implement, within a corresponding apparatus, the steps of the proposed map creation method and / or variations of the proposed positioning method. This embodiment of the invention, in the form of a device for a data processing apparatus, can also quickly and efficiently solve the task on which the invention is based.
[0039] Therefore, the device may have at least one computing unit for processing signals or data, at least one storage unit for storing signals or data, at least one interface for reading sensor signals from the sensor or outputting data signals or control signals to the actuator, and / or at least one communication interface for reading or outputting data embedded in a communication protocol. The computing unit may be, for example, a signal processor, a microcontroller, etc., and the storage unit may be a flash memory, EPROM, or magnetic storage unit. The communication interface may be configured to read or output data wirelessly and / or via wired connections, wherein a communication interface capable of reading or outputting wired data may, for example, read such data electrically or optically from a corresponding data transmission line or output data to a corresponding data transmission line.
[0040] In the current context, a device can be understood as an electrical device that processes sensor signals and outputs control signals and / or data signals accordingly. This device may have an interface that can be constructed in hardware and / or software. In the case of a hardware construction, the interface may, for example, be part of a so-called system ASIC that contains various functions of the device. However, the interface may also be a separate integrated circuit or at least partially composed of discrete devices. In the case of a software construction, the interface may be a software module that exists, for example, on a microcontroller along with other software modules.
[0041] A positioning system for an environment in which at least one mobile unit can be used is also proposed, wherein the positioning system has the following characteristics:
[0042] The at least one moving unit, wherein the moving unit has the above-described embodiment of the device for the moving unit;
[0043] A data processing apparatus, wherein the data processing apparatus includes embodiments of the device for the data processing apparatus described above, wherein the device for the mobile unit and the device for the data processing apparatus are interconnected to transmit data.
[0044] Specifically, the positioning system may include the data processing device and multiple mobile units. At least one of the mobile units may optionally have a device for the mobile unit configured to perform, manipulate, or implement steps of variations of the provided method presented herein within a corresponding device.
[0045] It is also advantageous to have a computer program product or computer program with program code, which can be stored on a machine-readable carrier or storage medium such as semiconductor memory, hard disk memory or optical memory and is used to execute, implement and / or manipulate the steps of the method according to one of the above embodiments, especially when the program product or program is running on a computer or device. Attached Figure Description
[0046] Embodiments of the proposed scheme are shown in the accompanying drawings and explained in more detail in the following description.
[0047] Figure 1 A schematic diagram of an embodiment of a positioning system for use in an environment is shown;
[0048] Figure 2 A schematic diagram of an embodiment of a device for a mobile unit is shown;
[0049] Figure 3 A schematic diagram of an embodiment of a device for a mobile unit is shown;
[0050] Figure 4 A schematic diagram of an embodiment of a device for a data processing apparatus is shown;
[0051] Figure 5 A schematic diagram of an embodiment of a device for a data processing apparatus is shown;
[0052] Figure 6 A flowchart illustrating an embodiment of a method for providing mapping data for a map of the usage environment of at least one mobile unit;
[0053] Figure 7 A flowchart illustrating an embodiment of a method for creating a map of the usage environment for at least one mobile unit;
[0054] Figure 8 A flowchart illustrating an embodiment of a method for determining location data for the positioning of at least one mobile unit in a usage environment is shown;
[0055] Figure 9 A flowchart illustrating an embodiment of a method for locating at least one mobile unit in a usage environment is shown;
[0056] Figure 10 A schematic diagram of the image and feature image regions is shown;
[0057] Figure 11 A schematic diagram of image 1123 and image feature 335 according to one embodiment is shown;
[0058] Figure 12 A schematic diagram of overlapping images with characteristic image regions is shown;
[0059] Figure 13 A schematic diagram of an overlapping image 1123 and image feature 335 according to one embodiment is shown;
[0060] Figure 14 A schematic diagram of reproducible conditions according to one embodiment is shown; and
[0061] Figure 15 A schematic diagram of the three-stage time flow of centralized localization based on image features is shown.
[0062] In the following description of advantageous embodiments of the invention, the same or similar reference numerals are used for elements shown in the various figures that have similar effects, and repeated descriptions of these elements are omitted. Detailed Implementation
[0063] Figure 1 A schematic diagram of an embodiment of a positioning system 110 for use environment 100 is shown. At least one mobile unit 120 can be used in use environment 100. Figure 1 The illustration shows four mobile units 120 only as an example in a usage environment 100. The usage environment 100 is, for example, the interior and / or exterior of at least one building, a surface that can be traversed by at least one mobile unit 120. The usage environment 100 has a base plate 102 on which at least one mobile unit 120 can move. At least one mobile unit 120 is a vehicle for highly automated driving, particularly a robot or robotic vehicle, etc.
[0064] The positioning system 110 includes at least one mobile unit 120 and a data processing device 140. The data processing device 140 is arranged inside and / or outside the operating environment 100. Figure 1 In the illustration, the data processing device 140 is shown only exemplarily within the operating environment 100. The data processing device 140 is configured to perform data processing on at least one mobile unit 120.
[0065] Each moving unit 120 includes an image capturing device 122 whose field of view is directed toward the base plate 102 of the usage environment 100. In this case, the image capturing device 122 is a camera. Optionally, each moving unit 120 includes at least one illumination device 124 for illuminating the field of view of the image capturing device 122. Figure 1In the illustrated embodiment, each mobile unit 120 exemplarily includes a ring-shaped lighting device 124. The image capture process of each image capture device 122 can image a sub-segment 104 of the base plate 102 of the usage environment 100. Furthermore, each mobile unit 120 includes a device 130 for the mobile unit. The device 130 for the mobile unit is connected to the image capture device 122 either data-transmittingly or signal-transmittingly, or may alternatively be designed as part of the image capture device 122. The device 130 for the mobile unit is configured to perform methods for providing mapping data 160 for a map 170 of the usage environment 100 and / or for determining positioning data 180 for the positioning of at least one mobile unit 120 in the usage environment 100. The device 130 for the mobile unit is discussed in more detail with reference to the following figures.
[0066] The data processing device 140 includes a device 150 for data processing. The device 150 is configured to perform a method for creating a map 170 of the usage environment 100 and / or a method for locating at least one mobile unit 120 in the usage environment 100. Here, the device 130 for the mobile unit and the device 150 for the data processing device are interconnected for data transmission, particularly by means of a radio connection, such as a WLAN or mobile radio. In this case, mapping data 160 and / or positioning data 180 or image features can be transmitted from at least one mobile unit 120 to the data processing device 140, and attitude information 190 or estimated robot attitude for positioning can be transmitted from the data processing device 140 to each mobile unit 120.
[0067] In other words, in the usage environment 100, multiple independent robots or mobile units 120 are wirelessly communicated with a data processing device 140, also referred to as a central server 170 with stored maps. Each mobile unit 120 is equipped with a downward-oriented camera or image capture device 122. Furthermore, the field of view or shooting area can be artificially illuminated, thus enabling reliable localization regardless of external light conditions. The mobile unit 120 captures images of the substrate 102 at regular time intervals to determine its own pose. For this purpose, features at arbitrary locations are extracted from the images, and these features are sequentially sent to the server or data processing device 140, where a substrate texture map is created and / or stored, which allows the pose of the mobile unit 120 to be estimated based on the sent features or localization data 180. According to one embodiment, for this feature extraction, communication, and pose estimation, they can be performed at least partially in parallel, thereby yielding runtime advantages over methods in which each of the three steps should be completely completed before the next step can begin. After localization is complete, the estimated pose is sent back to the mobile unit 120 in the form of pose information 190. The mobile unit 120 can use the pose information 155, for example, to accurately locate itself. For the radio connection between the server and the robot, WLAN or 5G can be used, for example.
[0068] Figure 2 A schematic diagram of an embodiment of a device 130 for a mobile unit is shown. The device 130 for the mobile unit corresponds to or is similar to... Figure 1 Devices used for mobile units. Figure 2 The illustrated device 130 for a mobile unit is configured to perform and / or manipulate, within a corresponding apparatus, steps of a method for providing mapping data 160 for the usage environment of at least one mobile unit. This providing method corresponds, for example, to or is similar to [the method described in the original text]. Figure 6 The method.
[0069] The device 130 for the mobile unit includes a reading device 232, an extraction device 234, and a generating device 236. The reading device 232 is configured to read reference image data 223 from an interface 231 with an image capture device of the mobile unit. The reference image data 223 represents multiple reference images, each of which is captured by means of the image capture device of a specific sub-segment of a substrate of the operating environment. Here, adjacent sub-segments partially overlap each other. Furthermore, the reading device 232 is configured to forward the reference image data 223 to the extraction device 234. The extraction device 234 is configured to extract multiple reference image features 235 for each reference image using the reference image data 223. The position of the reference image features 235 in each reference image is determined by means of a random process and / or according to a predefined distribution scheme. The extraction device 234 is also configured to forward the reference image features 235 to the generating device 236. The generating device 236 is configured to generate drawing data 160, wherein reference feature descriptors are determined at the location of each reference image feature 235 using reference image data. The drawing data 160 includes reference image data 223, the locations of the reference image features 235, and the reference feature descriptors. The device 130 for the moving unit is also configured to output the drawing data 160 to a separate interface 239 with the data processing device.
[0070] Figure 3 A schematic diagram of an embodiment of a device 130 for a mobile unit is shown. The device 130 for the mobile unit corresponds to or is similar to... Figure 1 or Figure 2 Devices used for moving units. Figure 3 The illustrated device 130 for a mobile unit is configured to perform and / or manipulate, in a corresponding apparatus, the steps of a method for determining positioning data 180 for positioning at least one mobile unit in a usage environment. This determination method corresponds, for example, to or is similar to [the method described in the original text]. Figure 8 The method.
[0071] The device 130 for the mobile unit includes an additional reading device 332, an additional extraction device 334, and a generation device 336. The reading device 332 is configured to read image data 323 from an interface 231 with an image capture device of the mobile unit. The image data 323 represents at least one image captured by the image capture device of a sub-segment of the substrate of the operating environment. The additional reading device 332 is also configured to forward the image data 323 to the additional extraction device 334. The additional extraction device 334 is configured to extract multiple image features 335 from the image using the image data 323. Here, the positions of the image features 335 in the image are determined by means of a random process and / or according to a predefined distribution scheme. The additional extraction device 334 is configured to forward the image features 335 to the generation device 336. The generation device 336 is configured to generate feature descriptors at the positions of each image feature 335 using the image data 323 to determine location data 180. The location data 180 includes the positions of the image features 335 and the feature descriptors.
[0072] The device 130 for the mobile unit is specifically configured to output positioning data 180 to an additional interface 239 with a data processing device. According to one embodiment, the positioning data 180 is output in multiple data packets, each data packet including at least one location of image feature 335 and at least one feature descriptor. More precisely, according to this embodiment, the location of each image feature 335 and its associated feature descriptor are output in the data packet as soon as the feature descriptor is generated. Therefore, data packets can be output at least partially in parallel, and additional image features 335 can be extracted and feature descriptors can be generated.
[0073] Figure 4 A schematic diagram of an embodiment of a device 150 for a data processing apparatus is shown. The device 150 for the data processing apparatus corresponds to or is similar to [the device from...]. Figure 1 Equipment used for data processing devices. Figure 4 The device 150 shown for a data processing apparatus is configured to perform and / or manipulate, within a corresponding apparatus, the steps of a method for creating a map 170 of a usage environment for at least one mobile unit. This creation method corresponds, for example, to or is similar to, from... Figure 7 The method.
[0074] The device 150 for the data processing apparatus includes a receiving device 452, a determining device 454, and a combining device 456. The receiving device 452 is configured to receive rendering data 160 from a communication interface 451 with at least one mobile unit. In this case, the rendering data 160 is provided by means of a device for the mobile unit. Furthermore, the receiving device 452 is configured to forward the rendering data 160 to the determining device 454. The determining device 454 is configured to use the rendering data 160 and, based on the correspondence between reference image features of overlapping reference images determined using a reference feature descriptor, determine the reference pose 455 of the image capturing device relative to a reference coordinate system for each reference image. The determining device 454 is also configured to forward the reference pose 455 to the combining device 456. The combining device 456 is configured to combine the reference image, the position of the reference image features, the reference feature descriptor, and the reference pose 455 according to the reference pose 455 to create a map 170 of the usage environment. The device 150 for the data processing apparatus is also specifically configured to output the map 170 to the memory interface 459 of the memory device of the data processing apparatus.
[0075] Figure 5 A schematic diagram of an embodiment of a device 150 for a data processing apparatus is shown. The device 150 for the data processing apparatus corresponds to or is similar to... Figure 1 or Figure 4 The equipment used for data processing. Figure 5 The device 150 shown for a data processing apparatus is configured to perform and / or manipulate steps of a method for positioning at least one mobile unit in a usage environment within a corresponding device. This positioning method corresponds to or is similar to, for example, a method for positioning a mobile unit in a usage environment. Figure 9 The method.
[0076] The device 150 for the data processing apparatus includes an additional receiving device 552, an obtaining device 554, an additional determining device 556, and an output device 558. The additional receiving device 552 is configured to receive positioning data 180 from a communication interface 451 with at least one mobile unit. In this case, the positioning data 180 is determined by means of a device for the mobile unit. The additional receiving device 552 is also configured to forward the positioning data 180 to the obtaining device 554. The obtaining device 554 is configured to determine a correspondence 555 between image features of the positioning data 180 and reference image features of the map 170 using feature descriptors of the positioning data 180 and reference feature descriptors of the map 170. Furthermore, the obtaining device 554 is configured to forward the correspondence 555 to the additional determining device 556. The additional determining device 556 is configured to determine the orientation of the image capturing device relative to a reference coordinate system based on the correspondence 555 and a reference orientation of the map 170, to generate orientation information 190. The orientation information 190 represents the determined orientation. The additional determining device 556 is also configured to output attitude information 190 via output device 558 to a communication interface 451 with at least one mobile unit to perform positioning.
[0077] Figure 6 A flowchart illustrating an embodiment of a method 600 for providing mapping data for a map of the usage environment of at least one mobile device is shown. The providing method 600 includes a reading step 632, an extraction step 634, and a generation step 636. In the reading step 632, reference image data is read from an interface with an image capture device of the mobile unit. The reference image data represents a plurality of reference images, which are captured by means of an image capture device of a specific sub-segment of the background of the usage environment for each reference image, wherein adjacent sub-segments partially overlap. In the extraction step 634, a plurality of reference image features are extracted from each reference image using the reference image data. Here, the location of the reference image features in each reference image is determined by means of a random process and / or according to a predefined distribution scheme. In the generation step 636, mapping data is generated. Here, reference feature descriptors are determined at the locations of each reference image feature using the reference image data. The mapping data includes the reference image data, the locations of the reference image features, and the reference feature descriptors.
[0078] Figure 7 A flowchart illustrating an embodiment of a method 700 for creating a map of a usage environment for at least one mobile device is shown. The creation method 700 includes a receiving step 752, a determining step 754, and a combining step 756. In the receiving step 752, data is received from a communication interface with at least one mobile unit according to... Figure 6The method shown or similar methods provide the drawing data. In determination step 754, the drawing data is used and the correspondence between reference image features of overlapping reference images determined using reference feature descriptors is used to determine the reference pose of the image capturing device relative to the reference coordinate system for each reference image. In combination step 756, the reference images, the positions of reference image features, the reference feature descriptors, and the reference pose are combined according to the reference pose to create a map of the usage environment.
[0079] According to an embodiment of the method 700 for creating a map, in the determination step 754, a reference pose is determined based on the correspondence between reference image features, wherein reference feature descriptors that satisfy similarity criteria among the reference image features in overlapping reference images have been determined.
[0080] Figure 8 A flowchart illustrating an embodiment of a method 800 for determining location data for locating at least one mobile unit in a usage environment is shown. The determination method 800 includes a reading step 832, an extraction step 834, and a generation step 836. In the reading step 832, image data is read from an interface with an image capture device of the mobile unit. The image data represents at least one image captured by the image capture device of a sub-segment of a base plate of the usage environment. In the extraction step 834, multiple image features of the image are extracted using the image data. Here, the location of the image features in the image is determined by means of a random process and / or according to a predefined distribution scheme. In the generation step 836, feature descriptors are generated at the location of each image feature using the image data to determine location data. The location data includes the location of the image feature and the feature descriptor.
[0081] Specifically, the method 800 for determining location data further includes a step 838 of outputting the location data to an interface with a data processing device. Here, the location data is output in multiple data packets, each data packet including at least one location of an image feature and at least one feature descriptor. For example, the output step 838 is repeated such that the location of each image feature 335 and its associated feature descriptor are output in the data packet as soon as the feature descriptor is generated. Thus, data packets can be output and additional image features 335 can be extracted and feature descriptors generated, at least in part in parallel with this.
[0082] According to one embodiment, the method 800 for determining positioning data further includes a step 842 of finding a correspondence between image features of the positioning data and reference image features of the previous image using feature descriptors of the positioning data and reference feature descriptors of the previous image, and a step 844 of determining the orientation of the image capturing device relative to a reference coordinate system based on the correspondence found in step 842 to perform positioning. In the event of a data transmission interruption in the usage environment, this embodiment can at least temporarily achieve independent positioning of at least one mobile unit.
[0083] refer to Figure 6 The method 600 and / or shown for providing map data Figure 8 The method 800 shown for determining location data uses a random process and / or a predefined distribution scheme in extraction steps 634 and / or 834, wherein a list of all possible image locations having reference image features or image features is generated, and locations are pseudo-randomly mixed from the list or pseudo-randomly selected from the list, and / or a fixed location pattern or one of multiple pseudo-randomly created location patterns is used. Additionally or alternatively, in extraction steps 634 and / or 834, a random process and / or a predefined distribution scheme is used, wherein a variable or predetermined number of locations are used, and / or wherein different location distribution densities are set for different sub-regions of the reference image or image.
[0084] Figure 9 A flowchart illustrating an embodiment of a method 900 for locating at least one mobile unit in a usage environment is shown. The positioning method 900 includes a receiving step 952, a determining step 954, a determining step 956, and an output step 958. In the receiving step 952, data is received from a communication interface with at least one mobile unit according to... Figure 8 The location data is determined by the method shown or a similar method. In determination step 954, the feature descriptor of the location data is used and based on... Figure 7 The method or similar method creates a reference feature descriptor for the map, and determines the correspondence between the image features of the positioning data and the reference image features of the map. In determination step 956, based on the correspondence determined in determination step 954 and using the reference pose of the map, the pose of the image capturing device relative to the reference coordinate system is determined to generate pose information representing the pose. In output step 958, the pose information is output to a communication interface with at least one mobile unit to perform positioning.
[0085] According to one embodiment, in determination step 956, weighted values and / or confidence values are applied to the correspondence determined in determination step 954 to generate an evaluated correspondence. In this case, the pose is then determined based on the evaluated correspondence in determination step 956.
[0086] Figure 10 A schematic diagram of image 1000 and feature image regions 1002 is shown. In this case, multiple feature image regions 1002 are extracted from image data representing image 1000 according to classical or conventional methods. Figure 10 This illustrates a scenario where using a classic feature detector might fail. The feature detector finds the feature image region 1002 by searching for the region with the strongest expression that meets a specific criterion or has a specific attribute throughout the entire image 1000. Figure 10 In this context, the feature detector can, for example, search for regions with the strongest contrast. The texture shown in image 1000 has a very regular structure, which is represented here as a grid. This regular structure now manifests itself in the fact that specific regions corresponding to the recurrence of this regular structure have a strong expression of a property that will be maximized by the feature detector; this strong expression is, in this case, strong contrast. These regions, correspondingly extracted by the detector as feature image region 1002, are... Figure 10 The feature is shown as a square and has the strongest contrast between its inner and outer regions. The difficulty now lies, for example, in the fact that feature image regions 1002 have very similar, in extreme cases identical, content, making it impossible to distinguish feature image regions and their corresponding descriptors from each other. This makes it difficult to find a correspondence between 1000 shown here and another feature image region in another overlapping image, because for each feature extracted from said other image, all features from image 1000, i.e., feature image regions 1002, either correspond to it or none correspond to it. In both cases, it may be difficult to obtain information usable for localization. This is merely an illustrative example. In principle, an image 1000 could be created for each feature detector, where the found feature image regions 1002 degenerate in a similar manner.
[0087] Figure 11 A schematic diagram of image 1123 and image feature 335 according to one embodiment is shown. By performing... Figure 8 The method or similar method used to determine the location data extracts image features 335 from image data representing image 1123. Therefore, the position of image feature 335 in image 1123 is determined by means of a random process and / or according to a predefined distribution scheme, in other words, it is independent of the specific image content of image 1123. Figure 11 The image content shown in this case corresponds to Figure 10 The image content of the image.
[0088] When using randomly distributed feature image regions or image feature 335, no occurrence will occur. Figure 10 The problem described in [the document]. Besides the features previously extracted by the feature detector... Figure 10 Beyond the feature image regions in the image, image 1123 contains features fully usable for localization, which are represented here by irregular symbols. Some of the randomly selected feature image regions or image features 335 also fall on portions of a regular structure, shown as a grid, and are associated with... Figure 10 Like the feature image regions in the image, they can only be used to form correspondences in a limited way. However, some arbitrarily located image features 335 also contain irregular image content that needs to be clearly identified between the stripes of the grid. At least some of these available image features 335 can also be found in the overlapping images used for localization, where the image features 335 do not need to have pixel-precise identical image content, thus allowing for successful localization.
[0089] Figure 12 A schematic diagram of overlapping images 1000 and 1200 with feature image regions 1002 is shown. The feature image regions 1002 in the two images 1000 and 1200 correspond only to those from... Figure 10 The feature image regions. In other words, the use of the same pattern in overlapping images such as 1000 and 1200 causes the feature image regions 1002 to be exactly shifted relative to each other, so that there are no overlapping features and therefore no correct correspondence.
[0090] Figure 13 A schematic diagram of overlapping images 1123, 1323 and image feature 335 according to one embodiment is shown. By performing... Figure 8 The method or similar method used to determine location data extracts image features 335 from image data representing images 1123 and 1323. Therefore, the position of image features 335 in images 1123 and 1323 is determined by means of a random process and / or according to a predefined distribution scheme; in other words, it is independent of the specific image content of images 1123 and 1323. Specifically, the positions of image features 335 in images 1123 and 1323 are distributed differently.
[0091] Using feature locations of different modes or the location of image feature 335 in overlapping images 1123 and 1323 can prevent [the occurrence of] [the following]. Figure 12 The problem is illustrated here. In this case, there are overlapping feature image regions, which allows for the determination of membership relationships for these overlapping feature image regions when searching for correspondences.
[0092] Figure 14A schematic diagram of reproducible conditions according to one embodiment is shown. For this purpose, an image 1123 represented by image data, two reference images 1423 or drawn images represented by reference image data, and reference image features 235 are shown. Images 1123 and 1423 at least partially overlap each other. The reproducible conditions indicate that, in Figure 7 In a method or similar method for creating a map, a reference pose is determined in the determination step based on the correspondence between reference image features 235, and reference feature descriptors that satisfy similarity criteria have been determined for the reference image features in the overlapping reference images.
[0093] In other words, the reference image feature 235 from the drawn image or reference image 1423 is stored only if the corresponding reference image feature 235 in the overlapping drawn image or reference image 1423 results in a similar feature descriptor. This increases the probability that the corresponding image feature in the location image or image 1123 is also evaluated as a similar feature descriptor.
[0094] Figure 15 A schematic diagram of the time flow 1500 and 1505 for three stages 1511, 1512, and 1513 based on image feature-based centralized localization is shown. For this purpose, a time axis t is plotted in the diagram. The first flow 1500 represents the conventional flow of the three stages 1511, 1512, and 1513 in a sequential or serial manner. According to one embodiment, the second flow 1505 represents the flow of the three stages 1511, 1512, and 1513 in a parallel or at least partially parallel manner. (Reference) Figure 8 Methods for determining location data and Figure 9 The localization method, at least partially parallelized in the second process 1505, is implemented specifically by executing the output steps of the determination method. The first stage 1511 represents image processing, typically on the mobile unit side; the second stage 1512 represents communication or data transmission between the mobile unit and the server; and the third stage 1513 represents localization, typically on the server side. In the second process 1505, the three stages 1511, 1512, and 1513 can be executed overlappingly, i.e., partially in parallel, thereby significantly shortening the duration of the localization process from start to finish.
[0095] Referring to the above figures, the embodiments, background, and advantages of the embodiments will be summarized and briefly explained below. According to the embodiments, positioning can be achieved based on the texture features of the base plate.
[0096] A common approach to solving this task is to determine corresponding features from images captured for localization and one or more reference images. These correspondences can then be used to determine the pose, which consists of the position and orientation of the camera or image capture device relative to the reference images at the time the localization image was captured. A typical approach can be divided into, for example, four stages:
[0097] 1. Feature Detection: First, during feature detection, a set of image regions (feature image regions) suitable for later correspondence searching is determined. This can be, for example, image regions that are particularly bright or dark compared to the local environment, or image regions that differ from the local environment in other ways, or image regions with specific structures (e.g., lines or corners). In this case, it is assumed that these regions of the substrate texture meet the selection criteria even when viewed from other camera poses, thereby finding identical (or at least overlapping) feature image regions in the localization and reference images.
[0098] 2. Feature Description: Then, feature descriptors for these image regions are calculated in the feature description stage.
[0099] 3. Finding Correspondences: Then use these descriptors to determine the corresponding features. It is assumed here that the corresponding features have already been described with similar descriptors, while descriptors for non-corresponding features should have less similarity.
[0100] 4. Attitude Determination: Finally, the proposed correspondences are used to determine the attitudes, where it is often meaningful to use a method that is robust to a certain proportion of incorrect correspondences.
[0101] Below, before discussing possible extensions or other embodiments, we will first describe an embodiment in which random feature positions are used.
[0102] For map-based positioning, first prepare or create a map 170 of the area or environment to be used, for example... Figure 2 and Figure 4 or Figure 6 and Figure 7 As shown in the diagram, map creation can be divided into five stages, for example:
[0103] 1. A vehicle or robot, or possibly a drone, as a mobile unit 120, travels entirely across the area or environment of use 100, and in the process continuously captures, in particular, superimposed reference images 1423 of the base plate 102.
[0104] 2. For each captured reference image 1423, extract a set of reference image features 235. The positions of the reference image features 235 in the reference image 1423 are determined using a random process. This random process can be as follows: First, create a list of all possible image positions; then, mix this list and use the first n entries of the mixed list of image positions to determine the set of image positions. Here, n represents the number of reference image features 235 to be extracted. As an alternative to mixing the list, a random number generator can be used n times to determine random list indices for the image positions and the corresponding list entries can be used as additional image feature positions. Compared to the first variation, the second variation has a lower computational workload but may use the same image positions multiple times. To prevent this, it is possible to check whether the random list index has been used after each determination, which may slightly increase the computational workload. Which variation is best suited depends on the application, and in particular on the number of reference image features 235 to be extracted.
[0105] 3. Calculate feature descriptors for each image feature location determined in the previous stage. The process here depends on the chosen feature description method. These feature description methods can fix the size of the viewed image portion, or allow the user to define a matching size, or determine the size using a suitable method based on the image content of the area surrounding the feature location. If the feature description method requires the orientation of the viewed image portion, typically to rotate the image portion accordingly, this orientation can be determined using a suitable method based on the area surrounding the feature location, such as the direction with the strongest gradient, or by using the current camera orientation, thereby assigning the same orientation to all features of the reference image 1423. The camera orientation here can be an orientation relative to the initial camera orientation from the first reference image 1423 taken for drawing, or an absolute orientation determined, for example, using a compass.
[0106] 4. Determine the reference pose 455 of the captured reference image 1423. Here, the reference pose 455 of the first reference image 1423 can form the origin of a coordinate system, or a coordinate system with a known reference can be used, such as a coordinate system defined by the ground view of the usage environment 100. Here, the image pose or reference pose 455 should be determined such that they are consistent with each other. For this purpose, for example, an image stitching process can be used to combine the individual photographs into a large image, so that the reference images 1423 are then correctly positioned relative to each other.
[0107] 5. Effectively store the extracted reference image features 235. It is meaningful to store the location of the reference image features 235 within the coordinate system of map 170. This creates a map 170 that can be used for localization. Essentially, map 170 comprises a set of reference images 1423 whose reference poses 455 have been optimized so that these reference images can be placed in a matching manner. Furthermore, a set of reference image features 235 at arbitrary or random locations is extracted from each reference image 1423. The pose of the features in the world, i.e., the pose relative to the origin of the coordinate system of map 170, is known. Additionally, a descriptor is stored for each feature image region, which can then be used to form correspondences during localization.
[0108] Subsequent map-based positioning, such as Figure 3 and Figure 5 or Figure 8 and Figure 9 As shown, it can be divided into six stages:
[0109] 1. Take images 1123 and 1323, which should be used for positioning, in the drawn usage environment 100.
[0110] 2. Determine the location of random or arbitrary image features or the location of image feature 335.
[0111] 3. If the camera position is already roughly known, this can be used to limit the search area for localization by, for example, considering only such reference image features from the vicinity of the estimated position 235.
[0112] 4. Similar to map creation, calculate feature descriptors at the locations of image features. If orientation is required in this case, it can be redetermined in an absolute manner, such as using a compass, or if the camera orientation relative to the coordinate system of map 170 is roughly known from the previous pose determination, it can be used as feature orientation.
[0113] 5. Then, a suitable method (e.g., nearest neighbor matching) is used to find the correspondence to determine the correspondence between the drawn reference image features 235 and the image features 335 extracted from the localized images 1123 and 1323.
[0114] 6. The correspondences found in this way (some of which may be incorrect) are then used for attitude estimation using appropriate methods, such as RANSAC-based Euclidean transform estimation and subsequent Levenberg-Marquardt optimization.
[0115] According to one embodiment, so-called incremental positioning can also be performed. In this case, it is also possible to perform... Figure 8The method involves estimating or determining the relative camera pose with respect to previous camera poses. This works similarly to map-based localization described above, but with the difference that there are no reference image features 235 from map 170 available for finding correspondences. Instead, reference image features 235 from previous images 1123, 1323, or a sequence of previous images 1123, 1323, along with their previously estimated poses, are used for incremental localization. The limitation of incremental localization over map-based localization is that inaccuracies propagate between images, causing the estimated pose to deviate increasingly from the actual pose. However, incremental localization can be particularly useful for areas of environment 100 that were not considered during rendering.
[0116] The proposed concept of using random image feature locations or the locations of reference image features 235 and image features 335 can be meaningfully extended. The use of random or pseudo-random locations is generally advantageous because the extracted reference image features 235 and image features 335 are uniformly distributed on average. This also applies to using fixed patterns, such as grid or mesh-shaped arrangements, at uniformly distributed locations, but here it may occur that the feature image regions of two overlapping reference images 1423 or images 1123, 1323 are exactly shifted relative to each other, resulting in no correct feature correspondence between these feature image regions (see also...). Figure 12 Determining random locations requires significantly more computation than using a fixed set of locations (e.g., uniformly distributed locations). Therefore, it may be meaningful to avoid using random location features. Consequently, possible alternatives exist depending on the application:
[0117] 1. For map-based localization: Random locations can be used when recreating the map, as this process is typically not time-critical. Predefined fixed location distributions can then be used during time-critical localization.
[0118] 2. For incremental localization: Feature extraction in each image 1123 and 1323 is time-critical, so it makes sense to use only a fixed set of feature locations.
[0119] To generally offset the aforementioned limitations—that is, the feature positions used in the two overlapping reference images 1423 or images 1123, 1323 are exactly shifted from one another so that there is no insufficient overlap between the feature image regions—alternating different feature position patterns can be used (see also...) Figure 13 Another approach to reduce computational workload is, for example, to pre-generate a large number of random location patterns, which can then be used sequentially during positioning.
[0120] Depending on the application, it may be meaningful to extract reference image features 235 or image features 335 at a higher density for a specific image region than for other image regions:
[0121] 1. Regarding map-based localization: The overlap of the reference images 1423 used to create the map is important. If these reference images do not overlap or barely overlap, the highest probability of finding the correct correspondence 555 during localization is achieved using a uniform distribution of features, because how subsequent localization images 1123, 1323 overlap with the drawn image or reference image 1423 is unknown when determining the reference image features 235. If the reference images 1423 overlap, reference image features 235 are extracted for the overlapping regions across multiple reference images 1423. In this case, it may be meaningful to extract fewer features in the overlapping regions or edges of the reference images 1423 than in the non-overlapping regions or centers of the reference images 1423, thus obtaining a uniform distribution of features or their locations across all reference images 1423.
[0122] 2. Regarding incremental localization: We are only interested in the image regions that overlap with previous or subsequent images 1123 and 1323. This overlap depends on the driving speed, driving direction, and shooting frequency. It is meaningful to extract features only in areas that can be assumed to overlap with previous or upcoming images 1123 and 1323, i.e., more features at the image edges.
[0123] Another meaningful extension is based on the concept also known as the reproducibility condition (see also...) Figure 14 This is a condition that reference image features 235 extracted during map creation must satisfy in order for the reference image features to be stored; otherwise, these reference image features are discarded and replaced by features that satisfy the condition. The reproducibility condition requires that feature image regions or reference image features 235 in two overlapping drawn images 1423 be evaluated as similar feature descriptors, thus providing robustness to image transformations such as camera translation and rotation, as well as photometric transformations applied to reference image 1423. It has been shown that using this condition increases the probability that corresponding feature image regions between drawn image 1423 and positioning images 1123, 1323 are also evaluated as similar feature descriptors, thereby increasing the probability of finding the correct correspondence 555.
[0124] According to another embodiment, a fully parallelized positioning system 110 is proposed, for example, for robot swarms. This is a cost-effective solution for high-precision positioning of mobile units 120 (e.g., autonomous vehicles or robots). In this case, two concepts are combined: (1) a server-based centralized positioning system 110, and (2) a positioning method based on re-identifying substrate texture features in a previously created map 170.
[0125] A typical application scenario for this and other embodiments is, for example, a warehouse where multiple or a group of autonomous robots, acting as mobile units 120, are responsible for transporting materials, goods, and tools. For the mobile units 120 to move autonomously, it is important that they know their orientation, i.e., their position and orientation. The requirements for positioning accuracy and robustness vary depending on the task the mobile unit 120 is performing. Thus, when a mobile unit 120 travels from one location to another, it is sufficient for the mobile unit to know its position accurately to 10 cm, especially as long as the mobile unit can temporarily avoid obstacles. Conversely, if the mobile unit 120 should, for example, automatically load materials at a specific location, millimeter-precision positioning may be required. In applications where the majority of goods transport in a warehouse should be handled by the mobile units 120, a large number of mobile units 120 operating simultaneously can be used. In this situation, a suitable technique for positioning the mobile units 120 is visual positioning using downward-oriented cameras or feature-based positioning. This positioning can achieve high accuracy without the need for infrastructure measures such as installing visual markers, reflectors, or radio units. Furthermore, this type of positioning works even under challenging conditions in dynamic environments, such as warehouses where there are no static landmarks for orientation, because, for example, shelves can be arranged differently at any time. This is because the substrate texture typically remains stable over long periods, especially in protected areas such as warehouses. Wear that occurs over time typically occurs only locally, so the affected areas can continue to be detected based on their environment in the map 170 of the application area or usage environment 100 and can then be updated accordingly.
[0126] Localization based on substrate texture, particularly visual features of substrate 102, can be used like a fingerprint to definitively identify a location on substrate 102. Here, typically it is not a single feature, such as bitumen, that is definitively re-identifiable, but rather a constellation of multiple such visual substrate texture features. Before actual localization can be performed, the environment 100 is mapped, i.e., reference images 1423 are captured during one or more mapping sessions. The relative poses of these reference images with respect to each other are determined in an optimization method, for example by means of so-called image stitching, so that the reference images 1423 can then be matched together. The mobile unit 120 can then be localized using the map 170 created in this way by re-finding the drawn reference image features 235 in the images 1123, 1323 captured for localization.
[0127] To cost-effectively implement baseplate texture-based localization of multiple mobile units 120 in the form of robot swarms, it may be meaningful to outsource some of the computational workload from the mobile units 120 to a central server or data processing unit 140. In a simple variant, images 1123, 1323 captured by the baseplate 102 for localization can be sent to the server unprocessed, thus completely outsourcing image processing and subsequent feature-based localization. However, this variant can be disadvantageous because the mobile units 120 can no longer operate independently in such a constellation, but must rely on a stable and fast connection to the server. Furthermore, the images require significant storage space and result in a correspondingly large amount of communication workload. Therefore, a sensible variant is to perform image processing on the mobile units 120 and only transmit the extracted image features 335 to the server for localization. In this variant, the communication workload is significantly reduced, and the mobile units 120 can optionally operate at least temporarily independently of the server by determining their current pose relative to a previous pose. The advantage over a completely decentralized variant, i.e., performed on the respective mobile unit 120, is that the map 170 is not stored on the mobile unit 120, but only on a server, and the higher workload of absolute positioning compared to relative positioning is outsourced. In this case, absolute positioning means that the orientation of the mobile unit 120 is determined based on the previously captured and optimized map 170.
[0128] According to at least one embodiment, an efficient implementation of centralized localization based on substrate texture features is proposed, wherein image processing and relative pose determination (visual ranging) are performed on the mobile unit 120, while absolute pose determination based on a previously captured map 170 is outsourced to a central server or data processing device 140. In this case, the three stages of localization, namely image processing 1511, communication 1512, and localization 1513, are performed in parallel or partially overlapping in time. To make this possible, arbitrary or pseudo-random feature image regions or their locations are used, or in other words, global optimality is abandoned. Using arbitrarily arranged feature image regions differs from conventional methods, in which the optimal feature image region best suited for finding correspondences with reference features is globally determined or determined throughout the entire image. However, in such conventional methods, the entire image must be fully considered to satisfy the criterion of global optimality. Therefore, the entire image must be fully processed before finding a suitable set of features and forwarding it to the server. However, according to the embodiment, in the case of localization based on substrate texture, the global optimality of the extracted features can be abandoned because, for example, the degrees of freedom of the robot pose can be reduced to two (x position and y position), since the distance to the substrate is known with high precision and the orientation can be estimated well approximated, either absolutely or relatively relative to the previous pose using a compass, and since using random feature image regions is sufficient, as the substrate texture has a very high information density, any image region can be used to definitively identify the substrate region like a fingerprint. The use of random or arbitrary positions, particularly for image features 335, allows one feature to be calculated after another and then directly sent to the server or data processing device 140. The server thus obtains a constant stream of extracted image features 335 from at least one mobile unit 120, allowing image processing 1511 and communication 1512 to be performed in parallel or partially overlapping in time. Subsequent localization 1513 on the server based on image features 335 obtained from localization images 1123, 1323 and a map 170 of the application area or usage environment 100 can also be performed in parallel or partially overlapping in time. For this purpose, a voting procedure can be applied, where each found feature correspondence or correspondence 555 votes for the camera's position at the time of capture. This method allows the correspondences 555 to be entered one after another in parallel or in partially overlapping time with communication 1512 and image processing 1511.
[0129] In the conventional concept of centralized localization systems—where corresponding image features are used for localization and some necessary computations are outsourced to a server (see, for example, Schmuck and Chli (2019); Kim et al. (2019))—the localization process is viewed as a sequential process: first, image processing 1511 is performed; after this image processing is completed, the information required for localization 1513 is fully transmitted to the server; and after communication with the server ends 1512, the server calculates the robot's pose estimate, such as... Figure 15 As shown in the first process 1500.
[0130] According to the embodiments, the conventional concept of centralized positioning systems can be improved by performing positioning not in sequential stages, but in parallel, that is, in processes or stages that partially overlap in time, thereby completing positioning significantly faster. See also Figure 15 The second process 1505 in the diagram. This can be achieved, in particular, by using a baseplate texture image instead of an image from a forward-oriented camera, and by discarding optimal image features of the global or image extent. It has been shown that any image feature region can be used in the case of baseplate texture-based localization. This is not the case for images from a forward-oriented camera, because it is a more complex problem in which the size of the image feature region used is decisive for finding the correct correspondence between the map and the localization image, and moreover, a large part of the observed environment is not suitable for forming a correspondence, for example, because they are not static or contain little visual information. In the case of a downward-oriented camera, a constant size of the image feature region used can be used because the distance to the baseplate 102 is essentially constant. Furthermore, it has been shown that typical baseplate textures contain sufficient information content everywhere to form a correspondence.
[0131] According to the embodiments, image-based centralized localization can be accelerated. This eliminates a drawback of conventional localization methods, where a long time can be required between capturing a localization image and completing pose estimation. This is particularly important when high levels of localization accuracy are required, because even if the pose of the localization image is determined very accurately, the mobile unit continues to move during this period, thus requiring a less accurate determination of the current pose by estimating the path covered during this time. The unit will also benefit from the embodiments if it relies so heavily on highly accurate pose estimation that it must remain stationary until it obtains the pose estimate from the server, as the stationary time can be reduced.
[0132] The advantage of the image-based positioning method implemented according to the embodiment is that it is an infrastructure-free solution, meaning the environment 100 does not (must) be adapted for it, but rather uses existing landmarks to determine the location. Furthermore, the camera advantageously operates (e.g., compared to radar) in both indoor and outdoor areas, which is not the case for GPS, and achieves highly accurate attitude determination, for example, compared to radio measurements. The advantage of the image-based positioning method implemented according to the embodiment, which uses a downward-oriented camera to capture images of the substrate texture, is that it can also operate in the environment 100 where the object can move freely or where observation of the environment can be limited, such as in a warehouse or among crowds. Moreover, the high-precision positioning according to the embodiment can be easily achieved because the observed visual features are very close to those of the camera compared to the features from a forward-oriented camera, and positioning is performed independently of external light conditions when using its own artificial lighting (e.g., by means of at least one lighting device 124).
[0133] The advantage of the substrate texture-based localization method implemented according to the embodiment, which uses any or random feature image regions or their locations to form correspondences, is that it reduces the computational workload of image processing because, unlike using (typical) feature detectors, this localization method does not require processing the entire image to identify near-optimal feature image regions. Instead, according to the embodiment, features are determined at any location in reference image 1423 or images 1123, 1323. Another advantage of this method is that image processing does not need to end before information can be used in the next processing step. Instead, image processing can be performed incrementally, i.e., feature-by-feature, with the information obtained so far already available in the next processing step. The advantage of the localization method implemented according to the embodiment, which outsources storage and computational workloads to a central server, is that these capabilities can be saved on individual game units 120 (e.g., autonomous vehicles, robots, etc.), thereby enabling cost-effective group size scaling.
[0134] According to the embodiments, a centralized localization method based on substrate texture is specifically proposed, in which any image region is used for feature extraction. Compared with conventional centralized localization methods, the advantage is that localization can be performed faster, thereby achieving higher localization accuracy.
[0135] Before discussing some possible alternatives and extensions, a simple embodiment is described below. Figure 1 As shown, system 110 is used to use data from... Figures 6 to 9Methods 600, 700, 800, and 900 describe a system comprising one or more mobile units 120 (e.g., robots or other vehicles to be located) and a central server or data processing unit 140. Each mobile unit is specifically equipped with a downward-oriented camera or image capture device 122, a computing unit, and a radio module (e.g., WLAN or mobile radio). Optionally, the camera's shooting area can be illuminated by artificial lighting, thus allowing the image to be captured independently of external light conditions and enabling shorter camera exposure times, thereby minimizing motion blur in the image. The server itself has computing power, a radio module for communicating with the mobile units 120, and memory, specifically a pre-drawn map 170 of the operating environment 100, located in this memory. The positioning method can be divided into two parts in principle: the creation of the map 170 and the actual positioning.
[0136] To map the environment 100, a specially designed mobile unit 120 can be used, for example, to capture wide stripes of the substrate 102 in a single shot, or at least one general-purpose mobile unit 120 or an application robot can be used. For example, the environment 100 is completely traversed and thus scanned. A long sequence of overlapping substrate texture images or reference images 1423 is then given. Reference image features 235 are extracted from these images. Each reference image feature 235 is defined, on the one hand, by its associated image region, for example, by the image coordinates and radius of its center and the orientation angle; on the other hand, a feature descriptor describing the associated feature image region is calculated for each reference image feature 235. Typically, an optimal feature image region in the reference image 1423 is found using a suitable feature detection method. In this case, a feature image region is well-suited when similar image regions can be found again with a high probability in the overlapping reference images 1423. Traditionally, the image region with the strongest performance (global optimization) of a particular attribute is determined here. An example of such an attribute is contrast with the local environment. However, according to the embodiment, a method without global optimization is used because it has been shown that random or arbitrary locations are sufficient for localization based on substrate texture. However, other intelligent methods can also be used as long as global optimization is not required. For example, the reference image 1423 can be systematically searched for areas with specific properties, such as areas that look like corners, edges, or intersections. These features can then be described using common methods such as SIFT (Lowe (2004)) or faster methods such as BRIEF (Calonder et al. (2010)). After feature extraction, correspondences between features of overlapping reference images 1423 are found. Typically, a distance metric between feature descriptors is computed for this purpose, suggesting that features with similar descriptors correspond. Incorrect correspondences are then filtered out, and the pose of the reference images 1423 is estimated based on the remaining correspondences, such that their overlap is correctly superimposed and a mosaic effect of the captured substrate 102 is obtained. The optimized image pose and extracted reference image features are stored in map 170 on the server, so that the pose of positioning images 1123 and 1323 in map 170 can be determined during actual positioning.
[0137] For positioning, the moving unit 120 to be positioned captures images 1123 and 1323 of the base plate 102. Then, image features 335 are extracted sequentially, and the following process is triggered for each image feature 335:
[0138] 1. Image processing on the robot: Determine the feature image region, which should be done using the same method as when drawing, such as through a random process.
[0139] 2. Image processing on the robot: Calculate feature descriptors for the feature image regions, using the same methods used for feature description during rendering.
[0140] 3. Communication: Location data 180, i.e., information about the selected feature image region, especially the image coordinates and descriptors, is sent to the server.
[0141] 4. Location on the server: Location data 180 arrives at the server. There, corresponding image features are searched in map 170. If the approximate orientation of mobile unit 120 is known, the search area can be narrowed down.
[0142] 5. Location on the server: The correspondences 555 found so far during the location process are used to determine the pose of the mobile unit 120. For example, this can be done using a voting method where each correspondence 555 votes for the correct pose on the map 170, where the set of votes expands with each image feature 335 processed. Once enough image features 335 have been processed, or once the confidence level of the current pose estimate is high enough, the server can notify the mobile device 120 of its pose, thereby terminating image processing and communication.
[0143] To maintain sufficiently high positioning accuracy, according to one embodiment, the attitude of the moving mobile unit 120 can be updated at shorter intervals. An update rate of 10 to 60 Hz is conceivable here. However, in a system 110 with a large number of mobile units 120, this could result in a significant communication workload, making it potentially meaningful to allow only every nth update to be performed on the server. For intermediate steps, the attitude (visual ranging) can be calculated separately relative to the previous step. Then, map-based absolute positioning via the server can only be used periodically to correct for accumulated errors (drift) in the local attitude estimation on the mobile units 120.
[0144] According to one embodiment, a confidence value for the current pose estimate is determined on a server or data processing unit 140 such that localization can be terminated once the confidence value exceeds a defined threshold.
[0145] Instead of a downward-oriented camera, it is conceivable that an upward-oriented camera could be used in a similar manner. However, this is only suitable for indoor settings where the ceiling height is known. In this case, instead of the base plate 102, the ceiling of the environment 100 is photographed.
[0146] In the case of system 110 described here, for example, it is assumed that there are independent mapping processes. It is also conceivable that map 170 is created online by mobile unit 120. For this purpose, mobile unit 120 can create its own local map, which is later unified with map 170 on a server into a large common map. This would then be a Simultaneous Localization and Mapping (SLAM) system.
[0147] Instead of always sending the extracted image features 335 one by one, it may be meaningful for communication to send the feature set all at once to reduce communication overhead.
[0148] Even in conceivable variations without a server—where computation runs entirely on the mobile unit 120—parallelizing image processing and localization can be meaningful. This can be used to better utilize available hardware, such as in multi-core processors or when some computation can be outsourced to dedicated hardware such as a graphics card. Furthermore, image processing can be stopped here as long as the confidence level for the current pose estimate is high enough.
Claims
1. A method (600) for providing drawing data (160) for a map (170) of an environment (100) for use of at least one mobile unit (120), wherein the method (600) comprises the following steps: Reference image data (223) is read (632) from the interface (231) of the image capture device (122) of the moving unit (120), wherein the reference image data (223) represents a plurality of reference images (1423) which are captured by means of the image capture device (122) of a sub-segment (104) of the base plate (102) of the use environment (100) specific to each reference image (1423), wherein adjacent sub-segments (104) partially overlap. Using the reference image data (223), multiple reference image features (235) are extracted (634) for each reference image (1423), wherein the position of the reference image features (235) in each reference image (1423) is determined by means of a random process and / or according to a predefined distribution scheme, independent of the specific image content of the reference image data (223); The drawing data (160) is generated (636), wherein a reference feature descriptor is determined using the reference image data (223) at the location of each reference image feature (235), wherein the drawing data (160) has the reference image data (223), the location of the reference image feature (235) and the reference feature descriptor.
2. A method (700) for creating a map (170) of a usage environment (100) for at least one mobile unit (120), wherein the method (700) comprises the following steps: Drawing data (160) is received (752) from the communication interface (451) with the at least one mobile unit (120), wherein the drawing data (160) is provided by the method (600) according to any one of the preceding claims; Using the drawing data (160) and based on the correspondence between the reference image features (235) of the overlapping reference images (1423) determined using the reference feature descriptor, the reference pose (455) of the image capturing device (122) of each reference image (1423) relative to the reference coordinate system is determined (754); and The reference image (1423) is combined (756) according to the reference pose (455), the position of the reference image feature (235), the reference feature descriptor and the reference pose (455) to create a map (170) of the usage environment (100).
3. The method (700) according to claim 2, wherein in the determining step (754), the reference pose (455) is determined based on the correspondence between reference image features (235), and reference feature descriptors that satisfy similarity criteria between each other have been determined for the reference image features in the overlapping reference image (1423).
4. A method (800) for determining positioning data (180) for positioning at least one mobile unit (120) in a usage environment (100), wherein the method (800) comprises the following steps: Image data (323) is read (832) from the interface (231) of the image capture device (122) of the moving unit (120), wherein the image data (323) represents at least one image (1123, 1323) which is captured by means of the image capture device (122) of a sub-segment (104) of the base plate (102) of the use environment (100); Using the image data (323), extract (834) multiple image features (335) of the images (1123, 1323), wherein the position of the image features (335) in the images (1123, 1323) is determined by means of a random process and / or according to a predefined distribution scheme, independent of the specific image content of the images (1123, 1323); The image data (323) is used to generate (836) feature descriptors at the location of each image feature (335) to determine the positioning data (180), wherein the positioning data (180) has the location of the image feature (335) and the feature descriptor.
5. The method (800) according to claim 4, having the step (838) of outputting the positioning data (180) to an interface (239) with a data processing device (140), wherein the positioning data (180) is output in a plurality of data packets, wherein each data packet includes at least one location of an image feature (335) and at least one feature descriptor, wherein the data packet is output once at least one feature descriptor is generated.
6. The method (800) according to any one of claims 4 to 5, comprising a step (842) of finding a correspondence between image features (335) of the positioning data (180) and reference image features (235) of the previous images (1123, 1323) using feature descriptors of the positioning data (180) and reference feature descriptors of the previous images (1123, 1323), and comprising a step (844) of determining the orientation of the image capturing device (122) of the images (1123, 1323) relative to a reference coordinate system based on the correspondence found in the finding step (842) to perform positioning.
7. The method (600; 800) according to any one of claims 1 and 4 to 5, wherein a random process and / or a predefined distribution scheme are used in the extraction step (634; 834), wherein a list of all possible image locations having reference image features (235) or image features (335) is generated, and the list is pseudo-randomly mixed or locations are pseudo-randomly selected from the list, and / or a fixed location pattern or one of a plurality of pseudo-randomly created location patterns is used.
8. The method (600; 800) according to any one of claims 1 and 4 to 5, wherein a random process and / or a predefined distribution scheme are used in the extraction step (634; 834), wherein a variable or set number of locations are used, and / or wherein different location distribution densities are set for different sub-regions of the reference image (1423) or the image (1123, 1323).
9. A method (900) for locating at least one mobile unit (120) in a usage environment (100), wherein the method (900) comprises the following steps: Location data (180) is received (952) from the communication interface (451) with the at least one mobile unit (120), wherein the location data (180) is determined by the method (800) according to any one of claims 4 to 8; Using the feature descriptor of the location data (180) and the reference feature descriptor of the map (170) created by the method (600; 700) according to any one of claims 1 to 3, determine (954) the correspondence (555) between the image features (335) of the location data (180) and the reference image features (235) of the map (170). Based on the correspondence (555) determined in step (954) and using the reference pose (455) of the map (170), the pose (956) of the image capturing device (122) of the images (1123, 1323) relative to the reference coordinate system is determined (956) to generate pose information (190) representing the pose; and The attitude information (190) is output (958) to the communication interface (451) with the at least one mobile unit (120) to perform the positioning.
10. The method (900) of claim 9, wherein in the determining step (956), a weighted value and / or a confidence value are applied to the correspondence (555) determined in the determining step (954) to generate an evaluated correspondence, wherein the pose is determined based on the evaluated correspondence.
11. A device (130) for a mobile unit (120), wherein the device (130) is configured to perform and / or manipulate the steps of the method (600) according to any one of claims 1, 7 and 8 and / or the steps of the method (800) according to any one of claims 4 to 8 in the corresponding unit (232, 234, 236; 332, 334, 336).
12. An apparatus (150) for a data processing device (140), wherein the apparatus (150) is configured to perform and / or manipulate the steps of the method (700) according to any one of claims 2 to 3 and / or the steps of the method (900) according to any one of claims 9 to 10 in corresponding units (452, 454, 456; 552, 554, 556, 558).
13. A positioning system (110) for an environment (100) in which at least one mobile unit (120) can be used, wherein the positioning system (110) has the following characteristics: The at least one mobile unit (120) has a device (130) for the mobile unit (120) according to claim 11. A data processing apparatus (140), wherein the data processing apparatus (140) has a device (150) for the data processing apparatus (140) according to claim 12, wherein the device (130) for the moving unit (120) and the device (150) for the data processing apparatus (140) are interconnected to transmit data.
14. A computer program product configured to perform and / or manipulate the steps of the method (600; 700; 800; 900) according to any one of the preceding claims.
15. A machine-readable storage medium having a computer program product according to claim 14 stored thereon.
Citation Information
Patent Citations
Method for automatically guiding a vehicle along a virtual rail system
DE102017220291A1
Method and apparatus for mapping an operational environment for at least one mobile unit and for localizing at least one mobile unit in an operational environment and localization system for an operational environment
DE102020213151A1