Method and system for generating an environment model and localization using cross-sensor feature point references
Patent Information
- Application Number
- CN202310268002.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-11-29
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2036-11-29
AI Technical Summary
在这种情况下,提供有用特征点的场景或环境的表示的数量可能很少,并且确定位置可能需要更长的时间、可能不太准确或不可能
[0039] One or more of modules one through fifteen can also be computer software programs that execute on a computer and provide the functionality of the corresponding module. Combinations of dedicated hardware modules and computer software programs are also conceivable.
Smart Images

Figure CN116147613B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 201680091159.6, entitled "Method and System for Generating and Localizing an Environment Model Using Cross Sensor Feature Point References", filed on November 29, 2016. Technical Field
[0002] This invention relates to the mapping or scanning of an environment and the determination of a location within that environment. Background Technology
[0003] Advanced driver assistance systems (ADAS) and autonomous vehicles require highly accurate maps of roads and other areas where vehicles can travel. Traditional satellite navigation systems, such as GPS, Galileo, GLONASS, or other known positioning techniques like triangulation, cannot determine a vehicle's position on the road, or even within its lane, with an accuracy of a few centimeters. However, especially when autonomous vehicles are traveling on multi-lane roads, they need to accurately determine their lateral and longitudinal positions within their lanes.
[0004] One known method for determining vehicle location with high accuracy involves capturing images of road markings once or multiple times with a camera and comparing the unique features of the road markings or objects along the road in the captured images with a corresponding reference image obtained from a database, in which the corresponding location of the road markings or objects is provided. This method of location determination only provides sufficiently accurate results if the database provides highly accurate location data in images and if the database is updated periodically or at appropriate time intervals. Road markings can be captured and registered by special-purpose vehicles that capture road images while in motion, or they can be extracted from aerial photographs or satellite imagery. The latter variation can be considered advantageous because road markings and other features shown in vertical or overhead view images on a substantially flat surface are rarely distorted. However, aerial photographs and satellite imagery may not provide sufficient detail to generate highly accurate maps of road markings and other road features. Furthermore, aerial photographs and satellite imagery are not well-suited for providing the details of objects and road features best observed from a ground perspective.
[0005] Most localization systems used today, such as Simultaneous Localization and Mapping (SLAM) and other machine vision algorithms, generate and use feature points. Feature points can be salient points or salient regions in a two-dimensional image generated by a 2D sensor, such as a camera, or salient points or salient regions in a two-dimensional representation of the environment generated by a scanning sensor. These salient points or regions may carry some information about the third dimension, but are usually defined and used as a two-dimensional representation because machine vision or robot vision is typically implemented using cameras that provide two-dimensional information.
[0006] A set of feature points can be used to determine a location within a specific range, such as along a road. In image-based localization, or typically in localization based on some representation of the environment, feature points of a specific part of the environment are provided in what are called key images, keyframes, or reference frames of that part of the environment. For clarity, the term "reference frame" will always be used in this document when referring to a reference representation of a part of the environment.
[0007] A reference frame can be a two-dimensional image, in which an image plane and one or more feature points are identified. Feature points can be identified by processing the camera image with filters or other optical processing algorithms to find suitable image content that can be used as feature points. The image content that can be used for feature points may be related to objects and markers in the image, but this is not a mandatory requirement. For a particular image processing algorithm used, salient points or salient regions in the image that match the feature points stand out substantially from other regions or points in the image, for example, by their shape, contrast, color, etc. Feature points can be independent of the shape or appearance of objects in the image, can be independent of each other, and can correspond to identifiable objects. Therefore, in this case, saliency does not simply refer to structures or the outlines, colors, etc., of structures that would be obvious to a human observer. Rather, saliency can refer to any property of a part of the scene that a particular algorithm “sees” or recognizes, which is applied to a representation of the scene captured by a particular type of sensor, and that property makes that part sufficiently different from the rest of the representation of the scene.
[0008] A set of reference frames, each containing one or more feature points, can be understood as a map that can be used for machine orientation. For example, a robot or autonomous vehicle can use this map to understand its environment, improve results by combining the results of several scans, and locate itself in that environment by using the reference frames.
[0009] Different types of sensors generate different types of scene representations. Images generated by a camera are drastically different from those generated by radar sensors, ultrasonic sensor arrays, or scanning laser sensors. Visible objects or features in the environment may appear at different resolutions and have different shapes, depending on how the sensor "sees" or captures the objects or features. Furthermore, a camera image can generate a color image, while a radar sensor, ultrasonic sensor array, scanning laser, or even another camera cannot. In addition, different cameras can have different resolutions, focal lengths, lens apertures, etc., which can also result in each camera having different feature points.
[0010] Furthermore, different algorithms for identifying feature points can be applied to different representations of a scene generated by different types of sensors. Each algorithm can be optimized for a specific representation generated by a particular type of sensor.
[0011] Therefore, as a result of different sensor types and their specific processing, the feature points found in each reference frame can be sensor-specific and / or algorithm-specific. Influenced by machine vision methods, representations of scenes or environments generated by different types of sensors can produce feature points in one representation that are not found in another, and can also produce different numbers of feature points that can be used for localization in representations originating from different sensors.
[0012] Machines or vehicles that need to determine their location or position may not be equipped with all possible types of sensors, and even if they are, different sensors may have imaging characteristics unsuitable for certain environmental conditions. For example, radar sensors can work unimpeded in fog or rain, while cameras may not produce useful results in such conditions, but underwater robots can “see much better” with sonar. In such cases, the number of representations of the scene or environment providing useful feature points may be limited, and determining the location may take longer, be less accurate, or be impossible. Summary of the Invention
[0013] One object of this invention is to improve localization using different types of sensors. The improvements provided by this invention lie in providing a large number of feature points for any given location across a wide range of sensor types, and also in accelerating the identification of feature points in different representations during the localization process. This object, as well as its variations and improvements, are achieved by the methods and systems claimed in the appended claims. The different types of sensors include those further mentioned in the discussion of the background art of this invention.
[0014] In a first aspect, the present invention addresses the aforementioned problems by generating a general description or environmental reference model of the environment, which includes feature points of a wide range of sensor types and / or environmental conditions. The general description can be interpreted as a high-level description of the environment, which can serve as the basis for deriving reference frames for various sensor types and / or environmental conditions. In a second aspect, the general description or model of the environment is used as a data source to provide such reference frames to a device moving within the environment (e.g., a vehicle or robot) for orientation and positioning. The reference frames include feature points that can be found in representations of the environment derived from different types of sensors, all of which can be used by the device. Hereinafter, a device moving within the environment is referred to as a moving entity.
[0015] According to the present invention, a high-level description of the environment is a three-dimensional vector model containing 3D feature points. These 3D feature points can be associated with, or linked to, three-dimensional objects in the 3D vector model. Because these 3D feature points are true 3D points, the information associated with them is more extensive than that associated with 2D feature points commonly used today—that is, a more detailed spatial description—so 3D feature points can be used as reference points for various representations generated by different types of sensors. In other words, each reference point can carry properties known to feature points for one sensor type, as well as additional properties characterizing that reference point for different sensor types and their respective processing.
[0016] As discussed further above, for the same scene, prominent feature points may differ depending on the sensor and the processing used. This may make a single feature point useful for one sensor representation but unavailable for another. Objects can be defined, for example, by their outlines and have properties associated with them in a more abstract way. Such objects can be identified fairly easily in different representations derived from different types of sensors, and this invention takes advantage of this. Once such an object is identified, one or more feature points associated with the identified object and the specific sensor type can be obtained, identified in the object's corresponding locally generated sensor representation, and used to determine its relative position relative to the object. Object identification can be aided by using a 3D vector model that may already include the object's position. In this case, only those positions in the sensor representation need to be analyzed, and the object's shape can also be more easily identified in the sensor representation when it is known in advance from the data provided by the 3D vector model. Identifying feature points of a specific sensor type in an object may include reference to instructions on how to find feature points associated with the object in the 3D vector model. The instructions may include filtering or processing parameters for filtering or processing representations of the scene generated by different types of sensors to locate feature points, or may include the position of the feature points relative to the object, such as the object's outline. If the absolute location of an object or one or more feature points associated with it is known, the absolute location can be determined simply. Single-dimensional feature points can be used to determine location. However, if multiple feature points from a single object, multiple feature points from multiple different objects, or feature points unrelated to the object form a specific and unique pattern, this pattern can be used to determine location. For example, this allows the use of feature points associated with multiple road markings that form a unique pattern to determine location.
[0017] According to an embodiment of a first aspect of the present invention, a method for generating an environmental reference model for positioning includes receiving a plurality of datasets representing a scanned environment. The datasets may be generated by multiple mobile entities moving within the environment. The data also includes information about objects and / or feature points identified in the scanned environment and the types of sensors used, and may further include sensor attributes and data for determining the absolute positions of the feature points and / or objects represented by the data. The data for determining the absolute positions of the feature points and objects may include geographic locations represented in a suitable format, such as latitude and longitude, or may simply consist of data describing the objects. In the latter case, once an object is identified, its position is determined by obtaining corresponding information about the object's position from a database. The absolute positions of the object's feature points can then be determined based on the known positions of the objects and sensors, according to the known positions of the objects.
[0018] The method also includes generating a three-dimensional vector representation of the scanned environment from the received dataset, which is aligned with a reference coordinate system, and objects / feature points are represented at their corresponding positions in the three-dimensional vector representation.
[0019] The received dataset can represent the scanned environment as a locally generated 3D vector model, which includes at least the objects and / or feature points at their respective locations, as well as information about the sensor types used to determine the objects and feature points. This 3D vector model representation of the received scanned environment can facilitate the integration of partial representations of the larger environment into a global environment reference model. At least if the data source generates a 3D vector model for determining its position within the environment, the need for some processing power at the data source may be irrelevant.
[0020] By matching objects and / or feature points, the received 3D vector model can be aligned with an existing 3D vector representation or a portion thereof of the environmental reference model. Where the received dataset does not share any parts with the existing environmental reference model, the position and / or orientation information provided in the received dataset can be used to non-continuously align the received 3D vector model within blank areas of the existing environmental reference model.
[0021] However, the received dataset can also represent the scanned environment in other forms, including but not limited to pictures or images of the environment, or other similar image representations enhanced by feature point indications. Other forms include processed abstract representations, such as machine-readable descriptions of objects, object features, etc., or identified feature points and their locations in the environment. One or more forms of representation may be preferable to others because they may require less data during transmission and storage.
[0022] The method also includes extracting objects and / or feature points from each element in the dataset. Particularly when the received dataset does not represent the scanned environment as a locally generated 3D vector model, extraction may include analyzing images, pictures, or abstract representations used to identify objects and / or feature points, or using identified feature points from the dataset. Once the feature points and / or objects are extracted, their positions in a reference coordinate system are determined. This reference coordinate system may be aligned with absolute geographic coordinates.
[0023] The dataset may optionally include information about the effective environmental conditions in which the data was generated, which may be useful, for example, for assigning confidence values to the data or for selecting an appropriate algorithm for identifying feature points.
[0024] The method also includes: creating links between objects and / or feature points in a 3D vector model using at least one type of sensor, which allows objects and / or feature points to be detected in the environment, and storing the 3D vector model representation and links in a retrievable manner.
[0025] According to an embodiment of a second aspect of the invention, a method for adaptively providing a reference frame of a first environment for positioning includes receiving a request for a reference frame of the first environment. The first environment may be part of a larger environment in which a vehicle or other moving entity is moving and its position needs to be determined. The request also includes information about at least one type of sensor that can be used to create a local representation of the environment. The reference frame is preferably a 3D vector model including objects and feature points, and any reference herein to a reference frame includes such a 3D vector model.
[0026] A request for a reference frame for the first environment may include an indication of the environment's location and / or direction of observation or travel, thereby enabling the identification of one or more candidate reference frames. This can be used when a relatively small number of reference frames for the first environment need to be transmitted, for example, due to limited storage capacity of the reference frames on the receiver side. When the receiver-side storage capacity is very large, a larger portion of the environment's reference frames can be transmitted. The indication of the environment's location may include coordinates from a satellite navigation system, or it may simply include one or more identifiable objects by which the location can be determined. If a unique identifiable object is identified (e.g., the Eiffel Tower), that unique identifiable object may be sufficient to roughly determine the location and provide a suitable reference frame.
[0027] The method also includes retrieving a three-dimensional vector model representation of the first environment from memory and generating a first reference frame from the three-dimensional vector model, the first reference frame including at least those feature points that are linked to or associated with at least one type of sensor. If the first reference frame includes only those feature points associated with the types of sensors available at the receiver, the amount of data to be transmitted can be reduced, although this may not be necessary depending on the amount of data required to describe the feature points and the data rate of the communication connection used in a particular implementation. The first reference frame may be a two-dimensional image or a graphical abstract image of the first environment, or it may be represented as a three-dimensional representation, such as a stereoscopic image, holographic data, or a 3D vector representation of the first environment. In response to the request, the first reference frame is transmitted to the requesting entity.
[0028] The method described above can be performed by a server or database located remotely from the mobile entity moving within the environment. In this case, receiving and sending can simply be transmissions within different components of the mobile entity. These components can be independent hardware components or independent software components executing on the same hardware.
[0029] To determine the location or position of a moving entity in an environment, objects found in a locally generated sensor representation of the environment are identified in a received first reference frame. The information provided in the received first reference frame can support the search for and identification of objects in the sensor representation of the scanned environment. This can be done for each type of sensor available locally. The reference frame can be received from a remote server or database, or formed using a server or database provided by the moving entity. Once an object is identified, data relating one or more feature points and their relationship to the object provided in the reference frame is used to locate one or more feature points in the sensor representation of the environment. Data that can be received with the reference frame may include filter settings or processing parameters that help locate feature points in the representation of the environment or simply limit the area for searching for feature points. Once feature points are found in the sensor representation of the environment, they can be used to determine their position relative to the object, taking into account sensor properties (e.g., field of view, orientation, etc.), for example, by matching them with feature points in the reference frame. If the absolute position of one or more feature points is provided in the reference frame, the absolute position in the environment can be determined. It's important to note that the reference frame can also be a 3D vector representation of the environment. The location can be determined by matching feature points between the received 3D vector model and a locally generated 3D vector model. The reference frame can be provided, for example, to an advanced driver assistance system (ADAS) in the vehicle, such as an accelerometer or braking system. This system uses the data provided to determine the vehicle's current position and generate control data for vehicle control.
[0030] If two or more sensors are used, they will not all be in the same location or facing the same direction. Furthermore, the rates at which representations of the environment are generated or updated may differ. Using information about the location of each sensor and its respective position and orientation, a reference frame, or even a realistic 3D vector graphics model, can be drawn for each sensor's updated representation of the environment. In this way, objects and, consequently, feature points can be quickly identified across different representations of the environment, thus reducing processing costs.
[0031] According to an embodiment of the second aspect of the invention, the environment is scanned by a first type of sensor and a second type of sensor, and an object is identified in a representation of the environment generated by the first type of sensor at a first distance between the object and the sensor. A reference frame of the environment including the identified object is received, and information about the object that can be detected by the first type of sensor and the second type of sensor and feature points relating to the identified object is extracted from the reference frame. The reference frame is preferably a 3D vector representation including the object and feature points. When approaching the object, at a second distance less than the first distance, the object is also identified in the representation of the environment generated by the second type of sensor using the extracted information. Similarly, the extracted information is used to identify feature points in the representation of the environment generated by the second sensor, and at least the feature points identified in the representation of the environment generated by the second sensor are used to determine the position in the environment. Since the object has been identified and the sensor type is known, the position of the feature points relative to the object can be found more easily and quickly in the representation of the environment generated by the second type of sensor, because the positions of the feature points relative to the object of both sensor types are provided in the reference frame of the environment. Therefore, once the object is identified in the representation of the environment generated by the second type of sensor, sensor data processing can be prepared to search for feature points only in those portions of the representation that are known to include feature points. The extracted information can include data about the absolute location of feature points in the environment, thus enabling the determination of their location within the environment.
[0032] With the second type of sensor providing a higher resolution representation of the environment, it becomes possible to determine one's position within the environment more accurately, making such high-precision positioning easier and faster. This, in turn, allows for movement within the environment at higher speeds without compromising safety.
[0033] This embodiment may also be useful when the second type of sensor is hampered by environmental conditions (e.g., fog, drizzle, heavy rain, snow, etc.) while the first type of sensor is not. For example, a radar sensor of a vehicle traveling along a road can detect objects on the roadside at a great distance, even if drizzle obscures the view provided by the vehicle's camera. Therefore, objects can be located in a reference frame of the environment based on data from the radar sensor, and the positions of feature points on both the radar and camera can be identified even if the camera has not yet provided a useful image. As the vehicle approaches the object, the camera image provides a useful image. Since the object's position is known from the reference frame and may have already been tracked using the radar image, feature points in the camera image can be identified more easily and quickly by referring to the information provided in the reference frame. High-precision vehicle positioning can be easier and faster than having to identify feature points anywhere in the camera image without any reference cues, because the camera image can have a higher resolution. In other words, a kind of cooperation between different types of sensors is achieved through a reference frame of the environment and the data provided therein.
[0034] A first apparatus for generating an environmental reference model for localization includes a first module adapted to receive a plurality of datasets representing a scanned environment, the datasets further including information about the type of sensors used and data for determining the absolute positions of objects and / or feature points represented by the datasets. The first apparatus also includes a second module adapted to extract one or more objects and / or feature points from each of the datasets and determine the positions of the objects and / or feature points in a reference coordinate system. The first apparatus also includes a third module adapted to generate a three-dimensional vector representation of the scanned environment aligned with the reference coordinate system, and the objects / feature points are represented at their respective positions in the three-dimensional vector representation. The first apparatus further includes: a fourth module adapted to create links between objects and / or feature points in the three-dimensional vector model using at least one type of sensor, by which objects and / or feature points can be detected in the environment; and a fifth module adapted to store the three-dimensional vector model representation and links in a retrievable manner.
[0035] A second apparatus for adaptively providing a reference frame for a first environment for localization includes a sixth module adapted to receive a request for the reference frame of the first environment, the request including at least one type of sensor that can be used to create a local representation of the environment. The second apparatus further includes: a seventh module adapted to retrieve a three-dimensional vector model representation including the first environment from a memory; and an eighth module adapted to generate a first reference frame from the three-dimensional vector model, the first reference frame including at least those feature points linked to at least one type of sensor. The second apparatus further includes: a ninth module adapted to send the first reference frame to the requesting entity.
[0036] A third apparatus for determining the position of a moving entity in an environment includes: a tenth module adapted to scan the environment using a first-type sensor; and an eleventh module adapted to identify an object in a representation of the scanned environment generated by the first-type sensor. The third apparatus further includes: a twelfth module adapted to receive a reference frame of the environment including the object, the reference frame including information about objects and / or feature points in the environment that can be detected by the first-type sensor; and a thirteenth module adapted to extract information from the reference frame about objects that can be detected by the first-type sensor and at least one feature point relating to the identified object. The third apparatus further includes: a fourteenth module adapted to use the information about the object and / or feature point from the reference frame to identify at least one feature point in the sensor representation of the environment; and a fifteenth module adapted to determine the position in the environment using at least one feature point identified in the sensor representation of the environment and information about the absolute position of the at least one feature point extracted from the reference frame.
[0037] In the improvement to the third device, in addition to using a first type of sensor, the tenth module is adapted to scan the environment using a second type of sensor, and the eleventh module is adapted to identify objects in a representation of the environment generated by the first type of sensor at a first distance between the sensor and the object. The twelfth module is adapted to receive a reference frame of the environment including the object, the reference frame including information about objects and / or feature points in the environment that can be detected by the first and second type of sensors. The thirteenth module is adapted to extract information from the reference frame about objects that can be detected by the first and second type of sensors and at least one feature point relating to the identified object. The fourteenth module is adapted to use the extracted information to identify objects in a representation of the environment generated by the second type of sensor at a second distance (the second distance is less than the first distance) between the sensor and the object, and is adapted to use the extracted information to identify one or more feature points in the representation of the environment generated by the second type of sensor. The fifteenth module is adapted to determine the position in the environment using at least one feature point identified in the representation of the environment generated by the second type of sensor and information about the absolute position of at least one feature point extracted from the reference frame.
[0038] One or more of the first to fifth modules of the first device, the sixth to ninth modules of the second device, and / or the tenth to fifteenth modules of the third or fourth device may be dedicated hardware modules. Each module includes one or more microprocessors, random access memory, non-volatile memory, and interfaces for inter-module communication and communication with data sources and data aggregations that are not part of the overall system. References to modules used for receiving or transmitting data, even those referred to above as separate modules, may be implemented in a single communication hardware device and are distinguished only by their role in the system or by the software used to control the communication hardware to perform the module's function or role.
[0039] One or more of modules one through fifteen can also be computer software programs that execute on a computer and provide the functionality of the corresponding module. Combinations of dedicated hardware modules and computer software programs are also conceivable.
[0040] This method and apparatus enable the determination of location or localization within an environment using a compact dataset, with updates requiring only a relatively small amount of data to be sent and received. Furthermore, a significant portion of image and data processing, as well as data fusion, is performed on a central server, reducing the processing power requirements of mobile devices or equipment. Additionally, a single 3D vector graphics model can be used to generate reference frames on demand for sensor type selection, excluding feature points not detected by the selected sensor type. Therefore, regardless of the type of sensor used to scan the environment and objects within it, by referencing the 3D vector graphics model and the information it provides, feature points can be found more easily and quickly in the scanned environment, and locations within the environment can be determined more easily and quickly.
[0041] If the vehicle is moving in an environment where a 3D vector model has already been generated, the reference frame provided to the vehicle is the 3D vector model, and the vehicle generates its own 3D vector model locally during scanning. The vehicle can easily align its locally generated 3D vector model with the 3D vector model received as the reference frame. Furthermore, the server that generates and provides the environment reference model can easily align the locally generated 3D vector model received in the dataset with its environment reference model using the identified feature points. Attached Figure Description
[0042] In the following sections, the invention will be described with reference to the accompanying drawings, wherein,
[0043] Figure 1 An exemplary simplified flowchart of a method according to one or more aspects of the present invention is shown;
[0044] Figure 2 An exemplary simplified block diagram of a mobile system according to one or more aspects of the present invention is shown; and
[0045] Figure 3 An exemplary simplified block diagram of a remote system according to one or more aspects of the present invention is shown. Detailed Implementation
[0046] In the accompanying drawings, the same or similar elements are indicated by the same reference numerals.
[0047] Figure 1 An exemplary simplified flowchart of a method 100 according to one aspect of the present invention is shown. In step 102, the environment is scanned using one or more different types of sensors. The scanning can be performed continuously or periodically at fixed or variable intervals. In step 104, the representation of the environment generated by the scanning is analyzed to identify objects and / or feature points. Optionally, step 104 may also include generating a three-dimensional vector model of the environment and the identified objects and / or feature points. The representation of the environment generated by the scanning in step 102 and / or the results of the analysis performed in step 104 are transmitted to a remote server or database in step 106. Steps 102 through 106 are performed by a mobile entity moving within the environment.
[0048] In step 108, the representation of the environment generated by scanning in step 102 and / or the results of the analysis performed in step 104 are received by a remote server or database, and may be further analyzed in optional step 110. Whether further analysis is performed may depend on the received data, i.e., whether the data needs to be used for analysis to identify objects and / or feature points, or whether such analysis is not required because the received data is already in the form of a 3D vector model. In step 112, a 3D reference model of the environment, including objects and / or feature points and their locations, is generated. This may include matching or aligning the received 3D vector model or the 3D vector model generated in step 110 to obtain a coherent global 3D reference model. Steps 108 through 112 are performed by a remote server or database.
[0049] Returning to the moving entity, in step 114, the moving entity determines its position in the environment while moving, for example, using a locally available environment model, satellite navigation, etc. At some point in time, in step 116, the moving entity requests a reference frame of its currently moving environment to determine its position. Requesting a reference frame can be done periodically or triggered by an event, depending on, for example, the distance traveled since the last reference frame was requested. In step 118, a remote server or database receives the request and generates the requested reference frame for the corresponding position in step 120. The requested reference frame is transmitted to the moving entity in step 122, and the moving entity receives the reference frame in step 124 and uses it in step 126 to determine its position in the environment by comparing locally generated scan data and / or a locally generated 3D vector model with the data provided in the reference frame.
[0050] The dashed lines that close the loops between steps 106 and 102, between steps 126 and 114, and between steps 126 and 102 indicate the repetition or continuous execution of the method or the individual loops. Other loops may also be conceived depending on the requirements and specific implementation of the method in the moving entity.
[0051] Figure 2 An exemplary simplified block diagram of a mobile system 200 according to one or more aspects of the present invention is shown. The following items are communicatively connected via one or more bus systems 216: a scanner 202 for scanning an environment in which the mobile system is moving using a first type of sensor; a device for determining a position 204 in the environment; a module 206 for identifying objects in a representation of the scanned environment; a module 208 for receiving a 3D vector graphic model of the environment, which includes objects; a module 210 for extracting information from the 3D vector graphic model about objects and at least one feature point relating to the identified objects, which objects and feature points are detectable by the first type of sensor; a module 212 for identifying at least one feature point in a sensor representation of the environment using information about the objects and / or feature points extracted from the 3D vector graphic model; and a module 214 for determining a position in the environment using at least one feature point identified in the sensor representation of the environment and information about the absolute position of at least one feature point extracted from the 3D vector graphic model.
[0052] Modules 206, 208, 210, 212, and / or 214 may include one or more microprocessors, random access memory, non-volatile memory, and software and / or hardware communication interfaces. The non-volatile memory may store computer program instructions that, when executed by one or more microprocessors in cooperation with the random access memory, perform one or more processing steps of the method described above.
[0053] Figure 3An exemplary simplified block diagram of a remote system 300 according to one or more aspects of the present invention is shown. The following items are communicatively connected via one or more bus systems 314: module 302 for communicating with a mobile entity, the communication including: receiving a plurality of datasets representing a scanned environment and receiving a request for a reference frame and sending the reference frame; module 304 for extracting one or more objects and / or feature points from each of the datasets and determining the positions of the objects and / or feature points in a reference coordinate system; module 306 for generating a three-dimensional vector representation of the scanned environment, the three-dimensional vector representation being aligned with the reference coordinate system, and the objects and / or feature points being represented at their respective positions in the three-dimensional vector representation; module 308 for creating links between objects and / or feature points in a three-dimensional vector model using at least one type of sensor, by which the objects and / or feature points can be detected in the environment; module 310 for storing the three-dimensional vector model representation and the links in a retrievable manner; and module 312 for retrieving the three-dimensional vector model representation and generating a first reference frame from the three-dimensional vector model based on a request for a reference frame from the memory, the first reference frame including at least those feature points that are linked to at least one type of sensor.
[0054] Modules 302, 304, 306, 308, 310, and / or 312 may include one or more microprocessors, random access memory, non-volatile memory, and software and / or hardware communication interfaces. The non-volatile memory may store computer program instructions that, when executed by one or more microprocessors in cooperation with the random access memory, perform one or more processing steps of the methods described above.
Claims
1. An apparatus for determining a position in an environment, comprising: Memory, which stores instructions; as well as At least one processor, communicatively coupled to the memory and configured to execute the instructions to perform the following operations: Receive scan data representing the environment being scanned from the first type of sensor; Determine the initial position in the environment; A request for a reference frame is sent based on the initial position in the environment, the request identifying the first type of sensor; In response to the request, the reference frame is received, the reference frame including information about at least one object in the environment that can be detected by the first type of sensor and the second type of sensor, and information about a first plurality of feature points relating to the at least one object; Receive additional scan data representing the scanned environment from the second type of sensor; Detecting the at least one object within the scan data and within the additional scan data; and The location in the environment is determined based at least in part on the at least one object detected in the scan data and the at least one object detected in the additional scan data.
2. The apparatus according to claim 1, wherein, The at least one processor is configured to execute the instructions to perform the following operations: A second plurality of feature points are determined based on the scan data; Generate a dataset, the dataset including information identifying the first type of sensor and the second plurality of feature points; as well as The location in the environment is also determined based on the dataset and the first plurality of feature points.
3. The apparatus according to claim 2, wherein, The at least one processor is configured to execute the instructions to send the dataset for generating a three-dimensional vector representation of the environment.
4. The apparatus according to claim 1, wherein, The reference frame includes information for identifying the at least one object within a search area, and wherein the at least one processor is configured to execute the instructions to detect the at least one object within the scan data based on the search area.
5. The apparatus according to claim 1, wherein, The reference frame includes information for identifying the first plurality of feature points within the search area, and wherein the at least one processor is configured to execute the instructions to detect the first plurality of feature points within the scan data based on the search area.
6. The apparatus according to claim 1, wherein, The reference frame identifies the absolute position of at least one of the first plurality of feature points, and wherein the at least one processor is configured to execute the instructions to perform the following operations: Extract the absolute position from the reference frame; and The location in the environment is also determined based on the absolute location of the at least one feature point.
7. The apparatus according to claim 1, wherein, The reference frame corresponds to a three-dimensional vector representation of the environment, including the first plurality of feature points and the at least one object.
8. The apparatus according to claim 7, wherein, The reference frame includes a link between (i) the at least one object and (ii) at least one of the first plurality of feature points and the first type of sensor.
9. The apparatus according to claim 1, wherein, The at least one processor is configured to: Based on the scan data, a second plurality of feature points are determined; and The location in the environment is also determined by matching the first plurality of feature points with the second plurality of feature points.
10. The apparatus according to claim 9, wherein, The reference frame includes a first three-dimensional vector representation of the first plurality of feature points, and wherein the at least one processor is configured to: A second three-dimensional vector representation is generated based on the second plurality of feature points; and The location in the environment is also determined by matching the first three-dimensional vector representation with the second three-dimensional vector representation.
11. The apparatus according to claim 1, wherein, The reference frame is a three-dimensional vector model representation of the scanned environment; and The at least one processor is configured to execute the instructions to perform the following operations: Generate a local 3D vector model representation based on the scan data; and The location in the environment is determined by matching the received 3D vector model representation with the local 3D vector model representation.
12. The apparatus according to claim 1, wherein, The reference frame includes filter settings, wherein the filter settings are used to limit the region in which the first plurality of feature points are searched.
13. A method for determining a location in an environment, comprising: Receive scan data representing the environment being scanned from the first type of sensor; Determine the initial position in the environment; A request for a reference frame is sent based on the initial position in the environment, the request identifying the first type of sensor; In response to the request, the reference frame is received, the reference frame including information about at least one object in the environment that can be detected by the first type of sensor and the second type of sensor, and information about a first plurality of feature points relating to the at least one object; Receive additional scan data representing the scanned environment from the second type of sensor; Detecting the at least one object within the scan data and within the additional scan data; and The location in the environment is determined based at least in part on the at least one object detected in the scan data and the at least one object detected in the additional scan data.
14. The method of claim 13, comprising: A second plurality of feature points are determined based on the scan data; Generate a dataset, the dataset including information identifying the first type of sensor and the second plurality of feature points; as well as The location in the environment is also determined based on the dataset and the first plurality of feature points.
15. The method of claim 14, comprising: Send the dataset to generate a three-dimensional vector representation of the environment.
16. The method according to claim 13, wherein, The reference frame identifies the absolute position of at least one feature point among the first plurality of feature points, and the method includes: Extract the absolute position from the reference frame; and The location in the environment is also determined based on the absolute location of the at least one feature point.
17. The method of claim 13, comprising: A second plurality of feature points are determined based on the scan data; as well as The location in the environment is also determined by matching the first plurality of feature points with the second plurality of feature points.
18. The method of claim 13, wherein The reference frame is a three-dimensional vector model representation of the scanned environment; and, The method further includes: A local 3D vector model representation is generated based on the scanned data; as well as The location in the environment is determined by matching the received 3D vector model representation with the local 3D vector model representation.
19. The method of claim 13, wherein The reference frame includes filter settings, wherein, The filter settings are used to limit the area in which the first plurality of feature points are searched.
20. A computer-readable medium for storing non-transitory instructions, wherein, When the instructions are executed by at least one processor, the at least one processor causes the processor to perform the following operations: Receive scan data representing the environment being scanned from the first type of sensor; Determine the initial position in the environment; A request for a reference frame is sent based on the initial position in the environment, the request identifying the first type of sensor; In response to the request, the reference frame is received, the reference frame including information about at least one object in the environment that can be detected by the first type of sensor and the second type of sensor, and information about a first plurality of feature points relating to the at least one object; Receive additional scan data representing the scanned environment from the second type of sensor; Detecting the at least one object within the scan data and within the additional scan data; and The location in the environment is determined based at least in part on the at least one object detected in the scan data and the at least one object detected in the additional scan data.
Citation Information
Patent Citations
Vision system and method of analyzing an image
CN102567449A
Method and apparatus for providing accurate localization for an industrial vehicle
CN103733084A
Indoor positioning method based on three-dimensional environment model matching
CN104574386A
Interactive calibration method and apparatus based on three dimensional reconstruction in three dimensional monitoring system
CN105678748A