Method and system for tracking mobile devices
By collecting environmental data during vehicle movement to generate an initial geometric model and combining it with visual tracking methods, the initialization and scaling factor determination challenges of handheld SLAM systems are solved, enabling more accurate camera pose estimation and environmental model generation, which is suitable for augmented reality and navigation applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2013-12-19
- Publication Date
- 2026-06-02
AI Technical Summary
In existing handheld augmented reality applications, the initialization and determination of the metric scaling factor of monocular SLAM systems are difficult, and users have difficulty in performing different camera movements to obtain sufficient displacement, resulting in insufficient accuracy of camera pose estimation and environmental geometry model.
By collecting data from vehicle sensors as it travels through the environment, an initial geometric model is generated. A vision-based tracking method is then used to track the mobile device on a handheld device, reducing reliance on manual user movement. Finally, a geometric model of the real environment is generated by combining vehicle sensor data and camera image information.
It simplifies the initialization process of SLAM systems, improves the accuracy of camera pose estimation and environment modeling, reduces reliance on user operations, and is suitable for augmented reality and navigation applications on mobile devices.
Smart Images

Figure CN116485870B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 201910437404.X, filed on December 19, 2013, entitled "Method and System for Tracking Mobile Devices". Patent application No. 201910437404.X is a divisional application of patent application No. 201380081725.1, filed on December 19, 2013, entitled "Method and System for Tracking Mobile Devices". Background Technology
[0002] This disclosure relates to a method for tracking a mobile device including at least one camera in a real environment, and to a method for generating a geometric model of at least a portion of a real environment using image information from at least one camera of the mobile device, the method comprising receiving image information associated with at least one image captured by at least one camera.
[0003] Camera pose estimation and / or digital reconstruction of real-world environments are challenging tasks common in many applications and fields, such as robot navigation, 3D object reconstruction, and augmented reality visualization. For example, known systems and applications, such as augmented reality (AR) systems and applications, enhance information about the real environment by providing visualizations that overlay computer-generated virtual information onto a view of the real environment. Virtual information can be any type of visually perceptible data, such as objects, text, images, videos, or combinations thereof. The view of the real environment can be perceived by the user's eyes as a visual impression and / or acquired as one or more images captured by a camera held by the user or attached to a device held by the user.
[0004] Camera pose estimation involves calculating the spatial relationship or transformation between the camera and a reference object (or environment). Camera motion estimation involves calculating the spatial relationship or transformation between the camera in one position and the camera in another position. Camera motion, also known as camera pose, describes the camera's orientation in one position relative to the same camera in another position. Camera pose or motion estimation is also called camera tracking. Spatial relationships or transformations describe translation, rotation, or combinations thereof in 3D space.
[0005] Vision-based methods are considered robust and popular for calculating camera pose or motion. They calculate the camera's pose (or motion) relative to the environment based on one or more images of the environment captured by the camera. These vision-based methods rely on the captured images and require detectable visual features within those images.
[0006] Simultaneous Localization and Mapping (SLAM) based on computer vision (CV) is a well-known technique for determining the pose and / or orientation of a camera relative to a real-world environment and creating a geometric model of that environment without any prior knowledge of it. The creation of this geometric model is also known as environment reconstruction. Vision-based SLAM can facilitate numerous applications, such as navigation for robotic systems or vehicles. Specifically, supporting mobile augmented reality (AR) in unknown real-world environments is a promising technology.
[0007] Many SLAM systems require initialization to obtain an initial part of the environment model. Initialization must be accomplished using different camera movements between two images captured of the real environment. These different movements require capturing two images from two different camera positions that have sufficient displacement relative to their distance from the environment. It's important to note that rotational camera movement alone produces undesirable results. One of the major limitations of using SLAM devices in AR, especially handheld or mobile AR, is that they are not easy to use and require the user to somehow move the device to make the system work. Rotational camera movement is a natural movement of a user looking around in a real environment and commonly occurs in many AR applications. However, rotational camera movement alone can produce undesirable results for monocular SLAM.
[0008] Furthermore, a single camera does not measure the metric scale. Another limitation of using monocular SLAM systems in AR lies in the reconstructed camera pose and the geometric model of the environment, which depend on the scale as an undetermined factor. The undetermined scaling factor poses a challenge to accurately overlaying virtual visual information onto the real environment in the camera image.
[0009] Currently, many cities or buildings have geometric models derived from 3D reconstructions or their blueprints. However, due to the frequent development and changes in urban construction, most of these models are not up-to-date. Specifically, parking lots often lack geometric models or up-to-date models because parked vehicles change over time.
[0010] Various monocular vision-based SLAM systems have been developed for vehicles, particularly for mobile handheld AR applications. Common challenges and limitations in their use include the initialization of the SLAM system and the determination of the metric scaling factor. Initializing a SLAM system requires different camera movements to acquire two images of the real-world environment, such that two images are captured from two different camera positions with sufficient displacement relative to their distance from the environment. The quality of camera pose estimation and any generated geometric models explicitly depends on the initialization.
[0011] Implementing different camera movements to achieve qualified SLAM initialization is challenging in handheld AR applications, where users holding the camera may be unaware of the importance of camera movement or even find it difficult to implement different movements. Therefore, it is desirable to simplify the startup process or even make it invisible to the user.
[0012] Furthermore, a single camera does not measure the metric scale. The camera pose and the reconstructed environment model from monocular vision-based SLAM depend on an undetermined scaling factor. An appropriate scaling factor defines the true camera pose and the size of the reconstructed environment model in the real world.
[0013] The first well-known monocular vision-based SLAM system was developed by Davison et al. They required the camera to have sufficient displacement between acquiring images of each newly observed part of the regional environment. To determine a suitable scaling factor, they introduced an additional calibration object with known geometric dimensions.
[0014] Lemaire et al. proposed using a stereo camera system to solve the problem of requiring camera movement and determining the scaling factor. However, using a stereo camera is only a partial solution, because the displacement between the two cameras relative to their distance from the environment must be significant in order to reliably calculate the depth of the environment. Therefore, a handheld stereo system cannot completely solve the problem, and additional different movements by the user remain indispensable.
[0015] Lieberknecht et al. incorporated depth information into monocular vision-based SLAM by employing an RGB-D camera that provides depth information associated with image pixels, allowing for accurate scaling of camera pose estimation. The scaling factor can be determined based on the known depth information. However, RGB-D camera devices are not commonly used in handheld devices such as mobile phones or PDAs compared to standard RGB cameras. Furthermore, common low-cost RGB-D cameras that should be considered as candidates for integration into handheld devices are typically based on infrared projection, such as the Kinect system from Microsoft or the Xtion Pro from Asus. These systems are readily available, inexpensive consumer devices.
[0016] US 8,150,142 B2 and US 7,433,024 B2 describe detailed possible implementations of RGB-D sensors. However, these systems present problems when used outdoors during the day due to the presence of sunlight.
[0017] Gauglitz et al. developed a system for camera pose estimation and environment model generation that can be used for both general camera motion and rotation-only camera motion. For rotation-only motion, their method creates a panoramic map of the real environment, rather than a 3D geometric model of the real environment. Summary of the Invention
[0018] One object of this disclosure is to provide a method for tracking a mobile device including at least one camera in a real environment, and a method for generating a geometric model of at least a portion of the real environment using image information from at least one camera of the mobile device, wherein the challenges and limitations of using SLAM methods (such as initialization) are reduced and startup is simplified for the user.
[0019] According to one aspect, a method is provided for tracking a mobile device including at least one camera in a real environment, the method comprising: receiving image information associated with at least one image captured by the at least one camera; generating a first geometric model of at least a portion of a real environment, distinct from the mobile device, based on environmental data or vehicle state data acquired during acquisition by at least one sensor of a vehicle; and performing a tracking process based on the image information associated with the at least one image and at least partially based on the first geometric model, wherein the tracking process determines at least one parameter of the pose of the mobile device relative to the real environment.
[0020] According to another aspect, a method is provided for generating a geometric model of at least a portion of a real environment using image information from at least one camera of a mobile device, the method comprising: receiving image information associated with at least one image captured by at least one camera; generating a first geometric model of at least a portion of a real environment based on environmental data or vehicle state data acquired during acquisition by at least one sensor of a vehicle, the vehicle being different from the mobile device; and generating a second geometric model of at least a portion of the real environment based on the image information associated with at least one image and at least in part according to the first geometric model.
[0021] According to the present invention, tracking a mobile device equipped with at least one camera in a real environment and / or generating a geometric model of the environment using at least one camera is performed by using image information associated with at least one image captured by at least one camera. Tracking the mobile device or generating a second geometric model is performed at least in part based on a first geometric model known to the real environment or a portion thereof. The first geometric model is created based on environmental data collected by at least one sensor of the vehicle. Specifically, the environmental data is collected while the vehicle is driving in the environment.
[0022] The mobile device can be transported by the vehicle during or part of the data acquisition process for collecting environmental data. Thus, the data acquisition process is at least partially performed while the mobile device is being transported by the vehicle. Tracking the mobile device or generating a second geometric model can be performed within a time period following or after the environmental data acquisition process. This time period can be 2 hours, 12 hours, or 24 hours.
[0023] A vehicle is specifically a mobile machine capable of transporting one or more people or goods. Vehicles can be, but are not limited to, bicycles, motorcycles, cars, trucks, forklifts, airplanes, or helicopters. Vehicles may or may not have an engine.
[0024] The acquisition of environmental data for creating the first geometric model can begin at any time or only when certain conditions are met, such as when the vehicle approaches a known destination known to the navigation system, or when the vehicle's speed is below a certain threshold. A condition can also be one of several vehicle states, such as speed, odometer readings, engine status, braking system, gear position, lighting conditions, or the state of an aircraft ejection seat. Similarly, a condition can be one of several mobile device states, such as whether the mobile device is inside or outside the vehicle, the distance of the mobile device to the destination, or sudden movements of the mobile device that are inconsistent with the vehicle's movement (e.g., sudden acceleration relative to the vehicle).
[0025] According to one implementation, at least a portion of the first geometric model can be generated based on one or more images captured by at least one camera.
[0026] According to one implementation, the generation of the second geometric model is performed within a set time period after the acquisition process or a part of the acquisition process, preferably within 24 hours.
[0027] According to another embodiment, the generation of the second geometric model is further based on received image information associated with at least one additional image captured by at least one camera, or further based on received depth information associated with at least one image.
[0028] According to one implementation scheme, the second geometric model is generated by extending the first geometric model.
[0029] Preferably, the acquisition process is performed at least partially while the vehicle is moving and the sensor data is acquired from at least one sensor of the vehicle at different vehicle locations.
[0030] According to one implementation scheme, environmental data is collected based on the vehicle's location and at least one designated destination. For example, environmental data may be collected after the vehicle arrives at at least one destination, or when the vehicle is within a distance of at least one destination, or environmental data may be collected based on the vehicle's location, vehicle speed, and at least one destination.
[0031] According to one implementation, the first geometric model is further generated based on image information associated with at least one image captured by another camera placed in a real environment, which is different from the camera of the mobile device.
[0032] According to one embodiment, at least one sensor of the vehicle includes at least two vehicle cameras with a known spatial relationship between them, and the metric scale of the first geometric model is determined based on the spatial relationship.
[0033] According to another embodiment, the generation of the first geometric model or a portion thereof is performed by the vehicle's processing equipment, and the first geometric model is transmitted from the vehicle to the mobile device. For example, the first geometric model is transmitted from the vehicle to the mobile device via a server computer, via point-to-point communication between the vehicle and the mobile device, or via broadcast or multicast communication (e.g., vehicle data).
[0034] According to one implementation, environmental data is transmitted from a vehicle to a mobile device, and the generation of a first geometric model or a portion thereof is performed on the mobile device. For example, the environmental data is transmitted from the vehicle to the mobile device via a server computer or via point-to-point communication between the vehicle and the mobile device.
[0035] According to another implementation, environmental data is transmitted from the vehicle to a server computer, and the generation of a first geometric model or a portion thereof is executed on the server computer.
[0036] According to one implementation, the first geometric model has a suitable metric determined by vehicle-mounted sensors such as radar, distance sensors and / or transit time sensors, and / or accelerometers, and / or gyroscopes, and / or GPS, and / or star trackers, and / or based on vehicle states such as vehicle speed.
[0037] For example, one or more routes to the destination are provided and environmental data is collected, and / or a first geometric model is generated based on one or more of the provided routes.
[0038] According to one implementation scheme, at least one of the first geometric model and the second geometric model describes at least the depth information of the real environment.
[0039] Preferably, the mobile device is a portable device for the user, especially a handheld device, mobile phone, head-mounted glasses or helmet, wearable device or implantable device.
[0040] In a preferred embodiment, the method is adapted for use in augmented reality and / or navigation applications running on mobile devices.
[0041] According to one implementation, vision-based tracking is performed during the tracking process or to generate a second geometric model. For example, vision-based tracking is vision-based simultaneous localization and mapping (SLAM). Vision-based tracking may include feature extraction, feature description, feature matching, and pose determination. For example, the features used may be at least one of the following or a combination thereof: intensity, gradient, edge, line segment, segment, corner, descriptive feature, primitive, histogram, polarity, and orientation.
[0042] Therefore, this invention describes a method for supporting vision-based tracking or environment reconstruction. The method disclosed in this invention also eliminates the requirement for different camera movements to initialize monocular SLAM, as described above.
[0043] According to another aspect, the present invention also relates to a computer program product comprising software code segments adapted to perform the methods described according to the present invention. Specifically, the software code segments are contained on a non-transitory computer-readable medium. The software code segments may be loaded into the memory of one or more processing devices described herein. Any processing device used may communicate via a communication network, such as via a server computer or point-to-point communication as described herein. Attached Figure Description
[0044] Aspects and embodiments of the invention will now be described with reference to the accompanying drawings, wherein:
[0045] Figure 1 A flowchart illustrating a method using SLAM according to an embodiment of the present invention is shown.
[0046] Figure 2 Exemplary implementations of detecting, describing, and matching features that can be used in tracking or reconstruction methods are shown.
[0047] Figure 3 A flowchart illustrating a method for generating a geometric model of the environment based on environmental data collected by vehicle sensors and for tracking equipment based on the generated environmental model, according to an embodiment of the present invention, is shown.
[0048] Figure 4 An exemplary application scenario of parking a vehicle according to an embodiment of the present invention is shown.
[0049] Figure 5A flowchart illustrating an implementation of a tracking method that matches a set of current features with a set of reference features based on camera images is shown.
[0050] Figure 6 The standard concept of triangulation is shown. Detailed Implementation
[0051] While various embodiments are described below with reference to certain components, any other configuration of the components described herein or obvious to those skilled in the art may be used in implementing any of these embodiments.
[0052] The following describes implementation schemes and exemplary scenarios, which should not be construed as limiting the invention.
[0053] Augmented Reality :
[0054] Augmented reality (AR) systems can augment the real-world environment with computer-generated information. The real-world environment can be enhanced by providing computer-generated audio information. One example is navigating a visually impaired person in a real-world environment based on computer-generated verbal instructions using GPS data or other tracking technologies. Computer-generated information can also be haptic feedback, such as the vibration of a mobile phone. In navigation applications, AR systems can generate vibrations to warn a user if they have gone astray.
[0055] Augmented reality (AR) is generally accepted to visually enhance a real environment by providing visualizations that overlay computer-generated virtual visual information onto visual impressions or images of the real environment. Virtual information can be any type of visually perceptual data, such as objects, text, images, videos, or combinations thereof. The real environment can be captured as a visual impression by the user's eyes or as one or more images captured by a camera worn by the user or attached to a device held by the user. Virtual visual information is overlaid or superimposed on the real environment or a portion of the real environment in camera images or visual impressions at the appropriate time, place, and manner to provide the user with a satisfactory visual experience.
[0056] For example, in a well-known optical see-through display with translucent glass, a user can see an overlay of virtual visual information with the real environment. The user then sees objects in the real environment enhanced by the virtual visual information integrated into the glass. Users can also see the overlay of virtual information with the real environment in a well-known video see-through display with a camera and a normal display device such as a screen. The real environment is captured by the camera, and the overlay of virtual data and the real environment is displayed to the user on the screen.
[0057] Virtual visual information should be superimposed on the real environment at desired pixel locations within an image or visual impression, for example, in a perspective-appropriate manner, i.e., derived from the observed real environment. To achieve this, the pose of the camera or the user's eye—that is, its orientation and position relative to the real environment or a portion thereof—must be known. Furthermore, it is preferable to superimpose virtual visual information on the real environment to achieve visually appropriate occlusion perception or depth perception between the virtual visual information and the real environment. For this, a geometric model or depth map of the real environment is typically required.
[0058] Monocular vision-based (i.e., single-camera-based) SLAM is a promising technique for generating camera poses and creating geometric environment models for AR applications. Monocular SLAM is particularly beneficial for mobile AR applications running on handheld devices equipped with a single camera, as capturing camera images of the real-world environment is typically the way to perform camera pose estimation and environment model generation. For optical perspective displays, where the camera and eye have a fixed relationship, the user's eye pose can be determined based on the camera pose.
[0059] An exemplary scenario of the present invention:
[0060] Currently, people typically drive to their destination using guidance from navigation systems, much like in an unfamiliar city. These navigation systems may have navigation software running on mobile computing devices or embedded systems within vehicles. The navigation system (or software) can calculate one or more routes to the destination. However, parking spaces are often unavailable at or near the destination. Therefore, people often have to park at a location different from the final destination and switch to other modes of transportation (e.g., walking) to reach their final destination. In unfamiliar environments, people may have difficulty or expend more effort finding their way from where they are parked to their destination. To address this, the present invention proposes navigation running on a camera-equipped handheld device based on a geometric model of the environment created from environmental data collected by the vehicle's sensors.
[0061] Typically, people drive to their destination where they may not find parking, and will likely continue driving until they find free parking. They then return to their destination from where they parked. The process of collecting environmental data (e.g., images, GPS data, etc.) can begin after the vehicle arrives at its destination and end when the vehicle is parked (e.g., the engine is turned off). A digital geometric model of the real-world environment between the destination and the actual location where the vehicle is parked can then be created based on the collected environmental data. This geometric model, along with a handheld device equipped with a camera, can be used to guide people to their destination.
[0062] As a further scenario, a user parks their car in a real-world environment and can then run a navigation or augmented reality (AR) application on their handheld device equipped with a camera. Navigation and AR applications may require a known pose of the device relative to its environment. For this purpose, a geometric model of the environment can be used to determine the device's pose, as described earlier in this paper.
[0063] A camera attached to a mobile device is an appropriate sensor for tracking the device and reconstructing a geometric model of the environment. Vision-based tracking often requires a known geometric model of the environment, and pose estimation can be based on the correspondence between the geometric model and the camera image. Monocular vision-based SLAM can perform camera tracking in a real-world environment while simultaneously generating a geometric model of the environment even without a prior geometric model. However, an initial model of the environment must be created to initialize monocular SLAM by moving the camera by different displacements.
[0064] Manually initializing monocular SLAM from scratch is challenging because it's not intuitive for users to move the camera of a handheld device by sufficient displacement. Users must initialize monocular SLAM manually. Specifically, scale-based and image-based tracking or reconstruction can present problems.
[0065] Returning to the exemplary scenario above and now referring to Figure 4 Suppose there are two cars, 421 and 422, parked in the parking lot (see...). Figure 4 According to an embodiment of the invention, when a vehicle is driving in environment 410 to find a parking space, the geometric model 409 of the real environment 410 is generated from one or more images of the vehicle camera 414 of the vehicle 411 (see description 401). Figure 4 (Descriptions 402, 403, and 404). 412 indicates the field of view of the vehicle camera 414. The extent of the generated geometry 409 is schematically represented by dots in description 406. After parking, the geometry 409 of the environment is available at the mobile device 408 equipped with a camera of the vehicle passenger 413. 407 shows the field of view of the camera attached to the mobile device 408. The passenger can then use the geometry 409, or a portion thereof, along with images captured by the camera of the mobile device 408, to track the mobile device 408 in the real environment, create another geometry of the real environment, and / or extend the geometry 409.
[0066] Figure 3 A flowchart illustrating a method according to an embodiment of the present invention for generating a geometric model of a real environment based on environmental data collected by the vehicle's sensors and for tracking a mobile device based on the generated environmental model. It is assumed that the vehicle is traveling in a real environment. Figure 3 (Step 301).
[0067] Environmental data (ED) can be collected by one or more sensors mounted on the vehicle while driving in or throughout the environment. The user can manually start, resume, pause, and / or stop the environmental data collection process (ED). The collection process can also be automatically started, resumed, paused, and / or stopped (step 302) under certain conditions, such as when the vehicle approaches a known destination from the navigation system, or when the vehicle's speed is below a certain threshold, etc. A condition can also be one of several vehicle states, such as speed, odometer reading, engine status, braking system, gear position, light, distance of another object to the front or rear of the vehicle, open / closed status of the driver's side door, steering wheel lock, handbrake, open / closed status of the trunk, status of an aircraft ejection seat (i.e., ejection seat), aircraft cabin pressure, or a combination thereof. A condition can also be one of several states of the mobile device 408, such as the mobile device being inside or outside the vehicle, the distance of the mobile device to the destination, or a sudden movement of the mobile device that is inconsistent with the vehicle's movement (e.g., a sudden acceleration relative to the vehicle), etc.
[0068] The acquisition of environmental data (ED) is initiated or resumed when one or more conditions for initiating or resuming the acquisition of environmental data are met, or when the user manually triggers the initiation or resumption (step 303). Then, if the acquisition of environmental data (ED) is automatically or manually triggered and must be stopped or paused (step 304), the acquisition process is stopped or paused (step 305). These steps are performed in the vehicle.
[0069] If environmental data ED is available to a handheld device equipped with a camera belonging to a user (e.g., the driver or passenger of the vehicle), any processor device of the vehicle (not shown in the figures) generates a geometric model Md of the environment based on the environmental data ED (step 307) and then transmits the model to the handheld device (step 308), or transmits the environmental data ED to the handheld device (step 311) and then generates the environmental model Md based on the environmental data ED in the handheld device (step 312).
[0070] The environmental data ED can also be transmitted to another computer, such as a server computer located remotely from the mobile device and vehicle, and an application running on the server computer can create a geometric model Md of the environment based on the environmental data ED on such a server computer. In this configuration, the server computer communicates with the mobile device and vehicle, which are acting as client devices, in a client-server architecture. The environmental data ED and / or the geometric model Md are then transmitted from the server computer to the mobile device.
[0071] The geometric model Md can be executed whenever environmental data or a portion of environmental data is available, such as online during the environmental data acquisition process or offline after environmental data acquisition. For example, in any case where new environmental data is available, new environmental data can be integrated to generate the geometric model Md.
[0072] In cases where tracking of a handheld device must be performed in the environment (step 309), assuming that the geometric model Md is available in the handheld device, tracking is performed at least in part based on the geometric model Md (step 310). Steps 309 and 310 can be performed in the handheld device.
[0073] One or more routes to the destination can be provided or calculated. These routes can be further updated based on the current location of the vehicle or handheld device. The destination can be manually provided by the user or defined in the navigation system. Environmental data (ED) and / or geometric model (MD) can be collected based on the routes. For example, relevant portions of the environmental data (ED) and / or relevant portions of the geometric model (MD) can be collected only at locations where the user is likely to travel along the route.
[0074] The geometric model of the environment uses data from the vehicle's sensors:
[0075] For example, when driving a vehicle in an environment, the geometry of the real-world environment can be generated from depth data provided by the vehicle's depth sensors, such as distance sensors or time-of-flight cameras mounted in the vehicle. Many methods can be used to reconstruct the 3D surface of the real-world environment based on the depth data. Pushbroom scanners can be used to create 3D surfaces.
[0076] A geometric model of a real-world environment (also referred to herein as an environment model) can be created or generated while a vehicle (e.g., an automobile) is driving in that environment, using vision-based SLAM and at least one camera mounted in the vehicle. Various vision-based SLAM methods have been developed and can be employed to create environment models using images captured by at least one camera of the vehicle. Other sensors of the vehicle can also be used to support the construction of the environment model.
[0077] The geometric model created from a monocular vision-based SLAM environment depends on an undetermined scaling factor.
[0078] The appropriate scaling factor required to bring an environmental model to a metric scale can be efficiently determined by using a camera mounted in a vehicle to capture images of two points in the environment that are a known distance apart, or images of real objects with known physical dimensions. For example, the scaling factor can be estimated using traffic lights, vehicles with known 3D models, or other road infrastructure (white line distances, roadside "pillars").
[0079] A suitable scaling factor can also be recovered based on the distance between the vehicle and the environment. This distance can be measured from sensors mounted on the vehicle, such as radar, distance sensors, or time-of-flight cameras, and can also be used to determine the scaling factor. A suitable scaling factor can also be determined if a reference distance between one (or both) cameras capturing two images is known. This reference distance can be obtained, for example, from odometer measurements using the rotational speed of the wheels or GPS coordinates. For stereo cameras, the baseline distance between the centers of the two cameras can be used as a reference distance.
[0080] If the vehicle's position in the environment is known, a suitable scaling factor can also be determined. The vehicle's position in the environment can be determined from GPS or from sensors fixed in the environment (such as security cameras).
[0081] See now Figure 1 Given at least one camera, the process of creating or generating a geometric model and / or calculating the camera pose based on images captured by at least one camera may consist of feature detection (step 102 or 105), feature description (step 102 or 105), feature matching (step 106), triangulation (step 107), and optional (global) map refinement, which adjusts the triangulation position and / or camera pose, and / or removes and / or adds points from the triangulation.
[0082] The process of creating geometric models and / or calculating camera poses can also be based on using a stereo camera system.
[0083] Optical flow from the camera can also be used to generate geometric models or support generative models.
[0084] To reconstruct the environment model, at least two images must be captured by the camera at different locations. For example, in step 101, the camera captures image IA in pose PA, and then the camera moves to different displacements M to capture image IB in a pose different from pose PB (steps 103 and 104).
[0085] Feature detection can be performed using highly repeatable methods to identify features in images IA and IB. In other words, this method has a high probability of selecting a portion of the image corresponding to the same physical 3D surface as a feature for different viewpoints, rotations, and / or lighting settings (e.g., as a local feature descriptor of SIFT, a shape descriptor, or other methods known to the art). Features are typically extracted at different scales in scale space. Therefore, each feature has a repeatable scale in addition to its two-dimensional location. Furthermore, the repeatable orientation (rotation) is calculated from the intensity of pixels surrounding the feature in the region, for example, as the dominant direction of the intensity gradient.
[0086] Feature descriptors transform detected image regions into typical feature descriptors that are robust or invariant to certain types of changes, such as (uniform) illumination, rotation, and occlusion. Feature descriptors are determined to allow for feature comparison and matching. Common methods use the computed scale and orientation of the features to transform the coordinates of the feature descriptors, providing invariance to rotation and scale. For example, a descriptor can be an n-dimensional real vector constructed by connecting a histogram of functions such as gradients that connect local image intensities. Alternatively, a descriptor can be an n-dimensional binary vector.
[0087] Furthermore, each detected feature may (optionally) be associated with a (local) position and orientation relative to the environment and / or a previous pose relative to the camera. The (local) position may be obtained from GPS sensor / receiver, IR or RFID triangulation, or positioning methods using broadband or wireless infrastructure. The (local) orientation may be obtained from a compass, accelerometer, gyroscope, or gravity sensor. Since the camera is mounted in the vehicle, the (local) position and orientation relative to the camera's previous pose may be obtained from the vehicle's speed or steering.
[0088] In an image, multiple features can be detected. Feature matching involves finding the feature with the most similar descriptor in another feature set for each feature in the first feature set and storing the two features as a correspondence (match). For example, given two feature sets FA and FB detected and described in images IA and IB, the goal is to find the feature with the most similar descriptor in feature set FB for each feature in feature set FA. See [link to relevant documentation] for more details. Figure 2 It shows an image CI with feature c and its corresponding descriptor d(c) and descriptor d(r) of reference feature r.
[0089] Matching a feature set FA with a feature set FB can be accomplished by determining the corresponding similarity measure between each corresponding feature descriptor in feature set FA and each corresponding feature descriptor in feature set FB. Common examples of image similarity measures include negative or inverse sum of squared differences (SSD), negative or inverse sum of absolute differences (SAD), (normalized) cross-correlation, and interaction information. The similarity result is a real number. The larger the similarity measure result, the more similar the two visual features are.
[0090] The simplest method for feature matching is to find the nearest neighbor of the current feature descriptor through exhaustive search and select the corresponding reference feature as the match. More advanced methods use spatial data structures in the descriptor domain to accelerate matching. Common methods use approximate nearest neighbor search instead, such as those supported by spatially partitioned data structures like kd-trees.
[0091] Following feature matching, a correspondence is created between features from feature sets FA and FB. This correspondence can be 2D-2D or 2D-3D. Based on the correspondence, the camera pose is determined relative to the environment or to one of the previous camera poses. Subsequently, a (global) refinement step is typically (but optional), which may re-evaluate the correspondences discarded in the initial stage. Various methods and heuristics exist for refinement.
[0092] Features may not have associated feature descriptors (e.g., SIFT) and may be represented by image patches. Feature comparison and matching can be performed by calculating the differences (e.g., pixel intensity differences) between image patches using methods such as sum of squared differences (SSD), normalized cross-correlation (NCC), sum of absolute differences (SAD), and interaction information (MI).
[0093] During the following triangulation, the geometric model (3D points) of the real environment and the camera pose are calculated based on feature correspondences.
[0094] Triangulation is the process of determining the location of a feature in 3D space given its projections onto two or more images (image features). See also Figure 6 For example, a 3D point P is projected onto two camera images Ix and Iy by lines Lx and Ly intersecting each camera focus O1 and O2, resulting in image points Px and Py (see...). Figure 6 Therefore, given the focal points O1 and O2 and the corresponding feature points Px and Py detected for the two camera images, lines Lx and Ly can be calculated, and the 3D position of point P can be determined from the intersection of Lx and Ly.
[0095] Models can be created by directly using their intensity or color values, i.e., registering images without using abstract concepts such as point, line, or blob features. Dense reconstruction methods can be built costly, where multiple different hypotheses are tested exhaustively for each pixel. They can also be based on previous sparse reconstructions. Dense methods are typically computationally expensive and run in real time on GPUs.
[0096] In vehicle-based scenarios, images from multiple cameras are available. A common setup includes four vehicle cameras, with each camera positioned away from the vehicle and aligned left, right, front, and back. The vehicle cameras can reference each other during rotation and translation. Bundle adjustment can be used to refine the reconstructed environment model based on images taken by multiple vehicle cameras, especially for multiple cameras with known spatial relationships.
[0097] To generate a geometric model of the environment, one or more cameras mounted on the vehicle can be calibrated or uncalibrated. Camera calibration calculates nonlinear and linear coefficients that map real-world objects with known appearance, geometry, and pose (relative to the camera) onto the image sensor. A calibration procedure is typically performed to calibrate one or more cameras before they are used for 3D reconstruction. Uncalibrated camera images can also be used simultaneously to perform 3D reconstruction. Camera parameters can also be changed during image acquisition (e.g., by zooming or focusing) for use in 3D reconstruction.
[0098] High-quality geometry model from vehicle-based reconstruction:
[0099] In most cases, geometric models of unknown environments created based on environmental data captured by the vehicle's sensors should outperform models created by handheld mobile devices in terms of accuracy and stability. This is because there are more sensors in a vehicle that can be used to cross-check the (intermediate) results of the reconstruction process. For example, correspondences can be verified by superimposing images from two cameras at the same time or by predicting object positions based on odometer readings; specifically, the steering angle of the front wheels and the vehicle's speed can be used to predict how an image of a real object can move from one camera frame (camera image) to another, where the prediction depends on the depth of the real object relative to the camera.
[0100] Furthermore, the motion of a car is more constrained than that of a handheld device. Due to its greater mass and thus stronger inertia (compared to a moving handheld device), it can be well approximated using fewer than six degrees of freedom (i.e., three for translation and three for rotation) and constrained motion. Since vehicles typically move on a 2D ground plane and do not “jump” or “roll,” usually two degrees of freedom are sufficient to simulate translational motion, and one degree of freedom is sufficient to simulate rotational motion. Of course, if necessary, the motion of a vehicle is often simulated using all six degrees of freedom.
[0101] Vision-based tracking (can be performed on mobile devices):
[0102] Standard vision-based tracking methods can be divided into four main components: feature detection, feature description, feature matching, and pose estimation. A known geometric model of the real environment or a portion thereof can support standard vision-based tracking to determine the camera's pose relative to the environment.
[0103] Additionally, optical flow from the camera can be used to calculate camera pose or motion in the environment or to support camera pose estimation.
[0104] Feature detection is also known as feature extraction. Features are, for example, but not limited to, intensity, gradient, edge, line segment, segment, corner, descriptive feature or any other type of feature, primitive, histogram, polarity or orientation, or a combination thereof.
[0105] To determine the camera's pose, the camera must capture the current image in the pose to be determined. First, feature detection is performed to identify features in the current image. Feature description transforms the detected image regions into typical feature descriptors. Feature descriptors are determined to allow feature comparison and matching. An important task is feature matching. Given current features detected in the current image and based on their descriptions, the goal is to find features from a set of provided features that correspond to the same physical 3D or 2D surface. Reference features can be obtained from a reference geometry model of the real environment. The reference geometry model is based on the vehicle's sensors. Reference features can also come from another image captured by the camera (e.g., an image captured by the camera in a pose different from the pose in which the current image was captured), or from a predefined list of features.
[0106] Matching current features with reference features is accomplished by determining the corresponding similarity measure between each corresponding current feature descriptor and each corresponding reference feature descriptor. After feature matching is completed, a correspondence is created between features from the current image and reference features. The correspondence can be 2D-2D or 2D-3D. Based on the correspondence, the camera pose is determined relative to one of the other camera poses, either from the environment or another camera pose.
[0107] Furthermore, a second geometric model of the real-world environment can be generated or extended from the reference geometric model via triangulation based on the feature correspondence between the current image and one of the other images captured by the camera. One of the current image and the other image must have an overlapping portion, which is then reconstructed based on triangulation. The reconstructed portion can then be added to the reference geometric model.
[0108] See now Figure 5 (combination) Figure 2 ), Figure 5A flowchart of a standard camera tracking method for matching a set of current features with a set of reference features is shown. In step 501, a current image CI of the real environment captured by the camera is provided. Then, in the next step 502, features in the current image CI are detected and described (optionally, based on selected extractions of estimated model feature locations), where each resulting current feature c has a feature descriptor d(c) and a 2D location in the camera image. In step 503, a set of reference features r is provided, each reference feature having a descriptor d(r) and optionally having a (local) position and / or orientation relative to the real environment or a previous camera pose. The (local) position can be obtained from GPS sensor / receiver, IR or RFID triangulation, or by using positioning methods with broadband or wireless infrastructure. The (local) orientation can be obtained from sensor devices such as a compass, accelerometer, gyroscope, and / or gravity sensor. The reference features can be extracted from a reference image or geometric model or other information about the real environment or a portion thereof. Note that in the case of visual search and classification tasks, the position and / or orientation relative to the real environment is optional. In step 504, the current feature c from step 502 and the reference feature r from step 503 are matched. For example, for each current feature, a reference feature is searched that has a descriptor that is closest to the current feature relative to a certain distance measurement. According to step 505, the camera's position and orientation can be determined based on feature matching (correspondence). This can support augmented reality applications that incorporate spatially registered virtual 3D objects into camera images.
[0109] Data switching from vehicle to mobile device (can be performed on both mobile devices and vehicles):
[0110] If environmental data captured by the vehicle's sensors while the vehicle is traveling through the environment is available, the environmental data can be transmitted to the mobile device of a user preparing to disembark (e.g., a passenger or driver). Transmission of data from the vehicle to the mobile device can be via a server computer or based on point-to-point communication. A geometric model of the environment can be reconstructed based on the environmental data received on the mobile device. The user can then use the geometric model for tracking.
[0111] It is also possible to generate a geometric model in the vehicle based on environmental data, and then transmit it to a mobile device via a server computer or based on point-to-point communication.
[0112] It is also possible to generate geometric models on a server computer based on environmental data. In this case, the captured environmental data is transferred from the vehicle to the server computer, and then the generated geometric model is transferred from the server computer to the mobile device.
[0113] Point-to-point communication can be via wireless or wired connections, push or pull based, unicast or broadcast. The latter allows a pool of mobile devices to simultaneously carry the same data from the vehicle. These mobile devices could be, for example, other passengers' mobile devices or users' assistive devices. For example, point-to-point communication could be Bluetooth-based or USB cable-based.
[0114] The transfer of data (geometric model or environmental data) from a vehicle or to a mobile device can be triggered manually or automatically by the user. Automatic triggering of data (e.g., model) transfer can be based on the distance to a destination known to the navigation system, the vehicle's speed, engine status, the vehicle's direction of travel (e.g., reversing), the vehicle's relative direction to the street (e.g., approaching a one-way parking space or lane), the distance of another object to the front or rear of the vehicle, the open / closed status of the driver's side door, the steering wheel lock, the handbrake, the open / closed status of the trunk, or a combination thereof. Automatic triggering can also be based on the status of the mobile device, for example, when it is removed from a wired connection or when an upward movement incompatible with the vehicle's normal movement is detected (e.g., when the driver exits the vehicle). Detection or determination of the conditions used to trigger data (e.g., model) transfer can be performed by the vehicle, the mobile device, or both. The data (model) transfer process can be initiated by the vehicle, the mobile device, or both.
[0115] In the described application example, one aspect of the invention is that tracking mobile devices in a parking lot can be supported by environmental data collected by the vehicle's sensors during the vehicle's journey to the parking lot.
[0116] like Figure 4 As shown, user 413 is driving vehicle 411 towards a parking lot, i.e., a real-world environment. Camera 414 (one or more sensors of the vehicle) is mounted at the front of the vehicle. While searching for and parking a parking space (see...),... Figure 4 Description 402), environmental data (e.g., an image of a parking lot) is captured by camera 414, and a geometric model of the parking lot is generated based on the environmental data. The geometric model can be created using an appropriate metric, such as based on odometer and GPS data provided by the vehicle. An appropriate scaling factor can define the realistic camera pose and the size of the reconstructed geometric model in the real world.
[0117] After parking (see) Figure 4 (Description 405) The geometric model is transmitted to the user's camera-equipped mobile device 408 (e.g., a smartphone) (or generated at the user's mobile device 408 from the transmitted environmental data). Tracking the mobile device in the parking lot can then ideally continue seamlessly based on at least one image captured by the camera of the mobile device 408 (see description 405). Figure 4(Description 405). Compared to existing methods, this invention provides the user with both an initial and an updated geometric model of the environment on their mobile device. This eliminates the need for an initial step requiring constrained camera movement.
[0118] Currently, geometric models of real-world environments, such as cities or buildings, are typically obtained through 3D reconstruction processes. However, due to the frequent development and changes in urban construction, most of these models are not up-to-date. Specifically, parking lots often lack geometric models or up-to-date models because parked vehicles change over time. In this regard, the present invention provides up-to-date models that will support more accurate tracking cameras.
[0119] A vehicle can be a mobile machine that can transport people or goods. A vehicle can be, for example, but not limited to, a bicycle, motorcycle, car, truck, or forklift. A vehicle may or may not have an engine.
[0120] The acquisition of environmental data for creating the first geometric model can begin at any time or only when certain conditions are met, such as when the vehicle approaches a destination known to the navigation system and / or when the vehicle's speed is below a certain threshold. A condition can also be one of several vehicle states, such as speed, engine status, braking system status, gear position, lighting conditions, etc. A condition can also be one of the states of a mobile device, such as whether the mobile device is inside or outside the vehicle, or the distance of the mobile device from the destination.
[0121] Environmental data can be collected based on the vehicle's location and a predefined route that includes at least one destination.
[0122] In one embodiment, environmental data is collected or at least a portion of the environmental data collection process is initiated after the vehicle reaches at least one destination. In another embodiment, environmental data is collected or at least a portion of the environmental data collection process is initiated while the vehicle is within a distance of at least one destination. Given the vehicle's speed, position, and at least one destination, the time period during which the vehicle will reach at least one destination can be estimated. In another embodiment, if the time period is less than a threshold, environmental data is collected or at least a portion of the environmental data collection process is initiated.
[0123] The generation of a first geometric model or a portion thereof may be performed at a processing device in the vehicle, and the first geometric model may be transferred from the vehicle to the mobile device. The first geometric model may be transferred from the vehicle to the mobile device via a server computer or based on point-to-point communication. The first geometric model or a portion thereof may also be generated by one or more processing devices of the mobile device after environmental data has been transferred from the vehicle to the mobile device via a server computer or based on point-to-point communication.
[0124] The generation of a first geometric model, or a portion thereof, can be performed by a server computer, and environmental data is transmitted from the vehicle to the server computer. The first geometric model is then transmitted from the server computer to a mobile device, for example, via the vehicle or based on point-to-point communication.
[0125] One or more sensors of the vehicle can be any type of sensor capable of capturing environmental data that can be used to generate a geometric model or a portion thereof. Examples of sensors can be, but are not limited to, optical cameras, cameras based on infrared spectroscopy, optical spectroscopy, ultraviolet spectroscopy, X-ray spectroscopy and / or gamma-ray spectroscopy, RGB-D cameras, depth sensors, time-of-pass cameras, and ultrasonic sensors.
[0126] Environmental data can be any type of data describing at least one visual or geometric feature of a real-world environment. Visual or geometric features can be one of the following: shape, color, distance, etc. Environmental data can be an optical image of the environment, the distance between the environment and the vehicle, the vehicle's orientation within the environment, or the vehicle's speed in the real-world environment.
[0127] The generation of the first geometric model can be performed whenever environmental data or a portion of environmental data is available, such as online during environmental data acquisition or offline after environmental data acquisition.
[0128] The first geometric model can have a suitable metric scale, which is determined based on sensors mounted on the vehicle, such as radar, distance sensors, or time-of-flight cameras. A suitable metric scale can be determined by using cameras mounted on the vehicle to capture images of two points in the environment that are a known distance apart, or images of real objects with known physical dimensions.
[0129] The first geometric model can also be generated based on the vehicle's pose relative to the environment. The vehicle's position relative to the environment can be obtained from GPS. However, GPS is not accurate, especially when the vehicle is inside a building. There are many different sensors, similar to cameras (e.g., security cameras), that can be mounted in the environment and have known locations within it. Object recognition or pose estimation can be performed based on vehicle images captured by security cameras to determine the vehicle's pose relative to the environment.
[0130] In addition, a vision system (e.g., one or more security cameras) positioned in a real environment can be used to capture environmental data that can be used to create at least a portion of the first geometric model.
[0131] In another specific implementation, a vision system (e.g., a security camera) positioned in a real-world environment can also be used to capture environmental data for creating a first geometric model, without using data acquired by at least one sensor of the vehicle. For example, a user may carry a mobile device containing a camera and enter a real-world environment equipped with one or more security cameras. Up-to-date environmental data of the real-world environment can be captured by the vision system (e.g., one or more security cameras) during the acquisition process, and the first geometric model can be created based on the captured environmental data. The mobile device can be tracked based on the created first geometric model and images captured by the mobile device's camera within a set time period (preferably within 24 hours) after the acquisition process or a portion thereof.
[0132] The vehicle's attitude relative to its environment can also be obtained based on the specific characteristics of vehicles traveling on the road. For example, it can be assumed that the vehicle's movement is parallel to the ground plane. The vehicle's orientation can be determined based on a 2D street map.
[0133] In one implementation, environmental data of only the locations a user is likely to pass through along their route can be collected and added to the first geometric model.
[0134] In another implementation, environmental data may be collected based on navigation data such as route, starting point, or destination. For example, environmental data may be collected only along the route, so the first geometric model can be based solely on environmental information along the route. Navigation data may be manually entered into the vehicle and / or device by the user.
[0135] In one implementation, the generation of the first geometric model may be based at least in part on vehicle state data acquired by at least one sensor of the vehicle during the acquisition process. For example, at least a portion of the environmental data used to generate the first geometric model may be acquired by a separate sensor that is not part of the vehicle. At least a portion of the acquired environmental data may be used in conjunction with vehicle state data acquired by one or more sensors of the vehicle to create at least a portion of the first geometric model. For example, a vehicle-independent camera held by a passenger sitting in the vehicle may capture images of the real environment, while the vehicle's odometer and / or speed sensor may be used to acquire mileage or speed data about the vehicle. Images of the real environment and mileage or speed data may be used together to create at least a portion of the first geometric model. This takes advantage of the passenger-held camera having a known movement or position relative to the vehicle. Movement may be that the camera is stationary (no motion) relative to the vehicle. The camera may also have movement relative to the vehicle. The camera may be tracked in the vehicle's coordinate system. For example, an image capture device may be mounted in the vehicle and determine the camera's orientation relative to the vehicle. In another example, images of at least a portion of the vehicle captured by the camera may be used to determine the camera's orientation relative to the vehicle.
[0136] In another embodiment, data acquired by at least one sensor of the vehicle may be insufficient to create at least a portion of the first geometric model. Similarly, data acquired by a camera of a mobile device may also be insufficient to create at least a portion of the first geometric model. However, at least a portion of the first geometric model can be created using data acquired by both at least one sensor of the vehicle and a camera of the mobile device. For example, the first geometric model may not be possible to create using only one image captured by the vehicle's camera or only one image captured by the mobile device's camera. However, the first geometric model can be created using images captured by both the vehicle's camera and the mobile device's camera.
[0137] Additionally, it may be impossible to create at least a portion of a first geometric model with a suitable metric scale using data acquired by at least one sensor of the vehicle or data acquired by a camera of a mobile device. However, it may be possible to create at least a portion of a first geometric model with a suitable metric scale using data acquired by both at least one sensor of the vehicle and a camera of a mobile device simultaneously. As in the example above, it may be impossible to create a first geometric model using speed or mileage captured by the vehicle's sensors. It may be impossible to create a first geometric model with a suitable metric scale using only images captured by a camera of a mobile device. However, it may be possible to create a first geometric model with a suitable metric scale using images captured by a camera of a mobile device and mileage or speed data related to the vehicle.
[0138] The transfer of the first geometric model to the mobile device can be triggered manually or automatically by the user. Automatic triggering can be based on the distance to the known destination of the navigation system, the vehicle's speed, the engine status, the vehicle's direction (e.g., moving backward), the vehicle's relative direction to the street (moving to a one-way parking space or lane), the distance of another object to the front or rear of the vehicle, the open / closed status of the driver's side door, the steering wheel lock, the handbrake, the open / closed status of the trunk, or a combination of the above.
[0139] This invention is particularly beneficial for mobile AR and navigation applications running on devices.
[0140] The mobile device according to the invention can come from a variety of devices that a user may carry, such as handheld mobile devices (e.g., mobile phones), head-mounted glasses or helmets, and wearable devices.
[0141] Tracking a mobile device equipped with at least one camera in an environment is for the purpose of determining the device's pose, i.e., its position and orientation relative to the environment, or for determining the device's motion, i.e., its position and orientation relative to another position of the device. Since the camera and the mobile device have a fixed spatial relationship, tracking the mobile device can be achieved by vision-based tracking, for example, using a first geometric model and images captured by the at least one camera to determine the camera's pose. The camera's pose relative to the environment can be calculated based on the correspondence between the geometric model and the camera images.
[0142] A second geometric model of the environment can be created using the first geometric model and at least two images captured by at least one camera that do not have depth data, or using at least one image that has depth data. To do this, the camera pose when capturing at least two images or at least one image can first be determined. Then, the second geometric model can be constructed based on triangulation using at least two images or at least one image and associated depth data.
[0143] A captured camera image without depth data is used to generate model information about the environment with undetermined metric scales. This model information with undetermined metric scales can be used to estimate camera pose when the camera undergoes pure rotation.
[0144] The first geometric model, or a portion thereof, may not cover the entire area of the real-world environment of interest. A second geometric model can be created by extending the first geometric model, or a portion thereof, to cover a larger area of the environment.
[0145] Standard methods for vision-based tracking include feature extraction, feature description, feature matching, and pose determination.
[0146] Features can be, for example, intensity, gradient, edge, line segment, section, corner, descriptive feature or any other type of feature, primitive, histogram, polarity or orientation.
[0147] Tracking mobile devices and / or generating a second geometric model can also be achieved using monocular vision-based simultaneous localization and mapping (SLAM). Generating a second geometric model can also include reconstruction algorithms that run at different times but use batch / quasi-offline reconstruction methods.
[0148] Monocular vision-based SLAM involves moving a single camera within a real-world environment to determine the camera's pose relative to the environment and creating a model of that environment. SLAM is typically used for tracking when at least a portion of the environment has an unknown geometry.
[0149] The camera's pose in a real-world environment describes its position and orientation relative to the environment or a portion of the environment. In 3D space, position is defined by three parameters, such as displacement along three orthogonal axes, and orientation is defined by three Euler angle parameters. Orientation can also be expressed in other mathematical formulas, such as axis angles and quaternions. Mathematical representations of rotation are generally interchangeable. The pose in this invention can be defined by at least one of six inherent parameters of position and orientation in three-dimensional space.
[0150] The proposed invention can be easily applied to any camera installed on a mobile device that provides image formats (color or grayscale). It is not limited to capture systems for color images in RGB format. It is also applicable to any other color format and to monochrome images, such as those provided by cameras in grayscale format.
[0151] A real-world environment can be any real-world scene, such as a natural scene, an indoor environment, or an urban scene. A real-world environment includes one or more real-world objects. Real-world objects, such as locations, buildings, trees, or mountains, are located in and occupy an area within the real-world environment.
[0152] A geometric model of the environment (also referred to as a model or map) describes at least the depth information of the environment. The model may also include, but is not limited to, at least one of the following attributes: shape, symmetry, planarity, geometric dimensions, color, texture, and density. A geometric model may include a variety of features.
[0153] The geometric model may also include information about the texture, color, and / or combinations thereof (i.e., materials) of a portion of the real environment. A common combination of representations of the first model is used to provide a sparse spatial description of the geometry of 3D points. The geometric model may also have associated feature descriptors that describe the texture (as part of the material) of features in an image patch surrounding the 3D points. Feature descriptors are mathematical representations describing local features in an image or image patch, such as SIFT (Scale Invariant Feature Transform), SURF (Fast Stable Feature Transform), and LESH (Local Energy Based Shape Histogram). Features are, for example, but not limited to, intensity, gradient, edges, line segments, sections, corners, descriptive features, or any other type of feature, primitive, histogram, polarity, or orientation.
[0154] The geometric model can be further represented as a model including 3D vertices and polygonal faces and / or edges extended from these vertices. The edges and faces of the model can also be represented as splines or NURBS surfaces. The geometric model can also be represented by a set of 3D points. These points can carry additional information about their color or intensity.
[0155] Throughout this document, it describes the capture of images and the provision or receipt of image information associated with those images. Those skilled in the art will recognize that this may include providing or receiving any processed or unprocessed information (versions) of an image, a portion of an image, and / or features of an image, which allows for pose estimation (tracking) or reconstruction. This invention does not require the provision or receipt of any unprocessed initial image data. Therefore, processing includes compression (e.g., JPEG, PNG, ZIP), encryption (e.g., RSA encryption, Schnorr signature, El-Gamal encryption, PGP), conversion to another color space or grayscale, cropping or scaling the image based on feature descriptors or converting it to a sparse representation, extraction, any of these, and combinations thereof. All these image processing methods may be performed optionally and are encompassed by the terminology of image information associated with an image.
Claims
1. A method for initializing an augmented reality system, comprising: Environmental data of the real environment captured by one or more sensors of a vehicle is obtained via a mobile device, wherein the environmental data is obtained when the mobile device is being transmitted by the vehicle, and wherein the environmental data is obtained based on determining that the mobile device is moving inconsistently with the vehicle. A reference geometric model of at least a portion of the real environment is generated based on the environmental data, wherein the reference geometric model is generated using the motion of the vehicle, which is more constrained than the motion of the mobile device. The attitude of the mobile device is determined based on the reference geometric model, and The mobile device is tracked via a vision-based tracking process using the determined pose and a reference geometric model, wherein the mobile device is tracked based on additional environmental data received by one or more sensors of the mobile device.
2. The method according to claim 1, wherein, The reference geometric model is generated based on environmental data from multiple sensors and vehicle state information associated with the environmental data.
3. The method according to claim 2, wherein, The vehicle status information includes at least one of the vehicle's speed or the vehicle's direction of steering.
4. The method according to claim 1, wherein, The vision-based tracking process is based on simultaneous localization and mapping ("SLAM").
5. The method according to claim 1, wherein, The reference geometric model is generated based on environmental data from multiple sensors and the predetermined spatial relationships between the multiple sensors.
6. The method according to claim 1, wherein, The vision-based tracking process is initialized as part of an augmented reality application on the mobile device.
7. A non-transitory computer-readable medium comprising computer-readable code, said computer-readable code being executable by one or more processors to perform the following operations: Environmental data of the real environment captured by one or more sensors of a vehicle is obtained via a mobile device, wherein the environmental data is obtained when the mobile device is being transmitted by the vehicle, and wherein the environmental data is obtained based on determining that the mobile device is moving inconsistently with the vehicle. A reference geometric model of at least a portion of the real environment is generated based on the environmental data, wherein the reference geometric model is generated using the motion of the vehicle, which is more constrained than the motion of the mobile device. The attitude of the mobile device is determined based on the reference geometric model, and The mobile device is tracked via a vision-based tracking process using the determined pose and a reference geometric model, wherein the mobile device is tracked based on additional environmental data received by one or more sensors of the mobile device.
8. The non-transitory computer-readable medium according to claim 7, wherein, The reference geometric model is generated based on environmental data from multiple sensors and vehicle state information associated with the environmental data.
9. The non-transitory computer-readable medium according to claim 8, wherein, The vehicle status information includes at least one of the vehicle's speed or the vehicle's direction of steering.
10. The non-transitory computer-readable medium according to claim 7, wherein, The vision-based tracking process is based on simultaneous localization and mapping ("SLAM").
11. The non-transitory computer-readable medium according to claim 7, wherein, The reference geometric model is generated based on environmental data from multiple sensors and the predetermined spatial relationships between the multiple sensors.
12. The non-transitory computer-readable medium according to claim 7, wherein, The vision-based tracking process is initialized as part of an augmented reality application on the mobile device.
13. A system for vision-based tracking, comprising: One or more processors; and One or more non-transitory computer-readable media, including computer-readable code that can be executed by one or more processors to perform the following operations: Environmental data of the real environment captured by one or more sensors of a vehicle is obtained via a mobile device, wherein the environmental data is obtained when the mobile device is being transmitted by the vehicle, and wherein the environmental data is obtained based on determining that the mobile device is moving inconsistently with the vehicle. A reference geometric model of at least a portion of the real environment is generated based on the environmental data, wherein the reference geometric model is generated using the motion of the vehicle, which is more constrained than the motion of the mobile device. The attitude of the mobile device is determined based on the reference geometric model, and The mobile device is tracked via a vision-based tracking process using the determined pose and a reference geometric model, wherein the system for vision-based tracking tracks the mobile device based on additional environmental data received by one or more sensors of the mobile device.
14. The system according to claim 13, wherein, The reference geometric model is generated based on environmental data from multiple sensors and vehicle state information associated with the environmental data.
15. The system according to claim 14, wherein, The vehicle status information includes at least one of the vehicle's speed or the vehicle's direction of steering.
16. The system according to claim 15, wherein, The vision-based tracking process is based on simultaneous localization and mapping ("SLAM").
17. The system according to claim 14, wherein, The reference geometric model is generated based on environmental data from multiple sensors and the predetermined spatial relationships between the multiple sensors.