Vision-based position and turn marker prediction

By combining data from vehicle-mounted cameras, GNSS, and IMU, and using artificial neural networks to predict road signs and turning points, the problem of accurate positioning and orientation of navigation systems in harsh environments has been solved, achieving real-time and accurate navigation prediction.

CN115917255BActive Publication Date: 2025-11-21HARMAN INT IND INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080102752.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-11-21
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

Existing vehicle navigation systems struggle to accurately determine vehicle location and direction of travel in tunnels, urban canyons, or in adverse weather conditions, especially when stationary. Furthermore, noise and drift exist in the IMU and dead reckoning processes, resulting in insufficient augmented reality capabilities.

Method used

By combining vehicle-mounted cameras, GNSS, IMU, and artificial neural networks, the vehicle's position and heading vector are determined in real time by calculating the angle changes between images and fusing multiple measurements. The trained neural network is then used to predict road signs and turning points.

Benefits of technology

It enables real-time and accurate determination of vehicle position and heading in various environments, accurately predicts road signs and turning points, and enhances the augmented reality capabilities of the navigation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115917255B_ABST
    Figure CN115917255B_ABST
Patent Text Reader

Abstract

A computer-implemented method and apparatus for accurately determining a vehicle position and predicting in real-time navigation key points. The method includes instructing a vehicle-mounted camera to capture a plurality of images; calculating an angular change between a first image and a subsequent second image; determining a vehicle position and heading vector from the calculated angular change; collecting at least one image of the plurality of images as a first training subset; obtaining image-related coordinates of a navigation key point related to the at least one image in the first training data subset as a second training subset; providing the first training data subset and the second training data subset to an artificial neural network as a training data set; and training the artificial neural network on the training data set to predict image-related coordinates of a navigation key point indicative of a road sign position and / or a turn point.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present disclosure relates to a computer-implemented method and apparatus for determining a vehicle position and predicting turn points and lane changes. BACKGROUND

[0002] Global Navigation Satellite Systems (GNSS) are often used in navigation systems of vehicles to determine the position of the vehicle. Typically, this is achieved by receiving signals from at least three satellites on the navigation system of the vehicle and calculating the position of the vehicle using the received signals through trilateration. Generally, it is possible to accurately determine the position of the vehicle from the received signals. However, this relies on an unobstructed path between the navigation system and the satellites. Thus, when the vehicle is in a tunnel / canyon or under adverse weather, it can be difficult for the navigation system to receive accurate signals from the satellites, thereby impairing its ability to accurately determine the position of the vehicle. To address this issue, current navigation systems include a dead reckoning process which calculates the position of the vehicle by using a previously determined position (by GNSS) and estimating the current position based on an estimate of the vehicle’s speed and the time elapsed since the previously determined position.

[0003] It has proven difficult to identify the precise heading and orientation of a vehicle using conventional GNSS devices, especially when the vehicle is stationary (e.g. at a red light). To address this issue, current navigation systems are equipped with an Inertial Measurement Unit (IMU) which measures and reports the specific force, angular rate and orientation of the vehicle using a combination of accelerometers, gyroscopes, magnetometers (sometimes).

[0004] Thus, current navigation systems can measure the position of the vehicle using GNSS trilateration and measure the orientation and heading of the vehicle using an IMU. Furthermore, current navigation systems can estimate the position of the vehicle when it is difficult for the navigation system to receive accurate signals from the satellites.

[0005] Despite being sufficiently accurate for basic navigation, current navigation systems still struggle to accurately determine the position, heading and orientation of the vehicle in real-time. For example, GNSS has an update rate in the range of 0.1 to 1.0 Hz which results in a slight delay in determining the position of the vehicle. Furthermore, IMUs are fast and have a refresh rate in the kHz range but are often subject to noise and drift. Still further, dead reckoning processes are also often subject to noise and experience drift, especially during longer signal outages. Thus, current navigation systems are not sufficiently accurate for navigation systems with augmented reality capabilities. Therefore, there is a need to accurately determine the position, heading and orientation of the vehicle in real-time in order to implement augmented reality capabilities, such as identifying road signs, intersections and lane changes, into current navigation systems. SUMMARY

[0006] To overcome the problems detailed above, the inventors have devised novel and inventive vehicle positioning apparatus and techniques.

[0007] More specifically, the present disclosure provides a computer-implemented method for determining a vehicle position and predicting navigation key points. A logic circuit within a vehicle navigation system can instruct a vehicle-mounted camera coupled to the logic to capture a plurality of images and transmit the images back to the logic circuit. The logic circuit can receive the plurality of images, wherein the plurality of images includes at least a first image captured at a first time interval and a second image captured at a second time interval. Upon receiving at least the first image and the second image, the logic circuit can analyze the images and calculate an angular change between the first image and the second image. With the calculated angular change, the logic circuit can accurately determine a position and a heading vector of the vehicle in real time.

[0008] Additionally, the computer-implemented method can receive GPS coordinates from a global navigation satellite system (GNSS) module coupled to the logic, measurements from an inertial measurement unit (IMU) coupled to the logic, and a speed of the vehicle to determine a confidence value of the determined position and heading vector of the vehicle based on the received GPS coordinates, the IMU measurements, and the speed of the vehicle. For example, if the logic determines that the position and heading vector determined from the images of the vehicle-mounted camera are significantly different from the GPS, IMU, and speed measurements, the logic can instruct the vehicle-mounted camera to capture a plurality of second images at a higher frequency such that a time interval between captured images of the plurality of second images is shorter than a time interval between captured images of the plurality of first images. The logic can then restart the process of calculating an angular change between two images of the plurality of second images and determine the position and heading vector of the vehicle in real time based on the calculated angular change.

[0009] In one embodiment, the method can collect at least one of any of the plurality of images as a first training data subset; obtain image-related coordinates of a navigation key point related to the at least one image of the first training data subset as a second training data subset; and provide the first training data subset and the second training data subset to an artificial neural network as a training data set. The method can then train the artificial neural network on the training data set to predict image-related coordinates of a navigation key point indicative of a road sign position and / or a turning point, and process an input data set through the artificial neural network to predict image-related coordinates of a navigation key point indicative of a road sign position and / or a turning point. This allows for converting depth information and upcoming turning points / intersections / lanes changes so that they can be displayed on a screen of a navigation system, or on a head-up display (HUD) to indicate to a user that a turning point / intersection / lanes change is upcoming.

[0010] In some examples, a second vehicle mounted camera can be coupled to the logic and the computer-implemented method can be operable to instruct the second vehicle mounted camera to capture a further plurality of images and transmit those images to the logic to more accurately determine the location and heading vector of the vehicle in real time. In certain examples, the first vehicle mounted camera and / or the second vehicle mounted camera can be a front-facing camera (FFC) that is used to capture images of the field of view directly in front of the vehicle.

[0011] The present disclosure also relates to an apparatus for determining a vehicle location and predicting navigation key points, wherein the apparatus comprises means for performing the methods of the present disclosure. For example, the apparatus comprises logic and a first vehicle mounted camera (FFC), wherein the logic can be used to instruct the vehicle mounted camera (FFC) to capture a plurality of images and the logic can calculate an angular change between a first image and a second image of the plurality of images to determine a location and heading vector of the vehicle in real time. Further, the logic can be operable to collect at least one image of the plurality of first images as a first training data subset; obtain image-related coordinates of a navigation key point related to the at least one image of the first training data subset as a second training data subset; provide the first training data subset and the second training data subset to an artificial neural network as a training data set; train the artificial neural network on the training data set to predict image-related coordinates of a navigation key point indicative of a road sign location and / or a turning point; and process an input data set through the artificial neural network to predict image-related coordinates of a navigation key point indicative of a road sign location and / or a turning point. BRIEF DESCRIPTION OF DRAWINGS

[0012] The embodiments are described by way of example with reference to the accompanying drawings, wherein the drawings are not drawn to scale and:

[0013] Figure 1 An overall system architecture for vision-based positioning with augmented reality capabilities is shown;

[0014] Figure 2 An apparatus for determining a vehicle location is shown;

[0015] Figure 3 An example of a deep neural network employed in the present disclosure is shown;

[0016] Figure 4 An example of a predicted key point marker location indicated in environmental data is shown;

[0017] Figure 5 Another example of a predicted key point marker location indicated in environmental data is shown;

[0018] Figure 6 An overlay of an exemplary image with an exemplary key marker generated by an artificial neural network is shown;

[0019] Figure 7 A series of images during processing by a convolutional neural network are shown;

[0020] Figure 8 An artificial neural network with HRNet architecture features is shown;

[0021] Figure 9 An exemplary image of a camera and a corresponding standard definition overhead map are shown; and

[0022] Figure 10 A flowchart of an embodiment of a method for determining a vehicle position is shown. DETAILED DESCRIPTION

[0023] Figure 1 An exemplary overall system architecture for vision-based localization with augmented reality capability is shown. The system architecture can be part of a vehicle on-board navigation system, where the navigation system can predict navigation key points indicative of road sign locations and / or turn points. Some examples of turn points include, but are not limited to, lane changes and intersections. The navigation system can be implemented in any type of road vehicle, such as a car, truck, bus, etc.

[0024] The system can include an automotive localization block 102 that can compute the position of the vehicle based on normative Global Navigation Satellite System (GNSS) and / or Inertial Measurement Unit (IMU) sensors. Additionally, the automotive localization block 102 can derive the position and heading vector of the vehicle from a Visual Odometry (VO). VO is a real-time computer vision algorithm that computes the angular change between one front-facing camera (FFC) frame and the next. When this is added to other signals from the GNSS subsystem, IMU, and vehicle speed, the position and heading vector of the vehicle can be more accurately determined. The automotive localization block 102 will be described in more detail in Figure 2

[0025] The position and heading vector can then be used by an aerial image landmark database 112 (offline database) to obtain suitable landmark locations that have been extracted offline by an artificial neural network, where the landmark locations correspond to navigation key points indicative of road sign locations and / or turn points. The aerial image landmark database 112 can be constructed by scanning aerial images of the satellite and / or the earth's surface. The database 112 can also include a landmark selector (not shown) that uses the current automotive's position and route information to select landmarks (e.g., road sign locations and / or turn points) that are relevant to the vehicle's current route. The construction of the offline aerial image landmark database 112 will be discussed in more detail in Figures 3 to 5

[0026] ​​The scene understanding 104 block can also use the position and heading vector to better determine the sign location within the FFC frame. Scene understanding 104 is a neural network that can run in real-time to place signs (e.g., road sign locations and / or turn points) relative to the scene from the vehicle front-facing camera (FFC) or similar camera using semantic image analysis. Scene understanding 104 can scan the FFC frame and use scene analysis and other information sources (e.g., standard definition maps) to determine the screen coordinates of the sign location as well as depth information and / or ground plane estimation of the sign location. Scene understanding 104 will be discussed in more detail in Figures 6 to 9

[0027] Scene understanding 104 and the sign locations extracted offline from the aerial image sign database 112 can be used individually or together to calculate the actual sign location from different sign location sources and predictions. Thus, the system architecture (i.e., the vehicle’s navigation system) can use any of the predicted sign locations from the aerial image sign database 112, the vehicle FFC frame from the scene understanding 104, or a combination of both to accurately predict upcoming road signs and / or turn points.

[0028] The sign localization 106 block can use information and normalized confidence scores from the aerial image sign database 112 and the scene understanding 104 components to determine where to place the signs based on the vehicle’s position and heading vector. The sign localization 106 block can implement smoothing techniques to avoid jittering from changing which source to use for each sign of a given route. The sign rendering engine 114 can work in sync with the scene understanding 104 and / or the aerial image sign database 112 to convert the geographic coordinates of the signs to screen coordinates so that the signs can be displayed on the vehicle’s navigation screen and visible to the user. The route planning 108, 110 can then determine the waypoints and maneuvers that the vehicle is expected to follow from the start point to the desired destination.

[0029] Figure 2 An overview of the system architecture is shown in FIG. 1. The aerial image sign database 112 can be used to store and retrieve sign locations from aerial images. The aerial image sign database 112 can be used to store and retrieve sign locations from aerial images. The aerial image sign database 112 can be used to store and retrieve sign locations from aerial images. The aerial image sign database 112 can be used to store and retrieve sign locations from aerial images. Figure 1 ​A more detailed example of an automotive localization block 102, which can be a device for determining a vehicle's position, such as part of a vehicle navigation system. The device can include a first on-board camera, such as a front-facing camera (FFC) 202, that captures a plurality of forward (relative to the vehicle) images. The FFC 202 can be coupled to a visual odometry (VO) unit 204 that can instruct the FFC 202 to capture the plurality of forward images and subsequently receive the plurality of first images from the FFC 202. The plurality of images can be taken in sequence, with an equal time period between each image. Upon receiving the plurality of images, the VO unit 204 can calculate the angular change between one image and the next image (one frame to the next). The VO unit 204 can then determine, in real-time, an updated version of the vehicle's position and heading vector based on the determined angular change and transmit the updated version of the vehicle's position and heading vector to a fusion engine 210.

[0030] The fusion engine 210 can receive additional measurements, such as GPS coordinates from a global navigation satellite system (GNSS) module 208, IMU measurements from an inertial measurement unit (IMU) module 206, and a speed of the vehicle. The GNSS module 208, the IMU module 206, and the VO unit 204 can each be coupled to the fusion engine 210 by physical or wireless means, such that they can transmit signals to one another. Additionally, the FFC 202 (and any other cameras) can be coupled to the VO unit 204 by physical or wireless means. The speed of the vehicle can be recorded from a speedometer of the vehicle or from an external source, such as by a calculation from the GNSS module 208. Upon receiving these additional measurements, the fusion engine 210 can compare the vehicle's position and heading vector calculated from the images captured by the FFC 202 to the measurements received from the GNSS module 208, the IMU module 206, and the speed of the vehicle to determine a confidence value for the calculated vehicle's position and heading vector. If it is determined that the confidence value is below a predetermined threshold, the fusion engine 210 can instruct the FFC 202 (through the VO unit 204) to increase the frequency at which images are captured, such that the FFC 202 captures an additional plurality of consecutive images with a shorter time period between each image. As described above, the VO unit 204 can receive the additional plurality of images taken at the increased frequency and can calculate the angular change between one image and the next image (one frame to the next) to determine an updated version of the vehicle's position and heading vector in real-time to transmit to the fusion engine 210. If it is still determined that the confidence value is below the predetermined threshold, the fusion engine 210 can repeat the method.

[0031] By using a monocular vision algorithm, the VO unit 204 can compute sufficient angles to overcome drift and noise of a typical IMU. Additionally, stereo vision can be implemented including a first vehicle camera and a second vehicle camera (e.g., two FFCs) allowing the VO unit 204 to compute even more precise angles and provide depth information. It should be noted that any number of vehicle cameras or FFCs can be used to provide more accurate results.

[0032] Figure 3 An example of a deep neural network (a type of artificial neural network) of the present disclosure is shown. The deep neural network can be a convolutional neural network 302, which can be trained and stored in the device 300 of the present disclosure after training. The convolutional neural network 302 can include a plurality of convolutional blocks 304, a plurality of deconvolutional blocks 306, and an output layer (not shown). Each block can include several layers. During training, a training data set is provided to a first convolutional block of the convolutional blocks 304. During inference, an input data set (i.e., a defined region of interest) is provided to the first convolutional block of the convolutional blocks 304. The convolutional blocks 304 and the deconvolutional blocks 306 can be two-dimensional. The deconvolutional blocks 306 after the output layer can transform the final output data of the convolutional blocks 304 into an output data set (an output prediction) that is subsequently output by the output layer. The output data set includes predicted key marker locations, i.e., predicted virtual road sign locations. The output data set of the convolutional neural network 302 can be given by a pixel map of possible intersection points in a predetermined region (training phase) or a defined region of interest (inference phase), where a probability value (a probability score) is associated with each pixel. Those pixels with high probability scores (i.e., exceeding a predefined threshold (e.g., 90% (0.9))) are subsequently identified as predicted key marker locations.

[0033] Figure 4 An example of predicted key marker locations 402 in a region of interest 400 that have been predicted by the deep neural network of the present disclosure is shown, which is provided as input data to the trained deep neural network during inference. As Figure 3 As shown on the right, the predicted key marker locations 402 represent turning point marker locations at a crossroad, such as an intersection or a T-junction, and can be located in the center of a lane or road leading into the crossroad.

[0034] The key marker locations can also be predicted, such as by the trained deep neural network, such that the key marker is located on a curved path connecting two adjacent potential key marker locations at the center of a road / lane. In this case, the key marker location (i.e., the visual road sign location) can be chosen, such as on the curved path, such that the key marker / virtual road sign is more intuitive / more visible / more discernible to a driver, for example, who is not obscured by a building but is located before the building. Figure 5An example is shown in which two adjacent potential key point marker locations 502 and 504 are connected by a curved path 506, and then a key point marker M (i.e., a virtual road sign) is placed on the curved path such that it can be better or more easily perceived by a driver than if the key point marker were placed at location 502 or 504.

[0035] Figure 6 An example overlay image 600 is shown with an example key marker generated by an artificial neural network. In this example, the FFC 202 can capture an image when the vehicle approaches an intersection. Figure 2 The fusion engine 210 can collect this image (or multiple captured images) as a first training data subset for the artificial neural network, and can further obtain image-related coordinates of navigation key points / markers related to the image (or multiple images) in the first training data subset as a second training data subset. In addition, the fusion engine 210 can provide both the first training data subset and the second data training subset to the artificial neural network as a training data set. Thus, the fusion engine 210 can train the artificial neural network on the training data set to predict image-related coordinates of navigation key points / markers that indicate road signs and / or turning points, and can process an input data set through the artificial neural network to predict image-related coordinates of navigation key points / markers that indicate road sign locations and / or turning points.

[0036] Overlayed on the image 600 are contour maps at two locations 602, 604 that indicate areas of high probability (white with black lines around) and very high probability (black fill) that a navigation key point is included therein. These contours are determined by the neural network and post-processed so that, for example, a maximum or individual centroid locations can be determined.

[0037] Figure 7 An example three-step process of front-facing camera image processing is shown in accordance with an embodiment of the disclosure. An image 700 from a front-facing camera is used as input. An artificial neural network then processes the image using a semantic approach. Different regions of the image are identified that are covered by different types of objects: street, other vehicles, stationary objects, sky, and lane / street delineation. An intermediate (processed) image is shown in image 702. Lines correspond to boundaries between regions related to different types of objects. In the next step, navigation key points are determined based on this semantic understanding of the image. This approach allows for determination of navigation key points even if its location is occluded by other objects. The result is shown in image 704, where marker 706 indicates the location of the navigation key point.

[0038] In some embodiments, the fusion engine 210 can determine a confidence value for the image-related coordinates. The confidence value can be determined for the image-related coordinates. The confidence value indicates a probability that the image-related coordinates are the correct coordinates of the navigation key point. It can thus indicate whether the coordinates have been determined correctly or not.

[0039] The first training data subset can comprise one or more images for which the artificial neural network determines a confidence value below a predefined threshold. Thus, the training of the neural network can be performed more efficiently. The artificial neural network can be a convolutional neural network, as discussed above in Figure 3 During the training, the weights of the convolutional neural network can be set such that the convolutional neural network predicts navigation key point marker positions as close as possible to the positions of the navigation key point markers contained in the second training data subset. At the same time, the coordinates of the second training data subset can be obtained by at least one of user input, one or more crowd-sourcing platforms, and providing intended geocentric positions of key points in the predetermined area.

[0040] Optionally, the fusion engine 210 can provide the first training data subset to a second artificial neural network as input data for predicting image-related coordinates of navigation key points by the second neural network based on the first training data subset. Thus, a second confidence value can be determined, which indicates a distance between the navigation key point positions predicted by the trained artificial neural network and the second artificial neural network. Thus, the effect of the training can be monitored and parameters, including a threshold for the confidence value, can be adjusted.

[0041] The image-related coordinates can indicate a position of the navigation key point on the image. They can be represented as a row and column number of pixels. The processing can comprise a generation of a heat map depicting a probability value for a point on the image being a navigation key point. The processing can further comprise additional image processing steps known in the art. The navigation key points can be used to display instructions correctly to a driver. This can be achieved by superimposing the instructions, street names and other outputs on a camera image or a similar depiction of the surroundings of the mobile device.

[0042] In embodiments, the method further comprises converting the image-related coordinates to geocentric coordinates, as described with respect to the landmark rendering engine 114. This is possible because the navigation key points are related to stationary objects, such as road signs, intersection corners, and lanes. This can be done in a two-step process: first, by projection, the image-related coordinates are converted to coordinates relative to the device, for example, expressed as a distance and a polar angle relative to predefined axes fixed to the device. In a second step, the position of the device is determined and the coordinates are converted to geocentric coordinates, for example, longitude and latitude. The geocentric coordinates can then be used by the same mobile device or other devices to identify the key point locations. Thus, other mobile devices that are not configured to perform the methods of the present disclosure can use these data. Furthermore, if camera data is not available due to bad weather, camera malfunction, or other conditions, these data can be used.

[0043] In embodiments, the artificial neural network and / or deep neural network as described above can utilize HRNet architecture features with modifications specific to the task, as described in Figure 8

[0044] Figure 9 Another exemplary camera image 900 is shown, as well as a schematic overhead map 902 generated from a standard definition (SD) road map. The schematic overhead map 902 can alternatively be generated from aerial or satellite images, and can be stored in memory in the vehicle and / or in a network-accessible server. The overhead map 902 can also be a pure topological map, in which the roads and their connections are presented as edges and vertices of a graph with known coordinates. Alternatively, satellite images, aerial images, or map images, such as standard definition (SD) maps, can be used directly, where the SD maps consist of roads and intersections. In exemplary embodiments, both images can be used as input for processing, such that the predicted navigation key points are superimposed onto the standard definition (SD) map, thereby improving the reliability and accuracy of determining the navigation key points. As an optional preprocessing step, a perspective transformation can be performed on the camera image 900 or the schematic overhead image 902, such that the image coordinates (pixel rows and columns) of both input images are related to the same geocentric coordinates. This is particularly useful when the camera image quality is low due to bad weather (e.g., fog) or other conditions (e.g., dirt on the camera).

[0045] Figure 10 ​A flowchart 1000 showing an embodiment of the method of the present disclosure is shown. In step 1002, a first vehicle mounted camera is instructed to capture a plurality of first images. The first vehicle mounted camera can be the FFC 202 as described above, and can be instructed by the VO unit 204 as described above. Step 1004 requires receiving the plurality of first images, wherein the plurality of first images comprises at least a first image captured at a first time interval and a second image captured at a second time interval after the first time interval. At step 1006, the method requires receiving GPS coordinates, inertial measurement unit (IMU) measurements, and vehicle speed. The GPS coordinates are received from the GNSS module 208 as described above, while the IMU measurements are received from the IMU module 206 as described above. As described above, all measurements are received at the fusion engine 210. At step 1008, the VO unit 204 calculates an angular change between the first image and the second image, and at step 1010, the VO unit 204 determines a position and heading vector of the vehicle from the calculated angular change. This allows the navigation system of the vehicle to determine the position and heading vector of the vehicle in real time, which allows the system to reliably predict turn points and lane changes using augmented reality, as discussed above. At step 1012, the fusion engine 210 collects at least one image of the plurality of first images as a first training subset. At step 1014, a second training data subset is obtained from image related coordinates of navigation key points related to at least one image of the first training data subset. The method continues with step 1016, wherein the first training data subset and the second training data subset are provided to an artificial neural network (such as a deep neural network) as a training data set. Thus, the fusion engine 210 can train the artificial neural network on the training data set to predict image related coordinates of navigation key points indicative of road sign locations and / or turn points, as described in step 1018. Finally, in step 1020, the artificial neural network processes an input data set to predict image related coordinates of navigation key points indicative of road sign locations and / or turn points. Figure 3 A flowchart 1000 showing an embodiment of the method of the present disclosure is shown. In step 1002, a first vehicle mounted camera is instructed to capture a plurality of first images. The first vehicle mounted camera can be the FFC 202 as described above, and can be instructed by the VO unit 204 as described above. Step 1004 requires receiving the plurality of first images, wherein the plurality of first images comprises at least a first image captured at a first time interval and a second image captured at a second time interval after the first time interval. At step 1006, the method requires receiving GPS coordinates, inertial measurement unit (IMU) measurements, and vehicle speed. The GPS coordinates are received from the GNSS module 208 as described above, while the IMU measurements are received from the IMU module 206 as described above. As described above, all measurements are received at the fusion engine 210. At step 1008, the VO unit 204 calculates an angular change between the first image and the second image, and at step 1010, the VO unit 204 determines a position and heading vector of the vehicle from the calculated angular change. This allows the navigation system of the vehicle to determine the position and heading vector of the vehicle in real time, which allows the system to reliably predict turn points and lane changes using augmented reality, as discussed above. At step 1012, the fusion engine 210 collects at least one image of the plurality of first images as a first training subset. At step 1014, a second training data subset is obtained from image related coordinates of navigation key points related to at least one image of the first training data subset. The method continues with step 1016, wherein the first training data subset and the second training data subset are provided to an artificial neural network (such as a deep neural network) as a training data set. Thus, the fusion engine 210 can train the artificial neural network on the training data set to predict image related coordinates of navigation key points indicative of road sign locations and / or turn points, as described in step 1018. Finally, in step 1020, the artificial neural network processes an input data set to predict image related coordinates of navigation key points indicative of road sign locations and / or turn points.

Claims

1. A computer-implemented method for determining vehicle position and predicting navigation key points, the method comprising the following steps: Instruct the first vehicle-mounted camera to capture multiple first images; The plurality of first images are received from the first vehicle-mounted camera, wherein the plurality of first images includes at least a first image captured at a first time interval and a second image captured at a second time interval after the first time interval; It receives GPS coordinates, inertial measurement unit (IMU) measurements, and vehicle speed; Calculate the angle change between the first image and the second image; The position and heading vector of the vehicle are determined by calculating the angle changes; Based on the received GPS coordinates, IMU measurements, and vehicle speed, determine the confidence level of the determined vehicle position and heading vector; Collect at least one image from the plurality of first images as a first subset of training data; Obtain the image-related coordinates of navigation key points associated with at least one image in the first training data subset as a second training data subset; The first training data subset and the second training data subset are provided to the artificial neural network as training datasets; The artificial neural network is trained on the training dataset to predict the image-related coordinates of navigation key points indicating landmark locations and / or turning points; as well as The artificial neural network processes the first image from the first vehicle-mounted camera as input data to predict the image-related coordinates of navigation key points indicating road sign locations and / or turning points.

2. The computer-implemented method according to claim 1, further comprising: If the confidence value is lower than a predetermined threshold, the first vehicle-mounted camera is instructed to capture multiple second images; as well as The plurality of second images are received from the first vehicle-mounted camera, wherein: The plurality of second images includes at least a third image captured at a third time interval and a fourth image captured at a fourth time interval after the third time interval; and The time difference between the third time interval and the fourth time interval is less than the time difference between the first time interval and the second time interval. The computer-implemented method further includes: Calculate the angular change between the third image and the fourth image; and The vehicle's position and heading vector are determined based on the calculated angular change between the third and fourth images.

3. The computer-implemented method according to claim 1, further comprising: The second vehicle-mounted camera was instructed to capture multiple third images; as well as The plurality of third images are received from the second vehicle-mounted camera.

4. The computer-implemented method according to claim 1, further comprising: The image-related coordinates are converted to geocentric coordinates.

5. The computer-implemented method according to claim 4, further comprising: The geocentric coordinates are stored in a memory device included in the mobile device and / or in a network-accessible server.

6. The method according to claim 1, The artificial neural network mentioned therein is a convolutional neural network.

7. The method of claim 1, wherein the coordinates of the second training data subset are obtained by at least one of user input, one or more crowdsourcing platforms, and providing predetermined geocentric locations of key points in a predetermined region.

8. The method according to claim 1, further comprising: The first subset of training data is provided to the second artificial neural network as input data; The second artificial neural network predicts the image-related coordinates of navigation key points based on the first training data subset; as well as A second confidence value is determined, which indicates the distance between the location of the navigation key point predicted by the trained artificial neural network and the location of the navigation key point predicted by the second artificial neural network.

9. The computer-implemented method according to claim 1, further comprising: The predicted navigation key points are overlaid onto a standard definition (SD) map, which consists of roads and intersections.

10. An apparatus for determining vehicle position and predicting navigation keypoints, the apparatus comprising: First vehicle-mounted camera; A Global Navigation Satellite System (GNSS) module coupled to logic, the GNSS module being operable to receive GPS coordinates; An inertial measurement unit (IMU) coupled to the logic, the IMU being operable to receive IMU measurements; as well as The logic coupled to the first vehicle-mounted camera is operable to: Receive vehicle speed; Instruct the first vehicle-mounted camera to capture multiple first images; The plurality of first images are received from the first vehicle-mounted camera, wherein the plurality of first images includes at least a first image captured at a first time interval and a second image captured at a second time interval after the first time interval; Calculate the angle change between the first image and the second image; The position and heading vector of the vehicle are determined by calculating the angle changes; Based on the received GPS coordinates, IMU measurements, and vehicle speed, determine the confidence level of the determined vehicle position and heading vector; Collect at least one image from the plurality of first images as a first subset of training data; Obtain the image-related coordinates of navigation key points associated with at least one image in the first training data subset as a second training data subset; The first training data subset and the second training data subset are provided to the artificial neural network as training datasets; The artificial neural network is trained on the training dataset to predict the image-related coordinates of navigation key points indicating landmark locations and / or turning points; and The artificial neural network processes the first image from the first vehicle-mounted camera as input data to predict the image-related coordinates of navigation key points indicating road sign locations and / or turning points.

11. The device of claim 10, wherein the logic is also operable to perform the steps of any one of claims 2 to 9.

12. The device of claim 10, further comprising a second vehicle-mounted camera coupled to the logic, the second vehicle-mounted camera being operable to capture a plurality of second images.

Citation Information

Patent Citations

  • Apparatus and method for generating training data to train neural network determining information associated with road included in image

    US20180164812A1

  • Semantic visual landmarks for navigation

    US20190114507A1

  • Estimation of a pose of a robot

    WO2020048623A1