Navigation system for adjusting the vehicle position when GPS fails by matching natural signs with points of interest on a map
By using cameras to capture natural landmarks in the vehicle navigation system and combining a visual-language alignment system with a multimodal embedding generative model, the problem of inaccurate GPS positioning in deep urban environments was solved, and accurate vehicle positioning was achieved when GPS malfunctioned.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GM GLOBAL TECHNOLOGY OPERATIONS LLC
- Filing Date
- 2025-03-06
- Publication Date
- 2026-07-07
AI Technical Summary
In deep urban environments, GPS positioning systems are less reliable, leading to inaccurate vehicle positioning.
By using cameras to capture natural landmarks in the vehicle navigation system, and combining a visual-language alignment system and a multimodal embedding generative model, the natural landmarks are matched with points of interest on the map to adjust the vehicle's position.
In the event of GPS failure, it can accurately adjust the vehicle's position, improving the positioning accuracy and reliability of the navigation system.
Smart Images

Figure CN122345407A_ABST
Abstract
Description
[0001] introduce
[0002] The information provided in this section is for the purpose of generally presenting the context of this disclosure. Within the scope described in this section, the work of the currently named inventors and aspects of this description that might not otherwise conform to the prior art at the time of filing are neither expressly nor implicitly acknowledged as prior art relative to this disclosure.
[0003] This disclosure relates to a navigation system for vehicles, and more particularly to a navigation system for adjusting the vehicle's position by matching natural landmarks with points of interest on a map when the output of the Global Positioning System (GPS) fails.
[0004] The Global Positioning System (GPS) provides location data for vehicles along a route. However, in some driving situations, such as deep urban environments, GPS may not be as reliable as expected. Summary of the Invention
[0005] The navigation system for the vehicle includes a Global Positioning System (GPS) configured to generate GPS data for the vehicle. The navigation system is configured to store a map including points of interest (POIs) and their locations. A camera is configured to generate images along the vehicle's path. An image processing module is configured to identify landmarks and their locations from the images and output cropped landmark images and their locations. An alignment module is configured to adjust the vehicle's position in response to the cropped landmark images and their locations, as well as the POIs and their locations from the map.
[0006] Among other features, the alignment module includes a POI identifier module configured to identify POIs corresponding to the image and their locations. The alignment module implements a visual-linguistic alignment system configured to estimate the match between the text associated with the POI and the cropped marker image. The visual-linguistic alignment system includes a multimodal embedding generative model configured to generate similarity values.
[0007] Among other features, the alignment module includes an association and registration module configured to selectively adjust the vehicle position in response to a selected similarity value in the similarity values and the corresponding location of the POI. The association and registration module also includes a thresholding module configured to compare similarity values with a predetermined threshold. The thresholding module selects selected similarity values that are greater than the predetermined threshold.
[0008] Among other features, the association and registration module includes a calculation module configured to calculate von Mises Fisher likelihood based on selected similarity values. The association and registration module also includes a corrected coherence point drift module configured to generate probabilities that associate a POI with a cropped sign image. Finally, the association and registration module includes an association matrix and rigid transformation module configured to adjust vehicle position based on the probabilities generated by the corrected coherence point drift module.
[0009] A method for determining vehicle location includes: generating Global Positioning System (GPS) data for the vehicle; storing a map including points of interest (POIs) and the locations of the POIs; generating images along the vehicle's path; identifying natural landmarks and their locations from the images, and outputting cropped landmark images and their locations; and adjusting the vehicle location in response to the cropped landmark images and their locations, as well as the POIs and their locations from the map.
[0010] Among other features, the method includes identifying the POIs corresponding to the image and the location of the POIs. The method includes using a visual-linguistic alignment system to estimate the match between the text associated with the POI and the cropped sign image.
[0011] Among other features, the visual-language alignment system includes a multimodal embedding generative model. This model is configured to generate similarity values. The method includes selectively adjusting the vehicle position in response to a selected similarity value and its corresponding location at the point of interest (POI).
[0012] Among other features, the method includes: comparing similarity values to a predetermined threshold; and selecting a subset of similarity values that are greater than the predetermined threshold. The method includes calculating a von Mises Fisher likelihood based on the selected similarity values. The method includes: generating a probability that a POI is associated with a cropped sign image using a corrected coherence point drift; and adjusting vehicle positions based on the probability.
[0013] Further applications of this disclosure will become apparent from the detailed description, claims, and accompanying drawings. The detailed description and specific examples are intended for illustrative purposes only and are not intended to limit the scope of this disclosure.
[0014] This disclosure provides the following examples:
[0015] Example 1. A navigation system for a vehicle, comprising:
[0016] A Global Positioning System (GPS) is configured to generate GPS data for the vehicle.
[0017] The navigation system is configured to store a map including points of interest (POIs) and the locations of the POIs;
[0018] A camera is configured to generate images of the vehicle's path;
[0019] An image processing module is configured to identify natural landmarks and their locations from the image, and output a cropped image of the landmarks and the locations of the cropped images; and
[0020] The alignment module is configured to adjust the position of the vehicle in response to the following:
[0021] The captured marker image and its location; and
[0022] The POI from the map and the location of the POI.
[0023] Example 2. The navigation system according to Example 1, wherein the alignment module includes a POI identifier module configured to identify the POI corresponding to the image and the location of the POI.
[0024] Example 3. The navigation system according to Example 1, wherein the alignment module implements a visual-language alignment system configured to estimate the match between the text associated with the POI and the captured sign image.
[0025] Example 4. The navigation system according to Example 3, wherein the visual-language alignment system includes a multimodal embedding generative model.
[0026] Example 5. The navigation system according to Example 4, wherein the multimodal embedding generative model is configured to generate similarity values.
[0027] Example 6. The navigation system according to Example 5, wherein the alignment module includes an association and registration module configured to selectively adjust the position of the vehicle in response to a selected similarity value in the similarity values and the corresponding position of the POI.
[0028] Example 7. The navigation system according to Example 6, wherein the association and registration module includes a threshold module configured to compare the similarity value with a predetermined threshold.
[0029] Example 8. The navigation system according to Example 7, wherein the threshold module selects a selected similarity value from the similarity values that is greater than the predetermined threshold.
[0030] Example 9. The navigation system according to Example 8, wherein the association and registration module includes a calculation module configured to calculate von Mises Fisher likelihood based on the selected similarity value among the similarity values.
[0031] Example 10. The navigation system according to Example 9, wherein the association and registration module includes a corrected coherence point drift module configured to generate a probability of associating the POI with the captured marker image.
[0032] Example 11. The navigation system according to Example 10, wherein the association and registration module includes an association matrix and rigid transformation module, the association matrix and rigid transformation module being configured to adjust the position of the vehicle based on the probability generated by the corrected coherence point drift module.
[0033] Example 12. A method for determining the location of a vehicle, comprising:
[0034] Generate Global Positioning System (GPS) data for the vehicle;
[0035] Store a map including points of interest (POIs) and the locations of the POIs;
[0036] Generate an image of the vehicle's path;
[0037] Identify natural landmarks and their locations from the image, and output a cropped image of the landmarks and their locations; and
[0038] The position of the vehicle is adjusted in response to the following:
[0039] The captured marker image and the location of the captured marker image; and
[0040] The POI from the map and the location of the POI.
[0041] Example 13. The method according to Example 12 further includes identifying the POI corresponding to the image and the location of the POI.
[0042] Example 14. The method according to Example 12 further includes using a visual-language alignment system to estimate the match between the text associated with the POI and the cropped sign image.
[0043] Example 15. The method according to Example 14, wherein the visual-language alignment system includes a multimodal embedding generative model.
[0044] Example 16. The method according to Example 15, wherein the multimodal embedding generative model is configured to generate similarity values.
[0045] Example 17. The method according to Example 16 further includes selectively adjusting the position of the vehicle in response to a selected similarity value in the similarity values and the corresponding position of the POI.
[0046] Example 18. The method according to Example 17 further includes:
[0047] The similarity value is compared with a predetermined threshold; and
[0048] Select a similarity value that is greater than the predetermined threshold from the similarity values.
[0049] Example 19. The method according to Example 18 further includes calculating a von-Mises Fisher likelihood based on the selected similarity value among the similarity values.
[0050] Example 20. The method according to Example 19 further includes:
[0051] The probability of the POI being associated with the cropped marker image is generated using a corrected coherence point drift; and
[0052] The vehicle's position is adjusted based on the probability. Attached Figure Description
[0053] This disclosure will be more fully understood based on the detailed description and accompanying drawings, in which:
[0054] Figure 1 This is a functional block diagram of an example vehicle including a navigation system according to the present disclosure;
[0055] Figure 2 This is a functional block diagram of an example alignment system according to the present disclosure, which is configured to match the positions of natural landmarks from an image generated by a camera with points of interest and their positions from a map;
[0056] Figure 3 This is a functional block diagram of an example alignment system according to the present disclosure, which is configured to match natural landmarks with a map;
[0057] Figure 4 This is a flowchart illustrating an example of a method according to the present disclosure, which is used to match the location of natural features from an image generated by a camera with points of interest and their locations on a map;
[0058] Figures 5A to 5I This is an example image of a sign that has been cropped and separated from an image generated by the vehicle's camera;
[0059] Figure 6 This is an example of a scoring matrix according to the present disclosure, which includes similarity values between business names in an image and possible businesses identified as points of interest; and
[0060] Figure 7 The illustration shows an example of estimating the translation error based on the root mean square of the GPS translation error, using only location information and using both location and similarity values, in accordance with this disclosure.
[0061] In the accompanying drawings, reference numerals may be reused to identify similar and / or identical elements. Detailed Implementation
[0062] While this disclosure relates to a navigation system for vehicles that determines vehicle location by matching natural landmarks (such as commercial signs, traffic signs, and / or street signs) cropped from images generated from a camera with map-based points of interest, the navigation system can be used for other types of traffic.
[0063] The vehicle includes a navigation system that receives vehicle location data from the Global Positioning System (GPS). The navigation system locates the vehicle relative to a stored map and provides route data from the vehicle's current location to the desired destination. However, due to interference, obstacles, and other factors, GPS may have difficulty accurately determining the vehicle's location in some locations, such as deep urban areas.
[0064] This disclosure relates to determining vehicle location by matching natural signs perceived in a driving environment with points of interest (POIs) and their corresponding locations provided by a stored map. Examples of POI data include commercial signs, traffic signs, and / or street signs and their locations. Matching the text of a sign from an image to a POI is not binary because variations in wording and / or appearance must be compensated for.
[0065] In some examples, visual-language alignment systems can be used, including multimodal embedding generative models such as Contrastive Language-Image Pretraining (CLIP), Bootstrapping Language-Image Pretraining (BLIP), ALIGN, Locked Image Text Tuning (LiT), VisualBERT, LXMERT, ViLBERT, UNITER, SimVLM, ALBEF, etc. For example, CLIP is a neural network trained on a variety of image and text pairs. Without directly optimizing for the task, given an image and candidate text, CLIP can be guided by natural language to predict the most relevant text fragments (e.g., the business name from a sign).
[0066] Process images from vehicle cameras to detect one or more signs in the images. Crop images around one or more signs. Generate embeddings for text and images using a multimodal embedding generative model and estimate the match between candidate text (e.g., POI business names or street names generated from maps) and the localized sign images. Determine the location of each of the one or more localized sign images (e.g., relative to the vehicle coordinate system and then convert it to longitude and latitude values).
[0067] The map identifies Points of Interest (POIs) (e.g., businesses) and their locations along the vehicle's path. For all signs in the field of view, a multimodal embedding generative model is used to generate embedding score matrices (e.g., including similarity values) for all candidate POI name and sign truncation pairs. In some examples, results for similarity values above a predetermined threshold are further analyzed.
[0068] Von Mises-Fisher likelihood values are calculated based on similarity scores above a threshold. Corrected coherence point drift values are generated using rigid coherence point drift, adjusted by the Von Mises-Fisher likelihood values and the corresponding locations of the POIs. If necessary, the correlation matrix and rigid transformation are used to selectively adjust vehicle positions. This method can be used to determine vehicle positions along a route in GPS-constrained environments without requiring precise HD maps.
[0069] Now for reference Figure 1 Example of vehicle 10 includes a Global Positioning System (GPS) 14 configured to generate the coordinates of vehicle 10. If vehicle 10 is an autonomous vehicle, vehicle 10 may include a radio detection and ranging (radar) system 18 and / or an optical detection and ranging (LiDAR) system that can be used to detect objects in the vehicle's path. Vehicle 10 includes one or more cameras 34 that generate images located in the vehicle's path.
[0070] Vehicle 10 includes a controller 42, which includes a navigation module 46, an alignment module 50, an image processing module 53, and an optional autonomous driving module 54. The navigation module 46 is configured to navigate the vehicle relative to a map 52 that includes points of interest (POIs) and their corresponding locations (e.g., longitude and latitude coordinates). The image processing module 53 is configured to identify and locate (or crop) landmarks (and their corresponding locations) in images generated by camera 34. LiDAR, radar, or other distance sensors may also be used for this purpose. The alignment module 50 is configured to determine the positions of natural landmarks and points of interest along the vehicle's path, and their locations (e.g., longitude and latitude coordinates).
[0071] The controller 42 may also include an autonomous driving module 54, which is configured to automatically operate the vehicle's steering, braking, and acceleration when enabled. The autonomous driving module 54 receives vehicle input, such as vehicle speed, wheel speed, steering angle, brake pedal position, accelerator pedal position, etc. When enabled, the autonomous driving module 54 is configured to control steering, braking, and / or acceleration along a route from the vehicle's current position to its destination position based on GPS data and / or selectively using natural landmarks for alignment.
[0072] Now for reference Figure 2 The coordinates output by GPS 114 are input to the Point of Interest (POI) labeling module 124 and the association and registration module 132. The map module 118 provides the POI labeler module 124 with POI data and location data of the current vehicle path based on the vehicle's position and heading.
[0073] Camera 136 outputs an image (in the direction of vehicle travel) to image processing module 140. Image processing module 140 identifies signs in the image, locates the image (finds the position of the sign in the image, zooms and / or crops the image around the sign), and forwards the located sign image to similarity score matrix calculation module 128.
[0074] The similarity score matrix calculation module 128 includes a multimodal embedding generative model 129 for generating the similarity score matrix. In some examples, the similarity score matrix is based on cosine similarity values, dot products, Euclidean distances, etc., between the POI business name, traffic information, or street name and the corresponding screenshot of the located sign image. The image processing module 140 outputs the location or coordinates of the located sign image to the association and registration module 132.
[0075] Now for reference Figure 3 The association and registration module 132 is shown in further detail. The similarity score matrix 214 is input to the thresholding module 217, which applies a threshold to the scores in the matrix. In other words, only flags with sufficiently high scores are output to the von Mises-Fisher likelihood calculation module 218. Each embedding score c between image m and POI n is calculated using the following relationship. mn Convert to von Mises-Fisher likelihood l mn :
[0076] l mn =C p (κ)exp(κ·c mn )
[0077] Where Cp(κ) is a normalization constant and κ is a tunable parameter.
[0078] The likelihood value output by the von Mises-Fisher likelihood calculation module 218 is input to the corrected coherence point drift module 220. The corrected coherence point drift module 220 is based on the likelihood value l mn (Where s = 1) The rigid coherence point drift module is modified to generate probability values, as shown below:
[0079]
[0080] Where c mn The input y corresponds to the cosine similarity between the label n (image) and the POI m (text). m The input is the 2D coordinates (e.g., ENU coordinates) corresponding to the POI m position in the Cartesian coordinate system, x n σ is the input corresponding to the 2D coordinates of the marker n in the vehicle coordinate system, w is the input corresponding to the outlier probability, and σ 2 It corresponds to y m The noise variance is the input in the data set. D is the dimension of the point set (e.g., D = 2). M and N correspond to the number of points in the point set.
[0081] p mn This is the output corresponding to the probability associated with m and n. R and t are the rotation matrix and translation vector used to align m and n. Additional details related to rigid coordinate point drift can be found in Myronenko, Andriy, and Xubo Song, "Point set registration: Coherentpoint drift", IEEE Transactions on Pattern Analysis and Machine Intelligence 32.12(2010):2262-2275, which is incorporated herein by reference in its entirety. As described above, the likelihood value l is used as described in this paper. mn The formulas in the paper have been revised.
[0082] Now for reference Figure 4 This paper illustrates a method for aligning a navigation system using natural landmarks. At 310, the method determines whether a landmark is detected in an image of the vehicle's path obtained from a camera. The method then crops and locates the landmark. If 310 is true, the method identifies the Point of Interest (POI) and its location from the map corresponding to the vehicle's path at 314. At 318, a similarity score matrix is generated. At 322, the location of the POI is used to determine position and appearance alignment. At 326, the final correspondence is determined. At 330, the vehicle position is aligned with the route.
[0083] Now for reference Figures 5A to 5IThis shows an example image of a sign cropped and located from a sample image (similar to an image acquired from a vehicle). Figure 5A In the image, a portion is separated and includes the Belle Tire logo. Figure 5B and Figure 5C In the image, a portion is separated and includes the Goodyear logo. Figure 5D and 5E In the image, a portion is separated and includes the Dollar General logo. Figure 5F In the image, a portion is separated and includes the AT&T logo. Figure 5G In the image, a portion is separated and includes the Starbucks logo. Figure 5H In the image, a portion is separated and includes the T-Mobile logo. Figure 5I In the image, a portion is separated and includes the Thai Kitchen logo. The cropped image (e.g., Figures 5A to 5I Those shown in the diagram are provided along with the first location data (longitude and latitude).
[0084] Now for reference Figure 6 Examples of similarity score matrices include cosine similarity values between located logo images and POI business names (e.g., generated using the CLIP model as a visual-language alignment system). Cosine similarity can be computed as a normalized dot product between text embedding vectors and image embedding vectors.
[0085] It is understood that the highest cosine similarity value in each column of the similarity score matrix corresponding to the actual business name would be as expected. While this is true in most cases, reliance on this information can lead to errors. For example, Thai Kitchen may not be accurately identified. To reduce errors, the association and registration module 132 uses POI location data to improve data accuracy.
[0086] Now for reference Figure 7The diagram illustrates the estimation of translation error based on the root mean square of the GPS translation error, using only location information and both location and cosine similarity values. Using the data shown above, the GPS error is randomly varied, and the registration error is measured both for location information (without a visual-language alignment system) and for location information using a visual-language alignment system. For GPS errors <18m, association using either method is feasible. For GPS errors <50m, association using both location and visual-language alignment data is significantly more robust (with alignment accuracy of 3 to 4m within these ranges). Zero alignment error is impossible due to biases in sign-POI locations. The dataset is not necessarily typical due to the large number of perceived signs. Association using both location and visual-language alignment data performs significantly better with fewer signs. In this example, the visual-language alignment system is CLIP.
[0087] The foregoing description is merely illustrative in nature and is in no way intended to limit this disclosure, its application, or use. The broad teachings of this disclosure can be implemented in many forms. Therefore, while this disclosure includes specific examples, its true scope should not be so limited, as other modifications will become apparent upon examination of the drawings, specification, and appended claims. It should be understood that one or more steps within a method may be performed in a different order (or simultaneously) without altering the principles of this disclosure. Furthermore, while each embodiment is described above as having specific features, any one or more of those features described with respect to any embodiment of this disclosure may be implemented in any other embodiment and / or combined with features of any other embodiment, even if such combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and the arrangement of one or more embodiments with respect to each other remains within the scope of this disclosure.
[0088] Spatial and functional relationships between components (e.g., between modules, circuit elements, semiconductor layers, etc.) are described using various terms, including “connection,” “joint,” “coupled,” “proximity,” “adjacent,” “on top of,” “above,” “below,” and “set.” Unless explicitly stated as “direct,” the relationship between the first and second components described in the above disclosure can be a direct relationship, where no other intermediate components exist between the first and second components, but it can also be an indirect relationship, where one or more intermediate components exist (spatially or functionally) between the first and second components. As used herein, the phrase “at least one of A, B, and C” should be interpreted as meaning logically (A or B or C) using the non-exclusive logical “OR,” and should not be interpreted as meaning “at least one of A, at least one of B, and at least one of C.”
[0089] In the accompanying drawings, the direction of the arrows, as indicated by the arrows, typically illustrates the flow of information (e.g., data or instructions) of interest to the illustration. For example, when components A and B exchange various types of information, but the information transmitted from component A to component B is relevant to the illustration, the arrow can point from component A to component B. This unidirectional arrow does not imply that no other information is transmitted from component B to component A. Furthermore, for information sent from component A to component B, component B can send a request for or confirmation of receipt of that information to component A.
[0090] In this application, including the following definitions, the term "module" or "controller" may be replaced by the term "circuit". The term "module" may refer to, be part of, or include the following: application-specific integrated circuit (ASIC); digital, analog, or mixed-signal analog / digital discrete circuit; digital, analog, or mixed-signal analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor circuitry (shared, dedicated, or grouped) that executes code; memory circuitry (shared, dedicated, or grouped) that stores code executed by the processor circuitry; other suitable hardware components that provide the described functionality; or combinations of some or all of the foregoing, such as in a system-on-a-chip.
[0091] This module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces connected to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module of this disclosure can be distributed among multiple modules connected via the interface circuits. For example, multiple modules can allow for load balancing. In a further example, a server (also referred to as a remote or cloud) module may perform some functions on behalf of a client module.
[0092] The term "code" as used above can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, data structures, and / or objects. The term "shared processor circuit" covers a single processor circuit that executes some or all of the code from multiple modules. The term "group processor circuit" covers a processor circuit that, in conjunction with additional processor circuitry, executes some or all of the code from one or more modules. References to multiple processor circuits cover multiple processor circuits on discrete dies, multiple processor circuits on a single die, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or a combination of the above. The term "shared memory circuit" covers a single memory circuit that stores some or all of the code from multiple modules. The term "group memory circuit" covers a memory circuit that, in conjunction with additional memory, stores some or all of the code from one or more modules.
[0093] The term "memory circuit" is a subset of the term "computer-readable medium." As used herein, the term "computer-readable medium" does not cover transient electrical or electromagnetic signals propagating through a medium (such as on a carrier wave); therefore, the term "computer-readable medium" can be considered tangible and non-transitory. Non-limiting examples of non-transitory, tangible computer-readable media are non-volatile memory circuits (such as flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (such as static random access memory circuits or dynamic random access memory circuits), magnetic storage media (such as analog or digital magnetic tape or hard disk drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs).
[0094] The apparatus and methods described in this application can be implemented, in part or in whole, by a special-purpose computer created by configuring a general-purpose computer to execute one or more specific functions embodied in a computer program. The aforementioned function blocks, flowchart components, and other elements serve as a software specification that can be routinely translated into a computer program by a skilled technician or programmer.
[0095] A computer program includes processor-executable instructions stored on at least one non-transitory tangible computer-readable medium. A computer program may also include or depend on stored data. A computer program may encompass a basic input / output system (BIOS) that interacts with the hardware of a special-purpose computer, device drivers that interact with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.
[0096] Computer programs may include: (i) descriptive text to be parsed, such as HTML (Hypertext Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation); (ii) assembly code; (iii) object code generated from source code by a compiler; (iv) source code for execution by an interpreter; (v) source code for compilation and execution by a just-in-time (JIT) compiler; and so on. As an example only, source code can be written using syntax from languages including: C, C++, C#, Objective C, Swift, Haskell, Go, SQL, R, Lisp, etc. Fortran, Perl, Pascal, Curl, OCaml, HTML5 (Hypertext Markup Language 5th Edition), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK and
Claims
1. A navigation system for a vehicle, comprising: A Global Positioning System (GPS) is configured to generate GPS data for the vehicle. The navigation system is configured to store a map including points of interest (POIs) and the locations of the POIs; A camera is configured to generate images of the vehicle's path; An image processing module is configured to identify natural landmarks and their positions from the image, and output a cropped landmark image and the position of the cropped landmark image; as well as The alignment module is configured to adjust the position of the vehicle in response to the following: The captured marker image and its location; and The POI from the map and the location of the POI.
2. The navigation system of claim 1, wherein the alignment module includes a POI identifier module configured to identify the POI corresponding to the image and the position of the POI.
3. The navigation system of claim 1, wherein the alignment module implements a visual-language alignment system configured to estimate the match between the text associated with the POI and the captured sign image.
4. The navigation system according to claim 3, wherein the vision-language alignment system comprises a multimodal embedding generative model.
5. The navigation system of claim 4, wherein the multimodal embedding generation model is configured to generate similarity values.
6. The navigation system of claim 5, wherein the alignment module includes an association and registration module configured to selectively adjust the position of the vehicle in response to a selected similarity value in the similarity values and the corresponding position of the POI.
7. The navigation system of claim 6, wherein the association and registration module includes a threshold module configured to compare the similarity value with a predetermined threshold.
8. The navigation system according to claim 7, wherein the threshold module selects a selected similarity value from the similarity values that is greater than the predetermined threshold.
9. The navigation system of claim 8, wherein the association and registration module includes a calculation module configured to calculate von Mises Fisher likelihood based on the selected similarity value among the similarity values.
10. The navigation system of claim 9, wherein the association and registration module includes a corrected coherence point drift module, the corrected coherence point drift module being configured to generate a probability of associating the POI with the captured marker image.