Method and system for generating navigational view images from street view representations

The method generates navigational view images from street view images by extracting and processing navigational clips with object removal and lane boundary detection, addressing the limitations of existing navigation systems to improve navigation accuracy and reduce confusion at complex intersections.

WO2026156821A1PCT designated stage Publication Date: 2026-07-30GRABTAXI HOLDINGS PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GRABTAXI HOLDINGS PTE LTD
Filing Date
2025-01-26
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing navigation systems struggle to provide accurate and intuitive representations of complex intersections and multi-lane interchanges, leading to increased driver confusion and potential accidents due to the lack of depth and context in 2D maps and text-based directions, and inefficient 3D modeling techniques.

Method used

A method and system that generates navigational view images by extracting a navigational image clip from a 360-degree street view image, removing objects using object detection and segmentation, detecting lane boundaries, and adding directional indicators based on user input, to create a clear and interactive navigation interface.

Benefits of technology

Enhances navigation by providing accurate and intuitive visual guidance at complex junctions, reducing driver confusion and potential accidents through precise lane and direction representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075262_30072026_PF_FP_ABST
    Figure CN2025075262_30072026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure relates to method and system for generating navigational view image from street view image. The method includes receiving the street view image centered at a junction. The method further includes extracting navigational image clip from the received street view image based on geographic coordinates derived from OpenStreetMap (OSM) node sequence. The method further includes detecting lane boundaries within the navigational image clip. Further, the method includes generating directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The method further includes combining the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR GENERATING NAVIGATIONAL VIEW IMAGES FROM STREET VIEW REPRESENTATIONSDESCRIPTIONTechnical Field

[0001] This disclosure relates generally to transportation, and more particularly to a method and system for generating a navigational view image from a street view image.Background

[0002] Modern navigation systems are critical in assisting drivers with real-time route guidance. However, navigating through complex intersections and highway interchanges has always been challenging for drivers, particularly in high-traffic or unfamiliar areas. At complex intersections and highway interchanges, accurate representation of the road layout is essential to reduce confusion and ensure safe driving.

[0003] Existing navigation systems primarily rely on conventional techniques such as, Two-Dimensional (2D) map views or text-based turn-by-turn directions to guide drivers. The conventional methods struggle to address the needs of complex road scenarios. In 2D map views, the representation of intersections and interchanges lacks depth and context, making it difficult for drivers to visualize the actual road structure or identify the correct lane to take. Similarly, text-based directions can be difficult to follow in real-time, especially when multiple road signs and lane options must be considered simultaneously. The limitations result in increased driver stress, a higher likelihood of errors, and a potential rise in traffic disruptions or accidents, particularly at junctions with intricate designs or multi-lane interchanges.

[0004] Further, the conventional techniques may include methods for displaying navigation guidance based on simplified 2D maps augmented with text-based instructions or generic lane arrows. The conventional techniques may include static images or diagrams for junctions but are often pre-rendered and lack customization for real-time conditions or specific user contexts. Additionally, the conventional techniques struggle to integrate lane markings, driving directions, and road signs cohesively while preserving an intuitive and accurate visual representation. Some conventional techniques may also attempt to overlay instructions on real-world images or videos. However, such conventional techniques either obscure essential details or require resource-intensive processing, limiting the scalability and efficiency of the existing navigation systems. Further, some conventional techniques may include generating Three-Dimensional (3D) model of the complex junction or intersection to represent the multi-lanes, flyovers, underpasses. However, generating 3D models usually require 3D modelling and simulation, that is not efficient in terms of time, resource, and scalability.

[0005] There is, therefore, a need in the present state of art for techniques to generate navigational view image of the complex intersections and the multi-lane interchanges.SUMMARY

[0006] In one embodiment, a method of generating a navigational view image from a street view image is disclosed. In one example, the method includes receiving the street view image cantered at a junction. The method further includes extracting a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The method further includes detecting lane boundaries within the navigational image clip. Further, the method includes generating directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The method further includes combining the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.

[0007] In an aspect, when the navigational image clip comprises one or more objects, the method includes processing the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique. The one or more objects includes pedestrians and vehicles within the navigational image clip.

[0008] In another aspect, to process the extracted navigational image clip, the method includes identifying pixels corresponding to the one or more objects within the navigational image clip using the object detection and segmentation model. Further, the method includes eliminating the identified pixels from the navigational image clip using the object removal technique.

[0009] In another aspect, to detect lane boundaries, the method includes identifying lane features from the processed navigational image clip, using an edge detection technique. Further, the method includes clustering the identified lane features to extract lane polylines and converting the lane features into defined lane boundaries.

[0010] In one aspect, the method includes displaying the navigational view image to a user through a user interface to assist with navigation decisions at the junction.

[0011] In another aspect, to processing the navigational image clip, the method includes detecting shadows in navigational image clip. Further, the method includes applying an object removal technique to remove pixels corresponding to the detected shadows.

[0012] In another aspect, to generate directional indicators, the method includes comprises determining the intended driving direction based on a user input regarding a desired route and the lane boundaries.

[0013] In another aspect, the method includes capturing user feedback regarding effectiveness of the navigational view image in aiding navigation decisions. The method further includes utilizing the feedback to enhance subsequent image generation processes.

[0014] In one aspect, the street view image is a 360-degree street view image.

[0015] In another aspect, to extract the navigational image clip, the method includes determining a heading direction and a center location, wherein a heading angle at which a user is approaching the junction is calculated. The heading angle is calculated based on geographic coordinates of a location of the user relative to the junction. Further, the method includes extracting the navigational image clip, based on the determined heading direction and the center location, by querying an associated database.

[0016] In one embodiment, a system for generating a navigational view image from a street view image is disclosed. In one example, the system includes a processor, and a memory communicatively coupled to the processor. The memory stores processor-executable instructions, which, on execution, cause the processor to receive the street view image cantered at a junction. The processor-executable instructions, on execution, further cause the processor to extract a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The processor-executable instructions, on execution, further cause the processor to detect lane boundaries within the navigational image clip. The processor-executable instructions, on execution, further cause the processor to generate directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The processor-executable instructions, on execution, further cause the processor to combine the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.

[0017] In one aspect, when the navigational image clip comprises one or more objects, the processor-executable instructions further cause the processor to, process the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique. The one or more objects include pedestrians and vehicles within the navigational image clip. The street view image is a 360-degree street view image

[0018] In another aspect, the processor-executable instructions further cause the processor to process the extracted navigational image clip by identifying pixels corresponding to the one or more of objects within the navigational image clip using the object detection and segmentation model and eliminating the identified pixels from the navigational image clip using the object removal technique.

[0019] In another aspect, the processor-executable instructions further cause the processor to detect lane boundaries by identifying lane features from the processed navigational image clip, using an edge detection technique and clustering the identified lane features to extract lane polylines and converting the lane features into defined lane boundaries.

[0020] In another aspect, the processor-executable instructions further cause the processor to display the navigational view image to a user through a user interface to assist with navigation decisions at the junction.

[0021] In another aspect, the processor-executable instructions further cause the processor to process the navigational image clip by detecting shadows in navigational image clip and applying the object removal technique to remove pixels corresponding to the detected shadows.

[0022] In another aspect, the processor-executable instructions further cause the processor to generate directional indicators by determining the intended driving direction based on a user input regarding a desired route and the lane boundaries.

[0023] In another aspect, the processor-executable instructions further cause the processor to capture user feedback regarding effectiveness of the navigational view image in aiding navigation decisions. Further, the processor-executable instructions further cause the processor to utilize the feedback to enhance subsequent image generation processes.

[0024] In another aspect, the processor-executable instructions further cause the processor to extract the navigational image clip by determining a heading direction and a center location. A heading angle at which a user is approaching the junction is calculated, and wherein the heading angle is calculated based on geographic coordinates of a location of the user relative to the junction and extracting the navigational image clip, based on the determined heading direction and the center location, by querying an associated database.

[0025] In another embodiment, a non-transitory computer-readable medium storing computer-executable instructions for generating a navigational view image from a street view image, the computer-executable instructions configured for receiving the street view image cantered at a junction. The computer-executable instructions further configured for extracting a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The computer-executable instructions further configured for detecting lane boundaries within the navigational image clip. Further, the computer-executable instructions configured for generating directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The computer-executable instructions further configured for combining the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.

[0026] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles.

[0028] FIG. 1 is a block diagram of an exemplary system for generating a navigational view image from a street view image, in accordance with some embodiments.

[0029] FIG. 2 illustrates a functional block diagram of a computing device of the exemplary system for generating the navigational view image from the street view image, in accordance with some embodiments.

[0030] FIG. 3 illustrates an exemplary process for extracting a navigational image clip from the street view image, in accordance with some embodiments.

[0031] FIGs. 4A –4E illustrate exemplary street view image and the navigational view image, in accordance with some embodiments.

[0032] FIG. 5A illustrates exemplary User Interfaces (UIs) of an existing navigation system, in accordance with some embodiments.

[0033] FIG. 5B illustrates exemplary User Interfaces (UIs) of an exemplary navigation system configured to generate the navigational view image, in accordance with some embodiments.

[0034] FIG. 6 illustrates an exemplary process of a method for generating the navigational view image from the street view image, in accordance with some embodiments.

[0035] FIG. 7 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.DETAILED DESCRIPTION

[0036] Exemplary embodiments are described with reference to the accompanying drawings. Wherever convenient, the same reference numbers are used throughout the drawings to refer to the same or like parts. While examples and features of disclosed principles are described herein, modifications, adaptations, and other implementations are possible without departing from the spirit and scope of the disclosed embodiments. It is intended that the following detailed description be considered as exemplary only, with the true scope and spirit being indicated by the following claims.

[0037] Referring now to FIG. 1, an exemplary system 100 for generating a navigational view image from a street view image is illustrated, in accordance with some embodiments. The system 100 may implement a computing device 102 (for example, server, desktop, laptop, notebook, netbook, tablet, smartphone, mobile phone, or any other computing device) , in accordance with some embodiments of the present disclosure. The computing device 102 may generate navigational view image from street view.

[0038] As will be described in greater detail in conjunction with FIGS. 2 -7, the computing device 102 receives the street view image cantered at a junction. The computing device 102 further extracts a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The navigational image clip includes one or more objects. Further, the one or more objects include at least pedestrians and vehicles within the navigational image clip, and the street view image is a 360-degree street view image. To extract the navigational image clip, the computing device 102 determines a heading direction and a center location. A heading angle at which a user is approaching the junction is calculated, and the heading angle is calculated based on geographic coordinates of a location of the user relative to the junction. Further, the computing device 102 extracts the navigational image clip, based on the determined heading direction and the center location, by querying an associated database.

[0039] Further, the computing device 102 processes the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique. To extract the navigational image clip, the computing device 102 may identify pixels corresponding to the one or more objects within the navigational image clip using the object detection and segmentation model. The computing device 102 detects shadows in navigational image clip. The computing device 102 further applies the object removal technique to remove pixels corresponding to the detected shadows.

[0040] Further, the computing device 102 detects lane boundaries within the processed navigational image clip. To detect lane boundaries, the computing device 102 identifies lane features from the processed navigational image clip, using an edge detection technique. The computing device 102 further clusters the identified lane features to extract lane polylines and convert the lane features into defined lane boundaries. The computing device 102 further generates directional indicators within the processed navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The computing device 102 determines the intended driving direction based on a user input regarding a desired route and the lane boundaries.

[0041] The computing device 102 further combines the processed navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image. Further, the computing device 102 displays the navigational view image to a user through a user interface to assist with navigation decisions at the junction. The computing device 102 further captures user feedback regarding effectiveness of the navigational view image in aiding navigation decisions. The computing device 102 utilizes the ufeedback to enhance subsequent image generation processes.

[0042] In some embodiments, the computing device 102 may include one or more processors 104 and a computer-readable medium 106 (for example, a memory) . The computer-readable medium 106 may include the database. Further, the computer-readable storage medium 106 may store instructions that, when executed by the one or more processors 104, cause the one or more processors 104 to generate navigational view image from street view image, in accordance with aspects of the present disclosure. The computer-readable storage medium 106 may also store various data (for example, demand and supply data, vehicle data, Geo-location, customer data, historical data, and the like) that may be captured, processed, and / or required by the system 100.

[0043] The system 100 may further include a display 108. The system 100 may interact with a user via a user interface 110 accessible via the display 108. The system 100 may also include one or more user devices 112. In some embodiments, the computing device 102 may interact with the one or more user devices 112 over a communication network 114 for sending or receiving various data. The user devices 112 may include, but may not be limited to, a remote server, a digital device, or another computing system.

[0044] Referring now to FIG. 2, functional block diagram of the computing device 102 of an exemplary system for generating a navigational view image from a street view image is illustrated, in accordance with some embodiments. In an embodiment, the system 200 may be analogous to the system 100. The system 200 may include the computing device 102. Further, the computing device 102 may include a receiving module 202, an extracting module 204, a processing module 206, a detecting module 208, a generating module 210, and a combining module 212.

[0045] In an embodiment, the receiving module 202 may be configured to receive the street view image centered at the junction. The street view image is a 360-degree street view image. The receiving module 202 is further configured to request the street view image based on the geographic coordinates of the junction, from an external database or repository, and parse the received street view image to isolate the relevant junction area. The receiving module 202 may utilize Application Programming Interfaces (APIs) or data streaming mechanisms to fetch 360-degree street view images in real time from third-party mapping or navigation services. The 360-degree street view image allows for capturing the complete spatial context of the junction, ensuring that all surrounding elements, such as roads, lanes, and signage, are included in the street view image.

[0046] In an embodiment, the extracting module 204 may extract a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The navigational image clip includes one or more objects. The one or more objects include at least one of pedestrians and vehicles within the navigational image clip. The extracting module 204 is further configured to determine a heading direction and a center location within the navigational image clip. The heading direction and the center location are derived based on the geographic coordinates of the user's position relative to the junction. Further, the extracting module 204 calculates a heading angle at which the user is approaching the junction. The heading angle may be computed using the geographic coordinates of the user's location and the relative position of the junction. Upon determining the heading direction and the center location, the extracting module 204 may query an associated database to extract the relevant navigational image clip, ensuring that the extracted navigational image clip is contextually aligned with the user's real-time position and orientation.

[0047] Upon extracting the navigational image clip, the processing module 206 may be configured to process the extracted navigational image clip to remove the one or more objects from the navigational image clip using an object detection and segmentation model and an object removal technique. In simpler words, the processing module 206 is configured to enhance the usability and clarity of the navigational image clip by removing unnecessary or distracting objects using advanced image processing techniques. In an embodiment, the processing module 206 is configured to identify and segment objects within the navigational image clip using an object detection and segmentation model. The object detection and segmentation model detect objects such as vehicles, pedestrians, road barriers, or temporary structures present in the navigational image clip. The object detection and segmentation model may be a machine learning model such as convolutional neural networks (CNNs) and pre-trained object detection frameworks such as You Only Look Once (YOLO) and Mask R-CNN. Further, the object detection and segmentation model segment the identified objects into distinct areas, isolating them from the background of the navigational image clip.

[0048] Once the objects are segmented, the processing module 206 may apply an object removal technique to eliminate the detected objects from the navigational image clip while preserving the surrounding visual context. In an embodiment, the processing module 206 reconstructs the background in the area of removed objects by implementing advanced inpainting techniques such as PatchMatch, diffusion-based methods, and diffusion models.

[0049] Upon processing the extracted navigational image clip, the detecting module 208 may be configured to detect lane boundaries within the processed navigational image clip. In an embodiment, the detecting module 208 is configured to identify lane features from the processed navigational image clip, using an edge detection technique. The edge detection technique may include algorithms such as Canny Edge Detection, Sobel Filters, etc. to highlight contrast boundaries indicative of lane markings in the navigational image clip. Further, the detecting module 208 may cluster the identified lane features to extract lane polylines and convert the lane features into defined lane boundaries using clustering techniques such as Density-Based Spatial Clustering (DBSCAN) , Sobel operator, Roberts cross operator, Laplacian of Gaussian (LoG) , wavelet transform, etc. The detecting module 208 further clusters the lane features, reducing noise and identifying continuous patterns representing lane structures. Further, the detecting module 208 is configured to extract lane polylines from the clustered lane features, representing the geometric paths of the lanes in the navigational image clip. The polylines are used to map the approximate trajectory of each lane with high accuracy. In an embodiment, the detecting module 208 is configured to convert the extracted polylines into defined lane boundaries, that are standardized and formatted to represent the lane configuration visually and mathematically.

[0050] In an embodiment, the generating module 210 is configured to generate directional indicators within the processed navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The directional indicators provide visual cues to drivers, helping them navigate complex junctions or interchanges effectively. In an embodiment, the generating module 210 may determine directional indicators, such as arrows or lane-specific markers, that visually represent the intended driving directions. Further, the generating module 210 may determine the intended driving direction by processing user input related to a desired route. The user input may include the final destination or specific waypoints along the route, enabling the generating module 210 to compute the appropriate driving directions. The generating module 210 is further configured to dynamically update the directional indicators based on real-time changes in the user’s route or road conditions.

[0051] In an embodiment, the combining module 212 is configured to combine the processed navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image. In an embodiment, the combining module 212 is configured to receive the processed navigational image clip, which has undergone enhancements such as object removal, resolution improvement, or contrast adjustments to eliminate distractions and ensure visual clarity. The combining module 212 is further configured to overlay detected lane boundaries onto the processed navigational image clip. The lane boundaries are identified through advanced image processing techniques, such as edge detection or machine learning models, ensuring an accurate representation of lane positions and road layouts. Further, the combining module 212 is configured to incorporate directional indicators, such as arrows, turn markers, or lane-specific guidance. The directional indicators are dynamically generated based on routing data, traffic rules, or user-specific navigation instructions. In an embodiment, the combining module 212 ensures precise alignment of the processed navigational image clip, lane boundaries, and directional indicators to generate the navigational image view.

[0052] It should be noted that all such aforementioned modules 202 -212 may be represented as a single module or a combination of different modules. Further, as will be appreciated by those skilled in the art, each of the modules 202 -212 may reside, in whole or in parts, on one device or multiple devices in communication with each other. In some embodiments, each of the modules 202 -212 may be implemented as dedicated hardware circuit comprising custom application-specific integrated circuit (ASIC) or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Each of the modules 202 -212 may also be implemented in a programmable hardware device such as a field programmable gate array (FPGA) , programmable array logic, programmable logic device, and so forth. Alternatively, each of the modules 202 -212 may be implemented in software for execution by various types of processors (e.g., processor 104) . An identified module of executable code may, for instance, include one or more physical or logical blocks of computer instructions, which may, for instance, be organized as an object, procedure, function, or other construct. Nevertheless, the executables of an identified module or component need not be physically located together but may include disparate instructions stored in different locations which, when joined logically together, include the module and achieve the stated purpose of the module. Indeed, a module of executable code could be a single instruction, or many instructions, and may even be distributed over several different code segments, among different applications, and across several memory devices.

[0053] As will be appreciated by one skilled in the art, a variety of processes may be employed for generating the navigational view image from the street view image. For example, the exemplary system 100 and the associated computing device 102 may generate the navigational view image from the street view image by the processes discussed herein. In particular, as will be appreciated by those of ordinary skill in the art, control logic and / or automated routines for performing the techniques and steps described herein may be implemented by the system 100 and the computing device 102 either by hardware, software, or combinations of hardware and software. For example, suitable code may be accessed and executed by the one or more processors on the system 100 to perform some or all of the techniques described herein. Similarly, application specific integrated circuits (ASICs) configured to perform some, or all of the processes described herein may be included in the one or more processors on the system 100.

[0054] Referring now to FIG. 3, an exemplary process 300 for extracting the navigational image clip from the street view image is illustrated, in accordance with some embodiments. In an embodiment, the process 300 may include receiving a 360-degree street view image and the corresponding geographical coordinate of a specific location such as a complex junction or multi-lane interchange. In an embodiment, the street view image may be mapped in a list of nodes 302. The list of nodes 302 represents specific points of interest in a map, likely derived from an OpenStreetMap (OSM) data or another mapping source. Further, each node from the list of nodes 302 is mapped to corresponding geographic coordinates stored in a list of geos 304. The geographic coordinates such as latitude and longitude indicate the physical location of each node.

[0055] Further, the process 300 may include calculating a heading direction by analyzing the sequence of geographic coordinates. The heading direction represents a direction the user needs to face to approach the junction or multi-lane interchange. The heading direction 306 ensures that the street view image aligns with the user perspective when approaching the junction.

[0056] Further, the process 300 may include determining a central geographic coordinate 308. The central geographic coordinate 308 acts as the focal point around which the 360° street view image is queried. Further, the central geographic coordinate 308 serves as the reference for locating the optimal street view data. In an embodiment, the process 300 may include querying the nearest street view image is queried based on the central geographic coordinate 308 and the heading direction 306, that best matches the location and orientation.

[0057] In an embodiment, the process 300 may include adjusting the street view image to align with the calculated heading direction using image heading 310. Further, the process 300 may include extracting a navigational image clip of the street view image to fit the desired perspective ensuring that the navigational image clip displays a clear view of the junction, highlighting lanes, signage, and relevant navigational features using a geoImageCentre 312. The geoImageCentre 312 is a coordinate such as latitude and longitude that represents the middle or focal point of the geographic area of interest, such as an intersection, highway exit, or complex junction. Further, the geoImageCentre 312 serves as an anchor point for aligning the street view image to ensure the navigational image clip focuses on the most relevant part of the junction.

[0058] In an exemplary embodiment, if a driver is approaching a multi-lane intersection at latitude (for example, 1.2833) and longitude (for example, 103.8198) , the geoImageCentre may be set to the coordinates. The process then retrieves a 360°image of that area and clips it to provide a clear, aligned view showing which lane to follow or where to turn.

[0059] Referring now to FIGs. 4A -4E, the exemplary street view image and the navigational view image are illustrated in accordance with some embodiments. In an embodiment, FIG. 4A depicts a 360-degree street view image 400A centered at a junction. The street view image 400A may be captured via a 360-degree camera mounted on a vehicle or handheld by a user. The street view image 400A is tagged with corresponding geographical coordinates to ensure the tracking, authenticity, and identification of the street view image 400A.

[0060] FIG. 4B depicts a navigational image clip 400B centered at the junction. The navigational image clip 400B is extracted from the street view image 400A based on the geographical coordinates. The geographical coordinates may be stored in an OpenStreetMap (OSM) node sequence along with the street view image. The OpenStreetMap (OSM) node sequence may be an ordered list of nodes that define the shape or path of an OSM way, which represents linear or polygonal features on a map, such as roads, paths, or boundaries. The navigational image clip 400B may be extracted, upon calculating the heading direction of the user and center location to query a nearest proper street view image. In an embodiment, the navigational image clip 400B is a perspective of the 360-degree street view image 400A in a way that the navigational image clip 400B is centered at the junction, and each lane and lane boundary of the junction are visually distinct.

[0061] FIG. 4C depicts a navigational image clip 400C with the one or more objects in the frame. The plurality of objects are the distractions in the navigational image clip 400C such as a pedestrian, vehicle, obstruction, etc. The one or more objects are removed from the navigational image clip 400C to ensure precise and noise free processing of the navigational image clip 400C. In an embodiment, the pixels corresponding to the plurality of objects are detected and extracted from the navigational image clip 400C using an object detection and object segmentation model. Further, the pixels corresponding to the shadow are detected and extracted. Finally, the extracted pixels are replaced with a mask to regenerate image on those areas using an image generation model such as stable diffusion XL. In an embodiment, the masking or re-generation of the removed pixel may be automated in the object removal techniques. In an alternate embodiment, the object removal technique may process the one or more objects by removing the pixels corresponding to the objects and filling the removed pixels with the pixels as similar to the background.

[0062] FIG. 4D depicts a pixelated navigational image clip 402 and a vectorized navigational image clip 404. The pixelated navigational image clip 402 and the vectorized navigational image clip 404 are extracted from the navigational image clip 400C. The pixelated navigational image clip 402 is a rasterized or bitmap representation of the navigational image clip 400C. The pixelated navigational image clip 402 maintains the visual appearance of the navigational image clip 400C in terms of pixel data, allowing for the preservation of complex details such as textures, colors, and photographic elements. Further, the vectorized navigational image clip 404 is a processed version of the navigational image clip 400C where key elements, such as lane boundaries, road signs, or directional indicators, are converted into vector graphics. Vector graphics use mathematical expressions such as lines, curves, and shapes to represent features, enabling high precision and scalability without loss of detail. In an embodiment, the pixelated navigational image clip 402 and the vectorized navigational image clip 404 are used to detect lane boundaries of the navigational image clip 400C. To detect the lane boundaries, lane features are identified from the pixelated navigational image clip 402 and the vectorized navigational image clip 404, using an edge detection technique. Further, the identified lane features are clustered to extract lane polylines and convert the lane features into defined lane boundaries.

[0063] FIG. 4E depicts the navigational view image 400E for accurate and precise navigation from complex junctions and intersections. The navigational view image 400E is generated by combining the navigational image clip 400C with the detected lane boundaries. The dual representation of the navigational image clip 400C with the detected lane boundaries integrate realistic imagery with clear, abstract guidance, improving the overall usability and visual appeal of the navigational interface. The navigational view image 400E ensures that users receive accurate and easily interpretable driving instructions while maintaining a sense of realism in the displayed scene.

[0064] Referring now to FIG. 5A, an exemplary User Interface (UI) 500A of the existing navigation system is illustrated, in accordance with some embodiments. The UI 500A may be rendered on the user device 112 of the user via an application. The user may interact with the UI 500A to navigate a destination. The UI 500A may provide an interactive, visually intuitive navigation platform for a user to reach a specified destination.

[0065] FIG. 5A depicts a navigation page of the UI 500A. The navigation page may include an overview bar 502, a speed tad 504, a map 506, and a navigation tab 508. The overview bar 502 may enable the user to view the entire route plan and make informed navigation decisions. The overview bar 502 may render a specified destination based on the user input and an estimated time of arrival to provide a summary of the journey to the user. Further, the overview bar 502 may include an overview tab 510. The user may interact with the overview tab 510 to view the complete route to the destination in a turn-by-turn fashion, along with the estimated time to reach the turn.

[0066] Further, the speed tab 504 may render a speed limit of the current road along with the speed of the user, ensuring that the user follows the specified traffic rules. In an embodiment, the speed limit of the current road may be a predefined speed set by an authority or government. Further, the speed of the user may be calculated by measuring a relative change in the geographical location of the user over a period of time, by a Global Positioning System (GPS) .

[0067] In an embodiment, the map 506 includes visual representation of roads and the surroundings, offering an intuitive and real-time navigation experience to the user. The map 506 dynamically updates to represent the user's current position and movement in real-time, ensuring the displayed location is always relevant to the user. Further, the map 506 may include a route 512, configured to represent a path to the specified destination based on the user input. The route 512 may be highlighted on the map 506 to ensure precise and accurate distinction between the route 512 and other roads and paths. The route 512 may also include an arrow 514 representing nearest turns and lane change. Further, the map 506 may include a location icon 516 corresponding to the real-time location of the user with respect to the map 506. In an embodiment, the location icon 516 may include an arrow sign pointing towards a direction of motion of the user.

[0068] In an embodiment, the navigation tab 508 may render the distance to the nearest deflection in route, nearest location of deflection in route and arrow pointing towards the nearest deflection in route to the user, ensuring accurate route guidance. However, at the complex intersection the navigation tab 508 may not precisely point out the changes due to many overlapping roads such as flyover, underpass, a busy intersection or interchange due to the 2D view of the navigation tab 508 and the map 506.

[0069] Referring now to FIG. 5B, an exemplary User Interface (UI) 500B of the exemplary navigation system configured to generate the navigational view image 518 is illustrated, in accordance with some embodiments. The UI 500B may be rendered on the user device 112 of the user via an application. The user may interact with the UI 500B to navigate to a destination. The UI 500B may provide an interactive, visually intuitive navigation platform for a user to reach a specified destination.

[0070] FIG. 5B depicts a navigation page of the UI 500B. The navigation page may include an overview bar (analogous to the overview bar 502) , a map (analogous to the map 506) , and a navigation tab 516. In an embodiment, the navigation tab 516 may include a 3D render of a navigational view image 518 to overcome the shortcomings of the existing navigation system as explained in FIG. 5A. The navigational view image 518 may display the junction or intersection in a 3D model including each lane and turns of the intersection as explained in detail in FIGs. 1 -4E. Further, the navigation tab 516 may include a highlighted route clearly visualizing the turn and the lane leading to the specified destination based on the user inputs.

[0071] Referring now to FIG. 6, an exemplary process of a method 600 for generating the navigational view image from the street view image is depicted via a flow chart, in accordance with some embodiments. The method 600 is implemented by the computing device 102 of the system 100. At step 602, a street view image centered at a junction is received. The street view image is a 360-degree street view image. The street view image may be received from the existing navigation system, the database, a repository or captured by a camera mounted on the vehicle. The 360-degree street view image offers a complete panoramic view of the junction and the surrounding environment, ensuring that all relevant elements such as lane markings, road signs, and the configuration of the intersection are included in the street view image.

[0072] At step 604, a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence is extracted. To extract the navigational image clip, a heading direction and a center location is determined. The heading direction may be a direction the user is traveling towards the junction. Further, a heading angle at which a user is approaching the junction is calculated based on geographic coordinates of the location of the user relative to the junction. Finally, the navigational image clip is extracted based on the determined heading direction and the center location, by querying an associated database.

[0073] In an embodiment, a sequence of nodes from the OpenStreetMap is identified. Each node in the sequence represents a specific geographic coordinate, outlining the road or the junction layout. Further, the heading direction is determined by analyzing the sequence of geographic coordinates leading to the junction. The center location is determined by aligning the user perspective approaching the junction. Further, the user’s geographic coordinates are compared with the junction's geographic coordinates to determine a heading angle of the user. The heading angle ensures that the extracted navigational image aligns with the direction from which the user is approaching the junction. Once the heading direction and the centre location are determined, a segment of the street view image is extracted. In an embodiment, the street view image may be queried from the database indexed by the corresponding geographic coordinates. The extracted segment of the street view image represents the specific perspective the user may encounter while approaching the junction, including road elements such as lane markings, signage, and intersection details.

[0074] In an embodiment, when the navigational image clip includes one or more objects, the extracted navigational image clip is processed to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique. The one or more objects may include pedestrians and vehicles within the navigational image clip. The object detection and segmentation model may identify and segment objects within an image or video, for example Region-Based Convolutional Neural Network (R-CNN) , You Only Look Once (YOLO) , Single Shot MultiBox Detector (SSD) , Detection Transformer (DETR) , Pyramid Scene Parsing Network (PSPNet) , etc. Further, the Object removal techniques are methods used in image processing to eliminate unwanted objects or elements from an image while maintaining a natural appearance of the background, for example inpainting, texture synthesis, Deep Learning-Based Image Completion, Seam Carving, Clone Stamping, Stable Diffusion Models, etc.

[0075] In an embodiment, to process the extracted navigational image clip, one or more pixels corresponding to the one or more objects within the navigational image clip are identified using the object detection and segmentation model. Further, the identified pixels from the navigational image clip are eliminated using the object removal technique. In an embodiment, pixels of one or more shadows corresponding to the one or more objects within the navigational image clip are identified and extracted using the object detection and segmentation model. Further, the object removal technique is applied to remove pixels corresponding to the one or more detected shadows.

[0076] At step 606, one or more lane boundaries are detected within the navigational image clip. To detect the lane boundaries, lane features are identified from the processed navigational image clip, using an edge detection technique. The edge detection technique are image processing methods used to identify significant boundaries or transitions in images, typically where there are abrupt changes in intensity or color, for example Sobel Operator, Canny Edge Detection, Prewitt Operator, Laplacian of Gaussian (LoG) , Roberts Cross Operator, Kirsch Operator, Marr-Hildreth (Gaussian Laplace) , Deep Learning-Based Edge Detection, etc. Further, the identified lane features are clustered to extract lane polylines and converted into defined lane boundaries.

[0077] In an embodiment, noise or irrelevant visual artifacts that may interfere with lane boundary detection are removed by applying image smoothing filters such as Gaussian blur or median filters to reduce high-frequency noise while preserving edges. Further, the navigational image clip is converted into a binary format, where pixel values represent either lane features or the background. The edges are further identified within the binary image corresponding to lane features or boundaries by using the edge detection technique. Further, the identified edges are clustered into one or more meaningful clusters that represent individual lane polylines by using clustering algorithms, such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN) or Hough transform-based line detection, to organize edges into coherent groups. Each of the one or more meaningful cluster corresponds to a single lane or road boundary. Finally, the lane polylines are structured into lane boundaries that may be used for navigation purposes using geometric fitting techniques, such as polynomial regression or Bezier curve fitting.

[0078] At step 608, directional indicators are generated within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The directional indicators guide users by visually representing the intended driving direction based on the user's input. The user input may be a route or destination specified by the user. The user input may determine the intended destination or next waypoint, forming the basis for calculating the driving direction. Further, the intended driving direction may be calculated by combining the user input with the detected lane boundaries, ensuring that the directional indicator aligns with the user's intended path and the physical road structure. To generate directional indicators, the intended driving direction is determined based on user input regarding a desired route and the lane boundaries. In some embodiments, an action is computed based on the user input and lane boundaries. The action may be the directional indicators such as turn left, turn right, or go straight. The directional indicators may dynamically change based on changes in the user input in real-time such as rerouting, lane shift, and lane merge.

[0079] At step 610, the navigational image clip is combined with the detected lane boundaries and the directional indicators to generate the navigational view image. The navigational view image may be a visual representation that assists the user by providing clear and intuitive guidance at complex junctions or interchanges. The processed navigational image clip acts as the background for the navigational view image. Further, the detected lane boundaries are overlaid onto the navigational image clip, ensuring alignment with the actual road layout. Finally, directional indicators are then added to specify the appropriate paths for the user to follow based on the user input such as straight arrows for continuing ahead, curved arrows for turns, exit specific markers for multi-lane highways or interchanges, etc.

[0080] In an embodiment, the method 600 may further display the navigational view image to the user through a user interface (UI) to assist with navigation decisions at the junction. The navigational view image is displayed on the user device 112 via an application. Further, the method 600 may capture user feedback regarding the effectiveness of the navigational view image in aiding navigation decisions. The method 600 may further utilize the feedback to enhance subsequent navigational view image.

[0081] As will be also appreciated, the above-described techniques may take the form of computer or controller implemented processes and apparatuses for practicing those processes. The disclosure can also be embodied in the form of computer program code containing instructions embodied in tangible media, such as floppy diskettes, solid state drives, CD-ROMs, hard drives, or any other computer-readable storage medium, wherein, when the computer program code is loaded into and executed by a computer or controller, the computer becomes an apparatus for practicing the invention. The disclosure may also be embodied in the form of computer program code or signal, for example, whether stored in a storage medium, loaded into and / or executed by a computer or controller, or transmitted over some transmission medium, such as over electrical wiring or cabling, through fiber optics, or via electromagnetic radiation, wherein, when the computer program code is loaded into and executed by a computer, the computer becomes an apparatus for practicing the invention. When implemented on a general-purpose microprocessor, the computer program code segments configure the microprocessor to create specific logic circuits.

[0082] The disclosed methods and systems may be implemented on a conventional or a general-purpose computer system, such as a personal computer (PC) or server computer. Referring now to FIG. 7, an exemplary computing system 700 that may be employed to implement processing functionality for various embodiments (e.g., as a SIMD device, client device, server device, one or more processors, or the like) is illustrated. Those skilled in the relevant art will also recognize how to implement the invention using other computer systems or architectures. The computing system 700 may represent, for example, a user device such as a desktop, a laptop, a mobile phone, personal entertainment device, DVR, and so on, or any other type of special or general-purpose computing device as may be desirable or appropriate for a given application or environment. The computing system 700 may include one or more processors, such as a processor 702 that may be implemented using a general or special purpose processing engine such as, for example, a microprocessor, microcontroller or other control logic. In this example, the processor 702 is connected to a bus 704 or other communication medium. In some embodiments, the processor 702 may be an Artificial Intelligence (AI) processor, which may be implemented as a Tensor Processing Unit (TPU) , or a graphical processor unit, or a custom programmable solution Field-Programmable Gate Array (FPGA) .

[0083] The computing system 700 may also include a memory 706 (main memory) , for example, Random Access Memory (RAM) or other dynamic memory, for storing information and instructions to be executed by the processor 702. The memory 706 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by the processor 702. The computing system 700 may likewise include a read only memory ( “ROM” ) or other static storage device coupled to bus 704 for storing static information and instructions for the processor 702.

[0084] The computing system 700 may also include storage devices 708, which may include, for example, a media drive 710 and a removable storage interface. The media drive 710 may include a drive or other mechanism to support fixed or removable storage media, such as a hard disk drive, a floppy disk drive, a magnetic tape drive, an SD card port, a USB port, a micro-USB, an optical disk drive, a CD or DVD drive (R or RW) , or other removable or fixed media drive. A storage media 712 may include, for example, a hard disk, magnetic tape, flash drive, or other fixed or removable medium that is read by and written to by the media drive 710. As these examples illustrate, the storage media 712 may include a computer-readable storage medium having stored there in particular computer software or data.

[0085] In alternative embodiments, the storage devices 708 may include other similar instrumentalities for allowing computer programs or other instructions or data to be loaded into the computing system 700. Such instrumentalities may include, for example, a removable storage unit 714 and a storage unit interface 716, such as a program cartridge and cartridge interface, a removable memory (for example, a flash memory or other removable memory module) and memory slot, and other removable storage units and interfaces that allow software and data to be transferred from the removable storage unit 714 to the computing system 700.

[0086] The computing system 700 may also include a communications interface 718. The communications interface 718 may be used to allow software and data to be transferred between the computing system 700 and external devices. Examples of the communications interface 718 may include a network interface (such as an Ethernet or other NIC card) , a communications port (for example, a USB port, a micro-USB port) , Near field Communication (NFC) , etc. Software and data transferred via the communications interface 718 are in the form of signals which may be electronic, electromagnetic, optical, or other signals capable of being received by the communications interface 718. These signals are provided to the communications interface 718 via a channel 720. The channel 720 may carry signals and may be implemented using a wireless medium, wire or cable, fiber optics, or another communications medium. Some examples of the channel 720 may include a phone line, a cellular phone link, an RF link, a Bluetooth link, a network interface, a local or wide area network, and other communications channels.

[0087] The computing system 700 may include Input / Output (I / O) devices 722. Examples may include, but are not limited to a display, keypad, microphone, audio speakers, vibrating motor, LED lights, etc. The I / O devices 722 may receive input from a user and also display an output of the computation performed by the processor 702. In this document, the terms “computer program product” and “computer-readable medium” may be used generally to refer to media such as, for example, the memory 706, the storage devices 708, the removable storage unit 714, or signal (s) on the channel 720. These and other forms of computer-readable media may be involved in providing one or more sequences of one or more instructions to the processor 702 for execution. Such instructions, generally referred to as “computer program code” (which may be grouped in the form of computer programs or other groupings) , when executed, enable the computing system 700 to perform features or functions of embodiments of the present invention.

[0088] In an embodiment where the elements are implemented using software, the software may be stored in a computer-readable medium and loaded into the computing system 700 using, for example, the removable storage unit 714, the media drive 710 or the communications interface 718. The control logic (in this example, software instructions or computer program code) , when executed by the processor 702, causes the processor 702 to perform the functions of the invention as described herein.

[0089] Thus, the disclosed method and system try to overcome the technical problem of generating a navigational view image from a street view image. The method and system eliminate the need for manual 3D modeling, which is time-consuming and costly. By leveraging automated image processing and machine learning (ML) models, the system significantly reduces the effort and time required to produce junction views at large scale. Further, the system and the method automate processes such as object detection, removal, and lane mapping, enabling large-scale generation of junction views with minimal human intervention and reducing production costs compared to traditional manual methods. By using high-resolution street view imagery and advanced image processing, the system and the method ensure precise rendering of lane markings, road signs, and directional arrows enhancing the usability and effectiveness of navigation systems. Further, the method and the system support customization based on specific intersections and geographic locations ensuring that the junction views are tailored to the user’s context. The method and system further provide clear and intuitive visualizations of complex junctions helping drivers make informed decisions, reducing confusion and stress at intersections or highway interchanges. Additionally, the method and the system use object segmentation and removal algorithms to ensure that distractions, such as pedestrians and vehicles, are eliminated from the imagery. Further, the use of object segmentation and removal algorithms make the junction view focused and clean, improving its interpretability. The method and system are capable of integrating real-world street view imagery into junction views while retaining background, enabling seamless integration with the existing navigation systems. In an aspect, the present disclosure offers a highly efficient, scalable, and user-friendly approach for generating junction views from street view imagery, addressing the limitations of existing 2D map views and improving navigation clarity and usability.

[0090] As will be appreciated by those skilled in the art, the techniques described in the various embodiments discussed above are not routine, or conventional, or well understood in the art. The techniques discussed above provide for generating the navigational view image from the street view image. The techniques first receive the street view image cantered at a junction. The techniques then extract a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence. The navigational image clip includes one or more objects. The techniques then process the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique. Further, the techniques detect lane boundaries within the processed navigational image clip. The techniques further generate directional indicators within the processed navigational image clip to represent intended driving directions, based on a user input and the lane boundaries. The techniques then combine the processed navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.

[0091] In light of the above-mentioned advantages and the technical advancements provided by the disclosed method and system, the claimed steps as discussed above are not routine, conventional, or well understood in the art, as the claimed steps enable the following solutions to the existing problems in conventional technologies. Further, the claimed steps clearly bring an improvement in the functioning of the system itself as the claimed steps provide a technical solution to a technical problem.

[0092] The specification has described a method and a system for generating the navigational view image from the street view image. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art (s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments.

[0093] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor (s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., be non-transitory. Examples include random access memory (RAM) , read-only memory (ROM) , volatile memory, nonvolatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0094] It is intended that the disclosure and examples be considered as exemplary only, with a true scope and spirit of disclosed embodiments being indicated by the following claims.

Claims

1.A method of generating a navigational view image from a street view image, the method comprising:receiving the street view image cantered at a junction;extracting a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence;detecting lane boundaries within the navigational image clip;generating directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries; andcombining the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.2.The method of claim 1, further comprising:when the navigational image clip comprises one or more objects:processing the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique, and wherein the one or more objects comprise pedestrians and vehicles within the navigational image clip.3.The method of claim 2, wherein processing the extracted navigational image clip further comprises:identifying pixels corresponding to the one or more objects within the navigational image clip using the object detection and segmentation model; andeliminating the identified pixels from the navigational image clip using the object removal technique.4.The method of claim 1, wherein detecting lane boundaries comprises:identifying lane features from the navigational image clip, using an edge detection technique; andclustering the identified lane features to extract lane polylines and converting the lane features into defined lane boundaries.5.The method of claim 1, further comprising displaying the navigational view image to a user through a user interface to assist with navigation decisions at the junction.6.The method of claim 1, wherein processing the navigational image clip further comprises:detecting shadows in navigational image clip; andapplying an object removal technique to remove pixels corresponding to the detected shadows.7.The method of claim 1, wherein generating directional indicators comprises determining the intended driving direction based on a user input regarding a desired route and the lane boundaries.8.The method of claim 1, further comprising:capturing user feedback regarding effectiveness of the navigational view image in aiding navigation decisions; andutilizing the feedback to enhance subsequent image generation processes.9.The method of claim 1, wherein the street view image is a 360-degree street view image.10.The method of claim 1, wherein extracting the navigational image clip further comprisesdetermining a heading direction and a center location, wherein a heading angle at which a user is approaching the junction is calculated, and wherein the heading angle is calculated based on geographic coordinates of a location of the user relative to the junction; andextracting the navigational image clip, based on the determined heading direction and the center location, by querying an associated database.11.A system for generating a navigational view image from a street view image, the system comprising:a processor; anda memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, cause the processor to:receive the street view image cantered at a junction;extract a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence;detect lane boundaries within the navigational image clip;generate directional indicators within the navigational image clip to represent intended driving directions, based on a user input and the lane boundaries; andcombine the navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.12.The system of claim 11, wherein the processor-executable instructions further cause the processor to:when the navigational image clip comprises one or more objects:process the extracted navigational image clip to remove the one or more objects from the navigational image clip, using an object detection and segmentation model, and an object removal technique, wherein the one or more objects comprise pedestrians and vehicles within the navigational image clip, and wherein the street view image is a 360-degree street view image.13.The system of claim 12, wherein the processor-executable instructions further cause the processor to process the extracted navigational image clip by:identifying pixels corresponding to the one or more objects within the navigational image clip using the object detection and segmentation model; andeliminating the identified pixels from the navigational image clip using the object removal technique.14.The system of claim 11, wherein the processor-executable instructions further cause the processor to detect lane boundaries by:identifying lane features from the navigational image clip, using an edge detection technique; andclustering the identified lane features to extract lane polylines and converting the lane features into defined lane boundaries.15.The system of claim 11, wherein the processor-executable instructions further cause the processor to display the navigational view image to a user through a user interface to assist with navigation decisions at the junction.16.The system of claim 11, wherein the processor-executable instructions further cause the processor to process the navigational image clip by:detecting shadows in navigational image clip; andapplying an object removal technique to remove pixels corresponding to the detected shadows.17.The system of claim 11, wherein the processor-executable instructions further cause the processor to generate directional indicators by determining the intended driving direction based on a user input regarding a desired route and the lane boundaries.18.The system of claim 11, wherein the processor-executable instructions further cause the processor to:capture user feedback regarding effectiveness of the navigational view image in aiding navigation decisions; andutilize the feedback to enhance subsequent image generation processes.19.The system of claim 11, wherein the processor-executable instructions further cause the processor to extract the navigational image clip by:determining a heading direction and a center location, wherein a heading angle at which a user is approaching the junction is calculated, and wherein the heading angle is calculated based on geographic coordinates of a location of the user relative to the junction; andextracting the navigational image clip, based on the determined heading direction and the center location, by querying an associated database.20.A non-transitory computer-readable medium storing computer-executable instructions for generating a navigational view image from a street view image, the computer-executable instructions configured for:receiving the street view image cantered at a junction;extracting a navigational image clip from the received street view image based on geographic coordinates derived from an OpenStreetMap (OSM) node sequence;detecting lane boundaries within the processed navigational image clip;generating directional indicators within the processed navigational image clip to represent intended driving directions, based on a user input and the lane boundaries; andcombining the processed navigational image clip with the detected lane boundaries, and the directional indicators to generate the navigational view image.