Lane indication system and method using a bird's-eye view map
The system uses infrastructure cameras to convert images to a bird's-eye view, generating segmentation maps and predicting lane occupancy, improving lane guidance and safety at merging points by overcoming visibility obstacles.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2025-10-09
- Publication Date
- 2026-05-13
AI Technical Summary
Drivers face confusion in complex road configurations, particularly at merging locations, due to obstructed visibility from obstacles and lighting conditions, which impairs their ability to make informed lane merging decisions.
A system utilizing infrastructure cameras to capture images, convert them to a bird's-eye view, generate static and dynamic segmentation maps, and predict lane occupancy, providing real-time lane guidance through indicator lights.
Enhances driving clarity by accurately indicating lane occupancy, reducing confusion and enhancing safety at merging points on complex roads.
Smart Images

Figure 2026077585000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to systems and methods for geospatial mapping and advanced driving assistance, and more particularly to systems and methods for improving vehicle driving assistance using real-time lane guidance.
Background Art
[0002] The complexity of road configurations can cause driver confusion and can impede the ability to make appropriate driving judgments, particularly in critical areas such as merging locations. In such situations, a driver may struggle to understand the lane availability, future traffic patterns, or how to merge into the moving traffic in a desired manner. Therefore, there is a need for systems and methods for real-time lane guidance that enable drivers to make timely driving judgments based on information.
Summary of the Invention
[0003] In one embodiment, a method for detecting one or more approaching vehicles includes converting one or more images of the surroundings of a host vehicle captured by an infrastructure camera into a surrounding bird's-eye view, the surroundings including a ramp, one or more lanes of a highway, and one or more approaching vehicles, the ramp and the lanes sharing at least one merging location; generating a static segmentation map and one or more dynamic objects in the static segmentation map based on the bird's-eye view; predicting the lane occupancy of the ramp and the lanes based on range estimation of the host vehicle and the one or more dynamic objects; and transmitting information about the predicted lane occupancy to the host vehicle.
[0004] In another embodiment, a system and method for detecting one or more approaching vehicles includes one or more processors and infrastructure cameras capable of capturing one or more images of the surroundings of the vehicle itself. The surroundings include a ramp, one or more lanes of a highway, and one or more approaching vehicles, wherein the ramp and lanes share at least one merging point. One or more processors capable of converting one or more images of the surroundings into a bird's-eye view of the surroundings, generating a static segmentation map and one or more dynamic objects in the static segmentation map based on the bird's-eye view, predicting lane occupancy of the ramp and lanes based on range estimation of the vehicle itself and one or more dynamic objects, and transmitting information about the predicted lane occupancy to the vehicle itself.
[0005] These and additional features provided by embodiments of this disclosure will be better understood by referring to the detailed description below in conjunction with the drawings. [Brief explanation of the drawing]
[0006] The embodiments shown in the drawings are for illustrative purposes only and are not intended to limit the disclosure. The subsequent detailed description of the exemplary embodiments can be understood when read in conjunction with the following drawings, in which similar structures are indicated by the same reference numbers. [Figure 1A] Figure 1A schematically illustrates an example of an improved lane indication system with infrastructure support using a bird's-eye view map for roads and ramps as described herein, according to one or more embodiments shown and described herein. [Figure 1B] Figure 1B schematically illustrates an example of an improved lane indication system with infrastructure support using a bird's-eye view map for roads and ramps as described herein, according to one or more embodiments shown and described herein. [Figure 2]Figure 2 schematically shows an example of components of an improved lane indication system with infrastructure support using a bird's-eye view map of this disclosure, according to one or more embodiments shown and described herein. [Figure 3] Figure 3 shows a block diagram of the improved lane indicator generation of this disclosure according to one or more embodiments shown and described herein. [Figure 4] Figure 4 schematically shows an example of a segmentation map of the present disclosure according to one or more embodiments shown and described herein. [Figure 5] Figure 5 shows a flowchart illustrating an exemplary process for improved lane indication with infrastructure assistance using the bird's-eye view map of this disclosure, according to one or more embodiments shown and described herein. [Modes for carrying out the invention]
[0007] Embodiments disclosed herein include systems and methods for improved lane indication and approaching vehicles with infrastructure assistance using bird's-eye view maps. The disclosed system uses one or more infrastructure cameras 208 to capture images of the surroundings around highway and ramp merging points. Highways and / or ramps may include one or more lanes. The system can convert the images into a bird's-eye view of the surroundings and, based on the bird's-eye view, can further generate a static segmentation map and one or more dynamic objects in the static segmentation map. The system can then predict lane occupancy of ramps and lanes based on range estimation and further provide lane occupancy information to vehicles of interest.
[0008] The visibility of drivers and vehicles on ramps may be intermittently obstructed or affected by various obstacles such as trees, road complexity (e.g., interchanges, elevated ramps), and lighting conditions (e.g., sunlight, streetlights). These factors can impair a driver's ability to clearly see other vehicles and make informed decisions, particularly regarding lane merging. For example, a driver may have difficulty identifying which lanes on the main road or ramp are occupied by approaching vehicles, making it difficult to make a desired merging decision.
[0009] Consider the scenario shown in Figure 1B, which has a complex road configuration featuring multiple ramps and highways. In this situation, both a vehicle on a lower ramp and another vehicle exiting a highway on an elevated ramp may be approaching the same merging point at the ramp connection. The driver of the vehicle on the lower ramp may be unable to detect the vehicle on the elevated ramp due to the ramp structure itself and visual obstacles such as sunlight. Furthermore, the presence of multiple vehicles on the main road may confuse the driver about whether the vehicle on the elevated ramp is still traveling on the highway or actively approaching the ramp. Similarly, in the scenario shown in Figure 1A, the driver of a vehicle traveling to merge from a ramp onto a highway may be confused about whether the vehicle approaching the merging point is in the right lane or the left lane.
[0010] The disclosed systems and methods address these challenges by utilizing improved lane indications supported by infrastructure-assisted technologies and bird's-eye mapping. Infrastructure cameras monitor lane occupancy and capture real-time images of lanes without interference from obstacles such as ramps or lights. This allows for more complete and accurate information on lane usage. Dynamic displays of lane activity are generated through image processing technologies, including bird's-eye mapping and segmentation, enabling drivers to better understand which lanes are occupied and make informed merging decisions. As a result, these systems and methods can significantly improve the driving experience on complex roads by providing clear lane indication information, reducing confusion, and enhancing the overall driving experience for vehicles traveling on ramps.
[0011] As used herein, the singular forms "a," "an," and "the" include multiple reference objects unless the context indicates otherwise. Thus, for example, a reference to "a component" includes two or more components, such as "components," unless the context indicates otherwise. Wherever possible, the same reference number is used throughout the drawings to refer to the same or similar parts.
[0012] Referring to the drawings, Figures 1A and 1B show an example of an improved lane indicating system 100 for improved lane indicating with infrastructure support using a bird's-eye view map of the present disclosure. The improved lane indicating system 100 may include an infrastructure camera 208 and lane indicator lights 105. The infrastructure camera 208 may capture one or more images of the surroundings 150 of an ego vehicle 101. The surroundings 150 may include, for example, a ramp 131, one or more lanes 121 of a highway 120, and one or more non-ego vehicles 103. The ramp 131 and lanes 121 may share at least one merging point 153. In some embodiments, the vehicle 101 may travel on the ramp 131 and merge onto the highway 120 at the merging point 153 when one or more other vehicles 103 are traveling on the highway 120, thereby causing potential interference and / or collision between the other vehicles 103 and the vehicle 101 at the merging point 153 (for example, as shown in Figure 1A). In some embodiments, one or more other vehicles 103 may exit the highway 120 and merge onto the ramp 131b or the road at the merging point 153 when the vehicle 101 is traveling on the ramp 131a or the road, thereby causing potential interference and / or collision at the merging point 153 (for example, as shown in Figure 1B).
[0013] In some embodiments, the infrastructure camera 208 may include, but is not limited to, a camera, proximity sensor, light detection and ranging (LIDAR) sensor, thermal imaging sensor, infrared sensor, ultrasonic sensor, and / or a combination thereof. The camera may be, but is not limited to, a red, green, and blue (RGB) camera, depth camera, infrared camera, wide-angle camera, or stereo camera. The infrastructure camera 208 may be any device having an arrangement of sensing devices capable of detecting radiation in the ultraviolet wavelength band, visible light wavelength band, or infrared wavelength band. The infrastructure camera 208 may have any resolution. In some embodiments, one or more optical components, such as a mirror, fisheye lens, or any other type of lens, may be optically connected to the infrastructure camera 208. The infrastructure camera 208 may be positioned around a merging point 153 to capture images of the surroundings 150. In some embodiments, but is not limited to, the infrastructure camera 208 may be mounted on fixed structures such as traffic light poles and lane indicator lights 105. In some embodiments, the infrastructure camera 208 may be mounted on a mobile object such as a drone, but is not limited to these embodiments. The images may include a perspective view of the surroundings, such as a two-dimensional (2D) representation.
[0014] In some embodiments, the improved lane indication system 100 may include one or more modules, such as a vision transformer module 222, a segmentation map module 232, and a lane occupancy module 242. The vision transformer module 222 can convert a perspective view into a bird's-eye view perception using a three-dimensional (3D) representation. The segmentation map module 232 may use bird's-eye view perception to generate a static segmentation map representing static objects in the surroundings 150 and to detect one or more dynamic objects, such as vehicles (e.g., the vehicle itself 101 and other vehicles 103), pedestrians, or other moving objects. The lane occupancy module 242 may determine and / or predict occupancy based on range estimations of the vehicle itself 101 and dynamic objects for each lane 121 of the highway 120 and ramp 131. In some embodiments, the occupancy state may include occupied and unoccupied states and may transmit information about the lane. In some embodiments, the improved lane indication system 100 may act on its own vehicle 101 and / or at least one other vehicle 103 to avoid potential interference / collision at merging points 153.
[0015] The vision transformer module 222, the segmentation map module 232, and / or the lane occupancy module 242 may include one or more machine learning (ML) algorithms. One or more vehicle modules may be pre-trained using range estimation and lane occupancy training data, such training data may include ground truth examples and scenarios, such scenarios may involve multiple entities (e.g., one or more self-vehicles 101, multiple other vehicles 103, and other objects) moving on a road, ramp, or other surface, taking into account the positions of one or more central entities and other entities, the operating conditions of the entities (e.g., speed, direction, acceleration, the entities' reactions to other entities), the distance between entities, and factors (e.g., environment, weather, road conditions, etc., but not limited to these). Pre-training may include labeling entities and lanes and ramps with desired lane occupancy prediction results in the examples and scenarios, and using one or more ML models to be trained to predict desired and undesired lane occupancy prediction results based on the training data. Pre-training may further include fine-tuning, evaluation, and testing phases. One or more modules may be continuously trained using real-world data to adapt to changing conditions and factors and to improve performance over time. One or more modules may be continuously trained during the operation of the improved lane indication system 100 using collected and generated data, such as historical static segmentation maps 227 and / or historical bird's-eye view images 237 (as shown in Figure 2).
[0016] In some embodiments, the improved lane indicator system 100 may include lane indicator lights 105. The lane indicator lights 105 may include a plurality of indicator lights, each of which represents a lane 121 on the highway 120 and / or ramp 131. Each indicator light may indicate the occupancy status of the corresponding lane. An occupancy status may refer to a state in which at least one vehicle 101 or 103 is approaching the area around the merging point 153, based on the results of analysis from the lane occupancy module 242, such as range estimation. It should be understood that when vehicles 101, 103 are moving away from the merging point 153 or passing the merging point 153 on one of the lanes 121, the lane indicator lights 105 may change the corresponding indicator light to a color or pattern that suggests that the corresponding lane 121 is unoccupied. The arrangement of the indicator lights may follow the same order as the arrangement of lanes on the highway 120 and / or ramp 131. For example, as shown in Figure 1A, the left indicator light of the lane indicator light 105 may indicate the occupancy status of the left lane 121a, the center indicator light of the lane indicator light 105 may indicate the occupancy status of the right lane 121b, and the right indicator light of the lane indicator light 105 may indicate the occupancy status of the ramp 131. The lane indicator lights 105 may be positioned above the highway 120 and / or ramp 131 near the merging point 153 so that the drivers of their own vehicle 101 and / or other vehicles 103 can see the occupancy status.
[0017] In one embodiment, one or more vehicles 101, 103 may travel on a highway 120 and / or ramp 131. Vehicles 101, 103 may be automobiles or any other passenger or non-passenger vehicles, such as land, water, and / or air vehicles. Each vehicle may be an autonomous or semi-autonomous vehicle that navigates its environment with or without limited human input. Vehicles 101, 103 may travel on roads and perform vision-based lane centering, for example, using sensors. Vehicles 101, 103 may include actuators for driving the vehicle, such as motors, engines, or any other powertrain. Vehicles 101, 103 may travel on a variety of surfaces, such as roads, highways, streets, expressways, bridges, tunnels, parking lots, garages, off-road trails, railroads, or any surface on which a vehicle can operate, but are not limited to these. As shown in Figures 1A and 1B, the vehicle 101 may travel on ramp 131, which is a non-limiting example of merging onto another road or ramp 131. In some embodiments, it should be understood that the vehicle 101 and / or other vehicles 103 may travel on a highway 120 that includes multiple lanes 121. For example, highway 120 may include a left lane 121a and a right lane 121b heading north. Another vehicle 103a may travel on the left lane 121a that does not interact with the merging point 153, and another vehicle 103b may travel on the right lane 121b that leads to the merging point. Vehicles 101 and 103 traveling on highway 120 may travel within a lane or change lanes to another lane traveling in the same direction. It should be understood that each vehicle 101 and 103 may be considered as the vehicle 101, and the other vehicles may be considered as other vehicles 103. Therefore, as applicable through this disclosure, the descriptions relating to our own vehicle 101 may apply to other vehicles 103.
[0018] The vehicle 101 may include one or more proximity sensors and steering sensors to acquire data for autonomous or semi-autonomous navigation and operation of the vehicle 101. The proximity sensors and steering sensors may be used to collect and generate environmental data and steering data, such environmental data and steering data may include, for example, the time difference and / or distance difference between the vehicle 101 and objects in its surroundings 150, the acceleration of the vehicle 101, the speed of the vehicle 101 and the speed of other vehicles 103, the current position of the vehicle 101, contextual information such as weather information, the type of road the vehicle 101 is traveling on, the surface conditions of the highway 120 and ramps 131 on which the vehicle 101 is traveling, and the degree of traffic on the highway 120 on which the vehicle 101 is traveling. Environmental data includes weather conditions (e.g., sunny, rainy, snowy, or foggy), road conditions (e.g., dry, wet, or icy road surface), traffic conditions, road infrastructure, obstacles (e.g., other vehicles 103 or pedestrians), lighting conditions, geographical features of highway 120, and other environmental conditions relating to driving on highway 120 and ramps 131.
[0019] In some embodiments, one or more proximity sensors of the vehicle 101 may include, but are not limited to, cameras, light detection and ranging (LIDAR) sensors, thermal imaging sensors, infrared sensors, ultrasonic sensors, and / or combinations thereof. Cameras may include, but are not limited to, red, green, and blue (RGB) cameras, depth cameras, infrared cameras, wide-angle cameras, or stereoscopic cameras. One or more proximity sensors may be any device having an arrangement of sensing devices capable of detecting radiation in the ultraviolet wavelength band, the visible light wavelength band, or the infrared wavelength band. One or more proximity sensors may have any resolution. In some embodiments, one or more optical components, such as mirrors, fisheye lenses, or any other type of lens, may be optically coupled to one or more proximity sensors. In some embodiments, one or more vehicle steering sensors of the vehicle 101 may include one or more speed sensors or motion sensors to detect and measure the motion of the vehicle 101 and changes in motion. Motion sensors may include inertial measurement units. Each of the one or more motion sensors may include one or more accelerometers and one or more gyroscopes. Each of the one or more motion sensors converts the recognized physical motion of the vehicle into a signal indicating the vehicle's direction, rotation, velocity, or acceleration. Data obtained from the vehicle steering sensor may be used to determine the vehicle kinematics of the vehicle 101. Therefore, the vehicle steering sensor may be used to collect and generate vehicle control data and vehicle kinematic data. Vehicle control data may include throttle position, brake position, steering angle, and gear selection. Vehicle kinematic data may include velocity, acceleration, position, and orientation.
[0020] Figure 2 schematically shows an example of the components of the improved lane indicator system 100. The improved lane indicator system 100 may include one or more processors 204. Each of the one or more processors 204 may be any device capable of executing machine-readable and executable instructions. The instructions may be in the form of a machine-readable instruction set stored in the data storage component 207 and / or the memory component 202. Thus, each of the one or more processors 204 may be a controller, integrated circuit, microchip, computer, or any other computing device. The one or more processors 204 are connected to a communication path 203 that provides signal interconnection between various modules of the system. Thus, the communication path 203 may connect any number of processors 204 to each other in a communicative manner, and the modules to which the communication path 203 is connected may operate in a distributed computing environment. In particular, each module may operate as a node capable of transmitting and / or receiving data. As used herein, the term “communicatively connected” means that connected components are capable of exchanging data with each other, such as electrical signals over a conductive medium, electromagnetic signals over air, optical signals over an optical waveguide, and so on.
[0021] Therefore, the communication path 203 can be formed from any medium capable of transmitting signals, such as conductive wires, conductive traces, and optical waves. In some embodiments, the communication path 203 can facilitate the transmission of wireless signals, such as Wi-Fi, Bluetooth®, and Near Field Communication (NFC). Furthermore, the communication path 203 can be formed from a combination of mediums capable of transmitting signals. In one embodiment, the communication path 203 comprises a combination of conductive traces, conductive wires, connectors, and a bus that cooperate to enable the transmission of electrical data signals to components such as processors, memory, sensors, input devices, output devices, and communication devices. Thus, the communication path 203 may comprise a vehicle bus, such as a LIN bus, CAN bus, or VAN bus. In addition, the term “signal” means a waveform (electrical, optical, magnetic, mechanical, or electromagnetic), such as DC, AC, sine wave, triangular wave, square wave, or vibration, that can propagate through a medium.
[0022] The improved lane indication system 100 may include one or more memory components 202 connected to a communication path 203. The one or more memory components 202 may comprise RAM, ROM, flash memory, a hard drive, or any device capable of storing machine-readable and executable instructions so that the machine-readable and executable instructions can be accessed by one or more processors 204. The machine-readable and executable instructions may comprise logic or algorithms described in any generation of any programming language (e.g., 1GL, 2GL, 3GL, 4GL, or 5GL), such as machine language directly executable by a processor, or assembly language, object-oriented programming (OOP), scripting language, microcode, etc., that has been compiled or assembled into machine-readable and executable instructions and stored in the one or more memory components 202. Alternatively, the machine-readable and executable instructions may be described in a hardware description language (HDL), such as logic implemented via either a field-programmable gate array (FPGA) configuration or an application-specific integrated circuit (ASIC), or an equivalent thereof. Thus, the methods described herein may be implemented in any conventional computer programming language, as a pre-programmed hardware element, or as a combination of hardware and software components. One or more processors 204 along the one or more memory components 202 may operate as a controller for the improved lane indication system 100.
[0023] One or more memory components 202 may include a vision transformer module 222, a segmentation map module 232, and a lane occupancy module 242. Each module 222, 232, and 242 may include, but are not limited to, routines, subroutines, programs, objects, components, data structures, etc., to perform specific tasks or execute specific data types, as described below. The data storage component 207 stores historical static segmentation maps 227, historical bird's-eye view images 237, sensor-generated data, and data about the operation of the lane indicator lights 105 and infrastructure cameras 208. The vision transformer module 222, the segmentation map module 232, and the lane occupancy module 242 may also be stored in the data storage component 207 during or after operation. Each module may include one or more machine learning algorithms. The vehicle module and server module may be trained via neural networks and given machine learning capabilities, as described herein. For example, but not limited to, a neural network may utilize one or more artificial neural networks (ANNs). In an ANN, the connections between nodes may form a directed acyclic graph (DAG). An ANN may include node inputs, one or more hidden activation layers, and node outputs, and one or more hidden activation layers may utilize activation functions such as linear functions, step functions, logistic (sigmoid) functions, hyperpolyc tangent (tanh) functions, rectified linear unit (ReLU) functions, or combinations thereof. An ANN is trained by applying such activation functions to a training dataset to determine an optimal solution from adjustable weights and biases applied to the nodes in the hidden activation layers, and to produce one or more outputs with minimized error as that optimal solution. In machine learning applications, new inputs (such as one or more generated outputs) may be given to the ANN model as training data to continuously improve the accuracy of the ANN model and minimize error.One or more ANN models may utilize one-to-one, one-to-many, many-to-one, and / or many-to-many (e.g., sequence-to-sequence) sequence modeling. One or more ANN models may utilize combinations of artificial intelligence techniques, such as, but are not limited to, deep learning, random forest classifiers, feature extraction from speech or images, clustering algorithms, or combinations thereof. In some embodiments, convolutional neural networks (CNNs) may be used. For example, in the field of machine learning, a convolutional neural network (CNN) may be used as a class of deep, feed-forward ANNs, applied, for example, to speech analysis of recorded material. CNNs may be shift-invariant or space-invariant and may utilize weight-sharing architectures and translation. Furthermore, each of the various modules may include generative artificial intelligence algorithms. Generative artificial intelligence algorithms may include generative adversarial networks (GANs) having two networks: a generative model and a discriminative model. Furthermore, generative artificial intelligence algorithms can be based on variational autoencoders (VAEs) or transformer-based models.
[0024] Continuing to refer to FIG. 2, the improved lane indication system 100 may include one or more infrastructure cameras 208. The one or more infrastructure cameras 208 may include options such as, but not limited to, proximity sensors, cameras, light detection and ranging (LIDAR) sensors, thermal image sensors, infrared sensors, ultrasonic sensors, and / or combinations thereof. The camera may be, but not limited to, a red, green, and blue (RGB) camera, a depth camera, an infrared camera, a wide-angle camera, or a stereo camera. The one or more infrastructure cameras 208 may be any device having an arrangement of detection devices capable of detecting radiation in the ultraviolet wavelength band, visible light wavelength band, or infrared wavelength band. The one or more infrastructure cameras 208 may have any resolution. In some embodiments, one or more optical components such as mirrors, fisheye lenses, or any other type of lens may be optically connected to the one or more infrastructure cameras 208. In the embodiments described herein, the one or more infrastructure cameras 208 may provide image data to another component communicatively connected to the one or more processors 204 or communication path 203. In some embodiments, the one or more infrastructure cameras 208 may also provide navigation support. That is, the data obtained by the one or more infrastructure cameras 208 may be used to navigate the vehicle automatically or semi-automatically.
[0025] The improved lane indicator system 100 may include network interface hardware 206 to enable communication of the improved lane indicator system 100 to vehicles 101, 103 and / or a server. The network interface hardware 206 may be any device that is communicatively connected to the communication path 203 and capable of transmitting and / or receiving data over the network. Thus, the network interface hardware 206 may include communication transceivers for transmitting and / or receiving any wired or wireless communication. For example, the network interface hardware 206 may include antennas, modems, LAN ports, WiFi cards, WiMAX cards, mobile communication hardware, near-field communication hardware, satellite communication hardware, and / or any wired or wireless hardware for communicating with other networks and / or devices. In one embodiment, the network interface hardware 206 includes hardware configured to operate according to the Bluetooth® wireless communication protocol.
[0026] The improved lane indication system 100 may include an infrastructure camera 208. The infrastructure camera 208 may be communicably connected to a communication path 203. The infrastructure camera 208 may include one or more image sensors configured to operate within the visible light spectrum and / or infrared spectrum in order to detect visible light and / or infrared light. In addition, while the specific embodiments described herein describe hardware for detecting light within the visible light spectrum and / or infrared spectrum, it should be understood that other types of sensors are also assumed. For example, the systems described herein may include one or more LIDAR sensors, radar sensors, sonar sensors, or other types of sensors for collecting data that can be integrated into or capture such data collections described herein. Rangefinders such as radar may be used to obtain coarse depth and velocity information about a surrounding viewpoint 150.
[0027] The improved lane indicator system 100 may include one or more lane indicator lights 105. The lane indicator lights 105 may be communicably connected to a communication path 203. The lane indicator lights 105 may include a plurality of indicator lights, each of which represents a lane 121 on the highway 120 and / or ramp 131. The indicator lights may include, but are not limited to, incandescent bulbs, light-emitting diode (LED) lights, and electroluminescent lights. The indicator lights may be monochromatic or polychromatic. The lane indicator lights 105 may use shape, color, or lighting pattern (e.g., always on and flashing) to indicate the occupancy status of the corresponding lane and ramp. The lane indicator lights 105 may further include digital display signs, traffic control signals, and other components / devices relating to lane indication.
[0028] Figure 3 shows an example block diagram of generating the improved lane indication of the present disclosure. The improved lane indication system 100 may use one or more infrastructure cameras 208 to capture one or more images of the surrounding area 150 near the merging point 153 and / or the vehicle 101, such images may include images and / or video frames of the surrounding area 150.
[0029] In block 301, the vision transformer module 222 (Figure 2) transforms the image from a 2D perspective view to a 3D bird's-eye view. The vision transformer module 222 can estimate depth information from the image, reconstruct the surrounding scene in 3D, and then change the viewpoint to a top-down view (i.e., a bird's-eye view). Depth estimation may be based on monocular depth prediction to estimate the distance of each pixel in the image. In some embodiments, depth information and data may be estimated from images generated by a single infrastructure camera. In some embodiments, depth information and data may be generated based on two or more infrastructure cameras 208 located at different positions. In some embodiments, depth information and data may be provided from a distance sensor such as a LIDAR sensor. The vision transformer module 222 then creates a depth map with a 3D point cloud, and each pixel in 2D may be projected into 3D space. The vision transformer module 222 can then perform a geometric transformation in block 303 to view the surrounding scenery from above (i.e., from a bird's-eye view) and generate a 3D bird's-eye view perception.
[0030] In blocks 305 and 307, the segmentation map module 232 may perform segmentation mapping on the generated bird's-eye view recognition image to generate one or more segmentation images 400 (e.g., Figure 4). The segmentation images may include the background 401 (e.g., in Figure 4), static objects 403 (e.g., in Figure 4), and dynamic objects 405 (e.g., in Figure 4) in the surrounding 150. In block 305, the segmentation map module 232 may perform semantic segmentation mapping on the bird's-eye view recognition to generate a static segmentation map by leveraging the bird's-eye view recognition. The segmentation map module 232 may classify each pixel in the bird's-eye view recognition into a category based on semantic features such as color or location. The segmentation map module 232 may compare the current bird's-eye view recognition with past bird's-eye view recognition and / or existing free-space maps of the surrounding area 150 in order to determine whether an object is a static object 403 or a dynamic object 405. The segmentation map module 232 may generate a surrounding free-space map if such a map does not exist, and / or continuously update the free-space map using multiple acquired images. The segmentation map module 232 may include one or more ML algorithms, such as deep learning models, such as Mask R-CNN, DeepLab, and U-Net. The ML algorithm may segment the bird's-eye view image into meaningful classes, such as, but not limited to, roads, sidewalks, buildings, vehicles, vegetation, and objects and structures around the confluence point 153. After segmentation, static objects 403, such as the background 401 and road surfaces and road boundaries, may be extracted, sorted, and / or classified (e.g., free space, roads, obstacles, etc.). The segmentation map module 232 can assign different values and / or colors to each class in block 305 and generate a static segmentation map (e.g., a semantic segmentation map).It should be understood that in some embodiments, other segmentation techniques, though not limited to these, may be used to generate static segmentation maps, such as instance segmentation, panoptic segmentation, depth segmentation, superpixel segmentation, region-based segmentation, and edge-based segmentation.
[0031] In some embodiments, the segmentation map module 232 may determine and classify the background 401 and static objects 403 as lanes 121, lane markings, and other relevant features (e.g., road boundaries and curbs). The segmentation map module 232 may generate a bindery mask with pixels corresponding to lane markings and lane boundaries, and classify regions in the segmentation map as lane regions and non-lane regions. The segmentation map module 232 may label lanes 121 and ramps 131 respectively as lane 121a, lane 121b (e.g., in Figure 1A), ramp 131a, and ramp 131b (e.g., in Figure 1B). The segmentation map module 232 may label regions representing static objects 403 and road structures (e.g., road boundaries). A labeled area may include one or more background areas 401, one or more lane mark areas, and one or more road boundary areas. In some embodiments, if lane marks are not found, the segmentation map module 232 may estimate the number of lanes and the lane width to determine the lane area for each lane 121.
[0032] In block 307, the segmentation map module 232 may further detect, classify, and track dynamic objects such as vehicles 101, 103, pedestrians, cyclists, motorcycles, and / or other moving objects in bird's-eye view recognition. The segmentation map module 232 may detect dynamic objects by comparing the bird's-eye view map with a static segmentation map generated in block 305. The segmentation map module 232 may include a neural network (e.g., YOLOv5) to generate a set of objects and object detections. In some embodiments, the segmentation map module 232 may assign one or more bounding boxes to dynamic objects detected in bird's-eye view recognition, each of which may be described with a bounding box, a class (e.g., a vehicle approaching merging point 153, a vehicle passing through merging point 153, etc.), and a detection confidence score ranging from 0 to 1.
[0033] In block 309, the lane occupancy module 242 can perform lane occupancy prediction based on the static segmentation map generated in block 305 and the detection and tracking of dynamic objects generated in block 307. The lane occupancy module 242 identifies and divides lanes 121 and ramps 131 in the bird's-eye view recognition, recognizes pixels corresponding to lanes 121 and ramps 131, and creates one or more binary lane marks that highlight lane boundaries. The lane occupancy module 242 may have a convolutional neural network (CNN) along with an encoder, decoder, and backbone (e.g., UNetFormer). The lane occupancy module 242 can extract pixels from the bird's-eye view recognition and assign each pixel to a corresponding label in the static segmentation map, thereby indicating whether the pixel belongs to a particular lane / ramp. Based on the corresponding lane mask, the lane occupancy module 242 may assign a lane index to each of the detected dynamic objects 405. For example, the improved lane indication system 100 can determine whether another vehicle 103 is traveling on one lane 121b in Figure 1A or on one ramp 131b in Figure 1B and approaching a merging point 153, thereby potentially causing interference with the vehicle 101, or whether the other vehicle 103 has passed the merging point 153 or is on lane 121 (e.g., left lane 121a) and does not cause a potential collision.
[0034] In block 311, the lane occupancy module 242 may output lane indications. The lane occupancy module 242 may predict lane occupancy of the ramp 131 and lane 121 based on range estimation of the vehicle 101 and one or more dynamic objects 405. The improved lane indication system 100 may transmit information about the predicted lane occupancy to the vehicle 101. In some embodiments, the improved lane indication system 100 may display lane occupancy on the lane indicator light 105.
[0035] In some embodiments, the lane occupancy module 242 may predict the speed of one or more dynamic objects 405 and the relative distance between the vehicle 101 and one or more dynamic objects 405, based on the current segmentation map and one or more past segmentation maps. Based on the lane occupancy, the speed of one or more dynamic objects 405, and the relative distance between the vehicle 101 and one or more dynamic objects 405, the lane occupancy module 242 may predict the probability of collision between the vehicle 101 and each of the one or more dynamic objects 405. The lane occupancy module 242 may then determine whether the collision probability exceeds a threshold probability. The threshold probability may be a preset value or a value determined based on past collision events and / or historical near collision events near the merging point 153. Depending on whether the collision probability is determined to have exceeded the threshold, the lane occupancy module 242 may warn the vehicle 101 of a potential collision with at least one of the one or more dynamic objects 405. In some embodiments, upon determining that the probability of collision exceeds a threshold, the improved lane indication system 100 may cause its vehicle to act in order to avoid a potential collision.
[0036] Figure 5 shows a descriptive example of Method 500 for an improved lane indication system with infrastructure support using a bird's-eye view map. In block 501, Method 500 includes converting one or more images of the surrounding area 150 (in Figures 1A and 1B) of the vehicle 101 (in Figures 1A and 1B) captured by an infrastructure camera 208 (in Figures 1A and 1B) into a bird's-eye view of the surrounding area 150. The surrounding area 150 may include a ramp 131 (in Figures 1A and 1B), one or more lanes 121 (in Figures 1A and 1B) of the highway 120, and one or more other vehicles 103 (in Figures 1A and 1B). Other vehicles 103 may include approaching vehicles. The ramp 131 and lanes 121 may share at least one merging point 153 (in Figures 1A and 1B). In block 502, method 500 may include generating a static segmentation map and one or more dynamic objects 405 (for example, in Figure 4) in the static segmentation map based on a bird's-eye view. In block 503, method 500 includes predicting lane occupancy of ramp 131 and lane 121 based on range estimation of the vehicle 101 and one or more dynamic objects 405. In block 504, method 500 includes transmitting information about the predicted lane occupancy to the vehicle 101.
[0037] In some embodiments, the static segmentation map may include labeled areas displaying static objects 403 and structures of the road. The static segmentation map may include one or more backgrounds 401, one or more lane mark areas, and one or more road boundary areas. Dynamic objects 405 may represent moving objects in the surrounding area 150. Moving objects may include moving vehicles 101, 103 and pedestrians on the highway 120 and ramp 131. An infrastructure camera 208 may be located at or near the merging point 153. The static segmentation map and dynamic objects may be generated based on past static segmentation maps of the surrounding area.
[0038] In some embodiments, method 500 may further include displaying lane occupancy on lane indicator lights 105. In some embodiments, method 500 may further include predicting the speed of movement of one or more dynamic objects 405 and the relative distance between the vehicle 101 and the one or more dynamic objects 405. In some embodiments, method 500 may further include predicting the probability of collision between the vehicle 101 and each of the one or more dynamic objects 405 based on lane occupancy, the speed of movement of one or more dynamic objects 405 and the relative distance between the vehicle 101 and the one or more dynamic objects 405. In some embodiments, method 500 may further include determining whether the collision probability exceeds a threshold probability and, in response to determining that the collision probability exceeds a threshold, warning the vehicle 101 of a potential collision with at least one of the one or more dynamic objects 405. In some embodiments, method 500 may further include causing the vehicle to act to avoid a potential collision in response to determining that the collision probability exceeds a threshold.
[0039] Furthermore, the terms “substantially” and “about” may be used herein to express the degree of inherent uncertainty that may arise from any quantitative comparison, value, measurement, or other expression. These terms may also be used herein to express the extent to which a quantitative expression may deviate from a described standard without resulting in a change in the fundamental function of the subject matter in question.
[0040] While specific embodiments have been described and documented herein, it should be understood that various other changes and modifications can be made without departing from the spirit and scope of the claimed subject matter. Furthermore, while various aspects of the claimed subject matter have been described herein, such aspects do not necessarily need to be used in combination. Accordingly, the appended claims are intended to encompass all such changes and modifications within the scope of the claimed subject matter.
Claims
1. A method for detecting one or more approaching vehicles, Converting one or more images of the area around the vehicle captured by an infrastructure camera into a bird's-eye view of the surrounding area, wherein the surrounding area includes a ramp, one or more lanes of a highway, and one or more approaching vehicles, and the ramp and the lanes share at least one merging point. Based on the aforementioned bird's-eye view, a static segmentation map and one or more dynamic objects in the static segmentation map are generated. Based on the range estimation of the vehicle itself and the one or more dynamic objects, predict the lane occupancy of the ramp and the lane, The information regarding the predicted lane occupancy is transmitted to the vehicle, Methods that include...
2. A method according to claim 1, wherein the static segmentation map comprises labeled regions representing static objects and structures of a road.
3. A method according to claim 2, wherein the static segmentation map comprises one or more background regions, one or more lane mark regions, and one or more road boundary regions.
4. A method according to claim 1, wherein the dynamic object represents a moving object in the surrounding area, and the moving object includes vehicles and pedestrians moving on the highway and the ramp.
5. A method according to claim 1, further comprising indicating lane occupancy on a lane indicator light.
6. The method according to claim 1, wherein the infrastructure camera is located at or near the confluence point.
7. A method according to claim 1, wherein the static segmentation map and the dynamic objects are generated based on the surrounding past static segmentation maps.
8. A method according to claim 1, further comprising predicting the speed of movement of one or more dynamic objects and the relative distance between the vehicle and the one or more dynamic objects.
9. A method according to claim 8, further comprising predicting the probability of collision between the vehicle and each of the one or more dynamic objects based on the lane occupancy, the speed of movement of the one or more dynamic objects, and the relative distance between the vehicle and the one or more dynamic objects.
10. The method according to claim 9, wherein the method is To determine whether the aforementioned collision probability exceeds a threshold probability, In response to determining that the collision probability exceeds the threshold probability, the vehicle is warned of a potential collision with at least one of the one or more moving objects. Methods that further include the above.
11. A method according to claim 10, further comprising, in response to determining that the probability of collision exceeds the threshold probability, causing the vehicle to act in a manner that avoids the potential collision.
12. A system for detecting one or more approaching vehicles, wherein the system is An infrastructure camera capable of capturing one or more images of the surroundings of its own vehicle, wherein the surroundings include a ramp, one or more lanes of a highway, and one or more approaching vehicles, and the ramp and the lanes share at least one merging point, One or more processors, Convert one or more of the surrounding images into a bird's-eye view of the surroundings, Based on the aforementioned bird's-eye view, a static segmentation map and one or more dynamic objects in the static segmentation map are generated. Based on the range estimation of the vehicle itself and the one or more dynamic objects, the lane occupancy of the ramp and the lane is predicted. The information regarding the predicted lane occupancy is transmitted to the vehicle. One or more processors capable of operating in this manner, A system that includes these features.
13. A system according to claim 12, wherein the static segmentation map comprises labeled areas representing static objects and structures of a road, one or more background areas, one or more lane mark areas, and one or more road boundary areas.
14. A system according to claim 12, wherein the dynamic object represents a moving object in the surroundings, and the moving object includes vehicles and pedestrians moving on the highway and the ramp.
15. A system according to claim 12, wherein one or more processors are further operable to display the lane occupancy on a lane indicator light.
16. A system according to claim 12, wherein the infrastructure camera is located at or near the confluence point.
17. A system according to claim 12, wherein the static segmentation map and the dynamic objects are generated based on the surrounding past static segmentation maps.
18. A system according to claim 12, wherein one or more processors are further operable to predict the speed of movement of the one or more dynamic objects and the relative distance between the vehicle and the one or more dynamic objects.
19. The system according to claim 18, wherein the one or more processors are Based on the lane occupancy, the speed of movement of the one or more moving objects, and the relative distance between the vehicle and the one or more moving objects, the probability of collision between the vehicle and each of the one or more moving objects is predicted. Determine whether the aforementioned collision probability exceeds the threshold probability. In response to determining that the collision probability exceeds the threshold probability, the vehicle is warned of a potential collision with at least one of the one or more moving objects. A system that is even more operational.
20. A system according to claim 19, wherein one or more processors are further operable to cause the vehicle to act to avoid the potential collision in response to determining that the collision probability exceeds the threshold probability.