Method and system for topology detection
The system employs deep learning and neural networks to detect and model road topology in real-time, addressing the challenge of accurate lane recognition for autonomous vehicles, enhancing safety and efficiency in dynamic environments.
Patent Information
- Application Number
- JP2025511473
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-23
- Filing Date
- 2023-08-23
- Publication Date
- 2025-09-09
AI Technical Summary
Autonomous vehicles face challenges in accurately and efficiently determining road topology for safe navigation, including lane recognition and object avoidance, due to reliance on sensors that require precise environmental understanding.
A system and method utilizing a combination of sensors and deep learning techniques, including neural networks, to detect and model road topology in real-time by identifying lane components and landmarks, enabling accurate lane representation and adaptive navigation.
Enhances the safety and efficiency of autonomous vehicle operation by providing a robust and adaptable framework for lane detection and navigation, reducing reliance on static maps and improving handling of dynamic road environments.
Smart Images

Figure 2025529870000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to methods and systems for topology detection. [Background technology]
[0002] An autonomous vehicle is a motor vehicle capable of performing one or more required driving functions without human driver input, and generally includes Level 2 or higher capabilities as generally described in SAE International's J3016 standard, and in certain embodiments, includes a self-driving truck that includes sensors, devices, and systems that may function together to generate sensor data indicative of various parameter values related to the vehicle's position, speed, operating characteristics, and state, including data generated in response to various objects, situations, and environments encountered by the autonomous vehicle during operation.
[0003] Autonomous vehicles may rely on sensors such as cameras, lidar, radar, and inertial measurement units (IMUs) to understand the road and the rest of the world around the vehicle without requiring user interaction. For example, accurate understanding and modeling of the road on which an autonomous vehicle operates is important so that the vehicle can safely navigate the road using sensor readings (i.e., sensor data). Accurate modeling or estimation of the road can be important for perception (computer vision), control, mapping, and other functions. Without a true understanding of the local environment in which it is operating, an autonomous vehicle may have additional problems such as staying within its lane, as well as steering, navigation, and object avoidance.
[0004] Therefore, a need exists for an efficient and robust system and method for accurately and efficiently determining or modeling road topology for the operation of autonomous vehicles. [Brief explanation of the drawings]
[0005] The features and advantages of the illustrative embodiments, and the manner in which they are achieved, will become more readily apparent from the following detailed description taken in conjunction with the accompanying drawings.
[0006] [Figure 1] FIG. 1 is an illustrative block diagram of a control system that may be deployed in a vehicle, according to an exemplary embodiment. [Figure 2A] 1 is an illustrative depiction of an exterior view of a semi-truck in accordance with an illustrative embodiment; [Figure 2B] 1 is an illustrative depiction of an exterior view of a semi-truck in accordance with an illustrative embodiment; [Figure 2C] 1 is an illustrative depiction of an exterior view of a semi-truck in accordance with an illustrative embodiment; [Figure 3] 1 is an illustrative depiction of a road including multiple lanes in which an autonomous vehicle may operate, in accordance with an illustrative embodiment; [Figure 4] 1 is an illustrative depiction of a set of defined lane geometries in accordance with an illustrative embodiment; [Figure 5A] 1 is an illustrative depiction of road lanes with each lane represented by N ordered points, according to an exemplary embodiment; [Figure 5B] 1 is an illustrative depiction of road lanes with each lane represented by N ordered points, according to an exemplary embodiment; [Figure 5C] 1 is an illustrative depiction of road lanes with each lane represented by N ordered points, according to an exemplary embodiment; [Figure 6A] 1 is an illustrative example of a lane represented by a combination of lane components at N ordered points, according to an exemplary embodiment. [Figure 6B] 1 is an illustrative example of a lane represented by a combination of lane components at N ordered points, according to an exemplary embodiment. [Figure 6C] 1 is an illustrative example of a lane represented by a combination of lane components at N ordered points, according to an exemplary embodiment. [Figure 6D]1 is an illustrative example of a lane represented by a combination of lane components at N ordered points, according to an exemplary embodiment. [Figure 7] FIG. 1 is an illustrative block diagram of a system in accordance with an exemplary embodiment. [Figure 8] FIG. 1 is an illustrative flow diagram of a process in accordance with an exemplary embodiment; [Figure 9] FIG. 1 is an illustrative block diagram of a computing system in accordance with an exemplary embodiment.
[0007] Throughout the drawings and detailed description, the same drawing reference numbers should be understood to refer to the same elements, features, and structures unless otherwise stated. The relative size and depiction of these elements may be exaggerated or adjusted for clarity, illustration, and / or convenience. DETAILED DESCRIPTION OF THE INVENTION
[0008] In the following description, specific details are set forth to provide a thorough understanding of various exemplary embodiments. It should be understood that various modifications to the embodiments will be readily apparent to those skilled in the art, and that the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the present disclosure. Moreover, in the following description, numerous details are set forth for purposes of explanation. However, those skilled in the art should understand that the embodiments may be practiced without the use of these specific details. In other instances, well-known structures and processes are not shown or described so as not to obscure the description with unnecessary detail. Thus, the present disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0009] For convenience and ease of description, certain terminology is used herein. For example, the term "semi-truck" is used to refer to a vehicle in which the system of the exemplary embodiments may be used. The terms "semi-truck," "truck," "tractor," "vehicle," and "semi" may be used interchangeably herein. Furthermore, as will be apparent to those skilled in the art upon reading this disclosure, embodiments of the present invention may be used in conjunction with other types of vehicles. In general, embodiments may be used with desirable results in conjunction with any vehicle that tows a trailer or carries cargo over long distances.
[0010] 1 illustrates a control system 100 that may be deployed, according to an example embodiment, with an autonomous vehicle (AV), such as, but not limited to, a semi-truck 200 depicted in FIGS. 2A-2C. Referring to FIG. 1, control system 100 may include sensors 110 that collect data and information that are provided to computer system 140 to perform actions, including, for example, control actions to control vehicle components via gateway 180. According to some embodiments, gateway 180 is configured to allow computer system 140 to control different components from different manufacturers.
[0011] Computer system 140, in conjunction with one or more central processing units (CPUs) 142, may be configured to perform processing, including processing for implementing features of embodiments of the present invention described elsewhere herein, and to receive sensor data from sensors 110 for use in generating control signals to control one or more actuators or other controllers associated with systems of the vehicle in which control system 100 is deployed (e.g., actuators or controllers that enable control of throttle 184, steering system 186, brakes 188, and / or other devices and systems). Generally, control system 100 may be configured to operate a vehicle (e.g., semi-truck 200) in an autonomous (or semi-autonomous) mode of operation.
[0012] For example, the control system 100 may be operated to capture images from one or more cameras 112 mounted at various locations on the semi-truck 200 and perform processing (e.g., image processing) on the captured images to identify objects proximate to or in the path of the semi-truck 200. In some aspects, one or more lidar 114 and radar 116 sensors may be positioned on the vehicle to sense or detect the presence and volume of objects proximate to or in the path of the semi-truck 200. Other sensors may also be positioned or mounted at various locations on the semi-truck 200 to capture other information, such as position data. For example, the sensors may include one or more satellite positioning sensors, such as the GNSS / IMU 118, and / or an inertial navigation system. The Global Navigation Satellite System (GNSS) is a space-based satellite system that provides location information (longitude, latitude, altitude) and time information anywhere on or near the Earth, in all weather conditions, to devices called GNSS receivers. GPS is the most used GNSS system in the world and may be used interchangeably herein with GNSS. An inertial measurement unit ("IMU") is an inertial navigation system. Generally, an inertial navigation system ("INS") measures and integrates the orientation, position, velocity, and acceleration of a moving object. The INS integrates the measured data, and the GNSS is used as a correction for integration errors in the INS orientation calculation. Any number of different types of GNSS / IMU 118 sensors may be used in conjunction with the features of the present invention.
[0013] Data collected by each of the sensors 110 may be processed by the computer system 140 to generate control signals that may be used to control the operation of the semi-truck 200. For example, image and location information may be processed to identify or detect objects around or within the path of the semi-truck 200, and control signals may be sent via the controller 182 to adjust the throttle 184, steering 186, and / or brakes 188 as needed to safely operate the semi-truck 200 in an autonomous or semi-autonomous manner. While illustrative example sensors, actuators, and other vehicle systems and devices are shown in FIG. 1 , those skilled in the art will understand, upon reading this disclosure, that other sensors, actuators, and systems may also be included in the system 100 consistent with the present disclosure. For example, in some embodiments, actuators may also be provided that provide mechanisms that allow for control of the transmission of the vehicle (semi-truck 200).
[0014] Control system 100 may include a computer system 140 (e.g., a computer server) configured to provide a computing environment in which one or more software, firmware, and control applications (e.g., items 160-182) may execute to perform at least a portion of the processes described herein. In some embodiments, computer system 140 includes components deployed on-board the vehicle (e.g., deployed in a system rack 240 positioned within a semi-truck berth 212, as shown in FIG. 2C). Computer system 140 may communicate with other computer systems (not shown) that may be local to semi-truck 200 and / or remote from semi-truck 200 (e.g., computer system 140 may communicate with one or more remote ground-based or cloud-based computer systems via a wireless communication network connection).
[0015] According to various embodiments described herein, computer system 140 may be implemented as a server. In some embodiments, computer system 140 may be configured using any of several computing systems, environments, and / or configurations, such as, but not limited to, a personal computer system, a cloud platform, a server computer system, a thin client, a thick client, a handheld or laptop device, a tablet, a smartphone, a database, a multiprocessor system, a microprocessor-based system, a set-top box, a programmable consumer electronics product, a network PC, a minicomputer system, a mainframe computer system, a distributed cloud computing environment, or the like, and may include any of the above systems or devices.
[0016] Different software applications or components may be executed by the computer system 140 and the control system 100. For example, as shown in active learning component 160, an application that performs active learning machine processing may be provided to process images captured by one or more cameras 112 and information obtained by the lidar 114. For example, image data may be processed using a deep learning segmentation model 162 to identify objects of interest in the captured images (e.g., other vehicles, construction signs, etc.). In some aspects herein, deep learning segmentation may be used to identify lane points in a lidar scan. As an example, the system may use an intensity-based voxel filter to identify lane points in a lidar scan. The lidar data may be processed by a machine learning application 164 to draw or identify bounding boxes on the image data to identify objects located by the lidar sensor.
[0017] Information output from the machine learning applications may be provided as input to object fusion 168 and vision map fusion 170 software components, which may perform processing to predict the behavior of other road users, fuse local vehicle pose with global map geometry in real time, and enable on-the-fly map correction. Output from the machine learning applications may be supplemented with information (as well as positioning data) from radar 116 and map localization 166 application data. In some aspects, these applications enable control system 100 to be less map-dependent and more capable of handling constantly changing road environments. Furthermore, by correcting any map errors on the fly, control system 100 may facilitate safer, more scalable, and more efficient operation compared to alternative map-centric approaches.
[0018] The information is provided to a prediction and planning application 172, which provides input to a trajectory planning 174 component that enables a trajectory to be generated in real time by a trajectory generation system 176 based on interactions and predicted interactions between the semi-truck 200 and other associated vehicles within the truck's operating environment. In some embodiments, for example, the control system 100 generates a 60-second planning period and analyzes the relevant actors and available trajectories. The plan that best meets multiple criteria (including safety, comfort, and route preference) may be selected, and any associated control inputs necessary to implement the plan are provided to a controller 182 to control the movement of the semi-truck 200.
[0019] In some embodiments, these disclosed applications or components (as well as other components or flows described herein) may be implemented in hardware, a computer program executed by a processor, firmware, or a combination of the above, unless otherwise specified. In some cases, the computer program may be embodied on a computer-readable medium, such as a storage medium or storage device. For example, the computer program, code, or instructions may reside in random access memory (“RAM”), flash memory, read-only memory (“ROM”), erasable programmable read-only memory (“EPROM”), electrically erasable programmable read-only memory (“EEPROM”), registers, a hard disk, a removable disk, a compact disk read-only memory (“CD-ROM”), or any other form of non-transitory storage medium known in the art.
[0020] The non-transitory storage medium may be coupled to the processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit ("ASIC"). In alternative embodiments, the processor and the storage medium may reside as separate components. For example, FIG. 1 illustrates an exemplary computer system 140 that may represent or be integrated with any of the components disclosed below, etc. Accordingly, FIG. 1 is not intended to suggest any limitation as to the scope of use or functionality of embodiments of the systems and methods disclosed herein. The computer system 140 may implement and / or perform any of the functions disclosed herein.
[0021] Computer system 140 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system 140 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be located in both local and remote computer system storage media, including non-transitory memory storage devices.
[0022] 1 , computer system 140 is shown in the form of a general-purpose computing device. Components of computer system 140 may include, but are not limited to, one or more processors (e.g., CPU 142 and GPU 144), a communications interface 146, one or more input / output interfaces 148, and one or more storage devices 150. Although not shown, computer system 140 may also include a system bus that couples various system components, including system memory, to CPU 142. In some embodiments, input / output (I / O) interface 148 may also include a network interface. For example, in some embodiments, some or all of the components of control system 100 may communicate via a controller area network ("CAN") bus or the like that interconnects various components within the vehicle in which control system 100 is deployed and associated.
[0023] In some embodiments, storage device 150 may include various types and forms of non-transitory computer-readable media. Such media may be any available media accessible by a computer system / server and may include both volatile and nonvolatile media, removable and non-removable media. In one embodiment, system memory implements processes represented by flow diagrams in other figures herein. System memory may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. As another example, storage device 150 may read from and write to non-removable, non-volatile magnetic media (not shown, typically referred to as a "hard drive"). Although not shown, storage device 150 may include one or more removable, non-volatile disk drives, such as magnetic, tape, or optical disk drives. In such an example, each may be connected to a bus by one or more data media interfaces. Storage device 150 may include at least one program product having program modules, codes, and / or sets of instructions (e.g., at least one) configured to perform the functions of various embodiments of an application.
[0024] 2A-2C are illustrative depictions of the exterior of a semi-truck 200 that may be associated with or used in accordance with exemplary embodiments. The semi-truck 200 is shown for illustrative purposes only. Accordingly, those skilled in the art will understand, upon reading this disclosure, that embodiments may be used in conjunction with several different types of vehicles and are not limited to the type of vehicle illustrated in FIGS. 2A-2C. The exemplary semi-truck 200 shown in FIGS. 2A-2C is one style of truck configuration common in North America, including an engine 206, a steering axle 214, and a drive axle 216 forward of a cab 202. A trailer (not shown) may be attached to the semi-truck 200 via a fifth-wheel trailer coupling, typically provided on a frame 218 and positioned above the drive axle 216. A sleeping compartment 212 may be positioned behind the cab 202, as shown in FIGS. 2A and 2C. FIGS. 2A-2C further illustrate several sensors positioned in different locations on the semi-truck 200. For example, one or more sensors may be mounted in a sensor rack 220 on the roof of the cab 202. Sensors may also be mounted on the side mirrors 210 and elsewhere on the semi-truck. Sensors may be mounted on the bumper 204, the side of the cab 202, or elsewhere. For example, rear-facing radar 236 is shown as being mounted on the side of the cab 202 in FIG. 2A . Embodiments may be used in other configurations of trucks and other vehicles (e.g., semi-trucks having cab-over or cab-forward configurations, etc.). Generally, and without limiting the embodiments of the present disclosure, features of the present invention may be used with desirable results in vehicles that haul cargo over long distances, such as long-distance semi-truck routes.
[0025] FIG. 2B is a front view of semi-truck 200 illustrating several sensors and sensor locations. A sensor rack 220 may secure and position several sensors on windshield 208, including long-range lidar 222, long-range camera 224, GPS antenna 234, and medium-range forward-facing camera 226. Side mirrors 210 may provide mounting locations for rear-facing camera 228 and medium-range lidar 230. Front radar 232 may be mounted on bumper 204. Other sensors (including some shown and some not shown) may be mounted or installed elsewhere on semi-truck 200. As such, the locations and mounting depicted in FIGS. 2A-2C are for illustrative purposes only.
[0026] 2C, there is shown a partial view of semi-truck 200 depicting some aspects of the interior of cab 202 and sleeping compartment 212. In some embodiments, portions of control system 100 of FIG. 1 may be deployed in a system rack 240 within sleeping compartment 212, allowing easy access to components of control system 100 for maintenance and operation.
[0027] Certain aspects of the present disclosure relate to methods and systems that provide a framework or architecture for generating an accurate representation of a road's topology, including one or more lanes, in real time as an autonomous vehicle AV (e.g., a truck similar to those disclosed in FIGS. 1 and 2A-2C) is operating (e.g., being driven). Aspects of the present disclosure generally provide a framework for accurately and efficiently determining the topology of lanes of a road near an AV, including, for example, different lane configurations.
[0028] As used herein, road topology refers to a map of the structure or configuration of a road, and some embodiments of the present disclosure may be interested in the topology of the road in front of a target AV (i.e., the ego vehicle).
[0029] FIG. 3 is an illustrative depiction of a road including multiple lanes on which an AV (or other) vehicle may operate, according to an exemplary embodiment. FIG. 3 illustrates road 300 including multiple lanes for accommodating vehicular traffic. In the example of FIG. 3, five lanes of traffic are depicted, including lanes 305, 310, 315, 320, and 325, with the direction of traffic flow as indicated therein. In FIG. 3, as in several other figures herein, the road may be represented as a single line, and the representation of the single line may correspond to the centerline of the traffic lane (i.e., the average between the boundaries of the traffic lane). In some instances herein, traffic lanes may also be referred to simply as lanes. Lanes of a road may have different configurations. For example, road 300 includes different configurations of lanes, where lanes 305, 310, and 315 may each be characterized as straight-through lanes, and lane 320 may be characterized as a branching lane because it branches or splits into another lane 325 at 330 (e.g., an exit or off-ramp).
[0030] In some aspects, lanes may be characterized as having particular lane geometries (i.e., structural shapes or configurations), as demonstrated in part by the example lanes of FIG. 3 . In some embodiments herein, lane geometries for road lanes may be represented by a finite set of predefined different shapes. FIG. 4 is an illustrative line-drawing depiction of a set 400 of predefined lane geometries that may be used in some embodiments herein to represent a variety of different lane configurations. The example set of predefined lane geometries 400 includes seven different lane geometries. With reference to the traffic flow direction shown in FIG. 4 , the seven lane geometries of set 400 include a straight lane geometry 405, a merging left lane geometry 410, a diverging right lane geometry 415, a merging left lane followed by a diverging right lane geometry 420, a merging right lane geometry 425, a diverging left lane geometry 430, and a merging right lane followed by a diverging left lane geometry 435.
[0031] In some embodiments, the set of predefined lane geometries 400 may be used, alone or in combination, to describe or otherwise represent substantially all road configurations that may be realistically used or otherwise encountered by an AV, i.e., real-world lane configurations may be fully represented by the set of predefined lane geometries 400 in some operational contexts and use cases of the AVs herein.
[0032] According to some embodiments of the present disclosure, each lane geometry (i.e., shape or configuration) in the set of predefined lane geometries 400 may be characterized or defined by a combination of a predefined number C of lane component types. In one exemplary embodiment, the set of predefined lane geometries 400 may be completely defined by a combination of three (i.e., C=3) predefined lane component types. In this embodiment, the three lane component types may include a continuing lane component, a merging lane component, and a diverging lane component. In some aspects, the set of predefined lane geometries 400 may be completely defined by a superposition (i.e., combination) of the three lane component types (i.e., a continuing lane component, a merging lane component, and a diverging lane component). Thus, because substantially all road topologies can be represented by the set of predefined lane geometries 400, the combination of the three lane component types disclosed herein may be used to characterize or define road lanes for substantially all road topologies.
[0033] In some aspects, the disclosed framework, which includes a predefined set of lane geometries and a predefined number of lane component types that can be combined to define road lanes, can be deployed on an AV to enable a computing system on the AV to characterize lanes of interest in real time near the AV as it travels on the road. In instances where a road includes more than one lanes and multiple lanes are of interest to the AV, the concepts and processes herein for determining the topology of a single lane can be applied to each lane of the multi-lane road, and an aggregation of multiple topologies, one for each lane, can be determined to generate an overall or global topology for the multi-lane road (e.g., FIG. 3, road 300).
[0034] In some embodiments, the systems and processes herein may leverage aspects of deep learning to detect road lane geometry. In particular, aspects and techniques of neural network landmark detection may be used in some embodiments herein to detect road geometry herein in a flexible and adaptable manner. For example, images captured by an AV's onboard camera as the vehicle crosses a road may be provided as input to a neural network configured to characterize or define the road lanes as a set of N ordered points (i.e., junctions / landmarks).
[0035] FIG. 5A is an illustrative depiction of road 500, for example, as captured in an image obtained by a camera on an AV as a vehicle travels the road. In the example of FIG. 5, the road includes two lanes, lane 505 (i.e., a continuing lane component) and lane 510 (i.e., a merging lane component). According to other aspects herein, roads may be configured to include diverging lane components, and roads may include one or more of the three lane component types disclosed herein in various combinations. Applying a landmark detection process to the image of road 500 may result in lanes 505 and 510, each defined by a set of N ordered points (i.e., junctions or landmarks), as depicted in FIGS. 5B and 5C. 5B and 5C, there are N=8 junctions or landmarks, each of which corresponds to approximately the same location within each lane (e.g., 515-530, 520-535, and 525-540), as demonstrated by representative junctions or landmarks 515, 520, and 525 in FIG. 5B and representative junctions or landmarks 530, 535, and 540 in FIG. 5B. In some embodiments, the value of N may be set based on, for example, rules, algorithms, operating conditions of the AV and its surroundings (environment), implementation constraints (e.g., available computing resources, etc.), and other factors.
[0036] In some embodiments, each set of N ordered points (i.e., landmarks) is vertically distributed within the image, and each landmark may represent each of the three (or other number of defined) lane component types (e.g., continuing lane components, merging lane components, and diverging lane components) disclosed herein. That is, each of the three lane component types will have N points describing it, such that a single lane herein may be defined by N×C points (C=3).
[0037] 6A-6D are illustrative examples of lanes represented by combinations of three lane component types at N ordered points defining the lanes, according to an exemplary embodiment. FIG. 6A demonstrates the three lane component types at each of the landmarks of a lane 605 that has no branches or merges. A junction 600 defines a graphical element representation for each of the three lane component types in FIGS. 6A-6D. Due to the strictly linear topology of the road 605, all of the landmarks are positioned on top of one another. FIG. 6B shows a road 610 with a merging lane topology, depicting the locations of the three lane component types at each of the landmarks. FIG. 6C demonstrates a road 615 with a diverging lane topology, depicting the locations of the three lane component types at each of the landmarks. FIG. 6D further illustrates a road 620 having a topology of merging lanes and then diverging lanes, which topology is also efficiently and completely represented by the locations of the three lane component types disclosed herein at each of the determined landmarks.
[0038] 6A-6D thus demonstrate how aspects disclosed herein can effectively and flexibly represent the topologies of a variety of different road lane configurations in an efficient manner. For example, in some respects, the topology of the lanes can be represented or defined by a minimum number N×C parameters.
[0039] 7 is an illustrative block diagram of a system according to an example embodiment. In some aspects, FIG. 7 is a high-level system architecture 700 for determining the topology of a single lane of a road. As illustrated, an input image 705 is received or otherwise acquired from a camera mounted on an AV. The image may typically include a forward-facing perspective of the road in the immediate vicinity of the AV.
[0040] An input image 705 is provided to or otherwise received by, for example, a U-Net-style (i.e., convolutional neural network) image segmentation network 710. In some embodiments, deep learning features are utilized to (1) interpret "objects as points" to detect lane junctions or landmarks within the input image and (2) use the neural network to generate a desired output 715 comprising N×C (e.g., C=3) images, which may be used to define the topology (i.e., structure) of the lanes detected within the input image 705. In some aspects, the desired output comprising N×C (e.g., C=3) images is a result of the problem representation used in some embodiments herein (i.e., how lanes herein can be represented or defined by C (e.g., 3) lane component types and N ordered landmarks).
[0041] In some embodiments, neural network 710 may include a U-Net-style implementation using a ResNet-50 backbone convolutional neural network to generate the desired output images. However, embodiments of the present disclosure are not limited to using conventional networks such as U-Net and ResNet-50, as well as other deep learning techniques and processes, as long as they are compatible with other aspects of the present disclosure that may be used. In some aspects, each output image includes one single strategic point or landmark, as described above. That is, all output images in output 715 correspond to one single strategic point or landmark. If the desired output is N ordered points, the order of images in output 715 of FIG. 7 corresponds to the order of strategic points or landmarks established or determined for the lanes in the input images. Thus, output 715 includes three different output type images for each strategic point (i.e., continuing lane component image 720, merging lane component image 725, and diverging lane component image 730), describing the merging lane, describing the continuing lane, and describing the diverging lane, respectively.
[0042] Recalling the principle of lane component superposition introduced above, images of a junction can be used to determine the locations of continuing lane components, merging lane components, and diverging lane components of each junction or landmark, which can then be combined to generate or determine the topology of a given lane.
[0043] In some embodiments, the systems and methodologies disclosed herein for determining the topology of a single lane in an image can be extended to determine the topology of additional (i.e., other) lanes in an image, i.e., the methods disclosed herein for representing lanes as specific lane component types, as well as the use and application of deep learning image segmentation and landmark detection techniques and neural networks, can be applied to all of the lanes in an image to generate the overall or complete topology of a road containing one or more lanes.
[0044] 8 is an illustrative flow diagram of one example of a road topology detection process 800, according to an example embodiment. In some embodiments, a framework or architecture disclosed herein (e.g., FIG. 7, architecture 700) may be used to implement some aspects of process 800. In some instances, certain aspects of process 800 are discussed in detail elsewhere in this disclosure and may not be repeated here in the following discussion of FIG. 8.
[0045] In operation 805, an image including a first lane of a road is received by a system, device, or apparatus implementing process 800. According to some example embodiments herein, process 800 may be performed by an AV as the vehicle travels down a road, with the images captured or otherwise obtained by one or more cameras located onboard the AV, and the images received by a computer processing unit, device, system, component, or module on the vehicle for analysis and processing. In some aspects, as disclosed above, each lane of a road may be represented by a set of predefined lane geometries.
[0046] In operation 810, a first lane in the image received in operation 805 may be defined by an ordered set of N (significant) points. As disclosed above (e.g., in FIG. 7), deep learning techniques, including certain types of neural networks, may be used to detect an ordered set of N significant points or landmarks in the first lane in the image, where these N ordered significant points define the first lane.
[0047] Following operation 815, each of the N ordered points defining the first lane may be represented by a combination of a predefined number of lane component types C. As disclosed in certain exemplary embodiments herein, the predefined number (C) of lane component types may be three, including a continuing lane component, a merging lane component, and a diverging lane component. In some aspects, as mentioned above, each lane geometry in the set of predefined lane geometries may be completely defined by a combination of a predefined number of lane component types C. Referring to the example of FIG. 4, as further illustrated in FIGS. 6A-6D, the set of predefined lane geometries may include seven lane geometries that may be defined by a combination of three predefined numbers of lane component types.
[0048] In operation 820, the aspect of defining a first lane as a set of N ordered points and representing the first lane by a combination of a predefined number of lane component types C may be used to generate (N×C) images, each representing one of the lane component types for one of the N ordered points. This aspect of operation is demonstrated, for example, by output 715 disclosed in FIG. 7.
[0049] Proceeding to operation 825, the images generated in operation 820 may be used to generate a topological representation for the first lane. In some aspects, the images generated in operation 820 may be combined (e.g., via superposition) to generate the topological representation for the first lane.
[0050] In some cases, process 800 may be performed for each lane of a multi-lane road to generate a complete topological representation for the road that includes all of the road's lanes. For example, process 800 may be performed for a first lane of the road and, subsequent to or at least partially in parallel with the processing of the first lane, may be repeatedly (e.g., iteratively) performed for each of one or more additional lanes of the road to generate a complete topological representation for the road that includes the first lane and one or more additional lanes.
[0051] FIG. 9 illustrates a computing system 900 that may be used in any of the architectures or frameworks (e.g., FIG. 1 , computer 140; FIG. 7 , architecture 700) and processes (e.g., FIG. 8) disclosed herein, according to an example embodiment. FIG. 9 is a block diagram of a server node 900 embodying an event processor, according to some embodiments. The computing system 900 may comprise a general-purpose computing device and may execute program code to perform any of the functions described herein. The computing system 900 may include other, not shown, elements, according to some embodiments.
[0052] The computing system 900 includes a processing unit 910 operably coupled to a communication device 920, a data storage device 930, one or more input devices 940, one or more output devices 950, and a memory 960. The communication device 920 may facilitate communication with external devices, such as an external network, a data storage device, or other data source. The input device 940 may comprise, for example, a keyboard, a keypad, a mouse or other pointing device, a microphone, a knob or switch, an infrared (IR) port, a docking station, and / or a touchscreen. The input device 940 may be used, for example, to input information into the computing system 900 (e.g., a manual request for a specific adjustment of an AV operation or function). The output device 950 may comprise, for example, a display (e.g., a display screen), a speaker, and / or a printer.
[0053] The data storage device 930 may comprise any suitable persistent storage device, including combinations of magnetic storage devices (e.g., magnetic tape, hard disk drives, and flash memory), optical storage devices, read-only memory (ROM) devices, etc., while the memory 960 may comprise random access memory (RAM).
[0054] The application servers 932 may each include program code executed by the processor 910 to cause the computing system 900 to perform any one or more of the processes described herein. Embodiments are not limited to the execution of these processes by a single computing device. The data storage device 930 may also store data and other program code to provide additional functionality and / or necessary for the operation of the computing system 900, such as device drivers, operating system files, etc. The topology detection engine 934 may include program code executed by the processor 910 to generate a topological representation of one or more lanes of a road traversed by the AV, as disclosed in various embodiments herein. The topology detection engine 934 may reference a database management system, DBMS, 936, as disclosed herein, to obtain an image of the road being navigated to determine the topology of the road's lanes. Results generated by the topology detection engine 934 may be stored in the DBMS node 936.
[0055] As will be understood based on the foregoing specification, the above-described examples of the present disclosure may be implemented using computer programming or engineering techniques, including computer software, firmware, hardware, or any combination or subset thereof. Any such resulting program having computer-readable code may be embodied or provided in one or more non-transitory computer-readable media, thereby creating a computer program product, i.e., an article of manufacture, according to the discussed examples of the present disclosure. For example, the non-transitory computer-readable medium may be, but is not limited to, a fixed drive, a diskette, an optical disk, a magnetic tape, flash memory, an external drive, semiconductor memory such as read-only memory (ROM), random access memory (RAM), and / or any other non-transitory transmission and / or reception medium, such as the Internet, cloud storage, the Internet of Things (IoT), or other communications network or link. An article of manufacture including the computer code may be created and / or used by executing the code directly from one medium, by copying the code from one medium to another, or by transmitting the code over a network.
[0056] A computer program (also referred to as a program, software, software application, "app," or code) may include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, cloud storage, Internet of Things, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. However, "machine-readable medium" and "computer-readable medium" do not include transitory signals. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or any other type of data to a programmable processor.
[0057] The above descriptions and illustrations of processes herein should not be construed as implying a fixed order for performing process steps. Rather, process steps may be performed in any order feasible, including simultaneous performance of at least some steps. Although the present disclosure has been described in conjunction with specific examples, it should be understood that various changes, substitutions, and alterations apparent to those skilled in the art may be made to the disclosed embodiments without departing from the spirit and scope of the present disclosure, as set forth in the appended claims.
Claims
1. 1. A vehicle computing system, comprising: a memory for storing computer instructions; a data storage device that stores data associated with operation of the vehicle, including data captured by at least the first sensor; a processor communicatively coupled to the memory for executing the computer instructions, during operation of the vehicle, receiving an image of a first lane of a road, the image being captured by the first sensor; defining the first lane as a set of N ordered points; representing the first lane for each of the N ordered points by a combination of a predefined number of lane component types C; generating an image based on associating one or more of the lane component types C with one or more of the N ordered points; and combining the generated images to generate a topological representation for the first lane.
2. The vehicle computing system of claim 1 , wherein each lane of the road is represented by a set of predefined lane geometries.
3. 3. The vehicle computing system of claim 2, wherein each lane geometry in the set of predefined lane geometries is completely defined by a combination of the predefined number of lane components of type C.
4. 4. The vehicle computing system of claim 3, wherein the set of predefined lane geometries includes at least seven lane geometries defined by a combination of a predefined number of lane component types, where C is equal to at least three components, and an (N×C) image is generated based on associating one or more of the lane component types C with one or more of the N ordered points.
5. The vehicle computing system of claim 1 , wherein the first sensor is a camera mounted on the vehicle.
6. The vehicle computing system of claim 1 , wherein the predefined number of lane components of type C comprises continuing lane components, merging lane components, and diverging lane components.
7. the processor: receiving an image of one or more additional lanes of the road, the image of the one or more additional lanes being captured by the first sensor; defining each of the one or more additional lanes as a set of N ordered points; representing the one or more additional lanes for each of the N ordered points by a combination of the predefined number of lane component types C; generating an image for each of the one or more additional lanes based on associating one or more of the lane component types C with one or more of the N ordered points; combining the generated images for each of the one or more additional lanes to generate a topological representation for each of the one or more additional lanes; 2. The vehicle computing system of claim 1, further comprising: combining the generated topological representation for each of the one or more additional lanes with the generated topological representation for the first lane to generate a complete topological representation for the road including the first lane and the one or more additional lanes.
8. 1. A method comprising: receiving an image of a first lane of a road, the image being captured by a first sensor; defining the first lane as a set of N ordered points; representing the first lane for each of the N ordered points by a combination of a predefined number of lane component types C; generating an image based on associating one or more of the lane component types C with one or more of the N ordered points; and combining the generated images to generate a topological representation for the first lane.
9. The method of claim 8 , wherein each lane of the road is represented by a set of predefined lane geometries.
10. The method of claim 9 , wherein each lane geometry in the set of predefined lane geometries is completely defined by a combination of the predefined number of lane components of type C.
11. 11. The method of claim 10, wherein the set of predefined lane geometries includes at least seven lane geometries defined by a combination of a predefined number of lane component types, where C is equal to at least three components, and an (N×C) image is generated based on associating one or more of the lane component types C with one or more of the N ordered points.
12. The method of claim 8 , wherein the first sensor is a camera mounted on a vehicle.
13. The method of claim 8 , wherein the predefined number of lane component types C includes continuing lane components, merging lane components, and diverging lane components.
14. receiving an image of one or more additional lanes of the road, the image of the one or more additional lanes being captured by the first sensor; defining each of the one or more additional lanes as a set of N ordered points; representing the one or more additional lanes for each of the N ordered points by a combination of a predefined number of lane component types C; generating an image for each of the one or more additional lanes based on associating one or more of the lane component types C with one or more of the N ordered points; combining the generated images for each of the one or more additional lanes to generate a topological representation for each of the one or more additional lanes; 9. The method of claim 8, further comprising combining the generated topological representation for each of the one or more additional lanes with the generated topological representation for the first lane to generate a complete topological representation for the road including the first lane and the one or more additional lanes.
15. A non-transitory medium having processor-executable instructions stored thereon, the non-transitory medium comprising: receiving an image of a first lane of a road, the image being captured by a first sensor; instructions for defining the first lane as a set of N ordered points; instructions for representing the first lane for each of the N ordered points by a combination of a predefined number of lane component types C; instructions for generating an image based on associating one or more of the lane component types C with one or more of the N ordered points; and instructions for combining the generated images to generate a topological representation for the first lane.
16. The non-transitory medium of claim 15 , wherein each lane of the road is represented by a set of predefined lane geometries.
17. 17. The non-transitory medium of claim 16, wherein each lane geometry in the set of predefined lane geometries is completely defined by a combination of the predefined number of lane component types C, and an (N×C) image is generated based on associating one or more of the lane component types C with one or more of the N ordered points.
18. The non-transitory medium of claim 15 , wherein the first sensor is a camera mounted on a vehicle.
19. The temporary medium of claim 15 , wherein the predefined number of type C lane components comprises continuing lane components, merging lane components, and diverging lane components.
20. instructions for receiving images of one or more additional lanes of the road, the images of the one or more additional lanes being captured by the first sensor; instructions for defining each of the one or more additional lanes as a set of N ordered points; instructions for representing the one or more additional lanes for each of the N ordered points by a combination of the predefined number of lane component types C; instructions for generating, for each of the one or more additional lanes, an image based on associating one or more of the lane component types C with one or more of the N ordered points; instructions for combining the generated images for each of the one or more additional lanes to generate a topological representation for each of the one or more additional lanes; 16. The temporary medium of claim 15, further comprising instructions for combining the generated topological representation for each of the one or more additional lanes with the generated topological representation for the first lane to generate a complete topological representation for the road including the first lane and the one or more additional lanes.