A trained network for identifying vehicle routes
By using a trained model to process vehicle images and generate target trajectory information, combined with sparse maps and crowdsourced data, the problems of data processing burden and map update difficulties in autonomous vehicle navigation are solved, thereby improving navigation efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOBILEYE VISION TECH LTD
- Filing Date
- 2024-08-23
- Publication Date
- 2026-05-26
Smart Images

Figure CN122094873A_ABST
Abstract
Description
[0001] Cross-reference to related applications This application claims the benefit of priority to U.S. Provisional Application No. 63 / 655,695, filed June 4, 2024, and U.S. Provisional Patent Application No. 63 / 534,683, filed August 25, 2023. The foregoing applications are incorporated herein by reference in their entirety. Technical Field
[0002] This disclosure relates generally to vehicle navigation, and more specifically to systems and methods for detecting drivable paths in the environment of a vehicle. Background Technology
[0003] With continuous technological advancements, the goal of fully autonomous vehicles capable of navigating on roads is fast approaching. Autonomous vehicles may need to consider a multitude of factors and make appropriate decisions based on those factors to safely and accurately reach their intended destination. For example, autonomous vehicles may need to process and interpret visual information (e.g., information captured from cameras) and may also use information from other sources (e.g., from GPS devices, rate sensors, accelerometers, suspension sensors, etc.). Simultaneously, to navigate to their destination, autonomous vehicles may also need to identify their position within a specific roadway (e.g., a specific lane in a multi-lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, obey traffic signals and signs, and move from one road to another at appropriate intersections or junctions. Utilizing and interpreting the vast amounts of information collected by the autonomous vehicle as it travels to its destination presents numerous design challenges. The sheer volume of data that autonomous vehicles may need to analyze, access, and / or store (e.g., captured image data, map data, GPS data, sensor data, etc.) presents challenges that can actually limit or even adversely affect autonomous navigation. Furthermore, if autonomous vehicles rely on traditional mapping techniques for navigation, the sheer volume of data required to store and update maps presents a formidable challenge. Summary of the Invention
[0004] Embodiments consistent with this disclosure provide systems and methods for vehicle navigation.
[0005] In one embodiment, a navigation system for a host vehicle may include at least one processor, the at least one processor including a circuit system and a memory. The memory may include instructions, when executed by the circuit system, to cause the at least one processor to receive at least one image captured by a camera mounted on the host vehicle, wherein the at least one image includes representations of two or more features; to provide the at least one image to a trained model configured to generate an output identifying two or more target trajectories associated with each of the two or more features; to determine location information of the two or more target trajectories associated with each of the two or more features based on the output generated by the trained model; to determine at least one navigation action of the host vehicle based on the location information determined for at least one target trajectory; and to cause the host vehicle to perform the at least one navigation action.
[0006] In one embodiment, a method for navigating a master vehicle may include receiving at least one image captured by a camera mounted on the master vehicle, wherein the at least one image includes representations of two or more features; providing the at least one image to a trained model configured to generate an output identifying two or more target trajectories associated with each of the two or more features; determining location information of the two or more target trajectories associated with each of the two or more features based on the output generated by the trained model; determining at least one navigation action of the master vehicle based on the location information determined for at least one target trajectory; and causing the master vehicle to perform the at least one navigation action.
[0007] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions that are executed by at least one processor and perform any of the methods described herein.
[0008] The foregoing general description and the following detailed description are merely exemplary and illustrative, and do not limit the scope of the claims. Attached Figure Description
[0009] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments. In the drawings: Figure 1 This is a schematic representation of an exemplary system consistent with the disclosed embodiments.
[0010] Figure 2A A schematic side view representation of an exemplary vehicle that includes a system consistent with the disclosed embodiments.
[0011] Figure 2B To be consistent with the disclosed embodiments Figure 2A The diagram shows a schematic top-view representation of the vehicle and system.
[0012] Figure 2C This is a schematic top-view representation of another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0013] Figure 2D This is a schematic top-view representation of yet another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0014] Figure 2E This is a schematic top-view representation of yet another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0015] Figure 2F This is a schematic representation of an exemplary vehicle control system consistent with the disclosed embodiments.
[0016] Figure 3A A schematic representation of the interior of a vehicle, including a rearview mirror and user interface for a vehicle imaging system, consistent with the disclosed embodiments.
[0017] Figure 3B An illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0018] Figure 3C To be viewed from a different perspective in accordance with the disclosed embodiments. Figure 3B The diagram shows the camera mount.
[0019] Figure 3D An illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0020] Figure 4 An exemplary block diagram of a memory configured to store instructions for performing one or more operations, consistent with the disclosed embodiments.
[0021] Figure 5A A flowchart is provided to illustrate an exemplary process for evoking one or more navigation responses based on monocular image analysis, consistent with the disclosed embodiments.
[0022] Figure 5B A flowchart illustrating an exemplary process for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments.
[0023] Figure 5CA flowchart illustrating an exemplary process for detecting road markings and / or lane geometry information in a set of images, consistent with the disclosed embodiments.
[0024] Figure 5D A flowchart illustrating an exemplary process for detecting traffic lights in a set of images, consistent with the disclosed embodiments.
[0025] Figure 5E A flowchart is provided to illustrate an exemplary process for evoking one or more navigation responses based on a vehicle path, consistent with the disclosed embodiments.
[0026] Figure 5F A flowchart illustrating an exemplary process for determining whether a vehicle ahead is changing lanes, consistent with the disclosed embodiments.
[0027] Figure 6 A flowchart illustrating an exemplary process for evoking one or more navigation responses based on stereo image analysis, consistent with the disclosed embodiments.
[0028] Figure 7 A flowchart illustrating an exemplary process consistent with the disclosed embodiments for evoking one or more navigation responses based on the analysis of three sets of images.
[0029] Figure 8 A sparse map for providing autonomous vehicle navigation is shown, consistent with the disclosed embodiments.
[0030] Figure 9A A polynomial representation of a portion of a road segment consistent with the disclosed embodiments is shown.
[0031] Figure 9B A curve representing a vehicle's target trajectory for a road segment is shown in three-dimensional space, which is included in a sparse map consistent with the disclosed embodiments.
[0032] Figure 10 Example landmarks that can be included in a sparse map consistent with the disclosed embodiments are shown.
[0033] Figure 11A A polynomial representation of the trajectory consistent with the disclosed embodiments is shown.
[0034] Figure 11B and Figure 11C A target trajectory along a multi-lane road, consistent with the disclosed embodiments, is shown.
[0035] Figure 11D An example road signature profile curve consistent with the disclosed embodiments is shown.
[0036] Figure 12A schematic illustration of a system for autonomous vehicle navigation using crowdsourced data received from multiple vehicles, consistent with the disclosed embodiments.
[0037] Figure 13 An example autonomous vehicle road navigation model, represented by multiple 3D splines, is shown, consistent with the disclosed embodiments.
[0038] Figure 14 A map skeleton generated from combined location information from multiple drivers is shown, consistent with the disclosed embodiments.
[0039] Figure 15 An example of longitudinal alignment of two driving vehicles with example signs as landmarks, consistent with the disclosed embodiments, is shown.
[0040] Figure 16 An example of longitudinal alignment of multiple drives with example signs as landmarks, consistent with the disclosed embodiments, is shown.
[0041] Figure 17 This is a schematic illustration of a system for generating driving data using cameras, vehicles, and servers, consistent with the disclosed embodiments.
[0042] Figure 18 A schematic illustration of a system for crowdsourcing sparse maps, consistent with the disclosed embodiments.
[0043] Figure 19 A flowchart illustrating an exemplary process for generating a sparse map for autonomous vehicle navigation along road segments, consistent with the disclosed embodiments.
[0044] Figure 20 A block diagram of a server consistent with the disclosed embodiments is shown.
[0045] Figure 21 A block diagram of a memory consistent with the disclosed embodiments is shown.
[0046] Figure 22 A process for clustering vehicle trajectories associated with vehicles, consistent with the disclosed embodiments, is demonstrated.
[0047] Figure 23 A navigation system for a vehicle, consistent with the disclosed embodiments, is shown, which can be used for autonomous navigation.
[0048] Figure 24A , Figure 24B , Figure 24C and Figure 24D An exemplary detectable lane marking consistent with the disclosed embodiments is shown.
[0049] Figure 24E An exemplary mapped lane marking consistent with the disclosed embodiments is shown.
[0050] Figure 24F An exemplary anomaly associated with detecting lane markings is shown, consistent with the disclosed embodiments.
[0051] Figure 25A An exemplary image of the vehicle's surroundings for navigation based on mapped lane markings, consistent with the disclosed embodiments, is shown.
[0052] Figure 25B Lateral positioning correction of a vehicle based on mapped lane markings is demonstrated in a road navigation model consistent with the disclosed embodiments.
[0053] Figure 25C and Figure 25D A conceptual representation of a localization technique is provided for locating a master vehicle along a target trajectory using mapped features included in a sparse map.
[0054] Figure 26A A flowchart illustrating an exemplary process for mapping lane markings used in autonomous vehicle navigation, consistent with the disclosed embodiments.
[0055] Figure 26B A flowchart illustrating an exemplary process for autonomously navigating a master vehicle along a road segment using mapped lane markings, consistent with the disclosed embodiments.
[0056] Figure 27 Example images consistent with the disclosed embodiments are shown that can be used to predict the trajectory of a target along a roadway.
[0057] Figure 28 An example target trajectory that can be identified using a trained model, consistent with the disclosed embodiments, is shown.
[0058] Figure 29 A block diagram illustrating an example process for training and implementing a trained model for predicting target trajectories, consistent with the disclosed embodiments.
[0059] Figure 30A Illustrations of example navigation actions that can be taken by the master vehicle based on locations of identified target trajectories, consistent with the disclosed embodiments.
[0060] Figure 30B Another illustration of example navigation actions that can be taken by the master vehicle based on the location of the identified target trajectory, consistent with the disclosed embodiments.
[0061] Figure 30CExample road segments consistent with the disclosed embodiments are shown, which can be applied to a trained model to determine a target trajectory.
[0062] Figure 30D An example intersection is shown where a trained model, consistent with the disclosed embodiments, can be applied to determine a target trajectory.
[0063] Figure 31 A flowchart illustrating an example process for navigating a master vehicle, consistent with the disclosed embodiments, is provided.
[0064] Figure 32 A flowchart illustrating an example process for determining a drivable path for a primary vehicle relative to a road segment, consistent with the disclosed embodiments.
[0065] Figure 33 A flowchart is provided to illustrate another example process for determining the drivable path of a primary vehicle relative to a road segment, consistent with the disclosed embodiments. Detailed Implementation
[0066] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar components. While several illustrative embodiments are described herein, modifications, adaptations, and other embodiments are possible. For example, components shown in the drawings may be replaced, added, or modified, and the illustrative methods described herein may be modified by replacing, reordering, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Rather, the appropriate scope is defined by the appended claims.
[0067] Overview of Autonomous Vehicles As used throughout this disclosure, the term "autonomous vehicle" means a vehicle capable of implementing at least one navigational change without driver input. "Navigational change" refers to a change in one or more of the vehicle's steering, braking, or acceleration. For autonomy to be achieved, the vehicle does not need to be fully automated (e.g., fully operational without a driver or driver input). Rather, autonomous vehicles include those that can operate under driver control during certain time periods and without driver control during other time periods. Autonomous vehicles may also include those that control only some aspects of vehicle navigation (such as steering (e.g., to maintain vehicle alignment between lane constraints)) but leave other aspects to the driver (e.g., braking). In some cases, an autonomous vehicle may handle some or all aspects of the vehicle's braking, rate control, and / or steering.
[0068] Since human drivers typically rely on visual cues and observation to control vehicles, traffic infrastructure is correspondingly established, with lane markings, traffic signs, and traffic lights all designed to provide visual information to the driver. Given these design characteristics of traffic infrastructure, autonomous vehicles can include cameras and processing units that analyze visual information captured from the vehicle's environment. Visual information can include, for example, components of traffic infrastructure (e.g., lane markings, traffic signs, traffic lights, etc.) and other obstacles (e.g., other vehicles, pedestrians, debris, etc.) observable by the driver. Additionally, autonomous vehicles can also use stored information, such as information providing a model of the vehicle's environment during navigation. For example, a vehicle can use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data to provide information relevant to its environment while driving, and the vehicle (and other vehicles) can use this information to locate itself on the model.
[0069] In some embodiments of this disclosure, the autonomous vehicle may use information obtained during navigation (e.g., from cameras, GPS devices, accelerometers, speed sensors, suspension sensors, etc.). In other embodiments, the autonomous vehicle may use information obtained from past navigation by the vehicle (or other vehicles) during navigation. In still other embodiments, the autonomous vehicle may use a combination of information obtained during navigation and information obtained from past navigation. The following sections provide an overview of a system consistent with the disclosed embodiments, followed by an overview of forward imaging systems and methods consistent with the system. The following sections disclose systems and methods for constructing, using, and updating sparse maps for autonomous vehicle navigation.
[0070] System Overview Figure 1This is a block diagram representation of system 100 consistent with the exemplary embodiments disclosed. System 100 may include various components depending on the requirements of a particular implementation. In some embodiments, system 100 may include a processing unit 110, an image acquisition unit 120, a position sensor 130, one or more memory units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. Processing unit 110 may include one or more processing means. In some embodiments, processing unit 110 may include an application processor 180, an image processor 190, or any other suitable processing means. Similarly, depending on the requirements of a particular application, image acquisition unit 120 may include any number of image acquisition means and components. In some embodiments, image acquisition unit 120 may include one or more image capture means (e.g., cameras), such as image capture means 122, image capture means 124, and image capture means 126. System 100 may also include a data interface 128 that communicatively connects processing means 110 to image acquisition means 120. For example, data interface 128 may include any one or more wired and / or wireless links for transmitting image data acquired by image acquisition device 120 to processing unit 110.
[0071] Wireless transceiver 172 may include one or more devices configured to exchange transmissions with one or more networks (e.g., cellular networks, the Internet, etc.) via an air interface using radio frequency, infrared frequency, magnetic field, or electric field. Wireless transceiver 172 may use any known standard to transmit and / or receive data (e.g., Wi-Fi, Bluetooth®, Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions may include communication from a master vehicle to one or more remotely located servers. Such transmissions may also include (one-way or two-way) communication between the master vehicle and one or more target vehicles in the master vehicle's environment (e.g., to facilitate coordination of navigation of the master vehicle in view of or with target vehicles in the master vehicle's environment), or even broadcast transmissions to unspecified receivers in the vicinity of the transmitting vehicle.
[0072] Both application processor 180 and graphics processor 190 may include various types of processing devices. For example, either or both of application processor 180 and graphics processor 190 may include a microprocessor, a preprocessor (such as an image preprocessor), a graphics processing unit (GPU), a central processing unit (CPU), supporting circuitry, a digital signal processor, an integrated circuit, memory, or any other type of device suitable for running applications and suitable for image processing and analysis. In some embodiments, application processor 180 and / or graphics processor 190 may include any type of single-core or multi-core processor, mobile device microcontroller, central processing unit, etc. Various processing devices may be used, including, for example, processors purchased from manufacturers such as Intel®, AMD®, etc., or GPUs purchased from manufacturers such as NVIDIA®, ATI®, etc., and may include various architectures (e.g., x86 processors, ARM®, etc.).
[0073] In some embodiments, application processor 180 and / or image processor 190 may include any processor chip from the EyeQ family of processor chips available from Mobileye®. These processor designs each include multiple processing units with local memory and instruction sets. Such processors may include video inputs for receiving image data from multiple image sensors and may also include video output capabilities. In one example, the EyeQ2® uses 90 nm micrometer technology operating at 332 MHz. The EyeQ2® architecture consists of two floating-point hyper-threaded 32-bit RISC CPUs (MIPS32® 34K® cores), five Visual Computing Engines (VCEs), three Vector Microcode Processors (VMPs®), a Denali 64-bit mobile DDR controller, a 128-bit internal Sonics interconnect, dual 16-bit video input and 18-bit video output controllers, a 16-channel DMA, and several peripherals. The MIPS34K CPU manages the five VCEs, three VMPs™ and DMA, a second MIPS34K CPU and multi-channel DMA, and other peripherals. Five VCEs, three VMP® processors, and a MIPS34K CPU can perform the intensive vision computing required for versatile bundled applications. In another example, an EyeQ3® processor (a third-generation processor with six times the performance of an EyeQ2®) can be used in the disclosed embodiments. In other examples, EyeQ4® and / or EyeQ5® can be used in the disclosed embodiments. Of course, any newer or future EyeQ processing devices can also be used with the disclosed embodiments.
[0074] Any of the processing devices disclosed herein can be configured to perform certain functions. Configuring a processing device (such as the described EyeQ processor or any other controller or microprocessor) to perform certain functions may include programming computer-executable instructions and making those instructions available for execution by the processing device during operation of the processing device. In some embodiments, configuring the processing device may include programming the processing device directly using architectural instructions. For example, processing devices such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may be configured using, for example, one or more hardware description languages (HDLs).
[0075] In other embodiments, configuring the processing means may include storing executable instructions on memory accessible to the processing means during operation. For example, the processing means may access the memory during operation to obtain and execute the stored instructions. In any case, a processing means configured to perform the sensing, image analysis, and / or navigation functions disclosed herein represents a specialized hardware-based system that controls multiple hardware-based components of a master vehicle.
[0076] Although Figure 1 Two separate processing devices included in processing unit 110 are depicted, but more or fewer processing devices may be used. For example, in some embodiments, a single processing device may be used to perform the tasks of application processor 180 and image processor 190. In other embodiments, these tasks may be performed by more than two processing devices. Furthermore, in some embodiments, system 100 may include one or more processing units 110 without other components such as image acquisition unit 120.
[0077] Processing unit 110 may include various types of devices. For example, processing unit 110 may include various devices such as controllers, image preprocessors, central processing units (CPUs), graphics processing units (GPUs), support circuitry, digital signal processors, integrated circuits, memory, or any other type of device for image processing and analysis. An image preprocessor may include a video processor for capturing, digitizing, and processing images from an image sensor. A CPU may include any number of microcontrollers or microprocessors. A GPU may also include any number of microcontrollers or microprocessors. Support circuitry may be any number of circuits well known in the art, including caches, power supplies, clocks, and input / output circuits. Memory may store software that controls the operation of the system when executed by the processor. Memory may include databases and image processing software. Memory may include any number of random access memories, read-only memories, flash memory, disk drives, optical storage devices, magnetic tape storage devices, removable storage devices, and other types of storage devices. In one instance, memory may be separate from processing unit 110. In another instance, memory may be integrated into processing unit 110.
[0078] Each memory unit 140, 150 may include software instructions that, when executed by a processor (e.g., application processor 180 and / or image processor 190), can control various aspects of the operation of system 100. For example, these memory units may include various database and image processing software, as well as trained systems such as neural networks or deep neural networks. Memory units may include random access memory (RAM), read-only memory (ROM), flash memory, disk drives, optical storage devices, magnetic tape storage devices, removable storage devices, and / or any other type of storage device. In some embodiments, memory units 140, 150 may be decoupled from application processor 180 and / or image processor 190. In other embodiments, these memory units may be integrated into application processor 180 and / or image processor 190.
[0079] The location sensor 130 may include any type of device adapted to determine a location associated with at least one component of the system 100. In some embodiments, the location sensor 130 may include a GPS receiver. Such a receiver can determine a user's location and speed by processing signals broadcast by Global Positioning System satellites. Location information from the location sensor 130 may be used by the application processor 180 and / or the image processor 190.
[0080] In some embodiments, system 100 may include components such as a speed sensor (e.g., a tachometer, speedometer) for measuring the rate of vehicle 200 and / or an accelerometer (single-axis or multi-axis) for measuring the acceleration of vehicle 200.
[0081] User interface 170 may include any means adapted to provide information to or receive input from one or more users of system 100. In some embodiments, user interface 170 may include user input devices, including, for example, a touchscreen, microphone, keyboard, pointer device, scroll wheel, camera, knob, button, etc. Using such input devices, users may be able to provide information input or commands to system 100 by typing instructions or information, providing voice commands, selecting menu options on the screen (using buttons, pointers, or eye-tracking capabilities), or via any other suitable technology for transmitting information to system 100.
[0082] User interface 170 may be equipped with one or more processing devices configured to provide and receive information from the user, and process the information for use by, for example, application processor 180. In some embodiments, such processing devices may perform instructions for recognizing and tracking eye movements, receiving and interpreting voice commands, recognizing and interpreting touches and / or gestures made on a touchscreen, responding to keyboard input or menu selection, etc. In some embodiments, user interface 170 may include a display, speakers, haptic devices, and / or any other means for providing output information to the user.
[0083] Map database 160 may include any type of database for storing map data useful to system 100. In some embodiments, map database 160 may include data relating to the location of various items in a reference coordinate system, including roads, water features, geographic features, businesses, points of interest, restaurants, gas stations, etc. Map database 160 may not only store the locations of such items but also descriptors associated with those items, including, for example, names associated with any of the stored features. In some embodiments, map database 160 may be physically located along with other components of system 100. Alternatively or additionally, map database 160 or a portion thereof may be remotely located relative to other components of system 100 (e.g., processing unit 110). In such embodiments, information from map database 160 may be downloaded via a wired or wireless data connection to a network (e.g., via cellular networks and / or the Internet, etc.). In some cases, map database 160 may store a sparse data model including a polynomial representation of certain road features (e.g., lane markings) or target trajectories for the primary vehicle. Reference will be made below. Figures 8 to 19 This paper discusses systems and methods for generating such maps.
[0084] Image capture devices 122, 124, and 126 may each include any type of means suitable for capturing at least one image from the environment. Furthermore, any number of image capture devices can be used to acquire images for input to an image processor. Some embodiments may include only a single image capture device, while other embodiments may include two, three, or even four or more image capture devices. Reference will be made below. Figures 2B to 2E The image capture devices 122, 124 and 126 are further described.
[0085] System 100 or its various components can be incorporated into a variety of different platforms. In some embodiments, system 100 may be included on vehicle 200, such as... Figure 2A As shown. For example, vehicle 200 may be equipped with processing unit 110 and any other components of system 100, as described above relative to... Figure 1 As described. Although in some embodiments, vehicle 200 may be equipped with only a single image capture device (e.g., a camera), in other embodiments (such as combined with...) Figures 2B to 2E Among those discussed, multiple image capture devices can be used. For example, such as Figure 2A As shown, either of the image capture devices 122 and 124 of the vehicle 200 may be part of an ADAS (Advanced Driver Assistance System) imaging set.
[0086] The image capture device, which is part of the image acquisition unit 120 and is included on the vehicle 200, can be positioned at any suitable location. In some embodiments, such as Figures 2A to 2E and Figures 3A to 3C As shown, the image capturing device 122 can be located in the vicinity of the rearview mirror. This location provides a line of sight similar to that of the driver of vehicle 200, which can help determine what is visible and invisible to the driver. The image capturing device 122 can be positioned anywhere near the rearview mirror, but placing the image capturing device 122 on the driver's side of the mirror can further assist in obtaining an image representing the driver's field of view and / or line of sight.
[0087] Other positioning of the image capture device for image acquisition unit 120 may also be used. For example, image capture device 124 may be positioned on or therein on the bumper of vehicle 200. Such positioning may be particularly suitable for image capture devices with a wide field of view. The line of sight of the image capture device positioned on the bumper may be different from that of the driver, and therefore the bumper image capture device and the driver may not always see the same object. Image capture devices (e.g., image capture devices 122, 124, and 126) may also be positioned in other locations. For example, the image capture device may be located on or therein on one or both of the side mirrors of vehicle 200, on the roof of vehicle 200, on the hood of vehicle 200, on the trunk of vehicle 200, on the side of vehicle 200, mounted on any of the windows of vehicle 200, positioned behind or in front of it, and mounted in or near the headlight graphics at the front and / or rear of vehicle 200, etc.
[0088] In addition to the image capture device, vehicle 200 may also include various other components of system 100. For example, processing unit 110 may be included on vehicle 200, or integrated with or separate from the vehicle's engine control unit (ECU). Vehicle 200 may also be equipped with position sensor 130 such as a GPS receiver, and may also include map database 160 and memory units 140 and 150.
[0089] As discussed above, the wireless transceiver 172 can receive data via one or more networks (e.g., cellular networks, the Internet, etc.) and / or. For example, the wireless transceiver 172 can upload data collected by the system 100 to one or more servers and download data from said one or more servers. Via the wireless transceiver 172, the system 100 can receive, for example, periodic or on-demand updates to data stored in map database 160, memory 140, and / or memory 150. Similarly, the wireless transceiver 172 can upload any data from the future system 100 (e.g., images captured by image acquisition unit 120, data received by position sensor 130 or other sensors, vehicle control system, etc.) and / or any data processed by processing unit 110 to one or more servers.
[0090] System 100 may upload data to a server (e.g., to the cloud) based on privacy level settings. For example, System 100 may implement privacy level settings to regulate or restrict the data types (including metadata) sent to a server that can uniquely identify the vehicle and / or the vehicle's driver / owner. Such settings may be configured by the user via, for example, wireless transceiver 172, through factory default settings, or through data received by wireless transceiver 172.
[0091] In some embodiments, system 100 may upload data according to a “high” privacy level, and under a setting, system 100 may transmit data (e.g., location information related to routes, captured images, etc.) without any details about specific vehicles and / or drivers / owners. For example, when uploading data under a “high” privacy setting, system 100 may not include the vehicle identification number (VIN) or the name of the driver or vehicle owner, and may instead transmit data such as captured images and / or location information related to route restrictions.
[0092] Other privacy levels are conceivable. For example, system 100 may transmit data to the server at a “medium” privacy level, including additional information not included at a “high” privacy level, such as the vehicle’s brand and / or model and / or vehicle type (e.g., passenger vehicle, SUV, truck, etc.). In some embodiments, system 100 may upload data at a “low” privacy level. At a “low” privacy level setting, system 100 may upload data and include information sufficient to uniquely identify a specific vehicle, owner / driver, and / or part or all of the route traveled by the vehicle. Such “low” privacy level data may include, for example, one or more of the following: VIN, driver / owner name, vehicle’s origin point before departure, vehicle’s intended destination, vehicle’s brand and / or model, vehicle type, etc.
[0093] Figure 2A A schematic side view representation of an exemplary vehicle imaging system consistent with the disclosed embodiments. Figure 2B for Figure 2A The illustrated embodiment is shown in a schematic top view. Figure 2B The illustrated and disclosed embodiments may include a vehicle 200, which includes a system 100 in its body, the system having a first image capture device 122 located in the vicinity of a rearview mirror and / or near the driver of the vehicle 200, a second image capture device 124 located on or therein in a bumper area of the vehicle 200 (e.g., one of the bumper areas 210), and a processing unit 110.
[0094] like Figure 2C As shown, both image capture devices 122 and 124 can be positioned in the vicinity of the rearview mirror and / or near the driver of vehicle 200. Additionally, although... Figure 2B and Figure 2C Two image capture devices 122 and 124 are shown, but it should be understood that other embodiments may include more than two image capture devices. For example, in Figure 2D and Figure 2E In the illustrated embodiment, the first, second, and third image capture devices 122, 124, and 126 are included in the system 100 of the vehicle 200.
[0095] like Figure 2D As shown, image capture device 122 can be positioned in the vicinity of the rearview mirror and / or near the driver of vehicle 200, and image capture devices 124 and 126 can be positioned on or within the bumper area of vehicle 200 (e.g., one of bumper areas 210). And as... Figure 2E As shown, image capturing devices 122, 124, and 126 may be positioned in the vicinity of the rearview mirror and / or near the driver's seat of vehicle 200. The disclosed embodiments are not limited to any particular number and configuration of image capturing devices, and the image capturing devices may be positioned within vehicle 200 and / or in any suitable location on said vehicle.
[0096] It should be understood that the disclosed embodiments are not limited to vehicles and can be applied to other situations. It should also be understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and can be applied to all types of vehicles, including cars, trucks, trailers and other types of vehicles.
[0097] The first image capture device 122 may include any suitable type of image capture device. The image capture device 122 may include an optical axis. In one example, the image capture device 122 may include an Aptina M9V024 WVGA sensor with a global shutter. In other embodiments, the image capture device 122 may provide a resolution of 1280 x 960 pixels and may include a rolling shutter. The image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included, for example, to provide a desired focal length and field of view for the image capture device. In some embodiments, the image capture device 122 may be associated with a 6mm lens or a 12mm lens. In some embodiments, the image capture device 122 may be configured to capture an image with a desired field of view (FOV) 202, such as... Figure 2DAs shown. For example, image capture device 122 can be configured to have a regular FOV, such as 46 degrees, 50 degrees, 52 degrees, or greater, in the range of 40 to 56 degrees. Alternatively, image capture device 122 can be configured to have a narrow FOV, such as 28 degrees or 36 degrees, in the range of 23 to 40 degrees. Furthermore, image capture device 122 can be configured to have a wide FOV in the range of 100 to 180 degrees. In some embodiments, image capture device 122 may include a wide-angle bumper camera or a camera with a maximum FOV of 180 degrees. In some embodiments, image capture device 122 may be a 7.2 M-pixel image capture device with an aspect ratio of approximately 2:1 (e.g., HxV = 3800 × 1900 pixels) and a horizontal FOV of approximately 100 degrees. Such an image capture device can be used to replace a three-image capture device configuration. Due to significant lens distortion, in embodiments where the image capture device uses radially symmetrical lenses, the vertical FOV of such image capture devices may be significantly less than 50 degrees. For example, such lenses may not be radially symmetrical, which would allow a vertical FOV greater than 50 degrees and a horizontal FOV of 100 degrees.
[0098] The first image capturing device 122 can acquire a plurality of first images relative to a scene associated with the vehicle 200. Each of the plurality of first images can be acquired as a series of image scan lines that can be captured using a rolling shutter. Each scan line may include a plurality of pixels.
[0099] The first image capture device 122 may have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate may refer to the rate at which the image sensor can acquire image data associated with each pixel included in a particular scan line.
[0100] For example, image capture devices 122, 124, and 126 may include any suitable type and number of image sensors, including CCD sensors or CMOS sensors. In one embodiment, a CMOS image sensor may be used in conjunction with a rolling shutter, such that each pixel in a row is read one at a time, and the scanning of rows is performed on a line-by-line basis until the entire image frame has been captured. In some embodiments, rows may be captured sequentially from top to bottom relative to the frame.
[0101] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124 and 126) may constitute a high-resolution imager and may have a resolution greater than 5M pixels, 7M pixels, 10M pixels or more.
[0102] The use of a rolling shutter can cause pixels in different rows to be exposed and captured at different times, which can lead to skew and other image artifacts in the captured image frame. On the other hand, when the image capture device 122 is configured to operate using a global or synchronous shutter, all pixels can be exposed for the same amount of time and during a common exposure period. Therefore, the image data in a frame collected from a system using a global shutter represents a snapshot of the entire FOV (e.g., FOV 202) at a specific time. Conversely, in a rolling shutter application, each row in the frame is exposed, and the data is captured at different times. Therefore, moving objects may appear distorted in an image capture device with a rolling shutter. This phenomenon will be described in more detail below.
[0103] The second image capture device 124 and the third image capture device 126 can be any type of image capture device. Similar to the first image capture device 122, each of the image capture devices 124 and 126 may include an optical axis. In one embodiment, each of the image capture devices 124 and 126 may include an Aptina M9V024WVGA sensor with a global shutter. Alternatively, each of the image capture devices 124 and 126 may include a rolling shutter. Similar to image capture device 122, image capture devices 124 and 126 may be configured to include various lenses and optical elements. In some embodiments, the lenses associated with image capture devices 124 and 126 may provide an FOV (such as FOV 202) that is the same as or narrower than that associated with image capture device 122 (such as FOV 204 and 206). For example, image capture devices 124 and 126 may have an FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less.
[0104] Image capture devices 124 and 126 can acquire multiple second and third images of a scene associated with vehicle 200. Each of these multiple second and third images can be acquired as a second series of image scan lines and a third series of image scan lines that can be captured using a rolling shutter. Each scan line or row can have multiple pixels. Image capture devices 124 and 126 can have a second scan rate and a third scan rate associated with the acquisition of each of the image scan lines included in the second and third series.
[0105] Each image capture device 122, 124, and 126 can be positioned relative to the vehicle 200 at any suitable location and orientation. The relative positioning of the image capture devices 122, 124, and 126 can be selected to facilitate the fusion of information acquired from the image capture devices. For example, in some embodiments, the field of view (FOV) associated with image capture device 124 (such as FOV 204) can partially or completely overlap with the FOV associated with image capture device 122 (such as FOV 202) and the FOV associated with image capture device 126 (such as FOV 206).
[0106] Image capture devices 122, 124, and 126 can be positioned at any suitable relative height on the vehicle 200. In one example, a height difference may exist between image capture devices 122, 124, and 126, which can provide sufficient parallax information for stereoscopic analysis. For example, as Figure 2A As shown, the two image capture devices 122 and 124 are at different heights. For example, a lateral displacement difference may also exist between image capture devices 122, 124, and 126, providing additional parallax information for stereo analysis by the processing unit 110. The difference in lateral displacement can be determined by d... x Indicates, such as Figure 2C and Figure 2D As shown. In some embodiments, there may be forward or backward displacement (e.g., range displacement) between image capture devices 122, 124, and 126. For example, image capture device 122 may be positioned 0.5 meters to 2 meters or more behind image capture devices 124 and / or 126. This type of displacement allows one of the image capture devices to cover potential blind spots of the other image capture devices.
[0107] Image capture device 122 may have any suitable resolution capability (e.g., the number of pixels associated with an image sensor), and the resolution of the image sensor associated with image capture device 122 may be higher, lower, or the same as the resolution of the image sensors associated with image capture devices 124 and 126. In some embodiments, the image sensors associated with image capture device 122 and / or image capture devices 124 and 126 may have a resolution of 640×480, 1024×768, 1280×960, or any other suitable resolution.
[0108] The frame rate (e.g., the rate at which an image capture device acquires a set of pixel data for an image frame before continuing to capture pixel data associated with the next image frame) can be controllable. The frame rate associated with image capture device 122 can be higher, lower, or the same as the frame rates associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 can depend on a variety of factors that may affect the timing of the frame rate. For example, one or more of image capture devices 122, 124, and 126 may include a selectable pixel delay period applied before or after the acquisition of image data associated with one or more pixels of the image sensors in image capture devices 122, 124, and / or 126. Generally, image data corresponding to each pixel can be acquired according to the clock rate for the device (e.g., one pixel per clock cycle). Additionally, in embodiments including a rolling shutter, one or more of the image capture devices 122, 124, and 126 may include a selectable horizontal blanking period applied before or after the acquisition of image data associated with pixel rows of the image sensors in the image capture devices 122, 124, and / or 126. Furthermore, one or more of the image capture devices 122, 124, and / or 126 may include a selectable vertical blanking period applied before or after the acquisition of image data associated with image frames of the image capture devices 122, 124, and 126.
[0109] These timing controls enable synchronization of the frame rates associated with image capture devices 122, 124, and 126, even when the line scan rates of each image capture device differ. Furthermore, as will be discussed in more detail below, these selectable timing controls and other factors (e.g., image sensor resolution, maximum line scan rate, etc.) can enable synchronization of image capture from areas where the field of view (FOV) of image capture device 122 overlaps with one or more FOVs of image capture devices 124 and 126, even when the field of view of image capture device 122 differs from the FOVs of image capture devices 124 and 126.
[0110] The frame rate timing in the image capture devices 122, 124, and 126 can depend on the resolution of the associated image sensor. For example, assuming similar line scan rates for two devices, if one device includes an image sensor with a resolution of 640×480 and the other device includes an image sensor with a resolution of 1280×960, it will take more time to acquire one frame of image data from the sensor with the higher resolution.
[0111] Another factor that can affect the timing of image data acquisition in image capture devices 122, 124, and 126 is the maximum line scan rate. For example, acquiring one line of image data from an image sensor included in image capture devices 122, 124, and 126 will require a minimum amount of time. Assuming no pixel delay period is added, this minimum amount of time for acquiring one line of image data will be related to the maximum line scan rate of a particular device. A device that provides a higher maximum line scan rate is likely to provide a higher frame rate than a device with a lower maximum line scan rate. In some embodiments, one or more of image capture devices 124 and 126 may have a higher maximum line scan rate than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture devices 124 and / or 126 may be 1.25 times, 1.5 times, 1.75 times, or 2 times or more of the maximum line scan rate of image capture device 122.
[0112] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may be operated at a scan rate less than or equal to its maximum scan rate. The system may be configured such that one or more of image capture devices 124 and 126 operate at a line scan rate equal to the line scan rate of image capture device 122. In other instances, the system may be configured such that the line scan rate of image capture device 124 and / or image capture device 126 may be 1.25 times, 1.5 times, 1.75 times, or 2 times or more of the line scan rate of image capture device 122.
[0113] In some embodiments, image capture devices 122, 124, and 126 may be asymmetric. That is, they may include cameras with different fields of view (FOV) and focal lengths. For example, the fields of view of image capture devices 122, 124, and 126 may include any desired area relative to the environment of vehicle 200. In some embodiments, one or more of image capture devices 122, 124, and 126 may be configured to acquire image data from the environment of the front of vehicle 200, the rear of vehicle 200, the sides of vehicle 200, or a combination thereof.
[0114] Furthermore, the focal length associated with each image capturing device 122, 124, and / or 126 can be selectable (e.g., by including appropriate lenses, etc.) so that each device acquires an image of an object relative to the vehicle 200 at a desired distance range. For example, in some embodiments, image capturing devices 122, 124, and 126 can acquire images of close-up objects within a few meters of the vehicle. Image capturing devices 122, 124, and 126 can also be configured to acquire images of objects at a greater distance from the vehicle (e.g., 25 m, 50 m, 100 m, 150 m, or more). Furthermore, the focal lengths of image capturing devices 122, 124, and 126 can be selected such that one image capturing device (e.g., image capturing device 122) can acquire images of objects relatively close to the vehicle (e.g., within 10 m or 20 m), while other image capturing devices (e.g., image capturing devices 124 and 126) can acquire images of objects further away from the vehicle (e.g., greater than 20 m, 50 m, 100 m, 150 m, etc.).
[0115] According to some embodiments, the field of view (FOV) of one or more image capture devices 122, 124, and 126 may have a wide angle. For example, having a 140-degree FOV may be advantageous, particularly for image capture devices 122, 124, and 126 that can be used to capture images of areas adjacent to the vehicle 200. For example, image capture device 122 may be used to capture images of areas to the right or left of the vehicle 200, and in such embodiments, it may be desirable for image capture device 122 to have a wide FOV (e.g., at least 140 degrees).
[0116] The field of view associated with each of the image capturing devices 122, 124, and 126 can depend on the corresponding focal length. For example, as the focal length increases, the corresponding field of view decreases.
[0117] Image capture devices 122, 124, and 126 can be configured to have any suitable field of view. In one particular example, image capture device 122 may have a horizontal FOV of 46 degrees, image capture device 124 may have a horizontal FOV of 23 degrees, and image capture device 126 may have a horizontal FOV between 23 degrees and 46 degrees. In another example, image capture device 122 may have a horizontal FOV of 52 degrees, image capture device 124 may have a horizontal FOV of 26 degrees, and image capture device 126 may have a horizontal FOV between 26 degrees and 52 degrees. In some embodiments, the ratio of the FOV of image capture device 122 to the FOV of image capture device 124 and / or image capture device 126 may vary between 1.5 and 2.0. In other embodiments, this ratio may vary between 1.25 and 2.25.
[0118] System 100 can be configured such that the field of view of image capture device 122 at least partially or completely overlaps with the field of view of image capture devices 124 and / or 126. In some embodiments, system 100 can be configured such that the field of view of image capture devices 124 and 126, for example, falls within the field of view of image capture device 122 (e.g., narrower than said field of view) and shares a common center with said field of view. In other embodiments, image capture devices 122, 124, and 126 can capture adjacent FOVs or can have partial overlap in their FOVs. In some embodiments, the field of view of image capture devices 122, 124, and 126 can be aligned such that the center of the narrower FOV image capture device 124 and / or 126 is located in the lower half of the field of view of the wider FOV device 122.
[0119] Figure 2F This is a schematic representation of an exemplary vehicle control system consistent with the disclosed embodiments. Figure 2F As indicated, vehicle 200 may include a throttle system 220, a braking system 230, and a steering system 240. System 100 may provide inputs (e.g., control signals) to one or more of the throttle system 220, braking system 230, and steering system 240 via one or more data links (e.g., any wired and / or wireless link or link used for data transmission). For example, based on analysis of images acquired by image capture devices 122, 124, and / or 126, system 100 may provide control signals to one or more of the throttle system 220, braking system 230, and steering system 240 to navigate vehicle 200 (e.g., by inducing acceleration, turning, lane changing, etc.). Furthermore, system 100 may receive inputs from one or more of the throttle system 220, braking system 230, and steering system 240 indicative of operating conditions of vehicle 200 (e.g., speed, whether vehicle 200 is braking and / or turning, etc.). The following is in conjunction with... Figures 4 to 7 Further details will be provided.
[0120] like Figure 3AAs shown, vehicle 200 may also include a user interface 170 for interaction with the driver or passengers of vehicle 200. For example, the user interface 170 in a vehicle application may include a touchscreen 320, a knob 330, a button 340, and a microphone 350. The driver or passengers of vehicle 200 may also interact with system 100 using handles (e.g., located on or near the steering column of vehicle 200, including, for example, a turn signal handle), buttons (e.g., located on the steering wheel of vehicle 200), etc. In some embodiments, microphone 350 may be located adjacent to rearview mirror 310. Similarly, in some embodiments, image capture device 122 may be located near rearview mirror 310. In some embodiments, user interface 170 may also include one or more speakers 360 (e.g., speakers of a vehicle audio system). For example, system 100 may provide various notifications (e.g., alarms) via speaker 360.
[0121] Figures 3B to 3D For the purpose of illustrating an exemplary camera mount 370 consistent with the disclosed embodiments, the camera mount is configured to be positioned behind a rearview mirror (e.g., rearview mirror 310) and against the vehicle's windshield. Figure 3B As shown, camera mount 370 may include image capture devices 122, 124, and 126. Image capture devices 124 and 126 may be positioned behind a glare shield 380 that is flush with a vehicle windshield and comprises a composition of a film and / or anti-reflective material. For example, glare shield 380 may be positioned such that it is aligned against a vehicle windshield having a matching slope. In some embodiments, each of image capture devices 122, 124, and 126 may be positioned behind glare shield 380, such as, for example... Figure 3D The embodiments described herein are not limited to any particular configuration of the image capture devices 122, 124 and 126, the camera mount 370 and the glare shield 380. Figure 3C From the front perspective Figure 3B The camera mount 370 is shown in the image.
[0122] As those skilled in the art who benefit from this disclosure will understand, various variations and / or modifications can be made to the disclosed embodiments. For example, not all components are essential for the operation of system 100. Furthermore, any component can be located in any suitable part of system 100, and these components can be rearranged into various configurations while providing the functionality of the disclosed embodiments. Thus, the foregoing configuration is illustrative, and regardless of the configuration discussed above, system 100 can provide a wide range of functions to analyze the environment surrounding vehicle 200 and navigate vehicle 200 in response to said analysis.
[0123] As discussed in further detail below and consistent with the various disclosed embodiments, system 100 can provide various features related to autonomous driving and / or driver assistance technologies. For example, system 100 can analyze image data, location data (e.g., GPS location information), map data, speed data, and / or data from sensors included in vehicle 200. System 100 can collect data for analysis from, for example, image acquisition unit 120, position sensor 130, and other sensors. Furthermore, system 100 can analyze the collected data to determine whether vehicle 200 should take certain actions, and then automatically take those actions without human intervention. For example, when vehicle 200 is navigating without human intervention, system 100 can automatically control the braking, acceleration, and / or steering of vehicle 200 (e.g., by sending control signals to one or more of throttle system 220, braking system 230, and steering system 240). Additionally, system 100 can analyze the collected data and issue warnings and / or alerts to vehicle occupants based on the analysis of the collected data. Further details regarding various embodiments provided by system 100 are provided below.
[0124] Forward multiple imaging system As discussed above, system 100 can provide driver assistance functions using a multi-camera system. The multi-camera system may use one or more cameras facing forward of the vehicle. In other embodiments, the multi-camera system may include one or more cameras facing the side or rear of the vehicle. In one embodiment, for example, system 100 may use a dual-camera imaging system, wherein a first camera and a second camera (e.g., image capture devices 122 and 124) may be positioned at the front and / or side of the vehicle (e.g., vehicle 200). The first camera may have a field of view larger than, smaller than, or partially overlapping with the field of view of the second camera. Furthermore, the first camera may be connected to a first image processor for monocular image analysis of the images provided by the first camera, and the second camera may be connected to a second image processor for monocular image analysis of the images provided by the second camera. The outputs of the first and second image processors (e.g., processed information) may be combined. In some embodiments, the second image processor may receive images from both the first and second cameras for stereo analysis. In another embodiment, system 100 may use a three-camera imaging system, wherein each of the cameras has a different field of view. Therefore, such systems can make decisions based on information derived from objects positioned at varying distances from both the front and sides of a vehicle. A reference to monocular image analysis can refer to instances where image analysis is performed based on images captured from a single viewpoint (e.g., from a single camera). Stereo image analysis can refer to instances where image analysis is performed based on two or more images captured using one or more variations of image capture parameters. For example, captured images suitable for stereo image analysis can include: images captured from two or more different locations, images captured from different fields of view, images captured using different focal lengths and parallax information, etc.
[0125] For example, in one embodiment, system 100 may use image capture devices 122, 124, and 126 to implement a three-camera configuration. In such a configuration, image capture device 122 may provide a narrow field of view (e.g., 34 degrees or other values selected from the range of about 20 to 45 degrees), image capture device 124 may provide a wide field of view (e.g., 150 degrees or other values selected from the range of about 100 to about 180 degrees), and image capture device 126 may provide an intermediate field of view (e.g., 46 degrees or other values selected from the range of about 35 to about 60 degrees). In some embodiments, image capture device 126 may act as a master or primary camera. Image capture devices 122, 124, and 126 may be positioned behind rearview mirror 310 and substantially side-by-side (e.g., spaced 6 cm apart). Further, in some embodiments, as discussed above, one or more of image capture devices 122, 124, and 126 may be mounted behind a glare shield 380 flush with the windshield of vehicle 200. Such shielding can minimize the impact of any reflections from inside the car on the image capture devices 122, 124 and 126.
[0126] In another embodiment, as described above... Figure 3B and Figure 3C The wide field-of-view camera discussed (e.g., image capture device 124 in the example above) can be mounted below the narrow field-of-view camera and the main field-of-view camera (e.g., image devices 122 and 126 in the example above). This configuration provides a free line of sight from the wide field-of-view camera. To reduce reflections, the camera can be mounted close to the windshield of the vehicle 200 and may include a polarizer on the camera to reduce reflected light.
[0127] A three-camera system can provide certain performance characteristics. For example, some embodiments may include verifying the ability of one camera to detect objects based on detection results from another camera. In the three-camera configuration discussed above, processing unit 110 may include, for example, three processing devices (e.g., three EyeQ series processor chips, as discussed above), wherein each processing device is dedicated to processing images captured by one or more of the image capture devices 122, 124, and 126.
[0128] In a three-camera system, the first processing unit can receive images from both the main camera and the narrow field-of-view camera, and perform visual processing on the narrow FOV camera to detect, for example, other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the first processing unit can calculate the pixel parallax between the images from the main camera and the narrow camera and create a 3D reconstruction of the vehicle 200's environment. The first processing unit can then combine the 3D reconstruction with 3D map data or with 3D information calculated based on information from the other camera.
[0129] The second processing unit can receive images from the main camera and perform visual processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the second processing unit can calculate camera displacement and, based on said displacement, calculate pixel parallax between consecutive images and create a 3D reconstruction of the scene (e.g., from moving structures). The second processing unit can then send the structure from the motion-based 3D reconstruction to the first processing unit for combination with the stereoscopic 3D image.
[0130] The third processing unit can receive images from a wide field of view (FOV) camera and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing unit can further execute additional processing instructions to analyze the images to identify moving objects in the images, such as vehicles changing lanes or pedestrians.
[0131] In some embodiments, enabling image-based information streams to be captured and processed independently can provide opportunities for redundancy in the system. Such redundancy may include, for example, using a first image capture device and images processed from said device to verify and / or supplement information obtained by capturing and processing image information from at least a second image capture device.
[0132] In some embodiments, system 100 may use two image capture devices (e.g., image capture devices 122 and 124) to provide navigation assistance to vehicle 200, and a third image capture device (e.g., image capture device 126) to provide redundancy and verify the analysis of data received from the other two image capture devices. For example, in such a configuration, image capture devices 122 and 124 may provide images for stereo analysis by system 100 for navigation of vehicle 200, while image capture device 126 may provide images for monocular analysis by system 100 to provide redundancy and verification based on information obtained from images captured by image capture devices 122 and / or 124. That is, image capture device 126 (and corresponding processing device) may be considered to provide a redundant subsystem for checking the analysis derived from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Furthermore, in some embodiments, the redundancy and verification of the received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors), information received from one or more transceivers outside the vehicle, etc.
[0133] Those skilled in the art will recognize that the camera configurations, camera placements, number of cameras, camera positioning, etc., described above are merely exemplary. These components, and other components described relative to the overall system, can be assembled and used in a variety of different configurations without departing from the scope of the disclosed embodiments. Further details regarding the use of a multi-camera system to provide driver assistance and / or autonomous vehicle functionality are as follows.
[0134] Figure 4 An exemplary functional block diagram of a memory 140 and / or 150, consistent with the disclosed embodiments, having instructions for performing one or more operations. Although memory 140 is mentioned below, those skilled in the art will recognize that instructions may be stored in memory 140 and / or 150.
[0135] like Figure 4 As shown, memory 140 may store monocular image analysis module 402, stereo image analysis module 404, velocity and acceleration module 406, and navigation response module 408. The disclosed embodiments are not limited to any particular configuration of memory 140. Further, application processor 180 and / or image processor 190 may execute instructions stored in any of the modules 402, 404, 406, and 408 included in memory 140. Those skilled in the art will understand that references to processing unit 110 in the following discussion may individually or collectively refer to application processor 180 and image processor 190. Therefore, steps in any of the following processes may be performed by one or more processing devices.
[0136] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from a set of images with additional sensor information (e.g., information from radar, lidar, etc.) to perform monocular image analysis. As described below... Figures 5A to 5D As described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, system 100 (e.g., via processing unit 110) may induce one or more navigation responses in vehicle 200, such as turning, lane changing, or changes in acceleration, as discussed below in conjunction with navigation response module 408.
[0137] In one embodiment, the stereo image analysis module 404 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform stereo image analysis on a first set of images and a second set of images acquired by a combination of image capture devices selected from any of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the first set of images and the second set of images with additional sensing information (e.g., information from radar) to perform stereo image analysis. For example, the stereo image analysis module 404 may include instructions for performing stereo image analysis based on the first set of images acquired by the image capture device 124 and the second set of images acquired by the image capture device 126. (The following is a continuation of the previous paragraph.) Figure 6 As described, the stereo image analysis module 404 may include instructions for detecting a set of features (such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, etc.) within a first set of images and a second set of images. Based on this analysis, the processing unit 110 may induce one or more navigation responses (such as turning, lane changing, changes in acceleration) in the vehicle 200, as discussed below in conjunction with the navigation response module 408. Furthermore, in some embodiments, the stereo image analysis module 404 may implement techniques associated with trained systems (such as neural networks or deep neural networks) or untrained systems (such as systems configured to use computer vision algorithms to detect and / or label objects in the environment from which it captures and processes sensor information). In one embodiment, the stereo image analysis module 404 and / or other image processing modules may be configured to use a combination of trained and untrained systems.
[0138] In one embodiment, the speed and acceleration module 406 may store software configured to analyze data received from one or more computing and electromechanical devices in the vehicle 200, which are configured to cause changes in the speed and / or acceleration of the vehicle 200. For example, the processing unit 110 may execute instructions associated with the speed and acceleration module 406 to calculate a target speed for the vehicle 200 based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include, for example, target position, speed and / or acceleration, position and / or speed of the vehicle 200 relative to nearby vehicles, pedestrians, or road objects, position information of the vehicle 200 relative to lane markings on the road, etc. In addition, the processing unit 110 may calculate the target speed for the vehicle 200 based on sensor inputs (e.g., information from radar) and inputs from other systems of the vehicle 200 (such as the vehicle 200's throttle system 220, braking system 230, and / or steering system 240). Based on the calculated target rate, the processing unit 110 can transmit electronic signals to the throttle system 220, braking system 230 and / or steering system 240 of the vehicle 200 to trigger a change in speed and / or acceleration by, for example, physically pressing down the brakes or releasing the accelerator of the vehicle 200.
[0139] In one embodiment, the navigation response module 408 may store software executable by the processing unit 110 to determine the desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include position and rate information associated with nearby vehicles, pedestrians, and road objects, target position information for vehicle 200, etc. Additionally, in some embodiments, the navigation response may be (partially or entirely) based on map data, the predetermined position of vehicle 200, and / or the relative velocity or relative acceleration between vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 may also determine the desired navigation response based on sensor inputs (e.g., information from radar) and inputs from other systems of vehicle 200 (such as the throttle system 220, braking system 230, and steering system 240 of vehicle 200). Based on the desired navigation response, the processing unit 110 can transmit electronic signals to the throttle system 220, braking system 230, and steering system 240 of the vehicle 200 to trigger the desired navigation response by, for example, turning the steering wheel of the vehicle 200 to achieve a predetermined angle of rotation. In some embodiments, the processing unit 110 can use the output of the navigation response module 408 (e.g., the desired navigation response) as input to the execution of the speed and acceleration module 406 to calculate the change in the rate of the vehicle 200.
[0140] Furthermore, any module disclosed herein (e.g., modules 402, 404, and 406) may implement techniques associated with trained systems (such as neural networks or deep neural networks) or untrained systems.
[0141] Figure 5A A flowchart illustrating an exemplary process 500A for inducing one or more navigation responses based on monocular image analysis, consistent with the disclosed embodiments, is provided. At step 510, processing unit 110 may receive multiple images via data interface 128 between processing unit 110 and image acquisition unit 120. For example, a camera included in image acquisition unit 120 (such as image capture device 122 having a field of view 202) may capture multiple images of an area in front of vehicle 200 (or, for example, the side or rear of vehicle) and transmit these multiple images to processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). At step 520, processing unit 110 may execute monocular image analysis module 402 to analyze the multiple images, as described below. Figures 5B to 5D As described in further details. By performing this analysis, the processing unit 110 can detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, etc.
[0142] At step 520, processing unit 110 may also execute monocular image analysis module 402 to detect various road hazards, such as truck tire components, fallen road signs, loose cargo, small animals, etc. Road hazards can have different structures, shapes, sizes, and colors, which can make the detection of such hazards more challenging. In some embodiments, processing unit 110 may execute monocular image analysis module 402 to perform multi-frame analysis on multiple images to detect road hazards. For example, processing unit 110 may estimate camera motion between consecutive image frames and calculate pixel parallax between frames to construct a 3D map of the road. Processing unit 110 can then use the 3D map to detect hazards on the road surface and above the road surface.
[0143] At step 530, processing unit 110 may execute navigation response module 408 based on the analysis performed at step 520 and as described above. Figure 4The described techniques induce one or more navigation responses in vehicle 200. Navigation responses may include, for example, turning, lane changing, changes in acceleration, etc. In some embodiments, processing unit 110 may induce one or more navigation responses using data derived from the execution of speed and acceleration module 406. Additionally, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof. For example, processing unit 110 may induce vehicle 200 to change lanes and then accelerate by, for example, sequentially transmitting control signals to steering system 240 and throttle system 220 of vehicle 200. Alternatively, processing unit 110 may induce vehicle 200 to brake while changing lanes by, for example, simultaneously transmitting control signals to braking system 230 and steering system 240 of vehicle 200.
[0144] Figure 5B A flowchart illustrating an exemplary process 500B for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500B. At step 540, processing unit 110 may determine a set of candidate objects representing possible vehicles and / or pedestrians. For example, processing unit 110 may scan one or more images, compare the images to one or more predetermined patterns, and identify possible locations within each image that may contain the object of interest (e.g., a vehicle, pedestrian, or a portion thereof). The predetermined patterns may be designed to achieve a high “false hit” rate and a low “missed” rate. For example, processing unit 110 may use a low threshold of similarity to the predetermined pattern to identify candidate objects as possible vehicles or pedestrians. Doing so allows processing unit 110 to reduce the probability of missing (e.g., not identifying) candidate objects representing vehicles or pedestrians.
[0145] At step 542, processing unit 110 may filter a set of candidate objects based on classification criteria to exclude certain candidates (e.g., irrelevant or less relevant objects). Such criteria can be derived from various attributes associated with object types stored in a database (e.g., a database stored in memory 140). Attributes may include object shape, size, texture, location (e.g., relative to vehicle 200), etc. Therefore, processing unit 110 may use one or more sets of criteria to reject erroneous candidates from a set of candidate objects.
[0146] At step 544, processing unit 110 may analyze multiple image frames to determine whether an object in a set of candidate objects represents a vehicle and / or a pedestrian. For example, processing unit 110 may track detected candidate objects across consecutive frames and accumulate frame-by-frame data associated with the detected objects (e.g., size, position relative to vehicle 200, etc.). Additionally, processing unit 110 may estimate parameters for the detected objects and compare the frame-by-frame position data of the objects with predicted positions.
[0147] At step 546, processing unit 110 can construct a set of measurements for the detected object. Such measurements may include, for example, position, velocity, and acceleration values (relative to vehicle 200) associated with the detected object. In some embodiments, processing unit 110 can construct the measurements based on estimation techniques using a series of time-based observations such as a Kalman filter or linear quadratic estimation (LQE) and / or based on available modeling data for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). The Kalman filter may be based on a measurement of the object's scale, where the scale measurement is proportional to the collision time (e.g., the amount of time it takes for vehicle 200 to arrive at the object). Thus, by performing steps 540 through 546, processing unit 110 can identify vehicles and pedestrians appearing within a set of captured images and derive information associated with the vehicles and pedestrians (e.g., position, speed, size). Based on this identification and derived information, processing unit 110 can elicit one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0148] At step 548, processing unit 110 may perform optical flow analysis on one or more images to reduce the probability of detecting "false hits" and missing candidate objects representing vehicles or pedestrians. Optical flow analysis may refer to, for example, analyzing motion patterns relative to vehicle 200 in one or more images associated with other vehicles and pedestrians, and said motion patterns are different from road surface motion. Processing unit 110 can calculate the motion of candidate objects by observing different positions of objects across multiple image frames captured at different times. Processing unit 110 can use position and time values as inputs to a mathematical model to calculate the motion of candidate objects. Therefore, optical flow analysis can provide another method for detecting vehicles and pedestrians near vehicle 200. Processing unit 110 may combine steps 540 to 546 to perform optical flow analysis to provide redundancy for detecting vehicles and pedestrians and increase the reliability of system 100.
[0149] Figure 5CA flowchart illustrating an exemplary process 500C for detecting road markings and / or lane geometry information in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500C. At step 550, processing unit 110 may detect a set of objects by scanning one or more images. To detect segments of lane markings, lane geometry information, and other relevant road markings, processing unit 110 may filter the group of objects to exclude those determined to be irrelevant (e.g., potholes, small rocks, etc.). At step 552, processing unit 110 may group segments detected in step 550 that belong to the same road marking or lane marking together. Based on this grouping, processing unit 110 may develop a model, such as a mathematical model, representing the detected segments.
[0150] At step 554, processing unit 110 may construct a set of measurements associated with the detected segment. In some embodiments, processing unit 110 may create a projection of the detected segment from the image plane to the real-world plane. This projection may be characterized using a cubic polynomial with coefficients corresponding to physical properties such as the detected road position, slope, curvature, and derivative of curvature. In generating the projection, processing unit 110 may consider changes in the road surface and the pitch and roll rates associated with vehicle 200. Furthermore, processing unit 110 may model the road elevation by analyzing position and motion cues present on the road surface. Additionally, processing unit 110 may estimate the pitch and roll rates associated with vehicle 200 by tracking a set of feature points in one or more images.
[0151] At step 556, processing unit 110 can perform multi-frame analysis, for example, by tracking detected segments across consecutive image frames and accumulating frame-by-frame data associated with the detected segments. As processing unit 110 performs multi-frame analysis, the set of measurements constructed at step 554 can become more reliable and correlated with increasingly higher confidence levels. Therefore, by performing steps 550, 552, 554, and 556, processing unit 110 can identify road markings appearing within a set of captured images and derive lane geometry information. Based on this identification and derived information, processing unit 110 can induce one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0152] At step 558, processing unit 110 may consider additional information sources to further develop a safety model for scenarios involving vehicle 200 in its surrounding environment. Processing unit 110 may use the safety model to define scenarios in which system 100 can safely perform autonomous control of vehicle 200. To develop the safety model, in some embodiments, processing unit 110 may consider the positions and movements of other vehicles, detected road edges and barriers, and / or general road shape descriptions extracted from map data (such as data from map database 160). By considering additional information sources, processing unit 110 can provide redundancy for detecting road markings and lane geometry features and increase the reliability of system 100.
[0153] Figure 5D A flowchart illustrating an exemplary process 500D for detecting traffic lights in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500D. At step 560, processing unit 110 may scan the set of images and identify objects appearing in the images at locations that may contain traffic lights. For example, processing unit 110 may filter the identified objects to construct a set of candidate objects, excluding those objects that are unlikely to correspond to traffic lights. This filtering may be done based on various attributes associated with traffic lights, such as shape, size, texture, location (e.g., relative to vehicle 200), etc. Such attributes may be based on multiple examples of traffic lights and traffic control signals and stored in a database. In some embodiments, processing unit 110 may perform multi-frame analysis on a set of candidate objects reflecting possible traffic lights. For example, processing unit 110 may track candidate objects across consecutive image frames, estimate the real-world location of the candidate objects, and filter out those objects that are moving (which are unlikely to be traffic lights). In some embodiments, the processing unit 110 may perform color analysis on candidate objects and identify the relative positions of detected colors that appear inside possible traffic lights.
[0154] At step 562, processing unit 110 may analyze the geometric features of the intersection. This analysis may be based on any combination of: (i) the number of lanes detected on either side of vehicle 200, (ii) markings detected on the road (such as arrow markings), and (iii) a description of the intersection extracted from map data (such as data from map database 160). Processing unit 110 may use information derived from the execution of monocular analysis module 402 for this analysis. In addition, processing unit 110 may determine the correspondence between traffic lights detected at step 560 and lanes appearing near vehicle 200.
[0155] At step 564, as vehicle 200 approaches the intersection, processing unit 110 can update the confidence level associated with the analyzed intersection geometry and detected traffic lights. For example, the estimated number of traffic lights at the intersection may affect the confidence level compared to the actual number present at the intersection. Therefore, based on the confidence level, processing unit 110 can delegate control to the driver of vehicle 200 to improve safety conditions. By performing steps 560, 562, and 564, processing unit 110 can identify traffic lights appearing in the set of captured images and analyze intersection geometry information. Based on the identification and analysis, processing unit 110 can trigger one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0156] Figure 5E A flowchart illustrating an exemplary process 500E for inducing one or more navigation responses in vehicle 200 based on a vehicle path, consistent with the disclosed embodiments, is provided. At step 570, processing unit 110 may construct an initial vehicle path associated with vehicle 200. The vehicle path can be represented using a set of points expressed in coordinates (x, z), and the distance d between any two points in this set is... i The distance can fall within the range of 1 to 5 meters. In one embodiment, processing unit 110 can use two polynomials (such as a left-road polynomial and a right-road polynomial) to construct an initial vehicle path. Processing unit 110 can calculate the geometric midpoint between the two polynomials and, if applicable, include an offset of a predetermined amount (e.g., a smart lane offset) at each point in the resulting vehicle path (zero offset may correspond to driving in the middle of the lane). The offset may be along a direction perpendicular to the segment between any two points in the vehicle path. In another embodiment, processing unit 110 can use a polynomial and an estimated lane width to offset each point of the vehicle path by half the estimated lane width plus a predetermined offset (e.g., a smart lane offset).
[0157] At step 572, processing unit 110 may update the vehicle path constructed at step 570. Processing unit 110 may use a higher resolution to reconstruct the vehicle path constructed at step 570, such that the distance d between any two points in the set of points representing the vehicle path... k The distance d is less than the distance mentioned above. i For example, distance d k It can fall within the range of 0.1 meters to 0.3 meters. The processing unit 110 can use a parabolic spline algorithm to reconstruct the vehicle path, which can produce a cumulative distance vector S (i.e., based on a set of points representing the vehicle path) corresponding to the total length of the vehicle path.
[0158] At step 574, processing unit 110 may determine the look-ahead point (expressed in coordinates as (x...)) based on the updated vehicle path constructed at step 572. l , z l The processing unit 110 can extract a look-ahead point from the cumulative distance vector S, and the look-ahead point can be associated with a look-ahead distance and a look-ahead time. The look-ahead distance (which may have a lower bound in the range of 10 to 20 meters) can be calculated as the product of the vehicle 200's speed and the look-ahead time. For example, as the vehicle 200's speed decreases, the look-ahead distance can also decrease (e.g., until it reaches the lower bound). The look-ahead time (which may be in the range of 0.5 to 1.5 seconds) can be inversely proportional to the gain of one or more control loops (such as a heading error tracking control loop) associated with causing the navigation response in the vehicle 200. For example, the gain of the heading error tracking control loop can depend on the bandwidth of the yaw rate loop, the steering actuator loop, the vehicle's lateral dynamics, etc. Therefore, the higher the gain of the heading error tracking control loop, the shorter the look-ahead time.
[0159] At step 576, processing unit 110 can determine the heading error and yaw rate commands based on the look-ahead point determined at step 574. Processing unit 110 can do this by calculating the arctangent of the look-ahead point, for example, arctan(x) l / z l The heading error is determined by the processing unit 110. The processing unit 110 can determine the yaw rate command as the product of the heading error and the high-level control gain. If the look-ahead distance is not at the lower limit, the high-level control gain can be equal to: (2 / look-ahead time). Otherwise, the high-level control gain can be equal to: (2 * vehicle speed 200 / look-ahead distance).
[0160] Figure 5F A flowchart is provided to illustrate an exemplary process 500F for determining whether a vehicle ahead is changing lanes, consistent with the disclosed embodiments. At step 580, processing unit 110 may determine navigation information associated with the vehicle ahead (e.g., a vehicle traveling in front of vehicle 200). For example, processing unit 110 may use the above-described combination of... Figure 5A and 5B The described technique determines the position, speed (e.g., direction and velocity), and / or acceleration of a vehicle ahead. Processing unit 110 may also use the techniques described above. Figure 5E The described technique determines one or more road polynomials, look-ahead points (associated with vehicle 200), and / or snail tracks (e.g., a set of points describing the path taken by the vehicle ahead).
[0161] At step 582, processing unit 110 may analyze the navigation information determined at step 580. In one embodiment, processing unit 110 may calculate the distance between the snail track and the road polynomial (e.g., along the track). If the variance of this distance along the track exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters on a straight road, 0.3 to 0.4 meters on a moderately curved road, and 0.5 to 0.6 meters on a road with sharp curves), processing unit 110 may determine that the vehicle ahead may be changing lanes. In the case of multiple vehicles detected traveling in front of vehicle 200, processing unit 110 may compare the snail track associated with each vehicle. Based on the comparison, processing unit 110 may determine that a vehicle whose snail track does not match the snail tracks of other vehicles may be changing lanes. Processing unit 110 may additionally compare the curvature of the snail track (associated with the vehicle ahead) with the expected curvature of the road segment the vehicle ahead is traveling on. The expected curvature can be extracted from map data (e.g., data from map database 160), from road polynomials, from the snail tracks of other vehicles, from prior knowledge about the road, etc. If the difference between the curvature of the snail track and the expected curvature of the road segment exceeds a predetermined threshold, the processing unit 110 can determine that the vehicle ahead may be changing lanes.
[0162] In another embodiment, processing unit 110 may compare the instantaneous position of the vehicle ahead with a look-ahead point (associated with vehicle 200) over a specific time period (e.g., 0.5 to 1.5 seconds). If the distance between the instantaneous position of the vehicle ahead and the look-ahead point changes over the specific time period, and the cumulative sum of the changes exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a moderately curved road, and 1.3 to 1.7 meters on a road with sharp curves), processing unit 110 may determine that the vehicle ahead may be changing lanes. In another embodiment, processing unit 110 may analyze the geometry of the track by comparing the lateral distance traveled along the snail's path with the expected curvature of the snail's path. The expected radius of curvature can be calculated based on (δ... z 2 + δ x 2 ) / 2 / (δ x ) is determined, where δ x δ represents the lateral distance traveled. zThis represents the longitudinal distance traveled. If the difference between the lateral distance traveled and the expected curvature exceeds a predetermined threshold (e.g., 500 meters to 700 meters), processing unit 110 can determine that the vehicle ahead may be changing lanes. In another embodiment, processing unit 110 can analyze the position of the vehicle ahead. If the position of the vehicle ahead obscures the road polynomial (e.g., the vehicle ahead is covered on top of the road polynomial), processing unit 110 can determine that the vehicle ahead may be changing lanes. If the position of the vehicle ahead is such that another vehicle is detected in front of the vehicle ahead and the snail tracks of the two vehicles are not parallel, processing unit 110 can determine that the (closer) vehicle ahead may be changing lanes.
[0163] At step 584, processing unit 110 may determine whether the vehicle 200 ahead is changing lanes based on the analysis performed at step 582. For example, processing unit 110 may make the determination based on a weighted average of the various analyses performed in step 582. In such an approach, for example, a decision made by processing unit 110 based on a particular type of analysis that the vehicle ahead may be changing lanes may be assigned a value "1" (and "0" to indicate a determination that the vehicle ahead is unlikely to be changing lanes). Different analyses performed at step 582 may be assigned different weights, and the disclosed embodiments are not limited to any particular combination of analysis and weights.
[0164] Figure 6 A flowchart illustrating an exemplary process 600 for evoking one or more navigation responses based on stereoscopic image analysis, consistent with the disclosed embodiments, is provided. At step 610, processing unit 110 may receive a first plurality of images and a second plurality of images via data interface 128. For example, a camera included in image acquisition unit 120 (such as image capture devices 122 and 124 having fields of view 202 and 204) may capture the first plurality of images and the second plurality of images of an area in front of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first plurality of images and the second plurality of images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0165] At step 620, processing unit 110 may execute stereo image analysis module 404 to perform stereo image analysis on the first plurality of images and the second plurality of images to create a 3D map of the road in front of the vehicle and detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. Stereo image analysis can be performed in conjunction with the above. Figures 5A to 5DThe steps described are performed in a similar manner. For example, processing unit 110 may execute stereo image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road markings, traffic lights, road hazards, etc.) in a first plurality of images and a second plurality of images, filter out subgroups of candidate objects based on each object, perform multi-frame analysis, construct measurement results, and determine confidence levels for the remaining candidate objects. In performing the steps described above, processing unit 110 may consider information from both the first plurality of images and the second plurality of images, rather than information from only one set of images. For example, processing unit 110 may analyze the differences in pixel-level data (or other subsets of data from the two streams of captured images) of candidate objects appearing in both the first plurality of images and the second plurality of images. As another example, processing unit 110 may estimate the position and / or velocity (e.g., relative to vehicle 200) of a candidate object by observing that the candidate object appears in one of the plurality of images but not in the other, or relative to other differences that may exist relative to objects appearing in both image streams. For example, the position, velocity, and / or acceleration relative to vehicle 200 can be determined based on the trajectory, position, motion characteristics, etc., of features associated with one or both objects appearing in the image stream.
[0166] At step 630, processing unit 110 may execute navigation response module 408 based on the analysis performed at step 620 and as described above. Figure 4 The described techniques are used to induce one or more navigation responses in vehicle 200. Navigation responses may include, for example, turning, lane changing, changes in acceleration, changes in speed, braking, etc. In some embodiments, processing unit 110 may use data derived from the execution of speed and acceleration module 406 to induce one or more navigation responses. Additionally, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.
[0167] Figure 7A flowchart illustrating an exemplary process 700 consistent with the disclosed embodiments for evoking one or more navigation responses based on the analysis of three sets of images is provided. At step 710, processing unit 110 may receive a first plurality of images, a second plurality of images, and a third plurality of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) may capture the first plurality of images, the second plurality of images, and the third plurality of images of areas in front of and / or to the sides of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first plurality of images, the second plurality of images, and the third plurality of images via three or more data interfaces. For example, each of image capture devices 122, 124, and 126 may have an associated data interface for transmitting data to processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0168] In step 720, processing unit 110 can analyze the first plurality of images, the second plurality of images, and the third plurality of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. The analysis can be performed in a manner similar to the combination described above. Figures 5A to 5D and Figure 6 The steps described are performed in the manner described. For example, processing unit 110 can perform monocular image analysis on each of the first plurality of images, the second plurality of images, and the third plurality of images (e.g., via execution by monocular image analysis module 402 and based on the above combination). Figures 5A to 5D The steps described above). Alternatively, processing unit 110 may perform stereoscopic image analysis on the first plurality of images and the second plurality of images, the second plurality of images and the third plurality of images and / or the first plurality of images and the third plurality of images (e.g., via execution by stereoscopic image analysis module 404 and based on the above). Figure 6The steps described herein. Processed information corresponding to the analysis of a first plurality of images, a second plurality of images, and / or a third plurality of images can be combined. In some embodiments, processing unit 110 may perform a combination of monocular image analysis and stereoscopic image analysis. For example, processing unit 110 may perform monocular image analysis on the first plurality of images (e.g., via execution of monocular image analysis module 402) and stereoscopic image analysis on the second plurality of images and the third plurality of images (e.g., via execution of stereoscopic image analysis module 404). The configuration of image capture devices 122, 124, and 126—including their respective positioning and fields of view 202, 204, and 206—can affect the type of analysis performed on the first plurality of images, the second plurality of images, and the third plurality of images. The disclosed embodiments are not limited to a specific configuration of image capture devices 122, 124, and 126, or the type of analysis performed on the first plurality of images, the second plurality of images, and the third plurality of images.
[0169] In some embodiments, processing unit 110 may test system 100 based on the images acquired and analyzed at steps 710 and 720. Such testing may provide an indicator of the overall performance of system 100 for certain configurations of image capture devices 122, 124, and 126. For example, processing unit 110 may determine the ratio of “false hits” (e.g., situations where system 100 incorrectly determines the presence of a vehicle or pedestrian) and “missed hits”.
[0170] At step 730, processing unit 110 may induce one or more navigation responses in vehicle 200 based on information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images. The selection of either the first plurality of images, the second plurality of images, or the third plurality of images may depend on various factors, such as, for example, the number, type, and size of objects detected in each of the plurality of images. Processing unit 110 may also base its responses on image quality and resolution, the effective field of view reflected in the images, the number of captured frames, the extent to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which the object appears, the proportion of objects appearing in each such frame, etc.), etc.
[0171] In some embodiments, processing unit 110 can select information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images by determining the degree to which information derived from one image source is consistent with information derived from other image sources. For example, processing unit 110 can combine processed information derived from each of image capture devices 122, 124, and 126 (whether by monocular analysis, stereo analysis, or any combination of both) and determine consistent visual indicators (e.g., lane markings, detected vehicles and their locations and / or paths, detected traffic lights, etc.) across the images captured from each of image capture devices 122, 124, and 126. Processing unit 110 can also exclude inconsistent information across the captured images (e.g., vehicles changing lanes, lane models indicating vehicles too close to vehicle 200, etc.). Therefore, processing unit 110 can select information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images based on the determination of consistent and inconsistent information.
[0172] Navigation responses may include, for example, turning, lane changing, and changes in acceleration. Processing unit 110 can base its responses on the analysis performed in step 720 and the above-mentioned factors. Figure 4 The described techniques are used to induce one or more navigation responses. Processing unit 110 may also use data derived from the execution of velocity and acceleration module 406 to induce one or more navigation responses. In some embodiments, processing unit 110 may induce one or more navigation responses based on the relative position, relative velocity, and / or relative acceleration between vehicle 200 and objects detected in any of a first plurality of images, a second plurality of images, and a third plurality of images. Multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.
[0173] Sparse road model for autonomous vehicle navigation In some embodiments, the disclosed systems and methods may use sparse maps for autonomous vehicle navigation. Specifically, sparse maps may be used for autonomous vehicle navigation along road segments. For example, sparse maps may provide sufficient information for navigating autonomous vehicles without requiring the storage and / or updating of large amounts of data. As discussed in further detail below, autonomous vehicles may use sparse maps to navigate one or more roads based on one or more stored trajectories.
[0174] Sparse maps for autonomous vehicle navigation In some embodiments, the disclosed systems and methods can generate sparse maps for autonomous vehicle navigation. For example, sparse maps can provide sufficient information for navigation without requiring excessive data storage or data transfer rates. As discussed in further detail below, vehicles (which may be autonomous vehicles) can use sparse maps to navigate one or more roads. For example, in some embodiments, sparse maps may include sufficient data related to roads and potential landmarks along those roads for vehicle navigation, but this data also exhibits a small data footprint. For example, sparse data maps, as described in detail below, may require significantly less storage space and data transfer bandwidth compared to digital maps that include detailed map information, such as image data collected along roads.
[0175] For example, sparse data maps can store a three-dimensional polynomial representation of preferred vehicle routes along a road, rather than a detailed representation of road segments. These routes may require very little data storage space. Furthermore, in the described sparse data maps, landmarks can be identified and included in the sparse map road model to aid navigation. These landmarks can be located at any spacing suitable for vehicle navigation, but in some cases, it is not necessary to identify and include such landmarks in the model with high density and short spacing. Instead, in some cases, navigation based on landmarks spaced at least 50 meters, at least 100 meters, at least 500 meters, at least 1 kilometer, or at least 2 kilometers apart is possible. As will be discussed in more detail in other sections, sparse maps can be generated based on data collected or measured by vehicles equipped with various sensors and devices, such as image capture devices, GPS sensors, motion sensors, etc., as the vehicles travel along a roadway. In some cases, sparse maps can be generated based on data collected during multiple drives of one or more vehicles along a particular roadway. Generating sparse maps using multiple drives of one or more vehicles can be referred to as “crowdsourced” sparse maps.
[0176] Consistent with the disclosed embodiments, autonomous vehicle systems can use sparse maps for navigation. For example, the disclosed systems and methods can distribute sparse maps to generate road navigation models for autonomous vehicles, and the sparse maps and / or the generated road navigation models can be used to navigate autonomous vehicles along road segments. Sparse maps consistent with this disclosure may include one or more three-dimensional shapes that represent predetermined trajectories that autonomous vehicles can traverse as they move along associated road segments.
[0177] Sparse maps consistent with this disclosure may also include data representing one or more road features. Such road features may include identified landmarks, road signature contours, and any other road-related features useful for navigating the vehicle. Sparse maps consistent with this disclosure can enable autonomous vehicle navigation based on a relatively small amount of data included in the sparse map. For example, instead of including detailed representations of roads, such as road edges, road curvature, images associated with road segments, or data detailing other physical features associated with road segments, the disclosed embodiments of sparse maps may require relatively small storage space (and relatively small bandwidth when portions of the sparse map are transferred to the vehicle), but still adequately provide autonomous vehicle navigation. The small data footprint of the disclosed sparse maps, discussed in further detail below, can be achieved in some embodiments by storing representations of road-related elements that require a small amount of data but still enable autonomous navigation.
[0178] For example, the disclosed sparse map can store a polynomial representation of one or more trajectories that a vehicle can follow along a road, rather than storing a detailed representation of various aspects of the road. Therefore, instead of storing (or necessarily transmitting) details about the physical properties of the road to enable navigation along it, using the disclosed sparse map, a vehicle can navigate along specific road segments, and in some cases, without needing to interpret the physical aspects of the road, but by aligning its travel path with trajectories (e.g., polynomial splines) along those specific road segments. In this way, vehicles can be navigated primarily based on stored trajectories (e.g., polynomial splines), which may require significantly less storage space compared to methods involving the storage of roadway images, road parameters, road layouts, etc.
[0179] In addition to a stored polynomial representation of the trajectory along a road segment, the disclosed sparse map may also include small data objects that can represent road features. In some embodiments, the small data objects may include digital signatures derived from digital images (or digital signals) acquired by sensors (e.g., cameras or other sensors, such as suspension sensors) mounted on a vehicle traveling along the road segment. The digital signature may have a reduced size relative to the signals acquired by the sensors. In some embodiments, the digital signature may be created to be compatible with a classifier function configured to detect and identify road features from signals acquired by sensors, for example, during subsequent driving. In some embodiments, digital signatures may be created such that they have the smallest possible footprint while maintaining the ability to correlate or match road features with images based on road features (or, if the stored signature is not based on images and / or includes other data, based on digital signals generated by sensors), images of which are subsequently captured by cameras mounted on vehicles traveling along the same road segment.
[0180] In some embodiments, the size of the data object may be further correlated with the uniqueness of the road feature. For example, for a road feature detectable by a camera mounted on a vehicle, and when the camera system mounted on the vehicle is coupled to a classifier capable of classifying the image data corresponding to that road feature into those associated with a specific type of road feature (e.g., road signs), and when such road signs are locally unique in the area (e.g., no identical road signs or road signs of the same type exist nearby), it may be sufficient to store data indicating the type and location of the road feature.
[0181] As will be discussed in further detail below, road features (e.g., landmarks along road segments) can be stored as small data objects that can represent road features with relatively few bytes, while providing sufficient information for identifying and using such features for navigation. In one example, road signs can be identified as identifiable landmarks upon which vehicle navigation can be based. The representation of road signs can be stored in a sparse map to include, for example, a few bytes of data indicating the type of landmark (e.g., a stop sign) and a few bytes of data indicating the location of the landmark (e.g., coordinates). Navigation based on such a data representation of landmarks (e.g., using a representation sufficient for landmark-based localization, identification, and navigation) can provide the desired level of navigation functionality associated with a sparse map without significantly increasing the data overhead associated with the sparse map. This concise representation of landmarks (and other road features) can utilize sensors and processors mounted on such vehicles and configured to detect, identify, and / or classify specific road features.
[0182] When, for example, a sign or even a sign of a particular type is locally unique in a given area (e.g., when there are no other signs, nor other signs of the same type), a sparse map can use data indicating a class of landmarks (signs or signs of a particular type), and during navigation (e.g., autonomous navigation), when a camera mounted on an autonomous vehicle captures an image of an area containing signs (or signs of a particular type), the processor can process the image, detect signs (if they are indeed present in the image), classify the image as signs (or signs of a particular type), and associate the location of the image with the location of the signs, as stored in the sparse map.
[0183] Sparse maps can include any suitable representation of objects identified along road segments. In some cases, objects may be referred to as semantic or non-semantic objects. Semantic objects may include, for example, objects associated with a pre-determined type classification. This type classification is useful in reducing the amount of data required to describe semantic objects identified in the environment, which is beneficial both during the acquisition phase (e.g., reducing the costs associated with bandwidth usage for transmitting driving information from multiple drive acquisition vehicles to a server) and during the navigation phase (e.g., reducing map data can accelerate the transmission of map tiles from the server to the navigation vehicle and also reduce the costs associated with bandwidth usage for such transmissions). Semantic object classification types can be assigned to any type of object or feature expected to be encountered along the roadway.
[0184] Semantic objects can be further divided into two or more logical groups. For example, in some cases, the semantic object type of a group can be associated with a pre-determined dimension. Such semantic objects may include specific speed limit signs, yield signs, lane merging signs, stop signs, traffic lights, directional arrows on a roadway, manhole covers, or any other type of object that can be associated with a normalized size. One benefit of such semantic objects is that very little data may be needed to represent / fully define the object. For example, if the normalized size of the speed limit is known, the acquisition vehicle may only need to identify (through analysis of captured images) the presence of the speed limit sign (the identified type) along with an indication of the location of the detected speed limit sign (e.g., the 2D location in the captured image of the center of the sign or a corner of the sign (or, alternatively, the 3D location in real-world coordinates)) to provide sufficient information for map generation on the server side. In cases where the 2D image location is transmitted to the server, the location associated with the captured image of the detected sign may also be transmitted, so that the server can determine the real-world location of the sign (e.g., by using a structure-in-motion technique from multiple captured images from one or more acquisition vehicles). Even with limited information (requiring only a few bytes to define each detected object), the server can construct a map containing a complete representation of speed limit signs based on the type classification (representing speed limit signs) received from one or more acquisition vehicles, along with location information for the detected signs.
[0185] Semantic objects may also include other identified object or feature types that are not associated with a specific standardized feature. Such objects or features may include potholes, tar joints, lampposts, non-standardized signs, curbs, trees, branches, or any other type of identified object with one or more variable features (e.g., variable dimensions). In such cases, in addition to transmitting to the server an indication of the detected object or feature type (e.g., potholes, poles, etc.) and location information for the detected object or feature, the acquisition vehicle may also transmit an indication of the size of the object or feature. Size can be expressed as 2D image dimensions (e.g., using bounding boxes or one or more size values) or real-world dimensions (determined through structural calculations in motion, based on LiDAR or radar system output, based on trained neural network output, etc.).
[0186] Non-semantic objects or features can include any detectable object or feature that falls outside the identified category or type but can still provide valuable information in map generation. In some cases, such non-semantic features may include a detected corner of a building or the corner of a detected window of a building, a single stone or object near a roadway, concrete splatter in a roadway shoulder, or any other detectable object or feature. When such an object or feature is detected, one or more acquisition vehicles can transmit the location of one or more points (2D image points or 3D real-world points) associated with the detected object / feature to the map generation server. Additionally, compressed or simplified image segments (e.g., image hashes) can be generated for the regions of the captured image that include the detected objects or features. This image hash can be calculated based on a pre-determined image processing algorithm and can form a valid signature for the detected non-semantic objects or features. Such signatures can be useful for navigation relative to sparse maps that include non-semantic features or objects, as vehicles crossing the roadway can apply an algorithm similar to the one used to generate the image hash to confirm / verify the presence of the mapped non-semantic features or objects in the captured image. Using this technique, non-semantic features can be added to the richness of sparse maps (e.g., to enhance their usefulness in navigation) without adding significant data overhead.
[0187] As noted, target trajectories can be stored in a sparse map. These target trajectories (e.g., 3D splines) can represent each available lane for a carriageway, each valid route through an intersection, preferred or recommended paths for merging and exiting, etc. In addition to target trajectories, other road features can also be detected, captured, and incorporated into the sparse map in the form of representational splines. Such features may include, for example, road edges, lane markings, curbs, guardrails, or any other objects or features extending along a carriageway or road segment.
[0188] Generating sparse maps In some embodiments, a sparse map may include at least one line representation of road surface features extending along a road segment and multiple landmarks associated with the road segment. In some aspects, the sparse map may be generated via “crowdsourcing,” for example, through image analysis of multiple images acquired as one or more vehicles pass through the road segment.
[0189] Figure 8 A sparse map 800 is shown that is accessible to one or more vehicles (e.g., vehicle 200, which may be an autonomous vehicle) to provide autonomous vehicle navigation. The sparse map 800 may be stored in a memory (such as memory 140 or 150). Such a memory device may include any type of non-transitory storage device or computer-readable medium. For example, in some embodiments, memory 140 or 150 may include a hard disk drive, optical disk, flash memory, magnetic-based memory device, optical-based memory device, etc. In some embodiments, the sparse map 800 may be stored in a database (e.g., map database 160) that may be stored in memory 140 or 150 or other types of storage devices.
[0190] In some embodiments, the sparse map 800 may be stored on a storage device or a non-transitory computer-readable medium (e.g., storage device included in a navigation system mounted on the vehicle 200) configured to be mounted on the vehicle 200. A processor (e.g., processing unit 110) provided on the vehicle 200 may access the sparse map 800 stored on the storage device or computer-readable medium configured to be mounted on the vehicle 200 in order to generate navigation instructions for guiding the autonomous vehicle 200 as it traverses road segments.
[0191] However, the sparse map 800 does not need to be stored locally relative to the vehicle. In some embodiments, the sparse map 800 may be stored on a storage device or computer-readable medium located on a remote server communicating with the vehicle 200 or a device associated with the vehicle 200. A processor (e.g., processing unit 110) located on the vehicle 200 may receive data included in the sparse map 800 from the remote server and may execute the data to guide autonomous driving of the vehicle 200. In such embodiments, the remote server may store all or only a portion of the sparse map 800. Therefore, a storage device or computer-readable medium configured to be mounted on the vehicle 200 and / or mounted on one or more additional vehicles may store the remaining portion of the sparse map 800.
[0192] Furthermore, in such embodiments, the sparse map 800 can be made accessible to multiple vehicles (e.g., tens, hundreds, thousands, or millions of vehicles) traversing various road segments. It should also be noted that the sparse map 800 may include multiple sub-maps. For example, in some embodiments, the sparse map 800 may include hundreds, thousands, millions, or more sub-maps (e.g., map tiles) available for navigation vehicles. Such sub-maps may be referred to as local maps or map tiles, and vehicles traveling along the roadway can access any number of local maps related to the location the vehicle is traveling on. Local map segments of the sparse map 800 may be stored along with Global Navigation Satellite System (GNSS) keys as an index to the database of the sparse map 800. Therefore, while the steering angle used for navigating the master vehicle in this system can be calculated without relying on the master vehicle's GNSS position, road features, or landmarks, such GNSS information can be used for retrieval of the relevant local maps.
[0193] Generally, sparse map 800 can be generated based on data (e.g., driving information) collected from one or more vehicles as they travel along a roadway. For example, using sensors (e.g., cameras, speedometers, GPS, accelerometers, etc.) mounted on one or more vehicles, the trajectories of one or more vehicles traveling along the roadway can be recorded, and a polynomial representation of the preferred trajectory for subsequent journeys along the roadway can be determined based on the collected trajectories traveled by one or more vehicles. Similarly, data collected by one or more vehicles can help identify potential landmarks along a particular roadway. Data collected from passing vehicles can also be used to identify road contour information, such as road width contours, road roughness contours, traffic line spacing contours, road conditions, etc. Using the collected information, sparse map 800 can be generated and distributed (e.g., for local storage or via real-time data transmission) for navigation of one or more autonomous vehicles. However, in some embodiments, map generation may not end at the initial generation of the map. As will be discussed in more detail below, the sparse map 800 can be continuously or periodically updated based on data collected from vehicles as those vehicles continue to cross the roadway included in the sparse map 800.
[0194] The data recorded in the sparse map 800 may include location information based on Global Positioning System (GPS) data. For example, location information may be included in the sparse map 800 for various map elements, including, for example, landmark locations, road outline locations, etc. The location of a map element included in the sparse map 800 can be obtained using GPS data collected from vehicles crossing a roadway. For example, a vehicle passing a marked landmark can determine the location of the marked landmark using GPS location information associated with the vehicle and a determination of the marked landmark's location relative to the vehicle (e.g., based on image analysis of data collected from one or more cameras mounted on the vehicle). Such location determinations of the marked landmark (or any other feature included in the sparse map 800) may be repeated as additional vehicles pass the location of the marked landmark. Some or all of these additional location determinations may be used to refine the location information relative to the marked landmark stored in the sparse map 800. For example, in some embodiments, multiple location measurements relative to a specific feature stored in the sparse map 800 may be averaged together. However, any other mathematical operations may also be used to refine the stored location of a map element based on multiple determined locations for that map element.
[0195] In a specific example, the acquisition vehicles may traverse specific road segments. Each acquisition vehicle captures images of its corresponding environment. Images can be collected at any suitable frame capture rate (e.g., 9 Hz, etc.). An image analysis processor mounted on each acquisition vehicle analyzes the captured images to detect the presence of semantic and / or non-semantic features / objects. At a high level, the acquisition vehicles transmit indications of the detection of semantic and / or non-semantic objects / features, along with the locations associated with those objects / features, to the mapping server. More specifically, type indicators, size indicators, etc., may be transmitted along with location information. Location information may include any suitable information that enables the mapping server to aggregate the detected objects / features into a sparse map useful for navigation. In some cases, location information may include one or more 2D image locations (e.g., XY pixel locations) in the captured images where semantic or non-semantic features / objects were detected. Such image locations may correspond to the center, corner, etc., of a feature / object. In this scenario, each acquisition vehicle may also provide the server with the location (e.g., GPS location) where each image was captured to help the mapping server reconstruct driving information and align driving information from multiple acquisition vehicles.
[0196] In other cases, the acquisition vehicle may provide a server with one or more 3D real-world points associated with the detected objects / features. These 3D points may be relative to a predetermined origin (such as the origin of a driving segment) and can be determined by any suitable technique. In some cases, structure-in-motion techniques can be used to determine the 3D real-world location of the detected objects / features. For example, a specific object, such as a specific speed limit sign, can be detected in two or more captured images. Using information such as the known self-motion of the acquisition vehicle between captured images (speed, trajectory, GPS location, etc.), along with observed changes in the speed limit sign in the captured images (changes in XY pixel location, size, etc.), the real-world location of one or more points associated with the speed limit sign can be determined and transmitted to a mapping server. This method is optional because it requires more computation on the components of the acquisition vehicle system. The sparse maps of the disclosed embodiments can enable autonomous navigation of the vehicle using a relatively small amount of stored data. In some embodiments, the sparse map 800 may have a data density of less than 2 MB per kilometer of road, less than 1 MB per kilometer of road, less than 500 kB per kilometer of road, or less than 100 kB per kilometer of road (e.g., including data representing target trajectories, landmarks, and any other stored road features). In some embodiments, the data density of the sparse map 800 may be less than 10 kB per kilometer of road or even less than 2 kB per kilometer of road (e.g., 1.6 kB per kilometer), or no more than 10 kB per kilometer of road, or no more than 20 kB per kilometer of road. In some embodiments, most (if not all) of the roadways in the United States may be autonomously navigated using a sparse map with a total of 4 GB or less of data. These data density values may represent averages across the entire sparse map 800, on local maps within the sparse map 800, and / or on specific road segments within the sparse map 800.
[0197] As noted, the sparse map 800 may include representations of multiple target trajectories 810 for guiding autonomous driving or navigation along road segments. Such target trajectories may be stored as three-dimensional splines. For example, target trajectories stored in the sparse map 800 may be determined based on two or more reconstructed trajectories previously traveled by the vehicle along a particular road segment. Road segments may be associated with a single target trajectory or multiple target trajectories. For example, on a two-lane road, a first target trajectory may be stored to represent the expected travel path along the road in a first direction, and a second target trajectory may be stored to represent the expected travel path along the road in another direction (e.g., opposite to the first direction). Additional target trajectories may be stored relative to a particular road segment. For example, on a multi-lane road, one or more target trajectories representing the expected travel path of a vehicle in one or more lanes associated with the multi-lane road may be stored. In some embodiments, each lane of the multi-lane road may be associated with its own target trajectory. In other embodiments, the number of target trajectories stored may be fewer than the number of lanes present on the multi-lane road. In such cases, a vehicle navigating on a multi-lane road can use any stored target trajectory to guide its navigation by taking into account the lane offset relative to the lane for which the target trajectory is stored (for example, if a vehicle is traveling in the leftmost lane of a three-lane highway and the target trajectory is stored only for the middle lane of the highway, when generating navigation instructions, the vehicle can use the target trajectory of the middle lane to navigate by taking into account the lane offset between the middle lane and the leftmost lane).
[0198] In some embodiments, the target trajectory may represent the ideal path that a vehicle should take as it travels. The target trajectory may be located, for example, at the approximate center of a driving lane. In other cases, the target trajectory may be located at other locations relative to a road segment. For example, the target trajectory may coincide approximately with the center of the road, the edge of the road, or the edge of a lane. In such cases, navigation based on the target trajectory may include a defined offset to be maintained relative to the location of the target trajectory. Furthermore, in some embodiments, the defined offset to be maintained relative to the location of the target trajectory may vary based on the type of vehicle (e.g., a passenger vehicle comprising two axles may have an offset along at least a portion of the target trajectory that differs from that of a truck comprising more than two axles).
[0199] The sparse map 800 may also include data associated with multiple pre-determined landmarks 820 linked to specific road segments, local maps, etc. These landmarks can be used for navigation of the autonomous vehicle, as discussed in more detail below. For example, in some embodiments, landmarks can be used to determine the vehicle's current position relative to a stored target trajectory. Using this positional information, the autonomous vehicle may be able to adjust its heading to match the direction of the target trajectory at the determined location.
[0200] Multiple landmarks 820 can be identified and stored in the sparse map 800 at any suitable spacing. In some embodiments, landmarks can be stored at a relatively high density (e.g., every few meters or more). However, in some embodiments, significantly larger landmark spacing values may be used. For example, in the sparse map 800, identified (or recognized) landmarks may be spaced 10 meters, 20 meters, 50 meters, 100 meters, 1 kilometer, or 2 kilometers apart. In some cases, identified landmarks may be located at distances even exceeding 2 kilometers.
[0201] Between landmarks, and therefore between the determination of the vehicle's position relative to a target trajectory, the vehicle can navigate based on dead reckoning, whereby the vehicle uses sensors to determine its own motion and estimate its position relative to the target trajectory. Since errors can accumulate during navigation by dead reckoning, the position determination relative to the target trajectory may become increasingly inaccurate over time. The vehicle can use landmarks (and their known locations) appearing in the sparse map 800 to remove dead reckoning-induced errors in position determination. In this way, the landmarks included in the sparse map 800 can act as navigation anchors from which the vehicle's accurate position relative to the target trajectory can be determined. Since a certain amount of error in the location is acceptable, the marked landmarks do not always need to be available to autonomous vehicles. Instead, suitable navigation can even be based on landmark spacing of 10 meters, 20 meters, 50 meters, 100 meters, 500 meters, 1 kilometer, 2 kilometers, or more, as noted above. In some embodiments, a density of one marked landmark per 1 km of road may be sufficient to maintain longitudinal position determination accuracy within 1 m. Therefore, not every potential landmark that appears along a road segment needs to be stored in the sparse map 800.
[0202] Furthermore, in some embodiments, lane markings can be used for vehicle positioning during landmark intervals. By using lane markings during landmark intervals, the accumulation of errors during navigation based on dead reckoning can be minimized.
[0203] In addition to target trajectories and marked landmarks, sparse maps 800 can also include information related to various other road features. For example, Figure 9A A representation of a curve along a specific road segment, which can be stored in a sparse map 800, is shown. In some embodiments, a single lane of the road can be modeled using a three-dimensional polynomial description of the left and right sides of the road. Such polynomials representing the left and right sides of a single lane are shown in... Figure 9A As shown in the diagram. Regardless of how many lanes a road can have, it can be similar to... Figure 9A The representation of roads uses polynomials. For example, the left and right sides of a multi-lane road can be represented by a polynomial similar to... Figure 9AThe polynomials shown are represented by polynomials, and also include the middle lane markings on multi-lane roads (e.g., dashed lines indicating lane boundaries, solid yellow lines indicating boundaries between lanes traveling in different directions, etc.) can also be represented using polynomials such as... Figure 9A The polynomials shown are used to represent this.
[0204] like Figure 9A As shown, lane 900 can be represented using a polynomial (e.g., a first-order, second-order, third-order, or any suitable order polynomial). For illustration, lane 900 is shown as a two-dimensional lane and the polynomial is shown as a two-dimensional polynomial. Figure 9A The lane 900 depicted includes a left lane 910 and a right lane 920. In some embodiments, more than one polynomial may be used to represent the location on each side of the road or lane boundary. For example, each of the left lane 910 and right lane 920 may be represented by multiple polynomials of any suitable length. In some cases, the polynomials may have a length of approximately 100 m, but other lengths greater than or less than 100 m may also be used. Additionally, the polynomials may overlap each other to facilitate seamless transitions when the primary vehicle is navigating along the lane based on subsequently encountered polynomials. For example, each of the left lane 910 and right lane 920 may be represented by multiple third-order polynomials divided into segments of approximately 100 meters in length (an example of a first predetermined range) and overlapping each other by approximately 50 meters. The polynomials representing the left lane 910 and right lane 920 may or may not have the same order. For example, in some embodiments, some polynomials may be second-order polynomials, some may be third-order polynomials, and some may be fourth-order polynomials.
[0205] exist Figure 9A In the example shown, the left side 910 of lane 900 is represented by two groups of third-order polynomials. The first group comprises polynomial segments 911, 912, and 913. The second group comprises polynomial segments 914, 915, and 916. The two groups are substantially parallel to each other, but follow the locations on their respective sides of the road. Polynomial segments 911, 912, 913, 914, 915, and 916 are approximately 100 meters long and overlap approximately 50 meters with adjacent segments in the series. However, as previously noted, polynomials of different lengths and different amounts of overlap can also be used. For example, polynomials can have lengths of 500 m, 1 km, or longer, and the overlap can vary from 0 m to 50 m, 50 m to 100 m, or greater than 100 m. Furthermore, although... Figure 9A These are shown as polynomials extending in 2D space (e.g., on the surface of paper), but it should be understood that these polynomials can represent curves extending in three dimensions (e.g., including a height component) to represent elevation changes in road segments in addition to XY curvature. Figure 9A In the example shown, the right side 920 of lane 900 is further represented by a first group having polynomial segments 921, 922 and 923 and a second group having polynomial segments 924, 925 and 926.
[0206] Return to the target trajectory on the sparse map 800. Figure 9B A three-dimensional polynomial representing the target trajectory of a vehicle traveling along a specific road segment is shown. The target trajectory represents not only the XY path the main vehicle should travel along the specific road segment, but also the elevation changes the main vehicle will experience as it travels along that road segment. Therefore, each target trajectory in the sparse map 800 can be represented by one or more three-dimensional polynomials, such as... Figure 9B The three-dimensional polynomial 950 is shown. The sparse map 800 may include multiple trajectories (e.g., millions or billions or more to represent the trajectories of vehicles along various road segments of a roadway throughout the world). In some embodiments, each target trajectory may correspond to a spline connecting the segments of the three-dimensional polynomial.
[0207] Regarding the data footprint of the polynomial curves stored in the sparse map 800, in some embodiments, each cubic polynomial can be represented by four parameters, each requiring four bytes of data. A suitable representation can be obtained using a third-order polynomial that requires approximately 192 bytes of data per 100 m. For a master vehicle traveling at approximately 100 km / hr, this translates to approximately 200 kB per hour in data usage / transmission requirements.
[0208] Sparse maps 800 can describe lane networks using a combination of geometric feature descriptors and metadata. Geometric features can be described using polynomials or splines as described above. Metadata can describe the number of lanes, special features (such as carpool lanes), and possibly other sparse labels. The total footprint of such indicators may be negligible.
[0209] Therefore, a sparse map according to embodiments of this disclosure may include at least one line representation of road surface features extending along road segments, each line representation representing a path along a road segment substantially corresponding to a road surface feature. In some embodiments, as discussed above, at least one line representation of a road surface feature may include a spline, a polynomial representation, or a curve. Furthermore, in some embodiments, the road surface feature may include at least one of a road edge or lane markings. Additionally, as discussed below with respect to “crowdsourcing,” road surface features may be identified through image analysis of multiple images acquired as one or more vehicles traverse a road segment.
[0210] As previously indicated, the sparse map 800 may include multiple pre-determined landmarks associated with road segments. Each landmark in the sparse map 800 can be represented and identified using less data than stored actual images, rather than storing actual images of the landmark and relying on image recognition analysis, for example, based on captured and stored images. The data representing the landmarks may still include sufficient information to describe or identify the landmarks along the road. Storing data describing the characteristics of the landmarks instead of actual images of the landmarks reduces the size of the sparse map 800.
[0211] Figure 10 Examples of the types of landmarks that can be represented in sparse map 800 are shown. Landmarks can include any visible and identifiable object along a road segment. Landmarks can be selected such that they are fixed and do not change frequently relative to their location and / or content. Landmarks included in sparse map 800 can be useful in determining the location of vehicle 200 relative to a target trajectory as a vehicle 200 traverses a particular road segment. Examples of landmarks can include traffic signs, directional signs, general signs (e.g., rectangular signs), roadside fixtures (e.g., lampposts, reflectors, etc.), and any other suitable categories. In some embodiments, lane markings on the road may also be included as landmarks in sparse map 800.
[0212] Figure 10 Examples of landmarks shown include traffic signs, directional signs, roadside fixtures, and general signs. Traffic signs may include, for example, speed limit signs (e.g., speed limit sign 1000), yield signs (e.g., yield sign 1005), route number signs (e.g., route number sign 1010), traffic light signs (e.g., traffic light sign 1015), and stop signs (e.g., stop sign 1020). Directional signs may include signs that include one or more arrows indicating one or more directions to different places. For example, directional signs may include a highway sign 1025 with arrows to guide vehicles to different roads or places, an exit sign 1030 with arrows to guide vehicles off the road, etc. Therefore, at least one of a plurality of landmarks may include a road sign.
[0213] General signs may not be related to traffic. For example, general signs may include billboards used for advertising, or welcome signs located at the border between two countries, states, counties, cities, or towns. Figure 10 The general sign 1040 (“Joe's Restaurant”) is shown. While the general sign 1040 may have a rectangular shape, such as... Figure 10 As shown, however, the general mark 1040 may have other shapes, such as square, circle, triangle, etc.
[0214] Landmarks may also include roadside fixtures. Roadside fixtures may not be the subject of the sign and may not be related to traffic or direction. For example, roadside fixtures may include lampposts (e.g., lamppost 1035), utility poles, traffic light poles, etc.
[0215] Landmarks may also include beacons specifically designed for autonomous vehicle navigation systems. For example, such beacons may include freestanding structures placed at predetermined intervals to aid in the navigation of the host vehicle. Such beacons may also include visual / graphical information added to existing road signs (e.g., icons, symbols, barcodes, etc.) that can be identified or recognized by vehicles traveling along road segments. Such beacons may also include electronic components. In such embodiments, electronic beacons (e.g., RFID tags, etc.) may be used to transmit non-visual information to the host vehicle. This information may include, for example, landmark identification and / or landmark location information that the host vehicle can use to determine its position along a target trajectory.
[0216] In some embodiments, landmarks included in the sparse map 800 may be represented by data objects of a predetermined size. The data representing a landmark may include any suitable parameters used to identify a particular landmark. For example, in some embodiments, landmarks stored in the sparse map 800 may include parameters such as the physical size of the landmark (e.g., to support estimation of distances to the landmark based on a known size / scale), distance to previous landmarks, lateral offset, height, type code (e.g., landmark type—what type of directional sign, traffic sign, etc.), GPS coordinates (e.g., to support Global Positioning), and any other suitable parameters. Each parameter may be associated with a data size. For example, 8 bytes of data may be used to store the landmark size. The distance to previous landmarks, lateral offset, and height may be specified using 12 bytes of data. The type code associated with a landmark, such as a directional sign or traffic sign, may require approximately 2 bytes of data. For general landmarks, 50 bytes of data storage may be used to store an image signature capable of identifying a general landmark. The landmark's GPS location may be associated with 16 bytes of data storage. These data sizes for each parameter are merely examples, and other data sizes may also be used. Representing landmarks in a sparse map 800 in this way provides a lean solution for efficiently representing landmarks in a database. In some embodiments, objects may be referred to as standard semantic objects or non-standard semantic objects. Standard semantic objects may include any kind of object for which there is a standardized set of characteristics (e.g., speed limit signs, warning signs, directional signs, traffic lights, etc., with known dimensions or other characteristics). Non-standard semantic objects may include any object not associated with a standardized set of characteristics (e.g., general advertising signs, signs identifying commercial establishments, potholes, trees, etc., which may have variable dimensions). Each non-standard semantic object may utilize 38 bytes of data (e.g., size 8 bytes; distance to previous landmarks, lateral offset, and height 12 bytes; type code 2 bytes; and location coordinates 16 bytes). Standard semantic objects may be represented using even less data because the map rendering server may not need size information to fully represent objects in a sparse map.
[0217] The sparse map 800 can use a labeling system to represent landmark types. In some cases, each traffic sign or directional sign can be associated with its own label, which is stored in a database as part of the landmark identifier. For example, the database may include nearly 1,000 different labels to represent various traffic signs and nearly 10,000 different labels to represent directional signs. Of course, any suitable number of labels can be used, and additional labels can be created as needed. In some embodiments, a generic sign can be used with less than about 100 bytes (e.g., about 86 bytes, including 8 bytes of size; 12 bytes for distance, lateral offset, and height to previous landmarks; 50 bytes for image signature; and 16 bytes for GPS coordinates).
[0218] Therefore, for semantic road signs that do not require image signatures, even at a relatively high landmark density of approximately one per 50 m, the data density impact on the sparse map 800 is likely to be close to approximately 760 bytes per kilometer (e.g., 20 landmarks per km x 38 bytes per landmark = 760 bytes). Even for generic signs that include image signature components, the data density impact is approximately 1.72 kB per km (e.g., 20 landmarks per km x 86 bytes per landmark = 1,720 bytes). For semantic road signs, this equates to approximately 76 kB of data usage per hour for a vehicle traveling at 100 km / hr. For generic signs, this equates to approximately 170 kB per hour for a vehicle traveling at 100 km / hr. It should be noted that in some environments (e.g., urban environments), the density of detectable objects available for inclusion in the sparse map may be much higher (potentially more than one per meter). In some embodiments, generally rectangular objects (such as rectangular signs) may be represented in the sparse map 800 by no more than 100 bytes of data. The representation of a roughly rectangular object (e.g., a general sign 1040) in the sparse map 800 may include a compressed image signature or image hash (e.g., compressed image signature 1045) associated with the roughly rectangular object. This compressed image signature / image hash can be determined using any suitable image hashing algorithm and can be used, for example, to help identify a general sign as a landmark. Such compressed image signatures (e.g., image information derived from actual image data representing the object) avoid the need for storing actual images of the object or for comparative image analysis of actual images to identify landmarks.
[0219] refer to Figure 10The sparse map 800 may include or store a compressed image signature 1045 associated with the general sign 1040, rather than an actual image of the general sign 1040. For example, after an image capture device (e.g., image capture device 122, 124, or 126) captures an image of the general sign 1040, a processor (e.g., image processor 190 or any other processor capable of processing images mounted on or remotely positioned relative to the host vehicle) may perform image analysis to extract / create the compressed image signature 1045, which includes a unique signature or pattern associated with the general sign 1040. In one embodiment, the compressed image signature 1045 may include a shape, color pattern, brightness pattern, or any other feature that can be extracted from the image of the general sign 1040 to describe the general sign 1040.
[0220] For example, in Figure 10 In the compressed image signature 1045, circles, triangles, and stars can represent areas of different colors. Patterns represented by circles, triangles, and stars can be stored in the sparse map 800, for example, within 50 bytes designated to include the image signature. It is important to note that circles, triangles, and stars do not necessarily indicate that such shapes are stored as part of the image signature. Rather, these shapes conceptually represent identifiable areas with discernible color differences, text areas, graphic shapes, or other variations of features that can be associated with general signs. Such compressed image signatures can be used to identify landmarks in the form of general signs. For example, compressed image signatures can be used for similarity and difference analysis based on comparisons between stored compressed image signatures and image data, for example, captured using cameras mounted on autonomous vehicles.
[0221] Therefore, multiple landmarks can be identified through image analysis of multiple images acquired when one or more vehicles pass through a road segment. As explained below regarding “crowdsourcing,” in some embodiments, the image analysis used to identify multiple landmarks may include accepting potential landmarks when the ratio of images in which landmarks appear to images in which landmarks do not appear exceeds a threshold. Furthermore, in some embodiments, the image analysis used to identify multiple landmarks may include rejecting potential landmarks when the ratio of images in which landmarks do not appear to images in which landmarks appear exceeds a threshold.
[0222] Returning to the main vehicle can be used to navigate a target trajectory for a specific road segment. Figure 11AThe diagram illustrates a polynomial representation of the trajectory captured during the process of building or maintaining the sparse map 800. The polynomial representation of the target trajectory included in the sparse map 800 may be determined based on two or more reconstructed trajectories of a vehicle's previous travel along the same road segment. In some embodiments, the polynomial representation of the target trajectory included in the sparse map 800 may be an aggregation of two or more reconstructed trajectories of a vehicle's previous travel along the same road segment. In some embodiments, the polynomial representation of the target trajectory included in the sparse map 800 may be the average of two or more reconstructed trajectories of a vehicle's previous travel along the same road segment. Other mathematical operations may also be used to construct the target trajectory along the road path based on the reconstructed trajectories collected from vehicles traveling along the road segment.
[0223] like Figure 11A As shown, road segment 1100 can be traveled by multiple vehicles 200 at different times. Each vehicle 200 can collect data related to the path traveled by the vehicle along the road segment. The path traveled by a particular vehicle can be determined based on camera data, accelerometer information, speed sensor information and / or GPS information, as well as other potential sources. Such data can be used to reconstruct the trajectory of the vehicle traveling along the road segment, and based on these reconstructed trajectories, a target trajectory (or multiple target trajectories) can be determined for a particular road segment. Such target trajectories can represent the preferred path of the master vehicle when the vehicle travels along the road segment (e.g., guided by an autonomous navigation system).
[0224] exist Figure 11A In the example shown, a first reconstructed trajectory 1101 may be determined based on data received from a first vehicle traversing road segment 1100 during a first time period (e.g., day 1), a second reconstructed trajectory 1102 may be obtained from a second vehicle traversing road segment 1100 during a second time period (e.g., day 2), and a third reconstructed trajectory 1103 may be obtained from a third vehicle traversing road segment 1100 during a third time period (e.g., day 3). Each trajectory 1101, 1102, and 1103 may be represented by a polynomial trajectory, such as a three-dimensional polynomial trajectory. It should be noted that in some embodiments, any of the reconstructed trajectories may be assembled and mounted on a vehicle traversing road segment 1100.
[0225] Alternatively or additionally, such reconstructed trajectories can be determined on the server side based on information received from vehicles traversing road segment 1100. For example, in some embodiments, vehicles 200 may transmit data related to their movement along road segment 1100 (e.g., steering angle, heading, time, position, speed, sensed road geometry and / or sensed landmarks, etc.) to one or more servers. Servers can reconstruct a trajectory for vehicle 200 based on the received data. Servers may also generate target trajectories for navigation of autonomous vehicles that will later travel along the same road segment 1100, based on a first trajectory 1101, a second trajectory 1102, and a third trajectory 1103. While target trajectories may be associated with a single previous travel along the road segment, in some embodiments, each target trajectory included in the sparse map 800 may be determined based on two or more reconstructed trajectories of vehicles traversing the same road segment. Figure 11A In this context, the target trajectory is represented by 1110. In some embodiments, the target trajectory 1110 may be generated based on the average of a first trajectory 1101, a second trajectory 1102, and a third trajectory 1103. In some embodiments, the target trajectory 1110 included in the sparse map 800 may be an aggregation (e.g., a weighted combination) of two or more reconstructed trajectories.
[0226] At the map rendering server, the server can receive actual trajectories for a specific road segment from multiple acquisition vehicles traversing that road segment. To generate a target trajectory for each valid path along the road segment (e.g., each lane, each driving direction, each path through an intersection, etc.), the received actual trajectories can be aligned. The alignment process may include correlating the actual, acquired trajectories with each other using detected objects / features identified along the road segment and the acquisition locations of those detected objects / features. Once aligned, an average or "best-fit" target trajectory for each available lane can be determined based on the aggregated, correlated / aligned actual trajectories.
[0227] Figure 11B and Figure 11C The concept of target trajectories associated with road segments existing within geographic region 1111 is further illustrated. For example... Figure 11BAs shown, a first road segment 1120 within geographic region 1111 may include a multi-lane road comprising two lanes 1122 designated for vehicles traveling in a first direction and two additional lanes 1124 designated for vehicles traveling in a second direction opposite to the first direction. Lanes 1122 and lanes 1124 may be separated by double yellow lines 1123. Geographic region 1111 may also include branch road segments 1130 intersecting with road segment 1120. Road segment 1130 may include a two-lane road, each lane designated for a different direction of travel. Geographic region 1111 may also include other road features such as stop lines 1132, stop signs 1134, speed limit signs 1136, and hazard signs 1138.
[0228] like Figure 11C As shown, the sparse map 800 may include a local map 1140, which includes a road model for assisting autonomous navigation of vehicles within geographic region 1111. For example, the local map 1140 may include target trajectories for one or more lanes associated with road segments 1120 and / or 1130 within geographic region 1111. For instance, the local map 1140 may include target trajectories 1141 and / or 1142 that the autonomous vehicle can access or rely on when crossing lane 1122. Similarly, the local map 1140 may include target trajectories 1143 and / or 1144 that the autonomous vehicle can access or rely on when crossing lane 1124. Furthermore, the local map 1140 may include target trajectories 1145 and / or 1146 that the autonomous vehicle can access or rely on when crossing road segment 1130. Target trajectory 1147 represents the preferred path that the autonomous vehicle should follow when transitioning from lane 1120 (and specifically, relative to target trajectory 1141 associated with the rightmost lane of lane 1120) to road segment 1130 (and specifically, relative to target trajectory 1145 associated with the first side of road segment 1130). Similarly, target trajectory 1148 represents the preferred path that the autonomous vehicle should follow when transitioning from road segment 1130 (and specifically, relative to target trajectory 1146) to a portion of road segment 1124 (and specifically, as shown, relative to target trajectory 1143 associated with the left lane of lane 1124).
[0229] The sparse map 800 may also include representations of other road-related features associated with geographic region 1111. For example, the sparse map 800 may also include representations of one or more landmarks identified in geographic region 1111. Such landmarks may include a first landmark 1150 associated with stop line 1132, a second landmark 1152 associated with stop sign 1134, a third landmark 1154 associated with speed limit sign, and a fourth landmark 1156 associated with hazard sign 1138. Such landmarks can be used, for example, to assist an autonomous vehicle in determining its current position relative to any of the indicated target trajectories, allowing the vehicle to adjust its heading to match the direction of the target trajectory at the determined location.
[0230] In some embodiments, the sparse map 800 may also include road signature profiles. Such road signature profiles may be associated with any identifiable / measurable variation in at least one parameter associated with a road. For example, in some cases, such profiles may be associated with variations in road surface information, such as variations in the surface roughness of a particular road segment, variations in the road width above a particular road segment, variations in the distance between dashed lines drawn along a particular road segment, variations in the road curvature along a particular road segment, etc. Figure 11D An example of a road signature profile 1160 is shown. While profile 1160 may represent any of the parameters mentioned above or other parameters, in one example, profile 1160 may represent a measure of road surface roughness, such as that obtained, for example, by monitoring one or more sensors that provide outputs indicating the amount of suspension displacement as the vehicle travels on a particular road segment.
[0231] Alternatively or simultaneously, contour 1160 may represent a change in road width, as determined based on image data obtained via a camera mounted on a vehicle traveling along a particular road segment. For example, such a contour can be used to determine a specific location of an autonomous vehicle relative to a particular target trajectory. That is, a contour that is measurable and associated with one or more parameters of the road segment as the autonomous vehicle traverses it. If the measured contour can be correlated / matched with a predetermined contour that plots the parameter changes relative to the position along the road segment, the measured contour and the predetermined contour (e.g., by overlaying corresponding segments of the measured contour and the predetermined contour) can be used to determine the current position along the road segment, and thus the current position relative to a target trajectory for the road segment.
[0232] In some embodiments, the sparse map 800 may include different trajectories based on different characteristics, environmental conditions, and / or other driving-related parameters associated with the user of the autonomous vehicle. For example, in some embodiments, different trajectories may be generated based on different user preferences and / or profiles. The sparse map 800 including such different trajectories may be provided to different autonomous vehicles of different users. For example, some users may prefer to avoid toll roads, while others may prefer to take the shortest or fastest route, regardless of whether there are toll roads on the route. The disclosed system may generate different sparse maps with different trajectories based on such different user preferences or profiles. As another example, some users may prefer to drive in fast-moving lanes, while others may prefer to always stay in the center lane.
[0233] Different trajectories can be generated and included in the sparse map 800 based on different environmental conditions (such as day and night, snow, rain, fog, etc.). The sparse map 800 generated based on these different environmental conditions can be provided to autonomous vehicles driving under different environmental conditions. In some embodiments, cameras mounted on the autonomous vehicle can detect environmental conditions and provide this information back to the server that generated and provided the sparse map. For example, the server can generate or update the already generated sparse map 800 to include trajectories that may be more suitable or safer for autonomous driving under the detected environmental conditions. The sparse map 800 can be dynamically updated based on environmental conditions as the autonomous vehicle travels along the road.
[0234] Other driving-related parameters can also be used as a basis for generating different sparse maps and providing them to different autonomous vehicles. For example, when an autonomous vehicle is traveling at high speed, cornering may be tighter. Trajectories associated with specific lanes rather than the road can be included in the sparse map 800, allowing the autonomous vehicle to remain within a specific lane while following a specific trajectory. When images captured by cameras mounted on the autonomous vehicle indicate that the vehicle has deviated from its lane (e.g., crossed lane markings), actions can be triggered within the vehicle to bring it back to the designated lane according to the specific trajectory.
[0235] Crowdsourcing of sparse maps The publicly available sparse maps can be generated efficiently (and passively) through the power of crowdsourcing. For example, any private or commercial vehicle equipped with a camera (e.g., a simple low-resolution camera typically included as OEM equipment in today's vehicles) and a suitable image analysis processor can be used as a data collection vehicle. No special equipment (e.g., high-resolution imaging and / or positioning systems) is required. Due to the publicly available crowdsourcing techniques, the generated sparse maps can be extremely accurate and can include extremely fine location information (achieving navigation error limits of 10 cm or less) without requiring any specialized imaging or sensing equipment as input to the map generation process. Crowdsourcing also enables faster (and cheaper) updates to the generated maps, as new driving information from any roads traversed by a private or commercial vehicle minimally equipped to also act as a data collection vehicle is continuously available to the mapping server system. No specific vehicle equipped with high-resolution imaging and mapping sensors is required. Therefore, the costs associated with building such specialized vehicles can be avoided. Furthermore, updating currently available sparse maps can be much faster than systems that rely on dedicated, specialized mapping vehicles (which are typically limited to a fleet of dedicated vehicles in numbers far fewer than the number of private or commercial vehicles already available for the publicly available data acquisition techniques due to their cost and specialized equipment).
[0236] Publicly available sparse maps generated through crowdsourcing can be extremely accurate because they can be generated based on numerous inputs from multiple (dozens, hundreds, millions, etc.) collection vehicles that have already collected driving information along a specific road segment. For example, each collection vehicle driving along a specific road segment can record its actual trajectory and determine the positional information of detected objects / features relative to the road segment. This information is passed from multiple collection vehicles to a server. The actual trajectories are aggregated to generate a refined target trajectory for each valid driving path along the road segment. Additionally, the positional information (semantic or non-semantic) of each detected object / feature collected from multiple collection vehicles for each road segment can also be aggregated. Therefore, the mapped position of each detected object / feature can constitute an average of hundreds, thousands, or millions of individually determined positions for each detected object / feature. Such techniques can produce extremely accurate mapped positions for detected objects / features.
[0237] In some embodiments, the disclosed systems and methods can generate sparse maps for autonomous vehicle navigation. For example, the disclosed systems and methods can use crowdsourced data to generate sparse maps that one or more autonomous vehicles can use to navigate along a road system. As used herein, “crowdsourcing” means receiving data from various vehicles (e.g., autonomous vehicles) traveling on a road segment at different times, and such data is used to generate and / or update a road model, including sparse map tiles. This model, or any of its sparse map tiles, can then be transmitted to vehicles or other vehicles traveling later along the road segment to assist autonomous vehicle navigation. The road model may include multiple target trajectories representing preferred paths that an autonomous vehicle should follow when traversing a road segment. The target trajectories may be identical to reconstructed actual trajectories collected from vehicles traversing the road segment, which may be transmitted from the vehicles to a server. In some embodiments, the target trajectories may differ from actual trajectories previously taken by one or more vehicles when traversing the road segment. The target trajectories may be generated based on actual trajectories (e.g., through averaging or any other suitable operation).
[0238] Vehicle trajectory data that a vehicle can upload to the server can correspond to either its actual reconstructed trajectory or a recommended trajectory. This recommended trajectory may be based on or related to the vehicle's actual reconstructed trajectory, but may differ from it. For example, a vehicle can modify its actual, reconstructed trajectory and submit (e.g., recommend) a modified actual trajectory to the server. The road model can then use the recommended, modified trajectory as a target trajectory for autonomous navigation of other vehicles.
[0239] In addition to trajectory information, other information that may be used in building a sparse data map 800 may include information related to potential landmark candidates. For example, through information crowdsourcing, the disclosed systems and methods can identify potential landmarks in the environment and refine landmark locations. Landmarks can be used by the navigation system of autonomous vehicles to determine and / or adjust the vehicle's position along the target trajectory.
[0240] A reconstructed trajectory that can be generated as a vehicle travels along a road can be obtained by any suitable method. In some embodiments, the reconstructed trajectory can be developed by stitching together segments of the vehicle's motion using, for example, self-motion estimation (e.g., three-dimensional translation and three-dimensional rotation of a camera, and thus the vehicle's body). Rotation and translation estimations can be determined based on analysis of images captured by one or more image capture devices, along with information from other sensors or devices such as inertial sensors and velocity sensors. For example, inertial sensors may include accelerometers or other suitable sensors configured to measure changes in the translation and / or rotation of the vehicle's body. The vehicle may include a velocity sensor that measures the vehicle's speed.
[0241] In some embodiments, the ego motion of the camera (and therefore the vehicle body) can be estimated based on optical flow analysis of captured images. Optical flow analysis of an image sequence identifies the movement of pixels within that sequence, and the vehicle's motion is determined based on the identified movement. The ego motion can be integrated over time and along road segments to reconstruct a trajectory associated with the road segments the vehicle has already followed.
[0242] Data collected by multiple vehicles at different times during multiple drives along a road segment (e.g., reconstructed trajectories) can be used to construct a road model (e.g., including target trajectories, etc.) included in a sparse data map 800. Data collected by multiple vehicles at different times during multiple drives along a road segment can also be averaged to improve model accuracy. In some embodiments, data regarding road geometry and / or landmarks can be received from multiple vehicles traveling through a common road segment at different times. Such data received from different vehicles can be combined to generate a road model and / or to update the road model.
[0243] The geometric features of the reconstructed trajectory (and target trajectory) along the road segment can be represented by a curve in three-dimensional space, which can be a spline connecting three-dimensional polynomials. The reconstructed trajectory curve can be determined from the analysis of a video stream or multiple images captured by a camera mounted on the vehicle. In some embodiments, a location a few meters ahead of the vehicle's current position is identified in each frame or image. This location is the position the vehicle is expected to reach within a predetermined time period. This operation can be repeated frame by frame, while the vehicle can simultaneously calculate the camera's self-motion (rotation and translation). At each frame or image, a short-range model of the desired path is generated by the vehicle in a reference frame attached to the camera. The short-range models can be stitched together to obtain a three-dimensional model of the road in a coordinate system, which can be arbitrary or predetermined. The three-dimensional model of the road can then be fitted using splines, which can include or connect one or more polynomials of suitable order.
[0244] To derive a short-range road model for each frame, one or more detection modules can be used. For example, a bottom-up lane detection module can be used. A bottom-up lane detection module can be useful when drawing lane markings on a road. This module finds edges in the image and assembles them to form lane markings. A second module can be used in conjunction with the bottom-up lane detection module. The second module is an end-to-end deep neural network that can be trained to predict the correct short-range path from the input image. In both modules, the road model can be detected in the image coordinate system and converted into a 3D space that can be virtually attached to the camera.
[0245] While the reconstructed trajectory modeling method may introduce accumulated errors due to the integration of self-motion over long time periods, which may include noise components, such errors are likely insignificant because the resulting model provides sufficient accuracy for navigation at local scales. Furthermore, integrated errors can be eliminated by using external information sources, such as satellite imagery or geodetic results. For example, the disclosed systems and methods may use a GNSS receiver to eliminate accumulated errors. However, GNSS positioning signals may not always be available and accurate. The disclosed systems and methods enable steering applications that are weakly dependent on the availability and accuracy of GNSS positioning. In such systems, the use of GNSS signals may be limited. For example, in some embodiments, the disclosed system may use GNSS signals solely for database indexing purposes.
[0246] In some embodiments, the distance scale (e.g., a local scale) associated with autonomous vehicle navigation and steering applications can be approximately 50 meters, 100 meters, 200 meters, 300 meters, etc. Such distances can be used because the geometric road model is primarily used for two purposes: planning the preceding trajectory and locating the vehicle on the road model. In some embodiments, when the control algorithm steers the vehicle based on a target point located 1.3 seconds ahead (or any other suitable distance, such as 1.5 seconds, 1.7 seconds, 2 seconds, etc.), the planning task can use a model within a typical range of 40 meters ahead (or any other suitable distance, such as 20 meters, 30 meters, 50 meters). According to a method called “rear alignment,” described in more detail in another section, the localization task uses a road model within a typical range of 60 meters behind the car (or any other suitable distance, such as 50 meters, 100 meters, 150 meters, etc.). The disclosed systems and methods can generate geometric models with sufficient accuracy within a specific range (such as 100 meters) such that the planned trajectory will not deviate from the lane center by more than, for example, 30 cm.
[0247] As explained above, a 3D road model can be constructed by detecting short road segments and stitching them together. This stitching can be achieved by calculating a six-degree-of-freedom self-motion model using video and / or images captured by cameras, data from inertial sensors reflecting vehicle motion, and the main vehicle's speed signal. The accumulated error may be small enough at certain local scales (such as approximately 100 meters). All of this can be accomplished in a single drive on a specific road segment.
[0248] In some embodiments, multiple drives can be used to average the resulting model and further improve its accuracy. The same car may drive the same route multiple times, or multiple cars may send their collected model data to a central server. In any case, a matching process can be performed to identify overlapping models and achieve averaging to generate the target trajectory. Once convergence criteria are met, the constructed model (e.g., including the target trajectory) can be used for steering. Subsequent drives can be used for further model improvements and to adapt to infrastructure changes.
[0249] If multiple vehicles are connected to a central server, sharing driving experiences (such as sensed data) among them becomes feasible. Each vehicle client can store a partial copy of a common road model, which may be related to its current location. Bidirectional updates between the vehicle and the server can be performed by both the vehicle and the server. The small footprint concept discussed above enables the disclosed systems and methods to perform bidirectional updates using very little bandwidth.
[0250] Information associated with potential landmarks can also be identified and forwarded to a central server. For example, the disclosed systems and methods can determine one or more physical attributes of potential landmarks based on one or more images including the landmark. Physical attributes may include the physical size of the landmark (e.g., height, width), the distance from the vehicle to the landmark, the distance between the landmark and previous landmarks, the lateral position of the landmark (e.g., the position of the landmark relative to a driving lane), the GPS coordinates of the landmark, the type of landmark, the identification of text on the landmark, etc. For example, a vehicle can analyze one or more images captured by a camera to detect potential landmarks, such as speed limit signs.
[0251] Vehicles can determine the distance from a landmark or the location associated with a landmark (e.g., any semantic or non-semantic object or feature along a road segment) based on the analysis of one or more images. In some embodiments, this distance can be determined based on the analysis of images of the landmark using suitable image analysis methods, such as scaling methods and / or optical flow analysis methods. As previously indicated, the location of an object / feature may include the 2D image location of one or more points associated with the object / feature (e.g., XY pixel locations in one or more captured images), or may include the 3D real-world location of one or more points (e.g., determined by structural techniques in motion / optical flow, lidar, or radar information, etc.). In some embodiments, the disclosed systems and methods can be configured to determine the type or classification of potential landmarks. If a vehicle determines that a potential landmark corresponds to a pre-determined type or classification stored in a sparse map, it may be sufficient for the vehicle to transmit an indication of the landmark's type or classification along with the landmark's location to a server. The server may store such indications. At a later time, during navigation, the navigation vehicle may capture an image including representations of landmarks, process the image (e.g., using a classifier), and compare the resulting landmarks to confirm the detection of the mapped landmarks and to use the mapped landmarks to locate the navigation vehicle relative to a sparse map.
[0252] In some embodiments, multiple autonomous vehicles traveling on a road segment can communicate with a server. Vehicles (or clients) can generate curves describing their driving in any coordinate system (e.g., through self-motion integration). Vehicles can detect landmarks and locate them within the same frame. Vehicles can upload curves and landmarks to the server. The server can collect data from the vehicles through multiple drives and generate a unified road model. For example, as described below... Figure 19 The server can use uploaded curves and landmarks to generate sparse maps with a uniform road model.
[0253] The server can also distribute the model to clients (e.g., vehicles). For example, the server can distribute a sparse map to one or more vehicles. The server can update the model continuously or periodically as new data is received from the vehicles. For example, the server can process the new data to evaluate whether it contains information that should trigger an update or creation of new data on the server. The server can then distribute the updated model or the update to the vehicles to provide autonomous vehicle navigation.
[0254] The server can use one or more criteria to determine whether new data received from a vehicle should trigger an update to the model or the creation of new data. For example, when new data indicates that a previously identified landmark at a specific location no longer exists or has been replaced by another landmark, the server can determine that the new data should trigger an update to the model. As another example, when new data indicates that a road segment has been closed, and this has been confirmed by data received from other vehicles, the server can determine that the new data should trigger an update to the model.
[0255] The server may distribute an updated model (or an updated portion of a model) to one or more vehicles traveling on a road segment associated with the update. The server may also distribute an updated model to vehicles about to travel on that road segment, or to vehicles whose planned journeys include road segments associated with the update. For example, when an autonomous vehicle is traveling along another road segment before reaching the road segment associated with the update, the server may distribute the updated or partially updated model to the autonomous vehicle before the vehicle reaches that road segment.
[0256] In some embodiments, a remote server may collect trajectories and landmarks from multiple clients, such as vehicles traveling along a public road segment. The server can use the landmarks to match curves and create an average road model based on the trajectories collected from the multiple vehicles. The server may also calculate the road map and the most probable path at each node or intersection of the road segment. For example, the remote server may align the trajectories to generate a crowdsourced sparse map from the collected trajectories.
[0257] The server can average landmark attributes (such as distances between one landmark and another (e.g., the previous landmark along the road segment) received from multiple vehicles traveling along a common road segment to determine arc length parameters and support positioning and speed calibration for each client vehicle along the path. The server can average the physical dimensions of landmarks measured by multiple vehicles traveling along the common road segment and identifying the same landmark. The averaged physical dimensions can be used to support distance estimation, such as the distance from a vehicle to a landmark. The server can average the lateral position of landmarks (e.g., from the lane a vehicle is traveling in to its position within the landmark) measured by multiple vehicles traveling along the common road segment and identifying the same landmark. The averaged lateral portion can be used to support lane assignment. The server can average the GPS coordinates of landmarks measured by multiple vehicles traveling along the same road segment and identifying the same landmark. The averaged GPS coordinates of landmarks can be used to support global localization or positioning of landmarks in the road model.
[0258] In some embodiments, the server may identify model changes, such as construction, detours, new signs, or sign removal, based on data received from the vehicle. The server may update the model continuously, periodically, or instantaneously as new data is received from the vehicle. The server may distribute the updated model or the updated model to the vehicle for use in providing autonomous navigation. For example, as discussed further below, the server may use crowdsourced data to filter out “ghost” landmarks detected by the vehicle.
[0259] In some embodiments, the server may analyze driver interventions during autonomous driving. The server may analyze data received from the vehicle at the time and location of the intervention and / or data received prior to the time of the intervention. The server may identify portions of the data that caused or is closely related to the intervention, such as data indicating temporary lane closures or data indicating pedestrians in the road. The server may update the model based on the identified data. For example, the server may modify one or more trajectories stored in the model.
[0260] Figure 12 A schematic diagram of a system that uses crowdsourcing to generate sparse maps (and uses crowdsourced sparse maps for distribution and navigation). Figure 12 Road segment 1200, comprising one or more lanes, is shown. Multiple vehicles 1205, 1210, 1215, 1220, and 1225 may travel on road segment 1200 at the same time or at different times (although...). Figure 12 (As shown in the diagram, they appear on road segment 1200 at the same time). At least one of vehicles 1205, 1210, 1215, 1220, and 1225 may be an autonomous vehicle. For the sake of simplicity in this example, all vehicles 1205, 1210, 1215, 1220, and 1225 are assumed to be autonomous vehicles.
[0261] Each vehicle may be similar to a vehicle disclosed in other embodiments (e.g., vehicle 200) and may include components or devices included in or associated with vehicles disclosed in other embodiments. Each vehicle may be equipped with an image capture device or camera (e.g., image capture device 122 or camera 122). Each vehicle may communicate with a remote server 1230 via one or more networks (e.g., via cellular networks and / or the Internet, etc.) through a wireless communication path 1235 as indicated by the dotted line. Each vehicle may transmit data to and from the server 1230. For example, the server 1230 may collect data from multiple vehicles traveling on road segment 1200 at different times and may process the collected data to generate an autonomous vehicle road navigation model or an update to that model. The server 1230 may transmit the autonomous vehicle road navigation model or an update to that model to the vehicle that transmitted the data to the server 1230. The server 1230 may transmit the autonomous vehicle road navigation model or an update to that model to other vehicles traveling on road segment 1200 at a later time.
[0262] As vehicles 1205, 1210, 1215, 1220, and 1225 travel on road segment 1200, navigation information collected (e.g., detected, sensed, or measured) by vehicles 1205, 1210, 1215, 1220, and 1225 may be transmitted to server 1230. In some embodiments, the navigation information may be associated with a common road segment 1200. The navigation information may include a trajectory associated with each of vehicles 1205, 1210, 1215, 1220, and 1225 as each vehicle travels on road segment 1200. In some embodiments, the trajectory may be reconstructed based on data sensed by various sensors and devices disposed on vehicle 1205. For example, the trajectory may be reconstructed based on at least one of accelerometer data, speed data, landmark data, road geometry or contour data, vehicle positioning data, and self-motion data. In some embodiments, the trajectory may be reconstructed based on data from inertial sensors such as accelerometers and the rate of vehicle 1205 sensed by a speed sensor. In addition, in some embodiments, the trajectory may be determined based on the sensed ego motion of the camera (e.g., by a processor on each of vehicles 1205, 1210, 1215, 1220, and 1225), which may indicate three-dimensional translation and / or three-dimensional rotation (or rotational motion). The ego motion of the camera (and therefore the vehicle body) may be determined from the analysis of one or more images captured by the camera.
[0263] In some embodiments, the trajectory of vehicle 1205 may be determined by a processor mounted on vehicle 1205 and transmitted to server 1230. In other embodiments, server 1230 may receive data sensed by various sensors and devices disposed in vehicle 1205 and determine the trajectory based on the data received from vehicle 1205.
[0264] In some embodiments, navigation information transmitted from vehicles 1205, 1210, 1215, 1220, and 1225 to server 1230 may include data about the road surface, road geometry, or road profile. The geometry of road segment 1200 may include lane structure and / or landmarks. Lane structure may include the total number of lanes in road segment 1200, lane type (e.g., one-way lane, two-way lane, driving lane, overtaking lane, etc.), lane markings, lane width, etc. In some embodiments, navigation information may include lane assignment, such as which lane a vehicle is traveling in among multiple lanes. For example, lane assignment may be associated with the value "3," indicating that the vehicle is traveling in the third lane from the left or right. As another example, lane assignment may be associated with the text value "center lane," indicating that the vehicle is traveling in the center lane.
[0265] Server 1230 may store navigation information on a non-transitory computer-readable medium, such as a hard disk drive, optical disk, magnetic tape, memory, etc. Server 1230 may generate (e.g., via a processor included in server 1230) at least a portion of an autonomous vehicle road navigation model for a public road segment 1200 based on navigation information received from multiple vehicles 1205, 1210, 1215, 1220, and 1225, and may store this model as part of a sparse map. Server 1230 may determine a trajectory associated with each lane based on crowdsourced data (e.g., navigation information) received from multiple vehicles (e.g., 1205, 1210, 1215, 1220, and 1225) traveling in lanes of the road segment at different times. Server 1230 may generate an autonomous vehicle road navigation model or a portion of that model (e.g., an updated portion) based on the multiple trajectories determined based on the crowdsourced navigation data. Server 1230 may transmit a model or an updated portion thereof to one or more of the autonomous vehicles 1205, 1210, 1215, 1220, and 1225 traveling on road segment 1200, or to any other autonomous vehicle traveling on the road segment at a later time, to update the existing autonomous vehicle road navigation model provided in the vehicle's navigation system. The autonomous vehicle road navigation model can be used by the autonomous vehicles when navigating autonomously along the public road segment 1200.
[0266] As explained above, autonomous vehicle road navigation models can be incorporated into sparse maps (e.g., Figure 8 The sparse map 800 depicted in the image is used to guide autonomous vehicles. The sparse map 800 may include sparse records of data relating to road geometry and / or landmarks along the road, which provide sufficient information for autonomous navigation of autonomous vehicles without requiring excessive data storage. In some embodiments, the autonomous vehicle road navigation model may be stored separately from the sparse map 800, and map data from the sparse map 800 may be used when the model is executed for navigation. In some embodiments, the autonomous vehicle road navigation model may use map data included in the sparse map 800 to determine a target trajectory along road segment 1200 for guiding autonomous vehicles 1205, 1210, 1215, 1220, and 1225, or other vehicles subsequently traveling along road segment 1200. For example, when the autonomous vehicle road navigation model is executed by a processor included in the navigation system of vehicle 1205, the model enables the processor to compare a trajectory determined based on navigation information received from vehicle 1205 with a pre-determined trajectory included in sparse map 800 to verify and / or correct the current driving route of vehicle 1205.
[0267] In an autonomous vehicle road navigation model, the geometric features of a road or target trajectory can be encoded by a curve in three-dimensional space. In one embodiment, the curve may be a three-dimensional spline comprising one or more connected three-dimensional polynomials. As those skilled in the art will understand, a spline may be a numerical function defined piecewise by a series of polynomials used to fit the data. The spline used to fit the three-dimensional geometric feature data of the road may include a linear spline (first order), a quadratic spline (second order), a cubic spline (third order), or any other spline (other orders), or a combination thereof. A spline may include one or more three-dimensional polynomials of different orders connecting (e.g., fitting) the data points of the three-dimensional geometric feature data of the road. In some embodiments, the autonomous vehicle road navigation model may include a three-dimensional spline corresponding to a target trajectory along a common road segment (e.g., road segment 1200) or a lane of road segment 1200.
[0268] As explained above, the autonomous vehicle road navigation model included in the sparse map may include other information, such as the identification of at least one landmark along road segment 1200. The landmark may be visible within the field of view of a camera (e.g., camera 122) mounted on each of vehicles 1205, 1210, 1215, 1220, and 1225. In some embodiments, camera 122 may capture an image of the landmark. A processor (e.g., processor 180, 190, or processing unit 110) located on vehicle 1205 may process the image of the landmark to extract identification information for the landmark. The landmark identification information, rather than the actual image of the landmark, may be stored in the sparse map 800. The landmark identification information may require significantly less storage space than the actual image. Other sensors or systems (e.g., a GPS system) may also provide some identification information for the landmark (e.g., the location of the landmark). Landmarks may include at least one of traffic signs, arrow markings, lane markings, dashed lane markings, traffic lights, stop lines, directional markings (e.g., highway exit signs with arrows indicating directions, highway signs with arrows pointing in different directions or places), landmark beacons, or lampposts. A landmark beacon refers to a device (e.g., an RFID device) installed along a road segment that transmits or reflects signals to a receiver installed on a vehicle, such that when a vehicle passes the device, the beacon received by the vehicle and the location of the device (e.g., determined from the device's GPS location) can be used as a landmark to be included in the autonomous vehicle road navigation model and / or sparse map 800.
[0269] The identification of at least one landmark may include the location of that landmark. The location of the landmark may be determined based on location measurements taken using sensor systems (e.g., GPS, inertial positioning systems, landmark beacons, etc.) associated with multiple vehicles 1205, 1210, 1215, 1220, and 1225. In some embodiments, the location of the landmark may be determined by averaging location measurements detected, collected, or received by sensor systems on different vehicles 1205, 1210, 1215, 1220, and 1225 through multiple driving actions. For example, vehicles 1205, 1210, 1215, 1220, and 1225 may transmit location measurement data to server 1230, which may average the location measurements and use the averaged location measurements as the location of the landmark. The location of the landmark may be continuously refined using measurements received from the vehicles in subsequent driving actions.
[0270] Landmark identification may include the size of the landmark. A processor located on a vehicle (e.g., 1205) may estimate the physical size of the landmark based on analysis of the image. Server 1230 may receive multiple estimates of the physical size of the same landmark from different vehicles driven by different drivers. Server 1230 may average the different estimates to arrive at the physical size of the landmark and store this landmark size in a road model. The physical size estimate may be used to further determine or estimate the distance from the vehicle to the landmark. The distance to the landmark may be estimated based on the vehicle's current speed and the extended scale based on the location of the landmark in the image relative to the extended focus of the camera. For example, the distance to the landmark may be estimated as Z = V*dt*R / D, where V is the vehicle's speed, R is the distance from the landmark in the image to the extended focus from time t1, and D is the change in distance of the landmark in the image from t1 to t2. dt represents (t2-t1). For example, the distance to a landmark can be estimated using Z = V * dt * R / D, where V is the vehicle speed, R is the distance between the landmark and the extended focal point in the image, dt is the time interval, and D is the image displacement of the landmark along the epipolar line. Other equations equivalent to the above, such as Z = V * ω / Δω, can be used to estimate the distance to a landmark. Here, V is the vehicle speed, ω is the image length (e.g., object width), and Δω is the change in that image length per unit time.
[0271] When the physical size of a landmark is known, the distance to the landmark can be determined based on the following equation: Z = f * W / ω, where f is the focal length, W is the size of the landmark (e.g., height or width), and ω is the number of pixels the landmark leaves the image. According to this equation, the change in distance Z can be calculated using ΔZ = f * W * Δω / ω² + f * ΔW / ω, where ΔW decays to zero through averaging, and Δω is the number of pixels representing the accuracy of the bounding box in the image. The estimated physical size of the landmark can be calculated by averaging multiple observations on the server side. The resulting distance estimate may have a very small error. There are two sources of error that may arise when using the above formula: ΔW and Δω. Their contribution to the distance error is given by ΔZ = f * W * Δω / ω² + f * ΔW / ω. However, ΔW decays to zero through averaging; therefore, ΔZ is determined by Δω (e.g., inaccuracies in the bounding box in the image).
[0272] For landmarks of unknown size, the distance to the landmark can be estimated by tracking feature points on the landmark between consecutive frames. For example, certain features appearing on a speed limit sign can be tracked between two or more image frames. Based on these tracked features, a distance distribution for each feature point can be generated. The distance estimate can be extracted from the distance distribution. For example, the most frequently occurring distance in the distance distribution can be used as the distance estimate. As another example, the average value of the distance distribution can be used as the distance estimate.
[0273] Figure 13 An example autonomous vehicle road navigation model is shown, represented by multiple 3D splines 1301, 1302, and 1303. Figure 13 The curves 1301, 1302, and 1303 shown are for illustrative purposes only. Each spline may include one or more three-dimensional polynomials connecting multiple data points 1310. Each polynomial may be a first-order polynomial, a second-order polynomial, a third-order polynomial, or any suitable combination of polynomials of different orders. Each data point 1310 may be associated with navigation information received from vehicles 1205, 1210, 1215, 1220, and 1225. In some embodiments, each data point 1310 may be associated with data related to landmarks (e.g., the size, location, and identification information of landmarks) and / or road signature profiles (e.g., road geometry, road roughness profile, road curvature profile, road width profile). In some embodiments, some data points 1310 may be associated with data related to landmarks, and other data points may be associated with data related to road signature profiles.
[0274] Figure 14 The diagram illustrates raw location data 1410 (e.g., GPS data) received from five individual drives. A drive can be separated from another drive if it is traversed by a separate vehicle at the same time, by the same vehicle at a separate time, or by a separate vehicle at a separate time. To account for errors in the location data 1410 and for different locations of vehicles within the same lane (e.g., one vehicle might be driving closer to the left side of the lane than another), server 1230 can use one or more statistical techniques to generate a map skeleton 1420 to determine whether variations in the raw location data 1410 represent actual deviations or statistical errors. Each path within the skeleton 1420 can be linked back to the raw data 1410 that formed that path. For example, the path between A and B within the skeleton 1420 is linked to the raw data 1410 from drives 2, 3, 4, and 5, but not from drive 1. The skeleton 1420 may not be detailed enough for navigating vehicles (e.g., because it combines drives from multiple lanes on the same road, unlike the splines described above), but it provides useful topological information and can be used to define intersections.
[0275] Figure 15 An example is shown that additional details can be generated for sparse maps within segments of a map skeleton (e.g., segments A to B within skeleton 1420). Figure 15 The data depicted (e.g., self-motion data, road marking data, etc.) can be shown as a function of the driver's position S (or S1 or S2). Server 1230 can identify landmarks for a sparse map by recognizing unique matches between landmarks 1501, 1503, and 1505 of driver 1510 and landmarks 1507 and 1509 of driver 1520. Such matching algorithms can result in the identification of landmarks 1511, 1513, and 1515. However, those skilled in the art will recognize that other matching algorithms can be used. For example, probabilistic optimization can be used instead of unique matching or in combination with unique matching. Server 1230 can longitudinally align the drivers to align with the matched landmarks. For example, server 1230 can select a driver (e.g., driver 1520) as a reference driver and then transform and / or elastically stretch other drivers (e.g., driver 1510) for alignment.
[0276] Figure 16 An example of landmark data used for alignment in sparse maps is shown. Figure 16 In the example, landmark 1610 includes road signs. Figure 16 The examples further depict data from multiple driving records 1601, 1603, 1605, 1607, 1609, 1611, and 1613. In Figure 16 In the example, the data from drive 1613 consists of “ghost” landmarks, and server 1230 may identify this landmark because none of drives 1601, 1603, 1605, 1607, 1609, and 1611 include the landmarks in the vicinity of the landmark identified in drive 1613. Therefore, server 1230 may accept a potential landmark when the ratio of images in which a landmark appears to images in which a landmark does not appear exceeds a threshold, and / or may reject a potential landmark when the ratio of images in which a landmark does not appear to images in which a landmark appears exceeds a threshold.
[0277] Figure 17 A system 1700 for generating driving data is described, which can be used for crowdsourcing of sparse maps. For example... Figure 17As depicted, system 1700 may include a camera 1701 and a positioning device 1703 (e.g., a GPS locator). Camera 1701 and positioning device 1703 may be mounted on a vehicle (e.g., one of vehicles 1205, 1210, 1215, 1220, and 1225). Camera 1701 may generate multiple types of data, such as self-motion data, traffic sign data, road data, etc. Camera data and location data may be segmented into driving segments 1705. For example, driving segments 1705 may each have camera data and location data from a drive of less than 1 km.
[0278] In some embodiments, system 1700 may remove redundancy in driving segment 1705. For example, if a landmark appears in multiple images from camera 1701, system 1700 may remove redundant data such that driving segment 1705 contains only a copy of the location of the landmark and any metadata associated with the landmark. By another example, if lane markings appear in multiple images from camera 1701, system 1700 may remove redundant data such that driving segment 1705 contains only a copy of the location of the lane markings and any metadata associated with the lane markings.
[0279] System 1700 also includes a server (e.g., server 1230). Server 1230 can receive driving segments 1705 from the vehicle and recombine the driving segments 1705 into a single drive 1707. This arrangement allows for reduced bandwidth requirements when transmitting data between the vehicle and the server, while also allowing the server to store data related to the entire driving process.
[0280] Figure 18 The text describes further configurations for crowdsourcing sparse maps. Figure 17 System 1700. For example, in... Figure 17 In this system 1700, a vehicle 1810 and a positioning device (e.g., a GPS locator) are included. The vehicle uses, for example, a camera (which generates, for example, self-motion data, traffic sign data, road data, etc.) to capture driving data. (As in...) Figure 17 In the process, vehicle 1810 divides the collected data into driving segments (in... Figure 18 The code describes these as "DS1 1", "DS2 1", and "DSN 1". Server 1230 then receives the driving segments and reconstructs the driving from the received segments (in...). Figure 18 (Described as "Driving 1" in the text).
[0281] As in Figure 18Further described, system 1700 also receives data from an attached vehicle. For example, vehicle 1820 also uses, for example, cameras (which generate, for example, self-motion data, traffic sign data, road data, etc.) and positioning devices (e.g., GPS locators) to capture driving data. Similar to vehicle 1810, vehicle 1820 segments the collected data into driving segments (in... Figure 18 The code describes these as "DS1 2", "DS2 2", and "DSN 2". Server 1230 then receives the driving segments and reconstructs the driving from the received segments (in...). Figure 18 (Described as "Driving 2" in the text). Any number of additional vehicles can be used. For example, Figure 18 This also includes "Sedan N," which captures driving data and segments it into driving segments (in... Figure 18 The text describes the process as "DS1 N", "DS2 N", "DSN N" and sends it to server 1230 for reconstruction as a driver (in...). Figure 18 It is described as "driving N" in the text.
[0282] like Figure 18 As depicted, server 1230 can use reconstructed driving data (e.g., "driving 1", "driving 2", and "driving N") collected from multiple vehicles (e.g., "car 1" (also labeled as vehicle 1810), "car 2" (also labeled as vehicle 1820), and "car N") to construct a sparse map (depicted as "map").
[0283] Figure 19 A flowchart illustrating an example process 1900 for generating a sparse map for autonomous vehicle navigation along road segments is provided. Process 1900 may be performed by one or more processing devices included in server 1230.
[0284] Process 1900 may include receiving multiple images acquired as one or more vehicles traverse a road segment (step 1905). Server 1230 may receive images from cameras included in one or more of vehicles 1205, 1210, 1215, 1220, and 1225. For example, as vehicle 1205 travels along road segment 1200, camera 122 may capture one or more images of the environment surrounding vehicle 1205. In some embodiments, server 1230 may also receive simplified image data from which redundancy has been removed by a processor on vehicle 1205, as described above regarding... Figure 17 The subject of discussion.
[0285] Process 1900 may further include identifying at least one line representation of a road surface feature extending along a road segment based on multiple images (step 1910). Each line representation may represent a path along a road segment substantially corresponding to a road surface feature. For example, server 1230 may analyze environmental images received from camera 122 to identify road edges or lane markings and determine a driving trajectory along a road segment 1200 associated with the road edge or lane marking. In some embodiments, the trajectory (or line representation) may include a spline, a polynomial representation, or a curve. Server 1230 may determine the driving trajectory of vehicle 1205 based on camera ego motion (e.g., three-dimensional translation and / or three-dimensional rotational motion) received in step 1905.
[0286] Process 1900 may also include identifying multiple landmarks associated with a road segment based on multiple images (step 1910). For example, server 1230 may analyze environmental images received from camera 122 to identify one or more landmarks, such as road signs along road segment 1200. Server 1230 may use analysis of multiple images acquired as one or more vehicles pass through the road segment to identify landmarks. To enable crowdsourcing, the analysis may include rules regarding the acceptance and rejection of potential landmarks associated with the road segment. For example, the analysis may include accepting a potential landmark when the ratio of images in which a landmark appears to images in which a landmark does not appear exceeds a threshold, and / or rejecting a potential landmark when the ratio of images in which a landmark does not appear to images in which a landmark appears exceeds a threshold.
[0287] Process 1900 may include other operations or steps performed by server 1230. For example, navigation information may include a target trajectory for vehicles traveling along a road segment, and process 1900 may include server 1230 clustering vehicle trajectories associated with multiple vehicles traveling on the road segment, and determining a target trajectory based on the clustered vehicle trajectories, as discussed in further detail below. Clustering vehicle trajectories may include server 1230 clustering multiple trajectories associated with vehicles traveling on the road segment into multiple clusters based on at least one of the vehicle's absolute heading or the vehicle's lane assignment. Generating the target trajectory may include server 1230 averaging the clustered trajectories. By another example, process 1900 may include aligning the data received in step 1905. As described above, other processes or steps performed by server 1230 may also be included in process 1900.
[0288] The disclosed systems and methods may include other features. For example, the disclosed systems may use local coordinates instead of global coordinates. For autonomous driving, some systems may present data in world coordinates. For example, longitude and latitude coordinates on the Earth's surface may be used. To use a map for steering, the host vehicle can determine its position and orientation relative to the map. It seems natural to use an onboard GPS device to locate the vehicle on the map and to find the rotational transformation between the host reference frame and the world reference frame (e.g., north, east, and down). Once the host reference frame is aligned with the map reference frame, the desired route can be expressed in the host reference frame and steering commands can be calculated or generated.
[0289] The disclosed systems and methods enable autonomous vehicle navigation (e.g., steering control) with a low-footprint model, which can be collected by the autonomous vehicle itself without the need for expensive survey equipment. To support autonomous navigation (e.g., steering applications), the road model may include a sparse map containing road geometry, lane structure, and landmarks that can be used to determine the location or position of the vehicle along a trajectory included in the model. As discussed above, the generation of the sparse map may be performed by a remote server that communicates with and receives data from vehicles traveling on the road. The data may include sensed data, a reconstructed trajectory based on the sensed data, and / or a recommended trajectory that may represent a modified reconstructed trajectory. As discussed below, the server may transmit the model back to the vehicle or other vehicles later traveling on the road to aid in autonomous navigation.
[0290] Figure 20 A block diagram of server 1230 is shown. Server 1230 may include a communication unit 2005, which may include both hardware components (e.g., communication control circuitry, a switch, and an antenna) and software components (e.g., communication protocols and computer code). For example, communication unit 2005 may include at least one network interface. Server 1230 can communicate with vehicles 1205, 1210, 1215, 1220, and 1225 via communication unit 2005. For example, server 1230 can receive navigation information transmitted from vehicles 1205, 1210, 1215, 1220, and 1225 via communication unit 2005. Server 1230 can distribute autonomous vehicle road navigation models to one or more autonomous vehicles via communication unit 2005.
[0291] Server 1230 may include at least one non-transitory storage medium 2010, such as a hard disk drive, optical disk, magnetic tape, etc. Storage device 1410 may be configured to store data, such as navigation information received from vehicles 1205, 1210, 1215, 1220, and 1225 and / or an autonomous vehicle road navigation model generated by server 1230 based on that navigation information. Storage device 2010 may be configured to store any other information, such as sparse maps (e.g., as described above regarding...). Figure 8 The sparse map discussed (800).
[0292] In addition to or in place of storage device 2010, server 1230 may include memory 2015. Memory 2015 may be similar to or different from memory 140 or 150. Memory 2015 may be non-transitory memory, such as flash memory, random access memory, etc. Memory 2015 may be configured to store data, such as computer code or instructions executable by a processor (e.g., processor 2020), map data (e.g., data from sparse map 800), autonomous vehicle road navigation models, and / or navigation information received from vehicles 1205, 1210, 1215, 1220, and 1225.
[0293] Server 1230 may include at least one processing unit 2020 configured to execute computer code or instructions stored in memory 2015 to perform various functions. For example, processing unit 2020 may analyze navigation information received from vehicles 1205, 1210, 1215, 1220, and 1225, and generate an autonomous vehicle road navigation model based on the analysis. Processing unit 2020 may control communication unit 1405 to distribute the autonomous vehicle road navigation model to one or more autonomous vehicles (e.g., one or more of vehicles 1205, 1210, 1215, 1220, and 1225, or any vehicle traveling on road segment 1200 at a later time). Processing unit 2020 may be similar to or different from processor 180, 190, or processing unit 110.
[0294] Figure 21 A block diagram of a memory 2015 is shown, which stores computer code or instructions for performing one or more operations to generate a road navigation model for use in autonomous vehicle navigation. Figure 21 As shown, memory 2015 may store one or more modules for performing operations for processing vehicle navigation information. For example, memory 2015 may include model generation module 2105 and model distribution module 2110. Processor 2020 may execute instructions stored in either module 2105 or 2110 included in memory 2015.
[0295] The model generation module 2105 may store instructions that, when executed by the processor 2020, may generate at least a portion of an autonomous vehicle road navigation model for a public road segment (e.g., road segment 1200) based on navigation information received from vehicles 1205, 1210, 1215, 1220, and 1225. For example, when generating the autonomous vehicle road navigation model, the processor 2020 may cluster vehicle trajectories along the public road segment 1200 into different clusters. The processor 2020 may determine a target trajectory along the public road segment 1200 based on the vehicle trajectories for each of the different clusters. Such operations may include finding the mean or average trajectory of the vehicle trajectories for each cluster (e.g., by averaging data representing the clusters). In some embodiments, the target trajectory may be associated with a single lane of the public road segment 1200.
[0296] Road models and / or sparse maps can store trajectories associated with road segments. These trajectories, referred to as target trajectories, are provided to autonomous vehicles for autonomous navigation. Target trajectories can be received from multiple vehicles, or generated based on actual trajectories received from multiple vehicles, or recommended trajectories (actual trajectories with some modifications). Target trajectories included in the road model or sparse map can be continuously updated (e.g., averaged) using new trajectories received from other vehicles.
[0297] Vehicles traveling on a road segment can collect data via various sensors. This data may include landmarks, road signature contours, vehicle motion (e.g., accelerometer data, speed data), vehicle position (e.g., GPS data), and may reconstruct the actual trajectory itself, or transmit the data to a server that reconstructs the actual trajectory for the vehicle. In some embodiments, the vehicle may transmit data related to the trajectory (e.g., a curve in an arbitrary reference frame), landmark data, and lane assignments along the driving path to server 1230. Various vehicles traveling along the same road segment under multiple driving conditions may have different trajectories. Server 1230 can identify routes or trajectories associated with each lane from the trajectories received from the vehicles through a clustering process.
[0298] Figure 22A process is illustrated for clustering vehicle trajectories associated with vehicles 1205, 1210, 1215, 1220, and 1225 to determine target trajectories for a common road segment (e.g., road segment 1200). The target trajectories or multiple target trajectories determined from the clustering process may be included in an autonomous vehicle road navigation model or a sparse map 800. In some embodiments, vehicles 1205, 1210, 1215, 1220, and 1225 traveling along road segment 1200 may transmit multiple trajectories 2200 to server 1230. In some embodiments, server 1230 may generate trajectories based on landmarks, road geometry features, and vehicle motion information received from vehicles 1205, 1210, 1215, 1220, and 1225. To generate an autonomous vehicle road navigation model, server 1230 may cluster vehicle trajectories 1600 into multiple clusters 2205, 2210, 2215, 2220, 2225, and 2230, as shown below. Figure 22 As shown.
[0299] Various criteria can be used for clustering. In some embodiments, the absolute headings of all drivers in a cluster with respect to road segment 1200 can be similar. Absolute headings can be obtained from GPS signals received by vehicles 1205, 1210, 1215, 1220, and 1225. In some embodiments, dead reckoning can be used to obtain absolute headings. As those skilled in the art will understand, dead reckoning can be used to determine the current position and thus the headings of vehicles 1205, 1210, 1215, 1220, and 1225 by using previously determined positions, estimated speeds, etc. Trajectories clustered by absolute headings can be useful for identifying routes along the roadway.
[0300] In some embodiments, the lane assignments of all driving in a cluster with respect to driving along road segment 1200 (e.g., in the same lane before and after an intersection) can be similar. The trajectory of the lane assignment cluster can be useful for identifying lanes along the carriageway. In some embodiments, two criteria (e.g., absolute heading and lane assignment) can be used for clustering.
[0301] Within each cluster 2205, 2210, 2215, 2220, 2225, and 2230, trajectories can be averaged to obtain a target trajectory associated with a specific cluster. For example, trajectories from multiple drives associated with the same lane cluster can be averaged. The averaged trajectory can be a target trajectory associated with a specific lane. To average the clustered trajectories, server 1230 can select a reference frame for any trajectory C0. For all other trajectories (C1, ..., Cn), server 1230 can find a rigid transformation mapping Ci to C0, where i = 1, 2, ..., n, where n is a positive integer corresponding to the total number of trajectories included in the cluster. Server 1230 can compute the mean curve or trajectory in the C0 reference frame.
[0302] In some embodiments, landmarks may define an arc length that matches different driving positions, which can be used for trajectory alignment with lanes. In some embodiments, lane markings before and after intersections may be used for trajectory alignment with lanes.
[0303] To assemble lanes from a trajectory, server 1230 can select a reference frame for any lane. Server 1230 can map partially overlapping lanes to the selected reference frame. Server 1230 can continue mapping until all lanes are in the same reference frame. Lanes adjacent to each other may be aligned as if they were the same lane, and then they may be transposed laterally.
[0304] Landmarks identified along road segments can be mapped to a common reference frame, first at lane level and then at intersection level. For example, the same landmark might be identified multiple times by multiple vehicles in multiple drives. The data received about the same landmark in different drives might differ slightly. Such data can be averaged and mapped to the same reference frame, such as the C0 reference frame. Alternatively or additionally, the variance of the data for the same landmark received in multiple drives can be calculated.
[0305] In some embodiments, each lane of road segment 120 may be associated with a target trajectory and certain landmarks. The target trajectory or multiple such target trajectories may be included in an autonomous vehicle road navigation model that may later be used by other autonomous vehicles traveling along the same road segment 1200. Landmarks identified by vehicles 1205, 1210, 1215, 1220, and 1225 as they travel along road segment 1200 may be recorded in association with the target trajectory. The target trajectory and landmark data may be continuously or periodically updated using new data received from other vehicles in subsequent driving.
[0306] For autonomous vehicle localization, the disclosed systems and methods may use an extended Kalman filter. The vehicle's location may be determined based on 3D position data and / or 3D orientation data, through the integral of self-motion, and prediction of future locations ahead of the vehicle's current location. The vehicle's localization may be corrected or adjusted through image observation of landmarks. For example, when the vehicle detects a landmark within an image captured by a camera, that landmark may be compared with known landmarks stored in a road model or sparse map 800. Known landmarks may have known locations (e.g., GPS data) along a target trajectory stored in the road model and / or sparse map 800. Based on the current speed and the image of the landmark, the distance from the vehicle to the landmark may be estimated. The vehicle's location along the target trajectory may be adjusted based on the distance to the landmark and the known location of the landmark (stored in the road model or sparse map 800). The location / location data of landmarks stored in the road model and / or sparse map 800 (e.g., the average from multiple drives) may be assumed to be accurate.
[0307] In some embodiments, the disclosed system may form a closed-loop subsystem, wherein estimation of the vehicle's six-DOF location (e.g., three-dimensional position data plus three-dimensional orientation data) can be used to navigate the autonomous vehicle (e.g., steer its steering wheel) to reach a desired point (e.g., 1.3 seconds in advance in storage). Data from steering and actual navigation measurements can then be used to estimate the six-DOF location.
[0308] In some embodiments, poles along the road (such as lampposts and utility or cable poles) can be used as landmarks for locating vehicles. Other landmarks such as traffic signs, traffic lights, arrows on the road, stop lines, and static features or signatures of objects along road segments can also be used as landmarks for locating vehicles. When using poles for location, the x-observation of the pole (i.e., from the vehicle's viewing angle) can be used instead of the y-observation (i.e., the distance to the pole) because the base of the pole may be obscured and sometimes they are not on the road surface.
[0309] Figure 23 A navigation system for a vehicle was demonstrated, which can be used for autonomous navigation using crowdsourced sparse maps. For demonstration purposes, the vehicle is referred to as Vehicle 1205. Figure 23 The vehicle shown can be any other vehicle disclosed herein, including, for example, vehicles 1210, 1215, 1220, and 1225, as well as vehicle 200 shown in other embodiments. Figure 12As shown, vehicle 1205 can communicate with server 1230. Vehicle 1205 may include image capture device 122 (e.g., camera 122). Vehicle 1205 may include navigation system 2300, which is configured to provide navigation guidance for vehicle 1205 when driving on a road (e.g., road segment 1200). Vehicle 1205 may also include other sensors, such as speed sensor 2320 and accelerometer 2325. Speed sensor 2320 may be configured to detect the speed of vehicle 1205. Accelerometer 2325 may be configured to detect acceleration or deceleration of vehicle 1205. Figure 23 The vehicle 1205 shown can be an autonomous vehicle, and the navigation system 2300 can be used to provide navigation guidance for autonomous driving. Alternatively, the vehicle 1205 can also be a non-autonomous, human-controlled vehicle, and the navigation system 2300 can still be used to provide navigation guidance.
[0310] The navigation system 2300 may include a communication unit 2305 configured to communicate with a server 1230 via a communication path 1235. The navigation system 2300 may also include a GPS unit 2310 configured to receive and process GPS signals. The navigation system 2300 may further include at least one processor 2315 configured to process data such as GPS signals, map data from a sparse map 800 (which may be stored on a storage device configured to be mounted on the vehicle 1205 and / or received from the server 1230), geometric features sensed by a road profile sensor 2330, images captured by a camera 122, and / or an autonomous vehicle road navigation model received from the server 1230. The road profile sensor 2330 may include different types of means for measuring different types of road profiles, such as road surface roughness, road width, road elevation, road curvature, etc. For example, the road profile sensor 2330 may include means for measuring the motion of the suspension of the vehicle 2305 to derive a road roughness profile. In some embodiments, the road profile sensor 2330 may include a radar sensor to measure the distance from the vehicle 1205 to the roadside (e.g., an obstacle on the roadside), thereby measuring the width of the road. In some embodiments, the road profile sensor 2330 may include means configured to measure the vertical elevation of the road. In some embodiments, the road profile sensor 2330 may include means configured to measure the curvature of the road. For example, a camera (e.g., camera 122 or another camera) may be used to capture images of the road showing its curvature. The vehicle 1205 may use such images to detect the road curvature.
[0311] At least one processor 2315 may be programmed to receive at least one environmental image associated with vehicle 1205 from camera 122. At least one processor 2315 may analyze the at least one environmental image to determine navigation information associated with vehicle 1205. The navigation information may include a trajectory associated with the vehicle 1205's travel along road segment 1200. At least one processor 2315 may determine the trajectory based on motions of camera 122 (and therefore the vehicle), such as three-dimensional translational and three-dimensional rotational motions. In some embodiments, at least one processor 2315 may determine the translational and rotational motions of camera 122 based on analysis of multiple images acquired by camera 122. In some embodiments, the navigation information may include lane assignment information (e.g., in which lane vehicle 1205 is traveling along road segment 1200). Navigation information transmitted from vehicle 1205 to server 1230 may be used by server 1230 to generate and / or update an autonomous vehicle road navigation model, which may be transmitted back from server 1230 to vehicle 1205 for providing autonomous navigation guidance for vehicle 1205.
[0312] At least one processor 2315 may also be programmed to transmit navigation information from vehicle 1205 to server 1230. In some embodiments, the navigation information may be transmitted to server 1230 along with road information. Road location information may include at least one of GPS signals received by GPS unit 2310, landmark information, road geometry, lane information, etc. At least one processor 2315 may receive an autonomous vehicle road navigation model or a portion thereof from server 1230. The autonomous vehicle road navigation model received from server 1230 may include at least one update based on the navigation information transmitted from vehicle 1205 to server 1230. The portion of the model transmitted from server 1230 to vehicle 1205 may include the updated portion of the model. At least one processor 2315 may induce at least one navigation maneuver (e.g., steering such as turning, braking, accelerating, overtaking another vehicle, etc.) by vehicle 1205 based on the received autonomous vehicle road navigation model or the updated portion thereof.
[0313] At least one processor 2315 may be configured to communicate with various sensors and components included in the vehicle 1205, including a communication unit 1705, a GPS unit 2315, a camera 122, a speed sensor 2320, an accelerometer 2325, and a road contour sensor 2330. The processor 2315 may collect information or data from the various sensors and components and transmit the information or data to the server 1230 via the communication unit 2305. Alternatively or additionally, the various sensors or components of the vehicle 1205 may also communicate with the server 1230 and transmit data or information collected by the sensors or components to the server 1230.
[0314] In some embodiments, vehicles 1205, 1210, 1215, 1220, and 1225 can communicate with each other and share navigation information, such that at least one of vehicles 1205, 1210, 1215, 1220, and 1225 can generate an autonomous vehicle road navigation model using crowdsourcing, for example, based on information shared by other vehicles. In some embodiments, vehicles 1205, 1210, 1215, 1220, and 1225 can share navigation information with each other, and each vehicle can update its own autonomous vehicle road navigation model set in the vehicle. In some embodiments, at least one of vehicles 1205, 1210, 1215, 1220, and 1225 (e.g., vehicle 1205) can be used as a hub vehicle. At least one processor 2315 of the hub vehicle (e.g., vehicle 1205) can perform some or all of the functions performed by server 1230. For example, at least one processor 2315 of the hub vehicle can communicate with and receive navigation information from other vehicles. At least one processor 2315 of the wheel hub vehicle can generate an autonomous vehicle road navigation model or an update thereto based on shared information received from other vehicles. The processor 2315 of the wheel hub vehicle can transmit the autonomous vehicle road navigation model or an update thereto to other vehicles for use in providing autonomous navigation guidance.
[0315] Navigation based on sparse maps As previously discussed, an autonomous vehicle road navigation model including a sparse map 800 may include multiple mapped lane markings and multiple mapped objects / features associated with road segments. These mapped lane markings, objects, and features can be used when the autonomous vehicle is navigating, as discussed in more detail below. For example, in some embodiments, the mapped objects and features can be used to locate the primary vehicle relative to the map (e.g., relative to a mapped target trajectory). The mapped lane markings can be used (e.g., as a check) to determine the lateral position and / or orientation relative to the planned or target trajectory. Using this positional information, the autonomous vehicle may be able to adjust its heading direction to match the orientation of the target trajectory at the determined location.
[0316] Vehicle 200 can be configured to detect lane markings in a given road segment. A road segment may include any markings on the road used to guide vehicle traffic on a carriageway. For example, lane markings may be continuous or dashed lines defining the edges of a driving lane. Lane markings may also include double lines, such as double continuous lines, double dashed lines, or a combination of continuous and dashed lines, indicating, for example, whether passage in an adjacent lane is permitted. Lane markings may also include highway entrance and exit markings indicating, for example, deceleration lanes such as exit ramps, or dotted dashed lines indicating lanes that are turning only or where a lane is ending. Markings may further indicate work zones, temporary lane changes, travel paths through intersections, medians, dedicated lanes (e.g., bicycle lanes, HOV lanes, etc.) or other miscellaneous markings (e.g., pedestrian crossings, speed bumps, railway crossings, stop lines, etc.).
[0317] Vehicle 200 may use cameras, such as image capture devices 122 and 124 included in image acquisition unit 120, to capture images of surrounding lane markings. Vehicle 200 may analyze the images to detect points associated with the lane markings based on features identified in one or more of the captured images. These points may be uploaded to a server to represent lane markings in sparse map 800. Depending on the camera position and field of view, lane markings may be detected simultaneously for both sides of the vehicle from a single image. In other embodiments, different cameras may be used to capture images of multiple sides of the vehicle. Instead of uploading actual images of the lane markings, the markings may be stored as splines or a series of points in sparse map 800, thereby reducing the size of sparse map 800 and / or the data that must be remotely uploaded by the vehicle.
[0318] Figures 24A to 24D An exemplary point location that can be detected by vehicle 200 to represent a specific lane marking is demonstrated. Similar to the landmarks described above, vehicle 200 can use various image recognition algorithms or software to identify points within a captured image. For example, vehicle 200 can identify a range of edge points, corner points, or various other points associated with a specific lane marking. Figure 24A A continuous lane marking 2410, detectable by vehicle 200, is shown. Lane marking 2410 may represent the outer edge of a carriageway, indicated by a continuous white line. For example... Figure 24A As shown, vehicle 200 can be configured to detect multiple edge points 2411 along lane markings. Points 2411 can be collected to represent lane markings at any interval sufficient to create mapped lane markings in a sparse map. For example, lane markings can be represented by one point per meter of detected edges, one point per five meters of detected edges, or at other suitable intervals. In some embodiments, the interval can be determined by other factors such as, for example, points where vehicle 200 has the highest confidence rating for the detected points, rather than at set intervals. Although Figure 24AThe edge positioning points on the inner edge of lane marking 2410 are shown, but points can be collected on the outer edge of the line or along both edges. Furthermore, although in Figure 24A The diagram shows a single line, but similar edge points can be detected for dual continuous lines. For example, point 2411 can be detected along the edge of one or more of the continuous lines.
[0319] Depending on the type or shape of the lane markings, vehicle 200 may also represent lane markings differently. Figure 24B An exemplary dashed lane marking 2420, detectable by vehicle 200, is shown. Not as... Figure 24A By marking edge points as shown in the diagram, the vehicle can detect a series of corner points 2421 representing the corners of lane short crosses to define the complete boundary of the short crosses. Although Figure 24B For each corner of a given short dashed line marking that has been positioned, vehicle 200 can detect or upload a subgroup of points shown in the diagram. For example, vehicle 200 can detect the leading edge or leading corner of a given short dashed line marking, or it can detect the two corner points closest to the interior of the lane. Furthermore, not every short dashed line marking can be captured; for example, vehicle 200 can capture and / or record points representing samples of short dashed line markings (e.g., every one, every three, every five, etc.) or points of short dashed line markings at predetermined intervals (e.g., every meter, every 5 meters, every 10 meters, etc.). Corner points can also be detected for similar lane markings, such as markings indicating that a lane is an exit ramp, markings indicating the end of a specific lane, or other various lane markings that may have detectable corner points. Corner points can also be detected for lane markings consisting of double dashed lines or a combination of continuous lines and dashed lines.
[0320] In some embodiments, the points uploaded to the server to generate mapped lane markings may represent points other than detected edge points or corner points. Figure 24C A series of points representing the centerline of a given lane marking is shown. For example, a continuous lane 2410 can be represented by centerline point 2441 along the centerline 2440 of the lane marking. In some embodiments, vehicle 200 may be configured to detect these center points using various image recognition techniques, such as convolutional neural networks (CNN), scale-invariant feature transform (SIFT), orientation gradient histogram (HOG) features, or other techniques. Alternatively, vehicle 200 may detect other points, such as Figure 24A The edge point 2411 is shown, and the centerline point 2441 can be calculated, for example, by detecting points along each edge and determining the midpoint between the edge points. Similarly, the dashed lane marking 2420 can be represented by the centerline point 2451 along the centerline 2450 of the lane marking. The centerline point can be located at the edge of the short horizontal line, such as... Figure 24CAs shown, or located at various other locations along the centerline. For example, each short horizontal line can be represented by a single point at the geometric center of the short horizontal line. These points can also be spaced along the centerline at predetermined intervals (e.g., every meter, 5 meters, 10 meters, etc.). Centerline point 2451 can be directly detected by vehicle 200, or can be based on other detected reference points (such as... Figure 24B The corner point 2421 shown is used for calculation. Using a similar technique, the center line can also be used to indicate other lane marking types, such as double lines.
[0321] In some embodiments, vehicle 200 may identify points representing other features, such as vertices between two intersecting lane markings. Figure 24D An exemplary point representing the intersection between two lane markings 2460 and 2465 is shown. Vehicle 200 can calculate a vertex 2466 representing the intersection between the two lane markings. For example, one of lane markings 2460 or 2465 may represent a train crossing area or other crossing areas in a road segment. Although lane markings 2460 and 2465 are shown as intersecting each other perpendicularly, various other configurations can be detected. For example, lane markings 2460 and 2465 may intersect at other angles, or one or both lane markings may terminate at vertex 2466. Similar techniques can also be applied to intersections between dashed lines or other lane marking types. In addition to vertex 2466, various other points 2467 can also be detected, providing further information about the orientation of lane markings 2460 and 2465.
[0322] Vehicle 200 can associate real-world coordinates with each detected point of the lane marking. For example, a location identifier, including coordinates for each point, can be generated and uploaded to a server for map rendering of the lane marking. The location identifier can further include other identifying information about the point, including whether the point represents a corner, edge, center point, etc. Vehicle 200 can thus be configured to determine the real-world location of each point based on analysis of an image. For example, vehicle 200 can detect other features in the image, such as the various landmarks described above, to locate the real-world location of the lane marking. This may involve determining the location of the lane marking in the image relative to detected landmarks, or determining the vehicle's position based on detected landmarks, and then determining the distance from the vehicle (or the vehicle's target trajectory) to the lane marking. When landmarks are unavailable, the location of the lane marking point can be determined relative to the vehicle's position determined based on dead reckoning. The real-world coordinates included in the location identifier can be represented as absolute coordinates (e.g., latitude / longitude coordinates), or can be relative to other features, such as based on longitudinal position along the target trajectory and lateral distance from the target trajectory. Location identifiers can then be uploaded to a server for use in generating mapped lane markings in a navigation model (such as a sparse map 800). In some embodiments, the server may construct splines representing lane markings of road segments. Alternatively, vehicle 200 may generate splines and upload them to the server for recording in the navigation model.
[0323] Figure 24E An exemplary navigation model or sparse map is shown for a corresponding road segment including mapped lane markings. The sparse map may include a target trajectory 2475 for a vehicle to travel along the road segment. As described above, the target trajectory 2475 may represent the ideal path for a vehicle to take when traveling along the corresponding road segment, or it may be located elsewhere on the road (e.g., the centerline of the road, etc.). The target trajectory 2475 may be calculated using various methods described above, such as an aggregation (e.g., a weighted combination) of two or more reconstructed trajectories of vehicles traversing the same road segment.
[0324] In some embodiments, target trajectories can be generated equally for all vehicle types and for all road, vehicle, and / or environmental conditions. However, in other embodiments, various other factors or variables may also be considered when generating target trajectories. Different target trajectories may be generated for different types of vehicles (e.g., passenger cars, light trucks, and full trailers). For example, a target trajectory with a relatively tighter turning radius may be generated for a small passenger car compared to a larger semi-trailer truck. In some embodiments, road, vehicle, and environmental conditions may also be considered. For example, different target trajectories may be generated for different road conditions (e.g., wet, snowy, icy, dry, etc.), vehicle conditions (e.g., tire conditions or estimated tire conditions, braking conditions or estimated braking conditions, fuel remaining, etc.), or environmental factors (e.g., time of day, visibility, weather, etc.). Target trajectories may also depend on one or more aspects or characteristics of a particular road segment (e.g., speed limits, turning frequency and size, gradient, etc.). In some embodiments, various user settings may also be used to determine the target trajectory, such as set driving modes (e.g., desired driving aggressiveness, economy mode, etc.).
[0325] The sparse map may also include mapped lane markings 2470 and 2480 representing lane markings along road segments. The mapped lane markings may be represented by multiple location identifiers 2471 and 2481. As described above, the location identifiers may include locations in real-world coordinates associated with the detected lane markings. Similar to the target trajectory in the model, lane markings may also include elevation data and may be represented as curves in three-dimensional space. For example, the curve may be a spline connecting three-dimensional polynomials of appropriate order, which may be computed based on the location identifiers. The mapped lane markings may also include other information or metadata about the lane markings, such as identifiers of the type of lane marking (e.g., between two lanes with the same direction of travel, between two lanes with opposite directions of travel, the edge of a carriageway, etc.) and / or other characteristics of the lane markings (e.g., continuous, dashed, single line, double line, yellow, white, etc.). In some embodiments, the mapped lane markings may be continuously updated within the model, for example, using crowdsourcing techniques. The same vehicle can upload location identifiers during multiple instances of traveling on the same road segment, or it can select data from multiple vehicles (such as 1205, 1210, 1215, 1220, and 1225) traveling on the same road segment at different times. The sparse map 800 can then be updated or refined based on subsequent location identifiers received from the vehicle and stored in the system. As the mapped lane markings are updated and refined, the updated road navigation model and / or sparse map can be distributed to multiple autonomous vehicles.
[0326] Generating mapped lane markings in sparse maps can also include detecting and / or reducing errors based on anomalies in the image or the actual lane markings themselves. Figure 24F An exemplary anomaly 2495 associated with the detection of lane marking 2490 is shown. Anomaly 2495 may appear in an image captured by vehicle 200, for example, from objects obstructing the camera's view of the lane markings, debris on the lens, etc. In some cases, the anomaly may be due to the lane markings themselves, which may be damaged or worn, or partially covered by dirt, debris, water, snow, or other materials on the road. Anomaly 2495 may cause vehicle 200 to detect error point 2491. Sparse map 800 can provide correctly mapped lane markings and eliminate errors. In some embodiments, vehicle 200 may detect error point 2491, for example, by detecting anomaly 2495 in the image or by identifying errors based on lane marking points detected before and after the anomaly. Based on the detected anomaly, vehicle may ignore point 2491 or adjust it to be consistent with other detected points. In other embodiments, errors may be corrected after points have been uploaded, for example, by determining that a point is outside an expected threshold based on other points uploaded during the same trip or based on aggregation of data from previous trips along the same road segment.
[0327] Mapped lane markings in the navigation model and / or sparse map can also be used for navigation by autonomous vehicles traversing the corresponding lanes. For example, a vehicle navigating along a target trajectory can periodically use mapped lane markings in the sparse map to align itself with the target trajectory. As mentioned above, vehicles can navigate between landmarks based on dead reckoning, where the vehicle uses sensors to determine its own motion and estimate its position relative to the target trajectory. Errors can accumulate over time, and the vehicle's position determination relative to the target trajectory can become increasingly inaccurate. Therefore, vehicles can use lane markings (and their known locations) appearing in the sparse map 800 to reduce dead reckoning-induced errors in position determination. In this way, lane markings included in the sparse map 800 can act as navigation anchors from which the vehicle's accurate position relative to the target trajectory can be determined.
[0328] Figure 25A An exemplary image 2500 of the vehicle's surrounding environment, which can be used for navigation based on mapped lane markings, is shown. Image 2500 can be captured, for example, by vehicle 200 via image capture devices 122 and 124 included in image acquisition unit 120. Image 2500 may include an image of at least one lane marking 2510, such as... Figure 25A As shown. Image 2500 may also include one or more landmarks 2521, such as road signs, which are used for navigation as described above. Figure 25ASome elements shown, such as elements 2511, 2530 and 2520 (which do not appear in the captured image 2500, but are detected and / or determined by vehicle 200), are also shown for reference.
[0329] Use the above about Figures 24A to 24D and Figure 24F The various techniques described allow the vehicle to analyze image 2500 to identify lane markings 2510. Various points 2511 corresponding to features of the lane markings in the image can be detected. For example, point 2511 may correspond to the edge of a lane marking, a corner of a lane marking, the midpoint of a lane marking, a vertex between two intersecting lane markings, or various other features or locations. Point 2511 can be detected as a location corresponding to a point stored in a navigation model received from a server. For example, if a sparse map containing points representing the centerline of the mapped lane markings is received, point 2511 can also be detected based on the centerline of lane marking 2510.
[0330] The vehicle can also determine its longitudinal position, represented by element 2520 and located along the target trajectory. The longitudinal position 2520 can be determined from image 2500, for example, by detecting landmarks 2521 within image 2500 and comparing the measured location with known landmark locations stored in the road model or sparse map 800. The vehicle's location along the target trajectory can then be determined based on the distance to the landmark and the known location of the landmark. The longitudinal position 2520 can also be determined from images other than those used to determine the position of lane markings. For example, the longitudinal position 2520 can be determined by detecting landmarks in images taken simultaneously or nearly simultaneously with image 2500 from other cameras within image acquisition unit 120. In some cases, the vehicle may not be close to any landmarks or other reference points used to determine the longitudinal position 2520. In such cases, the vehicle can navigate based on dead reckoning and thus use sensors to determine its own motion and estimate its longitudinal position 2520 relative to the target trajectory. The vehicle can also determine a distance 2530, which represents the actual distance between the vehicle and lane markings 2510 observed in the captured images. When determining the distance 2530, camera angle, vehicle speed, vehicle width, or various other factors can be considered.
[0331] Figure 25BThis demonstrates lateral positioning correction for a vehicle based on mapped lane markings in a road navigation model. As described above, vehicle 200 can use one or more images captured by vehicle 200 to determine the distance 2530 between vehicle 200 and lane markings 2510. Vehicle 200 may also have access to a road navigation model, such as a sparse map 800, which may include mapped lane markings 2550 and a target trajectory 2555. The mapped lane markings 2550 can be modeled using the techniques described above (e.g., using crowdsourced location identifiers captured by multiple vehicles). The target trajectory 2555 can also be generated using various techniques previously described. Vehicle 200 can also determine or estimate the longitudinal position 2520 along the target trajectory 2555, as described above regarding... Figure 25A The vehicle 200 can then determine the expected distance 2540 based on the lateral distance between the target trajectory 2555 and the mapped lane markings 2550 corresponding to the longitudinal position 2520. The lateral positioning of the vehicle 200 can be corrected or adjusted by comparing the actual distance 2530 measured using captured images with the expected distance 2540 from the model.
[0332] Figure 25C and Figure 25D An illustration is provided in association with another example of locating the primary vehicle during navigation based on mapped landmarks / objects / features in a sparse map. Figure 25C This conceptually represents a series of images captured from a vehicle navigating along road segment 2560. In this example, road segment 2560 comprises a straight section of a highway divided into two lanes by road edges 2561 and 2562 and a center lane marker 2563. As shown, the primary vehicle is navigating along lane 2564 associated with a mapped target trajectory 2565. Therefore, ideally (and without influencing factors such as the presence of a target vehicle or object in the carriageway), the primary vehicle should closely track the mapped target trajectory 2565 as it navigates along lane 2564 of road segment 2560. In practice, the primary vehicle may experience drift while navigating along the mapped target trajectory 2565. For effective and safe navigation, this drift should be maintained within acceptable limits (e.g., a lateral displacement of + / - 10 cm from the target trajectory 2565 or any other suitable threshold). In order to periodically account for drift and make any necessary route corrections to ensure that the master vehicle follows the target trajectory 2565, the disclosed navigation system may be able to use one or more mapped features / objects included in the sparse map to locate the master vehicle along the target trajectory 2565 (e.g., determine the lateral and longitudinal position of the master vehicle relative to the target trajectory 2565).
[0333] As a simple example, Figure 25CA speed limit sign 2566 is shown, which may appear in five different, sequentially captured images as the master vehicle travels along road segment 2560. For example, at the first time t0, sign 2566 may appear in a captured image near the horizon. As the master vehicle approaches sign 2566, in subsequent captured images at times t1, t2, t3, and t4, sign 2566 will appear at different 2D XY pixel locations in the captured images. For example, in the captured image space, sign 2566 will move downwards and to the right along curve 2567 (e.g., a curve extending through the center of the sign in each of the five captured image frames). As the master vehicle approaches sign 2566, the sign will also appear to increase in size (i.e., it will occupy a large number of pixels in subsequent captured images).
[0334] These changes in the image spatial representation of objects such as sign 2566 can be used to determine the local location of the master vehicle along the target trajectory. For example, as described in this disclosure, any detectable object or feature (such as semantic features like sign 2566 or detectable non-semantic features) can be identified by one or more acquisition vehicles of a previously traversed road segment (e.g., road segment 2560). A mapping server can collect acquired driving information from multiple vehicles, aggregate and correlate that information, and generate a sparse map that includes, for example, the target trajectory 2565 for lane 2564 of road segment 2560. The sparse map may also store the location of sign 2566 (along with type information, etc.). During navigation (e.g., before entering road segment 2560), map tiles including the sparse map for road segment 2560 can be supplied to the master vehicle. To navigate in lane 2564 of road segment 2560, the master vehicle can follow the mapped target trajectory 2565.
[0335] The mapped representation of sign 2566 can be used by the master vehicle to locate itself relative to the target trajectory. For example, a camera on the master vehicle will capture an image 2570 of the master vehicle's environment, and this captured image 2570 may include an image representation of sign 2566 with a specific size and a specific XY image location, such as... Figure 25DAs shown. The size and XY image location can be used to determine the position of the primary vehicle relative to the target trajectory 2565. For example, based on a sparse map including the representation of sign 2566, the primary vehicle's navigation processor can determine that, in response to the primary vehicle traveling along the target trajectory 2565, the representation of sign 2566 should appear in the captured image such that the center of sign 2566 will move along line 2567 (in image space). If the captured image (such as image 2570) shows that the center (or other reference point) is displaced from line 2567 (e.g., the expected image space trajectory), the primary vehicle navigation system can determine that it was not on the target trajectory 2565 at the time the image was captured. However, based on the image, the navigation processor can determine appropriate navigation corrections to return the primary vehicle to the target trajectory 2565. For example, if analysis shows that the image location showing sign 2566 is displaced on line 2567 in the image by a distance 2572 to the left of the expected image space location, the navigation processor can cause a change in the primary vehicle's heading (e.g., a change in the steering angle of the wheels) to move the primary vehicle a distance 2573 to the left. In this way, each captured image can be used as part of a feedback loop process, minimizing the difference between the observed image position of marker 2566 and the expected image trajectory 2567, ensuring that the main vehicle continues along the target trajectory 2565 with little or no deviation. Of course, the more mapped objects available, the more frequently the described localization technique can be used, which can reduce or eliminate drift-induced deviations from the target trajectory 2565.
[0336] The process described above can be used to detect the lateral orientation or displacement of the primary vehicle relative to the target trajectory. The positioning of the primary vehicle relative to the target trajectory 2565 may also include determining the longitudinal location of the target vehicle along the target trajectory. For example, the captured image 2570 includes a representation of a marker 2566 with a specific image size (e.g., a 2D XY pixel region). As the mapped marker 2566 travels along line 2567 through the image space (e.g., as the size of the marker gradually increases, such as...), the positioning of the primary vehicle relative to the target trajectory 2565 may be affected by the following process: Figure 25C As shown, this size can be compared to the expected image size of the mapped sign. Based on the image size of sign 2566 in image 2570, and based on the expected size progression in image space relative to the mapped target trajectory 2565, the host vehicle can determine its longitudinal position relative to the target trajectory 2565 (at the time when image 2570 is captured). This longitudinal position, coupled with any lateral displacement relative to the target trajectory 2565 as described above, allows for complete positioning of the host vehicle relative to the target trajectory 2565 as the host vehicle navigates along road 2560.
[0337] Figure 25C and Figure 25DThis is just one example of a disclosed localization technique using a single mapped object and a single target trajectory. In other examples, there may be more target trajectories (e.g., one target trajectory for each lane of a multi-lane highway, city street, complex intersection, etc.) and there may be more maps available for localization. For example, a sparse map representing an urban environment may include many objects available for localization per meter.
[0338] Figure 26A A flowchart illustrating an exemplary process 2600A for mapping lane markings for use in autonomous vehicle navigation, consistent with the disclosed embodiments, is provided. At step 2610, process 2600A may include receiving two or more location identifiers associated with detected lane markings. For example, step 2610 may be performed by server 1230 or one or more processors associated with that server. Location identifiers may include the location in real-world coordinates of the point associated with the detected lane markings, as described above regarding… Figure 24E As described. In some embodiments, the location identifier may also include other data, such as additional information about road segments or lane markings. Additional data, such as accelerometer data, speed data, landmark data, road geometry or contour data, vehicle positioning data, self-motion data, or various other forms of data described above, may also be received during step 2610. The location identifier may be generated by vehicles such as vehicles 1205, 1210, 1215, 1220, and 1225 based on images captured by the vehicles. For example, the identifier may be determined based on acquiring at least one image representing the environment of the main vehicle from a camera associated with the main vehicle, analyzing at least one image to detect lane markings in the environment of the main vehicle, and analyzing at least one image to determine the position of the detected lane markings relative to a location associated with the main vehicle. As described above, lane markings may include a variety of different marking types, and the location identifier may correspond to a variety of points relative to the lane markings. For example, in the case where the detected lane marking is part of a dashed line marking the boundary of a lane, the point may correspond to a detected corner of the lane marking. When the detected lane marking is part of a continuous line marking the lane boundary, the point may correspond to the detected edge of the lane marking, having various spacings as described above. In some embodiments, the point may correspond to the centerline of the detected lane marking, such as... Figure 24C As shown, or it may correspond to the vertex between two intersecting lane markings and at least two other points associated with the intersecting lane markings, such as... Figure 24D As shown.
[0339] At step 2612, process 2600A may include associating the detected lane markings with corresponding road segments. For example, server 1230 may analyze real-world coordinates or other information received during step 2610 and compare the coordinates or other information with location information stored in the autonomous vehicle road navigation model. Server 1230 may determine the road segment in the model corresponding to the real-world road segment of the detected lane markings.
[0340] At step 2614, process 2600A may include updating the autonomous vehicle road navigation model relative to the corresponding road segment based on two or more location identifiers associated with the detected lane markings. For example, the autonomous road navigation model may be a sparse map 800, and server 1230 may update the sparse map to include or adjust the mapped lane markings in the model. Server 1230 may update the sparse map based on the above regarding… Figure 24E Various methods or processes are described to update the model. In some embodiments, updating the autonomous vehicle road navigation model may include storing one or more indicators of the real-world coordinates of detected lane markings. The autonomous vehicle road navigation model may also include at least one target trajectory for the vehicle to travel along corresponding road segments, such as... Figure 24E As shown.
[0341] At step 2616, process 2600A may include distributing the updated autonomous vehicle road navigation model to multiple autonomous vehicles. For example, server 1230 may distribute the updated autonomous vehicle road navigation model to vehicles 1205, 1210, 1215, 1220, and 1225 that can use the model for navigation. The autonomous vehicle road navigation model may be distributed via one or more networks (e.g., via cellular networks and / or the Internet, etc.) through wireless communication path 1235, such as... Figure 12 As shown.
[0342] In some embodiments, lane markings can be mapped using data received from multiple vehicles, such as through crowdsourcing technologies, as described above. Figure 24EAs described. For example, process 2600A may include receiving a first communication from a first master vehicle, including a location identifier associated with the detected lane marking, and receiving a second communication from a second master vehicle, including an additional location identifier associated with the detected lane marking. For example, the second communication may be received from a subsequent vehicle traveling on the same road segment, or from the same vehicle in a subsequent journey along the same road segment. Process 2600A may further include refining the determination of at least one location associated with the detected lane marking based on the location identifier received in the first communication and based on the additional location identifier received in the second communication. This may include using an average of multiple location identifiers and / or filtering out “ghost” identifiers that may not reflect the real-world location of the lane marking.
[0343] Figure 26B A flowchart illustrating an exemplary process 2600B for autonomously navigating a host vehicle along a road segment using mapped lane markings is provided. Process 2600B may be performed, for example, by processing unit 110 of autonomous vehicle 200. At step 2620, process 2600B may include receiving an autonomous vehicle road navigation model from a server-based system. In some embodiments, the autonomous vehicle road navigation model may include a target trajectory for the host vehicle along a road segment and location identifiers associated with one or more lane markings associated with the road segment. For example, vehicle 200 may receive a sparse map 800 or another road navigation model developed using process 2600A. In some embodiments, the target trajectory may be represented as a three-dimensional spline, for example, as shown in the diagram. Figure 9B As shown above regarding... Figures 24A to 24F As described, the location identifier may include the location of the point associated with the lane marking in real-world coordinates (e.g., the corner point of a dashed horizontal line lane marking, the edge point of a continuous lane marking, the vertex between two intersecting lane markings and other points associated with intersecting lane markings, the center line associated with the lane marking, etc.).
[0344] At step 2621, process 2600B may include receiving at least one image representing the environment of the vehicle. The image may be received from an image capturing device of the vehicle, such as image capturing devices 122 and 124 included in image acquisition unit 120. The image may include images of one or more lane markings, similar to image 2500 described above.
[0345] At step 2622, process 2600B may include determining the longitudinal position of the master vehicle along the target trajectory. (As mentioned above regarding...) Figure 25A As described, this can be based on other information in the captured images (e.g., landmarks, etc.) or by dead reckoning of the vehicle between detected landmarks.
[0346] At step 2623, process 2600B may include determining the expected lateral distance to the lane markings based on the determined longitudinal position of the main vehicle along the target trajectory and based on two or more location identifiers associated with at least one lane marking. For example, vehicle 200 may use sparse map 800 to determine the expected lateral distance to the lane markings. Figure 25B As shown, the longitudinal position 2520 along the target trajectory 2555 can be determined in step 2622. Using the sparse map 800, the vehicle 200 can determine the expected distance 2540 to the mapped lane markings 2550 corresponding to the longitudinal position 2520.
[0347] At step 2624, process 2600B may include analyzing at least one image to identify at least one lane marking. For example, vehicle 200 may use various image recognition techniques or algorithms to identify lane markings within the image, as described above. For instance, lane marking 2510 may be detected through image analysis of image 2500, such as... Figure 25A As shown.
[0348] At step 2625, process 2600B may include determining the actual lateral distance to at least one lane marking based on analysis of at least one image. For example, the vehicle may determine a distance 2530 representing the actual distance between the vehicle and lane marking 2510, such as... Figure 25A As shown. When determining the distance 2530, factors such as camera angle, vehicle speed, vehicle width, camera position relative to the vehicle, and various other factors can be considered.
[0349] At step 2626, process 2600B may include determining an autonomous steering action for the master vehicle based on the difference between the expected lateral distance to at least one lane marking and the determined actual lateral distance to at least one lane marking. For example, as described above regarding Figure 25B As described, vehicle 200 can compare actual distance 2530 with expected distance 2540. The difference between the actual distance and the expected distance indicates the error (and its magnitude) between the vehicle's actual position and the target trajectory to be followed by the vehicle. Therefore, the vehicle can determine autonomous steering actions or other autonomous actions based on this difference. For example, if the actual distance 2530 is less than the expected distance 2540, such as... Figure 25B As shown, the vehicle can determine autonomous steering actions to guide it to the left away from lane marking 2510. Therefore, the vehicle's position relative to the target trajectory can be corrected. Process 2600B can be used, for example, to improve vehicle navigation between landmarks.
[0350] Processes 2600A and 2600B only provide examples of techniques that can be used to navigate a master vehicle using the disclosed sparse map. In other examples, compared to... Figure 25C and Figure 25DA process that is consistent with the described process can also be executed.
[0351] Path prediction network As described herein, navigation systems (including those for autonomous or semi-autonomous vehicles) can use computer vision solutions to detect features in the vehicle's environment to predict drivable paths along a roadway. For example, the primary vehicle's navigation system can identify lane markings or other road features to navigate safely and accurately along a roadway. To this end, the navigation system can analyze images captured from cameras mounted on the primary vehicle, detect features within the images, and estimate drivable paths along the roadway. In some embodiments, the primary vehicle may also use one or more navigation maps that include a target trajectory of the primary vehicle. For example, the navigation system can determine the position of the primary vehicle relative to a navigation map, and thus determine the position of the primary vehicle relative to a target trajectory within the navigation map, which may represent a drivable path for the primary vehicle.
[0352] These technologies can identify predicted target trajectories in the lane in which the primary vehicle is traveling, but may not necessarily provide information about other available trajectories. For example, these technologies may not be able to identify trajectories associated with adjacent lanes in areas where navigation maps are unavailable. To provide more comprehensive information about available target trajectories, the system can be trained to identify target trajectories using training images and / or navigation map data. Once trained, the system can predict target trajectories not only for the lane in which the primary vehicle is traveling, but also for one or more other lanes represented in the image. This system can be particularly useful when navigating in areas where navigation maps are unavailable. Therefore, the disclosed embodiments provide improved efficiency, accuracy, and performance compared to conventional navigation technologies.
[0353] To identify target trajectories in the environment of the primary vehicle, the navigation system for the primary vehicle can capture and analyze one or more images as described herein. Figure 27 Example image 2700, consistent with the disclosed embodiments, is shown and can be used to predict the trajectory of a target along a roadway. Image 2700 may be captured by a camera of the host vehicle (such as image capture devices 122, 124, and / or 126 of the host vehicle 200). Figure 27In the example shown, images can be captured from the front-facing camera of the main vehicle as it travels along the roadway. In this example, the roadway may include multiple lanes 2702, 2704, and 2706. Lane 2704 may be the lane in which the main vehicle is currently traveling, and lanes 2702 and 2706 may be lanes adjacent to the lane in which the main vehicle is currently traveling. In this example, lane 2706 may be in the same direction of travel as lane 2704, while lane 2702 may be associated with a direction of travel opposite to that of lanes 2704 and 2706. Therefore, vehicles traveling in lane 2702 may be oncoming traffic towards the main vehicle.
[0354] Image 2700 may include representations of various features within the environment of the main vehicle. For example, such as... Figure 27 As shown, image 2700 may include representations of lane markings 2710, 2712, 2714, and 2716, which may define various boundaries of lanes 2702, 2704, and 2706. In this example, as shown, lane markings 2710 and 2712 may define the boundary of lane 2802, lane markings 2712 and 2714 may define the boundary of lane 2804, and lane markings 2714 and 2716 may define the boundary of lane 2806. As discussed further below, the type of road feature represented in the image is not limited to lane markings, and any feature that can indicate the existence of a drivable path along the carriageway may be used.
[0355] Based on image 2700, the trained system can determine one or more target trajectories associated with a carriageway. In one example, based on image 2700, the trained system can determine a target trajectory for its own lane (i.e., the lane in which the vehicle equipped with the camera for capturing the image is currently located), and using the same image 2700, the trained system can determine a target trajectory for at least one other lane adjacent to its own lane. In yet another example, based on image 2700, the trained system can determine target trajectories for its own lane and for any other lane directly or indirectly adjacent to its own lane (e.g., another lane, two other lanes, three other lanes, etc.), wherein the one or more other lanes are included in the field of view of the camera used to capture image 2700. In this respect, a second lane indirectly adjacent to its own lane may be directly adjacent to a lane that is itself directly adjacent to its own lane (and the latter lane may be located between these two lanes). In order for the trained system to provide target trajectories for lanes appearing in image 2700, the trained system may require a minimum amount or a minimum range of information, features, context, etc., to provide the target trajectory for that lane. When this document refers to any other lane included in the field of view of the camera used to capture 2700 images, it is intended to refer to any lane where there is a minimum amount or range of information, features, context, etc.
[0356] Figure 28 Example target trajectories 2802, 2804, and 2806, consistent with the disclosed embodiments, are shown. As described in further detail above, target trajectories 2802, 2804, and 2806 can represent preferred or recommended drivable paths for each available lane of a roadway. In this example, target trajectories 2802, 2804, and 2806 can correspond to lanes 2702, 2704, and 2706, respectively. Therefore, target trajectories can be predicted or identified not only for the primary vehicle (i.e., its "self-trajectory," such as trajectory 2804) but also for other drivable paths. For example, target trajectories 2802 and 2806 can represent trajectories for vehicles traveling along lanes 2702 and 2706, which are adjacent to the primary vehicle's current driving lane. Target trajectories 2802, 2804, and 2806 can be represented in various ways. In some embodiments, target trajectories can be represented as a plurality of points along the surface of a road segment. In some embodiments, these points can be 3D points on the surface that traces the road. Alternatively or additionally, the target trajectory can be represented as a 3D spline, as described above.
[0357] In some embodiments, the system may determine the direction of travel associated with target trajectories 2802, 2804, and 2806. For example, target trajectory 2802 may be associated with the direction of travel indicating that lane 2702 is a lane for oncoming traffic, which is opposite to the direction of travel associated with target trajectories 2804 and 2806. The direction of travel may be derived from contextual "cues" within the image. For example, lane marking 2712 may be a double solid line, which may indicate the boundary between different directions of travel along the carriageway. Various other indicators of the direction of travel may include the presence of vehicles traveling along the lane (and the direction the vehicles are facing), the position of the lane relative to the carriageway, the position of the lane relative to the lane from which the image was captured, traffic signs, traffic signals or traffic lights, arrows or other lane markings, etc. As further described below, the system may not necessarily identify any single feature in the image that indicates the direction of travel, but the determination may be based on a combination of features identified using a trained model.
[0358] Figure 29 An example process for training and implementing a trained model 2930 for predicting target trajectories, consistent with the disclosed embodiments, is shown. The trained model 2930 can be trained using a set of training data 2910 that can be input into the training algorithm 2920, such as... Figure 29 As indicated herein. Training data 2910 may include at least one training image, such as training image 2912 used to train the trained model 2930. Training image 2912 may include any image representing features that can be associated with a drivable path along the roadway. For example, training image 2912 may be similar to image 2700 described above. Thus, as described herein, training image 2912 may represent an image captured by a vehicle as it traverses a road segment.
[0359] In some embodiments, training image 2912 may be labeled to indicate the target trajectory associated with the training image. For example, similar to target trajectories 2802, 2804, and 2806 shown relative to image 2700, training image 2912 may be labeled to indicate the target trajectory represented in training image 2912. Thus, by inputting the labeled image into training algorithm 2912, trained model 2930 can predict the relative location of the target trajectory based on other images. In some embodiments, the labels provided as part of the training data may include labels unrelated to the motion of any particular vehicle traversing the corresponding environment. For example, the labels may be based on trajectories determined according to the average path traveled by multiple vehicles along road segments. Alternatively or additionally, the labels may be based on data from other sources, such as the center of the carriageway, trajectories stored in a mapping database, etc. In some embodiments, the labels may be based on the actual path chosen by the vehicle when capturing images used in the training dataset.
[0360] Consistent with some embodiments, training data 2910 may include map information 2914, such as... Figure 29As indicated in the document. Map information 2914 may include any data representing one or more roads that can be navigated by the vehicle. In some embodiments, map information 2914 may include one or more AV maps. Autonomous vehicle (AV) maps (or AV maps) may include information for supporting and / or enabling one or more autonomous vehicle (AV) functions of the vehicle in a manner that enables the vehicle to operate safely and / or navigate accurately. AV functions supported and / or implemented by autonomous vehicles (or semi-autonomous vehicles) may include one or more autonomously controlled functions (e.g., functions determined, selected, and / or implemented based on instructions executed by at least one processor) such as steering, accelerating, and / or braking the vehicle. AV functions may be part of a driving strategy such as RSS, which is developed and implemented, for example, by Mobileye in Jerusalem, Israel. Information used for safe operation of a vehicle may include, but is not limited to, information relating to one or more regulations applicable to the location or jurisdiction of the vehicle (e.g., right-hand traffic or left-hand traffic jurisdiction), information relating to the environment in which the vehicle is located (e.g., information relating to drivable paths, stop signs, traffic lights, speed limits, lane markings, landmarks, free space, virtual or physical stop lines, traffic light-related information, etc.), and / or information relating to adjusting and / or adapting the vehicle's navigation to take into account one or more objects in the vehicle's environment (e.g., other vehicles, pedestrians, objects, obstacles, obstructions, hazards, road work zones, traffic cones, etc.). The vehicle may use one or more sensors (e.g., cameras, radar, lidar) to sense objects in the vehicle's environment, as discussed herein. AV maps may serve as a redundant source of information for the sensed information and, in some cases, may supplement the sensed information (e.g., providing locations with virtual stop lines where no stop lines are marked on the road). In some embodiments, safe operation of a vehicle may further include operating the vehicle to maintain a level of comfort for one or more passengers in the vehicle. Comfort levels may include one or more predetermined criteria (e.g., related to speed, acceleration, and / or cornering) selected to cause the vehicle to operate in a manner that meets or exceeds a specified or selected passenger comfort level. Comfort levels can be formalized and expressed using appropriate mathematical formulas. For example, mathematical formulas limiting the degree of abrupt movement or acceleration applied to passengers in different directions may be used. Information for operating the vehicle in an accurate manner may include, but is not limited to, information relating to a planned or specified navigation path (e.g., a trajectory, as discussed herein) or route (e.g., from a specific location such as a starting point to a destination). In some embodiments, the information included in the AV map may further include information on how to effectively support and / or implement one or more AV functions.Information used to operate the vehicle effectively may include information relating to speed, acceleration, lane changes, and / or lane positioning, and / or information on driving and / or selecting a path or route based on traffic conditions (e.g., driving a route that is longer than the shorter route that experiences the traffic conditions) and / or other factors or attributes associated with the potential route (e.g., weather conditions, road conditions, or other characteristics of the route, such as traveling on a highway without traffic lights instead of a road with traffic lights). In some embodiments, the AV map may further include at least some information from a high-resolution (HD) map. In some embodiments, the AV map may be a sparse map, as described above. In some embodiments, the AV map may be crowdsourced, as discussed herein. In this disclosure, the terms "AV map" and "sparse map" are used interchangeably.
[0361] In some embodiments, map information 2914 may provide information that associates a target trajectory with training image 2912. For example, training image 2912 may be associated with location data to indicate the position of the image relative to the AV map. For example, training image 2912 may include metadata or other information indicating the position of the captured image relative to the AV map and / or the camera orientation at the time of image capture. Therefore, the position of the target trajectory in the AV map relative to training image 2912 can be used to train trained model 2930 to determine similar trajectories for other images. In some embodiments, this may include converting location information of the target trajectory from road model space to image space. In other words, based on the position and orientation of the training image in the AV map and the position of one or more target trajectories in the AV map, a representation of the target trajectory can be projected onto the image, similar to... Figure 28 The target trajectories 2802, 2804, and 2806 are shown relative to image 2700. Therefore, training image 2912 can be labeled based on map information 2914. Alternatively or additionally, map information 2914 and training image 2912 can be provided separately to training algorithm 2920, and the labeling of target trajectories relative to training image 2912 can occur intrinsically during training.
[0362] In some embodiments, map information 2914 and / or training image 2912 can be generated from crowdsourced data, as described herein. For example, various vehicles can travel along road segments and collect image data, including training image 2912. The paths traveled by multiple vehicles along road segments can be averaged together to generate a target trajectory within the AV map. This may also include detecting and storing the locations of various other landmarks represented in the image data. Thus, data received from different vehicles can be combined to generate and / or update a road model representing the road segments.
[0363] In some embodiments, training data 2910 may include one or more deformed images, such as deformed image 2916. The deformed image may be appended to or used as a replacement for training image 2912, and may be used as input to training algorithm 2920, similar to training image 2912. Consistent with the disclosed embodiments, the deformed image may be generated based on an undeformed captured image, such as training image 2912. As used herein, "deformed image" may refer to an image that has been digitally processed to distort the location of features appearing within the image. For example, the locations of various features in the original image may be remapped to appear at different locations in the deformed image. The generated deformed image may simulate a view of features in the vehicle's environment from a simulated "drone view," which may be derived from a simulated vantage point elevated relative to the actual camera used to capture the image. This simulated view can allow for improved detection of lane markings and other features within the image, particularly at greater distances. For example, road features in the deformed view may appear larger than in the original image, making them easier to detect. Furthermore, deformable views can normalize the geometry of lane markings, making them easier to detect. For example, the geometry of lane markings or other road features can be detected with a higher confidence level than lane markings detected in the captured image. In other embodiments, lane markings or other features can be detected faster and / or more efficiently (e.g., using fewer computational resources) compared to detecting the same features in the captured image.
[0364] Deformed images can be generated using various transformation algorithms and techniques. For example, one or more transformation algorithms can be used to translate, scale, and / or rotate pixels in the original image to form a deformed image. Such transformations can include Prototype transforms, affine transforms, perspective transforms, bilinear transforms, polynomial transforms, elastic deformation, thin-plate spline techniques, Bayesian methods, mesh deformation techniques, or any other transformation. In some embodiments, a combination of one or more transformation techniques can be used.
[0365] In some embodiments, the representation of a feature in the deformed image may include more image pixels than the representation of a feature in the original captured image, which can make the feature easier to detect. For example, due to the deformation, pixels representing the same feature in the original image may be spaced further apart in the deformed image compared to the spacing in the original image. In some embodiments, additional pixels may be added as part of generating the deformed image 2916 to account for the increased spacing between pixels. Thus, the feature represented in the original image can have increased resolution in the deformed image.
[0366] When morphing an image, various algorithms can be used to generate intermediate pixels, including, for example, upsampling algorithms configured to increase the resolution of pixels in the image. Such algorithms may include, for example, nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, sine or Lanzos resampling, box sampling, Fourier transform methods, edge-guided interpolation, high-quality scaling (“hqx”), seam cropping, and / or vector extraction. In some embodiments, machine learning models, including models based on deep convolutional neural networks (such as waifu2x, Neural Enhance, Topaz Gigapixel A.I., etc.), can be used to increase the image resolution. Such algorithms can be applied to the entire morphed image or to certain regions depending on the degree and type of transformation performed in that region. Various denoising algorithms can also be applied to smooth the upsampled image or portions of the image. Therefore, features represented at higher resolution in the morphed image may be easier to detect, and thus the accuracy or efficiency of training the model can be improved.
[0367] As des...
Claims
1. A navigation system for a host vehicle, the navigation system comprising: At least one processor, the at least one processor including a circuit system and a memory, wherein the memory includes instructions that, when executed by the circuit system, cause the at least one processor to: Receive at least one image captured by a camera mounted on the main vehicle, wherein the at least one image includes a representation of two or more features; The at least one image is provided to a trained model, which is configured to generate outputs that identify two or more target trajectories associated with each of the two or more features; Based on the output generated by the trained model, determine the location information of the two or more target trajectories associated with each of the two or more features; At least one navigation action of the master vehicle is determined based on the location information determined for at least one of the target trajectories; as well as This enables the main vehicle to perform at least one navigation action.
2. The system of claim 1, wherein the two or more features are associated with the road surface.
3. The system of claim 2, wherein the two or more features comprise two or more lane markings.
4. The system of claim 1, wherein the two or more features include a first feature and a second feature, wherein the first feature and the second feature intersect to form an intersection.
5. The system of claim 1, wherein the two or more features are associated with lanes that extend together in the longitudinal direction of a road segment.
6. The system of claim 1, wherein the two or more features are arranged in series along the longitudinal direction.
7. The system of claim 1, wherein each of the two or more features comprises at least one of a main vehicle lane, an adjacent lane in the direction of travel of the main vehicle, an adjacent lane in the direction of travel opposite to the main vehicle, a shoulder, a parking space or parking lot vacancy.
8. The system of claim 1, wherein the two or more target trajectories comprise at least one self-trajectories and at least one other trajectory different from the self-trajectories, the self-trajectories representing the estimated trajectory of the master vehicle ahead of the current location.
9. The system of claim 1, wherein at least one of the target trajectories is associated with the current driving path of the master vehicle.
10. The system of claim 10, wherein the current travel path of the master vehicle includes the current travel lane associated with the master vehicle.
11. The system of claim 1, wherein the at least one navigation action includes changing the navigation path of the master vehicle from one of the two or more target trajectories to another target trajectories.
12. The system of claim 1, wherein the at least one navigation action includes abandoning the change of the navigation path of the master vehicle from one of the two or more target trajectories to another target trajectories.
13. The system of claim 12, wherein the other target trajectory in the target trajectory is associated with oncoming traffic relative to the master vehicle.
14. The system of claim 1, wherein the at least one navigation action includes modifying the heading direction of the master vehicle to intersect with and follow one of the two or more target trajectories.
15. The system of claim 1, wherein the at least one navigation action includes changing the master vehicle from its current lane to an adjacent lane to follow one of the two or more target trajectories.
16. The system of claim 1, wherein the at least one navigation action includes at least one of steering, braking, or accelerating the master vehicle.
17. The system of claim 1, wherein the trained model has undergone at least one training process based on at least one training image and map information.
18. The system of claim 17, wherein the map information includes autonomous vehicle (AV) maps.
19. The system of claim 17, wherein the map information is based on information received from multiple vehicles.
20. The system of claim 17, wherein the map information includes information for positioning accurate to at least 10 centimeters.
21. The system of claim 17, wherein the map information includes a sparse map.
22. The system of claim 21, wherein the sparse map has a data density of no more than 1 megabyte per kilometer.
23. The system of claim 17, wherein the at least one training image comprises a deformed image.
24. The system of claim 23, wherein the deformed image is deformed relative to the ground plane associated with the road surface.
25. The system of claim 24, wherein the deformed image simulates a viewpoint elevated at least eight meters above the road surface.
26. The system of claim 24, wherein the deformed image simulates a viewpoint elevated at least ten meters above the road surface.
27. The system of claim 24, wherein the deformed image simulates a viewpoint elevated between ten and twenty meters above the road surface.
28. The system of claim 23, wherein the deformed image includes a representation of multiple lanes.
29. The system of claim 17, wherein the at least one training image comprises a plurality of training images.
30. The system of claim 17, wherein the trained model comprises a machine learning model.
31. The system of claim 30, wherein the machine learning model includes at least one of a convolutional neural network (CNN) or a random forest model.
32. The system of claim 1, wherein the trained model has undergone at least one training process based on at least one training image.
33. The system of claim 32, wherein the at least one training image comprises a deformed image.
34. The system of claim 33, wherein the deformed image is deformed relative to the ground plane associated with the road surface.
35. The system of claim 34, wherein the deformed image simulates a viewpoint elevated at least eight meters above the road surface.
36. The system of claim 34, wherein the deformed image simulates a viewpoint elevated at least ten meters above the road surface.
37. The system of claim 34, wherein the deformed image simulates a viewpoint elevated between ten and twenty meters above the road surface.
38. The system of claim 33, wherein the deformed image includes a representation of multiple lanes.
39. The system of claim 32, wherein the at least one training image comprises a plurality of training images.
40. The system of claim 32, wherein the trained model comprises a machine learning model.
41. The system of claim 40, wherein the machine learning model comprises at least one of a convolutional neural network (CNN) or a random forest model.
42. The system of claim 32, wherein the at least one training process is further based on map information.
43. The system of claim 42, wherein the map information includes an autonomous vehicle (AV) map.
44. The system of claim 42, wherein the map information is based on information received from multiple vehicles.
45. The system of claim 1, wherein the output generated by the trained model further includes an indicator of whether each of the two or more features corresponds to a drivable path.
46. The system of claim 1, wherein the master vehicle is navigating in an area where no map has been drawn.
47. The system of claim 46, wherein the map information accessible to the main vehicle omits information about the unmapped areas.
48. The system of claim 1, wherein the location information of the two or more target trajectories includes location information for each target trajectory that is at least partially represented in the captured image.
49. A method for navigating a master vehicle, the method comprising: Receive at least one image captured by a camera mounted on the main vehicle, wherein the at least one image includes a representation of two or more features; The at least one image is provided to a trained model, which is configured to generate outputs that identify two or more target trajectories associated with each of the two or more features; Based on the output generated by the trained model, determine the location information of the two or more target trajectories associated with each of the two or more features; At least one navigation action of the master vehicle is determined based on the location information determined for at least one of the target trajectories; as well as This enables the main vehicle to perform at least one navigation action.
50. The method of claim 49, wherein the two or more features include a first feature and a second feature, wherein the first feature and the second feature intersect to form an intersection.
51. The method of claim 49, wherein the at least one navigation action comprises changing the navigation path of the master vehicle from one of the two or more target trajectories to another target trajectories.
52. The method of claim 49, wherein the master vehicle is navigating in an area where no map has been drawn.
53. The method of claim 52, wherein the map information accessible to the main vehicle omits information about the unmapped areas.
54. A non-transitory computer-readable medium storing instructions executable by at least one processor to perform a method for navigating a master vehicle, the method comprising: Receive at least one image captured by a camera mounted on the main vehicle, wherein the at least one image includes a representation of two or more features; The at least one image is provided to a trained model, which is configured to generate outputs that identify two or more target trajectories associated with each of the two or more features; Based on the output generated by the trained model, determine the location information of the two or more target trajectories associated with each of the two or more features; At least one navigation action of the master vehicle is determined based on the location information determined for at least one of the target trajectories; as well as This enables the main vehicle to perform at least one navigation action.
55. The non-transitory computer-readable medium of claim 54, wherein the trained model has undergone at least one training process based on at least one training image and map information.
56. The non-transitory computer-readable medium of claim 55, wherein the map information is based on information received from multiple vehicles.
57. The non-transitory computer-readable medium of claim 55, wherein the at least one training image comprises a deformed image that simulates a viewpoint elevated relative to the road surface.
58. A system for determining a drivable path for a primary vehicle relative to a road segment, the system comprising: At least one processor, the at least one processor including a circuit system and a memory, wherein the memory includes instructions that, when executed by the circuit system, cause the at least one processor to: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; as well as The representation of the at least one drivable path of the primary vehicle is provided to the primary vehicle navigation system, which is configured to navigate the primary vehicle relative to the at least one drivable path of the primary vehicle.
59. The system of claim 58, wherein the captured images are acquired by a camera mounted on a main vehicle traversing the road segment.
60. The system of claim 58, wherein the representation of the at least one drivable path of a primary vehicle comprises a plurality of points along the surface of the road segment.
61. The system of claim 60, wherein the plurality of points are 3D points.
62. The system of claim 58, wherein the representation of the at least one drivable path of a master vehicle comprises 3D splines.
63. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises a plurality of drivable paths for a primary vehicle inferred by the trained neural network based on the captured image received as input to the trained neural network.
64. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for the lane in which the primary vehicle is currently traveling.
65. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises one or more drivable paths generated for lanes other than the lane in which the primary vehicle is currently traveling.
66. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for each lane at least partially represented in the captured image.
67. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for lane merging features of the road segment represented in the captured image.
68. The system of claim 58, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for lane segmentation features of the road segment represented in the captured image.
69. The system of claim 58, wherein the at least one drivable path for the primary vehicle is provided by the trained neural network based on captured images excluding lane markings.
70. The system of claim 58, wherein the captured image provided as input to the trained neural network is deformed to simulate an advantageous point above the road surface that is higher than the camera that acquired the captured image.
71. The system of claim 58, wherein the trained neural network is trained based on an image representation generated based on a sparse map comprising a plurality of stored three-dimensional drivable paths.
72. The system of claim 71, wherein each of the stored three-dimensional drivable paths is associated with a different lane extending along a training road segment.
73. A method for determining and mapping a primary vehicle drivable path relative to a road segment, the method comprising: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; as well as The representation of the at least one drivable path of the primary vehicle is provided to the primary vehicle navigation system, which is configured to navigate the primary vehicle relative to the at least one drivable path of the primary vehicle.
74. The method of claim 73, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for lane merging features of the road segment represented in the captured image.
75. The method of claim 73, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for lane segmentation features of the road segment represented in the captured image.
76. The method of claim 73, wherein the at least one drivable path for the primary vehicle is provided by the trained neural network based on captured images excluding lane markings.
77. A system for determining and mapping a master vehicle drivable path relative to a road segment, the system comprising: At least one processor, the at least one processor including a circuit system and a memory, wherein the memory includes instructions that, when executed by the circuit system, cause the at least one processor to: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; The representation of the at least one drivable path of the main vehicle is stored in a map; as well as The map is distributed to at least one primary vehicle navigation system for navigating the primary vehicle along the road segment relative to the at least one drivable path of the primary vehicle stored in the map.
78. The system of claim 78, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for lane merging features of the road segment represented in the captured image.
79. The system of claim 78, wherein the at least one drivable path of the primary vehicle comprises a drivable path generated for lane segmentation features of the road segment represented in the captured image.
80. The system of claim 78, wherein the at least one drivable path for the primary vehicle is provided by the trained neural network based on captured images excluding lane markings.
81. A non-transitory computer-readable medium storing instructions executable by at least one processor for determining a master vehicle drivable path relative to a road segment according to a method comprising: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; as well as The representation of the at least one drivable path of the primary vehicle is provided to the primary vehicle navigation system, which is configured to navigate the primary vehicle relative to the at least one drivable path of the primary vehicle.
82. The non-transitory computer-readable medium of claim 81, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for a lane in which the primary vehicle is currently traveling.
83. The non-transitory computer-readable medium of claim 81, wherein the at least one drivable path for a primary vehicle comprises one or more drivable paths generated for lanes other than the lane in which the primary vehicle is currently traveling.
84. The non-transitory computer-readable medium of claim 81, wherein the at least one drivable path of the primary vehicle comprises a drivable path generated for each lane at least partially represented in the captured image.
85. A non-transitory computer-readable medium storing instructions executable by at least one processor for determining and mapping a primary vehicle-drivable path relative to a road segment according to a method comprising: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; The representation of the at least one drivable path of the main vehicle is stored in a map; as well as The map is distributed to at least one primary vehicle navigation system for navigating the primary vehicle along the road segment relative to the at least one drivable path of the primary vehicle stored in the map.
86. The non-transitory computer-readable medium of claim 85, wherein the at least one primary vehicle drivable path comprises a plurality of primary vehicle drivable paths inferred by the trained neural network based on the captured image received as input to the trained neural network.
87. The non-transitory computer-readable medium of claim 85, wherein the trained neural network is trained based on an image representation generated based on a sparse map comprising a plurality of stored three-dimensional drivable paths.
88. The non-transitory computer-readable medium of claim 87, wherein each of the stored three-dimensional drivable paths is associated with a different lane extending along a training road segment.
89. A method for determining and mapping a primary vehicle drivable path relative to a road segment, the method comprising: Receive captured images representing at least a portion of the road segment; The captured image is provided to a trained neural network, wherein the trained neural network is configured to receive the captured image as input and provide an output including a representation of at least one main vehicle drivable path relative to the road segment based on at least one road terrain feature represented in the captured image; The representation of the at least one drivable path of the main vehicle is stored in a map; as well as The map is distributed to at least one primary vehicle navigation system for navigating the primary vehicle along the road segment relative to the at least one drivable path of the primary vehicle stored in the map.
90. The method of claim 89, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for a lane in which the primary vehicle is currently traveling.
91. The method of claim 89, wherein the at least one drivable path for a primary vehicle comprises one or more drivable paths generated for lanes other than the lane in which the primary vehicle is currently traveling.
92. The method of claim 89, wherein the at least one drivable path for a primary vehicle comprises a drivable path generated for each lane at least partially represented in the captured image.