Attention-based anomaly detection
By acquiring images through cameras and generating embedded spatial representations, autonomous vehicles can determine whether an image exceeds a predetermined area, thus solving the challenges of information processing and map updates in autonomous vehicle navigation and improving navigation accuracy and efficiency.
Patent Information
- Application Number
- CN202510769859.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-13
- Filing Date
- 2025-06-10
- Publication Date
- 2025-12-23
AI Technical Summary
Autonomous vehicles need to process and interpret a large amount of visual information and sensor data during navigation. Traditional map-making technologies require a large amount of data, which leads to storage and update challenges and affects navigation efficiency.
The system uses a camera to acquire images, generates an embedding space representation of the images, determines whether a portion of the image is outside a predetermined embedding space region, determines navigation actions based on the determination result, and implements the navigation actions through a processor.
It improves the navigation accuracy and efficiency of autonomous vehicles in complex environments, reduces the data processing burden, and simplifies the map update process.
Smart Images

Figure CN121190806A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally pertains to autonomous vehicle navigation. Background Technology
[0002] With continuous technological advancements, the goal of fully autonomous vehicles capable of navigating roads is fast approaching. Autonomous vehicles may need to consider a multitude of factors and make appropriate decisions based on those factors to safely and accurately reach their intended destination. For example, autonomous vehicles may need to process and interpret visual information (e.g., information captured from cameras) and may also use information from other sources (e.g., from GPS devices, speed sensors, accelerometers, suspension sensors, etc.). Simultaneously, to navigate to their destination, autonomous vehicles may also need to identify their position within a specific road (e.g., a specific lane in a multi-lane road), navigate alongside other vehicles, avoid obstacles and pedestrians, obey traffic signals and signs, and navigate from one road to another at appropriate intersections or junctions. Utilizing and interpreting the vast amounts of information collected by the autonomous vehicle as it travels to its destination presents numerous design challenges. The sheer volume of data that autonomous vehicles may need to analyze, access, and / or store (e.g., captured image data, map data, GPS data, sensor data, etc.) presents challenges that can actually limit or even adversely affect autonomous navigation. Furthermore, if autonomous vehicles rely on traditional mapping techniques for navigation, the sheer volume of data required to store and update maps presents a formidable challenge. Summary of the Invention
[0003] A system for navigating a host vehicle relative to a road segment includes: at least one processor, the at least one processor comprising a circuit system and a memory, wherein the memory includes instructions that, when executed by the circuit system, cause the at least one processor to: receive captured images obtained by a camera on the host vehicle; generate a representation in an embedding space of at least a portion of the captured images; determine whether the representation in the embedding space of the at least a portion of the captured images falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region; determine a navigation action of the host vehicle based on the determination that the representation in the embedding space of the at least a portion of the captured images falls outside the predetermined embedding space region; and cause at least one system associated with the host vehicle to perform the navigation action.
[0004] A method for navigating a master vehicle relative to a road segment includes: receiving captured images acquired by a camera on the master vehicle; generating a representation in an embedding space of at least a portion of the captured images; determining whether the representation in the embedding space of the at least a portion of the captured images falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region; determining a navigation action of the master vehicle based on the determination that the representation in the embedding space of the at least a portion of the captured images falls outside the predetermined embedding space region; and causing at least one system associated with the master vehicle to perform the navigation action.
[0005] Consistent with other disclosed embodiments, a non-transitory computer-readable storage medium may store program instructions that can be executed by at least one processing device to perform any of the steps and / or methods described herein. Attached Figure Description
[0006] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments. In the drawings:
[0007] Figure 1 This is a schematic representation of an exemplary system consistent with the disclosed embodiments.
[0008] Figure 2A A schematic side view representation of an exemplary vehicle that includes a system consistent with the disclosed embodiments.
[0009] Figure 2B To be consistent with the disclosed embodiments Figure 2A The diagram shows a schematic top-view representation of the vehicles and systems.
[0010] Figure 2C This is a schematic top-view representation of another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0011] Figure 2D This is a schematic top-view representation of yet another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0012] Figure 2E This is a schematic top-view representation of yet another embodiment of a vehicle that includes a system consistent with the disclosed embodiments.
[0013] Figure 2F This is a schematic representation of an exemplary vehicle control system consistent with the disclosed embodiments.
[0014] Figure 3A A schematic representation of the interior of a vehicle, including a rearview mirror and user interface for a vehicle imaging system, consistent with the disclosed embodiments.
[0015] Figure 3B An illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0016] Figure 3C To be viewed from a different perspective in accordance with the disclosed embodiments. Figure 3B The diagram shows the camera mounting components.
[0017] Figure 3D An illustration of an example of a camera mount configured to be positioned behind a rearview mirror and against a vehicle windshield, consistent with the disclosed embodiments.
[0018] Figure 4 An exemplary block diagram of a memory configured to store instructions for performing one or more operations, consistent with the disclosed embodiments.
[0019] Figure 5A A flowchart is provided to illustrate an exemplary process for evoking one or more navigation responses based on monocular image analysis, consistent with the disclosed embodiments.
[0020] Figure 5B A flowchart illustrating an exemplary process for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments.
[0021] Figure 5C A flowchart illustrating an exemplary process for detecting road markings and / or lane geometry information in a set of images, consistent with the disclosed embodiments.
[0022] Figure 5D A flowchart illustrating an exemplary process for detecting traffic lights in a set of images, consistent with the disclosed embodiments.
[0023] Figure 5E A flowchart is provided to illustrate an exemplary process for evoking one or more navigation responses based on a vehicle path, consistent with the disclosed embodiments.
[0024] Figure 5F A flowchart illustrating an exemplary process for determining whether a vehicle ahead is changing lanes, consistent with the disclosed embodiments.
[0025] Figure 6 A flowchart illustrating an exemplary process for evoking one or more navigation responses based on stereo image analysis, consistent with the disclosed embodiments.
[0026] Figure 7A flowchart illustrating an exemplary process consistent with the disclosed embodiments for evoking one or more navigation responses based on the analysis of three sets of images.
[0027] Figure 8 A schematic representation of a trained model consistent with the disclosed embodiments is provided.
[0028] Figure 9 A schematic representation of a trained model consistent with the disclosed embodiments is provided.
[0029] Figure 10 A schematic representation of a trained model consistent with the disclosed embodiments is provided.
[0030] Figure 11 A representation of exceptional scenarios consistent with the disclosed embodiments is provided.
[0031] Figure 12 A representation of exceptional scenarios consistent with the disclosed embodiments is provided.
[0032] Figure 13 A representation of exceptional scenarios consistent with the disclosed embodiments is provided.
[0033] Figure 14 A representation of exceptional scenarios consistent with the disclosed embodiments is provided.
[0034] Figure 15 A flowchart illustrating an exemplary method for navigating a master vehicle, consistent with the disclosed embodiments. Detailed Implementation
[0035] The following detailed description refers to the accompanying drawings. Where possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other embodiments are possible. For example, components shown in the drawings may be replaced, added, or modified, and the exemplary methods described herein may be modified by replacing, reordering, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Rather, the appropriate scope is defined by the appended claims.
[0036] Overview of Autonomous Vehicles (AVs)
[0037] As used throughout this disclosure, the term "autonomous vehicle" means a vehicle capable of implementing at least one navigational change without driver input. "Navigational change" refers to a change in one or more of the vehicle's steering, braking, or acceleration. For autonomy to be achieved, the vehicle does not need to be fully automated (e.g., fully operational without a driver or driver input). Rather, autonomous vehicles include those that can operate under driver control during certain time periods and without driver control during other time periods. Autonomous vehicles may also include those that control only some aspects of vehicle navigation (such as steering (e.g., to maintain vehicle alignment between lane constraints)) but leave other aspects (e.g., braking) to the driver. In some cases, an autonomous vehicle may handle some or all aspects of the vehicle's braking, speed control, and / or steering.
[0038] Since human drivers typically rely on visual cues and observation to control vehicles, traffic infrastructure is correspondingly established, with lane markings, traffic signs, and traffic lights all designed to provide visual information to the driver. Given these design characteristics of traffic infrastructure, autonomous vehicles can include cameras and processing units that analyze visual information captured from the vehicle's environment. Visual information can include, for example, components of traffic infrastructure that can be observed by the driver (e.g., lane markings, traffic signs, traffic lights, etc.) and other obstacles (e.g., other vehicles, pedestrians, debris, etc.). Additionally, autonomous vehicles can also use stored information, such as information providing a model of the vehicle's environment during navigation. For example, a vehicle can use GPS data, sensor data (e.g., from accelerometers, speed sensors, suspension sensors, etc.), and / or other map data to provide information related to its environment while driving, and the vehicle (and other vehicles) can use said information to locate itself on the model.
[0039] In some embodiments of this disclosure, the autonomous vehicle may use information obtained during navigation (e.g., from cameras, GPS devices, accelerometers, speed sensors, suspension sensors, etc.). In other embodiments, the autonomous vehicle may use information obtained from past navigation by the vehicle (or other vehicles) during navigation. In still other embodiments, the autonomous vehicle may use a combination of information obtained during navigation and information obtained from past navigation. The following sections provide an overview of systems consistent with the disclosed embodiments, followed by an overview of forward imaging systems and methods consistent with the systems. The following sections disclose systems and methods for constructing, using, and updating sparse maps for autonomous vehicle navigation.
[0040] System Overview
[0041] Figure 1This is a block diagram representation of a system 100 consistent with the exemplary embodiments disclosed. System 100 may include various components depending on the requirements of a particular implementation. In some embodiments, system 100 may include a processing unit 110, an image acquisition unit 120, a position sensor 130, one or more memory units 140, 150, a map database 160, a user interface 170, and a wireless transceiver 172. Processing unit 110 may include one or more processing means. In some embodiments, processing unit 110 may include an application processor 180, an image processor 190, or any other suitable processing means. Similarly, depending on the requirements of a particular application, image acquisition unit 120 may include any number of image acquisition means and components. In some embodiments, image acquisition unit 120 may include one or more image capture means (e.g., cameras), such as image capture means 122, image capture means 124, and image capture means 126. System 100 may also include a data interface 128 that communicatively connects processing unit 110 to image acquisition unit 120. For example, data interface 128 may include any one or more wired and / or wireless links for transmitting image data acquired by image acquisition unit 120 to processing unit 110.
[0042] Wireless transceiver 172 may include one or more devices configured to exchange transmissions with one or more networks (e.g., cellular networks, the Internet, etc.) via an air interface using radio frequency, infrared frequency, magnetic field, or electric field. Wireless transceiver 172 may use any known standard to transmit and / or receive data (e.g., Wi-Fi, etc.). (Bluetooth Smart, 802.15.4, ZigBee, etc.). Such transmissions may include communication from a master vehicle to one or more remotely located servers. Such transmissions may also include (one-way or two-way) communication between the master vehicle and one or more target vehicles in the master vehicle's environment (e.g., to facilitate coordination of navigation of the master vehicle in view of or with the target vehicles in the master vehicle's environment), or even broadcast transmissions to unspecified receivers in the vicinity of the transmitting vehicle.
[0043] Both application processor 180 and image processor 190 can include various types of processing devices. For example, either or both of application processor 180 and image processor 190 can include a microprocessor, a preprocessor (such as an image preprocessor), a graphics processing unit (GPU), a central processing unit (CPU), support circuitry, a digital signal processor, an integrated circuit, memory, or any other type of device suitable for running applications and suitable for image processing and analysis. In some embodiments, application processor 180 and / or image processor 190 can include any type of single-core or multi-core processor, mobile device microcontroller, central processing unit, etc. Various processing devices can be used, including, for example, those available from manufacturers such as... Processors obtained from manufacturers such as GPUs can be obtained, and can include various architectures (e.g., x86 processors, etc.). wait).
[0044] In some embodiments, application processor 180 and / or image processor 190 may include components that can be accessed from... Any processor chip from the EyeQ series of processor chips obtained. These processor designs each include multiple processing units with local memory and instruction sets. Such processors may include video input for receiving image data from multiple image sensors, and may also include video output capabilities. In one example, It uses 90nm micrometer technology operating at 332MHz. The architecture consists of two floating-point hyper-threaded 32-bit RISC CPUs ( (core), five visual computing engines (VCE), and three vector microcode processors The Denali MIPS34K CPU comprises a 64-bit mobile DDR controller, a 128-bit internal Sonics interconnect, dual 16-bit video input and 18-bit video output controllers, a 16-channel DMA, and several peripherals. The CPU manages five VCEs and three VMPs. TM And DMA, a second MIPS34K CPU and multi-channel DMA, and other peripherals. Five VCEs, three The MIPS34K CPU can perform the intensive vision computations required by versatile bundled applications. In another example, (It is a third-generation processor and its performance is) (Six times) can be used in the disclosed embodiments. In other instances, and / or This can be used in the disclosed embodiments. Of course, any newer or future EyeQ processing device can also be used with the disclosed embodiments.
[0045] Any of the processing devices disclosed herein can be configured to perform certain functions. Configuring a processing device (such as the described EyeQ processor or any other controller or microprocessor) to perform certain functions may include programming computer-executable instructions and making those instructions available for execution by the processing device during operation of the processing device. In some embodiments, configuring the processing device may include programming the processing device directly using architectural instructions. For example, processing devices such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may be configured using, for example, one or more hardware description languages (HDLs).
[0046] In other embodiments, configuring the processing means may include storing executable instructions on memory accessible to the processing means during operation. For example, the processing means may access the memory during operation to obtain and execute the stored instructions. In any case, a processing means configured to perform the sensing, image analysis, and / or navigation functions disclosed herein represents a specialized hardware-based system that controls multiple hardware-based components of a master vehicle.
[0047] Although Figure 1 Two separate processing devices included in processing unit 110 are depicted, but more or fewer processing devices may be used. For example, in some embodiments, a single processing device may be used to perform the tasks of application processor 180 and image processor 190. In other embodiments, these tasks may be performed by more than two processing devices. Furthermore, in some embodiments, system 100 may include one or more processing units 110 without other components such as image acquisition unit 120.
[0048] Processing unit 110 may include various types of devices. For example, processing unit 110 may include various devices such as controllers, image preprocessors, central processing units (CPUs), graphics processing units (GPUs), support circuitry, digital signal processors, integrated circuits, memory, or any other type of device for image processing and analysis. An image preprocessor may include a video processor for capturing, digitizing, and processing images from an image sensor. A CPU may include any number of microcontrollers or microprocessors. A GPU may also include any number of microcontrollers or microprocessors. Support circuitry may be any number of circuits well known in the art, including caches, power supplies, clocks, and input / output circuits. Memory may store software that controls the operation of the system when executed by the processor. Memory may include databases and image processing software. Memory may include any number of random access memories, read-only memories, flash memory, disk drives, optical storage devices, magnetic tape storage devices, removable storage devices, and other types of storage devices. In one case, the memory may be separate from processing unit 110. In another case, the memory may be integrated into processing unit 110.
[0049] Each memory unit 140, 150 may include software instructions that, when executed by a processor (e.g., application processor 180 and / or image processor 190), can control various aspects of the operation of system 100. For example, these memory units may include various database and image processing software, as well as trained systems such as neural networks or deep neural networks. Memory units may include random access memory (RAM), read-only memory (ROM), flash memory, disk drives, optical storage devices, magnetic tape storage devices, removable storage devices, and / or any other type of storage device. In some embodiments, memory units 140, 150 may be decoupled from application processor 180 and / or image processor 190. In other embodiments, these memory units may be integrated into application processor 180 and / or image processor 190.
[0050] The location sensor 130 may include any type of device suitable for determining positioning associated with at least one component of the system 100. In some embodiments, the location sensor 130 may include a GPS receiver. Such a receiver can determine the user's location and rate by processing signals broadcast by Global Positioning System satellites. Location information from the location sensor 130 may be used by the application processor 180 and / or the image processor 190.
[0051] In some embodiments, system 100 may include components such as a speed sensor (e.g., a tachometer, speedometer) for measuring the speed of vehicle 200 and / or an accelerometer (single-axis or multi-axis) for measuring the acceleration of vehicle 200.
[0052] User interface 170 may include any means adapted to provide information to or receive input from one or more users of system 100. In some embodiments, user interface 170 may include user input devices, including, for example, a touchscreen, microphone, keyboard, pointer device, scroll wheel, camera, knob, button, etc. Using such input devices, users may be able to provide information input or commands to system 100 by typing instructions or information, providing voice commands, selecting menu options on the screen (using buttons, pointers, or eye-tracking capabilities), or by any other suitable technology for transmitting information to system 100.
[0053] User interface 170 may be equipped with one or more processing devices configured to provide and receive information from the user, and to process the information for use by, for example, application processor 180. In some embodiments, such processing devices may execute instructions to recognize and track eye movements, receive and interpret voice commands, recognize and interpret touches and / or gestures made on a touchscreen, respond to keyboard input or menu selection, etc. In some embodiments, user interface 170 may include a display, speakers, haptic devices, and / or any other means for providing output information to the user.
[0054] Map database 160 may include any type of database for storing map data useful to system 100. In some embodiments, map database 160 may include data relating to the location of various items (including roads, water features, geographic features, businesses, points of interest, restaurants, gas stations) in a reference coordinate system. Map database 160 may not only store the locations of such items but also descriptors associated with these items, including, for example, names associated with any of the stored features. In some embodiments, map database 160 may be physically located together with other components of system 100. Alternatively or additionally, map database 160 or a portion thereof may be remotely located relative to other components of system 100 (e.g., processing unit 110). In such embodiments, information from map database 160 may be downloaded via a wired or wireless data connection to a network (e.g., via cellular networks and / or the Internet, etc.). In some cases, map database 160 may store a sparse data model including a polynomial representation of certain road features (e.g., lane markings) or target trajectories for the primary vehicle. Systems and methods for generating such maps are referenced below. Figures 8-15Let's have a discussion.
[0055] Image capture devices 122, 124, and 126 may each include any type of means suitable for capturing at least one image from the environment. Furthermore, any number of image capture devices can be used to acquire images for input to an image processor. Some embodiments may include only a single image capture device, while other embodiments may include two, three, or even four or more image capture devices. Image capture devices 122, 124, and 126 will be referenced below. Figure 2B-2E To be further described.
[0056] System 100 or its various components can be incorporated into a variety of different platforms. In some embodiments, system 100 may be included on vehicle 200, such as... Figure 2A As shown. For example, vehicle 200 may be equipped with processing unit 110 and any other components of system 100, as described above relative to... Figure 1 As described. While in some embodiments, vehicle 200 may be equipped with only a single image capture device (e.g., a camera), in other embodiments, such as combined with... Figure 2B-2E The ones discussed can use multiple image capture devices. For example, such as Figure 2A As shown, either of the image capture devices 122 and 124 of the vehicle 200 can be part of an ADAS (Advanced Driver Assistance System) imaging set.
[0057] The image capture device, included on the vehicle 200 as part of the image acquisition unit 120, can be positioned at any suitable location. In some embodiments, such as Figures 2A-2E As shown in 3A-3C, the image capturing device 122 can be positioned in the vicinity of the rearview mirror. This position provides a line of sight similar to that of the driver of vehicle 200, which can help determine what is visible and invisible to the driver. The image capturing device 122 can be positioned anywhere near the rearview mirror, but placing the image capturing device 122 on the driver's side of the mirror can further assist in obtaining an image representing the driver's field of view and / or line of sight.
[0058] Other positioning of the image capture device for image acquisition unit 120 may also be used. For example, image capture device 124 may be positioned on or therein on the bumper of vehicle 200. Such positioning may be particularly suitable for image capture devices with a wide field of view. The line of sight of the image capture device positioned on the bumper may be different from that of the driver, and therefore the bumper image capture device and the driver may not always see the same object. Image capture devices (e.g., image capture devices 122, 124, and 126) may also be positioned in other locations. For example, the image capture device may be positioned on or therein on one or both of the side mirrors of vehicle 200, on the roof of vehicle 200, on the hood of vehicle 200, on the trunk of vehicle 200, on the side of vehicle 200, mounted on any of the windows of vehicle 200, positioned behind or in front of them, and mounted in or near the headlight graphics at the front and / or rear of vehicle 200, etc.
[0059] In addition to the image capture device, vehicle 200 may also include various other components of system 100. For example, processing unit 110 may be included on vehicle 200, or integrated with or separate from the vehicle's engine control unit (ECU). Vehicle 200 may also be equipped with position sensor 130 such as a GPS receiver, and may also include map database 160 and memory units 140 and 150.
[0060] As discussed above, the wireless transceiver 172 can receive data via one or more networks (e.g., cellular networks, the Internet, etc.) and / or receive data. For example, the wireless transceiver 172 can upload data collected by the system 100 to one or more servers and download data from said one or more servers. Via the wireless transceiver 172, the system 100 can receive, for example, periodic or on-demand updates to data stored in map database 160, memory 140, and / or memory 150. Similarly, the wireless transceiver 172 can upload any data from the system 100 (e.g., images captured by image acquisition unit 120, data received by position sensor 130 or other sensors, vehicle control system, etc.) and / or any data processed by processing unit 110 to one or more servers.
[0061] System 100 may upload data to a server (e.g., to the cloud) based on privacy level settings. For example, System 100 may implement privacy level settings to regulate or restrict the data types (including metadata) sent to a server that can uniquely identify the vehicle and / or the vehicle's driver / owner. Such settings may be configured by the user via, for example, wireless transceiver 172, through factory default settings, or through data received by wireless transceiver 172.
[0062] In some embodiments, system 100 may upload data according to a “high” privacy level, and under certain settings, system 100 may transmit data (e.g., location information related to routes, captured images, etc.) without any details about the specific vehicle and / or driver / owner. For example, when uploading data according to a “high” privacy setting, system 100 may not include the vehicle identification number (VIN) or the name of the vehicle driver or owner, and may instead transmit data such as captured images and / or restricted location information related to routes.
[0063] Other privacy levels are envisioned. For example, system 100 may transmit data to the server at a “medium” privacy level, including additional information not included at a “high” privacy level, such as the vehicle’s brand and / or model and / or vehicle type (e.g., passenger vehicle, SUV, truck, etc.). In some embodiments, system 100 may upload data at a “low” privacy level. At a “low” privacy level setting, system 100 may upload data and include information sufficient to uniquely identify a specific vehicle, owner / driver, and / or part or all of the route traveled by the vehicle. Such “low” privacy level data may include, for example, one or more of the following: VIN, driver / owner name, vehicle’s origin point before departure, vehicle’s intended destination, vehicle’s brand and / or model, vehicle type, etc.
[0064] Figure 2A A schematic side view representation of an exemplary vehicle imaging system consistent with the disclosed embodiments. Figure 2B for Figure 2A The illustrated embodiment is shown in a schematic top view. Figure 2B The illustrated and disclosed embodiments may include a vehicle 200 that includes a system 100 in its body, the system having a first image capture device 122 located in the vicinity of a rearview mirror and / or near the driver of the vehicle 200, a second image capture device 124 located on or therein in a bumper area of the vehicle 200 (e.g., one of the bumper areas 210), and a processing unit 110.
[0065] like Figure 2C As shown, both image capture devices 122 and 124 can be positioned in the vicinity of the rearview mirror and / or near the driver of vehicle 200. Additionally, although the two image capture devices 122 and 124 are shown in... Figure 2B and 2C However, it should be understood that other embodiments may include more than two image capture devices. For example, in Figure 2D and 2EIn the embodiment shown, the system 100 of the vehicle 200 includes a first image capture device 122, a second image capture device 124, and a third image capture device 126.
[0066] like Figure 2D As shown, image capture device 122 can be positioned in the vicinity of the rearview mirror and / or near the driver of vehicle 200, and image capture devices 124 and 126 can be positioned on or within the bumper area of vehicle 200 (e.g., one of bumper areas 210). And as... Figure 2E As shown, image capturing devices 122, 124, and 126 can be positioned in the vicinity of the rearview mirror and / or near the driver's seat of vehicle 200. The disclosed embodiments are not limited to any particular number and configuration of image capturing devices, and the image capturing devices can be positioned within vehicle 200 and / or in any suitable location on said vehicle.
[0067] It should be understood that the disclosed embodiments are not limited to vehicles and can be applied to other situations. It should also be understood that the disclosed embodiments are not limited to a specific type of vehicle 200 and can be applied to all types of vehicles, including cars, trucks, trailers and other types of vehicles.
[0068] The first image capture device 122 may include any suitable type of image capture device. The image capture device 122 may include an optical axis. In one embodiment, the image capture device 122 may include an Aptina M9V024 WVGA sensor with a global shutter. In other embodiments, the image capture device 122 may provide a resolution of 1280x960 pixels and may include a rolling shutter. The image capture device 122 may include various optical elements. In some embodiments, one or more lenses may be included, for example, to provide a desired focal length and field of view for the image capture device. In some embodiments, the image capture device 122 may be associated with a 6mm lens or a 12mm lens. In some embodiments, the image capture device 122 may be configured to capture an image with a desired field of view (FOV) 202, such as... Figure 2DAs shown. For example, image capture device 122 can be configured to have a regular FOV, such as 46 degrees, 50 degrees, 52 degrees, or greater, within the range of 40 to 56 degrees. Alternatively, image capture device 122 can be configured to have a narrow FOV, such as 28 degrees or 36 degrees, within the range of 23 to 40 degrees. Furthermore, image capture device 122 can be configured to have a wide FOV within the range of 100 to 180 degrees. In some embodiments, image capture device 122 may include a wide-angle bumper camera or a camera with a maximum FOV of 180 degrees. In some embodiments, image capture device 122 may be a 7.2M pixel image capture device with an aspect ratio of approximately 2:1 (e.g., HxV = 3800x1900 pixels) and a horizontal FOV of approximately 100 degrees. Such image capture devices can be used to replace a three-image capture device configuration. Due to significant lens distortion, in embodiments where the image capture device uses radially symmetrical lenses, the vertical FOV of such image capture devices may be significantly less than 50 degrees. For example, such lenses may not be radially symmetrical, which would allow a vertical FOV greater than 50 degrees and a horizontal FOV of 100 degrees.
[0069] The first image capturing device 122 can acquire a plurality of first images relative to a scene associated with the vehicle 200. Each of the plurality of first images can be acquired as a series of image scan lines that can be captured using a rolling shutter. Each scan line may include a plurality of pixels.
[0070] The first image capture device 122 may have a scan rate associated with the acquisition of each of the first series of image scan lines. The scan rate may refer to the rate at which the image sensor can acquire image data associated with each pixel included in a particular scan line.
[0071] For example, image capture devices 122, 124, and 126 may contain any suitable type and number of image sensors, including CCD sensors or CMOS sensors. In one embodiment, a CMOS image sensor may be used in conjunction with a rolling shutter, such that each pixel in a row is read one at a time, and the scanning of rows is performed on a line-by-line basis until the entire image frame has been captured. In some embodiments, rows may be captured sequentially from top to bottom relative to the frame.
[0072] In some embodiments, one or more of the image capture devices disclosed herein (e.g., image capture devices 122, 124, and 126) may constitute a high-resolution imager and may have a resolution greater than 5M pixels, 7M pixels, 10M pixels, or more pixels.
[0073] The use of a rolling shutter can cause pixels in different rows to be exposed and captured at different times, which can lead to skew and other image artifacts in the captured image frame. On the other hand, when the image capture device 122 is configured to operate using a global or synchronous shutter, all pixels can be exposed for the same amount of time and during a common exposure period. Therefore, the image data in a frame collected from a system using a global shutter represents a snapshot of the entire FOV (e.g., FOV 202) at a specific time. In contrast, in a rolling shutter application, each row in the frame is exposed, and the data is captured at different times. Therefore, moving objects may appear distorted in an image capture device with a rolling shutter. This phenomenon will be described in more detail below.
[0074] The second image capture device 124 and the third image capture device 126 can be any type of image capture device. Similar to the first image capture device 122, each of the image capture devices 124 and 126 may include an optical axis. In one embodiment, each of the image capture devices 124 and 126 may include an Aptina M9V024WVGA sensor with a global shutter. Alternatively, each of the image capture devices 124 and 126 may include a rolling shutter. Similar to image capture device 122, image capture devices 124 and 126 may be configured to include various lenses and optical elements. In some embodiments, the lenses associated with image capture devices 124 and 126 may provide an FOV (such as FOV 202) that is the same as or narrower than that associated with image capture device 122 (such as FOV 204 and 206). For example, image capture devices 124 and 126 may have an FOV of 40 degrees, 30 degrees, 26 degrees, 23 degrees, 20 degrees, or less.
[0075] Image capture devices 124 and 126 can acquire multiple second and third images of a scene associated with vehicle 200. Each of the multiple second and third images can be acquired as a second series and a third series of image scan lines that can be captured using a rolling shutter. Each scan line or row can have multiple pixels. Image capture devices 124 and 126 can have a second scan rate and a third scan rate associated with the acquisition of each of the image scan lines included in the second and third series.
[0076] Each image capture device 122, 124, and 126 can be positioned relative to the vehicle 200 at any suitable location and orientation. The relative positioning of the image capture devices 122, 124, and 126 can be selected to facilitate the fusion of information acquired from the image capture devices. For example, in some embodiments, the field of view (FOV) associated with image capture device 124 (e.g., FOV 204) can partially or completely overlap with the FOV associated with image capture device 122 (e.g., FOV 202) and the FOV associated with image capture device 126 (e.g., FOV 206).
[0077] Image capture devices 122, 124, and 126 can be positioned at any suitable relative height on the vehicle 200. In one case, there may be a height difference between image capture devices 122, 124, and 126, which can provide sufficient parallax information for stereoscopic analysis. For example, as Figure 2A As shown, the two image capture devices 122 and 124 are at different heights. For example, there may also be lateral displacement differences between image capture devices 122, 124, and 126, providing additional parallax information for stereo analysis by the processing unit 110. The difference in lateral displacement can be represented by dx, such as... Figure 2C and 2D As shown. In some embodiments, there may be forward or backward displacement (e.g., range displacement) between image capture devices 122, 124, and 126. For example, image capture device 122 may be positioned 0.5 meters to 2 meters or more behind image capture devices 124 and / or 126. This type of displacement allows one of the image capture devices to cover potential blind spots of the other image capture devices.
[0078] Image capture device 122 may have any suitable resolution capability (e.g., the number of pixels associated with an image sensor), and the resolution of the image sensor associated with image capture device 122 may be higher, lower, or the same as the resolution of the image sensors associated with image capture devices 124 and 126. In some embodiments, the image sensors associated with image capture device 122 and / or image capture devices 124 and 126 may have a resolution of 640x480, 1024x768, 1280x960, or any other suitable resolution.
[0079] The frame rate (e.g., the rate at which an image capture device acquires a set of pixel data for an image frame before continuing to capture pixel data associated with the next image frame) can be controllable. The frame rate associated with image capture device 122 can be higher, lower, or the same as the frame rates associated with image capture devices 124 and 126. The frame rates associated with image capture devices 122, 124, and 126 can depend on a variety of factors that may affect the timing of the frame rate. For example, one or more of image capture devices 122, 124, and 126 may include a selectable pixel delay period applied before or after the acquisition of image data associated with one or more pixels of the image sensors in image capture devices 122, 124, and / or 126. Generally, image data corresponding to each pixel can be acquired according to the clock rate for the device (e.g., one pixel per clock cycle). Additionally, in embodiments including a rolling shutter, one or more of the image capture devices 122, 124, and 126 may include a selectable horizontal blanking period applied before or after the acquisition of image data associated with pixel rows of the image sensors in the image capture devices 122, 124, and / or 126. Furthermore, one or more of the image capture devices 122, 124, and / or 126 may include a selectable vertical blanking period applied before or after the acquisition of image data associated with image frames of the image capture devices 122, 124, and 126.
[0080] These timing controls enable synchronization of the frame rates associated with image capture devices 122, 124, and 126, even when the line scan rates of each image capture device differ. Furthermore, as will be discussed in more detail below, these selectable timing controls and other factors (e.g., image sensor resolution, maximum line scan rate, etc.) can enable synchronization of image capture from areas where the field of view (FOV) of image capture device 122 overlaps with one or more FOVs of image capture devices 124 and 126, even when the field of view of image capture device 122 differs from the FOV of image capture devices 124 and 126.
[0081] The frame rate timing in the image capture devices 122, 124, and 126 can depend on the resolution of the associated image sensor. For example, assuming similar line scan rates for two devices, if one device includes an image sensor with a resolution of 640x480 and the other device includes an image sensor with a resolution of 1280x960, it will take more time to acquire one frame of image data from the sensor with the higher resolution.
[0082] Another factor that can affect the timing of image data acquisition in image capture devices 122, 124, and 126 is the maximum line scan rate. For example, acquiring one line of image data from an image sensor included in image capture devices 122, 124, and 126 will require a minimum amount of time. Assuming no pixel delay period is added, this minimum amount of time for acquiring one line of image data will be related to the maximum line scan rate of the particular device. A device that provides a higher maximum line scan rate is likely to provide a higher frame rate than a device with a lower maximum line scan rate. In some embodiments, one or more of image capture devices 124 and 126 may have a higher maximum line scan rate than the maximum line scan rate associated with image capture device 122. In some embodiments, the maximum line scan rate of image capture devices 124 and / or 126 may be 1.25 times, 1.5 times, 1.75 times, or 2 times or more of the maximum line scan rate of image capture device 122.
[0083] In another embodiment, image capture devices 122, 124, and 126 may have the same maximum line scan rate, but image capture device 122 may be operated at a scan rate less than or equal to its maximum scan rate. The system may be configured such that one or more of image capture devices 124 and 126 operate at a line scan rate equal to the line scan rate of image capture device 122. In other cases, the system may be configured such that the line scan rate of image capture device 124 and / or image capture device 126 can be 1.25 times, 1.5 times, 1.75 times, or 2 times or more of the line scan rate of image capture device 122.
[0084] In some embodiments, image capture devices 122, 124, and 126 may be asymmetrical. That is, they may include cameras with different fields of view (FOV) and focal lengths. For example, the fields of view of image capture devices 122, 124, and 126 may include any desired area relative to the environment of vehicle 200. In some embodiments, one or more of image capture devices 122, 124, and 126 may be configured to acquire image data from the environment of the front of vehicle 200, the rear of vehicle 200, the sides of vehicle 200, or a combination thereof.
[0085] Furthermore, the focal length associated with each image capturing device 122, 124, and / or 126 can be selectable (e.g., by including suitable lenses, etc.) so that each device acquires an image of an object relative to the vehicle 200 at a desired distance range. For example, in some embodiments, image capturing devices 122, 124, and 126 can acquire images of close-up objects within a few meters of the vehicle. Image capturing devices 122, 124, and 126 can also be configured to acquire images of objects at a greater distance from the vehicle (e.g., 25m, 50m, 100m, 150m, or more). Furthermore, the focal lengths of image capturing devices 122, 124, and 126 can be selected such that one image capturing device (e.g., image capturing device 122) can acquire images of objects relatively close to the vehicle (e.g., within 10m or 20m), while other image capturing devices (e.g., image capturing devices 124 and 126) can acquire images of objects further away from the vehicle (e.g., greater than 20m, 50m, 100m, 150m, etc.).
[0086] According to some embodiments, the field of view (FOV) of one or more image capture devices 122, 124, and 126 may have a wide angle. For example, an FOV of 140 degrees may be advantageous, particularly for image capture devices 122, 124, and 126 that can be used to capture images of areas adjacent to the vehicle 200. For example, image capture device 122 may be used to capture images of areas to the right or left of the vehicle 200, and in such embodiments, it may be desirable for image capture device 122 to have a wide FOV (e.g., at least 140 degrees).
[0087] The field of view associated with each of the image capturing devices 122, 124, and 126 can depend on the corresponding focal length. For example, as the focal length increases, the corresponding field of view decreases.
[0088] Image capture devices 122, 124, and 126 can be configured to have any suitable field of view. In one particular example, image capture device 122 may have a horizontal FOV of 46 degrees, image capture device 124 may have a horizontal FOV of 23 degrees, and image capture device 126 may have a horizontal FOV between 23 degrees and 46 degrees. In another case, image capture device 122 may have a horizontal FOV of 52 degrees, image capture device 124 may have a horizontal FOV of 26 degrees, and image capture device 126 may have a horizontal FOV between 26 degrees and 52 degrees. In some embodiments, the ratio of the FOV of image capture device 122 to the FOV of image capture device 124 and / or image capture device 126 may vary between 1.5 and 2.0. In other embodiments, this ratio may vary between 1.25 and 2.25.
[0089] System 100 can be configured such that the field of view of image capture device 122 at least partially or completely overlaps with the field of view of image capture device 124 and / or image capture device 126. In some embodiments, system 100 can be configured such that the field of view of image capture devices 124 and 126, for example, falls within the field of view of image capture device 122 (e.g., narrower than said field of view) and shares a common center with said field of view. In other embodiments, image capture devices 122, 124, and 126 can capture adjacent FOVs or can have partial overlap in their FOVs. In some embodiments, the field of view of image capture devices 122, 124, and 126 can be aligned such that the center of image capture device 124 and / or 126 with the narrower FOV can be located in the lower half of the field of view of image capture device 122 with the wider FOV.
[0090] Figure 2F This is a schematic representation of an exemplary vehicle control system consistent with the disclosed embodiments. Figure 2F As indicated, vehicle 200 may include a throttle system 220, a braking system 230, and a steering system 240. System 100 may provide inputs (e.g., control signals) to one or more of the throttle system 220, braking system 230, and steering system 240 via one or more data links (e.g., any wired and / or wireless link or link used for data transmission). For example, based on analysis of images acquired by image capture devices 122, 124, and / or 126, system 100 may provide control signals to one or more of the throttle system 220, braking system 230, and steering system 240 to navigate vehicle 200 (e.g., by inducing acceleration, turning, lane changes, etc.). Furthermore, system 100 may receive inputs from one or more of the throttle system 220, braking system 230, and steering system 240 indicative of operating conditions of vehicle 200 (e.g., speed, whether vehicle 200 is braking and / or turning, etc.). Further details are described below in conjunction with... Figure 4-7 supply.
[0091] like Figure 3AAs shown, vehicle 200 may also include a user interface 170 for interaction with the driver or passengers of vehicle 200. For example, the user interface 170 in a vehicle application may include a touchscreen 320, a knob 330, a button 340, and a microphone 350. The driver or passengers of vehicle 200 may also interact with system 100 using handles (e.g., located on or near the steering column of vehicle 200, including, for example, a turn signal handle), buttons (e.g., located on the steering wheel of vehicle 200), etc. In some embodiments, microphone 350 may be located adjacent to rearview mirror 310. Similarly, in some embodiments, image capture device 122 may be located near rearview mirror 310. In some embodiments, user interface 170 may also include one or more speakers 360 (e.g., speakers of a vehicle audio system). For example, system 100 may provide various notifications (e.g., alarms) via speaker 360.
[0092] Figure 3B-3D An illustration of an exemplary camera mount 370 configured to be positioned behind a rearview mirror (e.g., rearview mirror 310) and against the vehicle windshield, consistent with the disclosed embodiments. Figure 3B As shown, the camera mount 370 may include image capture devices 122, 124, and 126. Image capture devices 124 and 126 may be positioned behind a glare shield 380, which may be flush with a vehicle windshield and comprises a composition of a film and / or anti-reflective material. For example, the glare shield 380 may be positioned such that it is aligned against a vehicle windshield having a matching slope. In some embodiments, each of the image capture devices 122, 124, and 126 may be positioned behind the glare shield 380, such as, for example... Figure 3D The embodiments described herein are not limited to any particular configuration of the image capture devices 122, 124 and 126, the camera mount 370 and the glare shield 380. Figure 3C From the front perspective Figure 3B The illustration shows the camera mounting component 370.
[0093] As those skilled in the art who benefit from this disclosure will understand, various variations and / or modifications can be made to the disclosed embodiments. For example, not all components are essential for the operation of system 100. Furthermore, any component can be located in any suitable part of system 100, and these components can be rearranged into various configurations while providing the functionality of the disclosed embodiments. Thus, the foregoing configuration is an example, and regardless of the configuration discussed above, system 100 can provide a wide range of functions to analyze the environment surrounding vehicle 200 and navigate vehicle 200 in response to said analysis.
[0094] As discussed in further detail below and consistent with the various disclosed embodiments, system 100 can provide various features related to autonomous driving and / or driver assistance technologies. For example, system 100 can analyze image data, location data (e.g., GPS positioning information), map data, speed data, and / or data from sensors included in vehicle 200. System 100 can collect data for analysis from, for example, image acquisition unit 120, position sensor 130, and other sensors. Furthermore, system 100 can analyze the collected data to determine whether vehicle 200 should take certain actions, and then automatically take the determined actions without human intervention. For example, when vehicle 200 is navigating without human intervention, system 100 can automatically control the braking, acceleration, and / or steering of vehicle 200 (e.g., by sending control signals to one or more of throttle system 220, braking system 230, and steering system 240). Additionally, system 100 can analyze the collected data and issue warnings and / or alarms to vehicle occupants based on the analysis of the collected data. Further details regarding the various embodiments provided by system 100 are provided below.
[0095] Forward multiple imaging system
[0096] As discussed above, system 100 can provide driver assistance functions using a multi-camera system. The multi-camera system can use one or more cameras facing forward of the vehicle. In other embodiments, the multi-camera system may include one or more cameras facing the side or rear of the vehicle. In one embodiment, for example, system 100 may use a dual-camera imaging system, wherein a first camera and a second camera (e.g., image capture devices 122 and 124) may be positioned at the front and / or side of the vehicle (e.g., vehicle 200). The first camera may have a field of view larger than, smaller than, or partially overlapping with the field of view of the second camera. Furthermore, the first camera may be connected to a first image processor for monocular image analysis of the images provided by the first camera, and the second camera may be connected to a second image processor for monocular image analysis of the images provided by the second camera. The outputs of the first and second image processors (e.g., processed information) may be combined. In some embodiments, the second image processor may receive images from both the first and second cameras for stereo analysis. In another embodiment, system 100 may use a three-camera imaging system, wherein each of the cameras has a different field of view. Therefore, such systems can make decisions based on information derived from objects positioned at varying distances from both the front and sides of the vehicle. A reference to monocular image analysis can refer to cases where image analysis is performed based on images captured from a single viewpoint (e.g., from a single camera). Stereo image analysis can refer to cases where image analysis is performed based on two or more images captured using one or more variations of image capture parameters. For example, images suitable for stereo image analysis could include: images captured from two or more different locations, images captured from different fields of view, images captured using different focal lengths and parallax information, etc.
[0097] For example, in one embodiment, system 100 may use image capture devices 122, 124, and 126 to implement a three-camera configuration. In such a configuration, image capture device 122 may provide a narrow field of view (e.g., 34 degrees or other values selected from the range of about 20 to 45 degrees), image capture device 124 may provide a wide field of view (e.g., 150 degrees or other values selected from the range of about 100 to about 180 degrees), and image capture device 126 may provide an intermediate field of view (e.g., 46 degrees or other values selected from the range of about 35 to about 60 degrees). In some embodiments, image capture device 126 may act as a primary or main camera. Image capture devices 122, 124, and 126 may be positioned behind rearview mirror 310 and substantially side-by-side (e.g., spaced 6 cm apart). Furthermore, in some embodiments, as discussed above, one or more of image capture devices 122, 124, and 126 may be mounted behind a glare shield 380 flush with the windshield of vehicle 200. Such a mask can minimize the impact of any reflections from inside the car on the image capture devices 122, 124 and 126.
[0098] In another embodiment, as described above Figure 3B and 3C The wide field-of-view camera discussed (e.g., image capture device 124 in the examples above) can be mounted below the narrow field-of-view and main field-of-view cameras (e.g., image capture devices 122 and 126 in the examples above). This configuration provides a free line of sight from the wide field-of-view camera. To reduce reflections, the camera can be mounted close to the windshield of the vehicle 200 and may include a polarizer on the camera to reduce reflected light.
[0099] A three-camera system can provide certain performance characteristics. For example, some embodiments may include verifying the ability of one camera to detect an object based on detection results from another camera. In the three-camera configuration discussed above, processing unit 110 may include, for example, three processing devices (e.g., three EyeQ series processor chips, as discussed above), wherein each processing device is dedicated to processing images captured by one or more of the image capture devices 122, 124, and 126.
[0100] In a three-camera system, the first processing unit can receive images from both the main camera and the narrow field-of-view (FOV) camera, and perform visual processing on the narrow FOV camera to detect, for example, other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Furthermore, the first processing unit can calculate the pixel parallax between the images from the main camera and the narrow FOV camera and create a 3D reconstruction of the vehicle 200's environment. The first processing unit can then combine the 3D reconstruction with 3D map data or with 3D information calculated based on information from the other camera.
[0101] The second processing unit can receive images from the main camera and perform visual processing to detect other vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. Additionally, the second processing unit can calculate camera displacement and, based on said displacement, calculate pixel parallax between consecutive images and create a 3D reconstruction of the scene (e.g., from moving structures). The second processing unit can then send the structure from the motion-based 3D reconstruction to the first processing unit for combination with the stereoscopic 3D image.
[0102] The third processing unit can receive images from a wide field of view (FOV) camera and process the images to detect vehicles, pedestrians, lane markings, traffic signs, traffic lights, and other road objects. The third processing unit can further execute additional processing instructions to analyze the images to identify moving objects in the images, such as vehicles changing lanes or pedestrians.
[0103] In some embodiments, enabling image-based information streams to be captured and processed independently can provide opportunities for redundancy in the system. Such redundancy may include, for example, using a first image capture device and images processed from said device to verify and / or supplement information obtained by capturing and processing image information from at least a second image capture device.
[0104] In some embodiments, system 100 may use two image capture devices (e.g., image capture devices 122 and 124) to provide navigation assistance to vehicle 200, and a third image capture device (e.g., image capture device 126) to provide redundancy and verify the analysis of data received from the other two image capture devices. For example, in such a configuration, image capture devices 122 and 124 may provide images for stereo analysis by system 100 for navigating vehicle 200, while image capture device 126 may provide images for monocular analysis by system 100 to provide redundancy and verification based on information obtained from images captured by image capture devices 122 and / or 124. That is, image capture device 126 (and corresponding processing device) may be considered to provide a redundant subsystem for checking the analysis derived from image capture devices 122 and 124 (e.g., to provide an automatic emergency braking (AEB) system). Furthermore, in some embodiments, the redundancy and verification of the received data can be supplemented based on information received from one or more sensors (e.g., radar, lidar, acoustic sensors), information received from one or more transceivers outside the vehicle, etc.
[0105] Those skilled in the art will recognize that the camera configurations, placements, number of cameras, and positioning described above are merely examples. These components, and other components described relative to the overall system, can be assembled and used in a variety of different configurations without departing from the scope of the disclosed embodiments. Further details regarding the use of a multi-camera system to provide driver assistance and / or autonomous vehicle functionality are as follows.
[0106] Figure 4 An exemplary functional block diagram of a memory 140 and / or 150, consistent with the disclosed embodiments, that can store / program instructions for performing one or more operations. Although memory 140 is mentioned below, those skilled in the art will recognize that instructions can be stored in memory 140 and / or 150.
[0107] like Figure 4 As shown, memory 140 may store monocular image analysis module 402, stereo image analysis module 404, rate and acceleration module 406, and navigation response module 408. The disclosed embodiments are not limited to any particular configuration of memory 140. Furthermore, application processor 180 and / or image processor 190 may execute instructions stored in any of the modules 402, 404, 406, and 408 included in memory 140. Those skilled in the art will understand that references to processing unit 110 in the following discussion may refer individually or collectively to application processor 180 and image processor 190. Therefore, steps in any of the following processes may be performed by one or more processing devices.
[0108] In one embodiment, the monocular image analysis module 402 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform monocular image analysis on a set of images acquired by one of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from a set of images with additional sensor information (e.g., information from radar, lidar, etc.) to perform monocular image analysis. (The following is in conjunction with...) Figures 5A-5D As described, the monocular image analysis module 402 may include instructions for detecting a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, and any other features associated with the vehicle's environment. Based on this analysis, system 100 (e.g., via processing unit 110) may induce one or more navigation responses in vehicle 200, such as turning, lane changing, or changes in acceleration, as discussed below in conjunction with navigation response module 408.
[0109] In one embodiment, the stereo image analysis module 404 may store instructions (such as computer vision software) that, when executed by the processing unit 110, perform stereo image analysis on a first set of images and a second set of images acquired by a combination of image capture devices selected from any of the image capture devices 122, 124, and 126. In some embodiments, the processing unit 110 may combine information from the first set of images and the second set of images with additional sensing information (e.g., information from radar) to perform stereo image analysis. For example, the stereo image analysis module 404 may include instructions for performing stereo image analysis based on the first set of images acquired by the image capture device 124 and the second set of images acquired by the image capture device 126. (The following is a continuation of the previous paragraph.) Figure 6 As described, the stereo image analysis module 404 may include instructions for detecting a set of features (such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, hazardous objects, etc.) within a first set of images and a second set of images. Based on the analysis, the processing unit 110 may induce one or more navigation responses in the vehicle 200 (such as turning, lane changing, or changes in acceleration), as discussed below in conjunction with the navigation response module 408. Furthermore, in some embodiments, the stereo image analysis module 404 may implement techniques associated with trained systems (such as neural networks or deep neural networks) or untrained systems (such as systems that can be configured to use computer vision algorithms to detect and / or label objects in the environment from which they capture and process sensor information). In one embodiment, the stereo image analysis module 404 and / or other image processing modules may be configured to use a combination of trained and untrained systems.
[0110] In one embodiment, the rate and acceleration module 406 may store software configured to analyze data received from one or more computing and electromechanical devices in the vehicle 200, which are configured to cause changes in the rate and / or acceleration of the vehicle 200. For example, the processing unit 110 may execute instructions associated with the rate and acceleration module 406 to calculate a target speed for the vehicle 200 based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include, for example, target position, rate and / or acceleration, the position and / or speed of the vehicle 200 relative to nearby vehicles, pedestrians, or road objects, and position information of the vehicle 200 relative to lane markings on the road. In addition, the processing unit 110 may calculate the target speed for the vehicle 200 based on sensor inputs (e.g., information from radar) and inputs from other systems of the vehicle 200 (such as the vehicle 200's throttle system 220, braking system 230, and / or steering system 240). Based on the calculated target speed, the processing unit 110 can transmit electronic signals to the throttle system 220, braking system 230 and / or steering system 240 of the vehicle 200 to trigger a change in rate and / or acceleration by, for example, physically pressing down the brakes or releasing the accelerator of the vehicle 200.
[0111] In one embodiment, the navigation response module 408 may store software executable by the processing unit 110 to determine the desired navigation response based on data derived from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. Such data may include position and speed information associated with nearby vehicles, pedestrians, and road objects, target position information for vehicle 200, etc. Additionally, in some embodiments, the navigation response may be (partially or entirely) based on map data, the predetermined position of vehicle 200, and / or the relative rate or relative acceleration between vehicle 200 and one or more objects detected from the execution of the monocular image analysis module 402 and / or the stereo image analysis module 404. The navigation response module 408 may also determine the desired navigation response based on sensor inputs (e.g., information from radar) and inputs from other systems of vehicle 200 (such as the throttle system 220, braking system 230, and steering system 240 of vehicle 200). Based on the desired navigation response, the processing unit 110 can transmit electronic signals to the throttle system 220, braking system 230, and steering system 240 of the vehicle 200 to trigger the desired navigation response by, for example, turning the steering wheel of the vehicle 200 to achieve a predetermined angle of rotation. In some embodiments, the processing unit 110 can use the output of the navigation response module 408 (e.g., the desired navigation response) as input to the execution of the rate and acceleration module 406 to calculate the change in the speed of the vehicle 200.
[0112] Furthermore, any module disclosed herein (e.g., modules 402, 404, and 406) can implement techniques associated with trained systems (such as neural networks or deep neural networks) or untrained systems.
[0113] Figure 5A A flowchart illustrating an exemplary process 500A for inducing one or more navigation responses based on monocular image analysis, consistent with the disclosed embodiments, is provided. At step 510, processing unit 110 may receive multiple images via data interface 128 between processing unit 110 and image acquisition unit 120. For example, a camera included in image acquisition unit 120 (such as image capture device 122 with field of view 202) may capture multiple images of an area in front of vehicle 200 (or, for example, the side or rear of vehicle) and transmit the multiple images to processing unit 110 via a data connection (e.g., digital, wired, USB, wireless, Bluetooth, etc.). At step 520, processing unit 110 may execute monocular image analysis module 402 to analyze the multiple images, as described below. Figures 5B-5D In a further detailed description, by performing the analysis, the processing unit 110 can detect a set of features within the set of images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, etc.
[0114] At step 520, processing unit 110 may also execute monocular image analysis module 402 to detect various road hazards, such as parts of truck tires, fallen road signs, loose cargo, small animals, etc. Road hazards can have different structures, shapes, sizes, and colors, which can make the detection of such hazards more challenging. In some embodiments, processing unit 110 may execute monocular image analysis module 402 to perform multi-frame analysis on multiple images to detect road hazards. For example, processing unit 110 may estimate camera motion between consecutive image frames and calculate pixel parallax between frames to construct a 3D map of the road. Processing unit 110 can then use the 3D map to detect hazards on the road surface and above the road surface.
[0115] At step 530, processing unit 110 may execute navigation response module 408 based on the analysis performed at step 520 and as described above. Figure 4 The described techniques are used to induce one or more navigation responses in vehicle 200. Navigation responses may include, for example, turning, lane changing, changes in acceleration, etc. In some embodiments, processing unit 110 may induce one or more navigation responses using data derived from the execution of rate and acceleration module 406. Alternatively, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof. For example, processing unit 110 may induce vehicle 200 to change lanes and then accelerate by, for example, sequentially transmitting control signals to steering system 240 and throttle system 220 of vehicle 200. Alternatively, processing unit 110 may induce vehicle 200 to brake while changing lanes by, for example, simultaneously transmitting control signals to braking system 230 and steering system 240 of vehicle 200.
[0116] Figure 5B A flowchart illustrating an exemplary process 500B for detecting one or more vehicles and / or pedestrians in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500B. At step 540, processing unit 110 may determine a set of candidate objects representing possible vehicles and / or pedestrians. For example, processing unit 110 may scan one or more images, compare the images to one or more predetermined patterns, and identify possible locations within each image that may contain the object of interest (e.g., a vehicle, pedestrian, or a portion thereof). The predetermined patterns may be designed to achieve a high “false hit” rate and a low “miss” rate. For example, processing unit 110 may use a low threshold of similarity to the predetermined pattern to identify candidate objects as possible vehicles or pedestrians. Doing so allows processing unit 110 to reduce the probability of missing (e.g., not identifying) candidate objects representing vehicles or pedestrians.
[0117] At step 542, processing unit 110 may filter a set of candidate objects based on classification criteria to exclude certain candidates (e.g., irrelevant or less relevant objects). Such criteria can be derived from various attributes associated with object types stored in a database (e.g., a database stored in memory 140). Attributes may include object shape, size, texture, location (e.g., relative to vehicle 200), etc. Therefore, processing unit 110 may use one or more sets of criteria to reject erroneous candidates from a set of candidate objects.
[0118] At step 544, processing unit 110 may analyze multiple image frames to determine whether an object in a set of candidate objects represents a vehicle and / or a pedestrian. For example, processing unit 110 may track detected candidate objects across consecutive frames and accumulate frame-by-frame data associated with the detected objects (e.g., size, position relative to vehicle 200, etc.). Additionally, processing unit 110 may estimate parameters for the detected objects and compare the frame-by-frame position data of the objects with predicted positions.
[0119] At step 546, processing unit 110 can construct a set of measurements for the detected objects. Such measurements may include, for example, position, velocity, and acceleration values (relative to vehicle 200) associated with the detected objects. In some embodiments, processing unit 110 can construct the measurements based on estimation techniques using a series of time-based observations such as Kalman filters or linear quadratic estimation (LQE) and / or based on available modeling data for different object types (e.g., cars, trucks, pedestrians, bicycles, road signs, etc.). Kalman filters may be based on measurements of the object's scale, where the scale measurement is proportional to the collision time (e.g., the amount of time it takes for vehicle 200 to arrive at the object). Therefore, by performing steps 540 through 546, processing unit 110 can identify vehicles and pedestrians appearing within a set of captured images and derive information associated with the vehicles and pedestrians (e.g., position, velocity, size). Based on the identification and derived information, processing unit 110 can induce one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0120] At step 548, processing unit 110 may perform optical flow analysis on one or more images to reduce the probability of detecting "false hits" and missing candidate objects representing vehicles or pedestrians. Optical flow analysis may refer to, for example, analyzing motion patterns relative to vehicle 200 in one or more images associated with other vehicles and pedestrians, and said motion patterns are different from road surface motion. Processing unit 110 can calculate the motion of candidate objects by observing different positions of objects across multiple image frames captured at different times. Processing unit 110 can use position and time values as inputs to a mathematical model to calculate the motion of candidate objects. Therefore, optical flow analysis can provide another method for detecting vehicles and pedestrians near vehicle 200. Processing unit 110 may combine steps 540 to 546 to perform optical flow analysis to provide redundancy for detecting vehicles and pedestrians and increase the reliability of system 100.
[0121] Figure 5C A flowchart illustrating an exemplary process 500C for detecting road markings and / or lane geometry information in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500C. At step 550, processing unit 110 may detect a set of objects by scanning one or more images. To detect segments of lane markings, lane geometry information, and other relevant road markings, processing unit 110 may filter the set of objects to exclude those determined to be irrelevant (e.g., potholes, small rocks, etc.). At step 552, processing unit 110 may group segments detected in step 550 that belong to the same road marking or lane marking together. Based on this grouping, processing unit 110 may develop a model, such as a mathematical model, representing the detected segments.
[0122] At step 554, processing unit 110 may construct a set of measurements associated with the detected segment. In some embodiments, processing unit 110 may create a projection of the detected segment from the image plane to the real-world plane. The projection may be characterized using a cubic polynomial with coefficients corresponding to physical properties such as the detected road position, slope, curvature, and derivative of curvature. In generating the projection, processing unit 110 may consider changes in the road surface and the pitch and roll rates associated with vehicle 200. Furthermore, processing unit 110 may model the road elevation by analyzing position and motion cues present on the road surface. Additionally, processing unit 110 may estimate the pitch and roll rates associated with vehicle 200 by tracking a set of feature points in one or more images.
[0123] At step 556, processing unit 110 can perform multi-frame analysis, for example, by tracking the detected segments across consecutive image frames and accumulating frame-by-frame data associated with the detected segments. As processing unit 110 performs multi-frame analysis, the set of measurements constructed at step 554 can become more reliable and correlated with increasingly higher confidence levels. Therefore, by performing steps 550, 552, 554, and 556, processing unit 110 can identify road markings appearing within a set of captured images and derive lane geometry information. Based on the identified and derived information, processing unit 110 can induce one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0124] At step 558, processing unit 110 may consider additional information sources to further develop a safety model for scenarios involving vehicle 200 in its surrounding environment. Processing unit 110 may use the safety model to define scenarios in which system 100 can safely perform autonomous control of vehicle 200. To develop the safety model, in some embodiments, processing unit 110 may consider the positions and movements of other vehicles, detected road edges and barriers, and / or general road shape descriptions extracted from map data (such as data from map database 160). By considering additional information sources, processing unit 110 can provide redundancy for detecting road markings and lane geometry features and increase the reliability of system 100.
[0125] Figure 5D A flowchart illustrating an exemplary process 500D for detecting traffic lights in a set of images, consistent with the disclosed embodiments, is provided. Processing unit 110 may execute monocular image analysis module 402 to implement process 500D. At step 560, processing unit 110 may scan the set of images and identify objects appearing in the images at locations where traffic lights may be present. For example, processing unit 110 may filter the identified objects to construct a set of candidate objects, excluding those objects unlikely to correspond to traffic lights. This filtering may be performed based on various attributes associated with traffic lights, such as shape, size, texture, location (e.g., relative to vehicle 200), etc. Such attributes may be based on multiple instances of traffic lights and traffic control signals and stored in a database. In some embodiments, processing unit 110 may perform multi-frame analysis on a set of candidate objects reflecting possible traffic lights. For example, processing unit 110 may track candidate objects across consecutive image frames, estimate the real-world location of the candidate objects, and filter out those objects that are moving (which are unlikely to be traffic lights). In some embodiments, the processing unit 110 may perform color analysis on candidate objects and identify the relative positions of detected colors that appear inside possible traffic lights.
[0126] At step 562, processing unit 110 may analyze the geometric features of the intersection. This analysis may be based on any combination of: (i) the number of lanes detected on either side of vehicle 200, (ii) markings detected on the road (such as arrow markings), and (iii) a description of the intersection extracted from map data (such as data from map database 160). Processing unit 110 may use information derived from the execution of monocular analysis module 402 for this analysis. In addition, processing unit 110 may determine the correspondence between traffic lights detected at step 560 and lanes appearing near vehicle 200.
[0127] At step 564, as vehicle 200 approaches the intersection, processing unit 110 can update the confidence level associated with the analyzed intersection geometry and detected traffic lights. For example, the estimated number of traffic lights at the intersection may affect the confidence level compared to the actual number present at the intersection. Therefore, based on the confidence level, processing unit 110 can delegate control to the driver of vehicle 200 to improve safety conditions. By performing steps 560, 562, and 564, processing unit 110 can identify traffic lights appearing in the set of captured images and analyze intersection geometry information. Based on the identification and analysis, processing unit 110 can trigger one or more navigation responses in vehicle 200, as described above. Figure 5A As described.
[0128] Figure 5E A flowchart illustrating an exemplary process 500E for inducing one or more navigation responses in vehicle 200 based on a vehicle path, consistent with the disclosed embodiments, is provided. At step 570, processing unit 110 may construct an initial vehicle path associated with vehicle 200. The vehicle path may be represented using a set of points in coordinates (x, z), and the distance di between any two points in the set may fall within the range of 1 meter to 5 meters. In one embodiment, processing unit 110 may construct the initial vehicle path using two polynomials (such as a left-road polynomial and a right-road polynomial). Processing unit 110 may calculate the geometric midpoint between the two polynomials and, if applicable, include an offset of a predetermined amount (e.g., a smart lane offset) at each point in the resulting vehicle path (zero offset may correspond to driving in the middle of a lane). The offset may be along a direction perpendicular to the segment between any two points in the vehicle path. In another embodiment, processing unit 110 may use a polynomial and an estimated lane width to offset each point in the vehicle path by half the estimated lane width plus a predetermined offset (e.g., a smart lane offset).
[0129] At step 572, processing unit 110 may update the vehicle path constructed at step 570. Processing unit 110 may use a higher resolution to reconstruct the vehicle path constructed at step 570, such that the distance dk between any two points in the set of points representing the vehicle path is less than the distance di described above. For example, the distance dk may fall within the range of 0.1 meters to 0.3 meters. Processing unit 110 may use a parabolic spline algorithm to reconstruct the vehicle path, which may produce a cumulative distance vector S corresponding to the total length of the vehicle path (i.e., based on the set of points representing the vehicle path).
[0130] At step 574, processing unit 110 may determine a look-ahead point (represented by coordinates (xl, zl)) based on the updated vehicle path constructed at step 572. Processing unit 110 may extract the look-ahead point from the cumulative distance vector S, and the look-ahead point may be associated with a look-ahead distance and a look-ahead time. The look-ahead distance (which may have a lower limit in the range of 10 to 20 meters) may be calculated as the product of the vehicle 200's speed and the look-ahead time. For example, as the speed of vehicle 200 decreases, the look-ahead distance may also decrease (e.g., until it reaches the lower limit). The look-ahead time (which may be in the range of 0.5 to 1.5 seconds) may be inversely proportional to the gain of one or more control loops (such as a heading error tracking control loop) associated with the navigation response in vehicle 200. For example, the gain of the heading error tracking control loop may depend on the bandwidth of the yaw rate loop, steering actuator loop, vehicle lateral dynamics, etc. Therefore, the higher the gain of the heading error tracking control loop, the shorter the look-ahead time.
[0131] At step 576, processing unit 110 can determine the heading error and yaw rate command based on the look-ahead point determined at step 574. Processing unit 110 can determine the heading error by calculating the arctangent of the look-ahead point, for example, arctan(xl / zl). Processing unit 110 can determine the yaw rate command as the product of the heading error and the high-level control gain. If the look-ahead distance is not at the lower limit, the high-level control gain can be equal to: (2 / look-ahead time). Otherwise, the high-level control gain can be equal to: (2 * vehicle speed 200 / look-ahead distance).
[0132] Figure 5F A flowchart is provided to illustrate an exemplary process 500F for determining whether a vehicle ahead is changing lanes, consistent with the disclosed embodiments. At step 580, processing unit 110 may determine navigation information associated with the vehicle ahead (e.g., a vehicle traveling in front of vehicle 200). For example, processing unit 110 may use the above-described combination of... Figure 5A and 5BThe described technique determines the position, rate (e.g., direction and speed), and / or acceleration of a vehicle ahead. Processing unit 110 may also use the combination of the above. Figure 5E The described technique determines one or more road polynomials, look-ahead points (associated with vehicle 200), and / or snail tracks (e.g., a set of points describing the path taken by the vehicle ahead).
[0133] At step 582, processing unit 110 may analyze the navigation information determined at step 580. In one embodiment, processing unit 110 may calculate the distance between the snail track and the road polynomial (e.g., along the track). If the variance of this distance along the track exceeds a predetermined threshold (e.g., 0.1 to 0.2 meters on a straight road, 0.3 to 0.4 meters on a moderately curved road, and 0.5 to 0.6 meters on a road with sharp curves), processing unit 110 may determine that the vehicle ahead may be changing lanes. In the case of multiple vehicles detected traveling in front of vehicle 200, processing unit 110 may compare the snail track associated with each vehicle. Based on the comparison, processing unit 110 may determine that a vehicle whose snail track does not match the snail tracks of other vehicles may be changing lanes. Processing unit 110 may additionally compare the curvature of the snail track (associated with the vehicle ahead) with the expected curvature of the road segment the vehicle ahead is traveling on. The expected curvature can be extracted from map data (e.g., data from map database 160), from road polynomials, from the snail tracks of other vehicles, from prior knowledge about the road, etc. If the difference between the curvature of the snail track and the expected curvature of the road segment exceeds a predetermined threshold, the processing unit 110 can determine that the vehicle ahead may be changing lanes.
[0134] In another embodiment, processing unit 110 may compare the instantaneous position of the vehicle ahead with a look-ahead point (associated with vehicle 200) over a specific time period (e.g., 0.5 to 1.5 seconds). If the distance between the instantaneous position of the vehicle ahead and the look-ahead point changes over the specific time period, and the cumulative sum of the changes exceeds a predetermined threshold (e.g., 0.3 to 0.4 meters on a straight road, 0.7 to 0.8 meters on a moderately curved road, and 1.3 to 1.7 meters on a road with sharp curves), processing unit 110 may determine that the vehicle ahead may be changing lanes. In another embodiment, processing unit 110 may analyze the geometry of the track by comparing the lateral distance traveled along the snail's path with the expected curvature of the snail's path. The expected radius of curvature can be determined by calculating (δz² + δx²) / 2 / (δx), where δx represents the lateral distance traveled and δz represents the longitudinal distance traveled. If the difference between the traveled lateral distance and the expected curvature exceeds a predetermined threshold (e.g., 500 to 700 meters), processing unit 110 can determine that the vehicle ahead may be changing lanes. In another embodiment, processing unit 110 can analyze the position of the vehicle ahead. If the position of the vehicle ahead obscures the road polynomial (e.g., the vehicle ahead is covered on top of the road polynomial), processing unit 110 can determine that the vehicle ahead may be changing lanes. If the position of the vehicle ahead is such that another vehicle is detected in front of it and the snail tracks of the two vehicles are not parallel, processing unit 110 can determine that the (closer) vehicle ahead may be changing lanes.
[0135] At step 584, processing unit 110 may determine whether the vehicle 200 ahead is changing lanes based on the analysis performed at step 582. For example, processing unit 110 may make the determination based on a weighted average of the various analyses performed in step 582. In such an approach, for example, a decision made by processing unit 110 based on a particular type of analysis that the vehicle ahead may be changing lanes may be assigned a value "1" (and "0" to indicate a determination that the vehicle ahead is unlikely to be changing lanes). Different analyses performed at step 582 may be assigned different weights, and the disclosed embodiments are not limited to any particular combination of analysis and weights.
[0136] Figure 6A flowchart illustrating an exemplary process 600 for evoking one or more navigation responses based on stereoscopic image analysis, consistent with the disclosed embodiments, is provided. At step 610, processing unit 110 may receive a first plurality of images and a second plurality of images via data interface 128. For example, a camera included in image acquisition unit 120 (such as image capture devices 122 and 124 having fields of view 202 and 204) may capture the first plurality of images and the second plurality of images of an area in front of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first plurality of images and the second plurality of images via two or more data interfaces. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0137] At step 620, processing unit 110 may execute stereo image analysis module 404 to perform stereo image analysis on the first plurality of images and the second plurality of images to create a 3D map of the road in front of the vehicle and detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. Stereo image analysis can be combined with the above. Figures 5A-5D The described steps are performed in a similar manner. For example, processing unit 110 may execute stereo image analysis module 404 to detect candidate objects (e.g., vehicles, pedestrians, road markings, traffic lights, road hazards, etc.) in a first plurality of images and a second plurality of images, filter out subgroups of candidate objects based on each object, perform multi-frame analysis, construct measurement results, and determine confidence levels for the remaining candidate objects. In performing the steps described above, processing unit 110 may consider information from both the first plurality of images and the second plurality of images, rather than information from only one set of images. For example, processing unit 110 may analyze differences in pixel-level data (or other subsets of data from the two streams of captured images) of candidate objects appearing in both the first plurality of images and the second plurality of images. As another example, processing unit 110 may estimate the position and / or rate (e.g., relative to vehicle 200) of a candidate object by observing that the candidate object appears in one of the plurality of images but not the other, or relative to other differences that may exist relative to objects appearing in both image streams. For example, the position, velocity, and / or acceleration relative to vehicle 200 can be determined based on the trajectory, position, motion characteristics, etc., of features associated with one or both objects appearing in the image stream.
[0138] At step 630, processing unit 110 may execute navigation response module 408 based on the analysis performed at step 620 and as described above. Figure 4The described techniques are used to induce one or more navigation responses in vehicle 200. Navigation responses may include, for example, turning, lane changing, acceleration change, speed change, braking, etc. In some embodiments, processing unit 110 may use data derived from the execution of speed and acceleration module 406 to induce one or more navigation responses. Additionally, multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.
[0139] Figure 7 A flowchart illustrating an exemplary process 700 for inducing one or more navigation responses based on the analysis of three sets of images, consistent with the disclosed embodiments, is provided. At step 710, processing unit 110 may receive a first plurality of images, a second plurality of images, and a third plurality of images via data interface 128. For example, cameras included in image acquisition unit 120 (such as image capture devices 122, 124, and 126 having fields of view 202, 204, and 206) may capture the first plurality of images, the second plurality of images, and the third plurality of images of a region in front of and / or to the sides of vehicle 200 and transmit them to processing unit 110 via a digital connection (e.g., USB, wireless, Bluetooth, etc.). In some embodiments, processing unit 110 may receive the first plurality of images, the second plurality of images, and the third plurality of images via three or more data interfaces. For example, each of image capture devices 122, 124, and 126 may have an associated data interface for transmitting data to processing unit 110. The disclosed embodiments are not limited to any particular data interface configuration or protocol.
[0140] In step 720, processing unit 110 can analyze the first plurality of images, the second plurality of images, and the third plurality of images to detect features within the images, such as lane markings, vehicles, pedestrians, road signs, highway exit ramps, traffic lights, road hazards, etc. This analysis can be combined with the above description. Figures 5A-5D and Figure 6 The steps described are performed in a similar manner. For example, processing unit 110 can perform monocular image analysis on each of the first plurality of images, the second plurality of images, and the third plurality of images (e.g., via execution by monocular image analysis module 402 and based on the above). Figures 5A-5D The steps described above). Alternatively, processing unit 110 may perform stereoscopic image analysis on the first plurality of images and the second plurality of images, the second plurality of images and the third plurality of images and / or the first plurality of images and the third plurality of images (e.g., via execution by stereoscopic image analysis module 404 and based on the above). Figure 6The steps described herein. Processed information corresponding to the analysis of a first plurality of images, a second plurality of images, and / or a third plurality of images can be combined. In some embodiments, processing unit 110 may perform a combination of monocular image analysis and stereoscopic image analysis. For example, processing unit 110 may perform monocular image analysis on the first plurality of images (e.g., via execution of monocular image analysis module 402) and stereoscopic image analysis on the second plurality of images and the third plurality of images (e.g., via execution of stereoscopic image analysis module 404). The configuration of image capture devices 122, 124, and 126—including their respective positioning and fields of view 202, 204, and 206—can affect the type of analysis performed on the first plurality of images, the second plurality of images, and the third plurality of images. The disclosed embodiments are not limited to the specific configuration of image capture devices 122, 124, and 126, or the type of analysis performed on the first plurality of images, the second plurality of images, and the third plurality of images.
[0141] In some embodiments, processing unit 110 may test system 100 based on the images acquired and analyzed at steps 710 and 720. Such testing may provide an indicator of the overall performance of system 100 for certain configurations of image capture devices 122, 124, and 126. For example, processing unit 110 may determine the ratio of “false hits” (e.g., situations where system 100 incorrectly determines the presence of a vehicle or pedestrian) to “missed hits”.
[0142] At step 730, processing unit 110 may induce one or more navigation responses in vehicle 200 based on information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images. The selection of either the first plurality of images, the second plurality of images, or the third plurality of images may depend on various factors, such as the number, type, and size of objects detected in each of the plurality of images. Processing unit 110 may also base its responses on image quality and resolution, the effective field of view reflected in the images, the number of captured frames, and the extent to which one or more objects of interest actually appear in the frames (e.g., the percentage of frames in which the object appears, the proportion of objects appearing in each such frame, etc.).
[0143] In some embodiments, processing unit 110 can select information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images by determining the degree to which information derived from one image source is consistent with information derived from other image sources. For example, processing unit 110 can combine processed information derived from each of image capture devices 122, 124, and 126 (whether by monocular analysis, stereo analysis, or any combination of both) and determine consistent visual indicators (e.g., lane markings, detected vehicles and their locations and / or paths, detected traffic lights, etc.) across the images captured from each of image capture devices 122, 124, and 126. Processing unit 110 can also exclude inconsistent information across the captured images (e.g., vehicles changing lanes, lane models indicating vehicles too close to vehicle 200, etc.). Therefore, processing unit 110 can select information derived from both of the first plurality of images, the second plurality of images, and the third plurality of images based on the determination of consistent and inconsistent information.
[0144] Navigation responses may include, for example, turning, lane changes, and changes in acceleration. Processing unit 110 can base its responses on the analysis performed in step 720 and the above-mentioned factors. Figure 4 The described techniques are used to induce one or more navigation responses. Processing unit 110 may also use data derived from the execution of rate and acceleration module 406 to induce one or more navigation responses. In some embodiments, processing unit 110 may induce one or more navigation responses based on the relative position, relative rate, and / or relative acceleration between vehicle 200 and objects detected in any of a first plurality of images, a second plurality of images, and a third plurality of images. Multiple navigation responses may occur simultaneously, sequentially, or in any combination thereof.
[0145] Anomaly detection
[0146] Autonomous vehicle (AV) systems, whether partially or fully autonomous, can be configured to navigate based on individual information sources or combinations thereof. For example, an AV can navigate based on mapping information (e.g., REM maps, which store 3D representations of drivable paths for each available lane, such as road segments, intersections, and navigable areas; representations and locations of traffic lights, traffic signs, road edges, lane markings, and various other types of road terrain; and navigation aids such as traffic light relevance, lane priority, and road geometry / sizes relative to drivable paths).
[0147] The AV can also navigate relative to information collected by sensors on the host vehicle. These sensors can include one or more cameras (e.g., a surround-view camera system), LiDAR, radar, etc. The collected sensing information can be aggregated, refined, etc., to provide the host vehicle's navigation system with a sensing state at a specific time snapshot. This sensing state can indicate object detection in the host vehicle's environment (e.g., target vehicles, pedestrians, road debris, road terrain, etc.) and can also include spatial information associated with the detected objects. This spatial information can include any or all of the following: object size, bounding box size, height above the road surface, relative spacing between detected objects, depth / range information, 3D point coordinates of one or more points associated with the detected object, etc.
[0148] To make navigation decisions, available sensing conditions (and mapping information, if available) can be provided to a driving strategy, one or more end-to-end systems, or other types of systems configured to generate planned navigation actions based on the vehicle's sensing conditions. Such driving strategies, end-to-end systems, etc., can generate planned navigation actions to be implemented by vehicle systems (e.g., steering, braking, accelerator, etc.) to advance the AV's navigation objective (e.g., moving from the current location to the planned destination) while operating safely relative to the object and the situation represented by the sensing conditions. In some cases, such driving strategies or end-to-end systems can be configured to impose certain navigation constraints (e.g., a minimum safe distance maintained relative to a target vehicle, a minimum buffer distance relative to pedestrians, etc.) when generating planned navigation actions.
[0149] Modern AV can incorporate various types of deep learning models (e.g., trained models, trained neural networks, trained GNNs, transformer-based networks, etc.) to support a range of functions, including object identification, situation identification / recognition, identifying or interpreting semantic relationships in a scene, generating planned trajectories based on sensed targets and road terrain, validating planned trajectories, and many other tasks. These models can also help generate appropriate navigation actions in response to detected vehicles, pedestrians, traffic light states, traffic signs, etc.
[0150] Under normal conditions, trained models can operate with extremely high accuracy and extremely low failure rates. For example, trained models can become highly proficient at performing tasks on the training datasets used to train them. Even if the model's input falls outside the data represented by the training dataset used to train the model, if the input is similar to the training dataset, the trained model can skillfully "infer" to predict the correct answer based on the similarity between the new input and the datasets it experienced during training.
[0151] Edge cases can be more difficult to handle for trained models. For example, identifying and responding appropriately to anomalies is a significant challenge for decision-making systems that rely on sensor input. While trained models (e.g., neural networks) can be trained relative to anomalous edge cases, collecting or generating training data representing the virtually limitless number of possible anomalies that might be encountered is difficult or impossible (e.g., towed vehicles, rear-facing vehicles carrying cargo, pedestrians lying on the road surface, pedestrians standing on top of vehicles, mattresses tied to vehicle roofs swaying in the wind, sunlight patterns in acquired images, motorcycles traveling between lanes, etc.).
[0152] One approach to addressing the challenge of improving the performance of trained models in recognizing anomalous situations is to simply train the model with a robust method that involves an ever-increasing number of edge cases indicative of these anomalous situations. New training data can be generated to simulate these types of features or situations and their variations whenever the AV system fails to identify a road feature or situation or fails to respond to it as expected. However, regardless of the number of training instances provided, there will always be anomalous situations that fall outside the network's training scope, potentially limiting the performance of the trained model.
[0153] Currently disclosed embodiments address this challenge from a different angle. Instead of training a system to directly identify what constitutes anomalies (which would require iteratively training the network to recognize virtually countless specific types of anomalies / objects), the system of the present invention can identify, based on its training (e.g., using training samples representing normal situations / objects), whether a particular sample falls into what is considered a normal situation, or whether the sample exhibits one or more features that would cause it to fall outside what is considered a normal situation. For example, a normal region can be represented by a predetermined embedding space distribution containing encodings (e.g., feature vectors) associated with normally occurring objects, scenes, etc. If the encoding of a particular input sample falls outside this predetermined embedding space distribution, the system can determine that the input sample represents an anomalous object, scene, etc. In this way, the system of the present invention can be trained on readily available “normal” training dataset samples, yet can identify far more than a limited number of predetermined anomalies / objects. Conversely, based on determining or inferring that these situations / objects do not fall into situations that the network has been trained to identify as “normal,” the system of the present invention has the potential to identify / identify countless anomalies / objects.
[0154] One example embodiment includes a system for navigating a host vehicle relative to road segments (e.g., any road section, intersection, roundabout, navigable area, etc.). The system includes at least one processor, which includes a circuit system and a memory, wherein the memory includes instructions executable by the circuit system. The processor and memory may include any computing hardware and memory units described in the preceding sections. When executed by the processor (or at least one processor), the instructions are configured to cause the processor to receive captured images obtained by a camera on the host vehicle; generate a representation in an embedding space of at least a portion of the captured images; determine whether the representation in the embedding space of the at least a portion of the captured images falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region; determine a navigation action for the host vehicle based on the determination that the representation in the embedding space of the at least a portion of the captured images falls outside the predetermined embedding space region; and cause at least one system associated with the host vehicle to perform the navigation action.
[0155] An embedding space, or latent space, can be viewed as a mathematical structure in which real-world data can be encoded, compressed, simplified, etc., using feature vector representations of real-world data. The latent space is a low-dimensional, continuous space in which input data is encoded. In the embedding space, similar items / features / object representations are positioned closer to each other than less similar items, giving each item a specific location based on its characteristics. In this way, embedding spaces can be used to efficiently organize information. Because embedding spaces focus on similarity and distance, mapping high-dimensional data to a low-dimensional space, using embedding spaces can reduce computational and memory requirements while improving model efficiency. Embedding spaces also preserve semantic and syntactic relationships, enabling deep learning models to more effectively grasp the data domain of the real world.
[0156] The process of creating an embedding space involves transforming raw data into a structured format that facilitates efficient analysis and pattern recognition. This transformation typically converts input elements (such as images or portions of images) into numerical vectors (feature vectors) that capture the characteristics of the image or image portion. Each image (or image portion) is fed to an encoder that maps image features into high-dimensional vectors, where semantic relationships are preserved. This mapping allows algorithms, trained models, and the like to interpret image data based on proximity and orientation within the embedding space.
[0157] The similarity between images or image segments can be indicated by the proximity (or distance) between their vector representations in the embedding space. Distance in the embedding space can be determined using various techniques. For example, the Euclidean distance between vectors can be determined. This distance represents the "straight-line" distance between two vectors. Cosine similarity between vectors can also be determined based on the angular distance between two vectors. Cosine similarity offers the advantage of being less sensitive to differences in vector magnitudes. A smaller distance in the embedding space between encoded vectors generally indicates a higher degree of similarity and / or semantic relationship between the corresponding real-world data (e.g., images or image segments).
[0158] In some cases, dedicated encoders can be used to generate feature vectors based on input images or image segments. Representations of images or image segments can also be performed using trained models. Various types of encoders also exist, such as variational autoencoders (VAEs), which are a type of trained model / neural network that can be particularly useful for learning efficient data representations for dimensionality reduction, feature learning, etc.
[0159] An autoencoder can consist of an encoder and a decoder. The encoder compresses the input data into a lower-dimensional latent space, while the decoder reconstructs the original data from this compressed representation. During training of variational autoencoders or other types of models, the goal is to minimize the difference between the input and the reconstructed output. For example, as... Figure 8 The training of the VAE can be described as follows: Image 810 (or a portion of a captured image) is presented to the model as input 812. The model, including encoder 814, is forced to encode the input image (such as input image 810) into a compressed embedding space 816. The model, including decoder 818, is forced to decode the compressed image or image segment from embedding space 816 to generate a corresponding image output, such as output image 820. The output image can be compared with the input image. If a difference exists, the model can be penalized, and the weights can be adjusted. This process can be iterative until the output image (after decoding) matches the input image (before encoding). In this way, the network learns to generate an accurate compressed representation of the input image or image segment. The input image / image segment can be compressed into a vector space representation according to the following:
[0160] Z i ~N(μ) i ,σ i 2 I)
[0161] A sparse vector space represents a distribution formed across the vector space (e.g., a multi-dimensional distribution such as a Gaussian distribution). The model can learn this distribution so that, during field runs, the trained model can use the learned distribution as a reference to evaluate whether a received image or image segment includes a scene or object represented by feature vectors, for example, falling outside the "normal" region of the distribution when compressed into the sparse vector space. That is, if the encoded input image / image segment is represented by feature vectors, scores, or values falling outside a predetermined region representing a normal scene or object in the embedding space distribution, the model can return an output indicating that the input image / image segment represents an anomalous or abnormal scene.
[0162] The learned distribution of encountered images / image segments in vector space can represent the type of scenes or objects the model expects to encounter during runtime. For example, for autonomous vehicle applications, the learned distribution (generated through model training) may be associated with the type of scenes the autonomous vehicle navigation system expects to encounter. These scenes can include representations of the environment in front of, to the sides of, behind, or around the vehicle in a 360° radius. Furthermore, such scenes can include representations of faces (e.g., the driver's face), joint visibility information of people, road signs, road markings, road edges, road surfaces, road surface markings, and various other types of scenes or objects typically found in the environment of a navigation vehicle. The trained neural network learns to represent one or more distributions of "normal" scenes or objects, such that when the generated embeddings fall outside one or more normal distributions, the system can indicate that an abnormal scene or object has been encountered.
[0163] As a simple example of how a trained model can operate, if the model's training dataset includes input image 810, then when the model is running in the field (e.g., as part of a vehicle navigation system), if the model receives input including a similar scene containing the sign represented in input image 810, the model will almost certainly return an indication that the input image does not represent an anomalous scene. The model's training may also include hundreds or thousands (or more) other types of traffic signs that are typically encountered during vehicle navigation. Therefore, input images including representations of any of these signs may also cause the output from the trained model to indicate that the input image does not represent an anomalous object / scene. Furthermore, the input image may represent road signs not included in the dataset used to train the model. Even in this case, as long as the new road sign is sufficiently similar to the model's training data (e.g., if the encoding of the new sign falls within a predetermined range considered to represent a "normal" range by the mathematical distribution of the embedding space), the trained model may indicate that the new sign is not anomalous. On the other hand, if the input received by the trained model is very similar to the input image 810, except that the markings include yellow spray paint graffiti, the model may determine that the encoding of such image input falls outside the “normal” range of the embedding space distribution and may return an output indicating that the input image represents an anomalous or unusual scene.
[0164] As described above, the trained model may be forced to generate an embedding space represented by a predetermined distribution (e.g., a Gaussian distribution or other types of probability distribution). The size of the predetermined embedding space region used to specify normal and abnormal embeddings (and therefore, the input image, features represented by the input image, etc.) can be optional and can correspond to a predetermined sub-region of the probability distribution.
[0165] Using thresholding, a trained model can determine whether a particular input image falls within the normal distribution range of the represented subject (e.g., a normal face, sign, etc.) or whether the represented subject in the input image falls outside the normal distribution range. This normal distribution range can be set by one or more thresholds. The thresholds can be selected (e.g., by the model creator, user, etc.) to fine-tune the normal / abnormal output sensitivity of the trained model. In some cases, the performance of the trained model can be fine-tuned so that a pedestrian on a crosswalk returns a normal indicator, while the same pedestrian lying on the crosswalk returns an abnormal indicator. Using thresholding, the performance of the trained model can be fine-tuned so that the model returns a normal indicator (i.e., not abnormal) in response to both input images representing real people and input images representing crash test dummies. In other words, while a trained model may be able to distinguish between real people and dummies (e.g., based on the spacing in the embedding space of their respective embeddings), there may be situations where it might be desirable to classify a non-human driver / passenger as "normal" (e.g., in vehicle safety testing).
[0166] As described above, a trained model can identify such objects / scenes based on whether the embeddings of images or image segments representing anomalous objects, scenes, etc., fall outside a predetermined embedding space region, which is defined as a non-anomalous embedding space region. Various thresholding techniques exist for defining non-anomalous embedding space regions. In some cases, the chosen thresholding technique may depend on whether the embedding space is represented as a probability distribution, and if so, on the type of probability distribution employed. In instances where the embedding space is represented as a Gaussian distribution, such as... Figure 9 As shown, an anomaly threshold can be selected to capture any desired percentage of the embedding space within the "normal" range of the distribution. This "normal" range can be set by selecting a threshold for any desired standard deviation σ associated with the probability distribution. For example, as... Figure 9 As shown, a threshold of σ (one standard deviation from the center) will establish approximately 68.2% of the embedding space as an indication of "normal" objects / scenes, etc. A threshold of two standard deviations (2σ) can capture approximately 95.4% of the embedding space within the "normal" region, and a threshold of three standard deviations (3σ) can capture approximately 99.6% of the embedding space within the "normal" region. The chosen threshold does not have to be limited to the entire standard deviation. Instead, any desired percentage or y-value associated with the probability distribution can be chosen as the "normal" range threshold. For example, a threshold of 99.9% can be established (e.g., during model generation or model use, etc.) to establish the normal region of the embedding space. In such cases, besides Figure 9 Except for the extreme tails of the Gaussian distribution shown, all others represent the normal embedding space. Any encoded output falling outside the established “normal” threshold (e.g., threshold 918 indicates that the normal embedding space sub-region covers approximately 99.8% of the embedding space) will be indicated as including anomalies / abnormal scenes, etc. On the other hand, an image input including a representation of a typical road sign indicating two-way traffic (such as image 912) may result in an encoder output falling within the first standard deviation of the Gaussian distribution 910 (e.g., indicated by arrow 916), which will be indicated by the trained model as an indication of a normal object / scene.
[0167] on the contrary, Figure 10This indicates that the trained model identifies the input image 1010 as an instance including a flag that is identified as an anomaly. For example, in response to receiving the flag 1010 as input, the encoder 814 generates encoded embedding space values (e.g., feature vectors, scores, % values, etc.) near the tail end of the Gaussian distribution 1016, as indicated by arrow 1020. In this case, the encoded embedding space value indicated by arrow 1020 falls outside the threshold 1018 that defines the "normal" sub-region of the embedding space. Therefore, the trained model returns an output value (e.g., a Boolean value) indicating that the input image 1010 represents an anomalous object, scene, etc.
[0168] The sensitivity of a trained model in identifying anomalies can be controlled by varying the "normal" threshold. For example, as the threshold increases (e.g., higher standard deviation, lower percentage value, etc.), the trained model becomes less sensitive to anomalies represented by the input image. Conversely, as the threshold increases (e.g., less standard deviation, lower percentage value, etc.), the trained model becomes more sensitive to anomalies represented by the input image. Therefore, a predetermined threshold for detecting anomalous events can be fine-tuned to provide the desired level of sensitivity. Higher sensitivity may result in only a relatively concentrated region in the sparse vector space being associated with "normal" events. Lower sensitivity may result in a larger region in the sparse vector space being associated with "normal" events. In some cases, it may be desirable to fine-tune the threshold to provide relatively low sensitivity, thereby limiting the number of false alarms (e.g., "normal" situations being incorrectly identified as anomalous). In a particular instance, the described system can be fine-tuned so that test dummies are not identified as anomalous, which is beneficial for testing a primary vehicle navigation system. That is, while the system can distinguish between real people and test dummies, its sensitivity threshold can be fine-tuned so that, for navigation system testing purposes, the described system can classify real people and test dummies within a "normal" distribution space. In other words, test dummies are not identified as anomalous, which can affect the results of navigation system testing.
[0169] Encoder 814 can be configured to generate any suitable output associated with an established or referenced embedding space. In some cases, encoder 814 can be configured to generate feature vectors, and the embedding space can be defined by a threshold distance relative to an embedding space reference (e.g., the embedding space origin). For example, in such cases, “normal” feature vectors (and their corresponding input images) will correspond to those feature vectors that fall within a predetermined distance from the origin (e.g., determined using Euclidean distance, cosine similarity, or other techniques). Feature vectors that fall outside the predetermined distance will indicate an anomaly or anomalous scene represented in the corresponding input image. In other cases, encoder 814 can be configured to directly output a score or other indicator (e.g., a percentage value, etc.) relative to a probability distribution representing the embedding space. In still other cases, the encoder can encode the input image, compare the encoded values to a predetermined embedding space (e.g., an embedding space determined during training), and then return an output that simply indicates whether the encoded values fall within or outside the “normal” embedding space. One or more additional software-based modules can also be used to help determine whether the encoded values, feature vectors, etc., representing the input image fall into the “normal” region of the embedding space.
[0170] As described above, the threshold used to fine-tune the sensitivity of the trained model can be user-selectable. For example, a user can set the threshold defining the "normal range" of the embedding space to any value that provides the expected performance level relative to the encountered objects / scenes. If the system indicates that certain objects or scenes (e.g., crash test dummies in a test vehicle) are anomalous, but the user wants to classify these objects / scenes as non-anomalous (normal), the user can increase the threshold to expand the "normal" embedding space region. In this way, the normal embedding space region can correspond to a predetermined sub-region of the embedding space defined by at least one user-selectable parameter value. In some cases, the user-selectable parameter value is an embedding space distance indicator (e.g., Euclidean distance, angular cosine similarity value, etc.). In other cases, the user-selectable parameter value is a percentage value (or other type of value) associated with a probability distribution representation of the embedding space.
[0171] The described master vehicle navigation system can determine the master vehicle's navigation actions based on the output of a trained model indicating whether an anomalous object or scene has been encountered. For example, a determined navigation action can be generated in response to the determination that at least a portion of the captured image's representation in the embedding space falls outside a predetermined embedding space region (e.g., the embedding space is defined as a sub-region representing a "normal" object / scene). Various navigation actions and types can be taken in response to detected anomalous objects, scenes, etc. In some cases, in response to detected anomalous objects or scenes, the master vehicle navigation system can cause the braking system to decelerate or stop the master vehicle, the acceleration system to increase the vehicle's speed, or the steering system to change the master vehicle's heading. In some cases, planned navigation actions in response to detected anomalous objects / scenes may include maintaining the master vehicle's current speed and heading. Alternatively or additionally, in response to detected anomalous objects / scenes, navigation actions may include generating auditory or visual alarms to provide to the master vehicle's operator or passengers (e.g., via one or more speakers, displays, interactive screens, etc.).
[0172] The detection of anomalous objects or scenes can serve as a trigger to cause the primary vehicle navigation system (e.g., a policy component of the primary vehicle navigation system) to take one or more additional actions. For example, the primary vehicle navigation system can perform additional analysis relative to the output of one or more sensors on the primary vehicle. For instance, acquired images from cameras or point clouds generated by one or more lidar or radar systems can be analyzed to determine the motion characteristics of one or more anomalous objects in the scene. Such additional analysis may enable the primary vehicle navigation system to generate appropriate planned navigation actions (e.g., maintaining heading and speed when it is determined that an anomalous object is moving away from the primary vehicle).
[0173] Any input image or portion thereof can be encoded by a publicly disclosed trained model and represented in an embedding space. In some cases, the encoded image represented in the embedding space may correspond to an image acquired representing the full field of view (FOV) of a camera on the host vehicle. However, in other cases, a region of interest can be identified in the acquired image, and image data from the identified region of interest can be encoded into the embedding space. Such regions of interest can be identified by one or more trained models or image segmentation techniques, and among other things, may include representations of traffic signs, road signs, traffic lights, target vehicles, debris, objects on road surfaces, objects adjacent to roads, road barrier structures, pedestrians, cyclists, etc.
[0174] The disclosed embodiments can be used to enable a primary vehicle navigation system to identify and respond to virtually any type of anomalous object, situation, or scenario that may be encountered. In fact, the disclosed embodiments, including the described trained neural network, can be configured to identify virtually any type of anomalous situation or object that exceeds normal expectations. The following sections will discuss several examples of anomalous scenario types that the disclosed embodiments may encounter.
[0175] Figure 11 This illustrates a scenario where an autonomous navigation system may struggle to properly identify and navigate. In this example, the primary vehicle is traveling along a highway and detects two target vehicles ahead of it. The first target vehicle, 1110, is directly in front of the primary vehicle and traveling in the same lane. The second target vehicle, 1120, is to the right front of the primary vehicle. The navigation system has identified target vehicle 1120 (as indicated by the dashed bounding box), but target vehicle 1120 includes a carrier vehicle and a transported vehicle traveling in the opposite direction. Because the transported vehicle faces the primary vehicle (e.g., features typically associated with the front of the vehicle (headlights, windshield, grille, etc.) are visible to the primary vehicle), the primary vehicle's navigation system can make the preliminary determination that vehicle 1120 is a single vehicle traveling towards the primary vehicle. While such scenarios are rare on highways, they are possible and potentially dangerous. If the primary vehicle encounters a target vehicle traveling in the wrong direction on a highway, a reasonable response could include steer the primary vehicle toward the shoulder, reduce its speed, or bring it to a stop. In such situations, maintaining the course and speed of the main vehicle can be very tricky and even potentially dangerous.
[0176] However in Figure 11 In the described scenario, the transported vehicle moves together with the carrier vehicle and in the same direction as the main vehicle. Therefore, the risk posed by the target vehicle 1120 to the main vehicle may not be greater than that posed by the target vehicle 1110. In such cases, the main vehicle does not expect to perform the type of evasive maneuver that would be appropriate if the target vehicle 1120 were traveling in the opposite direction to the main vehicle.
[0177] The disclosed embodiments enable a primary vehicle navigation system to determine and implement appropriate responses to such unusual situations. For example, upon receiving an acquired image frame 1130, the encoder of the disclosed trained model can encode image frame 1130 (or a sub-portion of image 1130, such as a region of interest including a representation of target vehicle 1120) into an embedding space and determine whether image 1130 contains a representation of an unusual object or scene. In this example, the trained model can return an indication that the representation of target vehicle 1120 (or the obvious scene represented by image 1130, i.e., a large truck potentially approaching the primary vehicle along a split-level highway) indicates an unusual situation. In response, the primary vehicle navigation system can determine one or more navigation actions for the primary vehicle based on the determination that the representation of at least a portion of the captured image 1130 in the embedding space falls outside a predetermined “normal” embedding space region.
[0178] In some cases, in response to indications of abnormal objects / scenes in the primary vehicle's environment, the primary vehicle system can perform additional analysis to better assess the situation. Figure 11 In some instances, the primary vehicle navigation system can access the output of one or more other systems / sensors to assess the motion characteristics of objects in the primary vehicle's environment. Such assessments can be performed in several ways. In some cases, this motion characteristic information can be determined and stored while a series of image frames are captured relative to the primary vehicle's environment (e.g., using structures in motion computing techniques; performing forward torsion on the captured images while subtracting the effects of the primary vehicle's self-motion to compare with subsequently captured images and reveal object motion separated from the effects of the primary vehicle's self-motion; inferring the motion characteristics of objects using one or more trained models based on known primary vehicle self-motion and image localization of objects in two or more captured image frames, etc.). In such cases, upon receiving an indication of an anomaly, the primary vehicle navigation system can access the stored motion information to determine whether the motion of the target vehicle 1120 poses a threat to the primary vehicle. In other cases, such analysis can be performed as needed in response to receiving an anomaly indication.
[0179] In other cases, the sensor outputs available to the primary vehicle navigation system can be analyzed to determine the motion characteristics of the target vehicle 1120 and whether the target vehicle poses a threat to the primary vehicle (e.g., a frontal collision threat). Such sensors may include lidar information, radar information, etc. In the actual situation depicted in image frame 1130, the trained system's indication of abnormal scenarios and review of the motion characteristics of the target vehicle 1120 (assuming the target vehicle 1120 is moving in the same direction and at a similar speed as the primary vehicle) may lead to planned navigation actions for the primary vehicle, including maintaining the primary vehicle's route and speed. On the other hand, if the review of the target vehicle 1120's motion characteristics indicates that the target vehicle 1120 is moving in the same direction as the primary vehicle but at a significantly slower speed, the primary vehicle navigation system may decelerate the primary vehicle and / or slightly steer to the left to increase the buffer distance between the primary vehicle and the target vehicle 1120.
[0180] Figure 12 This illustrates another anomalous situation in which the disclosed embodiments can assist in determining an appropriate navigation response. In this example, the captured image 1210 can be provided as input to the disclosed trained model. The trained model may return an indication that the generated embedding associated with the captured image 1210 falls outside a predetermined "normal" embedding space sub-region (e.g., because image 1210 includes a representation of the motorcycle 1220 traveling between lanes and between target vehicle lines on either side of the motorcycle). In response, and perhaps based on analysis of sensor outputs or motion determination techniques as described above, the primary vehicle navigation system can determine that not only is the motorcycle 1220 traveling between lanes, but the motorcycle 1220 is moving towards the primary vehicle from behind. In such a case, based on the anomalous scenario determination and the motion characteristics of the motorcycle 1220, the primary vehicle navigation system can determine the navigation actions of the primary vehicle. In this example, navigation actions may include maintaining the speed and heading of the primary vehicle, slightly turning the primary vehicle away from the direction of the motorcycle to increase the buffer distance between the primary vehicle and the motorcycle, etc.
[0181] Figure 13This illustrates another anomalous situation where the disclosed embodiments can assist in determining an appropriate navigation response. In this example, the captured image 1310 can be provided as input to the disclosed trained model. The trained model may return an indication that the generated embedding associated with the captured image 1310 falls outside a predetermined "normal" embedding space sub-region (e.g., because image 1310 includes a representation of the sun 1320 and some sun-related image effects 1330, 1340 that may obscure one or more target objects / vehicles in front of the host vehicle). If such effects cannot be identified, there may be missed detection and / or incorrect identification of objects (and other adverse effects) represented in one or more input images. In response, the host vehicle navigation system can analyze sensor outputs (e.g., radar outputs, etc.) to assess whether a target object / vehicle may be present in front of the host vehicle. If it can be confirmed (e.g., based on analysis of additional sensor outputs) that there is no hazard in front of the host vehicle, the host vehicle navigation system can generate a planned navigation action for the host vehicle, including maintaining the current speed and heading. On the other hand, based on the presence of image effects 1330 and 1340 in the captured image 1310, the main vehicle navigation system can determine the navigation actions of the main vehicle, including slowing down (or even stopping) the main vehicle.
[0182] In the above examples, the trained model's indication of whether a scene or object falls outside a predetermined "normal" sub-region of the embedding space can be based on the analysis of a single input image. Alternatively, however, the trained model's determination of whether an object or scene is anomalous can be based on a series of acquired images representing, for example, the environment of a primary vehicle over a specific duration. For example, the model can be trained on a sequence of images such that the trained network can identify anomalous situations that occur over time (e.g., a target vehicle on a collision trajectory with the primary vehicle, an object falling from or at risk of falling from a vehicle, a pedestrian walking along a road segment, a person gesturing to traffic, etc.). That is, the disclosed embodiments may include a trained model configured to encode multiple images (e.g., a video frame stream) and detect temporal anomalies using predetermined sub-regions in the embedding space. Such a model can be trained to fill the embedding space distribution based on what is normal only in terms of features and also based on the normal evolution of driving scenarios, etc.
[0183] Figure 14This illustrates another anomalous situation in which the disclosed embodiments can assist in determining an appropriate navigation response. In this example, the disclosed trained model can be provided with multiple captured images representing the environment of the master vehicle 1410 during the time period in which the images were captured. The trained model can return, for each (or any) individual image frame, an indication that the generated embedding associated with the captured image falls outside a predetermined "normal" embedding space sub-region. Furthermore, however, the trained model can also return, for any group of two or more images, an indication that the anomalous scenario has evolved over time. For example, in Figure 14 In the illustrated scenario, the primary vehicle 1410 navigates from left to right in the right lane of a highway segment. The primary vehicle 1410 follows the target vehicle 1420. The primary vehicle 1410 includes an onboard camera 1421, and the target vehicle 1420 occupies a portion of the field of view of the camera 1421, as indicated by line of sight 1422. Therefore, there exists an occlusion region 1424 obscured by the target vehicle 1420, such that the image captured by the camera 1421 will not include representations of objects located within the occlusion region 1424.
[0184] exist Figure 14 In the scenario, the second target vehicle 1440 is traveling in the right lane of the highway and is directly in front of the target vehicle 1420. The third target vehicle 1430 enters the highway from the entrance ramp 1432 and navigates to location 1451, which is in front of target vehicle 1420 and behind target vehicle 1440. At a similar time, target vehicle 1440 changes lanes from the right lane of the highway to the left lane and moves to location 1452.
[0185] Due to the movement of the target vehicle relative to camera 1421 and line of sight 1422, as the target vehicle 1430 approaches along the entrance ramp 1432, the target vehicle 1430 will become visible to camera 1421 and represented in the image captured by camera 1421. As part of the target vehicle identification and tracking process, the main vehicle navigation system of vehicle 1410 can assign a tracking ID to the target vehicle 1430. However, when the target vehicle 1430 moves to area 1424, it becomes hidden from camera 1421. But almost simultaneously, the target vehicle 1440 moves from a location in area 1424 that is not visible to camera 1421 to a location 1452 where camera 1421 becomes visible. Upon detection of vehicle 1440, the navigation system of vehicle 1410 can assign a tracking ID to vehicle 1440. Because vehicles 1430 and 1440 are spatially adjacent, and considering that the time it takes for vehicle 1430 to move to location 1451 (outside the line of sight of camera 1421) is close to the time it takes for vehicle 1440 to move to location 1452 (within the line of sight of camera 1421), the main vehicle navigation system of vehicle 1410 is very likely to incorrectly identify vehicle 1440 as vehicle 1430 and assign the tracking ID of vehicle 1430 to vehicle 1440. This could lead to tracking problems, especially when vehicle 1430 later appears in front of vehicle 1420 and enters the field of view of camera 1421.
[0186] In this scenario, a trained model might return a time-based indication of scene evolution across multiple captured image frames. For example, embeddings associated with individually captured images and / or changes occurring within captured images over a period of time might represent anomalies (such as mismatched trajectories, colors, body shapes, etc., between vehicles 1430 and 1440). Such anomalies can inform a trained model to return an indication of a time-based abnormal scene evolution. In response, the primary vehicle navigation system can reacquire the target vehicle ID before an erroneous vehicle ID transfer event occurs and then track forward the visible target vehicles, excluding erroneous vehicle ID transfers that led to the detection of the anomaly.
[0187] The disclosed embodiments may have many other uses. In some cases, the trained model can encode images captured from cameras facing the interior of the vehicle. In such instances, the generated embeddings and predetermined “normal” sub-regions of the embedding space can be used to detect the presence of anomalies relative to the AV passenger or operator. For example, one or more cameras can be arranged to capture images representing the driver, operator, or passenger of the host vehicle. The captured images can be provided to the described system, including the described trained neural network, and the trained neural network can return an indication of whether the captured individual’s face or any other aspect of the interior of the host vehicle is determined to be abnormal. Such situations may exist, for example, where the driver / operator is facing the correct direction, but the driver / operator’s face exhibits signs of health conditions such as stroke, heart attack, or whether the driver / operator is asleep.
[0188] In many other instances, the disclosed trained model may return indications of anomalies based on one or more captured images, including representations of: a pedestrian standing on top of a target vehicle (where the determined navigation action may include changing the course of the main vehicle or slowing it down); or cargo attached to the top of the target vehicle (e.g., a rocking mattress or a loose ladder) (where the determined navigation action may include changing the course of the main vehicle or slowing it down).
[0189] The disclosed embodiments can also return indications of normal objects / scenes, wherein the input image includes representations of recognized traffic sign types (e.g., "STOP" signs, "YIELD" signs, etc.) or representations of unrecognized traffic sign types (e.g., newly established traffic signs), but wherein the unrecognized traffic signs are close enough in the embedding space to be identified as non-abnormal. However, as described above, the disclosed trained model can return indications of anomalous objects / scenes, wherein the captured image includes representations of recognized traffic signs (e.g., STOP signs), but wherein the recognized traffic signs include graffiti markings.
[0190] Various types of actions can be taken in response to indications of detected anomalous objects or scenes. For example, the primary vehicle navigation system can use the output of a trained model as a trigger to cause the policy system to implement navigational changes (e.g., change the vehicle's heading, brake the primary vehicle, etc.). Alternatively or concurrently, the primary vehicle navigation system can also use the output of a trained model as a trigger to cause the policy system to change its operating mode, such as requiring a greater safety distance or margin, or requiring driver intervention. Additionally, indicators of anomalous situations (e.g., a rollover vehicle blocking a lane segment of a road) can trigger the primary vehicle system to automatically report the detected anomalous situation and its location to one or more authorities, traffic reporting systems, etc. (e.g., using REM map driving information acquisition technology and vehicle localization technology). The trained model can be used to filter training datasets (e.g., to identify anomalous data) or to identify potential anomalies in point clouds generated by LiDAR or radar systems, VIDAR depth maps (e.g., depth information of captured images generated pixel-by-pixel by a trained model), etc. The trained model can be used to enhance REM map driving information collection (e.g., by using an auxiliary collection system to identify anomalous objects or scenes that can be excluded from the collected driving information used to generate the REM map).
[0191] The disclosed system may offer several potential benefits. For example, the described trained model can eliminate the need to train the model on datasets that include edge cases (which may be numerous). By training on data representing objects / scenes considered "normal," the disclosed trained model can still correctly identify anomalous objects / scenes, even if the training data used to train the model does not include instances of such anomalous objects / scenes.
[0192] Figure 15 This illustrates an example method 1510 for navigating a master vehicle relative to a road segment. The method may include: receiving captured images obtained by a camera on the master vehicle (step 1520); generating a representation in an embedding space of at least a portion of the captured images (step 1530); determining whether the representation in the embedding space of at least a portion of the captured images falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region (step 1540); determining a navigation action for the master vehicle based on the determination that the representation in the embedding space of at least a portion of the captured images falls outside the predetermined embedding space region (step 1550); and causing at least one system associated with the master vehicle to perform the navigation action (step 1560).
[0193] The foregoing description has been presented for illustrative purposes. It is not exhaustive and is not limited to the precise forms or embodiments disclosed. Modifications and adaptations will be apparent to those skilled in the art in light of the practice of the specification and the disclosed embodiments. Additionally, while aspects of the disclosed embodiments are described as being stored in memory, those skilled in the art will understand that these aspects may also be stored on other types of computer-readable media (such as secondary storage devices, such as hard disks or CD-ROMs, or other forms of RAM or ROM, USB media, DVDs, Blu-rays, 4K Ultra HD Blu-rays) or other optical drive media.
[0194] Computer programs based on the written description and disclosed methods are within the skill level of an experienced developer. Various programs or program modules can be created using any techniques known to those skilled in the art, or can be designed in conjunction with existing software. For example, program segments or modules can be designed using or with the aid of the .NET Framework, .NET CompactFramework (and related languages such as Visual Basic, C, etc.), Java, C++, Objective-C, HTML, HTML / AJAX combination, XML, or HTML with included Java applets.
[0195] Furthermore, while exemplary embodiments have been described herein, those skilled in the art will understand based on this disclosure that the scope of any and all embodiments has equivalent elements, modifications, omissions, combinations (e.g., combinations across aspects of various embodiments), adaptations, and / or alterations. Limitations in the claims should be interpreted broadly based on the language used in the claims, and not limited to the examples described herein or during the application process. These examples should be interpreted as non-exclusive. Moreover, the steps of the disclosed methods can be modified in any way, including by reordering and / or inserting or deleting steps. Therefore, the specification and embodiments are considered to be illustrative only, wherein the true scope and spirit are expressed by the full scope of the appended claims and their equivalents.
[0196] The following are further examples.
[0197] Example 1. A device for navigating a master vehicle relative to a road segment, the device comprising:
[0198] Device for receiving captured images obtained by a camera on the main vehicle;
[0199] A means for generating a representation in an embedding space of at least a portion of the captured image;
[0200] A means for determining whether the representation of at least a portion of a captured image in an embedding space falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region.
[0201] A means for determining navigation actions of a master vehicle based on the determination that the representation of at least a portion of a captured image in an embedding space falls outside a predetermined embedding space region; and
[0202] A means for enabling at least one system associated with the main vehicle to perform navigation actions.
[0203] Example 2. The device according to Example 1, wherein the representation of at least a portion of the captured image in the embedding space is performed by a trained model.
[0204] Example 3. Based on the device in Example 2, where the trained model is a variational autoencoder.
[0205] Example 4. The device according to Example 2, wherein the trained model is further configured to output an indicator of whether the representation of at least a portion of the captured image in the embedding space falls outside a predetermined embedding space region, and wherein the determination of the navigation action of the master vehicle is based on the output of the trained model.
[0206] Example 5. The device according to Example 1, wherein the generation of representations in the embedding space is performed by a trained model configured to map at least a portion of the captured image to a probability distribution on latent variables associated with the embedding space.
[0207] Example 6. The device according to Example 1, wherein a predetermined embedding space region corresponds to a predetermined sub-region of a probability distribution.
[0208] Example 7. Based on the device in Example 6, where the probability distribution is a Gaussian distribution.
[0209] Example 8. The device according to Example 6, wherein the predefined sub-region is defined by at least one user-selectable parameter value.
[0210] Example 9. According to the device of Example 8, at least one user-selectable parameter value includes a percentage value associated with a probability distribution.
[0211] Example 10. The device according to Example 1, wherein the predetermined embedding space region is defined by an embedding space distance relative to an embedding space reference.
[0212] Example 11. The device according to Example 10, wherein the embedding space reference corresponds to the embedding space origin.
[0213] Example 12. The device according to Example 1, wherein the representation in the embedding space includes at least one feature vector.
[0214] Example 13. The device according to Example 1, wherein at least a portion of the captured image includes a representation of a road sign.
[0215] Example 14. The device according to Example 1, wherein at least a portion of the captured image includes a representation of the target vehicle.
[0216] Example 15. The device according to Example 1, wherein navigation actions include generating visual or auditory alarms.
[0217] Example 16. The device according to Example 1, wherein the navigation action includes changing the heading of the master vehicle.
[0218] Example 17. The device according to Example 1, wherein the navigation action includes slowing down the main vehicle.
[0219] Example 18. The device according to Example 1, wherein navigation actions include maintaining the speed and heading of the master vehicle.
[0220] Example 19. The device according to Example 1, wherein at least a portion of the captured image includes a representation of a first target vehicle being transported as cargo by a second target vehicle, wherein the first target vehicle and the second target vehicle are facing opposite directions, and wherein the determined navigation actions include maintaining the speed and heading of the master vehicle.
[0221] Example 20. The device according to Example 1, wherein at least a portion of the captured image includes a representation of a pedestrian standing on top of the target vehicle, and wherein the determined navigation action includes changing the course of the master vehicle or slowing down the master vehicle.
[0222] Example 21. The device according to Example 1, wherein at least a portion of the captured image includes a representation of cargo attached to the top of the target vehicle, and wherein the determined navigation action includes changing the course of the master vehicle or slowing down the master vehicle.
Claims
1. A system for navigating a master vehicle relative to a road segment, the system comprising: At least one processor, the at least one processor comprising a circuit system and a memory, wherein the memory includes instructions that, when executed by the circuit system, cause the at least one processor to: Receive images captured by cameras on the main vehicle; Generate a representation in the embedding space of at least a portion of the captured image; Determine whether the representation of at least a portion of the captured image in the embedding space falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region; The navigation action of the master vehicle is determined based on the determination that the representation of at least a portion of the captured image in the embedding space falls outside the predetermined embedding space region; and The navigation action is performed by at least one system associated with the main vehicle.
2. The system of claim 1, wherein the representation of at least a portion of the captured image in the embedding space is performed by a trained model.
3. The system according to claim 2, wherein the trained model is a variational autoencoder.
4. The system of claim 2, wherein the trained model is further configured to output an indicator of whether the representation of at least a portion of the captured image in the embedding space falls outside the predetermined embedding space region, and wherein the determination of the navigation action of the master vehicle is based on the output of the trained model.
5. The system of claim 1, wherein the generation of the representation in the embedding space is performed by a trained model configured to map the at least portion of the captured image to a probability distribution on latent variables associated with the embedding space.
6. The system according to claim 1, wherein the predetermined embedding space region corresponds to a predetermined sub-region of the probability distribution.
7. The system according to claim 6, wherein the probability distribution is a Gaussian distribution.
8. The system of claim 6, wherein the predetermined sub-region is defined by at least one user-selectable parameter value.
9. The system of claim 8, wherein the at least one user-selectable parameter value includes a percentage value associated with the probability distribution.
10. The system of claim 1, wherein the predetermined embedding space region is defined by an embedding space distance relative to an embedding space reference.
11. The system of claim 10, wherein the embedding space reference corresponds to the embedding space origin.
12. The system of claim 1, wherein the representation in the embedding space comprises at least one feature vector.
13. The system of claim 1, wherein the at least portion of the captured image includes a representation of road signs.
14. The system of claim 1, wherein the at least portion of the captured image includes a representation of the target vehicle.
15. The system of claim 1, wherein the navigation action includes generating a visual or auditory alarm.
16. The system of claim 1, wherein the navigation action includes changing the heading of the master vehicle.
17. The system of claim 1, wherein the navigation action includes decelerating the master vehicle.
18. The system of claim 1, wherein the navigation action includes maintaining the speed and heading of the master vehicle.
19. The system of claim 1, wherein the at least portion of the captured image includes a representation of a first target vehicle being transported as cargo by a second target vehicle, wherein the first target vehicle and the second target vehicle are facing opposite directions, and wherein the determined navigation action includes maintaining the speed and heading of the master vehicle.
20. The system of claim 1, wherein the at least portion of the captured image includes a representation of a pedestrian standing on top of the target vehicle, and wherein the determined navigation action includes changing the course of the master vehicle or slowing down the master vehicle.
21. The system of claim 1, wherein the at least portion of the captured image includes a representation of cargo attached to the top of the target vehicle, and wherein the determined navigation action includes changing the course of the master vehicle or slowing down the master vehicle.
22. A method for navigating a master vehicle relative to a road segment, the method comprising: Receive images captured by cameras on the main vehicle; Generate a representation in the embedding space of at least a portion of the captured image; Determine whether the representation of at least a portion of the captured image in the embedding space falls outside a predetermined embedding space region, wherein the predetermined embedding space region is defined as a non-abnormal embedding space region; The navigation action of the master vehicle is determined based on the determination that the representation of at least a portion of the captured image in the embedding space falls outside the predetermined embedding space region; as well as The navigation action is performed by at least one system associated with the main vehicle.