Estimating object attributes using visual image data

A machine learning model using visual data and auxiliary sensor correlations in autonomous vehicles addresses sensor optimization challenges, enhancing accuracy and reducing complexity and cost by predicting object distances without dedicated sensors.

JP7777620B2Active Publication Date: 2025-11-28TESLA INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024043716
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-02-19
Filing Date
2024-03-19
Publication Date
2025-11-28
Estimated Expiration
2040-02-07

AI Technical Summary

Technical Problem

Autonomous driving systems face challenges in optimizing sensor configurations to reduce complexity and cost while maintaining accurate environmental description, as additional sensors increase system complexity, cost, and bandwidth requirements.

Method used

A system using a vehicle camera and a trained machine learning model, correlated with auxiliary sensor data, to determine object distances, reducing the need for dedicated distance measurement sensors by leveraging visual data and auxiliary sensor correlations.

Benefits of technology

Accurately predicts object distances and attributes using visual data, potentially eliminating the need for expensive auxiliary sensors, thereby reducing system complexity, cost, and bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007777620000001
    Figure 0007777620000001
  • Figure 0007777620000002
    Figure 0007777620000002
  • Figure 0007777620000003
    Figure 0007777620000003
Patent Text Reader

Abstract

To provide a system that utilizes image data as a basis of input data to a trained machine learning model to at least in part identify a distance from a vehicle to a certain object.SOLUTION: A system comprises: multiple processors and a memory coupled to one or more processors. The system is configured to: receive image data based on an image captured using a camera of a vehicle; and utilize the image data as a basis of input data to a trained machine learning model to at least in part identify a distance from the vehicle to a certain object. The trained machine learning model is trained using a training image and a correlated output of a distance measurement sensor with irradiation.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a continuation of and claims priority to U.S. patent application Ser. No. 16 / 279,657, filed Feb. 19, 2019, entitled "Estimating Object Attributes Using Visual Image Data," the disclosure of which is incorporated herein by reference in its entirety. [Background technology]

[0002] Autonomous driving systems typically rely on a large number of sensors onboard vehicles, including a collection of vision and distance-based sensors (e.g., radar, lidar, ultrasonic, etc.). Data captured by each sensor is collected to help the vehicle understand its surroundings and determine how to control the vehicle. Vision sensors can be used to identify objects from captured image data, and distance-based sensors can be used to measure the distance to the detected objects. Steering and speed adjustments can be made based on obstacle detection and a possible obstacle-free route. However, as the number and type of sensors increase, the complexity and cost of the system also increase. For example, distance-based sensors such as lidar are often expensive to install in mass-produced vehicles. Furthermore, each additional sensor increases the input bandwidth requirements of the autonomous driving system. Therefore, an optimal configuration of sensors on a vehicle must be found. This configuration must limit the total number of sensors while accurately describing the surrounding environment and not limiting the amount and type of data captured to safely control the vehicle. Summary of the Invention [Means for solving the problem]

[0003] One embodiment includes a system comprising one or more processors configured to receive image data based on images captured using a vehicle camera and use the image data as a basis of input data for a trained machine learning model to at least partially determine a distance from the vehicle to an object, the trained machine learning model being trained using the training images and a correlation output of an illumination distance measuring sensor, and a memory coupled to the one or more processors.

[0004] Another embodiment includes a computer program product, the computer program product embodied in a non-transitory computer-readable storage medium, comprising computer instructions for receiving image data based on images captured using a camera of the vehicle and using the image data as a basis of input data for a trained machine learning model used to at least partially determine a distance from the vehicle to an object, the trained machine learning model having been trained using training images and a correlation output of an illumination distance measuring sensor.

[0005] Yet another embodiment includes a method including receiving selected images based on images captured using a vehicle camera, receiving distance data based on an illumination distance measurement sensor of the vehicle, using the selected images as input data to a trained machine learning model to identify an object, extracting distance estimates for the identified object from the received distance data, and annotating the selected images with the extracted distance estimates to generate an image of the training image. creating an image of the vehicle, training a second machine learning model to predict distance measurements using a training dataset that includes the training images, and providing the trained second machine learning model to a second vehicle equipped with a second camera. [Brief explanation of the drawings]

[0006] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings.

[0007] [Figure 1] FIG. 1 is a block diagram illustrating one embodiment of a deep learning system for use in autonomous driving.

[0008] [Figure 2] FIG. 1 is a flow diagram illustrating one embodiment of a process for creating training data for predicting object attributes.

[0009] [Figure 3] FIG. 1 is a flow diagram illustrating one embodiment of a process for training and applying machine learning models for use in autonomous driving.

[0010] [Figure 4] FIG. 1 is a flow diagram illustrating one embodiment of a process for training and applying machine learning models for use in autonomous driving.

[0011] [Figure 5] FIG. 1 illustrates an example of capturing auxiliary sensor data for training a machine learning network.

[0012] [Figure 6] FIG. 10 is a diagram illustrating an example of predicting an object attribute. DETAILED DESCRIPTION OF THE INVENTION

[0013] The present invention may be implemented in numerous ways, including as a process, such as an apparatus, a system, a composition of matter, a computer program product embodied on a computer-readable storage medium, and / or a processor, e.g., a processor configured to execute instructions stored in and / or provided by a memory coupled thereto. These implementations, or any other form the present invention may take, may be referred to herein as techniques. The order of steps in disclosed processes may be varied generally within the scope of the present invention. Unless otherwise specified, components, such as processors or memories, described as being configured to perform a task may be implemented as general-purpose components temporarily configured to perform the task at a given time, or as specialized components manufactured to perform the task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0014] A detailed description of one or more embodiments of the present invention follows, along with accompanying drawings that illustrate the principles of the invention. While the present invention has been described in connection with such embodiments, the invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description to provide a thorough understanding of the present invention. These details are provided for the purpose of example, and the present invention may be practiced according to the claims without some or all of these specific details. For purposes of clarity, technical material known in the art related to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.

[0015] Machine learning training techniques to generate highly accurate machine learning results from visual data The present disclosure discloses a method for predicting object attributes, such as object distance, by associating auxiliary data with objects identified from visual data using auxiliary sensor data, such as radar or lidar results. In various embodiments, the collection of auxiliary data and its association with visual data is automated, requiring little, if any, human intervention. For example, objects identified using vision techniques do not need to be manually labeled, greatly improving the efficiency of machine learning training. Instead, training data is automatically generated and can be used to train a machine learning model to predict object attributes with high accuracy. For example, this data can be automatically collected from a fleet of vehicles by collecting snapshots of associated data, such as visual and radar data. In some embodiments, only a subset of associated targets for the visual and radar data are sampled. Fusion data collected from these fleets of vehicles is automatically collected, and a neural net is trained using this fusion data to mimic the captured data. This trained machine learning model can then be deployed to vehicles to accurately predict object attributes, such as distance, direction, and speed, using only visual data. For example, once a machine learning model is trained to measure object distances using camera images without the need for a dedicated distance measurement sensor, it may no longer be necessary to equip an autonomous vehicle with a dedicated distance measurement sensor. When used with a dedicated distance measurement sensor, the machine learning model may be used as a redundant or secondary distance data source to improve accuracy and / or provide fault tolerance. The identified objects and corresponding attributes may be used to perform autonomous driving functions, such as autonomous driving of a vehicle or driver assistance operations. For example, an autonomous vehicle may be controlled to avoid a merging vehicle identified using the disclosed techniques.

[0016] A system including one or more processors coupled to a memory is configured to receive image data based on images captured using a vehicle camera. For example, a processor, such as an artificial intelligence (AI) processor onboard an autonomous vehicle, receives image data from a camera, such as a front-facing camera on the vehicle. Other cameras, such as a side-facing camera or a rear-facing camera, may also be used. The image data is used as input data to a machine learning model trained to at least partially determine the distance of an object from the vehicle. For example, the captured image is used as input data to a machine learning model, such as a deep learning network model running on the AI ​​processor. The model is used to predict the distance to an object identified in the image data. Surrounding objects, such as vehicles and pedestrians, may be identified from the image data, and their accuracy and direction may be inferred using a deep learning system. In various embodiments, the trained machine learning model is trained using training images and a correlation output of an illumination ranging sensor. The illumination ranging sensor may emit a signal (e.g., a radio signal, an ultrasonic signal, an optical signal, etc.) upon detecting the distance of an object from the sensor. For example, a radar sensor mounted on a vehicle transmits radar to determine the distance and direction to surrounding obstacles. These distances are then correlated with identified objects in training images captured from the vehicle's camera. The associated training images are annotated with the distance measurements and used to train a machine learning model. In some embodiments, the model is used to predict additional attributes, such as the object's velocity. For example, the object's velocity, as measured by radar, is associated with the object in the training images when training a machine learning model to predict the object's velocity and direction.

[0017] In some embodiments, a vehicle captures the vehicle's surrounding environment and vehicle operating parameters. The vehicle includes sensors for capturing and analyzing vehicle motion data. The captured data includes visual data (e.g., video and / or still images) and additional auxiliary data, such as radar sensor data, lidar sensor data, inertial sensor data, audio sensor data, odometry sensor data, position sensor data, and / or other forms of sensor data. For example, the sensor data may capture vehicles, pedestrians, lane markings, vehicle traffic, obstacles, traffic signs, traffic sounds, etc. Odometry sensors and other similar sensors capture vehicle operation parameters such as vehicle speed, steering, heading, direction change, position change, altitude change, and speed change. The captured visual and auxiliary data are transmitted from the vehicle to a training server to create a training dataset. In some embodiments, the transmitted visual and auxiliary data are correlated and used to automatically generate training data. This training data is used to train a machine learning model to generate highly accurate machine learning results. In some embodiments, a time series of captured data is used to generate the training data. A ground truth is designated based on a group of time series elements, and at least one of the elements, such as a single image, from the group is annotated using the ground truth. For example, a series of image and radar data are captured at time intervals, such as 30 seconds. Vehicles identified from the image data and tracked over the time series are associated with corresponding radar ranges and radar directions from the time series. Associated auxiliary data, such as radar range data, is associated with the vehicles after analyzing the image data and range data captured over the time series. By analyzing the image data and auxiliary data over the time series, ambiguities such as multiple objects having similar distances can be resolved with high accuracy to designate the ground truth. For example, when using only a single captured image, there may be insufficient corresponding radar data to accurately estimate the individual distances to the two vehicles when one vehicle obscures another or when two vehicles are approaching each other.However, by tracking the vehicles over the time series, the distances determined by the radar can be properly associated with the correct vehicle even though the vehicles may be moving away from each other, in different directions, and / or at different speeds, etc. In various embodiments, once the auxiliary data has been properly associated with an object, one or more images from the time series are converted into training images, and the images are annotated with the corresponding ground truth, such as distance, speed, and / or other suitable object attributes.

[0018] In various embodiments, a machine learning model trained using auxiliary sensor data can accurately predict the results of an auxiliary sensor without requiring a physical auxiliary sensor. For example, a training vehicle can be equipped with auxiliary sensors, including sensors that are expensive and / or difficult to operate, to collect training data. This training data can then be used to train a machine learning model to predict the results of an auxiliary sensor, such as a radar sensor, a lidar sensor, or another sensor. This trained model is then deployed to vehicles, such as production vehicles, that require only a visual sensor. Auxiliary sensors are not required, but can also be used as a secondary data source. There are many benefits to reducing the number of sensors, including, among other things, the difficulty of sensor recalibration, sensor maintenance, the cost of adding sensors, and / or the additional bandwidth and computational requirements for additional sensors. In some embodiments, the trained model is used in the event of an auxiliary sensor failure. Instead of relying on additional auxiliary sensors, this trained machine learning model uses input data from one or more visual sensors to predict the results of the auxiliary sensor. The results predicted here can be used to perform automated driving functions that require the detection of objects (e.g., pedestrians, stationary vehicles, moving vehicles, curbs, obstacles, road barriers, etc.) and their distance and direction. The results predicted here can be used to detect the distance and direction to traffic control objects such as traffic lights, traffic signs, road markings, etc. Although the previous examples use visual sensors and object distance, alternative sensors and predicted attributes can be used as well.

[0019] 1 is a block diagram illustrating one embodiment of a deep learning system for use in automated driving. The deep learning system includes various components that can be used together to perform automated driving and / or driver assistance operations of a vehicle and to collect and process data for training machine learning models. In various embodiments, the deep learning system is installed in a vehicle, and data captured from the vehicle can be used to train and improve deep learning systems in the vehicle or other similar vehicles. The deep learning system may be used to perform automated driving functions, including identifying objects and predicting object attributes, such as distance and direction, using visual data as input data.

[0020] In the illustrated example, deep learning system 100 is a deep learning network that includes visual sensor 101, additional sensor 103, image preprocessor 105, deep learning network 107, artificial intelligence (AI) processor 109, vehicle control module 111, and network interface 113. In various embodiments, the various components are communicatively coupled. For example, image data captured from visual sensor 101 is provided to image preprocessor 105. Sensor data processed in image preprocessor 105 is provided to deep learning network 107 running on AI processor 109. In some embodiments, sensor data from additional sensor 103 is used as input data to deep learning network 107. Output data of deep learning network 107 running on AI processor 109 is provided to vehicle control module 111. In various embodiments, vehicle control module 111 is connected to and controls vehicle operations, such as vehicle speed, braking, and / or steering. In various embodiments, sensor data and / or machine learning results may be transmitted to a remote server (not shown) via the network interface 113. For example, sensor data, such as data captured from the visual sensor 101 and / or the additional sensor 103, may be transmitted to a remote training server via the network interface 113 to collect training data for enhancing vehicle performance, comfort, and / or safety. In various embodiments, the network interface 113 is used to communicate with the remote server, make calls, send and / or receive text messages, transmit sensor data based on vehicle operation, and the like. In some embodiments, the deep learning system 100 may include additional or fewer components as needed. For example, in some embodiments, the image preprocessor 105 is an optional component. As another example, in some embodiments, a post-processing component (not shown) is used to perform post-processing on the output data of the deep learning network 107 before the output data is provided to the vehicle control module 111.

[0021] In some embodiments, the visual sensor 101 includes one or more camera sensors for capturing image data. In various embodiments, the visual sensor 101 may be mounted on a vehicle at various locations and / or oriented in one or more different directions. For example, the visual sensor 101 may be mounted on the front, side, rear, and / or roof of the vehicle in a forward-facing, rear-facing, side-facing, or other orientation. In some embodiments, the visual sensor 101 may be an image sensor, such as a high-dynamic range camera and / or a camera with a variety of fields of view. For example, in some embodiments, eight surround cameras are mounted on the vehicle, providing a 360-degree view around the vehicle with a range of up to 250 meters. In some embodiments, the camera sensors include a wide-angle forward camera, a narrow-angle forward camera, a rear-view camera, a forward-view side camera, and / or a rear-view side camera.

[0022] In some embodiments, the visual sensor 101 is a vehicle control module 111. The visual sensor 101 may not be mounted on either the vehicle or the surrounding vehicle. For example, the visual sensor 101 may be mounted on a surrounding vehicle and / or on the road or surrounding environment, and may even be part of a deep learning system for capturing sensor data. In various embodiments, the visual sensor 101 includes one or more cameras that capture the vehicle's environment, including the road along which the vehicle is traveling. For example, one or more forward-facing cameras and / or pillar cameras capture images of objects in the environment surrounding the vehicle, such as vehicles, pedestrians, traffic control objects, roads, curbs, and obstacles. As another example, the cameras capture a time series of image data that includes image data of surrounding vehicles, including vehicles attempting to cross into the lane in which the vehicle is traveling. The visual sensor 101 may include an image sensor capable of capturing still images and / or video. The data may be captured over a period of time, such as a series of captures over a period of time, and may be synchronized with other vehicle data, including other sensor data. For example, image data used to identify an object may be captured along with radar and odometry data over a 15 second period or another suitable period.

[0023] In some embodiments, the additional sensors 103 include additional sensors in addition to the visual sensors 101 for capturing sensor data. In various embodiments, the additional sensors 103 may be mounted on a vehicle at various locations and / or oriented in one or more different directions. For example, the additional sensors 103 may be mounted on the front, sides, rear, and / or roof of the vehicle in a forward-facing, rear-facing, side-facing, or other orientation. In some embodiments, the additional sensors 103 may be radar sensors, ultrasonic sensors, and / or illumination sensors such as lidar sensors. In some embodiments, the additional sensors 103 include non-visual sensors. The additional sensors 103 may include radar sensors, acoustic sensors, lidar sensors, inertial sensors, odometry sensors, position sensors, and / or ultrasonic sensors. For example, by mounting 12 ultrasonic sensors on a vehicle, both hard and soft objects may be detected. In some embodiments, forward-facing radar is used to capture data about the surrounding environment. In various embodiments, the radar sensor is able to capture details of the surrounding environment despite heavy rain, fog, dust accumulation, and the proximity of other vehicles.

[0024] In some embodiments, the additional sensors 103 are not mounted on the vehicle with the vehicle control module 111. Similar to the visual sensors 101, for example, the additional sensors 103 may be mounted on surrounding vehicles and / or on the road or surrounding environment, and may even be included as part of a deep learning system for capturing sensor data. In some embodiments, the additional sensors 103 include one or more cameras that capture the vehicle's surrounding environment, including the road on which the vehicle is traveling. For example, a forward-facing radar sensor captures distance data to objects in the vehicle's forward field of view. The additional sensors may capture odometry information, including information about the vehicle's trajectory, location information, and / or vehicle control information. The sensor data may be captured over a period of time, such as a series of captures over a period of time, and may be associated with image data captured from the visual sensors 101. In some embodiments, the additional sensors 103 include a location sensor, such as a global positioning system (GPS) sensor, for determining the vehicle's location and / or changes in location. In various embodiments, one or more of the additional sensors 103 are optional and are only included in vehicles designed to capture training data. Vehicles without one or more of the additional sensors 103 can simulate the results of the additional sensors 103 by predicting outputs using trained machine learning models and techniques disclosed herein. For example, a vehicle without a forward-facing radar sensor or a forward-facing lidar sensor can use image data to predict the results of the optional sensors by applying trained machine learning models, such as models from deep learning network 107.

[0025] In some embodiments, an image preprocessor 105 is used to preprocess the sensor data of the visual sensor 101. For example, the image preprocessor 105 may be used to preprocess the sensor data, to split the sensor data into one or more components, and / or to post-process the one or more components. In some embodiments, the image preprocessor 105 may be implemented using a graphics processing unit (GPU), a central processing unit (CPU), or a processor. processing unit), image signal processor, or dedicated image processor. In various embodiments, the image preprocessor 105 is a tone mapper processor that processes high dynamic range data. In some embodiments, the image preprocessor 105 is implemented as part of the artificial intelligence (AI) processor 109. For example, the image preprocessor 105 may be a component of the AI ​​processor 109. In some embodiments, the image preprocessor 105 may be used to normalize or transform an image. For example, an image captured with a fisheye lens may be warped, and the image preprocessor 105 may be used to transform the image to remove or correct this warping. In some embodiments, noise, distortion, and / or blur are removed or reduced during the preprocessing step. In various embodiments, the image is adjusted or normalized to improve the results of the machine learning analysis. For example, the white balance of the image is adjusted to account for different lighting operating conditions, such as daylight conditions, sunny conditions, cloudy conditions, twilight conditions, sunrise conditions, sunset conditions, and nighttime conditions.

[0026] In some embodiments, deep learning network 107 is a deep learning network used to identify vehicle control parameters, including analyzing the driving environment to identify objects and their corresponding attributes, such as distance, speed, or another suitable parameter. For example, deep learning network 107 may be an artificial neural network, such as a convolutional neural network (CNN), trained with input data, such as sensor data, and whose output data is provided to vehicle control module 111. As one example, this output data may include at least distance estimates of detected objects. As another example, this output data may include at least potential vehicles that may merge into the vehicle's lane, the distance between these vehicles, and the speed of these vehicles. In some embodiments, deep learning network 107 receives at least image sensor data as input data, identifies objects in the image sensor data, and predicts distances to the objects. Additional input data may include scene data describing the environment around the vehicle and / or vehicle specifications, such as the vehicle's operating characteristics. The scene data may include scene tags describing the environment around the vehicle, such as rainy roads, wet roads, snowy traffic, muddy traffic, high-density traffic, highway areas, urban areas, school routes, etc. In some embodiments, the output data of the deep learning network 107 is a three-dimensional representation of the vehicle's environment, including cuboids representing objects, such as the identified object. In some embodiments, the output data of the deep learning network 107 is used for automated driving, including navigating the vehicle towards a target destination.

[0027] In some embodiments, the artificial intelligence (AI) processor 109 is a hardware processor for running the deep learning network 107. In some embodiments, the AI ​​processor 109 is a dedicated AI processor for performing inference on the sensor data using a convolutional neural network (CNN). The AI ​​processor 109 may be optimized for the bit depth of the sensor data. In some embodiments, the AI ​​processor 109 is optimized for deep learning operations, such as neural network operations including, among others, convolution, dot product, vector, and / or matrix operations. In some embodiments, the AI ​​processor 109 is a graphics processing unit (GPU). In various embodiments, the AI ​​processor 109 is implemented using a deep learning processing unit (DSM). In various embodiments, the AI ​​processor 109 is coupled to a memory configured to provide the AI ​​processor with instructions that, when executed, cause the AI ​​processor to perform deep learning analysis of received input sensor data and determine machine learning results, such as object distance, for use in autonomous driving. In some embodiments, the AI ​​processor 109 is used to process the sensor data to make the data usable as training data.

[0028] In some embodiments, vehicle control module 111 is used to process output data from artificial intelligence (AI) processor 109 and convert the output data into vehicle control calculations. In some embodiments, vehicle control module 111 is used to control a vehicle for autonomous driving. In various embodiments, vehicle control module 111 can regulate vehicle speed, acceleration, steering, braking, etc. For example, in some embodiments, vehicle control module 111 is used to control a vehicle, for example, to maintain the vehicle's position within a lane, merge the vehicle into another lane, and adjust the vehicle's speed or lane positioning to account for merging vehicles.

[0029] In some embodiments, the vehicle control module 111 is used to control vehicle lighting, such as brake lights, turn signals, and headlights. In some embodiments, the vehicle control module 111 is used to control vehicle acoustics, such as the vehicle's sound system, playing audio alerts, enabling the microphone, and enabling the horn. In some embodiments, the vehicle control module 111 is used to control notification systems, including warning systems for notifying the driver and / or passengers of driving events, such as a potential collision or approaching a planned destination. In some embodiments, the vehicle control module 111 is used to adjust sensors, such as the visual sensor 101 and additional sensors 103, of a vehicle. For example, the vehicle control module 111 may be used to change parameters of one or more sensors, such as changing the orientation, changing the output resolution and / or format type, increasing or decreasing the capture rate, adjusting the captured dynamic range, adjusting the focus of a camera, enabling and / or disabling sensors, etc. In some embodiments, the vehicle control module 111 may be used to change parameters of the image preprocessor 105, such as changing the frequency range of a filter, adjusting feature detection parameters and / or edge detection parameters, adjusting channels and bit depth, etc. In various embodiments, vehicle control module 111 is used to perform automated driving and / or driver assistance control of a vehicle. In some embodiments, vehicle control module 111 is implemented using a processor coupled with memory. In some embodiments, vehicle control module 111 is implemented using an application specific integrated circuit (ASIC), a programmable logic device (PLD), or other suitable processing hardware.

[0030] In some embodiments, network interface 113 is a communications interface for transmitting and / or receiving data, including training data. In various embodiments, network interface 113 includes a cellular or wireless interface for interfacing with a remote server to transmit sensor data, transmit potential training data, receive updates to the deep learning network, including updated machine learning models, connect to make voice calls, send and / or receive text messages, etc. For example, network interface 113 may be used to transmit captured sensor data for use as potential training data to a remote training server for training a machine learning model. As another example, network interface 113 may be used to transmit commands to vision sensor 101, additional sensors 103, image preprocessor 105, deep learning network 107, AI processor 109, and / or vehicle control module 111. Updates to instructions and / or operational parameters may be received. The machine learning models of the deep learning network 107 may be updated using the network interface 113. As another example, the network interface 113 may be used to update firmware of the vision sensor 101 and additional sensors 103 and / or operational parameters of the image pre-processor 105, such as image processing parameters.

[0031] FIG. 2 is a flow diagram illustrating one embodiment of a process for creating training data for predicting object attributes. For example, image data is annotated with sensor data from additional auxiliary sensors to automatically create training data. In some embodiments, time series elements consisting of sensor data and associated auxiliary data are collected from a vehicle, and the training data is automatically created using the time series elements. In various embodiments, the training data is automatically labeled with corresponding ground truth using the process of FIG. 2. The ground truth and image data are packaged as training data for predicting object attributes identified from the image data. In various embodiments, the sensor data and associated auxiliary data are captured using the deep learning system of FIG. 1. For example, in various embodiments, the sensor data is captured from the visual sensor 101 of FIG. 1, and associated data is captured from the additional sensor 103 of FIG. 1. In some embodiments, the process of FIG. 2 is performed to automatically collect data when existing predictions are inaccurate or can be improved. For example, predictions are made by an autonomous vehicle to identify one or more object attributes, such as distance or direction, from visual data. The prediction is compared to distance data received from an illumination distance measurement sensor. A determination may be made whether the prediction is within an acceptable accuracy threshold. In some embodiments, a determination is made that the prediction can be improved. If the prediction is not sufficiently accurate, the process of FIG. 2 may be applied to the prediction scenario to create a curated set of training examples to improve the machine learning model.

[0032] At 201, visual data is received. This visual data may be image data, such as video and / or still images. In various embodiments, this visual data is captured at a vehicle and then transmitted to a training server. This visual data may be captured over a period of time to create time series elements. In various embodiments, these elements include timestamps to preserve the order of the elements. By capturing time series elements, objects in the time series are tracked over time, making it easier to distinguish objects that are difficult to identify from a single input sample, such as a single input image and corresponding associated data. For example, a pair of oncoming headlights may initially appear to both belong to one vehicle, but as the headlights move away from each other, each headlight is recognized as belonging to a separate motorcycle. In some scenarios, objects in the image data are more distinguishable than objects in the associated auxiliary data received at 203. For example, it may be difficult to determine the estimated distance of a van from the wall it is next to using distance data alone. However, by tracking the van across the corresponding time series of image data, the correct distance data may be associated with the identified van. In various embodiments, the sensor data captured as a time series is captured in a format that the machine learning model uses as input data, for example, the sensor data may be raw image data or processed image data.

[0033] In various embodiments, when time series data is received, the time series may be organized by associating a timestamp with each element of the time series. For example, a timestamp may be associated with at least the first element in the time series. This timestamp may be used to calibrate the time series element with the associated data, such as the data received at 203. In various embodiments, the length of the time series may be 10 seconds, 3 The time series may be a fixed length of time, such as 0 seconds, or another suitable length. This length of time may be configurable. In various embodiments, these time series may be based on the vehicle speed, such as the vehicle's average speed. For example, at slower speeds, increasing the time length of the time series may capture data over a longer distance traveled than would be possible if a shorter length of time were used for the same speed. In some embodiments, the number of elements in these time series is configurable. The number of elements may be based on the distance traveled. For example, a vehicle moving faster over a given period of time will include more elements in the time series than a vehicle moving slower. Adding these elements can increase the fidelity of the captured surrounding environment and improve the accuracy of the predicted machine learning results. In various embodiments, the number of elements is adjusted by adjusting the frames per second at which the sensors capture data and / or by discarding unnecessary intermediate frames.

[0034] At 203, data related to the received visual data is received. In various embodiments, this related data is received at the training server along with the visual data received at 201. In some embodiments, this related data is sensor data received from additional sensors on the vehicle, such as ultrasonic sensors, radar sensors, lidar sensors, or other suitable sensors. This related data may be captured by the additional sensors on the vehicle, such as data related to distance, direction, speed, position, orientation, change in position, change in orientation, and / or other related data. This related data may be used to specify ground truth for features identified in the visual data received at 201. For example, distance and direction measurements from a radar sensor may be used to determine object distance and object direction for objects identified in the visual data. In some embodiments, the related data received here is time series data corresponding to the time series of visual data received at 201.

[0035] In some embodiments, data associated with this visual data includes map data. For example, offline data such as road-level and / or satellite-level map data may be received at 203. This map data may be used to identify features such as roads, lanes, intersections, speed limits, school districts, etc. For example, this map data may describe lane paths. An estimated position of an identified vehicle within a lane may be used to determine or corroborate an estimated distance to a detected vehicle. As another example, this map data may describe speed limits associated with various roads on the map. In some embodiments, this speed limit data may be used to validate a speed vector of an identified vehicle.

[0036] At 205, objects are identified within the visual data. In some embodiments, the visual data is used as input data to identify objects within the vehicle's environment. For example, vehicles, pedestrians, obstacles, etc. are identified from the visual data. In some embodiments, these objects are identified using a deep learning system with trained machine learning models. In various embodiments, a bounding box is created for the identified object. This bounding box may be a two-dimensional bounding box, such as a rectangular prism, or a three-dimensional bounding box that outlines the outer contour of the identified object. In some embodiments, additional data, such as the data received at 203, is used to assist in identifying the object. This additional data may be used to improve accuracy in identifying the object.

[0037] At 207, a ground truth is determined for the identified object. At 205, a ground truth is determined for the identified object from the visual data received at 201 using the associated data received at 203. In some embodiments, this The associated data is depth (and / or distance) data for the identified object. By associating distance data with the identified object, a machine learning model can be trained and object distance can be estimated by using the associated distance data as ground truth for the detected object. In some embodiments, these distances are distances to detected objects such as obstacles, barriers, moving vehicles, stationary vehicles, traffic signals, pedestrians, etc., and are used as ground truth for training. In addition to distance, ground truth for other object parameters such as direction, velocity, and acceleration may also be specified. For example, precise distance and direction are specified as ground truth for the identified object. As another example, precise velocity vectors are specified as ground truth for identified objects such as vehicles and pedestrians.

[0038] In various embodiments, the visual data and the associated data are organized by timestamp, and corresponding timestamps are used to synchronize the two data sets. In some embodiments, timestamps are used to synchronize time series data, such as a series of images and a corresponding series of associated data. This data may be synchronized at the time of capture. For example, as each element of the time series is captured, a corresponding set of associated data is captured and stored with the time series element. In various embodiments, the time interval of this associated data is configurable and / or matches the time interval of the time series element. In some embodiments, this associated data is sampled at the same rate as the time series element.

[0039] In various embodiments, ground truth may be specified solely by examining the time series data. For example, analyzing only a subset of the visual data may result in misidentification of objects and / or their attributes. Expanding the analysis to the entire time series removes ambiguity. For example, the presence of an occluded vehicle may be revealed back and forth in the time series. Once identified, this occasionally occluded vehicle may be tracked throughout the time series, even when occluded. Similarly, by associating object attributes from related data with objects identified in the visual data, the object attributes of this occasionally occluded vehicle may be tracked throughout the time series. In some embodiments, this data may be played backward (and / or forward) to identify any ambiguities in correlating the related data with the visual data. The state of an object at various times in the time series may be used to assist in identifying the object attributes of the object throughout the time series.

[0040] In various embodiments, a threshold is used to determine whether an object attribute is associated as ground truth for an identified object. For example, associated data with a high degree of confidence is associated with the identified object, while associated data with a confidence below the threshold is not associated with the identified object. In some embodiments, this associated data may be conflicting sensor data. For example, ultrasound data output may be conflicting with radar data output. As another example, range data may be conflicting with map data. Range data may estimate that a school district begins within 30 meters, while information from map data may state that the same school district begins within 20 meters. If the associated data has a low confidence, it may be discarded and not used to specify the ground truth.

[0041] In some embodiments, this ground truth is specified to predict a semantic label. For example, a detected vehicle may be labeled as being in the left lane or the right lane based on the predicted distance and direction. In some embodiments, the detected vehicle may be labeled as being in a blind spot, or as being a vehicle that should be given priority, or with another appropriate semantic label. In some embodiments, vehicles are assigned to roads or lanes in a map based on the designated ground truth. As another example, this designated ground truth may be used to label traffic lights, lanes, drivable spaces, or other features that aid in autonomous driving.

[0042] At 209, training data is packaged. For example, elements of the visual data received at 201 are selected and associated with the specified ground truth at 207. In some embodiments, the selected elements are time series elements. The selected elements represent sensor data inputs to the machine learning model, such as training images, and the ground truth represents the predicted results. In various embodiments, the selected data is annotated and created as training data. In some embodiments, the training data is packaged into training, validation, and testing data. The training data is packaged based on the specified ground truth and the selected training elements to train a machine learning model to predict results associated with one or more associated auxiliary sensors. For example, the trained model can be used to accurately predict the distance and direction to an object, achieving results similar to measurements using sensors such as radar or lidar sensors. In various embodiments, the machine learning results are used to implement functionality used in autonomous driving. The packaged training data is then available for use in training the machine learning model.

[0043] FIG. 3 is a flow diagram illustrating one embodiment of a process for training and applying a machine learning model for use in autonomous driving. For example, input data including primary and secondary sensor data is received and processed to create training data for training the machine learning model. In some embodiments, the primary sensor data corresponds to image data captured by the autonomous driving system, and the secondary sensor data corresponds to sensor data captured from an illumination-based distance measurement sensor. The secondary sensor data may be used to annotate the primary sensor data to train the machine learning model to predict an output based on the secondary sensor. In some embodiments, the sensor data corresponds to sensor data captured based on a specific use case, such as when a user manually disengages autonomous driving or when distance estimates from the visual data differ significantly from distance estimates from the secondary sensor. In some embodiments, the primary sensor data is sensor data from visual sensor 101 of FIG. 1, and the secondary sensor data is sensor data from one or more sensors in additional sensors 103 of FIG. 1. In some embodiments, this process is used to create and deploy a machine learning model for use in deep learning system 100 of FIG. 1.

[0044] At 301, training data is created. In some embodiments, a training data set is created by receiving sensor data including image data and auxiliary data. This image data may include still images and / or video from one or more cameras. Additional sensors, such as radar, lidar, or ultrasonic sensors, may be used to provide associated auxiliary sensor data. In various embodiments, this image data is paired with corresponding auxiliary data to aid in identifying attributes of objects detected in the sensor data. For example, distance and / or velocity data obtained from the auxiliary data may be used to accurately estimate the distance and / or velocity to objects identified in the image data. In some embodiments, this sensor data is a time series element, and this data is used to specify a ground truth. A group ground truth is then associated with a subset of the time series, such as a frame of image data. Selected time series elements and this ground truth are used to create training data. In some embodiments, this training data may include objects such as vehicles, pedestrians, obstacles, or other objects. The training data generated here may include data used for training, validation, and testing. In various embodiments, the sensor data may be in different formats. For example, the sensor data may be still image data, video data, radar data, ultrasonic data, audio data, position data, odometry data, etc. The odometry data may include vehicle operation parameters such as applied acceleration, applied braking, applied steering, vehicle position, vehicle heading, changes in vehicle position, and changes in vehicle heading. In various embodiments, the training data is curated and annotated to create a training dataset. In some embodiments, some of the training data creation work may be performed by human curators. In various embodiments, some of the training data is automatically generated from data captured from the vehicle, significantly reducing the effort and time required to build a robust training dataset. In some embodiments, the format of the data is compatible with the machine learning model used in the deployed deep learning application. In various embodiments, this training data includes validation data for testing the accuracy of the trained model. In some embodiments, the process of Figure 2 is performed at 301 in Figure 3.

[0045] At 303, a machine learning model is trained. For example, the machine learning model is trained using the data created at 301. In some embodiments, the model is a neural network, such as a convolutional neural network (CNN). In various embodiments, the model includes multiple hidden layers. In some embodiments, the neural network may include multiple layers, including multiple convolutional layers and pooling layers. In some embodiments, the training model is validated using a validation dataset created from the received sensor data. In some embodiments, the machine learning model is trained to predict the output of a sensor, such as a distance and illumination measurement sensor, from a single input image. For example, distance and direction attributes of an object can be inferred from a single image captured from a camera. As another example, the velocity vector of a surrounding vehicle, including whether the vehicle is attempting to merge, is predicted from a single image captured from a camera.

[0046] At 305, the trained machine learning model is deployed. For example, the trained machine learning model is installed on the vehicle as an update to a deep learning network, such as deep learning network 107 of FIG. 1. In some embodiments, the newly trained machine learning model is installed using an over-the-air update. For example, the over-the-air update may be received through a network interface of the vehicle, such as network interface 113 of FIG. 1. In some embodiments, the update is a firmware update transmitted using a wireless network, such as a WiFi network or a cellular network. In some embodiments, the new machine learning model may be installed when the vehicle is serviced.

[0047] At 307, sensor data is received. For example, the sensor data is captured from one or more sensors in the vehicle. In some embodiments, the sensor is the visual sensor 101 of FIG. 1. The sensor may include an image sensor, such as a fisheye camera mounted behind the windshield, a pillar-mounted forward-facing or side-facing camera, a rear-facing camera, or the like. In various embodiments, the sensor data is in or has been converted to a format that the machine learning model trained at 303 utilizes as input data. For example, the sensor data may be raw image data or processed image data. In some embodiments, the sensor data is preprocessed using an image preprocessor, such as image preprocessor 105 of FIG. 1, during a preprocessing step. For example, the image may be normalized to remove distortion, noise, etc. In some alternative embodiments, the received sensor data here is data captured from an ultrasonic sensor, a radar sensor, a LiDAR sensor, a microphone, or other suitable technology, and this data is used as potential input data for the trained machine learning model deployed in 305.

[0048] At 309, the trained machine learning model is applied. For example, the machine learning model trained at 303 is applied to the received sensor data at 307. In some embodiments, the application of the model is performed by an AI processor, such as AI processor 109 of FIG. 1, using a deep learning network, such as deep learning network 107 of FIG. 1. In various embodiments, the trained machine learning model is applied to predict one or more object attributes, such as object distance, object direction, and / or object velocity, from the image data. For example, as different objects are identified in the image data, the object distance and object direction of each identified object are inferred using the trained machine learning model. As another example, for a vehicle identified in the image data, a velocity vector for the vehicle is inferred. This velocity vector may be used to determine whether a nearby vehicle is likely to cut into the current lane and / or whether the vehicle is likely to pose a safety risk. In various embodiments, vehicles, pedestrians, obstacles, lanes, traffic signals, map features, speed limits, available space, etc., and their associated attributes are identified by applying the machine learning model. In some embodiments, these features are identified in three dimensions, such as three-dimensional velocity vectors.

[0049] At 311, the automated vehicle is controlled. For example, one or more automated driving functions are performed by controlling various aspects of the vehicle. Examples may include controlling the vehicle's steering, speed, acceleration, and / or braking; maintaining the vehicle's position within its lane; maintaining the vehicle's position relative to other vehicles and / or obstacles; providing notifications or warnings to the occupants; etc. Based on the analysis performed at 309, the vehicle's steering and speed may be controlled to safely maintain the vehicle between two lane lines and at a safe distance from other objects. For example, distances and directions to surrounding objects are predicted, and corresponding drivable spaces and driving paths are identified. In various embodiments, a vehicle control module, such as vehicle control module 111 of FIG. 1 , controls the vehicle.

[0050] FIG. 4 is a flow diagram illustrating one embodiment of a process for training and applying machine learning models used in automated driving. In some embodiments, the process of FIG. 4 is used to collect and maintain sensor data for training machine learning models used in automated driving. In some embodiments, the process of FIG. 4 is performed in a vehicle with automated driving enabled, regardless of whether automated driving control is enabled. For example, sensor data may be collected while a vehicle is being driven by a human driver and / or while the vehicle is being driven autonomously, immediately after automated driving is disengaged. In some embodiments, the technique described by FIG. 4 is performed using the deep learning system of FIG. 1. In some embodiments, portions of the process of FIG. 4 are performed at 307, 309, and / or 311 of FIG. 3 as part of a process for applying machine learning models used in automated driving.

[0051] At 401, sensor data is received. For example, a vehicle equipped with sensors captures the sensor data and feeds the sensor data to a neural network running on the vehicle. In some embodiments, the sensor data may be visual data, ultrasonic data, radar data, LiDAR data, or other suitable sensor data. For example, images may be captured from a high dynamic range forward-facing camera. As another example, images may be captured from a side-facing super camera. Ultrasonic data is captured from sonic sensors. In some embodiments, multiple sensors capturing data are installed on the vehicle. For example, in some embodiments, eight surround cameras are installed on the vehicle, providing a 360-degree view around the vehicle with a range of up to 250 meters. In some embodiments, the camera sensors include a wide-angle forward camera, a narrow-angle forward camera, a rear-view camera, a forward-view side camera, and / or a rear-view side camera. In some embodiments, ultrasonic and / or radar sensors are used to capture details of the surrounding environment. For example, 12 ultrasonic sensors may be installed on the vehicle to detect both hard and soft objects.

[0052] In various embodiments, data captured from different sensors is associated with captured metadata so that the data captured from the different sensors can be associated together. For example, direction, field of view, frame rate, resolution, timestamp, and / or other captured metadata may be received along with the sensor data. This metadata may be used to correlate sensor data of different formats to better capture the vehicle's surroundings. In some embodiments, the sensor data includes odometry data, including the vehicle's position, heading, change in position, and / or change in heading. For example, position data may be captured and then correlated with other sensor data captured during the same time frame. As one example, position data captured at the time of image data capture may be used to correlate the position information with the image data. In various embodiments, the received sensor data is provided for deep learning analysis purposes.

[0053] At 403, the sensor data is preprocessed. In some embodiments, one or more preprocessing passes may be performed on the sensor data. For example, the data may be preprocessed to remove noise, correct for alignment issues and / or blurring, etc. In some embodiments, one or more different filtering passes may be performed on the data. For example, a high-pass filter may be performed on the sensor data to isolate individual components, and a low-pass filter may be performed on the data. In various embodiments, the preprocessing steps performed at 403 are optional and / or may be incorporated into a neural network.

[0054] At 405, deep learning analysis of the sensor data is initiated. In some embodiments, this deep learning analysis is performed on the sensor data received at 401 and optionally preprocessed at 403. In various embodiments, this deep learning analysis is performed using a neural network, such as a convolutional neural network (CNN). In various embodiments, this machine learning model is trained offline using the process of FIG. 3 and deployed to the vehicle to perform inference on the sensor data. For example, the model may be trained to predict object attributes such as distance, direction, and / or speed. In some embodiments, the model is trained to identify pedestrians, moving vehicles, parked vehicles, obstacles, lane lines, drivable space, etc., as needed. In some embodiments, a bounding box is specified for each object identified in the image data, and distance and direction are predicted for each identified object. In some embodiments, these bounding boxes are three-dimensional bounding boxes, such as rectangular prisms. The bounding boxes outline the outer surfaces of the identified objects and may be adjusted based on the size of the objects. For example, vehicles of different sizes are represented using bounding boxes (or cuboids) of different sizes. In some embodiments, object attributes estimated by deep learning analysis are compared to attributes measured by sensors and received as sensor data. In various embodiments, the neural network comprises multiple neural networks including one or more hidden layers. The sensor data is analyzed using one or more different neural networks, including layers of neural networks. In various embodiments, the sensor data and / or the results of the deep learning analysis are retained and transmitted 411 for automatic generation of training data.

[0055] In various embodiments, this deep learning analysis is used to predict additional features. These predicted features may be used to assist automated driving. For example, the detected vehicle may be assigned to a lane or road. As another example, the detected vehicle may be determined to be in a blind spot, a vehicle that should have priority, a vehicle in the adjacent lane to the left, a vehicle in the adjacent lane to the right, or a vehicle with another suitable attribute. Similarly, the deep learning analysis may identify traffic lights, clearance, pedestrians, obstacles, or other suitable driving features.

[0056] At 407, results of the deep learning analysis are provided to vehicle control. For example, these results are provided to a vehicle control module to control the autonomous vehicle and / or perform autonomous driving functions. In some embodiments, the results of the deep learning analysis at 405 are passed through one or more additional deep learning passes using one or more different machine learning models. For example, a drivable space may be designated using the identified objects and their attributes (e.g., distance, direction, etc.). This drivable space is then used to designate a drivable path for the vehicle. Similarly, in some embodiments, a predicted speed vector for the vehicle is determined. A path for the vehicle designated based at least in part on this predicted speed vector is used to predict cut-ins and avoid potential collisions. In some embodiments, various outputs from the deep learning are used to build a three-dimensional representation of the autonomous vehicle, including speed limits, obstacles to avoid, identified objects including road conditions, distances and directions to the identified objects, the predicted path for the vehicle, identified traffic signals, etc. In some embodiments, a vehicle control module uses these findings to control the vehicle along the specified route, which in some embodiments is vehicle control module 111 of FIG.

[0057] At 409, the vehicle is controlled. In some embodiments, an autonomous driving enabled vehicle is controlled using a vehicle control module, such as vehicle control module 111 of FIG. 1 . The vehicle control may, for example, adjust the vehicle's speed and / or steering to maintain a safe distance from other vehicles and within a lane at an appropriate speed that takes into account the vehicle's surroundings. In some embodiments, these results are used to make adjustments to the vehicle in anticipation of a nearby vehicle merging into the same lane. In various embodiments, the vehicle control module uses the results of the deep learning analysis to determine an appropriate way to drive the vehicle along a specified path at an appropriate speed, for example. In various embodiments, the results of vehicle control, such as changes in speed, application of braking, and steering adjustments, are retained and used to automatically generate training data. In various embodiments, these vehicle control parameters may be retained and transmitted at 411 so that training data can be automatically generated.

[0058] At 411, the sensor data and associated data are transmitted. For example, the sensor data received at 401 is transmitted to a computer server along with the results of the deep learning analysis at 405 and / or the vehicle control parameters used at 409 so that training data is automatically generated. In some embodiments, this data is time series data, and these various collected data are correlated together by a remote training computer server. For example, ground truth is generated by correlating the image data with auxiliary sensor data, such as distance data, direction data, and / or speed data. In various embodiments, these collected data are transmitted wirelessly from the vehicle to a training data center, for example, via a WiFi or cellular connection. In some embodiments, Metadata is transmitted along with the sensor data. For example, the metadata may include vehicle control and / or operation parameters such as time, timestamp, location, vehicle type, speed, acceleration, braking, whether autonomous driving was enabled, steering angle, odometry data, etc. Additional metadata may include the time since the most recent sensor data transmission, vehicle type, weather conditions, road conditions, etc. In some embodiments, this transmitted data is anonymized, for example, by removing the vehicle's unique identifier. As another example, data from similar vehicle models is merged to prevent identification of individual users and their use of the vehicles.

[0059] In some embodiments, this data is transmitted only in response to a trigger. For example, in some embodiments, an inaccurate prediction triggers the transmission of image sensor data and auxiliary sensor data to automatically collect data to create a curated set of examples that improve the deep learning network's predictions. For example, a prediction performed at 405 using only image data to estimate the distance and direction to the vehicle is determined to be inaccurate by comparing the prediction with distance data from an illumination distance measurement sensor. If the prediction differs from the actual sensor data by more than a certain threshold, the image sensor data and associated auxiliary data are transmitted and used to automatically generate training data. In some embodiments, the trigger may be used to identify individual scenarios, such as sharp turns, road forks, lane merges, sudden stops, intersections, or other suitable scenarios where additional training data would be useful and may be difficult to collect. For example, a trigger may be based on the sudden termination or disengagement of an autonomous driving function. As another example, vehicle operating characteristics such as changes in speed or acceleration may form the basis of a trigger. In some embodiments, a prediction with accuracy below a certain threshold triggers the transmission of sensor data and associated auxiliary data. For example, in certain scenarios, a prediction does not have a Boolean true or false outcome and is instead evaluated by determining the accuracy value of the prediction.

[0060] In various embodiments, the sensor data and associated auxiliary data are captured over a period of time, and the entire time series is transmitted together. This time interval may be set and / or may be based on one or more factors, such as the vehicle's speed, distance traveled, changes in speed, etc. In some embodiments, the sampling rate of the captured sensor data and / or associated auxiliary data is configurable. For example, the sampling rate may be higher or faster during hard braking, hard acceleration, hard steering, or another suitable scenario where more repeatability is required.

[0061] FIG. 5 illustrates an example of capturing supplemental sensor data for training a machine learning network. In the illustrated example, autonomous vehicle 501 includes at least sensors 503 and 553 and captures sensor data used to measure object attributes of surrounding vehicles 511, 521, and 561. In some embodiments, the captured sensor data is captured and processed using a deep learning system installed on autonomous vehicle 501, such as deep learning system 100 of FIG. 1. In some embodiments, sensors 503 and 553 are additional sensors 103 of FIG. 1. In some embodiments, the captured data is data related to a portion of the visual data received at 203 of FIG. 2 and / or the sensor data received at 401 of FIG. 4.

[0062] In some embodiments, sensors 503 and 553 of autonomous vehicle 501 are illumination-based distance measurement sensors, such as radar, ultrasonic, and / or lidar sensors. Sensor 503 is a forward-facing sensor and sensor 553 is a right-side facing sensor. Additional sensors, such as a rear-facing sensor and a left-side facing sensor (not shown), may be attached to autonomous vehicle 501. Axes 505 and 507, indicated by long dashed arrows, represent the direction of travel of autonomous vehicle 501. Axis 505 and 507 are reference axes of vehicle 501 and may be used as reference axes for data captured using sensor 503 and / or sensor 553. In the illustrated example, axes 505 and 507 are centered forward of sensor 503 and autonomous vehicle 501. In some embodiments, an additional elevation axis (not shown) is used to track attributes in three dimensions. In various embodiments, another axis may be used. For example, this reference axis may be the center of autonomous vehicle 501. In some embodiments, each sensor of sensors 503 and 553 may use its own reference axis and coordinate system. Data captured and analyzed using the respective local coordinate systems of sensors 503 and 553 may be transformed to the local (or world) coordinate system of autonomous vehicle 501 so that data captured from different sensors may be shared using the same frame of reference.

[0063] In the illustrated example, the fields of view 509 and 559 of sensors 503 and 553, respectively, are indicated by dotted arcs between dotted arrows. The illustrated fields of view 509 and 559 represent overhead views of the areas measured by sensors 503 and 553, respectively. Attributes of objects within field of view 509 may be captured by sensor 503, and attributes of objects within field of view 559 may be captured by sensor 553. For example, in some embodiments, distance, direction, and / or speed measurements to objects within field of view 509 are captured by sensor 503. In the illustrated example, sensor 503 captures the distance and direction to peripheral vehicles 511 and 521. Sensor 503 does not measure peripheral vehicle 561 because peripheral vehicle 561 is outside of field of view 509. Instead, the distance and direction to peripheral vehicle 561 is captured by sensor 553. In various embodiments, objects not captured by one sensor may be captured by another sensor within the vehicle. Although shown in FIG. 5 with only sensors 503 and 553, autonomous vehicle 501 may be equipped with multiple surround sensors (not shown) that provide a 360-degree view around the vehicle.

[0064] In some embodiments, sensors 503 and 553 capture distance and direction measurements. Distance vector 513 indicates the distance and direction to surrounding vehicle 511, distance vector 523 indicates the distance and direction to surrounding vehicle 521, and distance vector 563 indicates the distance and direction to surrounding vehicle 561. In various embodiments, the actual distance and direction values ​​captured are sets of values ​​corresponding to the exterior surfaces detected by sensors 503 and 553. In the illustrated example, the set of distances and directions measured for each surrounding vehicle is approximated by distance vectors 513, 523, and 563. In some embodiments, sensors 503 and 553 detect velocity vectors (not shown) of objects within their respective fields of view 509 and 559. In some embodiments, these distance and velocity vectors are three-dimensional vectors. For example, these vectors include a height (or elevation) component (not shown).

[0065] In some embodiments, detected objects, including detected surrounding vehicles 511, 521, and 561, are approximated by bounding boxes. These bounding boxes approximate the exterior of the detected objects. In some embodiments, these bounding boxes are three-dimensional bounding boxes, such as rectangular prisms, or other solid representations of the detected objects. In the example of FIG. 5, these bounding boxes are shown as rectangles around surrounding vehicles 511, 521, and 561. In various embodiments, distance and direction from autonomous vehicle 501 may be measured for each point on the edge (or surface) of the bounding box.

[0066] In various embodiments, distance vectors 513, 523, and 563 are data related to visual data captured at the same time. Distance vectors 513, 523, and 563 are used to calculate the distances to surrounding vehicles 511, 521, and 561 identified in the corresponding visual data. The distances and directions of the surrounding vehicles 511, 521, and 561 are annotated. For example, distance vectors 513, 523, and 563 may be used as ground truth to annotate training images that include surrounding vehicles 511, 521, and 561. In some embodiments, the training images corresponding to the captured sensor data of FIG. 5 are captured from sensors with overlapping fields of view and utilize data captured at matching times. For example, if the training images are image data captured from a forward-facing camera that captures only surrounding vehicles 511 and 521, but not surrounding vehicle 561, then only surrounding vehicles 511 and 521 are identified in the training images and their corresponding distances and directions are annotated. Similarly, a right-side image capturing surrounding vehicle 561 includes annotations for the distance and direction to surrounding vehicle 561 only. In various embodiments, the annotated training images are sent to a training server to train a machine learning model to predict the annotated object attributes. 5 is transmitted to a training platform where the data is analyzed and training images are selected and annotated. For example, the data captured here may be time series data that is analyzed to associate relevant data with objects identified in the visual data.

[0067] FIG. 6 illustrates an example of predicting object attributes. In the illustrated example, analyzed visual data 601 represents a field of view of image data captured from a visual sensor, such as a forward-facing camera, of an autonomous vehicle. In some embodiments, the visual sensor is one of visual sensors 101 of FIG. 1. In some embodiments, the environment in front of the vehicle is captured and processed using a deep learning system, such as deep learning system 100 of FIG. 1. In various embodiments, the process illustrated in FIG. 6 is performed at 307, 309, and / or 311 of FIG. 3 and / or 401, 403, 405, 407, and / or 409 of FIG. 4.

[0068] In the illustrated example, analyzed visual data 601 captures the environment ahead of an autonomous vehicle. The analyzed visual data 601 includes lane lines 603, 605, 607, and 609 of a detected vehicle. In some embodiments, these lane lines are identified using a deep learning system, such as deep learning system 100 of FIG. 1 , trained to identify driving functions. The analyzed visual data 601 includes bounding boxes 611, 613, 615, 617, and 619 corresponding to detected objects. In various embodiments, the detected objects represented by bounding boxes 611, 613, 615, 617, and 619 are identified by analyzing the captured visual data. The captured visual data is used as input to a trained machine learning model to predict object attributes, such as distance and direction to the detected objects. In some embodiments, velocity vectors are predicted. In the illustrated example, the detected objects in bounding boxes 611, 613, 615, 617, and 619 correspond to surrounding vehicles. Bounding boxes 611, 613, and 617 correspond to vehicles in the lane defined by lane boundaries 603 and 605. Bounding boxes 615 and 619 correspond to vehicles in the merging lane defined by lane boundaries 607 and 609. In some embodiments, the bounding boxes are used to represent the detected objects as three-dimensional bounding boxes (not shown).

[0069] In various embodiments, the predicted object attributes for bounding boxes 611, 613, 615, 617, and 619 are predicted by applying a machine learning model trained using the process of FIGS. 2-4. The predicted object attributes may also be captured using an auxiliary sensor, as shown in the diagram of FIG. 5. While FIGS. 5 and 6 illustrate different driving scenarios, FIG. 5 shows a more accurate representation of the detected objects compared to FIG. 6. The figures show that the number and location of objects differs, and a trained machine learning model, when trained with sufficient training data, can accurately predict object attributes of objects detected in the scenario of FIG. 6. In some embodiments, distance and direction are predicted. In some embodiments, velocity is predicted. The predicted attributes may be predicted in two or three dimensions. By automating the generation of training data using the process described with respect to FIGS. 1-6, training data for making accurate predictions is generated in an efficient and appropriate manner. In some embodiments, the identified objects and corresponding attributes may be used to perform automated driving functions, such as automated driving or driver assistance operations of a vehicle. For example, the steering and speed of a vehicle may be controlled to safely keep the vehicle between two lane lines and at a safe distance from other objects.

[0070] Although the foregoing embodiments have been described in some detail for clarity of understanding, the present invention is not limited to the details shown. There are many alternative ways of implementing the present invention. The disclosed embodiments are illustrative and not limiting.

Claims

1. 1. A system comprising one or more processors, the one or more processors comprising: receiving sensor data representative of at least one object in an environment of the vehicle; transmitting the sensor data to a trained machine learning model and causing the trained machine learning model to generate an output representative of at least one feature of the at least one object in the environment, the at least one feature comprising a velocity vector of the at least one object in the environment; the trained machine learning model has been trained using training images and a correlation output of an illumination distance measurement sensor, the training images and the correlation output of the illumination distance measurement sensor being captured by a training vehicle different from the vehicle, and the correlation output including a velocity vector of an object detected by the illumination distance measurement sensor of the training vehicle; 10. The system of claim 1, wherein the trained machine learning model of the vehicle is configured to generate the output without the use of any illumination-based distance measurement sensors.

2. The system of claim 1 , wherein the at least one characteristic of the at least one object in the environment includes at least one of a distance of the object relative to the sensor, a direction of the object relative to the environment, or a velocity of the object relative to the environment.

3. the at least one object includes at least one object or at least one identified object; the at least one feature includes at least one feature of the at least one object or at least one feature of the at least one identified object, the at least one identified object including a pedestrian or a second vehicle moving relative to the vehicle; the one or more processors 10. The system of claim 1, further configured to transmit the sensor data to the trained machine learning model and cause the trained machine learning model to generate a representation of the at least one feature of the at least one object in the environment or the at least one feature of the at least one identified object in the environment.

4. the one or more processors determining the velocity vector of the at least one object based on the output of the trained machine learning model; The system of claim 1 , further configured to determine a future steering based on the velocity vector of the at least one object.

5. 10. The system of claim 1, wherein the one or more processors receive the sensor data based on generation of the sensor data by at least one visual sensor, at least one camera, at least one fisheye camera, at least one LiDAR sensor, at least one ultrasonic sensor, or at least one radar sensor.

6. the one or more processors further configured to normalize the sensor data; The one or more processors that transmit the sensor data to the trained machine learning model: The system of claim 1 , configured to transmit the sensor data to the trained machine learning model based on a normalization of the sensor data.

7. the one or more processors 4. The system of claim 3, further configured to cause a vehicle control module to control operation of the vehicle based on the at least one characteristic of the at least one object in the environment or the at least one characteristic of the at least one identified object in the environment.

Citation Information

Patent Citations

  • Method, apparatus and computer program for a vehicle

    EP3438872A1

  • Object detection device and object detection method and program

    JP2019008460A

  • Collision avoidance system for autonomous vehicle

    JP2019008796A

  • High resolution 3D point clouds generation from upsampled low resolution lidar 3D point clouds and camera images

    US20190004534A1

  • Predicting depth from image data using a statistical model

    WO2018046964A1