Estimation of Object Attributes Using Visual Image Data
By employing a machine learning model trained with visual image data and correlated sensor outputs, the autonomous driving system effectively estimates object distances, addressing the challenges of sensor complexity and cost while achieving high precision.
Patent Information
- Application Number
- JP2021547712
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-19
- Filing Date
- 2020-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-02-07
AI Technical Summary
Autonomous driving systems face challenges in accurately determining object attributes like distance without relying on expensive and complex sensors, while also managing increased input bandwidth and system complexity.
A system that uses image data from a vehicle camera as input to a trained machine learning model to estimate the distance to objects, leveraging correlated outputs from distance measurement sensors by irradiation for training, potentially reducing the need for dedicated distance measurement sensors.
This approach allows for accurate prediction of object attributes with high precision, potentially reducing system complexity and cost by minimizing the number of sensors required, while also providing fault tolerance and improved accuracy.
Smart Images

Figure 0007685953000001 
Figure 0007685953000002 
Figure 0007685953000003
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application is a continuation of U.S. Patent Application No. 16 / 279,657, entitled "Estimation of Object Attributes Using Visual Image Data", filed on February 19, 2019, and claims the benefit of its priority. The entire disclosure of this application is incorporated herein by reference in its entirety.
Background Art
[0002] An autonomous driving system typically relies on mounting a number of sensors on a vehicle, including a set of visual sensors and distance - measuring sensors by irradiation (e.g., radar sensors, lidar sensors, ultrasonic sensors, etc.). In that case, the data captured by each sensor is collected to assist in understanding the surrounding environment of the vehicle and in determining a method for controlling the vehicle. Using a visual sensor, an object can be identified from the captured image data, and using a distance - measuring sensor by irradiation, the distance to the detected object can be measured. Steering adjustments and speed adjustments can be made based on the detection of obstacles and the drivable path without obstacles. On the other hand, as the number and types of sensors increase, the complexity and cost of the system also increase. For example, distance - measuring sensors by irradiation such as lidar are often expensive to mount on mass - produced vehicles. Also, by adding each sensor, the input bandwidth requirements of the autonomous driving system increase. Therefore, it is necessary to find an optimal configuration of sensors on the vehicle. In this configuration, it is necessary to limit the total number of sensors without restricting the amount and type of data captured to accurately describe the surrounding environment and safely control the vehicle.
Summary of the Invention
Means for Solving the Problems
[0003] One embodiment includes a system. The system receives image data based on an image captured using a vehicle camera and uses this image data as a basis for input data to a trained machine learning model to at least partially determine the distance from the vehicle to an object. The one or more processors are configured to do this, and the trained machine learning model has been trained using training images and correlated outputs of a distance measurement sensor by irradiation. The system includes one or more processors and a memory coupled to the one or more processors.
[0004] Another embodiment includes a computer program product embodied on a non-transitory computer-readable storage medium and comprising computer instructions. The computer instructions receive image data based on an image captured using a vehicle camera and use this image data as a basis for input data to a trained machine learning model to be used to at least partially determine the distance from the vehicle to an object. The trained machine learning model has been trained using training images and correlated outputs of a distance measurement sensor by irradiation.
[0005] Yet another embodiment includes a method. The method includes receiving a selected image based on an image captured using a vehicle camera, receiving distance data based on a distance measurement sensor by vehicle irradiation, using this selected image as input data to a trained machine learning model to identify an object, extracting an estimated distance of the identified object from this received distance data, creating a training image by annotating the selected image with this extracted estimated distance, training a second machine learning model to predict distance measurements using a training data set including this training image, and supplying this trained second machine learning model to a second vehicle equipped with a second camera.
Brief Description of the Drawings
[0006] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
[0007]
Figure 1
[0008]
Figure 2
[0009]
Figure 3
[0010]
Figure 4
[0011]
Figure 5
[0012]
Figure 6
Embodiments for Carrying Out the Invention
[0013] The present invention can be implemented by a number of means, including an apparatus, a system, a composition, a computer program product embodied on a computer-readable storage medium, and / or a processor, for example, a processor configured to execute instructions stored in a memory coupled to itself and / or instructions provided by this memory. In this specification, these implementation forms, or any other form that the present invention can take, may also be referred to as techniques. The order of the steps of the disclosed process can generally be changed within the scope of the present invention. Unless otherwise specified, components such as processors or memories described as being configured to perform tasks are implemented as general-purpose components temporarily configured to perform tasks at a given time or as dedicated components manufactured to perform tasks. As used in this specification, the term "processor" refers to one or more devices, circuits, and / or processing cores configured to process data such as computer program instructions.
[0014] A detailed description of one or more embodiments of the present invention is set forth below along with the accompanying drawings that illustrate the principles of the invention. Although the invention is described in connection with such embodiments, the invention is not limited to any particular embodiment. The scope of the invention is defined only by the claims, and the invention includes numerous alternative, modifications, and equivalents. To fully understand the invention, numerous specific details are set forth in the following description. These details are provided for illustrative purposes, and the invention can be practiced without some or all of these specific details, in accordance with the claims. For the sake of clarity, well-known technical materials related to the technical field of the present invention are not described in detail so that the invention is not unnecessarily obscured.
[0015] Disclosed are techniques for machine learning training to generate high-precision machine learning results from visual data. To accurately estimate object attributes such as object distance, auxiliary sensor data such as radar results and lidar results are used, and auxiliary data is associated with objects identified from visual data. In various embodiments, the collection of auxiliary data and its association with visual data are performed automatically and require little to no human intervention. For example, since objects identified using vision technology do not need to be manually labeled, the efficiency of machine learning training is greatly improved. Instead, training data is automatically generated, and a machine learning model can be trained to predict object attributes with high precision using this training data. For example, this data may be automatically collected from a fleet of vehicles by collecting snapshots of associated relevant data such as visual data and radar data. In some embodiments, only a subset of the association targets for visual data and radar data are sampled. The fusion data collected from these vehicle fleets is automatically collected, and a neural net is trained using this fusion data to mimic the data being captured. Once this trained machine learning model is deployed in a vehicle to accurately predict object attributes such as distance, direction, and speed using only visual data, for example, it may no longer be necessary to equip an autonomous vehicle with a dedicated distance measurement sensor. If used with a dedicated distance measurement sensor, this machine learning model can be used as a redundant distance data source or a secondary distance data source to improve accuracy and / or provide fault tolerance. Using the identified objects and corresponding attributes, autonomous driving functions such as the autonomous driving of a vehicle or driver assistance operations can be performed. For example, an autonomous vehicle can be controlled to avoid a merging vehicle identified using the disclosed techniques.
[0016] A system comprising one or more processors coupled to a memory is configured to receive image data based on an image captured using a vehicle camera. For example, a processor such as an artificial intelligence (AI) processor mounted on an autonomous vehicle receives image data from a camera such as a front-facing camera of the vehicle. Other cameras such as side-facing cameras and rear-facing cameras can also be used similarly. The image data is used as a basis for input data to a model that has been machine learning trained to at least partially identify the distance from the vehicle to an object. For example, this captured image is used as input data to a machine learning model such as a model of a deep learning network executed on an AI processor. Using this model, the distance to an object identified in the image data is predicted. Peripheral objects such as vehicles and pedestrians can be identified from the image data, and their accuracy and direction are inferred using a deep learning system. In various embodiments, the trained machine learning model has been trained using training images and the correlated output of a distance measurement sensor by irradiation. The distance measurement sensor by irradiation can transmit a signal (e.g., a radio wave signal, an ultrasonic signal, an optical signal, etc.) when detecting the distance from the sensor to an object. For example, a radar sensor mounted on a vehicle transmits radar to identify the distance and direction to surrounding obstacles. These distances are then correlated with the identified objects in the training images captured by the vehicle's camera. The associated training images are annotated with distance measurements, and this image is used to train the machine learning model. In some embodiments, additional attributes such as the speed of an object are predicted using this model. For example, the speed of an object measured by radar is correlated with the object in the training image when training the machine learning model to predict the speed and direction of the object.
[0017] In some embodiments, a vehicle comprises sensors for capturing the vehicle's surrounding environment and vehicle operation parameters. This captured data includes visual data (such as videos and / or still images), and additional auxiliary data such as radar sensor data, lidar sensor data, inertial sensor data, audio sensor data, odometry sensor data, position sensor data, and / or other forms of sensor data. For example, this sensor data may capture vehicles, pedestrians, lane boundaries, vehicle traffic volume, obstacles, traffic signs, traffic sounds, etc. Odometry sensors and other similar sensors capture vehicle operation parameters such as vehicle speed, steering, orientation, direction change, position change, altitude change, speed change, etc. The visual data and auxiliary data captured here are transmitted from the vehicle to a training server to create a training data set. In some embodiments, these transmitted visual data and auxiliary data are correlated, and using these data, training data is automatically generated. Using this training data, a machine learning model is being trained to generate highly accurate machine learning results. In some embodiments, time-series captured data is used to generate this training data. Based on a group of time-series elements, ground truth is specified, and using this ground truth, at least one of the elements such as a single image from the group is annotated. For example, at time intervals such as 30 seconds, a series of images and radar data are captured. The vehicle identified from this image data and tracked over time is associated with the corresponding radar distance and radar direction from this time series. Auxiliary data to be associated such as radar distance data is associated with the vehicle after analyzing the image data and distance data captured in time series. By analyzing the image data and auxiliary data over time series, ambiguities such as multiple objects with similar distances can be resolved with high accuracy, and ground truth can be specified.For example, when using only a single captured image, when one vehicle obscures another vehicle or when two vehicles are approaching, there may be insufficient radar data to accurately estimate the individual distances to the two vehicles. However, by tracking the vehicles over time, the distances identified by the radar can be properly associated with the correct vehicles even if the vehicles move away from each other, move in different directions, and / or move at different speeds. In various embodiments, when auxiliary data is appropriately associated with an object, one or more images in the time series are converted into training images, and corresponding ground truths such as Data , speed Data , and / or other appropriate object attributes are annotated in this image.
[0018] In various embodiments, a machine learning model trained using auxiliary sensor data can accurately predict the results of an auxiliary sensor without the need for a physical auxiliary sensor. For example, a training vehicle can be equipped with an auxiliary sensor, including an expensive and / or difficult-to-operate sensor, to collect training data. This training data can then be used to train a machine learning model to predict the results of auxiliary sensors such as radar sensors, lidar sensors, or other sensors. Subsequently, this trained model is deployed to vehicles such as mass-produced vehicles that require only vision sensors. The auxiliary sensor is not essential but can also be used as a secondary data source. There are many advantages to reducing the number of sensors, including, among other things, the difficulty of recalibrating the sensors, sensor maintenance, the cost of adding sensors, and / or the additional bandwidth and computational requirements for additional sensors. In some embodiments, the trained model is used when the auxiliary sensor fails. Instead of relying on additional auxiliary sensors, this trained machine learning model uses input data from one or more vision sensors to predict the results of the auxiliary sensor. The predicted results can be used here to perform an autonomous driving function that requires the detection of objects (e.g., pedestrians, stationary vehicles, moving vehicles, curbs, obstacles, road barriers, etc.), as well as the distance and direction to them. The predicted results can be used here to detect the distance and direction to traffic control objects such as traffic lights, traffic signs, road signs, etc. Although vision sensors and object distances were used in the previous examples, alternative sensors and predicted attributes can be used as well.
[0019] FIG. 1 is a block diagram showing an embodiment of a deep learning system used for autonomous driving. This deep learning system includes various components that can be used together to perform autonomous driving and / or driver assistance operations of a vehicle, as well as to collect and process data for training a machine learning model. In various embodiments, this deep learning system is installed in a vehicle, and using data captured from the vehicle, the deep learning system in the vehicle or other similar vehicles can be trained and improved. Using this deep learning system, autonomous driving functions may be performed, including identifying objects and predicting object attributes such as distance and direction using visual data as input data.
[0020] In the illustrated embodiment, the deep learning system 100 is a deep learning network that includes a vision sensor 101, an additional sensor 103, an image preprocessor 105, a deep learning network 107, an artificial intelligence (AI) processor 109, a vehicle control module 111, and a network interface 113. In various embodiments, the various components are communicatively coupled. For example, the image data captured from the vision sensor 101 is supplied to the image preprocessor 105. The sensor data processed in the image preprocessor 105 is supplied to the deep learning network 107 running on the AI processor 109. In some embodiments, the sensor data from the additional sensor 103 is used as input data to the deep learning network 107. The output data of the deep learning network 107 running on the AI processor 109 is supplied to the vehicle control module 111. In various embodiments, the vehicle control module 111 is connected to the operations of the vehicle, such as the speed, braking, and / or steering of the vehicle, to control these operations of the vehicle. In various embodiments, the sensor data and / or the machine learning results can be transmitted to a remote server (not shown) via the network interface 113. For example, sensor data, such as the data captured from the vision sensor 101 and / or the additional sensor 103, can be transmitted to a remote training server via the network interface 113 to collect training data for enhancing the performance, comfort, and / or safety of the vehicle. In various embodiments, the network interface 113 is used to communicate with, call, send and / or receive text messages from a remote server, and transmit sensor data based on the operations of the vehicle. In some embodiments, the deep learning system 100 may add or reduce components as needed. For example, in some embodiments, the image preprocessor 105 is an optional component. As another example, in some embodiments, a post-processing component (not shown) is used to perform post-processing on the output data of the deep learning network 107 before the output data is supplied to the vehicle control module 111.
[0021] In some embodiments, the visual sensor 101 includes one or more camera sensors for capturing image data. In various embodiments, the visual sensor 101 may be attached to a vehicle at various positions of the vehicle and / or may be oriented in one or more different directions. For example, the visual sensor 101 may be attached to the front, side, rear, and / or roof of the vehicle in directions such as the front-facing direction, the rear-facing direction, and the side-facing direction. In some embodiments, the visual sensor 101 may be an image sensor such as a high-dynamic range camera and / or a camera having various fields of view. For example, in some embodiments, eight surround cameras are attached to the vehicle, and these cameras provide a 360-degree field of view around the vehicle up to a range of 250 meters. In some embodiments, the camera sensors include a wide-angle front camera, a narrow-angle front camera, a rear view camera, a front side view camera, and / or a rear side view camera.
[0022] In some embodiments, the vision sensor 101 is not attached to a vehicle that includes the vehicle control module 111. For example, the vision sensor 101 may be attached to a surrounding vehicle and / or may be attached to a road or the surrounding environment, and may further be incorporated as part of a deep learning system for capturing sensor data. In various embodiments, the vision sensor 101 includes one or more cameras that capture the surrounding environment of the vehicle, including the road on which the vehicle is traveling. For example, one or more front-facing cameras and / or pillar cameras capture images of objects in the environment surrounding the vehicle, such as the vehicle, pedestrians, traffic control objects, roads, curbs, obstacles, etc. As another example, the camera captures time-series image data that includes image data of surrounding vehicles, including vehicles that are attempting to cut into the lane in which the vehicle is traveling. The vision sensor 101 may include an image sensor that can capture still images and / or videos. The data may be captured over a period of time, such as a series of captured data over a certain period, and may also be synchronized with other vehicle data including other sensor data. For example, the image data used to identify an object may be captured together with radar data and odometry data over a period of 15 seconds or another appropriate time.
[0023] In some embodiments, additional sensor 103 includes additional sensors for capturing sensor data in addition to vision sensor 101. In various embodiments, additional sensor 103 may be attached to a vehicle at various positions of the vehicle and / or may be oriented in one or more different directions. For example, additional sensor 103 may be attached in a direction such as a front-facing direction, a rear-facing direction, a side-facing direction, etc. to the front, side, rear, and / or roof of the vehicle. In some embodiments, additional sensor 103 may be an irradiation sensor such as a radar sensor, an ultrasonic sensor, and / or a lidar sensor. In some embodiments, additional sensor 103 includes non-visual sensors. Additional sensor 103 may include a radar sensor, a voice sensor, a lidar sensor, an inertial sensor, an odometry sensor, a position sensor, and / or an ultrasonic sensor, etc. For example, by attaching 12 ultrasonic sensors to a vehicle, both hard objects and soft objects may be detected. In some embodiments, a front-facing radar is used to capture data on the surrounding environment. In various embodiments, the radar sensor can capture details of the surrounding environment despite heavy rain, fog, dust generation, and the proximity of other vehicles.
[0024] In some embodiments, the additional sensor 103 is not attached to a vehicle equipped with the vehicle control module 111. Similar to the visual sensor 101, for example, the additional sensor 103 may be attached to a surrounding vehicle and / or attached to a road or the surrounding environment, and may further be incorporated as part of a deep learning system for capturing sensor data. In some embodiments, the additional sensor 103 includes one or more cameras that capture the surrounding environment of the vehicle, including the road on which the vehicle is traveling. For example, a front-facing radar sensor captures distance data to an object within the forward field of view of the vehicle. The additional sensor may capture odometry information, position information, and / or vehicle control information including information regarding the trajectory of the vehicle. The sensor data may be captured over a period of time, such as a series of captured data over a period of time, and may also be associated with the image data captured from the visual sensor 101. In some embodiments, the additional sensor 103 includes a position sensor such as a global positioning system (GPS) sensor for identifying the position and / or change in position of the vehicle. In various embodiments, one or more of the additional sensors 103 are optional and are only incorporated in vehicles designed to capture training data. A vehicle without one or more of the additional sensors 103 can simulate the results by the additional sensor 103 by predicting the output using a trained machine learning model and the techniques disclosed herein. For example, a vehicle without a front-facing radar sensor or a front-facing lidar sensor can predict the results of an optional sensor using image data by applying a trained machine learning model such as the model of the deep learning network 107.
[0025] In some embodiments, the image preprocessor 105 is used to preprocess the sensor data of the vision sensor 101. For example, the image preprocessor 105 may be used to preprocess this sensor data, this sensor data may be split into one or more components, and / or these one or more components may be post-processed. In some embodiments, the image preprocessor 105 is a graphics processing unit (GPU), a central processing unit (CPU), an image signal processor, or a dedicated image processor. In various embodiments, the image preprocessor 105 is a tone mapper processor that processes high dynamic range data. In some embodiments, the image preprocessor 105 is implemented as part of the artificial intelligence (AI) processor 109. For example, the image preprocessor 105 may be a component of the AI processor 109. In some embodiments, the image preprocessor 105 may be used to normalize or transform an image. For example, an image captured with a fish-eye lens may be warped, and the image may be transformed using the image preprocessor 105 to remove or correct this warping. In some embodiments, noise, distortion, and / or blur are removed or reduced during the preprocessing step. In various embodiments, the image is adjusted or normalized to improve the results of machine learning analysis. For example, the white balance of the image is adjusted to take into account different lighting operating conditions such as daylight conditions, sunny conditions, cloudy conditions, dim conditions, sunrise conditions, sunset conditions, and nighttime conditions.
[0026] In some embodiments, the deep learning network 107 is a deep learning network used for identifying vehicle control parameters, including analyzing the driving environment to identify objects and their corresponding attributes, such as distance, speed, or another appropriate parameter. For example, the deep learning network 107 may be an artificial neural network such as a convolutional neural network (CNN) that is trained with input data such as sensor data and whose output data is supplied to the vehicle control module 111. As an example, this output data may at least include an estimated distance to the detected object. As another example, this output data may at least include potential vehicles that may merge into the vehicle's lane, the distances between these vehicles, and the speeds of these vehicles. In some embodiments, the deep learning network 107 at least receives image sensor data as input data, identifies objects within this image sensor data, and predicts the distance to the objects. The additional input data may include scene data that describes vehicle specifications such as the environment around the vehicle and / or the operating characteristics of the vehicle. The scene data may include scene tags that describe the environment around the vehicle, such as a road during rain, a wet road, traffic during snowfall, traffic during the occurrence of mud, high-density traffic, an arterial road area, an urban area, a school route, etc. In some embodiments, the output data of the deep learning network 107 is a three-dimensional representation of the vehicle's surrounding environment, including a cuboid representing an object such as the identified object. In some embodiments, the output data of the deep learning network 107 is used for autonomous driving, including navigating the vehicle towards the target destination.
[0027] In some embodiments, the artificial intelligence (AI) processor 109 is a hardware processor for executing the deep learning network 107. In some embodiments, the AI processor 109 is a dedicated AI processor for performing inferences on sensor data using a convolutional neural network (CNN). The AI processor 109 may be optimized for the bit depth of the sensor data. In some embodiments, the AI processor 109 is optimized for deep learning operations, such as neural network operations including, among other things, convolutional operations, dot product operations, vector operations, and / or matrix operations. In some embodiments, the AI processor 109 is implemented using a graphics processing unit (GPU). In various embodiments, the AI processor 109 is coupled to a memory configured to provide the AI processor with instructions that, when executed, cause the AI processor to perform a deep learning analysis of the received input sensor data and determine machine learning results, such as object distances, for use in autonomous driving. In some embodiments, the sensor data is processed using the AI processor 109 to generate the data as available for use as training data.
[0028] In some embodiments, the vehicle control module 111 is used to process the output data of the artificial intelligence (AI) processor 109 and convert this output data into vehicle control operations. In some embodiments, the vehicle control module 111 is used to control a vehicle that performs autonomous driving. In various embodiments, the vehicle control module 111 can adjust the speed, acceleration, steering, braking, etc. of the vehicle. For example, in some embodiments, by using the vehicle control module 111 to control the vehicle, the position of the vehicle within a certain lane is maintained, the vehicle merges into another lane, and the speed and lane positioning of the vehicle are adjusted taking into account the merging vehicle.
[0029] In some embodiments, the vehicle control module 111 is used to control vehicle lighting such as brake lights, turn signals, headlights, etc. In some embodiments, the vehicle control module 111 is used to control the acoustic state of the vehicle, such as the vehicle's audio system, playback of voice alerts, activation of the microphone, activation of the horn, etc. In some embodiments, the vehicle control module 111 is used to control a notification system including an alert system for notifying the driver and / or passengers of driving events such as potential collision possibilities or the approach situation to the assumed destination. In some embodiments, the vehicle control module 111 is used to adjust sensors such as the visual sensor 101 and additional sensor 103 of a vehicle. For example, using the vehicle control module 111, parameters of one or more sensors may be changed, such as changing the orientation, changing the output resolution and / or format type, varying the capture rate up and down, adjusting the captured dynamic range, adjusting the focus of the camera, activating and / or deactivating the sensor. In some embodiments, using the vehicle control module 111, parameters of the image pre-processor 105 may be changed, such as changing the frequency range of the filter, adjusting the feature detection parameters and / or edge detection parameters, adjusting the channels and bit depth. In various embodiments, the vehicle control module 111 is used to perform the autonomous driving and / or driver assistance control of a vehicle. In some embodiments, the vehicle control module 111 is implemented using a processor coupled to a memory. In some embodiments, the vehicle control module 111 is implemented using an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or other suitable processing hardware.
[0030] In some embodiments, network interface 113 is a communication interface for sending and / or receiving data including training data. In various embodiments, network interface 113 includes a cellular interface or a wireless interface for interfacing with a remote server, thereby sending sensor data, sending potential training data, receiving updates to a deep learning network including an updated machine learning model, making and receiving voice calls, sending and / or receiving text messages, and the like. For example, using network interface 113, sensor data captured for use as potential training data may be sent to a remote training server for training a machine learning model. As another example, using network interface 113, updates to instructions and / or operating parameters for vision sensor 101, additional sensor 103, image preprocessor 105, deep learning network 107, AI processor 109, and / or vehicle control module 111 may be received. The machine learning model of deep learning network 107 may be updated using network interface 113. As another example, using network interface 113, the firmware of vision sensor 101 and additional sensor 103, and / or the operating parameters of image preprocessor 105 such as image processing parameters may be updated.
[0031] FIG. 2 is a flowchart showing one embodiment of a process for creating training data for predicting object attributes. For example, the image data is annotated with sensor data from additional auxiliary sensors to automatically create training data. In some embodiments, time series elements composed of sensor data and related auxiliary data are collected from a vehicle, and the training data is automatically created using these time series elements. In various embodiments, the training data is automatically labeled with the corresponding ground truth using the process of FIG. 2. The ground truth and the image data are packaged as training data for predicting object attributes identified from the image data. In various embodiments, the sensor data and related auxiliary data are captured using the deep learning system of FIG. 1. For example, in various embodiments, this sensor data is captured from the visual sensor 101 of FIG. 1, and the related data is captured from the additional sensor 103 of FIG. 1. In some embodiments, the process of FIG. 2 is executed to automatically collect data when existing predictions are inaccurate or can be improved. For example, predictions for identifying one or more object attributes such as distance and direction from visual data are made by an autonomous vehicle. The predictions are compared with distance data received from a distance measurement sensor by irradiation. It can be determined whether the prediction is within an acceptable accuracy threshold. In some embodiments, a determination is made that the prediction can be improved. If the prediction is not accurate enough, applying the process of FIG. 2 to the prediction scenario can create a set of curated training examples for improving the machine learning model.
[0032] At 201, visual data is received. This visual data may be image data such as video and / or still images. In various embodiments, this visual data is captured by a vehicle and then transmitted to a training server. This visual data may be captured over a period of time to create time-series elements. In various embodiments, these elements include timestamps for maintaining the order of the elements. By capturing time-series elements, objects within the time series are tracked over time, making it easier to distinguish objects that are difficult to identify from a single input sample, such as a single input image and corresponding associated data. For example, in the case of a pair of headlights of an oncoming vehicle, initially both may appear to belong to one vehicle, but when these headlights move apart from each other, each headlight is recognized as belonging to a separate motorcycle. In some scenarios, objects in the image data are more distinguishable than objects in the associated auxiliary data received at 203. For example, it may be difficult to determine the estimated distance from a wall where vans are parked to the vans using only distance data. However, by tracking this van over the corresponding time-series image data, the correct distance data can be associated with the identified van. In various embodiments, sensor data captured as a time series is captured in a format that a machine learning model can use as input data. For example, this sensor data may be raw image data or processed image data.
[0033] In various embodiments, when time series data is received, these time series may be compiled by associating a timestamp with each element of these time series. For example, one timestamp is associated with at least a first element in these time series. This timestamp may be used to calibrate time series elements with associated data, such as the data received at 203. In various embodiments, the length of these time series may be a fixed length of time, such as 10 seconds, 30 seconds, or another appropriate length. This length of time may be settable. In various embodiments, these time series may be based on a vehicle speed, such as an average speed of a vehicle. For example, at a slower speed, by increasing the time length of the time series, data may be captured over a longer travel distance than the travel distance that would be possible when using a shorter time length for the same speed. In some embodiments, the number of elements within these time series is settable. The number of these elements may be based on the travel distance. For example, a vehicle moving faster over a certain period of time will have more elements in the included time series than a vehicle moving slower. Adding these elements can enhance the reproducibility of the captured surrounding environment and improve the accuracy of the predicted machine learning results. In various embodiments, the number of these elements is adjusted by adjusting the frames per second at which the sensor captures data and / or by discarding unnecessary intermediate frames.
[0034] At 203, data related to the received visual data is received. In various embodiments, this related data is received at the training server along with the visual data received at 201. In some embodiments, this related data is sensor data received from additional sensors of the vehicle, such as ultrasonic sensors, radar sensors, lidar sensors, or other suitable sensors. This related data may be data captured by the additional sensors of the vehicle, such as data related to distance, direction, speed, position, orientation, change in position, and change in orientation, and / or other related data. This related data may be used to specify the ground truth for the features identified in the visual data received at 201. For example, distance measurements and direction measurements from a radar sensor may be used to determine the object distance and object direction of an object identified in the visual data. In some embodiments, the related data received here is time-series data corresponding to the time-series visual data received at 201.
[0035] In some embodiments, the data related to this visual data includes map data. For example, at 203, offline data such as road-level and / or satellite-level map data may be received. This map data may be used to identify features such as roads, lanes, intersections, speed limits, school zones, etc. For example, this map data can describe the lane path. Using the estimated position of the vehicle identified within the lane, the estimated distance to the detected vehicle may be determined or verified. As another example, this map data can describe the speed limits associated with various roads on the map. In some embodiments, this speed limit data may be used to verify the speed vector of the identified vehicle.
[0036] At 205, an object within the visual data is identified. In some embodiments, this visual data is used as input data to identify objects within the surrounding environment of the vehicle. For example, vehicles, pedestrians, obstacles, etc. are identified from the visual data. In some embodiments, these objects are identified using a deep learning system with a trained machine learning model. In various embodiments, a bounding box is created for the identified object. This bounding box may be a two-dimensional bounding box such as a rectangular cuboid or a three-dimensional bounding box that depicts the outer contour of the identified object. In some embodiments, the identification of the object is assisted by using additional data such as the data received at 203. The accuracy at the time of object identification may be increased by using such additional data.
[0037] At 207, a ground truth for the identified object is specified. Using the related data received at 203, a ground truth for the object identified at 205 is specified from the visual data received at 201. In some embodiments, this related data is depth (and / or distance) data of the identified object. By associating distance data with this identified object, the machine learning model can be trained, and the object distance can be estimated by using this related distance data as the ground truth of the detected object. In some embodiments, these distances are the distances to detected objects such as obstacles, barriers, moving vehicles, stationary vehicles, traffic signals, pedestrians, etc., and are used as the ground truth for training. Distance Data In addition to, direction Data , speed Data , acceleration DataA ground truth for other object parameters such as etc. may be specified. For example, as the ground truth for a specified object, an exact distance or direction is specified. As another example, as the ground truth for a specified object such as a vehicle and a pedestrian, an exact velocity vector is specified.
[0038] In various embodiments, the visual data and the associated data are organized by timestamps, and the corresponding timestamps are used to synchronize two data sets. In some embodiments, the timestamps are used to synchronize time series data such as a series of images and a corresponding series of associated data. This data may be synchronized at the time of capture. For example, when each element of the time series is captured, a corresponding set of associated data is captured and stored together with the time series element. In various embodiments, the time interval of this associated data is configurable and / or matches the time interval of the time series elements. In some embodiments, this associated data is sampled at the same rate as the time series elements.
[0039] In various embodiments, the ground truth can be specified only by verifying the time series data. For example, when only a subset of the visual data is analyzed, the objects and / or their attributes may be mis-identified. By expanding the analysis target to the entire time series, the ambiguity is removed. For example, the presence of a blocking vehicle may be revealed before and after in the time series. Once identified, even if it is blocked, this sometimes-blocked vehicle can be tracked over the entire time series. Similarly, by associating the object attributes obtained from the associated data with the objects identified in the visual data, the object attributes of this sometimes-blocked vehicle can be tracked over the entire time series. In some embodiments, this data is played back in the reverse direction (and / or forward direction) to identify any ambiguity points when associating the associated data with the visual data. Using the state of the objects at various times within the time series may assist in identifying the object attributes of the objects over the entire time series.
[0040] In various embodiments, a threshold is used to determine whether to associate a certain object attribute as the ground truth of a certain identified object. For example, highly accurate related data is associated with the identified object, while related data with a probability below the threshold is not associated with the identified object. In some embodiments, this related data may be conflicting sensor data. For example, the output results of ultrasonic data and radar data may conflict with each other. As another example, the distance data may conflict with the map data. In the distance data, it may be estimated that a certain school district starts within 30 meters, while the information from the map data may describe that the same school district starts within 20 meters. When the accuracy of the related data is low, this related data may be discarded and may not be used to specify the ground truth.
[0041] In some embodiments, this ground truth is specified to predict a semantic label. For example, based on the predicted distance and direction, a detected vehicle can be labeled as being in the left lane or the right lane. In some embodiments, this detected vehicle can be labeled as a vehicle in a blind spot, or as a vehicle to be prioritized, or with another appropriate semantic label. In some embodiments, based on the specified ground truth, vehicles are assigned to roads or lanes within the map. As another example, this specified ground truth can be used to label traffic lights, lanes, drivable spaces, or other mechanisms that assist in autonomous driving.
[0042] At 209, the training data is packaged. For example, elements of the visual data received at 201 are selected and associated with the ground truth specified at 207. In some embodiments, the elements selected here are time series elements. The elements selected here represent sensor data input to a machine learning model such as a training image, and the ground truth represents its prediction result. In various embodiments, the data selected here is annotated and created as training data. In some embodiments, this training data is packaged into training, validation, and test data. This training data is packaged based on the specified ground truth and the selected training elements to train a machine learning model to predict the results related to one or more associated auxiliary sensors. For example, here the trained model can be used to accurately predict the distance and direction to an object while obtaining results similar to measurements using sensors such as radar sensors or lidar sensors. In various embodiments, this machine learning result is used to implement functions used in autonomous driving. The packaged training data is available here when training the machine learning model.
[0043] FIG. 3 is a flowchart showing an embodiment of a process for training and applying a machine learning model used for autonomous driving. For example, input data including primary sensor data and secondary sensor data is received and processed to create training data for training the machine learning model. In some embodiments, the primary sensor data corresponds to image data captured by an autonomous driving system, and the secondary sensor data corresponds to sensor data captured from a distance measurement sensor by irradiation. This secondary sensor data may be used to annotate the primary sensor data in order to train the machine learning model to predict an output based on the secondary sensor. In some embodiments, this sensor data corresponds to sensor data captured based on a specific use case, such as when a user manually disengages autonomous driving or when the distance estimate from visual data is significantly different from the distance estimate from the secondary sensor. In some embodiments, the primary sensor data is the sensor data of the visual sensor 101 in FIG. 1, and the secondary sensor data is the sensor data of one or more sensors in the additional sensor 103 in FIG. 1. In some embodiments, using this process, a machine learning model used in the deep learning system 100 of FIG. 1 is created and deployed.
[0044] At 301, training data is created. In some embodiments, sensor data including image data and auxiliary data is received to create a training data set. This image data may include still images and / or videos from one or more cameras. Additional sensors such as radar sensors, lidar sensors, ultrasonic sensors, etc. may be used to supply relevant auxiliary sensor data. In various embodiments, this image data is paired with corresponding auxiliary data to assist in identifying the attributes of the objects detected within the sensor data. For example, using distance data and / or velocity data obtained from the auxiliary data, the distance and / or velocity to the objects identified within the image data can be accurately estimated. In some embodiments, this sensor data is a time-series element, and using this data, ground truth is specified. Then, the ground truth of the group is associated with a time-series subset such as a frame of the image data. Using the selected time-series elements and this ground truth, training data is created. In some embodiments, this training data is created to train a machine learning model to estimate only the object attributes identified within the image data, such as the distance and direction to vehicles, pedestrians, obstacles, etc. The training data created here may include data used for training, validation, and testing. In various embodiments, this sensor data may be in different formats. For example, this sensor data may be still image data, video data, radar data, ultrasonic data, audio data, position data, odometry data, etc. This odometry data may include vehicle motion parameters such as applied acceleration, applied braking, applied steering, vehicle position, vehicle orientation, change in vehicle position, change in vehicle orientation, etc. In various embodiments, the training data is curated and annotated to create a training data set. In some embodiments, part of the training data creation work may be performed by human curators.In various embodiments, since a portion of this training data is automatically generated from data captured by a vehicle, the effort and time required to construct a robust training data set are significantly reduced. In some embodiments, the format of this data is compatible with the machine learning model used in the deployed deep learning application. In various embodiments, this training data includes validation data for testing the accuracy of the trained model. In some embodiments, the process of FIG. 2 is executed at 301 of FIG. 3.
[0045] At 303, a machine learning model is trained. For example, using the data created at 301, a machine learning model is trained. In some embodiments, the model is a neural network such as a convolutional neural network (CNN). In various embodiments, the model includes a plurality of intermediate layers. In some embodiments, the neural network may include multiple layers including a plurality of convolutional layers and pooling layers. In some embodiments, this training model is verified using a validation data set created from the received sensor data. In some embodiments, the machine learning model is trained to predict the output of a sensor such as a distance irradiation measurement sensor from a single input image. For example, the distance attribute and direction attribute of an object can be inferred from one image captured by a camera. As another example, the speed vector of surrounding vehicles, including whether a vehicle is attempting to merge, is predicted from one image captured by a camera.
[0046] At 305, a trained machine learning model is deployed. For example, this trained machine learning model is installed in a vehicle as an update to a deep learning network such as the deep learning network 107 of FIG. 1. In some embodiments, a newly trained machine learning model is installed using a wireless update. For example, a wireless update can be received via a vehicle network interface such as the network interface 113 of FIG. 1. In some embodiments, this update is a firmware update transmitted using a wireless network such as a WiFi network or a cellular network. In some embodiments, this new machine learning model may be installed during vehicle servicing.
[0047] At 307, sensor data is received. For example, the sensor data is captured from one or more sensors of the vehicle. In some embodiments, this sensor is the vision sensor 101 of FIG. 1. This sensor may include an image sensor such as a fisheye camera mounted behind the front windshield, a front-facing camera or a side-facing camera mounted on a pillar, or a rear-facing camera. In various embodiments, the sensor data is in a format that the machine learning model trained at 303 uses as input data or is converted to that format. For example, this sensor data may be raw image data or processed image data. In some embodiments, this sensor data is preprocessed using an image processor such as the image processor 105 of FIG. 1 during a preprocessing step. For example, this image may be normalized to remove distortion, noise, etc. In some alternative embodiments, the received sensor data here is data captured from an ultrasonic sensor, a radar sensor, a LiDAR sensor, a microphone, or other suitable technology, and this data is used as a candidate for input data to the trained machine learning model deployed at 305.
[0048] At 309, a trained machine learning model is applied. For example, the machine learning model trained at 303 is applied to the sensor data received at 307. In some embodiments, the application of this model is performed by an AI processor such as AI processor 109 in FIG. 1 using a deep learning network such as deep learning network 107 in FIG. 1. In various embodiments, by applying this trained machine learning model, one or more object attributes such as object distance, object direction, and / or object speed are predicted from the image data. For example, when different objects are identified within the image data, the object distance and object direction of each identified object are inferred using the trained machine learning model. As another example, for a certain vehicle identified within the image data, the speed vector of that vehicle is inferred. Using this speed vector, it may be determined whether surrounding vehicles may cut into the current lane and / or whether the vehicle may pose a safety risk. In various embodiments, vehicles, pedestrians, obstacles, lanes, traffic signals, map features, speed limits, drivable spaces, etc. and their related attributes are identified by applying this machine learning model. In some embodiments, these features are identified in three dimensions, such as a three-dimensional speed vector.
[0049] At 311, the autonomous vehicle is controlled. For example, one or more autonomous driving functions are executed by controlling various aspects of the vehicle. Such examples may include controlling the steering, speed, acceleration, and / or braking of the vehicle, maintaining the position of the vehicle within a lane, maintaining the position of the vehicle relative to other vehicles and / or obstacles, and providing notifications or warnings to passengers. Based on the analysis performed at 309, the steering and speed of a vehicle may be controlled to safely maintain the vehicle between two lane boundary lines at a safe distance from other objects. For example, the distance and direction to surrounding objects are predicted, and the corresponding drivable space and driving route are identified. In various embodiments, a vehicle control module, such as vehicle control module 111 of FIG. 1, controls the vehicle.
[0050] FIG. 4 is a flowchart showing one embodiment of a process for training and applying a machine learning model used for autonomous driving. In some embodiments, sensor data for training a machine learning model used for autonomous driving is collected and retained using the process of FIG. 4. In some embodiments, the process of FIG. 4 is executed in a vehicle in which autonomous driving is available, regardless of whether autonomous driving control is enabled. For example, sensor data may be collected while a vehicle is being driven by a human driver and / or immediately after autonomous driving is deactivated while the vehicle is being autonomously driven. In some embodiments, the techniques described by FIG. 4 are executed using the deep learning system of FIG. 1. In some embodiments, a portion of the process of FIG. 4 is executed at 307, 309, and / or 311 of FIG. 3 as a portion of the process for applying a machine learning model used for autonomous driving.
[0051] At 401, sensor data is received. For example, a vehicle equipped with sensors captures sensor data and supplies this sensor data to a neural network running on the vehicle. In some embodiments, this sensor data may be visual data, ultrasonic data, radar data, LiDAR data, or other suitable sensor data. For example, an image is captured from a high-dynamic-range front-facing camera. As another example, ultrasonic data is captured from a side-facing ultrasonic sensor. In some embodiments, a plurality of sensors for capturing data are attached to the vehicle. For example, in some embodiments, eight surround cameras are attached to the vehicle, and these cameras provide a 360-degree field of view around the vehicle up to a range of 250 meters. In some embodiments, the camera sensors include a wide-angle front camera, a narrow-angle front camera, a rear-view camera, a front-side view camera, and / or a rear-side view camera. In some embodiments, ultrasonic sensors and / or radar sensors are used to capture details of the surrounding environment. For example, by attaching twelve ultrasonic sensors to the vehicle, both hard and soft objects may be detected.
[0052] In various embodiments, data captured from different sensors is associated with the captured metadata so that data captured from different sensors is associated with each other. For example, direction, field of view, frame rate, resolution, timestamp, and / or other capture metadata is received along with the sensor data. Using this metadata, sensor data in different formats can be associated with each other so that the surrounding environment of the vehicle is more easily captured. In some embodiments, this sensor data includes odometry data such as the position, orientation, change in position, and / or change in orientation of the vehicle. For example, position data is captured and then associated with other sensor data captured within the same time frame. As an example, position information and image data are associated using the position data captured at the time of capturing the image data. In various embodiments, the received sensor data is supplied for deep learning analysis.
[0053] At 403, the sensor data is preprocessed. In some embodiments, one or more preprocessing passes may be performed on this sensor data. For example, this data may be preprocessed to remove noise and correct alignment issues and / or blurring. In some embodiments, one or more different filtering passes are performed on this data. For example, a high-pass filter may be performed on the data to separate individual components of the sensor data, or a low-pass filter may be performed on the data. In various embodiments, the preprocessing steps performed at 403 are optional and / or may be incorporated into a neural network.
[0054] At 405, deep learning analysis of the sensor data is initiated. In some embodiments, this deep learning analysis is performed on the sensor data received at 401 and preprocessed at 403 as needed. In various embodiments, this deep learning analysis is performed using a neural network such as a convolutional neural network (CNN). In various embodiments, this machine learning model is trained offline using the process of FIG. 3 and deployed to the vehicle to make inferences on the sensor data. For example, the model may be trained to predict object attributes such as distance, direction, and / or speed. In some embodiments, the model is trained to identify pedestrians, moving vehicles, parked vehicles, obstacles, lane boundaries, drivable space, etc. as needed. In some embodiments, a bounding box is specified for each object identified within the image data, and the distance and direction are predicted for each identified object. In some embodiments, these bounding boxes are three-dimensional bounding boxes such as rectangular parallelepipeds. This bounding box outlines the outer surface of the identified object and may be adjusted based on the size of the object. For example, vehicles of different sizes are represented using bounding boxes (or rectangular parallelepipeds) of different sizes. In some embodiments, the object attributes estimated by the deep learning analysis are measured by the sensor and compared to the attributes received as sensor data. In various embodiments, the neural network includes multiple layers including one or more intermediate layers and / or one or more different neural networks are used to analyze the sensor data. In various embodiments, the sensor data and / or the results of the deep learning analysis are retained and transmitted at 411 such that automatic generation of training data is performed.
[0055] In various embodiments, this deep learning analysis is used to predict additional features. The features predicted here may be used to assist in autonomous driving. For example, a detected vehicle may be assigned to a lane or a road. As another example, the detected vehicle may be determined as a vehicle in a blind spot, or as a vehicle to be prioritized, or as a vehicle in the left adjacent lane, or as a vehicle in the right adjacent lane, or as a vehicle having another suitable attribute. Similarly, in the deep learning analysis, traffic lights, drivable space, pedestrians, obstacles, or other suitable features related to driving can be identified.
[0056] At 407, the results of the deep learning analysis are provided for vehicle control. For example, these results are provided to a vehicle control module to control a vehicle performing autonomous driving and / or to execute an autonomous driving function. In some embodiments, the results of the deep learning analysis at 405 are passed through one or more additional deep learning paths using one or more different machine learning models. For example, the drivable space may be specified using the identified objects and their attributes (such as distance, direction, etc.). Then, using this drivable space, the drivable route of the vehicle is specified. Similarly, in some embodiments, the predicted speed vector of the vehicle is detected. Using the route of the vehicle specified at least partially based on this predicted speed vector, interruptions are predicted and potential collisions are avoided. In some embodiments, various output results from deep learning are used to construct a three-dimensional representation of a vehicle performing autonomous driving, and this three-dimensional representation includes specified objects such as a speed limit, obstacles to be avoided, road conditions, the distance and direction to this specified object, the predicted route of the vehicle, identified traffic lights, etc. In some embodiments, the vehicle control module uses these findings to control the vehicle along the specified route. In some embodiments, this vehicle control module is the vehicle control module 111 of FIG. 1.
[0057] At 409, the vehicle is controlled. In some embodiments, a vehicle with active autonomous driving is controlled using a vehicle control module such as vehicle control module 111 of FIG. 1. With this vehicle control, for example, the speed and / or steering of a vehicle can be adjusted so that a vehicle maintains a safe distance from other vehicles and is kept within a single lane at an appropriate speed considering the surrounding environment of the vehicle. In some embodiments, these results are used to make adjustments to the vehicle anticipating that surrounding vehicles will merge into the same lane. In various embodiments, the vehicle control module uses the results of deep learning analysis to determine an appropriate method, such as driving the vehicle along a specified route at an appropriate speed. In various embodiments, the results of vehicle control, such as speed changes, application of brakes, adjustment of steering, etc., are retained and used for automatic generation of training data. In various embodiments, these vehicle control parameters may be retained and transmitted at 411 so that training data is automatically generated.
[0058] At 411, the sensor data and related data are transmitted. For example, the sensor data received at 401 is transmitted to a computer server together with the results of the deep learning analysis at 405 and / or the vehicle control parameters used at 409 so that training data is automatically generated. In some embodiments, this data is time-series data, and these variously collected data are associated together by a remote training computer server. For example, ground truth is generated by associating image data with auxiliary sensor data such as distance data, direction data, and / or speed data. In various embodiments, these collected data are wirelessly transmitted from the vehicle to a training data center via, for example, a WiFi connection or a cellular connection. In some embodiments, metadata is transmitted together with the sensor data. For example, the metadata may include vehicle control and / or operation parameters such as time, timestamp, position, vehicle type, speed, acceleration, braking, whether autonomous driving was enabled, steering angle, odometry data, etc. Additional metadata includes the time since the most recent sensor data was transmitted, vehicle type, weather conditions, road conditions, etc. In some embodiments, this transmitted data is anonymized, for example, by removing the unique identifier of the vehicle. As another example, data obtained from similar vehicle models is merged so that individual users and the use of the vehicle by the users are not identified.
[0059] In some embodiments, this data is transmitted only in response to a trigger. For example, in some embodiments, when an inaccurate prediction is made, the transmission of image sensor data and auxiliary sensor data is triggered to automatically collect data to create a curated set of examples to improve the predictions of the deep learning network. For example, a prediction performed at 405 using only image data to estimate the distance and direction to a vehicle is determined to be inaccurate by comparing this prediction to distance data from a distance measurement sensor by irradiation. If this prediction and the actual sensor data differ by more than a certain threshold, the image sensor data and associated auxiliary data are transmitted and used to automatically generate training data. In some embodiments, this trigger is used to identify individual scenarios such as sharp curves, road bifurcations, lane merges, sudden stops, intersections, or other suitable scenarios where additional training data may be useful and difficult to collect. For example, a trigger may be based on a sudden stop or release of an autonomous driving function. As another example, vehicle motion characteristics such as a change in speed or acceleration may form the basis of the trigger. In some embodiments, when a prediction is made with an accuracy below a certain threshold, the transmission of sensor data and associated auxiliary data is triggered. For example, in certain scenarios, a prediction does not have a boolean true / false result, so instead it is evaluated by obtaining the accuracy value of the prediction.
[0060] In various embodiments, these sensor data and associated auxiliary data are captured over a period of time and the entire time series data is transmitted together. This time interval may be set based on one or more factors such as the speed of the vehicle, the distance traveled, the change in speed, and / or may be based on. In some embodiments, the sampling rate of the sensor data and / or associated auxiliary data to be captured is configurable. For example, this sampling rate may be higher or faster during hard braking, hard acceleration, sharp steering, or other suitable scenarios where greater reproducibility is required.
[0061] FIG. 5 is a diagram showing an example of capturing auxiliary sensor data for training a machine learning network. In the illustrated example, the autonomous vehicle 501 includes at least sensors 503 and 553, and also captures sensor data used to measure object attributes of surrounding vehicles 511, 521, and 561. In some embodiments, the sensor data captured here is captured and processed using a deep learning system such as the deep learning system 100 of FIG. 1 installed in the autonomous vehicle 501. In some embodiments, sensors 503 and 553 are the additional sensors 103 of FIG. 1. In some embodiments, the data captured here is data related to a part of the visual data received at 203 in FIG. 2 and / or the sensor data received at 401 in FIG. 4.
[0062] In some embodiments, sensors 503 and 553 of the autonomous vehicle 501 are distance measurement sensors by irradiation such as radar sensors, ultrasonic sensors, and / or lidar sensors. Sensor 503 is a front-facing sensor, and sensor 553 is a right-side-facing sensor. Additional sensors such as a rear-facing sensor and a left-side-facing sensor (not shown) may be attached to the autonomous vehicle 501. Axes 505 and 507 indicated by long dashed arrows are reference axes of the autonomous vehicle 501 and may be used as reference axes for data captured using sensor 503 and / or sensor 553. In the illustrated embodiment, axes 505 and 507 are at the center in front of sensor 503 and the autonomous vehicle 501. In some embodiments, an additional height axis (not shown) is used to track the attributes in three dimensions. In various embodiments, other axes may be used. For example, this reference axis may be the center of the autonomous vehicle 501. In some embodiments, each of sensors 503 and 553 may use its own reference axis and coordinate system. The data captured and analyzed using the respective local coordinate systems of sensors 503 and 553 may be converted into the local (or world) coordinate system of the autonomous vehicle 501 so that the data captured from different sensors can be shared using the same reference frame.
[0063] In the illustrated embodiment, the fields of view 509 and 559 of sensors 503 and 553 are respectively indicated by the dotted arcs between the dotted arrows. The illustrated fields of view 509 and 559 show the overhead views of the areas measured by sensors 503 and 553 respectively. The attributes of the objects within the field of view 509 may be captured by sensor 503, and the attributes of the objects within the field of view 559 may be captured by sensor 553. For example, in some embodiments, distance measurements, direction measurements, and / or speed measurements to the objects within the field of view 509 are captured by sensor 503. In the illustrated embodiment, sensor 503 captures the distance and direction to the surrounding vehicles 511 and 521. Sensor 503 does not measure the surrounding vehicle 561 because this surrounding vehicle 561 is outside the field of view area 509. Instead, the distance and direction to the surrounding vehicle 561 are captured by sensor 553. In various embodiments, objects not captured by one sensor may be captured by another sensor within the vehicle. Although only sensors 503 and 553 are shown in FIG. 5, the autonomous vehicle 501 may be equipped with a plurality of surround sensors (not shown) that provide a 360-degree field of view around the vehicle.
[0064] In some embodiments, sensors 503 and 553 capture distance measurement values and direction measurement values. Distance vector 513 indicates the distance and direction to surrounding vehicle 511, distance vector 523 indicates the distance and direction to surrounding vehicle 521, and distance vector 563 indicates the distance and direction to surrounding vehicle 561. In various embodiments, the actual distance values and direction values captured are a set of values corresponding to the outer surfaces detected by sensors 503 and 553. In the illustrated example, the set of distances and directions measured for each surrounding vehicle is approximated by distance vectors 513, 523, and 563. In some embodiments, sensors 503 and 553 detect velocity vectors (not shown) of objects within their respective fields of view 509 and 559. In some embodiments, these distance vectors and velocity vectors are three-dimensional vectors. For example, these vectors include a height (or elevation) component (not shown).
[0065] In some embodiments, detected objects including detected surrounding vehicles 511, 521, and 561 are approximated by bounding boxes. These bounding boxes approximate the outside of the detected objects. In some embodiments, these bounding boxes are three-dimensional bounding boxes such as rectangular parallelepipeds, or other solid representations of the detected objects. In the example of FIG. 5, these bounding boxes are shown as rectangles around surrounding vehicles 511, 521, and 561. In various embodiments, the distance and direction from the autonomous vehicle 501 can be measured for each point on the edge (or surface) of the bounding box.
[0066] In various embodiments, the distance vectors 513, 523, and 563 are data related to visual data captured at the same time point. The distance vectors 513, 523, and 563 are used to annotate the distances and directions to the surrounding vehicles 511, 521, and 561 identified in the corresponding visual data. For example, the distance vectors 513, 523, and 563 may be used as ground truth for annotating training images including the surrounding vehicles 511, 521, and 561. In some embodiments, the training images corresponding to the captured sensor data of FIG. 5 are captured from sensors with overlapping fields of view and utilize data captured at the matching time. For example, if the training image is image data captured from a front-facing camera that captures only the surrounding vehicles 511 and 521 but not the surrounding vehicle 561, only the surrounding vehicles 511 and 521 are identified in the training image, and the corresponding distances and directions to them are annotated. Similarly, the right-side image capturing the surrounding vehicle 561 includes annotations for the distances and directions only to the surrounding vehicle 561. In various embodiments, the annotated training images are transmitted to a training server to train a machine learning model to predict the annotated object attributes. In some embodiments, the captured sensor data and the corresponding visual data of FIG. 5 are transmitted to a training platform where these data are analyzed, training images are selected, and annotated. For example, the data captured here may be time-series data, and this time series is analyzed to associate relevant data with the objects identified in the visual data.
[0067] FIG. 6 is a diagram showing an example of predicting object attributes. In the illustrated example, the analyzed visual data 601 represents the field of view of image data captured from a visual sensor such as a front-facing camera of an autonomous vehicle. In some embodiments, the visual sensor is one of the visual sensors 101 of FIG. 1. In some embodiments, the front environment of the vehicle is captured and processed using a deep learning system such as the deep learning system 100 of FIG. 1. In various embodiments, the process shown in FIG. 6 is executed at 307, 309, and / or 311 of FIG. 3 and / or at 401, 403, 405, 407, and / or 409 of FIG. 4.
[0068] In the illustrated embodiment, the analyzed visual data 601 captures the front-facing environment of an autonomous vehicle. The analyzed visual data 601 includes lane boundaries 603, 605, 607, and 609 of the detected vehicle. In some embodiments, these lane boundaries are identified using a deep learning system, such as the deep learning system 100 of FIG. 1, that has been trained to identify driving functions. The analyzed visual data 601 includes bounding boxes 611, 613, 615, 617, and 619 corresponding to detected objects. In various embodiments, the detected objects represented by the bounding boxes 611, 613, 615, 617, and 619 are identified by analyzing the captured visual data. The captured visual data is used as input data to a trained machine learning model to predict object attributes, such as the distance and direction to the detected object. In some embodiments, a velocity vector is predicted. In the illustrated embodiment, the detected objects in the bounding boxes 611, 613, 615, 617, and 619 correspond to surrounding vehicles. The bounding boxes 611, 613, and 617 correspond to vehicles within the lane defined by the lane boundaries 603 and 605. The bounding boxes 615 and 619 correspond to vehicles within the merging lane defined by the lane boundaries 607 and 609. In some embodiments, the use of the bounding box indicates that the detected object is a three-dimensional bounding box (not shown).
[0069] In various embodiments, the object attributes predicted for the bounding boxes 611, 613, 615, 617, and 619 are predicted by applying a machine learning model trained using the processes of FIGS. 2-4. The object attributes predicted herein may be captured using an auxiliary sensor as shown in the figure of FIG. 5. FIGS. 5 and 6 show different driving scenarios, where FIG. 5 shows a different number of detected objects and these are in different positions compared to FIG. 6, and the trained machine learning model can accurately predict the object attributes of the objects detected in the scenario of FIG. 6 when trained with sufficient training data. In some embodiments, distance and direction are predicted. In some embodiments, speed is predicted. The attributes predicted herein may be predicted in two or three dimensions. By automating the generation of training data using the processes described with respect to FIGS. 1-6, training data for making accurate predictions is generated in an efficient and appropriate manner. In some embodiments, an automated driving function such as the automated driving or driver assistance operation of a vehicle may be performed using the identified objects and corresponding attributes. For example, the steering and speed of a vehicle may be controlled to safely maintain the vehicle between two lane boundaries at a safe distance from other objects.
[0070] For the sake of clarity of understanding of the foregoing embodiments, they have been described in some detail, but the present invention is not limited to the details shown. There are many alternative ways of implementing the present invention. The disclosed embodiments are exemplary and not limiting.
Claims
1. One or more processors, wherein the one or more processors receive image data captured using a camera of a vehicle, the image data representing an object in the surrounding environment of the vehicle, and are configured to utilize the time-series image data as an input to a trained machine learning model, the trained machine learning model using only the time-series image data to output a distance from the vehicle to the object and a velocity vector of the object relative to the vehicle, the trained machine learning model having been trained using time-series training image data, corresponding velocity data of the object, and a correlated output of a distance measurement sensor by irradiation, one or more processors; a memory coupled to the one or more processors; A system comprising.
2. The system according to claim 1, wherein the trained machine learning model outputs a direction of the object relative to the vehicle.
3. The system according to claim 1 or 2, wherein the trained machine learning model is trained to predict the velocity vector of the object relative to the vehicle.
4. The system according to any one of claims 1 to 3, wherein the distance measurement sensor by irradiation includes a radar sensor.
5. The system according to any one of claims 1 to 3, wherein the distance measurement sensor by irradiation includes an ultrasonic sensor.
6. The system according to any one of claims 1 to 3, wherein the distance measurement sensor by irradiation includes a lidar sensor.
7. The system according to any one of claims 1 to 6, wherein the object is one of a pedestrian, a vehicle, an obstacle, a barrier, or a traffic control object.
8. The system according to any one of claims 1 to 7, wherein the training image data and the correlated output of the distance measurement sensor by irradiation are captured by a certain training vehicle.
9. The correlation output of the distance measurement sensor by the irradiation includes an estimated distance from the training vehicle to a second object and an estimated direction of the second object with respect to the training vehicle, and the second object is specified within the training image data. The system according to claim 8.
10. The system according to claim 9, wherein the training image data is one of the time-series image data captured using an image sensor of the training vehicle.
11. The system according to claim 10, wherein the second object is at least partially specified by analyzing a plurality of image data included in the time-series image data.
12. By using the time-series image data, a portion of the correlation output of the distance measurement sensor by the irradiation corresponding to the second object among a plurality of objects detected by the distance measurement sensor by the irradiation is discriminated. The system according to claim 10.
13. The system according to any one of claims 1 to 12, wherein a semantic label associated with the object is predicted using the distance from the vehicle to the object.
14. The system according to any one of claims 1 to 13, wherein the trained machine learning model is trained to predict the distance to the object using the correlation output of the distance measurement sensor by the irradiation based on input image data.
15. The relative position of the object with respect to the vehicle is obtained by using the distance to the object specified using the trained machine learning model in combination with the correlation output of the distance measurement sensor by the irradiation. The system according to any one of claims 1 to 14.
16. A computer program embodied in a non-transitory computer-readable storage medium and comprising computer instructions, the instructions being receiving image data captured using a camera of a vehicle, the image data representing an object in the surrounding environment of the vehicle, and Use the time-series image data as an input to a trained machine learning model, and the trained machine learning model uses only the time-series image data to output the distance from the vehicle to the object and the velocity vector of the object relative to the vehicle. A computer program, wherein the trained machine learning model has been trained using time-series training image data, corresponding velocity data of the object, and a correlated output of a distance measurement sensor by irradiation.
17. Receiving selected image data based on time-series image data captured using a camera of a vehicle; Receiving distance data based on a distance measurement sensor by irradiation of the vehicle; Using the selected image data as input data to a trained machine learning model to identify an object; Extracting an estimated distance from the vehicle to the identified object from the received distance data; Creating time-series training image data by annotating the selected image data with the extracted estimated distance; Training a second machine learning model to predict the distance data and the velocity vector of the object relative to the vehicle using the time-series training image data and a training data set including corresponding velocity data of the object; Supplying the trained second machine learning model to a second vehicle equipped with a second camera; A method comprising.
18. The method according to claim 17, wherein the selected image data is a part of the time-series image data.
19. The method according to claim 18, further comprising the step of tracking the identified object over the time-series image data, wherein the extracted estimated distance is based on correlating the identified object to be tracked with the received distance data.
20. The method according to any one of claims 17 to 19, wherein the identified object is a detected vehicle or a pedestrian.
21. The system of claim 1, wherein the trained machine learning model outputs a plurality of distances corresponding to points on a bounding box associated with the object as represented in the image data.
22. The computer program of claim 16, wherein the trained machine learning model outputs a plurality of distances corresponding to points on a bounding box associated with the object as represented in the image data.
23. A method executed by a processor included in a vehicle, the method comprising: receiving image data captured using a camera of the vehicle, the image data representing an object in the surrounding environment of the vehicle; utilizing the image data in a time series as an input to a trained machine learning model, the trained machine learning model outputting a distance from the vehicle to the object and a velocity vector of the object relative to the vehicle without using distance information from a distance measurement sensor by a first irradiation; the method, wherein the trained machine learning model has been trained using time series of training image data, corresponding velocity data of the object, and a correlation output of a distance measurement sensor by a second irradiation.
Citation Information
Patent Citations
Vehicle drive assist system
JP2010083312A
Image processing apparatus, imaging apparatus, moving body device control system, image processing method, and program
JP2017207874A
Object detection device
JP2018072105A
Vehicle collision avoidance device and vehicle collision avoidance method
JP2018097648A
Offline combination of convolution / de-convolution layer and batch normalization layer of convolution neutral network model used for automatic driving vehicle
JP2018173946A