Generating ground truth for machine learning from time series elements
By using time-series elements from vehicle sensors to create a training dataset with ground truth, the method addresses the labor-intensive challenge of data curation in autonomous driving, improving lane detection and path prediction accuracy and reducing computational resources.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2026-03-13
AI Technical Summary
The performance of deep learning systems in autonomous driving is limited by the quality of the training data, which is often manually curated and labor-intensive, making it difficult to collect and accurately label data for machine learning models.
A method is introduced to generate training data using time-series elements captured by vehicle sensors, creating a training dataset that includes ground truth based on a group of images over a period, allowing for accurate labeling and reducing manual effort by using sensor data such as images and odometry to determine three-dimensional representations of features like lane lines, which are then used to train machine learning models for autonomous driving.
This approach improves the accuracy of lane detection and path prediction, reducing computational resources and enhancing the safety of autonomous vehicles by providing highly accurate training data for machine learning models.
Smart Images

Figure 0007829617000001 
Figure 0007829617000002 
Figure 0007829617000003
Abstract
Description
Background Art
[0004] , ,
[0005] , ,
[0001] [Cross - Reference to Related Applications] This application is a continuation of U.S. Patent Application No. 16 / 265729, filed on February 1, 2019, entitled "GENERATING GROUND TRUTH FOR MACHINE LEARNING FROM TIME SERIES ELEMENTS", claims the priority thereof, and the entire disclosure thereof is incorporated herein by reference.
Summary of the Invention
Problems to be Solved by the Invention
[0002] Deep learning systems used in applications such as autonomous driving are developed by training machine learning models. Typically, the performance of a deep learning system is at least partially limited by the quality of the training set used to train the model. In many cases, a significant amount of resources are invested in the collection, curation, and annotation of training data. Conventionally, much of the effort to curate a training data set is done manually by considering potential training data and appropriately labeling the features associated with the data. The effort required to create a training set with accurate labels can be significant and is often cumbersome. Additionally, it is often difficult to collect and accurately label data that requires improvement for a machine learning model. Therefore, there is a need to improve the process for generating training data with accurately labeled features.
Brief Description of the Drawings
[0003] Various embodiments of the present invention are disclosed in the following detailed description and the accompanying drawings.
[0004] [Figure 1] A block diagram showing one embodiment of a deep learning system for autonomous driving.
[0005] [Figure 2] This flowchart illustrates one embodiment of the process for training and applying machine learning models for autonomous driving.
[0006] [Figure 3] This flowchart illustrates one embodiment of the process of creating training data using time-series elements.
[0007] [Figure 4] This flowchart illustrates one embodiment of the process for training and applying machine learning models for autonomous driving.
[0008] [Figure 5] This figure shows an example of an image captured by a vehicle sensor.
[0009] [Figure 6] This figure shows an example of an image captured by a vehicle sensor that has a predicted 3D trajectory of the lane line. [Modes for carrying out the invention]
[0010] The present invention can be implemented in various ways, including processes, apparatus, systems, compositions of materials, computer program products embodied on computer-readable storage media, and / or processors, such as processors configured to execute instructions stored in memory coupled to the processor and / or instructions provided by memory. In this specification, these implementations, or any other forms the present invention may take, may be referred to as the "Technology." Generally, the order of the steps of the disclosed process can be changed within the scope of the invention. Unless otherwise specified, the steps described are configured to perform the task. The components, such as processors or memory, may be implemented as general components temporarily configured to perform a task in a given time, or as specific components manufactured to perform a task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0011] A detailed description of one or more embodiments of the present invention is provided below, along with accompanying drawings illustrating the principles of the present invention. While the present invention is described in relation to such embodiments, it is not limited to any embodiment. The scope of the present invention is limited solely by the claims, and the present invention encompasses numerous alternative forms, variations, and equivalents. Numerous specific details are described below to provide a complete understanding of the present invention. These details are provided for illustrative purposes only, and the present invention can be carried out in accordance with the claims without some or all of these specific details. For clarity, known technical materials in the art relating to the present invention are not described in detail so as not to unnecessarily obscure the present invention.
[0012] Machine learning training techniques for generating highly accurate machine learning results are disclosed. A training dataset is created using data captured by sensors on the vehicle to capture the vehicle's environment and vehicle operation parameters. For example, sensors mounted on the vehicle capture data such as image data of the road and surrounding environment on which the vehicle is traveling. Sensor data can capture vehicle lane lines, vehicle lanes, other vehicle traffic, obstacles, traffic control signs, etc. Odometry and other similar sensors capture vehicle operation parameters such as vehicle speed, steering, orientation, change of direction, change of position, change of altitude, and change of speed. The captured dataset is sent to a training server to create a training dataset. The training dataset is used to train a machine learning model that generates highly accurate machine learning results. In some embodiments, time-series captured data is used to generate training data. For example, ground truth is determined based on a group of time-series elements and associated with a single element from the group. As an example, a series of images over a period such as 30 seconds is used to determine the actual path of the vehicle lane line over the period the vehicle is traveling. Vehicle lane lines are determined by using the most accurate images of the vehicle lane over a given period. Different portions (or locations) of the lane line can be identified from different image data in a time series. As the vehicle travels along the lane line, more accurate data is captured for different portions of the lane line. In some examples, obstructed portions of the lane appear when the vehicle travels, for example, along a hidden curve or over the top of a hill. The most accurate portion of the lane line from each image in the time series can be used to identify the lane line across the entire group of image data. Image data of distant lane lines is typically less detailed than image data of lane lines near the vehicle. By capturing time-series image data as the vehicle travels along the lane, accurate image data and corresponding odometry data for all portions of the corresponding lane line are collected.
[0013] In some embodiments, a three-dimensional representation of a feature, such as a lane line, is created from a group of time-series elements corresponding to ground truth. This ground truth is then associated with a subset of time-series elements, such as a single image frame from a group of captured image data. For example, the first image in the image group is associated with the ground truth of a lane line represented in three-dimensional space. The ground truth is determined based on the group of images, but the selected first frame and ground truth are used to create training data. As an example, training data is created to predict a three-dimensional representation of a vehicle lane using only a single image. In some embodiments, any element or group of elements in the group of time-series elements corresponds to ground truth. It is associated with and used to create training data. For example, ground truth can be applied to an entire video sequence to create training data. Another example is when the middle or last element of a group of time-series elements is associated with ground truth and used to create training data.
[0014] In various embodiments, selected images and ground truth can be applied to various features such as lane lines, vehicle path prediction including adjacent vehicles, object depth distances, and traffic signs. For example, a series of images of a vehicle in an adjacent lane can be used to predict the vehicle's path. Using a time-series of images taken by adjacent vehicles and the actual path, a group of images taken and a single image of the actual path can be used as training data for predicting the vehicle's path. This information can also be used to predict whether an adjacent vehicle will cut into the autonomous vehicle's path. For example, path prediction can predict whether an adjacent vehicle will merge in front of the autonomous vehicle. The autonomous vehicle can be controlled to minimize the likelihood of a collision. For example, the autonomous vehicle can slow down to prevent a collision, adjust the vehicle's speed and / or steering to prevent a collision, initiate warnings to adjacent vehicles and / or occupants of the autonomous vehicle, and / or change lanes. In various embodiments, the ability to accurately infer path predictions, including vehicle path prediction, greatly improves the safety of the autonomous vehicle.
[0015] In some embodiments, a trained machine learning model is used to predict a three-dimensional representation of one or more features for autonomous driving, including lane lines. For example, instead of identifying two-dimensional lane lines from image data by segmenting images of lane lines, a three-dimensional representation is generated using time-series elements and corresponding odometry data. The three-dimensional representation includes a high degree of variation that significantly improves the accuracy of lane line detection and the detection of corresponding lanes and identified drivable paths. In some embodiments, lane lines are represented using one or more splines or another parameterized representation format. Using piecewise polynomials to represent lane lines significantly reduces the computational resources required to evaluate three-dimensional objects. This reduction in computational resources corresponds to improvements in processing speed and efficiency without significantly sacrificing the accuracy of the representation. In various embodiments, lane lines, particularly those including curves, can be represented using piecewise polynomials, three-dimensional point sets, or another suitable representation. For example, a piecewise polynomial interpolates the actual lane lines using high-precision sections of lane lines identified from groups of elements captured over time using sensor data.
[0016] In some embodiments, sensor data is received. Sensor data may include images (such as video and / or still images), radar, audio, LiDAR, inertia, odometry, position, and / or other forms of sensor data. Sensor data may include a group of time-series elements. For example, a group of time-series elements may include a group of images captured by a vehicle's camera sensor over a period of time. In some embodiments, the training dataset is determined for at least selected time-series elements within a group of time-series elements, for example, by determining the corresponding ground truth based on multiple time-series elements within the group of time-series elements. For example, the ground truth is determined by examining the most relevant portion of each element in a group of time-series elements, including preceding and / or succeeding time-series elements within the group. In some scenarios, only preceding and / or succeeding time-series elements include data not present in preceding time-series elements, such as vehicle lane lines that first disappear around a curve and appear only in succeeding time-series elements. The determined ground truth may be a three-dimensional representation of vehicle lane lines, a predicted vehicle path, or another similar prediction. Elements from a group of time-series elements are selected and associated with the ground truth. The selected elements and the ground truth are part of the training dataset. In some embodiments, the processor is used to train a machine learning model using a training dataset. For example, the training dataset is used to train a machine learning model to infer features used for autonomous driving or driver assistance actions of a vehicle. Using the trained machine learning model, the neural network can infer features related to autonomous driving, such as vehicle lanes, drivable space, objects (e.g., pedestrians, stationary vehicles, moving vehicles, etc.), weather (e.g., rain, hail, fog, etc.), traffic control objects (e.g., traffic lights, traffic signs, road signs, etc.), and traffic patterns.
[0017] In some embodiments, the system comprises a processor and memory coupled to the processor. The processor is configured to receive image data based on images captured by a camera on the vehicle. For example, a camera sensor mounted on the vehicle captures images of the vehicle's environment. The camera may be a forward-facing camera, a pillar camera, or another appropriately positioned camera. The image data captured from the camera is processed using a processor such as a GPU or AI processor on the vehicle. In some embodiments, the image data is used as the basis for input to a trained machine learning model trained to predict the three-dimensional trajectory of the vehicle lane. For example, the image data is used as input to a neural network trained to predict the vehicle lane. The machine learning model infers the three-dimensional trajectory of the detected lane. Instead of segmenting the image into lane and non-lane segments of a two-dimensional image, a three-dimensional representation is inferred. In some embodiments, the three-dimensional representation is a spline, a parametric curve, or another representation that can describe a curve in three dimensions. In some embodiments, the three-dimensional trajectory of the vehicle lane is provided when automatically controlling the vehicle. For example, the three-dimensional trajectory is used to determine the lane line and the corresponding drivable space.
[0018] Figure 1 is a block diagram showing one embodiment of a deep learning system for autonomous driving. The deep learning system includes various components that can be used together for autonomous driving and / or driver assistance operations of a vehicle, as well as for collecting and processing data to train machine learning models for autonomous driving. In various embodiments, the deep learning system is installed in a vehicle. Data from the vehicle can be used to train and improve the autonomous driving capabilities of the vehicle or other similar vehicles.
[0019] In the illustrated example, the deep learning system 100 is a deep learning network comprising a sensor 101, an image preprocessor 103, a deep learning network 105, an artificial intelligence (AI) processor 107, a vehicle control module 109, and a network interface 111. In various embodiments, different components are connected in a communicative manner. For example, sensor data from sensor 101 is supplied to the image preprocessor 103. The processed sensor data from the image preprocessor 103 is supplied to the deep learning network 105 operating on the AI processor 107. The output of the deep learning network 105 operating on the AI processor 107 is supplied to the vehicle control module 109. In various embodiments, the vehicle control module 109 is connected to and controls vehicle behavior such as vehicle speed, braking, and / or steering. In various embodiments, sensor data and / or machine learning results can be transmitted to a remote server via the network interface 111. For example, sensor data can be transmitted to a remote server via the network interface 111 to collect training data to improve vehicle performance, comfort, and / or safety. In various embodiments, the network interface 111 is used, among other reasons, to communicate with a remote server, make phone calls, send and / or receive text messages, and transmit sensor data based on vehicle operation. In some embodiments, the deep learning system 100 may include additional or fewer components as needed. For example, in some embodiments, the image preprocessor 103 is an optional component. As another example, In one embodiment, a post-processing component (not shown) is used to perform post-processing on the output of the deep learning network 105 before the output is provided to the vehicle control module 109.
[0020] In some embodiments, sensor 101 includes one or more sensors. In various embodiments, sensor 101 may be attached to the vehicle at different positions of the vehicle and / or oriented in one or more different directions. For example, sensor 101 may be attached in a direction such as forward, backward, or sideways, to the front, side, rear, and / or roof of the vehicle. In some embodiments, sensor 101 may be an image sensor such as a high dynamic range camera. In some embodiments, sensor 101 includes non-visual sensors. In some embodiments, sensor 101 includes, among other things, radar, audio, LiDAR, inertial, odometry, position, and / or ultrasonic sensors. In some embodiments, sensor 101 is not attached to a vehicle having vehicle control module 109. For example, sensor 101 may be attached to an adjacent vehicle and / or attached to a road or environment and included as part of a deep learning system for capturing sensor data. In some embodiments, sensor 101 includes one or more cameras that capture the road surface on which the vehicle is traveling. For example, one or more forward and / or pillar cameras capture the lane markings of the lane in which the vehicle is traveling. As another example, the camera captures adjacent vehicles including a vehicle attempting to cut into the lane in which the vehicle is traveling. Additional sensors capture vehicle control information including information regarding odometry, position, and / or vehicle trajectory. Sensor 101 can include both image sensors that can capture still images and / or videos. The data can be captured over a period of time, such as a sequence of data captured over a period of time. For example, an image of lane markings may be captured along with vehicle odometry data over a period of 15 seconds or another appropriate period. In some embodiments, sensor 101 includes a position sensor such as a global positioning system (GPS) sensor for determining the position and / or change in position of the vehicle.
[0021] In some embodiments, the image preprocessor 103 is used to preprocess sensor data from the sensor 101. For example, the image preprocessor 103 can be used to preprocess sensor data, split the sensor data into one or more components, and / or postprocess one or more components. In some embodiments, the image preprocessor 103 is a graphics processing unit (GPU), a central processing unit (CPU), an image signal processor, or a dedicated image processor. In various embodiments, the image preprocessor 103 is a tone mapper processor for processing high dynamic range data. In some embodiments, the image preprocessor 103 is implemented as part of an artificial intelligence (AI) processor 107. For example, the image preprocessor 103 may be a component of the AI processor 107. In some embodiments, the image preprocessor 103 can be used to normalize or transform an image. For example, an image captured with a fisheye lens may be distorted, and the image preprocessor 103 can be used to transform the image and remove or correct the distortion. In some embodiments, noise, distortion, and / or blur are removed or reduced during the preprocessing step. In various embodiments, images are adjusted or normalized to improve the results of machine learning analysis. For example, the white balance of the images is adjusted to take into account different lighting operating conditions, such as daylight, sunny, cloudy, twilight, sunrise, sunset, and nighttime conditions.
[0022] In some embodiments, the deep learning network 105 is used to determine vehicle control parameters, including analyzing the driving environment to determine lane markers, lanes, drivable space, obstacles, and / or potential vehicle paths. It is a workpiece. For example, the deep learning network 105 may be an artificial neural network such as a convolutional neural network (CNN) that is trained with inputs such as sensor data and whose output is provided to the vehicle control module 109. As an example, the output can include at least a three-dimensional representation of the lane markers. As another example, the output can include at least potential vehicles that may merge into the vehicle's lane. In some embodiments, the deep learning network 105 receives at least sensor data as an input. The additional input can include scene data that describes vehicle specifications such as the environment around the vehicle and / or the operating characteristics of the vehicle. The scene data can include scene tags that describe the environment around the vehicle such as rainfall, wet roads, snowfall, mud, high-density traffic, highways, cities, school zones, etc. In some embodiments, the output of the deep learning network 105 is the three-dimensional trajectory of the vehicle's lane. In some embodiments, the output of the deep learning network 105 is a potential vehicle interruption. For example, the deep learning network 105 identifies adjacent vehicles that are likely to enter the lane in front of the vehicle.
[0023] In some embodiments, the artificial intelligence (AI) processor 107 is a hardware processor for running a deep learning network 105. In some embodiments, the AI processor 107 is a dedicated AI processor for performing inference using a convolutional neural network (CNN) on sensor data. The AI processor 107 can be optimized for the bit depth of the sensor data. In some embodiments, the AI processor 107 is optimized for deep learning operations, such as neural network operations including convolution, dot product, vector, and / or matrix operations, among others. In some embodiments, the AI processor 107 is implemented using a graphics processing unit (GPU). In various embodiments, the AI processor 107 is coupled to memory configured, when executed, to provide the AI processor with instructions to perform deep learning analysis on received input sensor data and determine machine learning results to be used for autonomous driving. In some embodiments, the AI processor 107 is used to process the sensor data in preparation for making it available as training data.
[0024] In some embodiments, the vehicle control module 109 is used to process the output of the artificial intelligence (AI) processor 107 and convert the output into vehicle control actions. In some embodiments, the vehicle control module 109 is used to control the vehicle for autonomous driving. In various embodiments, the vehicle control module 109 can adjust the vehicle's speed, acceleration, steering, braking, etc. For example, in some embodiments, the vehicle control module 109 is used to control the vehicle to maintain its position in a lane, to merge a vehicle into another lane, and to adjust the vehicle's speed and lane alignment to account for merging vehicles, etc.
[0025] In some embodiments, the vehicle control module 109 is used to control vehicle lighting such as brake lights, turn signals, and headlights. In some embodiments, the vehicle control module 109 is used to control vehicle audio states such as the vehicle's sound system, playback of audio warnings, microphone activation, and horn activation. In some embodiments, the vehicle control module 109 is used to control notification systems, including warning systems for notifying the driver and / or passengers of driving events such as potential collisions or approaching an intended destination. In some embodiments, the vehicle control module 109 is used to adjust sensors such as vehicle sensors 101. For example, the vehicle control module 109 can be used to change the parameters of one or more sensors, such as changing orientation, changing output resolution and / or format type, increasing or decreasing capture rate, adjusting captured dynamic range, adjusting camera focus, and enabling and / or disabling sensors. In some embodiments, the vehicle control module 109 can be used to change the frequency range of filters, feature and / or edge detection. The parameters of the image preprocessor 103 can be modified, such as adjusting parameters, channels, and bit depth. In various embodiments, the vehicle control module 109 is used to implement autonomous driving and / or driver assistance control of the vehicle. In some embodiments, the vehicle control module 109 is implemented using a processor coupled with memory. In some embodiments, the vehicle control module 109 is implemented using an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or other suitable processing hardware.
[0026] In some embodiments, the network interface 111 is a communication interface for transmitting and / or receiving data, including voice data. In various embodiments, the network interface 111 includes a cellular or wireless interface for interface-connecting with a remote server, making and making voice calls, sending and / or receiving text messages, transmitting sensor data, receiving updates to a deep learning network including an updated machine learning model, and retrieving environmental conditions, including weather conditions and forecasts, traffic conditions, etc. For example, the network interface 111 may be used to receive updates to instructions and / or operating parameters of a sensor 101, an image preprocessor 103, a deep learning network 105, an AI processor 107, and / or a vehicle control module 109. The machine learning model of the deep learning network 105 can be updated using the network interface 111. As another example, the network interface 111 can be used to update the firmware of the sensor 101 and / or operating parameters of the image preprocessor 103, such as image processing parameters. As yet another example, the network interface 111 can be used to send potential training data to a remote server to train a machine learning model.
[0027] Figure 2 is a flowchart illustrating one embodiment of the process for training and applying a machine learning model for autonomous driving. For example, input data, including sensor and odometry data, is received and processed to create training data for training a machine learning model. In some embodiments, the sensor data corresponds to image data captured via an autonomous driving system. In some embodiments, the sensor data corresponds to sensor data captured based on a specific use case, such as a user manually disengaging autonomous driving. In some embodiments, the process is used to create and deploy a machine learning model for the deep learning system 100 in Figure 1.
[0028] In 201, training data is prepared. In some embodiments, sensor data, including image data and odometry data, is received to create a training dataset. The sensor data may include still images and / or videos from one or more cameras. Additional sensors, such as radar, LiDAR, and ultrasound, may be used to provide relevant sensor data. In various embodiments, the sensor data is paired with corresponding odometry data to help identify features of the sensor data. For example, position and change of position data can be used to identify the location of relevant features in the sensor data, such as lane lines, traffic control signals, and objects. In some embodiments, the sensor data is a time series element used to determine ground truth. The group ground truth is then associated with a subset of the time series, such as a first frame of image data. The selected elements of the time series and ground truth are used to prepare the training data. In some embodiments, the training data is prepared to train a machine learning model to identify only features from the sensor data, such as lane lines, vehicle paths, and traffic patterns. The prepared training data may include data for training, validation, and testing. In various embodiments, the sensor data may be in different formats. For example, the sensor data may include still images and videos. The data may include images, audio, etc. Odometry data may include vehicle operation parameters such as applied acceleration, applied braking, applied steering, vehicle position, vehicle orientation, change in vehicle position, and change in vehicle orientation. In various embodiments, training data is curated and annotated to create a training dataset. In some embodiments, part of the preparation of the training data may be performed by a human curator. In various embodiments, part of the training data is automatically generated from data captured from the vehicle, significantly reducing the effort and time required to build a robust training dataset. In some embodiments, the data format is compatible with machine learning models used in deployed deep learning applications. In various embodiments, the training data includes validation data to test the accuracy of the trained model.
[0029] In 203, a machine learning model is trained. For example, the machine learning model is trained using the data prepared in 201. In some embodiments, the model is a neural network, such as a convolutional neural network (CNN). In various embodiments, the model includes multiple hidden layers. In some embodiments, the neural network may include multiple layers, including multiple convolutional and pooling layers. In some embodiments, the trained model is validated using a validation dataset created from received sensor data. In some embodiments, the machine learning model is trained to predict a three-dimensional representation of features from a single input image. For example, a three-dimensional representation of lane lines can be inferred from an image captured by a camera. As another example, the predicted path of adjacent vehicles, including whether or not a vehicle is about to merge, can be predicted from an image captured by a camera.
[0030] In 205, a trained machine learning model is deployed. For example, the trained machine learning model is installed in the vehicle as an update to a deep learning network, such as the deep learning network 105 in Figure 1. In some embodiments, a wireless update is used to install the newly trained machine learning model. In some embodiments, the update is a firmware update transmitted using a wireless network such as WiFi or a cellular network. In some embodiments, the new machine learning model may be installed when the vehicle is serviced.
[0031] In 207, sensor data is received. For example, sensor data is captured from one or more sensors on the vehicle. In some embodiments, the sensor is sensor 101 in Figure 1. The sensor may include image sensors such as a fisheye camera mounted behind the windshield, a forward-facing or side-facing camera mounted on a pillar, or a rear-facing camera. In various embodiments, the sensor data is in a format that can be used as input by a machine learning model trained in 203, or is converted to such a format. For example, the sensor data may be raw or processed image data. In some embodiments, the data is data captured from an ultrasonic sensor, radar, LiDAR sensor, microphone, or other suitable technology. In some embodiments, the sensor data is preprocessed using an image preprocessor such as image preprocessor 103 in Figure 1 during the preprocessing step. For example, the image may be normalized to remove distortion, noise, etc.
[0032] In 209, the trained machine learning model is applied. For example, the machine learning model trained in 203 is applied to the sensor data received in 207. In some embodiments, the application of the model is performed by an AI processor, such as the AI processor 107 in Figure 1, using a deep learning network, such as the deep learning network 105 in Figure 1. In various embodiments, by applying the trained machine learning model, a three-dimensional representation of a feature, such as a lane line, is identified and / or predicted. For example, two splines representing the lane line of the lane in which a vehicle is traveling are inferred. As another example, adjacent vehicles The predicted paths of adjacent vehicles are inferred, including whether both vehicles are likely to cut into the current lane. In various embodiments, a machine learning model is applied to identify vehicles, obstacles, lanes, traffic control signals, map features, object distances, speed limits, drivable space, etc. In some embodiments, features are identified in three dimensions.
[0033] In 211, the autonomous vehicle is controlled. For example, one or more autonomous driving functions are implemented by controlling various aspects of the vehicle. Examples include controlling the vehicle's steering, speed, acceleration, and / or braking, maintaining the vehicle's position in a lane, maintaining the vehicle's position relative to other vehicles and / or obstacles, and providing notifications or warnings to the occupants. Based on the analysis performed in 209, the vehicle's steering and speed are controlled to maintain the vehicle between two lane lines. For example, the left and right lane lines are predicted, and the corresponding vehicle lanes and drivable space are identified. In various embodiments, a vehicle control module, such as the vehicle control module 109 in Figure 1, controls the vehicle.
[0034] Figure 3 is a flowchart illustrating one embodiment of the process for creating training data using time-series elements. For example, time-series elements consisting of sensor and odometry data are collected from a vehicle and used to automatically create training data. In various embodiments, the process in Figure 3 is used to automatically label the training data with the corresponding ground truth. Results corresponding to the time series are associated with the time-series elements. The results and selected elements are packaged as training data for predicting future outcomes. In various embodiments, the sensor and associated data are captured using the deep learning system in Figure 1. For example, in various embodiments, sensor data is captured from sensor 101 in Figure 1. In some embodiments, the process in Figure 3 is performed at 201 in Figure 2. In some embodiments, the process in Figure 3 is performed to automatically collect data when existing predictions are incorrect or can be improved. For example, a prediction is made by an autonomous vehicle to determine whether a vehicle is cutting in the autonomous vehicle's path. After waiting for a certain period and analyzing the captured sensor data, a determination can be made as to whether the prediction is correct or incorrect. In some embodiments, a determination is made as to whether the prediction can be improved. If the prediction was incorrect or could be improved, the process in Figure 3 can be applied to the data related to the prediction to create a curated set of examples for improving the machine learning model.
[0035] In 301, time series elements are received. In various embodiments, elements are sensor data, such as image data captured by a vehicle and sent to a training server. Sensor data is captured over a period of time to create time series elements. In various embodiments, elements are timestamps to maintain the order of the elements. As the elements progress through the time series, later events in the time series are used to help predict results from earlier elements in the time series. For example, the time series may capture a vehicle in an adjacent lane signaling to merge, accelerating and positioning itself near the nearby lane line. The entire time series and the results can be used to determine that the vehicle has merged into the shared lane. This result can be used to predict that the vehicle will merge based on selected elements of the time series, such as one of the initial images in the time series. As another example, the time series captures the curves of the lane line. The time series captures various dips, bends, ridges, etc., in the lane that are not apparent from a single element of the time series alone. In various embodiments, elements are sensor data in a format that a machine learning model uses as input. For example, the sensor data may be raw or processed image data. In some embodiments, the data is data captured from an ultrasonic sensor, radar, LiDAR sensor, or other suitable technology.
[0036] In various embodiments, time series are organized by associating a timestamp with each element of the time series. For example, a timestamp is associated with at least a first element in the time series. The timestamp may be used to calibrate the time series elements with related data, such as odometry data. In various embodiments, the length of the time series may be a fixed length of time, such as 10 seconds, 30 seconds, or another appropriate length. The length of time may be configurable. In various embodiments, the time series may be based on the speed of a vehicle, such as the average speed of a vehicle. For example, slower speeds may result in an increased time length in the time series to capture data over longer distances than would be possible when using shorter time lengths for the same speed. In some embodiments, the number of elements in the time series is configurable. For example, the number of elements may be based on distance traveled. For example, a vehicle traveling faster over a given period of time will have more elements in its time series than a vehicle traveling slower. Additional elements can increase the fidelity of the captured environment and improve the accuracy of the predicted machine learning results. In various embodiments, the number of elements is adjusted by adjusting the number of frames per second that the sensor captures data and / or by discarding unnecessary intermediate frames.
[0037] In 303, data related to the time-series elements is received. In various embodiments, the related data is received at the training server along with the elements received in 301. In some embodiments, the related data is vehicle odometry data. Position, orientation, change in position, change in orientation, and / or other related vehicle data can be used to label the positional data of features identified by the time-series elements. For example, by examining the time-series of lane line elements, lane lines can be labeled with very precise position. Typically, the lane line closest to the vehicle camera is accurate and closely related to the vehicle's position. On the other hand, the XYZ position of the line furthest from the vehicle is difficult to determine. Distant portions of lane lines may be obstructed (e.g., behind a bend or hill) and / or difficult to capture accurately (e.g., due to distance or lighting). The data related to the elements is used to label portions of features identified in the time-series with high precision. In various embodiments, thresholds are used to determine whether to associate identified portions of features (e.g., portions of lane lines) with the related data. For example, portions of the lane line identified with high accuracy (such as those close to the vehicle) are associated with relevant data, while portions identified with accuracy below a threshold (such as those far from the vehicle) are not associated with relevant data for that element. Instead, another element in the time series, such as a subsequent element with higher accuracy, and its relevant data are used. In some embodiments, the relevant data is the output of a neural network, such as the output of the deep learning network 105 in Figure 1. In some embodiments, the relevant data is the output of a vehicle control module, such as the vehicle control module 109 in Figure 1. The relevant data may include vehicle operation parameters such as speed, speed change, acceleration, acceleration change, steering, steering change, braking, and braking change. In some embodiments, the relevant data is radar data for estimating the distance to an object, such as an obstacle.
[0038] In some embodiments, data related to time-series elements include map data. For example, in 303, offline data such as road and / or satellite-level map data is received. Map data can be used to identify features such as roads, vehicle lanes, intersections, speed limits, and school zones. For example, map data can describe the routes of vehicle lanes. As another example, map data can describe speed limits associated with various roads on the map.
[0039] In various embodiments, data related to time-series elements is organized by associating timestamps with the related data. The two datasets can be synchronized using the corresponding timestamps from the time-series elements and the related data. In some embodiments, the data is synchronized upon capture. For example, when each element of a time series is captured, a corresponding set of related data is captured and stored together with the time series element. In various embodiments, the duration of the related data is configurable and / or matches the duration of the time series of the element. In some embodiments, the related data is sampled at the same rate as the time series element.
[0040] In 305, ground truth is determined for a time series. In various embodiments, the time series is analyzed to determine ground truth associated with machine learning features. For example, lane lines are identified from the time series corresponding to the ground truth of those lane lines. As another example, the ground truth for the paths of moving objects (e.g., vehicles, pedestrians, bicycles, animals, etc.) is the path identified for the detected moving objects from the time series. In some embodiments, if a moving vehicle enters the lane of an autonomous vehicle over time, the moving vehicle is annotated as an interrupting vehicle. In some embodiments, ground truth is represented as a three-dimensional representation, such as a three-dimensional trajectory. For example, ground truth associated with lane lines may be represented as a three-dimensional parameterized spline or curve. As another example, the predicted path of a detected vehicle is determined and represented as a three-dimensional trajectory. The predicted path may be used to determine whether the vehicle is merging into occupied space. In various embodiments, ground truth can be determined solely by examining elements of the time series. For example, analyzing only a subset of the time series may leave parts of the lane lines obscured. Extending the analysis across elements of the time series reveals the obscured parts of the lane lines. Furthermore, data captured near the end of the time series captures details of more distant parts of the lane lines more accurately (e.g., with higher fidelity). Moreover, the relevant data is also more accurate because it is based on data captured at closer proximity (both in distance and time). In various embodiments, simultaneous localization and mapping techniques are applied to different parts of a detected object, such as lane lines identified by different elements of the time series of the elements, to map different parts of the object to precise three-dimensional locations, including altitude. The set of mapped three-dimensional locations represents the ground truth of the object, such as segments of lane lines captured over the time series.In some embodiments, localization and mapping techniques yield a precise set of points, for example, a set of points corresponding to different points along a vehicle lane line. This set of points can be converted into a more efficient format, such as a spline curve or a parametric curve. In some embodiments, ground truth is determined to detect objects in three dimensions, such as lane lines, drivable space, traffic control systems, and vehicles.
[0041] In some embodiments, ground truth is determined to predict semantic labels. For example, a detected vehicle may be labeled as being in the left or right lane. In some embodiments, a detected vehicle may be labeled as being in a blind spot, as a vehicle that should yield, or with another appropriate semantic label. In some embodiments, vehicles are assigned to roads or lanes in the map based on the determined ground truth. As an additional example, the determined ground truth can be used to label traffic lights, lanes, drivable space, or other features that assist in autonomous driving.
[0042] In some embodiments, the relevant data is depth (or distance) data of the detected object. By associating the distance data with the identified object in a time series of elements, a machine learning model can be trained to estimate the object distance by using the relevant distance data as ground truth for the detected object. In some embodiments, the distance is an obstacle, barrier, moving vehicle, stationary vehicle, traffic control signal, pedestrian This is the distance to the detected object, such as a person.
[0043] In 307, the training data is packaged. For example, time-series elements are selected and associated with the ground truth determined in 305. In various embodiments, the selected elements are initial elements in the time series. The selected elements represent sensor data input to the machine learning model, and the ground truth represents the predicted results. In various embodiments, the training data is packaged and prepared as training data. In some embodiments, the training data is packaged into training, validation, and test data. Based on the determined ground truth and the selected time-series elements, the training data can be packaged to train a machine learning model to identify lane lines, predicted vehicle paths, speed limits, vehicle interruptions, object distances, and / or drivable spaces, among other useful features for autonomous driving. The packaged training data is thus made available for training the machine learning model.
[0044] Figure 4 is a flowchart illustrating one embodiment of the process for training and applying a machine learning model for autonomous driving. In some embodiments, the process in Figure 4 is used to collect and retain sensor and odometry data for training a machine learning model for autonomous driving. In some embodiments, the process in Figure 4 is performed in a vehicle with autonomous driving enabled, regardless of whether autonomous driving control is enabled or not. For example, sensor and odometry data can be collected immediately after autonomous driving is deactivated, while the vehicle is being driven by a human driver, and / or while the vehicle is autonomous driving. In some embodiments, the technique described by Figure 4 is performed using the deep learning system in Figure 1. In some embodiments, parts of the process in Figure 4 are performed in Figures 207, 209, and / or 211 as part of the process of applying a machine learning model for autonomous driving.
[0045] In 401, sensor data is received. For example, a vehicle equipped with sensors captures sensor data and provides it to a neural network operating on the vehicle. In some embodiments, the sensor data may be visual data, ultrasonic data, LiDAR data, or other suitable sensor data. For example, images are captured from a high dynamic range forward-facing camera. Another example is the capture of ultrasonic data from a lateral ultrasonic sensor. In some embodiments, the vehicle is fitted with multiple sensors for capturing data. For example, in some embodiments, eight surround cameras are fitted to the vehicle to provide a 360-degree view around the vehicle with a range of up to 250 meters. In some embodiments, the camera sensors include a wide-angle forward camera, a narrow-angle forward camera, a rear-view camera, a forward-view side camera, and / or a rear-view side camera. In some embodiments, ultrasonic and / or radar sensors are used to capture details of the surroundings. For example, twelve ultrasonic sensors can be fitted to the vehicle to detect both hard and soft objects. In some embodiments, forward-facing radar is used to capture data of the surrounding environment. In various embodiments, radar sensors can capture surrounding details despite heavy rain, fog, dust, and other vehicles. Various sensors are used to capture the environment around the vehicle, and the captured data is provided for deep learning analysis.
[0046] In some embodiments, the sensor data includes odometry data such as the vehicle's position, orientation, change in position, and / or change in orientation. For example, position data is captured and associated with other sensor data captured during the same time frame. As an example, position data captured when image data is acquired is used to associate position information with image data.
[0047] In 403, the sensor data is preprocessed. In some embodiments, the sensor data One or more preprocessing passes can be performed on the data. For example, the data may be preprocessed to remove noise and correct alignment problems and / or blurring, etc. In some embodiments, one or more different filtering passes are performed on the data. For example, a high-pass filter may be applied to the data and a low-pass filter to the data to separate different components of the sensor data. In various embodiments, the preprocessing steps performed in 403 are optional and / or may be incorporated into a neural network.
[0048] In 405, deep learning analysis of the sensor data is initiated. In some embodiments, the deep learning analysis is performed on sensor data that has been optionally preprocessed in 403. In various embodiments, the deep learning analysis is performed using a neural network such as a convolutional neural network (CNN). In various embodiments, the machine learning model is trained offline using the process in Figure 2 and deployed in the vehicle to perform inference on the sensor data. For example, the model may be trained to appropriately identify road lane lines, obstacles, pedestrians, moving vehicles, parked vehicles, drivable space, etc. In some embodiments, multiple trajectories of the lane line are identified. For example, several potential trajectories of the lane line are detected, and each trajectory has a corresponding probability of occurrence. In some embodiments, the predicted lane line is the lane line with the highest probability of occurrence and / or the highest associated confidence value. In some embodiments, the predicted lane line from the deep learning analysis is required to exceed a minimum confidence threshold. In various embodiments, the neural network includes multiple layers, each containing one or more hidden layers. In various embodiments, sensor data and / or the results of deep learning analysis are retained for the automatic generation of training data and transmitted in 411.
[0049] In various embodiments, deep learning analysis is used to predict additional features. The predicted features can be used to assist autonomous driving. For example, detected vehicles can be assigned to lanes or roads. As another example, a detected vehicle can be determined to be in a blind spot, a vehicle that should yield, a vehicle in the lane to the left, a vehicle in the lane to the right, or to have other appropriate attributes. Similarly, deep learning analysis can identify traffic lights, drivable space, pedestrians, obstacles, or other appropriate features for driving.
[0050] In 407, the results of the deep learning analysis are provided to the vehicle control. For example, the results are provided to the vehicle control module to control the vehicle for autonomous driving and / or to implement autonomous driving functions. In some embodiments, the results of the deep learning analysis in 405 go through one or more additional deep learning paths using one or more different machine learning models. For example, the predicted path of the lane lines can be used to determine the vehicle lane, and the determined vehicle lane can be used to determine the drivable space. The drivable space is then used to determine the vehicle's path. Similarly, in some embodiments, predicted vehicle interrupts are detected. The determined path of the vehicle takes the predicted interrupts into account to avoid potential collisions. In some embodiments, the various outputs of the deep learning are used to construct a three-dimensional representation of the vehicle environment for autonomous driving, including the predicted path of the vehicle, identified obstacles, identified traffic control signals including speed limits, etc. In some embodiments, the vehicle control module utilizes the determined results to control the vehicle along the determined path. In some embodiments, the vehicle control module is the vehicle control module 109 in Figure 1.
[0051] In 409, the vehicle is controlled. In some embodiments, a vehicle with autonomous driving activated is controlled using a vehicle control module, such as the vehicle control module 109 in Figure 1. Vehicle control, for example, maintains the vehicle in the lane at an appropriate speed, taking the surrounding environment into consideration. Therefore, the vehicle's speed and / or steering can be adjusted. In some embodiments, the results are used to adjust the vehicle in anticipation of adjacent vehicles merging into the same lane. In various embodiments, using the results of deep learning analysis, the vehicle control module determines the appropriate way to operate the vehicle along a determined path, for example, at an appropriate speed. In various embodiments, the results of vehicle control, such as changes in speed, application of braking, and steering adjustments, are retained and used for the automatic generation of training data. In various embodiments, vehicle control parameters are retained for the automatic generation of training data and transmitted in 411.
[0052] In 411, sensor data and related data are transmitted. For example, sensor data received in 401 is sent to a computer server for the automatic generation of training data, along with the results of deep learning analysis in 405 and / or vehicle control parameters used in 409. In some embodiments, the data is time-series data, and various collected data are associated together by the computer server. For example, odometry data is associated with captured image data to generate ground truth. In various embodiments, the collected data is transmitted wirelessly from the vehicle to the training data center, for example, via WiFi or a cellular connection. In some embodiments, metadata is transmitted along with the sensor data. For example, metadata may include time, timestamp, location, vehicle type, and vehicle control and / or operating parameters such as speed, acceleration, braking, whether autonomous driving is enabled, steering angle, and odometry data. Additional metadata may include the time since the last previous sensor data was transmitted, vehicle type, weather conditions, road conditions, etc. In some embodiments, the transmitted data is anonymized, for example, by removing the vehicle's unique identifier. As another example, data from similar vehicle models is merged so that individual users and their vehicle usage cannot be identified.
[0053] In some embodiments, data is transmitted only in response to a trigger. For example, in some embodiments, an incorrect prediction triggers the transmission of sensor data and related data to automatically collect data in order to create a curated set of examples for improving the predictions of the deep learning network. For example, a prediction made in 405 in relation to whether a vehicle is about to merge is determined to be incorrect by comparing the prediction with the observed actual result. Data, including sensor data and related data associated with the incorrect prediction, is then transmitted and used to automatically generate training data. In some embodiments, triggers can be used to identify specific scenarios, such as sharp curves, road junctions, lane merging, sudden stops, or other suitable scenarios where additional training data would be useful and may be difficult to collect. For example, a trigger may be based on the sudden stop or disengagement of an autonomous driving function. As another example, vehicle operating characteristics such as changes in speed or acceleration may form the basis of a trigger. In some embodiments, a prediction with accuracy below a certain threshold triggers the transmission of sensor data and related data. For example, in certain scenarios, a prediction may not have a Boolean true or false result, but is instead evaluated by determining an accuracy value for the prediction.
[0054] In various embodiments, sensor data and associated data are captured over a period of time, and the entire time-series data is transmitted together. The period can be configured and / or based on one or more factors such as vehicle speed, distance traveled, and changes in speed. In some embodiments, the sampling rate of the captured sensor data and / or associated data is configurable. For example, the sampling rate is increased at high speeds, during sudden braking, during sudden acceleration, during sudden steering, or in other appropriate scenarios where additional fidelity is required.
[0055] Figure 5 shows an example of an image captured by a vehicle sensor. In the illustrated example, the image in Figure 5 includes image data 500 captured from a vehicle traveling in the lane between two lane lines. The positions of the vehicle and sensor used to capture the image data 500 are as follows: Represented by Bell A. Image data 500 is sensor data that can be captured by a camera sensor, such as a forward-facing camera of the vehicle, while driving. Image data 500 captures portions of lane lines 501 and 511. Lane lines 501 and 511 curve to the right as they approach the horizontal line. In the illustrated example, lane lines 501 and 511 are visible, but detection becomes increasingly difficult as they curve far away from the camera sensor's position. The white lines drawn on top of lane lines 501 and 511 approximate the detectable portions of lane lines 501 and 511 from image data 500 without additional input. In some embodiments, the detected portions of lane lines 501 and 511 can be detected by segmenting image data 500.
[0056] In some embodiments, labels A, B, and C correspond to different locations on the road and different times in a time series. Label A corresponds to the time and location of the vehicle when image data 500 was captured. Label B corresponds to a location on the road ahead of the location of label A, but at a later time than the time of label A. Similarly, label C corresponds to a location on the road ahead of the location of label B, but at a later time than the time of label B. As the vehicle travels, it passes through the locations of labels A, B, and C (from label A to label C), capturing a time series of sensor data and associated data during the journey. The time series includes elements captured at the locations (and times) of labels A, B, and C. Label A corresponds to the first element of the time series, label B corresponds to an intermediate element of the time series, and label C corresponds to an intermediate (or potentially last) element of the time series. At each label, additional data, such as odometry data of the vehicle at the label location, is captured. Depending on the length of the time series, additional or less data may be captured. In some embodiments, a timestamp is associated with each element in a time series.
[0057] In some embodiments, the ground truth (not shown) of lane lines 501 and 511 is determined. For example, using the process disclosed herein, the locations of lane lines 501 and 511 are identified by identifying different portions of lane lines 501 and 511 from different elements of a time series. In the illustrated example, portions 503 and 513 are identified using image data 500 and associated data (such as odometry data) acquired at location and time of label A. Portions 505 and 515 are identified using image data (not shown) and associated data (such as odometry data) acquired at location and time of label B. Portions 507 and 517 are identified using image data (not shown) and associated data (such as odometry data) acquired at location and time of label C. By analyzing the elements of the time series, the locations of different portions of lane lines 501 and 511 can be identified, and the ground truth can be determined by combining the different identified portions. In some embodiments, portions are identified as points along each portion of the lane line. In the illustrated example, only three parts of each lane line are highlighted to illustrate the process (parts 503, 505, and 507 of lane line 501 and parts 513, 515, and 517 of lane line 511), but additional parts may be captured over time to determine the position of the lane lines with higher resolution and / or higher precision.
[0058] In various embodiments, the position of the portion of the image data capturing lane lines 501 and 511 closest to the sensor location is determined with high precision. For example, the positions of portions 503 and 513 are identified with high precision using image data 500 labeled A and associated data (such as odometry data). The positions of portions 505 and 515 are identified with high precision using image data and associated data labeled B. The positions of portions 507 and 517 are identified with high precision using image data and associated data labeled C. By utilizing time-series elements, various aspects of the captured lane lines 501 and 511 can be determined by time series. The location of a specific part can be identified in three dimensions with high accuracy and used as the basis for the ground truth of the lane lines 501 and 511. In various embodiments, the determined ground truth is associated with selected elements of a time series, such as image data 500. The ground truth and selected elements can be used to create training data for predicting the lane lines. In some embodiments, the training data is created automatically without human labeling. The training data can be used to train a machine learning model to predict the three-dimensional trajectory of the lane lines from captured image data, such as image data 500.
[0059] Figure 6 shows an example of an image captured from a vehicle sensor having a predicted three-dimensional trajectory of lane lines. In the illustrated example, the image in Figure 6 includes image data 600 captured from a vehicle traveling in the lane between two lane lines. The positions of the vehicle and sensor used to capture the image data 600 are represented by label A. In some embodiments, label A corresponds to the same position as label A in Figure 5. The image data 600 is sensor data and may be captured from a camera sensor, such as a forward-facing camera of the vehicle, while driving. The image data 600 captures portions of lane lines 601 and 611. Lane lines 601 and 611 curve to the right as they approach the horizontal line. In the illustrated example, lane lines 601 and 611 are visible, but they become increasingly difficult to detect as they curve far away from the camera sensor's position. The red lines drawn over lane lines 601 and 611 are the predicted three-dimensional trajectories of lane lines 601 and 611. Using the process disclosed herein, a three-dimensional trajectory is predicted using image data 600 as input to a trained machine learning model. In some embodiments, the predicted three-dimensional trajectory is represented as a three-dimensional parameterized spline or another parameterized representation format.
[0060] In the illustrated example, portion 621 of lane lines 601 and 611 is a portion of lane lines 601 and 611 that is at a distance from each other. The three-dimensional location (i.e., longitude, latitude, and altitude) of portion 621 of lane lines 601 and 611 is determined with high accuracy using the process disclosed herein and is included in the predicted three-dimensional trajectory of lane lines 601 and 611. Using a trained machine learning model, the three-dimensional trajectory of lane lines 601 and 611 can be predicted using image data 600 and without requiring location data at the location of portion 621 of lane lines 601 and 611. In the illustrated example, image data 600 is captured at location and time of label A.
[0061] In some embodiments, label A in Figure 6 corresponds to label A in Figure 5, and the predicted 3D trajectories of lane lines 601 and 611 are determined using only image data 600 as input to a trained machine learning model. By training the machine learning model with ground truth determined using time-series image data and related data, including elements acquired at the locations of labels A, B, and C in Figure 5, the 3D trajectories of lane lines 601 and 611 can be predicted with high accuracy, even for distant parts of lane lines such as part 621. Although image data 600 and image data 500 in Figure 5 are related, trajectory prediction does not require image data 600 to be included in the training data. By training with sufficient training data, lane lines can be predicted even in newly encountered scenarios. In various embodiments, the predicted 3D trajectories of lane lines 601 and 611 are used to maintain the vehicle's position within the detected lane line and / or to autonomously navigate the vehicle along the detected lane of the predicted lane line. Predicting lane lines in 3D significantly improves navigation performance, safety, and accuracy.
[0062] The embodiments described above have been explained in some detail to clarify understanding, but the present invention The disclosed embodiments are not limited to the details provided. There are many alternative ways to carry out the invention. The disclosed embodiments are illustrative and not limiting.
Claims
1. A step of acquiring sensor data generated at multiple points in time within a first period by one or more processors, wherein the sensor data is generated by sensors on the vehicle and captures the environment of the vehicle. A step of determining ground truth based on sensor data using one or more processors, wherein the ground truth is associated with a three-dimensional representation of the characteristics of the environment using sensor data over a time series of at least two time points within a subset of the plurality of time points. A step of training a machine learning model using a training dataset with one or more processors, wherein the training dataset includes portions of the sensor data captured during a subset of the plurality of time points within a first period, associated with the three-dimensional representation of the features. A method wherein the machine learning model is trained to output the identification of the three-dimensional representation of the feature based on the input of sensor data associated with a point in time within a second period.
2. The method according to claim 1, wherein the three-dimensional representation of the aforementioned features is associated with lane lines.
3. The sensor data includes a plurality of images captured at a plurality of time points within the first period. The method according to claim 2, wherein the lane line portion is depicted by a set of images selected from the plurality of images.
4. The method according to claim 3, further comprising the step of selecting an image depicting the lane lines from the set of images based on means associated with the image, wherein the means is associated with a confidence value indicating the presence of the lane lines in each image of the set of images.
5. The method according to claim 1, wherein the three-dimensional representation of the aforementioned features reflects the trajectory of the lane line.
6. The method according to claim 1, wherein the three-dimensional representation of the features reflects a route associated with a vehicle.
7. The sensor data includes a plurality of images captured at a plurality of time points within the first period. The method according to claim 6, wherein the machine learning model is trained to output the identification of the route based on a second individual image of a vehicle separate from the plurality of images.
8. The method according to claim 6, wherein the vehicle is in a first lane adjacent to a second lane at some point in time, and at that point in time, the sensor data is captured by a sensor of a different vehicle located in the second lane.
9. The method according to claim 1, wherein the training dataset further includes scene data describing the real-world environment around the sensors of the vehicle that captured the sensor data.
10. It is a system, Sensor data generated at multiple points in time within the first period is acquired, and the sensor data is generated by sensors on the vehicle and captures the environment of the vehicle. The ground truth is determined based on the sensor data, and the ground truth is associated with a three-dimensional representation of the characteristics of the environment using the sensor data over a time series of at least two time points within a subset of the multiple time points. A machine learning model is trained using a training dataset, the training dataset including portions of the sensor data captured during a subset of multiple time points within a first period, associated with the three-dimensional representation of the features. Equipped with one or more processors for, The machine learning model is a system trained to output the identification of the three-dimensional representation of the feature based on the input of sensor data associated with a point in time within a second period.
11. The system according to claim 10, wherein the three-dimensional representation of the aforementioned features is associated with lane lines.
12. The sensor data includes a plurality of images captured at a plurality of time points within the first period. The system according to claim 11, wherein the lane line portion is depicted by a set of images selected from the plurality of images.
13. The one or more processors further include: The system according to claim 12, wherein, in order to train the machine learning model, it is configured to select an image depicting the lane lines from the set of images based on means associated with the image, the means being associated with a confidence value indicating the presence of the lane lines in each image of the set of images.
14. The system according to claim 10, wherein the three-dimensional representation of the aforementioned features reflects the trajectory of the lane line.
15. The system according to claim 10, wherein the three-dimensional representation of the aforementioned features reflects a route associated with a vehicle.
16. The sensor data includes a plurality of images captured at a plurality of time points within the first period. The system according to claim 15, wherein the machine learning model is trained to output the identification of the route based on a second individual image of a vehicle separate from the plurality of images.
17. The system according to claim 15, wherein the vehicle is in a first lane adjacent to a second lane at some point in time, and at that point in time, the sensor data is captured by a sensor of a different vehicle located in the second lane.
18. The system according to claim 10, wherein the training dataset further includes scene data describing the real-world environment around the sensors of the vehicle that captured the sensor data.
19. A non-temporary computer-readable storage medium containing computer instructions, wherein when the computer instructions are executed by one or more processors, the system provides to the one or more processors: Sensor data generated at multiple points in time within a first period is acquired, the sensor data generated at multiple points in time is represented by the sensor data captured during a subset of the multiple points in time, the sensor data is generated by sensors on the vehicle and captures the environment of the vehicle, The ground truth is determined based on the sensor data, and the ground truth is associated with a three-dimensional representation of the characteristics of the environment using the sensor data over a time series of at least two time points within a subset of the multiple time points. A machine learning model is trained using a training dataset, the training dataset comprising portions of the sensor data captured during a subset of the multiple time points within a first period, associated with the three-dimensional representation of the features, A non-temporary computer-readable storage medium, the machine learning model being trained to output the identification of the three-dimensional representation of the feature based on the input of sensor data associated with a point in time within a second period.
20. The non-temporary computer-readable storage medium according to claim 19, wherein the three-dimensional representation of the aforementioned features is associated with lane lines.
Citation Information
Patent Citations
Information processor, vehicle, information processing method and program
JP2017215940A
Mobile body detection method, mobile body learning method, mobile body detector, mobile body learning device, mobile body detection system, and program
JP2019008519A
Method and system for controlling a vehicle
JP2019519851A
Automated image labeling for vehicles based on maps
US20180283892A1
Method and system for controlling vehicle
WO2018084324A1