3D feature prediction for autonomous driving
By using vehicle-mounted sensors to capture time-series data and employing piecewise polynomial representations, the training of accurate machine learning models for autonomous driving is improved, addressing the inefficiencies in existing data collection and labeling methods, resulting in enhanced lane detection and path prediction for autonomous vehicles.
Patent Information
- Application Number
- JP2023182986
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-01
- Filing Date
- 2023-10-25
- Publication Date
- 2025-11-10
- Estimated Expiration
- 2040-01-28
AI Technical Summary
Deep learning systems for autonomous driving face challenges in training accurate machine learning models due to the significant effort and resources required for data collection, curation, and manual labeling, which limits the quality of training sets and hinders efficient model training.
A machine learning training technique that utilizes vehicle-mounted sensors to capture vehicle operating parameters and generate training data sets, incorporating time-series image data to create highly accurate models for predicting three-dimensional features, such as lane lines and vehicle paths, by determining ground truth from a group of time series elements and using piecewise polynomial representations.
This approach enhances the accuracy and efficiency of machine learning models for autonomous driving by reducing computational resources and improving processing speed while maintaining high accuracy in lane detection and path prediction, thereby enhancing vehicle safety and operational efficiency.
Smart Images

Figure 0007766661000001 
Figure 0007766661000002 
Figure 0007766661000003
Abstract
Description
[Background technology]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a part of the "PREDICTING THREE- DIMENSIONAL FEATURES FOR AUTONOMOUS DRIV U.S. Patent Application No. 16 / 2657 entitled "3D Feature Prediction for Autonomous Driving" This application is a continuation of and claims priority to application No. 20, the entire disclosure of which is incorporated herein by reference. will be incorporated into Summary of the Invention [Problem to be solved by the invention]
[0002] Deep learning systems used in applications such as autonomous driving train machine learning models. Typically, the performance of a deep learning system is measured by the number of iterations used to train the model. Limited at least in part by the quality of the training set used. Significant resources are invested in data collection, curation, and annotation. Traditionally, much of the effort to curate training datasets has focused on the potential training data. This is done manually by examining the data and appropriately labeling the features relevant to the data. The effort required to create a training set with accurate labels can be significant and often Furthermore, it is tedious to collect and accurately label data that needs improvement for machine learning models. Therefore, it is often difficult to train a training model with accurately labeled features. The process for generating training data needs to be improved. [Brief explanation of the drawings]
[0003] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings.
[0004] [Figure 1] FIG. 1 is a block diagram illustrating one embodiment of a deep learning system for autonomous driving.
[0005] [Figure 2] 1 is a flow diagram illustrating an embodiment of a process for training and applying machine learning models for autonomous driving.
[0006] [Figure 3] 1 is a flow diagram illustrating an embodiment of a process for creating training data using elements of a time series.
[0007] [Figure 4] 1 is a flow diagram illustrating an embodiment of a process for training and applying machine learning models for autonomous driving.
[0008] [Figure 5] FIG. 2 is a diagram illustrating an example of an image captured from a vehicle sensor.
[0009] [Figure 6] FIG. 1 illustrates an example of an image captured from a vehicle sensor with a predicted three-dimensional trajectory of lane lines. DETAILED DESCRIPTION OF THE INVENTION
[0010] The present invention relates to processes, apparatus, systems, compositions of matter, and computer-readable storage media. The embodied computer program product and / or processor, e.g., instructions stored in and / or provided by a memory coupled to the The present invention may be implemented in various ways, including by a processor configured to execute the method. These implementations, or any other form that the invention may take, may be referred to as technology. In general, the order of steps in the disclosed processes may be varied within the scope of the present invention. Unless otherwise specified, the software is described as being configured to perform a task. Components such as processors or memory are designed to perform tasks at a given time. A general component that is temporarily configured, or a specific component that is manufactured to perform a task. As used herein, the term "processor" is one or more programs configured to process data such as computer program instructions. Refers to a number of devices, circuits, and / or processing cores.
[0011] A detailed description of one or more embodiments of the invention is provided below, along with accompanying drawings that illustrate the principles of the invention. Both are provided below. While the invention will be described in connection with such embodiments, the invention The scope of the present invention is limited only by the claims. However, the present invention encompasses numerous alternatives, modifications and equivalents. For purposes of providing an understanding of the present invention, numerous specific details are set forth in the following description. and the present invention may not be patentable without some or all of these specific details. It is understood that the present invention may be practiced in accordance with the scope of the claims. Technical materials known in the art will not be described in detail in order to avoid unnecessarily obscuring the present invention. It has not been done.
[0012] A machine learning training technique is disclosed for generating highly accurate machine learning results. and uses data captured by sensors on the vehicle to capture vehicle operating parameters. For example, a vehicle-mounted sensor is used to generate a training data set. Captures data such as image data of the road and surrounding environment on which the vehicle is traveling. Includes vehicle lane lines, vehicle lanes, other vehicle traffic, obstacles, traffic control signs, etc. Odometry and other similar sensors can capture Vehicle movements, such as vehicle speed, steering, heading, change in direction, change in position, change in altitude, and change in speed The operational parameters are captured. The captured data set is used to create a training data set. The training dataset is sent to the training server for generating highly accurate machine learning results. In some embodiments, the time series of captured data is used to train a machine learning model. The data is used to generate training data. For example, ground truth The time series element (or elements) is determined based on a group of time series elements, and a single As an example, a series of images over a period of time, such as 30 seconds, can be used to identify a vehicle. Determine the actual path of the vehicle lane lines over the period of time the vehicle is traveling. is determined by using the most accurate image of the vehicle lane over that period. Different parts (or positions) of the lane lines are identified from different image data over time. As the vehicle travels along the lane, the lane lines More accurate data is captured for different parts of the lane. In some cases, lane occlusions The minute appears when the vehicle travels, for example, around a hidden curve or over the crest of a hill. The most accurate portion of the lane line from each image in the time series is calculated for the entire group of image data. Image data of distant lane lines can be used to identify lane lines over a wide area. This data is usually less detailed than the image data of lane lines near the vehicle. By capturing time-series image data as the vehicle travels along the lane, the corresponding lane line is Accurate image data and corresponding odometry data of all parts of the vehicle are collected.
[0013] In some embodiments, a three-dimensional representation of features such as lane lines is generated based on ground truth. The ground truth is then generated from a group of time series elements corresponding to the A subsection is a time series element, such as a single image frame, of a group of captured image data. For example, the first image in a group of images is represented in three-dimensional space. It is related to the ground truth of the lane lines. The first frame and the ground truth are selected based on the group. The images are used to create training data. For example, a car can be trained using only a single image. Training data is created to predict the three-dimensional representation of both lanes. indicates whether any element or group of elements in a group of time series elements is in the ground truth. are associated with each other and used to create training data, e.g., ground truth can be applied to the entire video sequence to create the training data. The middle or last element of a group of sequence elements is associated with the ground truth. and used to create training data.
[0014] In various embodiments, the selected image and ground truth may include lane lines, It is applied to various features such as vehicle path prediction including adjacent vehicles, object depth distance, and traffic signs. For example, a series of images of a vehicle in an adjacent lane can be used to predict the path of that vehicle. It uses a time series of images taken by adjacent vehicles and the actual route. Then, the captured group and a single image of the actual path are used to predict the path of the vehicle. This information can be used to train the autonomous vehicle. For example, route prediction can predict whether an adjacent vehicle will cut into the route of another vehicle. Autonomous vehicles can predict whether they will merge in front of a non-autonomous vehicle. For example, autonomous vehicles can be controlled to minimize collisions. slow down and adjust vehicle speed and / or steering to prevent a collision with adjacent vehicles and / or or initiate warnings to occupants of autonomous vehicles and / or change lanes, etc. In various embodiments, the ability to accurately infer path predictions, including vehicle path predictions, is enhanced by This will significantly improve the safety of motor vehicles.
[0015] In some embodiments, a trained machine learning model may be used to identify lane lines. It is used to predict a 3D representation of one or more features for autonomous driving, including For example, a two-dimensional image can be extracted from the image data by segmenting the image of the lane lines. Instead of identifying timelines, the time series elements and the corresponding odometry data are The 3D representation is generated using the lane line detection and the corresponding Including elevation changes that significantly improve the accuracy of lane and identified drivable path detection. In some embodiments, the lane lines are formed by one or more splines or other parabolic elements. It is represented using a metered representation, where a piecewise polynomial is used to represent the lane lines. The use of the formula significantly reduces the computational resources required to evaluate a three-dimensional object. This reduction in computational resources results in a significant improvement in processing speed and accuracy without significantly sacrificing the accuracy of the representation. In various embodiments, lane lines, particularly those including curves in the lane lines, are used to improve efficiency. The function can be expressed using a piecewise polynomial, a 3D point set, or another suitable representation. For example, a piecewise polynomial is a function of a group of elements captured over time using sensor data. Interpolating the actual lane lines using high-precision sections of the identified lane lines from do.
[0016] In some embodiments, sensor data is received. The sensor data may include images (video and and / or still images), radar, audio, LiDAR, inertial, odometry, position The sensor data may include time-series, location-based, and / or other forms of sensor data. A group of time series elements may contain a group of column elements. For example, a group of time series elements may contain a group of vehicle In some embodiments, the image may include a group of images captured from a camera sensor. The training dataset is a set of time series elements that are at least selected in the group of time series elements. For each time series element, the corresponding ground truth is based on multiple time series elements within the group of time series elements. For example, the ground truth is determined by determining the group The most relevant element of a group of time series elements, including the previous and / or next time series elements in the group. In some scenarios, the preceding and / or following is a car that disappears first around the curve and appears only in later elements of the time series. Contains data not present in the previous time series element, such as both lane lines. The predicted truth may be a 3D representation of the vehicle lane lines, the predicted path of the vehicle, or another similar prediction. An element of a group of time series elements is selected and associated with the ground truth. The selected elements and the ground truth are part of the training dataset. In some embodiments, the processor uses the training dataset to generate a machine learning model. For example, the training dataset is used to train a vehicle for autonomous driving or driving. It is used to train machine learning models to infer features used in assistive actions. Using a trained machine learning model, the neural network detects vehicle lanes, driving Vehicles, objects (e.g., pedestrians, stationary vehicles, moving vehicles), weather (e.g., rain, hail) , fog, etc.), traffic control objects (e.g., traffic lights, traffic signs, road signs, etc.), traffic patterns It is possible to infer features related to autonomous driving, such as:
[0017] In some embodiments, the system includes a processor and a memory coupled to the processor. The processor generates image data based on an image captured by a camera of the vehicle. For example, a camera sensor mounted on a vehicle may receive the vehicle's environment. The camera may be a forward-facing camera, a pillar camera, or another suitably positioned The image data captured from the camera can be processed by a GPU or APU on the vehicle. In some embodiments, the image data is processed using a processor such as an image processor. The data is fed to a pre-trained machine learning model that is trained to predict the 3D trajectory of the vehicle lane. Used as the basis for input. For example, image data is trained to predict vehicle lanes. The machine learning model uses the detected The 3D trajectory of the lane is inferred from the image. Instead of segmenting into 3D representations, a 3D representation is inferred. The original representation can be a spline, a parametric curve, or a curve described in 3 dimensions. In some embodiments, the three-dimensional trajectory of the vehicle lane is used to automatically locate the vehicle. For example, the 3D trajectory is provided to the driver when controlling the lane line and the corresponding driving position. It is used to determine the performance space.
[0018] FIG. 1 is a block diagram illustrating an embodiment of a deep learning system for autonomous driving. The multi-layer learning system is used for automated driving and / or driver assistance operation of vehicles and for autonomous driving. To collect and process data to train machine learning models for dynamic driving, In various embodiments, the deep learning system includes various components that can be used together. The system is installed in the vehicle. Data from the vehicle is used to control the autonomous driving functions of the vehicle or other similar vehicles. can be used to train and improve
[0019] In the illustrated example, the deep learning system 100 includes a sensor 101, an image preprocessor 103, and a , Deep Learning Network 105, Artificial Intelligence (AI) Processor 107, Vehicle Control Module 109, and a deep learning network including a network interface 111. In various embodiments, different components are communicatively connected. For example, sensors The sensor data from 101 is supplied to an image preprocessor 103. The sensor data processed by the sensor 103 is fed to a deep learning network running on the AI processor 107. The deep learning network 105 operates on the AI processor 107. The output of clock 105 is provided to a vehicle control module 109. In various embodiments, the vehicle The control module 109 controls the operation of the vehicle, such as the vehicle's speed, braking, and / or steering. In various embodiments, sensor data and / or mechanical The learning results are transmitted to a remote server via the network interface 111. For example, training to improve vehicle performance, comfort, and / or safety. To collect data, the sensor data is transmitted via a network interface 111. In various embodiments, the network interface may be used to transmit the data to a remote server. The interface 111 communicates with remote servers, places phone calls, among other reasons, Send and / or receive text messages and sensor data based on vehicle operation In some embodiments, the deep learning system 100 may Depending on the implementation, additional or fewer components may be included. For example, some implementations In some embodiments, the image preprocessor 103 is an optional component. In some embodiments, the deep learning network is run before the output is provided to the vehicle control module 109. A post-processing component (not shown) is used to perform post-processing on the output of network 105. It is used.
[0020] In some embodiments, the sensor 101 includes one or more sensors. In an embodiment, the sensors 101 are mounted on the vehicle at different locations on the vehicle, and / or The sensor 101 may be oriented in one or more different directions. For example, the sensor 101 may be oriented in the front of the vehicle. , side, rear, and / or roof, etc., in forward, rear, sideways, etc. orientations. In some embodiments, the sensor 101 may be a high dynamic range camera. In some embodiments, the sensor 101 is a non-visual sensor. In some embodiments, the sensor 101 may include, among other things, radar, audio, sensors, including audio, LiDAR, inertial, odometry, position, and / or ultrasonic sensors. In some embodiments, the sensor 101 is mounted on a vehicle having a vehicle control module 109. For example, the sensor 101 may be mounted on an adjacent vehicle and / or or attached to the road or environment, and uses deep learning to capture sensor data. In some embodiments, the sensor 101 detects when the vehicle is traveling. One or more cameras that capture the road surface being viewed. For example, one or more forward-facing The headlight and / or pillar camera captures the lane markings of the lane the vehicle is traveling in. As another example, the camera may detect a vehicle attempting to cut into the lane in which the vehicle is traveling. Additional sensors are used to measure odometry, position, and / or vehicle trajectory. The sensor 101 captures still images and / or vehicle control information including information about the vehicle's track. The camera can include both image sensors capable of capturing video. The data may be captured over a period of time, such as a sequence of data captured over a period of time. For example, images of lane markings may be displayed on the vehicle for a period of 15 seconds or another suitable period. In some embodiments, the sensor 101 may be captured along with vehicle Global Positioning System (GPS) sensors to determine the location and / or change in location of both vehicles. This includes position sensors such as
[0021] In some embodiments, the image preprocessor 103 processes the sensor data of the sensor 101. For example, the image preprocessor 103 can be used to preprocess the sensor Preprocessing the data and splitting the sensor data into one or more components; and / or One or more components can be post-processed. In some embodiments, the image processor The reprocessor 103 includes a graphics processing unit (GPU), a central processing unit (CPU), An image signal processor, or a dedicated image processor. The preprocessor 103 is a tone mapper processor that processes high dynamic range data. In some embodiments, the image preprocessor 103 is an artificial intelligence (AI) processor. For example, the image preprocessor 103 is implemented as part of the AI processor 107. In some embodiments, the image preprocessor 107 may be a component of the image preprocessor 107. The image 103 can be used to normalize or transform the image. For example, ,Images captured by a fisheye lens may be distorted, and the image preprocessor 103 can be used to transform the image to remove or correct distortions. In image processing, noise, distortion, and / or blurring are removed or reduced during the preprocessing step. In various embodiments, the images are adjusted or refined to improve the results of the machine learning analysis. For example, the white balance of an image is normalized for daylight, sunny, cloudy, dusk, among other conditions. and adjusted to account for different lighting operating conditions such as sunrise, sunset, and nighttime conditions. do.
[0022] In some embodiments, the deep learning network 105 detects lane markers, lanes, and driving Scan the driving environment to determine maneuverability, obstacles, and / or potential vehicle paths Deep learning networks used to determine vehicle control parameters, including analyzing For example, the deep learning network 105 is trained on inputs such as sensor data. and the output of which is provided to the vehicle control module 109. The output may be an artificial neural network such as a neural network (CNN). In another example, the output may include at least a three-dimensional representation of the line marker. This includes potential vehicles that may merge into the lane of the vehicle. In an embodiment, the deep learning network 105 receives at least sensor data as input. Additional inputs may include vehicle specifications such as the vehicle's surrounding environment and / or vehicle operating characteristics. The scene data may include scene data describing rain, wet roads, snow, etc. , mud, high-density traffic, highways, cities, school zones, and other environments around the vehicle. In some embodiments, the deep learning network may include scene tags that describe the scene. The output of the CL 105 is a three-dimensional trajectory of the vehicle's vehicle lane. The output of the deep learning network 105 is a potential vehicle interruption. The network 105 identifies adjacent vehicles that are likely to enter the lane ahead of the vehicle.
[0023] In some embodiments, the artificial intelligence (AI) processor 107 may implement a deep learning network. In some embodiments, the A The I processor 107 performs a convolutional neural network (CNN) on the sensor data. The AI processor 107 is a dedicated AI processor for performing inference using It can be optimized for the bit depth of the sensor data. The AI processor 107 performs, among other things, convolution, dot product, vector, and / or matrix operations. It is optimized for deep learning operations such as neural network operations, including In an embodiment, the AI processor 107 is implemented using a graphics processing unit (GPU). In various embodiments, the AI processor 107, when executed, The sensor performs deep learning analysis on the received input sensor data to be used for autonomous driving. a memory configured to provide instructions to an AI processor that causes the AI processor to determine a machine learning outcome based on the In some embodiments, the AI processor 107 combines the data as training data. It is used to process sensor data in preparation for making it available for use.
[0024] In some embodiments, the vehicle control module 109 includes an artificial intelligence (AI) processor 107 and converts the output into vehicle control actions. In an embodiment, the vehicle control module 109 is used to control the vehicle for autonomous driving. In various embodiments, the vehicle control module 109 may monitor the vehicle's speed, acceleration, Steering, braking, etc. may be adjusted. For example, in some embodiments, the vehicle control module The module 109 controls the vehicle to maintain its position within the lane and move the vehicle into another lane. Used to merge into other vehicles and adjust vehicle speed and lane placement to account for merging vehicles, etc. will be done.
[0025] In some embodiments, the vehicle control module 109 controls brake lights, turn signals, In some embodiments, the vehicle may be equipped with a power supply that is connected to the power supply and is used to control vehicle lighting, such as headlights. The control module 109 controls the vehicle's sound system, audio alert playback, microphone Used to control vehicle audio states such as enable horn, enable headset, etc. In some embodiments, the vehicle control module 109 may detect a potential collision or an intent alarms to notify the driver and / or passengers of driving events such as approaching a scheduled destination In some embodiments, the system is used to control a notification system, including a notification system. Vehicle control module 109 is used to regulate sensors such as sensor 101 in the vehicle. For example, the vehicle control module 109 may be used to control orientation changes, output resolution and / or Or change the format type, increase or decrease the acquisition rate, or increase or decrease the acquired dynamic range. one or more of the following: adjustment, adjusting camera focus, enabling and / or disabling sensors, etc. In some embodiments, the vehicle control module may change the parameters of the sensors. Use module 109 to change the frequency range of the filter, feature and / or edge detection Image preprocessor 103, such as adjusting parameters, adjusting channels and bit depth In various embodiments, the vehicle control module 109 , used to implement automated driving and / or driver assistance control of vehicles. In this embodiment, the vehicle control module 109 uses a processor coupled to a memory. In some embodiments, the vehicle control module 109 is implemented as an application specific aggregate. Application Specific Integrated Circuit (ASIC), Programmable Logic Device (PLD), or other suitable processing hardware It is implemented using hardware.
[0026] In some embodiments, the network interface 111 includes A communication interface for transmitting and / or receiving data. In this state, the network interface 111 interfaces with a remote server. connect and make voice calls and send and / or receive text messages It transmits sensor data and runs a deep learning network containing updated machine learning models. receive updates to the network and check environmental conditions, including weather conditions and forecasts, traffic conditions, etc. Includes a cellular or wireless interface for searching. For example, a network interface. The face 111 is composed of a sensor 101, an image preprocessor 103, and a deep learning network 1 05, the AI processor 107, and / or the vehicle control module 109 instructions and / or or to receive updates on operational parameters. The machine learning model in Work 105 is uploaded using the network interface 111. As another example, the network interface 111 can be used to , and image preparation, such as the firmware and / or image processing parameters of the sensor 101. As yet another example, the operating parameters of the processor 103 can be updated. ,Network interface 111 is used to train machine learning models. Potential training data can be sent to a remote server.
[0027] Figure 2 shows an example of the process for training and applying machine learning models for autonomous driving. 1 is a flow chart illustrating an embodiment in which input data, including, for example, sensor and odometry data, The data is received and processed to create training data for training a machine learning model. In some embodiments, the sensor data may be image data captured via an automated driving system. In some embodiments, the sensor data corresponds to a user manually disabling autonomous driving. Some respond to captured sensor data based on specific use cases, such as In this embodiment, the process generates a machine learning model for the deep learning system 100 of FIG. Used to create and deploy
[0028] At 201, training data is prepared. In some embodiments, image data and The sensor data, including the vehicle speed and odometry data, is received to create a training data set. The data can include still images and / or video from one or more cameras. Additional sensors such as radar, LiDAR, and ultrasonic sensors can be used to provide relevant sensor data. In various embodiments, the sensor data can be The data is paired with the corresponding odometry data to help identify the characteristics of the data. For example, position and change-of-position data can be used to identify lane lines, traffic control signals, objects, etc. In some embodiments, the location of relevant features within the sensor data can be identified. The sensor data is a time series of elements that are used to determine the ground truth. The ground truth for the group is then calculated from the first frame of image data, e.g. Associated with a subset of the time series. Selected elements of the time series and ground truth The source is used to prepare training data. In some embodiments, the training data It only identifies features from sensor data such as lane lines, vehicle paths, and traffic patterns. The prepared training data is used to train a machine learning model. In various embodiments, the sensor may include data for verification and testing. The data may be in different formats. For example, the sensor data may be still images, motion images, The odometry data may be an applied acceleration, an applied braking applied, steering applied, vehicle position, vehicle heading, change in vehicle position, change in vehicle heading In various embodiments, the training data may include vehicle operating parameters such as the are curated and annotated to create a dataset. In embodiments, some of the training data preparation may be performed by human curators. In various embodiments, a portion of the training data is generated automatically from data captured from a vehicle. significantly reducing the effort and time required to build a robust training dataset. In some embodiments, the format of the data is In various embodiments, the training data is compatible with the machine learning models used in the application. The data includes validation data for testing the accuracy of the trained model.
[0029] In 203, a machine learning model is trained. For example, the data prepared in 201 is In some embodiments, the machine learning model is trained using convolutional neural networks. In various embodiments, the neural network is a neural network such as a neural network (CNN). In some embodiments, the neural network comprises a plurality of hidden layers. , which may include multiple layers, including multiple convolutional layers and pooling layers. In an embodiment, the training model uses a validation data set created from the received sensor data. In some embodiments, the machine learning model is validated using a single input image. For example, the 3D representation of lane lines is trained to predict the 3D representation of the car. As another example, the image quality can be inferred from images captured from a camera. From the captured images, the predicted paths of adjacent vehicles are predicted, including whether the vehicle is about to merge. do.
[0030] At 205, the trained machine learning model is deployed. For example, the trained machine learning model The learning model is an update of a deep learning network, such as deep learning network 105 in Figure 1. In some embodiments, the software is installed in the vehicle using over-the-air updates. In some embodiments, the newly trained machine learning model is installed using Updates are performed using a wireless network such as WiFi or a cellular network. In some embodiments, a firmware update is sent to a new machine. The learning model may be installed when the vehicle is serviced.
[0031] At 207, sensor data is received. For example, the sensor data may be from one or more of the vehicle's In some embodiments, the sensor is sensor 10 of FIG. The sensors are a fisheye camera mounted behind the windshield, a pillar-mounted It can include imaging sensors such as front-facing or side-facing cameras or rear-facing cameras attached to the In various embodiments, the sensor data may be input to the machine learning model trained in 203. For example, sensor data is in a format that can be used as a The data may be raw or processed image data. In some embodiments, the data may be: From ultrasonic sensors, radar, LiDAR sensors, microphones, or other suitable technologies In some embodiments, the sensor data is captured during a pre-processing step. The image is preprocessed using an image preprocessor such as image preprocessor 103 of FIG. For example, the image may be normalized to remove distortion, noise, etc.
[0032] In 209, the trained machine learning model is applied. For example, The machine learning model is applied to the received sensor data at 207. In this case, the application of the model is performed using a deep learning network, such as deep learning network 105 in Figure 1. The algorithm is executed by an AI processor, such as AI processor 107 in FIG. 1, using the algorithm described above. In various embodiments, a trained machine learning model is applied to determine lane line 3D representations of features such as the vehicle moving are identified and / or predicted. Two splines are inferred to represent the lane lines of the lane. The predicted paths of adjacent vehicles are inferred, including whether they are likely to cut into the current lane. In various embodiments, a machine learning model is applied to identify vehicles, obstacles, and radar. Traffic patterns, traffic control signals, map features, object distances, speed limits, drivable spaces, etc. are identified. In some embodiments, the features are identified in three dimensions.
[0033] At 211, an autonomous vehicle is controlled. For example, one or more self-driving functions are , by controlling various aspects of the vehicle. Examples include vehicle steering, speed, and acceleration. control speed and / or braking, maintain vehicle position within lane, and avoid other vehicles maintain the vehicle's position relative to obstacles and / or obstacles, provide notification or warnings to occupants, Based on the analysis performed at 209, the vehicle steering and The speed and direction of the vehicle are controlled to keep it between two lane lines. Lane lines are predicted and corresponding vehicle lanes and drivable spaces are identified. In an embodiment, a vehicle control module, such as vehicle control module 109 in FIG. do.
[0034] FIG. 3 is a flow diagram illustrating one embodiment of a process for creating training data using elements of a time series. For example, the time series elements consisting of sensor and odometry data are and used to automatically create training data. Therefore, the process in Figure 3 automatically labels the training data with the corresponding ground truth. The results corresponding to a time series are associated with the elements of the time series. The results and selected elements are packaged as training data for predicting future outcomes. In various embodiments, the sensors and associated data are analyzed using the deep learning system of FIG. For example, in various embodiments, sensor data is captured using Sensor 1 of FIG. 01. In some embodiments, the process of FIG. 3 is performed at 201 of FIG. In some embodiments, the process of FIG. 3 is performed when an existing prediction is incorrect or This is done to automatically collect data when improvements can be made, e.g., when a vehicle is autonomous. Predictions are made by the autonomous vehicle to determine whether it is cutting into the vehicle's path. After waiting for a period of time and analyzing the captured sensor data, the prediction is determined to be correct or incorrect. In some embodiments, a determination can be made that the prediction may be improved. If the prediction is incorrect or can be improved, the process shown in Figure 3 is repeated. Curation of examples to improve machine learning models by applying the process to data relevant to the prediction. You can create customized sets.
[0035] At 301, a time series of elements is received. In various embodiments, the elements are captured by a vehicle. The sensor data is image data captured and sent to the training server. , are captured over a period of time to create a time series of elements. The elements are timestamps to maintain the order of the elements. As we move forward, later events in the time series are the result of events from earlier elements in the time series. Used to aid in prediction. For example, time series can signal confluence, accelerate, and It can capture vehicles in adjacent lanes that position themselves close to the nearby lane line. Using the entire time series, the results can be used to determine when a vehicle has merged into a shared lane. The result can be a selected image of the time series, such as one of the initial images in the time series. It can be used to predict when a vehicle will merge based on the factors it has identified. The time series captures the curve of the lane line. The various dips, bends, peaks, etc. in the lane that are not obvious to the user are captured. In this case, the elements are sensor data in a format that the machine learning model uses as input. For example, the sensor data may be raw or processed image data. In terms of form, the data may be acquired from ultrasonic sensors, radar, LiDAR sensors, or other suitable technologies. The data was captured from
[0036] In various embodiments, the time series may be organized by associating a timestamp with each element of the time series. For example, the timestamps are organized by the first element in the time series. The timestamp is used to identify time series elements in related data such as odometry data. In various embodiments, the length of the time series may be 10 seconds, 30 seconds, or It may be a fixed length of time, such as seconds, or another suitable length. The length of time is configurable. In various embodiments, the time series may be based on the speed of the vehicle, such as the average speed of the vehicle. For example, at slower speeds, the time length of the time series can be increased to The data is then transferred over a longer distance than would be possible if a shorter time period were used for the In some embodiments, the number of elements in the time series is configurable. For example, the number of elements can be based on the distance traveled, e.g., for a certain period of time: Faster moving vehicles contain more elements in the time series than slower moving vehicles The additional elements increase the fidelity of the captured environment and improve the accuracy of predicted machine learning results. In various embodiments, the number of elements can be determined by the number of sensors that capture data. By adjusting the frames per second and / or breaking unnecessary intermediate frames It is adjusted by discarding.
[0037] At 303, data relating to elements of the time series is received. The relevant data is received at the training server along with the elements received at 301. In the case of a vehicle, the relevant data is the vehicle's odometry data: position, orientation, change in position, orientation Use time series data to identify features identified in the time series. For example, the time series of lane line elements can be labeled as By inspecting the lane lines, they can be labeled with very precise positions. The lane line closest to the vehicle camera is usually accurate and closely related to the vehicle's position. On the other hand, the XYZ position of the line furthest from the vehicle is difficult to determine. The side section may be blocked (e.g., around a corner or behind a hill) and / or may be inaccurate. It may be difficult to accurately capture (e.g., due to distance or lighting). The data related to the elements is labeled with the parts of the features identified in the time series that were identified with high accuracy. In various embodiments, a threshold is used to determine the identified portion of the feature. Determine whether to associate related data (such as part of a lane line) with related data. The parts of the lane line that are identified with high confidence (e.g., the parts closest to the vehicle) are associated with relevant data. and the portion of the lane line that was identified with less than a certainty threshold (e.g., the portion far away from the vehicle). is not associated with the relevant data of that element. Instead, it is associated with the subsequent element with higher accuracy. In some embodiments, another element of the time series and its associated data is used, such as The relevant data can be generated by a neural network, such as the output of the deep learning network 105 in Figure 1. In some embodiments, the relevant data is output from the vehicle control module 10 of FIG. 9. Related data includes speed, speed change, acceleration, Vehicle operating parameters such as acceleration changes, steering, steering changes, braking, and braking changes may be included. In some embodiments, the relevant data may include estimating the distance of an object, such as an obstacle. This is radar data for the purpose of
[0038] In some embodiments, the data relating to the elements of the time series includes map data. For example, in 303, offline data such as road and / or satellite level map data may be used. The map data includes roads, vehicle lanes, intersections, speed limits, and school zones. For example, map data can be used to identify features such as vehicle lane As another example, map data may describe the route of a road or roads on a map. The speed limit imposed can be described.
[0039] In various embodiments, data associated with elements of a time series may include a timestamp associated with the associated data. The time series elements and the corresponding data from the related data are organized. The timestamps can be used to synchronize the two data sets. In an embodiment, the data is synchronized as it is captured. For example, as each element of the time series is captured, A corresponding set of related data is captured and stored with the time series element. In this case, the period of the associated data is configurable and / or matches the period of the element's time series. In some embodiments, the associated data is sampled at the same rate as the time series elements. can be.
[0040] At 305, a ground truth is determined for the time series. In this paper, time series are used as machine learning features. e) are analyzed to determine the ground truth related to the lane line. The lane lines are identified from the time series corresponding to the ground truth of the lane lines. For example, the graphs for the paths of moving objects (e.g., vehicles, pedestrians, bicycles, animals, etc.) The sound truth is the path identified for a moving object detected from a time series. In some embodiments, when a moving vehicle enters the lane of an autonomous vehicle over time, In some embodiments, the moving vehicle is annotated as an interrupting vehicle. The end truth is expressed as a 3D representation such as a 3D trajectory. The ground truth associated with the model is a 3D parameterized spline or curve. As another example, the predicted path of a detected vehicle may be determined and represented as a three-dimensional trajectory. The predicted path is used to determine whether the vehicle is merging into an occupied space. In various embodiments, by examining only the elements of the time series, Ground truth can be determined, e.g., by analyzing only a subset of the time series. This may leave part of the lane line blocked. By extending the analysis, the blocked lane lines are revealed. Data captured towards the end of the sequence provides more detail on the more distant parts of the lane lines. more accurately (e.g., with higher fidelity); and the relevant data is more closely The associated data is also more accurate because it is based on data captured over time (both distance and time). In various embodiments, different parts of an object can be mapped to precise 3D positions, including altitude. To identify detected objects, such as lane lines, identified in different elements of the element time series, Simultaneous localization and mapping techniques are applied to different parts of the body. A set of 3D positions is a set of objects, such as segments of lane lines, captured over a time series. In some embodiments, localization and mapping are performed to represent the ground truth of the body. The mapping technique involves generating an accurate set of points, e.g., points corresponding to different points along a vehicle lane line. The point set can be used to create more efficient curves such as splines or parametric curves. In some embodiments, the ground truth is used to detect objects such as lane lines, driving spaces, traffic control measures, and vehicles in 3D. It is decided for the purpose.
[0041] In some embodiments, the ground truth is determined to predict the semantic label. For example, a detected vehicle can be labeled as being in the left or right lane. In some embodiments, a detected vehicle may be in a blind spot and given way. It can be labeled as a vehicle or with another appropriate semantic label. In this embodiment, the vehicle navigates along roads or are assigned to lanes. As an additional example, to label traffic lights, lanes, drivable spaces, or other features that aid automated driving. It is possible.
[0042] In some embodiments, the associated data is depth (or distance) data of the detected object. By associating the distance data with the objects identified in the time series of elements, The machine learning model uses the relevant distance as ground truth for the detected object. It can be trained to estimate object distance by using the data. In some embodiments, distance may be measured by obstacles, barriers, moving vehicles, stationary vehicles, traffic control signals, pedestrians, etc. The distance to a detected object, such as a person.
[0043] At 307, the training data is packaged. For example, elements of the time series are selected. , is associated with the ground truth determined in 305. The element selected is the initial element in the time series. The ground truth represents the expected results, while the sensor data represents the input data. In various embodiments, the training data is packaged and prepared as training data. In some embodiments, the training data is packaged into training, validation, and test data. Based on the determined ground truth and the selected time series elements, The training data includes lane lines, vehicle position, among other useful features for autonomous driving. Identifying projected routes, speed limits, vehicle cut-offs, object distances, and / or drivable spaces It can be packaged to train machine learning models for The collected training data is then available for training machine learning models.
[0044] FIG. 4 illustrates one embodiment of a process for training and applying machine learning models for autonomous driving. In some embodiments, the process of FIG. 4 is a flow chart illustrating a mechanism for autonomous driving. Collecting and maintaining sensor and odometry data for training machine learning models In some embodiments, the process of FIG. 4 is utilized to enable autonomous driving control. Autonomous driving is implemented in vehicles that are enabled, regardless of whether they are equipped with sensors and Odometry data is used to measure the vehicle's speed immediately after autonomous driving is disengaged and the vehicle is being driven by a human driver. The data may be collected while the vehicle is in operation and / or while the vehicle is in autonomous driving mode. In some embodiments, the technique illustrated by FIG. 4 may be implemented using the deep learning system of FIG. In some embodiments, part of the process of FIG. 4 is implemented as a machine for autonomous operation. As part of the process of applying the learning model, 207, 209, and / or 2 It will be executed at 11.
[0045] At 401, sensor data is received. For example, a vehicle equipped with a sensor It captures data and provides sensor data to a neural network running on the vehicle. In some embodiments, the sensor data may include visual data, ultrasonic data, LiDAR data, or other suitable sensor data. As another example, images are captured from a camera. In some embodiments, the vehicle includes multiple sensors for capturing data. For example, in some embodiments, eight surround cameras are installed in the vehicle. When installed, it provides a 360-degree view around the vehicle with a reach of up to 250 meters. In some embodiments, the camera sensor may be a wide-angle front camera, a narrow-angle front camera, a rear-view camera, camera, forward-looking side cameras, and / or rear-looking side cameras. In embodiments, ultrasonic and / or radar sensors are used to capture details of the surroundings. For example, 12 ultrasonic sensors can be installed on a vehicle to detect both hard and soft objects. In some embodiments, forward-looking radar can be used to obtain data on the surrounding environment. In various embodiments, the radar sensor can detect heavy rain, fog, dust, and other vehicles. Regardless of the vehicle's surroundings, various sensors can capture details of the surroundings. The captured data is then provided for deep learning analysis.
[0046] In some embodiments, the sensor data may include vehicle position, orientation, change in position, and / or or odometry data including changes in orientation, etc. For example, position data is captured, It is correlated with other sensor data captured during the same time frame. The location data captured when the data was captured is used to associate the location information with the image data. I can.
[0047] At 403, the sensor data is preprocessed. One or more preprocessing passes can be performed on the data. For example, the data can be pre-processed to remove noise and correct alignment issues and / or blurriness In some embodiments, one or more different filtering passes may be performed on the data. For example, to separate different components of sensor data, You can perform a high-pass filter on the data and a low-pass filter on the data. In various embodiments, the pre-processing step performed in 403 is optional, and and / or may be incorporated into neural networks.
[0048] At 405, deep learning analysis of the sensor data is initiated. ,Deep learning analysis is performed on the optionally preprocessed sensor data at 403. In various embodiments, the deep learning analysis is performed using a convolutional neural network (CNN). In various embodiments, the machine learning model is implemented using a neural network such as The driver is trained offline using the process in Figure 2 to make inferences on sensor data. For example, the model can detect road lane lines, obstacles, and pedestrians. It may be trained to appropriately identify people, moving vehicles, parked vehicles, drivable spaces, etc. In some embodiments, multiple trajectories of lane lines are identified. Several potential trajectories of the pattern are detected, each with a corresponding probability of occurrence. In some embodiments, the predicted lane lines are those with the highest probability of occurrence and / or is the lane line with the highest associated confidence value. The predicted lane lines from the analysis must exceed a minimum confidence threshold. In
[0003] , the neural network includes multiple layers, including one or more hidden layers. In various embodiments, the sensor data and / or the results of the deep learning analysis are used to generate training data. The data is stored for automatic generation of the data and transmitted at 411.
[0049] In various embodiments, deep learning analysis is used to predict additional features. The predicted features can be used to assist automated driving, for example by identifying detected vehicles. As another example, a detected vehicle can be assigned to a lane or road. There is a vehicle that should give way, there is a vehicle in the lane to the left, there is a vehicle in the lane to the right or another suitable attribute. Similarly, deep learning analytics can be used to Identifying vehicles, drivable spaces, pedestrians, obstacles, or other pertinent features for operation. This can be done.
[0050] At 407, the results of the deep learning analysis are provided to vehicle control. For example, the results are used to automate To control the vehicle for driving and / or to perform automated driving functions, In some embodiments, the results of the deep learning analysis at 405 are provided to the control module. uses one or more different machine learning models to generate one or more additional deep learning For example, the predicted path of lane lines can be used to determine the vehicle lane. The determined vehicle lane can then be used to determine the drivable space. The space is used to determine the path of the vehicle. Vehicle intrusions are detected and the determined path of the vehicle is adjusted to avoid potential collisions. In some embodiments, the various outputs of the deep learning are: Includes the vehicle's predicted path, identified obstacles, and identified traffic control signals, including speed limits. It is used to build a 3D representation of the vehicle's environment for autonomous driving. In one embodiment, the vehicle control module controls a vehicle along the determined path. In some embodiments, the vehicle control module utilizes the results of the vehicle control This is module 109.
[0051] At 409, the vehicle is controlled. In some embodiments, autonomous driving is activated. The vehicle is controlled using a vehicle control module, such as vehicle control module 109 of FIG. Vehicle control, for example, is used to maintain a vehicle in a lane at an appropriate speed while taking into account the surrounding environment. In some embodiments, the vehicle's speed and / or steering may be adjusted to accommodate the The results are used to adjust vehicles in anticipation of adjacent vehicles merging into the same lane. In various embodiments, using the results of the deep learning analysis, the vehicle control module: For example, determining an appropriate way to operate the vehicle along the determined path at an appropriate speed. In various embodiments, the results of vehicle control such as changes in speed, application of braking, and adjustments in steering are retained. and used to automatically generate training data. The data is kept and transmitted at 411 for automatic generation of training data.
[0052] At 411, the sensor data and related data are transmitted. For example, at 401, The acquired sensor data is used in 409 and / or as a result of deep learning analysis in 405. Along with the vehicle control parameters, it is sent to a computer server for automatic generation of training data. In some embodiments, the data is time series data, and various collected data The data are linked together by a computer server. For example, the odometry data is It is associated with the captured image data to generate round truth. In an embodiment, the collected data is transmitted from the vehicle, for example, via a WiFi or cellular connection. In some embodiments, the metadata is transmitted wirelessly to a training data center. For example, metadata may include time, timestamp, location, vehicle type, etc. type, speed, acceleration, braking, whether autonomous driving is enabled, steering angle, odometry data Additional metadata may include vehicle control and / or operational parameters such as The time since the last previous sensor data was sent, the type of vehicle, weather conditions, and road conditions In some embodiments, the data transmitted may include, for example, a unique identifier for the vehicle. As another example, data from similar vehicle models is anonymized by removing , are merged so that individual users and their vehicle usage are not identified.
[0053] In some embodiments, data is transmitted only in response to a trigger. In some embodiments, the incorrect predictions are used as an example key to improve the predictions of the deep learning network. Sensor data collection for automatically collecting data to create an emulated set triggers the transmission of data and related data, e.g., whether the vehicle is about to merge The predictions made in relation to 405 are based on the comparison of the predictions with the actual results observed. The sensor data and the Data, including associated data, is submitted and used to automatically generate training data. In some embodiments, triggers can be used to detect sharp turns, road forks, lane merges, sudden stopping, or another suitable scenario where additional training data is useful and may be difficult to collect. For example, a trigger can identify a specific scenario, such as an autonomous driving feature crash. Another example is based on a change in speed or a change in acceleration. In some embodiments, vehicle operating characteristics such as the speed at which the vehicle is moving can form the basis for a trigger. ,Predictions with accuracy below a certain threshold trigger the transmission of sensor data and related data. For example, in certain scenarios, predictions may not have a Boolean true or false outcome. Instead, it is evaluated by determining the accuracy of the prediction.
[0054] In various embodiments, the sensor data and associated data are captured over a period of time; The entire time series is sent together, including the period, vehicle speed, distance traveled, and speed change. The method may be configured and / or based on one or more factors such as: In some embodiments, the sampling level of the captured sensor data and / or associated data is The sampling rate is configurable. For example, the sampling rate is set to It may be increased during high speeds, sudden maneuvers, or another appropriate scenario where additional fidelity is required.
[0055] FIG. 5 is a diagram showing an example of an image captured from a vehicle sensor. In the illustrated example, The image is image data 50 captured from a vehicle traveling in a lane between two lane lines. The location of the vehicle and sensors used to capture the image data 500 include The image data 500 is sensor data, and is represented by a signal A. The image data 500 can be captured from a camera sensor such as a lane direction camera. The lane lines 501 and 511 are captured. In the example shown, the lanes 501 and 511 curve to the right as they approach the horizon. Lines 501 and 511 are visible, but they are far away from the camera sensor position. As the lane curves, it becomes increasingly difficult to detect. The white lines drawn on the screen are generated from the image data 500 by deriving the lane lines 501 and 502 without any additional input. 11. In some embodiments, the lane lines 501 and The detected portion of 511 is detected by segmenting the image data 500. It is possible.
[0056] In some embodiments, the labels A, B, and C represent different locations and timescales on the road. Label A corresponds to the time of the vehicle when the image data 500 was captured. Label B corresponds to the time and location on the road ahead of the location of label A, Label C corresponds to the position at a time after the time of label A. Similarly, label C corresponds to the position at a time after the time of label B. The position on the road ahead of the device corresponds to the position at a time after the time of label B. As the vehicle travels, it will track the locations of labels A, B, and C (from label A to The vehicle passes through the label C and captures time series of sensor data and related data during the journey. The sequence contains elements captured at positions (and times) labeled A, B, and C. Labels Label A corresponds to the first element of the time series, label B corresponds to the middle element of the time series, and label C corresponds to the middle (or potentially the last) element of the time series. Additional data such as vehicle odometry data at the location is captured. Depending on the length of the time series, Depending on the circumstances, additional or less data may be captured. A time stamp is associated with each element of the time series.
[0057] In some embodiments, the ground truth ( For example, using the processes disclosed herein, lane lines can be determined. The positions of the ins 501 and 511 are determined by the lane lines 501 and 512 from different elements of the time series. and 511 by identifying different parts. The image data 500 and 513 are acquired at the position and time of label A. and associated data (such as odometry data). represents the image data (not shown) and related data ( Parts 507 and 517 are identified using the label C. Image data (not shown) and related data (odometry data) acquired at location and time By analyzing the time series elements, lane lines 50 The locations of the different parts of 1 and 511 are identified, and the different identified parts can be combined. In some embodiments, the ground truth may be determined by The points along each portion of the line are identified. In the example shown, Only three parts of each lane line are highlighted for clarity (parts of lane line 501). 503, 505, and 507 and portions 513, 515, and and 517) determine the location of lane lines with higher resolution and / or greater accuracy. Additional portions may be captured over the time series to capture the entire image.
[0058] In various embodiments, the lane lines 501 and 511 are captured in image data. The position of the part closest to the sensor position is determined with high accuracy. For example, parts 503 and 5 Position 13 is the image data 500 of label A and related data (odometry data, etc.) The locations of parts 505 and 515 are identified with high accuracy using the image data of label B. The locations of portions 507 and 517 are identified with high accuracy using the data and associated data. Label C is identified with high accuracy using image data and related data. By utilizing this, various lane lines 501 and 511 captured by the time series can be The position of the target area is identified in three dimensions with high accuracy, and the ground of the lane lines 501 and 511 is In various embodiments, the determined The ground truth is associated with selected elements of a time series, such as image data 500. The ground truth and selected elements are used as training data for predicting lane lines. In some embodiments, the training data may be used to generate training data for a human. The training data is automatically generated without labeling. To train a machine learning model to predict the 3D trajectory of lane lines from image data can be used for.
[0059] Figure 6 shows an image captured from a vehicle sensor with the predicted 3D trajectory of the lane lines. In the illustrated example, the image of FIG. 6 shows a lane between two lane lines. The image data 600 includes image data captured from a vehicle traveling in the area. The location of the vehicle and the sensor used for this is represented by label A. In the figure, the label A corresponds to the same position as the label A in FIG. 5. data, which may be captured from a camera sensor such as a forward-facing camera in the vehicle while driving. Image data 600 captures portions of lane lines 601, 611. 01 and 611 bay to the right as lane lines 601 and 611 approach the horizon. In the example shown, lane lines 601 and 611 are visible, but As they curve farther away from the camera sensor location, they become increasingly difficult to detect. The red lines painted above lane lines 601 and 611 indicate the lane lines 601 and 611. 11 predicted 3D trajectories. The trajectory is predicted using the image data 600 as input to a trained machine learning model. In some embodiments, the predicted three-dimensional trajectory is calculated using a three-dimensional parameterized spline. or as another parameterized representation format.
[0060] In the illustrated example, portions 621 of lane lines 601 and 611 are spaced apart. Part of lane lines 601 and 611. Part 6 of lane lines 601 and 611 The three-dimensional location (i.e., longitude, latitude, and altitude) of the 21 is determined by the process disclosed herein. The predicted 3D trajectory of lane lines 601 and 611 is determined with high accuracy using a The trained machine learning model is used to generate a 3D image using the image data 600. without needing position data at the location of portion 621 of line lines 601 and 611. The three-dimensional trajectory of the lane lines 601 and 611 can be predicted. Image data 600 is captured at a location and time labeled A.
[0061] In some embodiments, label A in FIG. 6 corresponds to label A in FIG. 5, and lane line 6 The predicted 3D trajectories of 01 and 611 were used as input to a trained machine learning model. The data 600 is determined using only the data 600. The data 600 is obtained at the positions labeled A, B, and C in FIG. Ground-to-ground measurements determined using time-series image data and related data containing By training a machine learning model using Ruth, we can see that lane lines 601 and 61 The 3D trajectory of 1 is predicted with high accuracy even for distant lane line sections such as section 621. The image data 600 and the image data 500 of FIG. 5 are related, but the trajectory prediction It is not necessary that the image data 600 be included in the training data. With practice, you can predict lane lines in new scenarios. In various embodiments, the predicted three-dimensional trajectory of the lane lines 601 and 611 is detected. to maintain the vehicle's position within the lane lines and / or to detect and forecast the lane lines. It is used to autonomously navigate the vehicle along the lanes that are provided. Predicting online navigation significantly improves navigation performance, safety, and accuracy. It will be improved.
[0062] Although the foregoing embodiments have been described in some detail for clarity of understanding, the present invention The invention is not limited to the details provided. There are many alternative ways of implementing the invention. The embodiments are illustrative and not limiting.
Claims
1. 1. A system comprising a non-transitory computer-readable storage medium storing instructions, the instructions, when executed by one or more processors, causing the one or more processors to: acquiring first sensor data generated by one or more cameras of the vehicle at a first time; determining, based on a machine learning model, three-dimensional features associated with the first sensor data, the three-dimensional features including at least one path associated with a predicted position for a first object in an environment along which the vehicle will be operated after the first time point; the machine learning model is trained using a training dataset comprising the determined ground truth of positions for a second object during the time period and the second sensor data captured during the time period to output a determined ground truth based on input of at least a portion of second sensor data generated at a second time during the time period; the determined ground truth indicates three-dimensional features including a path associated with the second sensor data; The system adjusts operation of the vehicle based on the three-dimensional characteristics.
2. The system of claim 1 , wherein the three-dimensional features are three-dimensional trajectories of real-world features.
3. The system of claim 2 , wherein the real-world features are vehicle lane lines.
4. The system of claim 2 , wherein the real-world features are different vehicles.
5. The system of claim 1 , wherein the instructions cause the one or more processors to adjust the speed or steering of the vehicle to adjust the operation of the vehicle.
6. The system of claim 1 , wherein the three-dimensional characteristics are determined based on the first sensor data and odometry data associated with the vehicle.
7. 2. The system of claim 1, wherein the training data set comprises the determined ground truth and the second sensor data including a plurality of time series elements, a portion of the time series elements representing a particular time series element of the plurality of time series elements.
8. 8. The system of claim 7, wherein the ground truth is determined through selecting respective portions of the elements of the time series that are associated with the highest likelihood in depicting the respective portions of real-world features.
9. The system of claim 8 , wherein the real-world feature is a vehicle lane line and at least one unselected portion occludes the vehicle lane line.
10. The system of claim 7 , wherein the ground truth is determined based on elevation information associated with the selected portion and vehicle lane lines.
11. 1. A method implemented by a system of one or more processors, comprising: acquiring first sensor data generated by one or more cameras of the vehicle at a first time; determining, based on a machine learning model, three-dimensional features associated with the first sensor data at the first time point, the three-dimensional features including at least one path associated with a predicted position for a first object in an environment along which the vehicle will be operated after the first time point; the machine learning model is trained using a training dataset comprising the determined ground truth of positions for a second object during the time period and the second sensor data captured during the time period to output a determined ground truth based on input of at least a portion of second sensor data generated at a second time during the time period; the determined ground truth indicates three-dimensional features including a path associated with the second sensor data; and adjusting operation of the vehicle based on the three-dimensional characteristics.
12. The method of claim 11 , wherein the three-dimensional features are three-dimensional trajectories of real-world features.
13. The method of claim 12 , wherein the real-world features are vehicle lane lines.
14. The method of claim 12 , wherein the real-world features are different vehicles.
15. The method of claim 11 , wherein adjusting the operation of the vehicle includes adjusting the speed or steering of the vehicle.
16. 12. The method of claim 11 , wherein the training data set comprises the determined ground truth and the second sensor data including a plurality of time series elements, a portion of the time series elements representing a particular time series element of the plurality of time series elements.
17. 17. The method of claim 16, wherein the ground truth is determined through selecting respective portions of the elements of the time series that are associated with the highest likelihood in depicting the respective portions of real-world features.
18. The method of claim 17 , wherein the real-world feature is a vehicle lane line and at least one unselected portion occludes the vehicle lane line.
19. The method of claim 16 , wherein the ground truth is determined based on elevation information associated with the selected portion and vehicle lane lines.
20. 1. A computer program product embodied in a non-transitory computer-readable storage medium and comprising computer instructions, which when executed by a system of one or more processors, cause the one or more processors to: acquiring first sensor data generated by one or more cameras of the vehicle at a first time; determining, based on a machine learning model, three-dimensional features associated with the first sensor data, the three-dimensional features including at least one path associated with a predicted position for a first object in an environment along which the vehicle will be operated after the first time point; the machine learning model is trained using a training dataset comprising the determined ground truth of positions for a second object during the time period and the second sensor data captured during the time period to output a determined ground truth based on input of at least a portion of second sensor data generated at a second time during the time period; the determined ground truth indicates three-dimensional features including a path associated with the second sensor data; A computer program product that adjusts operation of the vehicle based on the three-dimensional characteristics.
Citation Information
Patent Citations
Simulation system, simulation program, and simulation method
JP2018060511A
Travel control method and travel control device
WO2018235171A1