Systems and methods for controlling autonomous driving of a vehicle
The system uses satellite and infrastructure cameras with machine learning to predict virtual lanes at unlabeled intersections, addressing navigation challenges and ensuring safe autonomous vehicle operation.
Patent Information
- Application Number
- DE102025102286
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2026-02-05
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Autonomous vehicles face challenges in navigating unlabeled lane intersections due to the absence of lane markings, which complicates high-resolution mapping and mapless driving scenarios.
A system utilizing satellite and infrastructure-mounted cameras, combined with vehicle cameras, processes aerial and infrastructure images through machine learning models to predict virtual lanes, enabling autonomous control of steering, acceleration, and braking.
Enables reliable navigation through unlabeled intersections by generating accurate virtual lane predictions, allowing for safe and efficient autonomous vehicle operation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
IntroductionThe information provided in this section is for the purpose of generally presenting the context of the disclosure. The work of the present inventors, insofar as described in this section, as well as aspects of the description that cannot be qualified as prior art at the time of filing, are neither expressly nor silently accepted as prior art against the present disclosure.The present disclosure relates generally to autonomous vehicle control at unlabeled lane intersections that includes predicting virtual lane positions based on aerial images, infrastructure-mounted cameras, and vehicle cameras.Autonomous vehicles may use vehicle cameras to detect lane markings on road surfaces automated steering control. Intersection scenarios may present a challenge for autonomous vehicle navigation if lane markings are missing at the intersection.DE 10 2018 131 477 A1 discloses systems and methods including an artificial neural network for classifying and locating lane features. DE 10 2021 102 426 A1 discloses a method for operating a vehicle, comprising receiving a number of images with a view from above of a surrounding area of the vehicle. Based on the images, a digital map of the surroundings and based thereon a trajectory from a current position of the vehicle to a target position is ascertained, and driving along the ascertained trajectory is initiated by a control unit of the vehicle.SummaryThe present invention relates to systems and a method for controlling autonomous driving of a vehicle according to the independent claims. Embodiments are given in the dependent claims, the description and the drawings.An example system for controlling autonomous driving of a vehicle includes at least one satellite configured to obtain one or more aerial images of an intersection, the intersection including at least one unlabeled lane, an infrastructure-mounted camera configured to obtain one or more infrastructure camera images of the intersection, the infrastructure-mounted camera oriented with a viewing angle toward the intersection, a front vehicle camera configured to capture images from a front field of view of a vehicle, a wireless interface of the vehicle, the wireless interface configured to wirelessly receive the one or more aerial images and the one or more infrastructure camera images, and a vehicle control module configured to receive the one or more aerial images and the one or more infrastructure camera images, accessing the one or more aerial images and the one or more infrastructure camera images, obtaining one or more vehicle camera images from the front vehicle camera, providing the one or more aerial images, the one or more infrastructure camera images, and the one or more vehicle camera images to at least one machine learning model, generating a virtual lane prediction according to an output of the at least one machine learning model, wherein the virtual lane prediction indicates one or more unlabeled lanes present in the intersection, and automatically controlling steering, acceleration, and braking of the vehicle in the intersection according to the virtual lane prediction.In some examples, providing the one or more aerial images includes providing the one or more aerial images to a machine learning model for a transformer encoder. In some examples, providing includes providing the one or more infrastructure camera images and the one or more vehicle camera images to a machine learning model for a bird's eye view (BEV) encoder.In some examples, the vehicle control module is configured to provide an output of the machine learning model to a transformer encoder as an input to a machine learning model to a map decoder, provide map query data and map loss data as inputs to the machine learning model to a map decoder, and generate the virtual lane prediction based at least in part on an output of the machine learning model to a map decoder.In some examples, the vehicle control module is configured to provide an output of the machine learning model to a transformer encoder as input to a machine learning model to an auxiliary actuator decoder, provide actuator sample data and actuator loss data as inputs to the machine learning model to an auxiliary actuator decoder, and generate the virtual lane prediction based at least in part on an output of the machine learning model to an auxiliary actuator decoder.In some examples, the vehicle control module is configured to associate a subset of actuator features of the machine learning model for an auxiliary actuator decoder with a list of map feature inputs of the machine learning model for a map decoder. In some examples, the vehicle control module is configured to provide the one or more aerial images to a convolutional neural network and provide an output of the convolutional neural network to a machine learning model for a transformer encoder.In some examples, automatically controlling steering, acceleration, and braking of the vehicle includes controlling operation of the vehicle via map-free autonomous driving.In some examples, the vehicle control module is configured to receive the one or more vehicle camera images from the front vehicle camera in real-time in response to the vehicle approaching the intersection, and access the one or more aerial images and the one or more infrastructure camera images from a stored memory, wherein access to the one or more aerial images and the one or more infrastructure camera images are previously stored images.In some examples, the vehicle control module is configured to determine an Intersection Over Union (IoU) loss for bounding boxes of predicted lanes and ground truth labels for the machine learning model for a map decoder. In some examples, the vehicle control module is configured to determine a cosine of an angle between principal alignments of the predicted actuator alignments and the ground truth alignment labels for the machine learning model for an auxiliary actuator decoder.An example method for controlling autonomous driving of a vehicle includes obtaining, via a wireless interface of a vehicle, one or more aerial images of an intersection, wherein the intersection includes at least one unlabeled lane and the one or more aerial images are captured using at least one satellite, obtaining, via the wireless interface of the vehicle, one or more infrastructure camera images of the intersection from an infrastructure-mounted camera, wherein the infrastructure-mounted camera is oriented with a viewing angle toward the intersection, receiving one or more vehicle camera images from a front vehicle camera of the vehicle, providing the one or more aerial images, the one or more infrastructure camera images, and the one or more vehicle camera images for at least one machine learning model, generating a virtual lane prediction according to an output of the at least one machine learning model, the virtual lane prediction indicating one or more unlabeled lanes present in the intersection, and automatically controlling steering, acceleration, and braking of the vehicle in the intersection according to the virtual lane prediction.In some examples, providing the one or more aerial images includes providing the one or more aerial images to a machine learning model for a transformer encoder. In some examples, providing includes providing the one or more infrastructure camera images and the one or more vehicle camera images to a machine learning model for a bird's eye view (BEV) encoder.In some examples, the method includes providing an output of the machine learning model to a transformer encoder as input to a machine learning model to a map decoder, providing map query data and map loss data as inputs to the machine learning model to a map decoder, and generating the virtual lane prediction based at least in part on an output of the machine learning model to a map decoder.In some examples, the method includes providing an output of the machine learning model to a transformer encoder as input to a machine learning model to an auxiliary actor decoder, providing actor sample data and actor loss data as inputs to the machine learning model to an auxiliary actor decoder, and generating the prediction of virtual lanes based at least in part on an output of the machine learning model to an auxiliary actor decoder.In some examples, the method includes associating a subset of actor features of the machine learning model for an auxiliary actor decoder with a list of map feature inputs of the machine learning model for a map decoder. In some examples, the method includes providing the one or more aerial images to a convolutional neural network and providing an output of the convolutional neural network to a machine learning model to a transformer encoder.In some examples, the method includes automatically controlling steering, acceleration, and braking of the vehicle, controlling operation of the vehicle via map-free autonomous driving.An example system for controlling autonomous driving of a vehicle includes a front vehicle camera configured to capture images from a front field of view of a vehicle, a wireless interface configured to wirelessly receive one or more aerial images of an intersection, the intersection including at least one unlabeled lane and the one or more aerial images being captured using at least one satellite, and wirelessly receive one or more infrastructure camera images of the intersection from an infrastructure mounted camera, the infrastructure mounted camera oriented with a viewing angle toward the intersection, and a vehicle control module configured to access the one or more aerial images and the one or more infrastructure camera images, one or more vehicle camera images are obtained from the front vehicle camera, the one or more aerial images providing the one or more infrastructure camera images and the one or more vehicle camera images to at least one machine learning model generate a virtual lane prediction according to an output of the at least one machine learning model, wherein the virtual lane prediction indicates one or more unlabeled lanes present in the intersection, and automatically controls steering, acceleration, and braking of the vehicle in the intersection according to the virtual lane prediction.Further areas of applicability of the present disclosure will become apparent from the detailed description, claims and drawings. The detailed description and specific examples are intended for purposes of illustration only and are not intended to limit the scope of the disclosure.Brief Description of the DrawingsThe present disclosure will be more fully understood from the detailed description and the accompanying drawings. FIG. 1 is a diagram of an example vehicle including a vehicle control module configured to determine virtual lanes at unlabeled intersections. FIG. 2 is an exemplary top view of a road intersection with unlabeled lanes. FIG. 3 is a block diagram of an example system for predicting virtual lanes at an intersection using aerial images and a transformer encoder. FIG. 4 is a block diagram of an example system for predicting virtual lanes at an intersection using infrastructure camera images and vehicle camera images using a bird's eye view generator. FIG. 5 is a flowchart illustrating an example process for predicting virtual lanes in an intersection using aerial images and a transformer encoder. FIG. 6 is a flowchart illustrating an example process for predicting virtual lanes at an intersection using infrastructure camera images and vehicle camera images using a bird's-eye view generator. FIG. 7 is a flow diagram illustrating an example process for modifying inputs of map input features of a map decoder based on corresponding actuator features. FIG. 8 is a flowchart illustrating an example process for determining virtual lanes at unlabeled intersections. FIGS. 9A and 9B are graphical representations of example recurrent neural networks for predicting virtual lanes at unlabeled intersections. FIG. 10 is a graphical representation of layers of an example long short-term memory (LSTM) machine learning model. FIG. 11 is a flowchart illustrating an example process for training a machine learning model.In the drawings, reference numerals may be reused to identify similar and / or identical elements.Detailed DescriptionAutonomous vehicles may use vehicle cameras to detect lane markings on road surfaces for automated steering control. Intersection scenarios may present a challenge for autonomous vehicle navigation when lane markings are missing at the intersection, particularly for high resolution mapping processes and high resolution mapless autonomous driving. In some example embodiments herein, geometric and topological structures of lanes in an intersection are derived via a vehicle control module parent scene understanding.For example, a uniform learning-based method may be implemented to derive regular lanes and road features (e.g., pedestrian crossings) and virtual lanes at intersections. A vehicle control module may determine virtual lanes using one or more trained machine learning models that receive inputs of aerial images of an intersection, vehicle camera images, and infrastructure-based cameras (e.g., cameras mounted on street lights, buildings, etc., in the area of an intersection).In some examples, the virtual lane system may be configured based on other implicit indications in the scene. For example, a consistent and versatile decoder design for detecting lane instances may receive based on aerial images, vehicle camera images, and / or infrastructure camera images. A detailed transformer-based multi-head decoder design may make maximum use of implicit scene knowledge for detailed vectored detection of lane copies in ongoing operation, which may include lane types, lane edge types, centerlines, lane edge offsets, etc. as an integral lane representation.In some examples, a specific attention mechanism may implicitly integrate relevant information about actuator alignment into the detection of lane copies. Specific matching criteria may be used to associate prediction results with ground truth labels to facilitate training a proposed deep neural network design. Detailed loss functions may be used to train map and actor decoders with detailed separable attention mechanisms. In some examples, information may be dragged from the actor decoder into the map decoder for enhanced map derivations.Referring now to FIG. 1, a vehicle 10 includes front wheels 12 and rear wheels 13. In FIG. 1, a propulsion unit 14 selectively outputs torque to the front wheels 12 and the rear wheels 13, respectively, via powertrains 16, 18. The vehicle 10 may include various types of propulsion units. For example, the vehicle may be an electric vehicle such as a battery electric vehicle (BEV), a hybrid vehicle or a fuel cell vehicle, an internal combustion engine (ICE) vehicle, or another type of vehicle.Some examples of the drive unit 14 may include any suitable electric motor, a power inverter, and a motor controller configured to control power switches within the power inverter to adjust motor speed and torque during drive and / or regeneration. During propulsion or regeneration, a battery system provides power to or receives power from the electric motor of the propulsion unit 14 via the power inverter.While the vehicle 10 in FIG. 1 includes a propulsion unit 14, the vehicle 10 may have other configurations. For example, two separate propulsion units may propel the front wheels 12 and the rear wheels 13, one or more individual propulsion units may propel individual wheels, etc. As can be envisioned, other vehicle configurations and / or propulsion units may be used.The vehicle control module 20 may be configured to control operation of one or more vehicle components, such as the propulsion unit 14 (e.g., by commanding torque settings of an electric motor of the propulsion unit 14). The vehicle control module 20 may receive inputs for controlling components of the vehicle, such as signals received from a steering wheel, an accelerator pedal, a brake pedal, etc. The vehicle control module 20 may monitor telematics of the vehicle such as vehicle speed, vehicle location, vehicle braking and acceleration, etc. for safety reasons.The vehicle control module 20 may receive signals from any suitable components for monitoring one or more aspects of the vehicle including one or more vehicle sensors (such as cameras, microphones, pressure sensors, steering wheel position sensors, brake sensors, location sensors such as global positioning system (GPS) antennas, wheel height and / or position sensors, accelerometers, etc.). Some sensors may be configured to monitor current motion of the vehicle, acceleration of the vehicle, braking of the vehicle, current steering direction of the vehicle, current height and / or position of one or more wheels, etc.In the example of FIG. 1, the vehicle 10 includes a front vehicle camera 22, an optional side vehicle camera 24, and an optional rear vehicle camera 26. In some examples, images from vehicle cameras may be used for object detection, automated driving, virtual lane determination at unlabeled intersections, etc. Other example embodiments may include more or fewer cameras or cameras at other locations on the vehicle 10. Other systems, such as lidar, may be used to determine images or information about the environment of the vehicle.The vehicle control module 20 may communicate with another device via a wireless communication interface 28, which may include one or more wireless antennas for transmitting and / or receiving wireless communication signals. For example, the wireless communication interface 28 may communicate via any suitable wireless communication protocols including, but not limited to, vehicle-to-everything (V2X) communication, Wi-Fi communication, wireless area network (WAN) communication, cellular communication, personal area network (PAN) communication, short-range wireless communication (e.g., Bluetooth), etc. The wireless communication interface 28 may communicate with a remote computing device via one or more wireless and / or wired networks. With respect to vehicle-to-vehicle (V2X) communication, the vehicle 10 may include one or more V2X transceivers (e.g., V2X signal transmitting and / or receiving antennas).As shown in FIG. 1, the wireless communication interface 28 is configured to receive images from aerial cameras 30. For example, one or more satellites may obtain images of an intersection that are transmitted to the vehicle or stored in a vehicle memory (or on a server) to facilitate lane determination at an unlabeled intersection.Similarly, the wireless communication interface 28 may be configured to receive images from infrastructure-mounted cameras 32. For example, cameras mounted on a street lamp of the intersection, a building adjacent the intersection, other infrastructure features proximate the intersection, etc., may be configured to capture images of the intersection that may be used by the vehicle control module 20 to determine virtual lane markings for the intersection.The vehicle control module 20 may determine virtual lane markings for the intersection in real time based on images of the vehicle cameras using real-time or previously stored images from the infrastructure-mounted cameras 32 or aerial cameras 30. For example, as will be further described below, images from aerial cameras 30, infrastructure-mounted cameras 32, and vehicle cameras may be provided to one or more machine learning models to determine virtual lane markings at an intersection. The vehicle control module 20 may be configured to use the predicted virtual lane markings of the intersection to automatically control the acceleration of the vehicle 10 (e.g., via an accelerator pedal or controlling an engine of the propulsion unit 14 to provide more power to the front wheels 12 and the rear wheels 13), to control the braking of the vehicle 10 (e.g., via brakes applied to the front wheels 12 and the rear wheels 13, or via engine braking on an engine of the propulsion unit 14), to control the automated steering of the vehicle 10 (e.g., by turning a steering mechanism or directly changing an orientation of the front wheels 12), etc.FIG. 2 is an exemplary top view of a road intersection with unlabeled lanes. As shown in FIG. 2, a host vehicle 200 is traveling in a marked lane 204 approaching an intersection 202. The intersection 202 does not have lane markings (or only partial lane markings).For example, the marked lane 204 may have lines painted onto the road (e.g., in white or yellow) to mark the boundaries of the lane. The intersection 202 may not have painted lane lines or other markings, or may have only some lane lines, while other portions of the intersection 202 may not have lane markings.As further described herein, a vehicle control module of the host vehicle 10 may determine virtual lanes 208 of the intersection 202 based on vehicle camera images, aerial images of the intersection 202, and images of infrastructure-mounted cameras proximate the intersection 202. As shown in FIG. 2, the vehicle control module has determined that two virtual lanes are straight across the intersection 202 while the rightmost lane is a right turn virtual lane. As described herein, a virtual lane or a virtual lane marker may refer to predicting lane lines or boundaries of a virtual lane, a centerline of the path of a virtual lane, etc.FIG. 3 is a block diagram of an example system for predicting virtual lanes at an intersection using aerial images and a transformer encoder. As shown in FIG. 3, aerial images (e.g., from one or more satellites capturing images of an intersection) may be provided to a machine learning model, such as a convolutional neural network backbone 306.An output of convolutional neural network backbone 306 is provided to another machine learning model, such as transformer encoder 308. An output of the transformer encoder 308 is then provided to two different models, including an auxiliary actor decoder 310 and a map decoder 312.Auxiliary actor decoder (Auxiliary actor decoder) 310 is configured to receive data from actor queries 314 and actor losses 316. The auxiliary actuators may include other vehicles in or near the intersection, other vehicles passing the intersection, pedestrians walking on a sidewalk or a zebra strip of the intersection, etc.The auxiliary actor decoder 310 may be configured to generate a prediction output of virtual lanes of the intersection based on the actor input. Additional details regarding the auxiliary actuator decoder 310 are described below with reference to FIG. 6.The card decoder 312 is configured to receive data from card queries 318 and card losses 320. The map input data may include features present in an aerial view of the intersection from above. The map decoder 312 may be configured to generate a virtual lane prediction output of the intersection based on the map input. Additional details regarding the map decoder 312 are described below with reference to FIG. 5.As shown in FIG. 3, the models may implement a relationship between the auxiliary actor decoder 310 and the map decoder 312 according to map actor orientation interactions 322 (such as associating some actor features with the map features when the actor features meet a confidence threshold and / or a distance threshold). Additional details regarding the map-actuator orientation interactions 322 are described below with reference to FIG. 7.FIG. 4 is a block diagram of an example system for predicting virtual lanes at an intersection using infrastructure camera images and vehicle camera images using a bird's eye view generator. As shown in FIG. 4, infrastructure camera images (e.g., from one or more street lights at the intersection, buildings near the intersection, road features adjacent the intersection, etc.) may be provided to a machine learning model such as a convolutional neural network backbone 406.Vehicle camera images 404 such as images captured by the front vehicle camera 22 of the vehicle 10 in FIG. 1, the optional side camera 24, or the optional rear camera 26 may also be provided to the convolutional neural network backbone 406.An output of the convolutional neural network backbone 406 is provided to another machine learning model, such as a bird's eye view encoder 408 (e.g., a BEV shaper). The bird's-eye view encoder 408 may be configured to convert (convert) the images of the infrastructure cameras and the vehicle cameras captured at horizon perspective angles into a top bird's-eye view similar to aerial images (which may be more suitable for predicting lane lines of the intersection from a top-down perspective). An output of bird's-eye view encoder 408 is then provided to two different models, including an auxiliary actor decoder 410 and a map decoder 412.The auxiliary actor decoder 410 is configured to receive data from actor queries 414 and actor losses 416. The auxiliary actuators may include other vehicles in or near the intersection, other vehicles passing the intersection, pedestrians walking on a sidewalk or a zebra strip of the intersection, etc. The auxiliary actor decoder 410 may be configured to generate a prediction output of virtual lanes of the intersection based on the actor input.The card decoder 412 is configured to receive data from card queries 418 and card losses 420. The map input data may include features present in an aerial view of the intersection from above. The map decoder 412 may be configured to generate a virtual lane prediction output of the intersection based on the map input. As shown in FIG. 4, the models may implement a relationship between the auxiliary actor decoder 410 and the map decoder 412 according to map actor orientation interactions 422 (such as associating some actor features with the map features when the actor features meet a confidence threshold and / or a distance threshold).FIG. 5 is a flowchart illustrating an example process for predicting virtual lanes in an intersection using aerial images and a transformer encoder. The process may be performed by, for example, the vehicle control module 20 of FIG. 1. At 504, the process begins by obtaining input values for a current lane instance, such as a lane in which a host vehicle is currently traveling.The vehicle control module is configured to obtain, at 508, average feature input values for all other lanes, such as lanes on a right side and a left side of a lane in which the host vehicle is currently traveling. The vehicle control module then provides 512 the input values to a transformer model that includes cross-attention and self-attention.The vehicle control module is configured to output classifications for a current lane, a left lane, and a right lane in 516. The output may have focus losses for the lane type and the edge type. The vehicle control module is configured to determine the intersection over union (IoU) loss for bounding boxes of predicted lanes and ground truth labels at 520.The vehicle control module is configured to determine a cosine of an angle between a principal orientation of predicted lanes and ground truth labels at 524. Control then determines predicted virtual lanes of the intersection based on the model output at 528.FIG. 6 is a flowchart illustrating an example process for predicting virtual lanes in an intersection using infrastructure camera images and vehicle camera images using a bird's-eye view generator. The process may be performed by, for example, the vehicle control module 20 of FIG. 1. At 604, the process begins by obtaining input values for actuators near a current lane, such as other vehicles, pedestrians, etc., near a current lane of a host vehicle.The vehicle control module is configured to provide the input values to a transformer model that includes cross-attention and self-attention at 612. The vehicle control module is configured to output classifications for actuator types that may have focus losses at 616.The vehicle control module is configured to determine losses for actuator position and orientation regression at 620. The controller then determines a cosine of an angle between a main orientation of the predicted actuator orientation and the ground truth orientation labels at 624. The vehicle control module is configured to determine predicted virtual lanes of the intersection based on the model output at 628.FIG. 7 is a flow diagram illustrating an example process for modifying inputs of map input features of a map decoder based on corresponding actuator features. The process may be performed by, for example, the vehicle control module 20 of FIG. 1. At 704, the process begins by obtaining map-based model features (such as input features corresponding to aerial images or geographical positions of features with respect to the intersection).The vehicle control module is configured to obtain actuator-based model features, such as features related to other vehicles, pedestrians at the intersection, etc., at 708. The controller then selects a first actor feature from the list of actor features in 712 and compares the actor feature to a specified confidence score threshold.If the confidence score of the selected actor feature is not above the specified confidence score threshold at 716, control passes to 728 to determine if there are any additional actor features remaining in the list. If the confidence score is greater than the specified confidence score threshold at 716, control continues to 720 to compare a distance score of the actuator feature to a specified distance score threshold.If the distance score of the actor features is less than the distance score threshold at 720, control passes to 728 to determine if there are any additional actor features remaining in the list. If the distance score is greater than the specified distance score threshold at 720, control passes to 724 to add the actor feature (such as by associating the actor feature with a list of inputs for processing by the map decoder) to the list of map features.If any actor features are left from the list at 728, control passes to 732 to select a next actor feature from the list. Once all actuator features are processed at 728, control continues to 736 to process the linked list (e.g., the actuator features added to the map feature inputs) using multilayer perceptron (MLP) networks and a feed-forward neural network.FIG. 8 is a flowchart illustrating an example process for determining virtual lanes at unlabeled intersections. The process may be performed by, for example, the vehicle control module 20 of FIG. 1. At 804, the process begins by obtaining a lane marker status for an upcoming intersection. For example, the vehicle control module may determine whether an upcoming intersection has physical lane markings painted onto the roadway surface of the intersection based on stored map data and / or vehicle camera images.If, at 808, the vehicle control module determines that the lanes of the intersection are marked on the road (e.g., fully marked, with all lanes visible on the roadway), control continues to 832 to automatically control steering, braking, and acceleration of the vehicle based on the marked lanes in the intersection. If the intersection is not fully marked at 808, control continues to 812 to obtain aerials, infrastructure camera images, and vehicle camera images of the intersection.The vehicle control module may obtain one or more images of any type that may be captured in real time or obtained from previously stored images. The vehicle control module is configured to provide the aerial images to a transformer encoder to generate a prediction output at 816.The vehicle control module is configured to provide infrastructure and vehicle images to a bird's eye view encoder (e.g., BEV shaper) to generate a prediction output at 820. Control then determines virtual intersection lanes based on model prediction outputs in 824. The vehicle control module is configured to automatically control steering, acceleration, and braking of the vehicle based on the determined or predicted virtual lane lines in the intersection at 828.FIGS. 9A and 9B show an example of a recurrent neural network used to generate models such as those described above using machine learning techniques. Machine learning is a method used to design complex models and algorithms that are suitable for prediction (e.g., patient and provider matching predictions). The models generated using machine learning, such as those described above, may generate reliable, repeatable decisions and results and reveal hidden findings in the data by learning from historical relationships and trends.The purpose of using the recurrent neural network-based model and training the model using machine learning as described above may be to directly predict dependent variables without bringing the relationships between the variables into a mathematical form. The neural network model includes a large number of virtual neurons that operate in parallel and are arranged in layers. The first layer is the input layer 903 and receives raw input data 901. Each subsequent layer modifies outputs of a previous layer and sends them to a next layer. The last layer is the output layer 907 and generates the output 909 of the system.FIG. 9A shows a fully connected neural network in which each neuron in a given layer is connected to each neuron in a next layer. In the input layer, each input node is assigned a numerical value, which can be any desired real number. In each layer, each link originating from an input node has a weight associated with it, which may also be any real number (see Figure 9B). In the input layer, the number of neurons is equal to the number of features (columns) in a data set. The output layer may have a plurality of continuous outputs.The layers between the input layers 903 and the output layers 907 are hidden layers 905. The number of hidden layers may be one or greater (one hidden layer may be sufficient for most applications). A neural network without hidden layers may represent linearly separable functions or decisions. A hidden layer neural network can perform continuous mapping from one finite space to another. A two hidden layer neural network can approximate any smooth mapping with any accuracy.The number of neurons can be optimized. At the beginning of training, a network configuration is more likely to have excess nodes. Some of the nodes may be removed from the network during training, which would not appreciably degrade network performance. For example, nodes with weights that approach zero after training may be removed (this process is referred to as pruning). The number of neurons may cause underfitting (inability to adequately capture signals in a dataset) or overfitting (insufficient information to train all neurons; the network works well on a training dataset but not on a test dataset).Various methods and criteria may be used to measure the performance of a neural network model. For example, root mean squared error (RMSE) measures the average distance between observed values and model predictions. The coefficient of certainty (R2) measures a correlation (not the accuracy) between the observed and predicted results. This method may not be reliable if the data has a large variance. Other performance metrics include irreducible noise, model bias, and model variance. A high model bias for a model indicates that the model is unable to grasp an actual relationship between the predictors and the result. The model variance may indicate whether a model is stable (a minor perturbation in the data will significantly change the model fit). The neural network may receive inputs, e.g., vectors, that may be used to generate models that may be used to predict virtual lanes in unlabeled intersections based on aerial images, infrastructure camera images, and vehicle images.FIG. 10 illustrates an example of a long short-term memory (LSTM) neural network 1002. The neural network LSTM is an example of a machine learning model, and various example implementations may use other machine learning models such as a combination of transformers and a multilayer perceptron (MLP) set prediction (e.g., MapTR) in the decoding process. For example, while LSTM may be used to output polylines modeling derived virtual lanes within intersections, other model types may be used to output desired virtual lanes, such as, for example, a combination of transformers and an MLP set predictions to decode the encoded virtual lane queries into a vector of x, y coordinates.The generic example of an LSTM neural network 1002 may be used to implement a machine learning model, and various implementations may use other types of machine learning networks (such as transformer layers, MLP set projections such as MapTR, other model topologies or architectures, etc.). The LSTM neural network 1002 includes an input layer 1004, a hidden layer 1008, and an output layer 1012. The input layer 1004 includes inputs 1004 a, 1004 b... 1004n, input data 1001a, 1001a... 1001n. Hidden layer 1008 includes neurons 1008 a, 1008 b... 1008n. The output layer 1012 includes outputs 1012 a, 1012 b... 1012n.Each neuron of hidden layer 1008 receives an input from input layer 1004 and outputs a value to the corresponding output in output layer 1012. For example, neuron 1008a receives an input from input 1004a and outputs a value to output 1012a. Each neuron, except neuron 1008a, also receives an output of a previous neuron as an input. For example, neuron 1008b receives inputs from input 1004b and output 1012a. In this manner, the output of each neuron is forwarded to the next neuron in hidden layer 1008. The last output 1012n in the output layer 1012 outputs a probability 1016 associated with the inputs 1004a-1004n. Although the input layer 1004, the hidden layer 1008, and the output layer 1012 are shown as each comprising three elements, each layer may include any number of elements.In various implementations, each layer of the LSTM neural network 1002 must include the same number of elements as each of the other layers of the LSTM neural network 1002. In some example embodiments, a convolutional neural network may be implemented. Similar to LSTM neural networks, convolutional neural networks include an input layer, a hidden layer, and an output layer. In a convolutional neural network, however, the output layer contains one output less than the number of neurons in the hidden layer and each neuron is connected to each output. In addition, each input in the input layer is connected to each neuron in the hidden layer. In other words, input 1004 ais coupled to each of neurons 1008 a, 1008 b... 1008n.In various implementations, each input node in the input layer may be assigned a numerical value, which may be any real number. In each layer, each link originating from an input node is assigned a weight, which can also be any real number. In the input layer, the number of neurons is equal to the number of features (columns) in a data set. The output layer may have a plurality of continuous outputs.As mentioned above, the layers between the input and output layers are hidden layers. The number of hidden layers may be one or greater (one hidden layer may be sufficient for many applications). A neural network without hidden layers may represent linearly separable functions or decisions. A hidden layer neural network can perform continuous mapping from one finite space to another. A two hidden layer neural network can approximate any smooth mapping with any accuracy. The neural network of FIG. 10 may receive inputs, e.g., vectors, that may be used to generate models to predict, for example, virtual lanes in unlabeled intersections based on aerial images, infrastructure camera images, and vehicle images.FIG. 11 illustrates an example process for generating a machine learning model. At 1107, the controller receives data from a database 1102 (e.g., a data warehouse). The data may include any suitable data for developing machine learning models.At 1111, control separates the data obtained from database 1102 into training data 1115 and test data 1119. Training data 1115 is used to train the model at 1123, and test data 1119 is used to test the model at 1127. Typically, the set of training data 1115 is selected to be larger than the set of test data 1119 depending on the desired parameters for model development. For example, training data 1115 may include about seventy percent of the data obtained from database 1102, about eighty percent of the data, about ninety percent, etc. The remaining thirty percent, twenty percent, or ten percent are then used as the test data 1119.Separating a portion of the obtained data as test data 1119 allows testing of the trained model against actual output data to allow more accurate training and development of the model at 1123 and 1127. The model may be trained at 1123 using any suitable techniques for machine learning models, including those described herein, such as random forest, generalized linear models, decision trees, and neural networks.At 1131, the controller evaluates the model test results. For example, the trained model may be tested at 1127 using the test data 1119, and the results of the output data from the tested model may be compared to actual outputs of the test data 1119 to determine a level of accuracy. The model results may be evaluated using any suitable analysis of machine learning models, such as the example techniques described below.After evaluating the model test results at 1131, if the model test results are satisfactory, the model may be employed at 1135. The use of the model may include using the model to make predictions for a large input dataset with unknown outputs. If the evaluation of the model test results at 1131 is not satisfactory, the model may be advanced using different parameters, using different modeling techniques, using different model types, etc. The machine learning model method of FIG. 11 may receive inputs, e.g., vectors, that may be used to generate models that may be used to predict, for example, virtual lanes at unlabeled intersections based on aerial images, infrastructure camera images, and vehicle images.The foregoing description is merely illustrative in nature and is intended to limit the disclosure, its application, or uses. The broad teachings of the disclosure may be embodied in a variety of forms. Therefore, while this disclosure includes particular examples, the true scope of the disclosure should not be so limited since other modifications will become apparent upon a study of the drawings, the specification, and the following claims. It should be understood that one or more steps within a method may be performed in different orders (or concurrently) without altering the principles of the present disclosure. Further, although each of the embodiments is described above as having particular features, one or more of those features described with respect to any embodiment of the disclosure may be implemented in one of the other embodiments and / or combined with features of any of the other embodiments, even if that combination is not expressly described. In other words, the described embodiments are not mutually exclusive, and permutations of one or more embodiments with each other remain within the scope of this disclosure.Spatial and functional relationships between elements (e.g., between modules, circuit elements, semiconductor layers, etc.) are described using various terms, including "connected," "engaged," "coupled," "adjacent," "near," "on," "above," "below," and "disposed.". Unless expressly described as "direct", when a relationship between first and second elements is described in the above disclosure, this relationship may be a direct relationship in which no other intervening elements are present between the first and second elements, but may also be an indirect relationship in which one or more intervening elements (either spatially or functionally) are present between the first and second elements. As used herein, the phrase "at least one of A, B, and C" is intended to mean a logical (A OR B OR C) using a non-exclusive logical OR, and should not be understood to mean "at least one of A, at least one of B, and at least one of C.".In the figures, the direction of an arrow as indicated by the arrow head generally illustrates the flow of information (e.g., data or instructions) of interest for the illustration. For example, if element A and element B exchange a variety of information, but information transmitted from element A to element B is relevant for illustration, the arrow may point from element A to element B. This unidirectional arrow does not imply that no other information is transmitted from element B to element A. Moreover, for information transmitted from element A to element B, element B may transmit requests for, or acknowledgments of, the information to element A.In this application, including the definitions below, the term "module" or the term "controller" may be replaced by the term "circuit". "Module" may refer to a circuit, part of which, or may include an application specific integrated circuit (ASIC); a digital, analog, or mixed analog / digital discrete circuit; a digital, analog, or mixed analog / digital integrated circuit; a combinational logic circuit; a field programmable gate array (FPGA); a processor circuit (shared, dedicated, or group) that executes code; a memory circuit (shared, dedicated, or group) that stores code executed by the processor circuit; other suitable hardware components that provide the described functionality; or a combination of some or all of the above components, such as in a system-on-chip.The module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces connected to a local area network (LAN), the Internet, a wide area network (WAN), or combinations thereof. The functionality of any given module of the present disclosure may be distributed among multiple modules connected via interface circuits. For example, multiple modules may allow load balancing. In another example, a server module (also known as a remote or cloud module) may perform some functions for a client module.The term code as used above may include software, firmware, and / or microcode, and may refer to programs, routines, functions, classes, data structures, and / or objects. The term shared processor circuit includes a single processor circuit that executes some or all of the code from multiple modules. The term group processor circuit includes a processor circuit that, in combination with additional processor circuits, executes some or all of the code from one or more modules. References to multiple processor circuits include multiple processor circuits on single chips, multiple processor circuits on a single chip, multiple cores of a single processor circuit, multiple threads of a single processor circuit, or a combination of the foregoing. The term shared memory circuit includes a single memory circuit that stores some or all of the code from multiple modules. The term group memory circuit includes a memory circuit that, in combination with additional memories, stores some or all of the code from one or more modules.The term memory circuitry is a subset of the term computer readable medium. The term computer-readable medium as used herein does not include transitory electrical or electromagnetic signals propagating through a medium (such as on a carrier wave); the term computer-readable medium may therefore be considered tangible and non-transitory. Non-limiting examples of a non-transitory, tangible computer readable medium are nonvolatile memory circuits (such as a flash memory circuit, an erasable programmable read only memory circuit, or a mask read only memory circuit), volatile memory circuits (such as a static random access memory circuit or a dynamic random access memory circuit), magnetic storage media (such as an analog or digital magnetic tape or a hard disk drive), and optical storage media (such as a CD, a DVD, or a Blu-ray disk).The apparatuses and methods described in this application may be partially or fully implemented by a special purpose computer created by configuring a general purpose computer to perform one or more special functions embodied in computer programs. The above-described functional blocks, components of flowcharts, and other elements serve as software specifications that can be translated into the computer programs by the routine work of a person skilled in the art or programmer.The computer programs include processor-executable instructions stored on at least one non-transitory, tangible computer-readable medium. The computer programs may also contain or rely on stored data. The computer programs may include a basic input / output system (BIOS) that interacts with the special purpose computer hardware, device drivers that interact with particular special purpose computer devices, one or more operating systems, user applications, background services, background applications, etc.The computer programs may include: (i) a descriptive text to be analyzed, such as hypertext markup language (HTML), extensible markup language (XML), or javascript object notation (JSON), (ii) assembler code, (iii) object code generated from the source code by a compiler, (iv) source code for execution by an interpreter, (v) source code for compilation and execution by a just-in-time compiler, etc. As examples only, source code may be written using syntax of languages that include C, C++, C#, ObjectiveC-, Swift, Haskell, Go, SQL, R, Lisp, Java® Fortran, Perl, Pascal, Curl, OCaml, Javascript® HTML5 (Hypertext Markup Language 5th Revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Flash® Visual Basic® Lua, MATLAB, SIMULINK, and Python® may be used.LegendIn the drawing figures, N represents No and Y represents Yes.
Claims
A system for controlling autonomous driving of a vehicle (10), comprising: at least one satellite configured to obtain one or more aerial images of an intersection (202), the intersection (202) comprising at least one unlabeled lane; an infrastructure-mounted camera configured to obtain one or more infrastructure camera images of the intersection (202), the infrastructure-mounted camera oriented with a viewing angle toward the intersection (202); a front vehicle camera (22) configured to capture images from a front field of view of a vehicle (10); a wireless interface of the vehicle (10), the wireless interface configured to wirelessly receive the one or more aerial images and the one or more infrastructure camera images; and a vehicle control module (20) configured to: access the one or more aerial images and the one or more infrastructure camera images; obtain one or more vehicle camera images (404) from the front vehicle camera (22); provide the one or more aerial images, the one or more infrastructure camera images, and the one or more vehicle camera images (404) to at least one machine learning model; generate a virtual lane prediction according to an output of the at least one machine learning model, wherein the virtual lane prediction indicates one or more unlabeled lanes present in the intersection (202); and automatically control steering, acceleration, and braking of the vehicle (10) in the intersection (202) according to the virtual lane prediction; wherein providing the one or more aerial images comprises providing the one or more aerial images to a machine learning model for a transformer encoder (308); wherein the vehicle control module (20) is configured to: provide an output of the machine learning model to a transformer encoder (308) as an input to a machine learning model for a map decoder (312); provide map query data and map loss data as inputs to the machine learning model for a map decoder (312); and generate the virtual lane prediction based at least in part on an output of the machine learning model for a map decoder (312); wherein the vehicle control module (20) is configured to: provide an output of the machine learning model to a transformer encoder (308) as an input to a machine learning model to an auxiliary actuator decoder (310); provide actuator query data and actuator loss data as inputs to the machine learning model to an auxiliary actuator decoder (310); and generate the virtual lane prediction based at least in part on an output of the machine learning model to an auxiliary actuator decoder.The autonomous driving control system of a vehicle (10) of claim 1, wherein the vehicle control module (20) is configured to associate a subset of actuator features of the machine learning model for an auxiliary actuator decoder (310) with a list of map feature inputs of the machine learning model for a map decoder (312).The autonomous driving control system of a vehicle (10) of claim 1, wherein the vehicle control module (20) is configured to: provide the one or more aerial images to a convolutional neural network; and provide an output of the convolutional neural network to a machine learning model for a transformer encoder (308).The autonomous driving control system of a vehicle (10) of claim 1, wherein automatically controlling steering, acceleration, and braking of the vehicle (10) comprises controlling operation of the vehicle (10) via map-free autonomous driving.A method of controlling autonomous driving of a vehicle (10), the method comprising: obtaining, via a wireless interface of a vehicle (10), one or more aerial images of an intersection (202), wherein the intersection (202) comprises at least one unlabeled lane and the one or more aerial images are captured using at least one satellite; obtaining, via the wireless interface of the vehicle (10), one or more infrastructure camera images of the intersection (202) from an infrastructure-mounted camera, wherein the infrastructure-mounted camera is oriented with a viewing angle towards the intersection (202); receiving one or more vehicle camera images (404) from a front vehicle camera (22) of the vehicle (10); providing the one or more aerial images, the one or more infrastructure camera images, and the one or more vehicle camera images (404) for at least one machine learning model; wherein providing the one or more aerial images comprises providing the one or more aerial images for a machine learning model to a transformer encoder (308); generating a virtual lane prediction according to an output of the at least one machine learning model, wherein the virtual lane prediction indicates one or more unlabeled lanes present in the intersection (202); providing an output of the machine learning model to a transformer encoder (308) as an input to a machine learning model to a map decoder (312); providing map query data and map loss data as machine learning model inputs to a map decoder (312); generating a virtual lane prediction based at least in part on an output of the machine learning model to a map decoder; providing an output of the machine learning model to a transformer encoder (308) as an input to a sub-actuator decoder (310); providing actuator query data and actuator loss data as machine learning model inputs to a sub-decoder; generating a virtual lane prediction based at least in part on an output of the machine learning model to a sub-decoder; and automatically controlling steering, acceleration and braking of the vehicle (10) at the intersection (202) according to the prediction of virtual lanes.A system for controlling autonomous driving of a vehicle (10), comprising: a front vehicle camera (22) configured to capture images from a front field of view of a vehicle (10); a wireless interface configured to wirelessly receive one or more aerial images of an intersection (202), the intersection comprising at least one unlabeled lane and the one or more aerial images being captured using at least one satellite; and wirelessly receive one or more infrastructure camera images of the intersection (202) from an infrastructure-mounted camera, the infrastructure-mounted camera oriented with a viewing angle toward the intersection (202); and a vehicle control module (20) configured to: access the one or more aerial images and the one or more infrastructure camera images; one or more vehicle camera images (404) are obtained from the front vehicle camera (22); the one or more aerial images providing the one or more infrastructure camera images and the one or more vehicle camera images (404) to at least one machine learning model; generate a virtual lane prediction according to an output of the at least one machine learning model, wherein the virtual lane prediction indicates one or more unlabeled lanes present in the intersection (202); and automatically control steering, acceleration, and braking of the vehicle (10) in the intersection (202) according to the virtual lane prediction; wherein providing the one or more aerial images comprises providing the one or more aerial images for a machine learning model to a transformer encoder (308); wherein the vehicle control module (20) is configured to: provide an output of the machine learning model to a transformer encoder (308) as input to a machine learning model to a map decoder (312); provide map query data and map loss data as inputs to the machine learning model to a map decoder (312); and generate the virtual lane prediction based at least in part on an output of the machine learning model to a map decoder (312); wherein the vehicle control module (20) is configured to: provide an output of the machine learning model to a transformer encoder (308) as input to a machine learning model to an auxiliary actuator decoder (310); Actor query data and actor loss data as inputs to the machine learning model for an auxiliary actor decoder (310); and generates the prediction of virtual lanes based at least in part on an output of the machine learning model for an auxiliary actor decoder.
Citation Information
Patent Citations
ARTIFICIAL NEURAL NETWORK FOR CLASSIFYING AND LOCALIZING LANE CHARACTERISTICS
DE102018131477A1
Recording and classifying street attributes for map expansion
DE102020130513A1
METHOD FOR OPERATING A VEHICLE, COMPUTER PROGRAM PRODUCT, SYSTEM AND VEHICLE
DE102021102426A1
METHOD AND DEVICE FOR DETERMINING AN ENVIRONMENTAL MODEL OF A MOTOR VEHICLE
DE102021125582A1
Methods for creating a digital map
DE102023002030B3