Motion prediction in autonomous vehicles using machine learning models trained with cyclic consistency loss
By introducing cyclic consistency losses and forward and reverse motion prediction during the training process, the machine learning model is improved, and the problem of insufficient accuracy of object motion prediction is solved, and the safety and effectiveness of autonomous vehicles are improved.
Patent Information
- Application Number
- CN202380082315.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-22
- Filing Date
- 2023-09-28
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, when using machine learning models to predict objects, the accuracy is insufficient, resulting in problems with autonomous vehicles in terms of safety and effectiveness.
By introducing cyclic consistency losses during training, combining forward and reverse motion prediction, the training methods of machine learning models are improved to improve prediction accuracy.
Improves the accuracy of machine learning models when predicting object movement, thus enabling autonomous vehicles to operate safer and more efficiently.
Smart Images

Figure CN120303670A_ABST
Abstract
Description
BRIEF DESCRIPTION OF THE DRAWINGS
[0001] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system;
[0002] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0003] Figure 3 is Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems of;
[0004] Figure 4 is a diagram of certain components of an autonomous system;
[0005] Figure 4B is a diagram of an implementation of a neural network;
[0006] Figure 4C and Figure 4D is a diagram illustrating an example operation of a CNN;
[0007] Figure 5A and 5B is an illustrative visualization of training data for training a machine learning model using a cycle consistency loss that includes both forward-time data and backward-time data;
[0008] Figure 6 is a block diagram showing an example environment in which a system determines to train a machine learning model using a cycle consistency loss for object motion prediction;
[0009] Figure 7 is a visualization of an example training process for training a machine learning model using a cycle consistency loss for object motion prediction;
[0010] Figure 8 is a flowchart of an example process for training a machine learning model using a cycle consistency loss for object motion prediction. DETAILED DESCRIPTION
[0011] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the embodiments described herein may be practiced without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0012] In the drawings, for ease of description, a specific arrangement or order of illustrative elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of the illustrative elements in the drawings is not intended to imply a required processing order or sequence, or a separation of processes. Additionally, unless explicitly described, the inclusion of illustrative elements in the drawings is not intended to imply that such elements are required in all embodiments, nor is it intended to imply that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.
[0013] Furthermore, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows, etc.) are used to illustrate connections, relationships, or associations between or among two or more other illustrative elements. The absence of any such connecting element is not intended to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of a signal, data, or instruction (e.g., “software instruction”), those skilled in the art should understand that such an element can represent one or more signal paths (e.g., a bus) that may be required to affect the communication.
[0014] Although terms such as “first,” “second,” and / or “third,” etc. are used to describe various elements, these elements should not be limited by these terms. The terms “first,” “second,” and / or “third” are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact can be referred to as a second contact, and similarly, a second contact can be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.
[0015] The terms used in the description of the various embodiments described herein are included for the purpose of describing particular embodiments only and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and may be used interchangeably with "one or more than one" or "at least one", unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprise", "include", "have", and / or "with" are used in this specification, it specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0016] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receipt, transmission, conveyance, and / or provision of information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is inherently wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted between the first unit and the second unit is modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message can refer to a network packet (e.g., a data packet, etc.) that includes data.
[0017] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "upon", "in response to determining that", and / or "in response to detecting", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [stated condition or event] is detected" is optionally interpreted to mean "when determining...", "in response to determining that", or "when [stated condition or event] is detected" and / or "in response to detecting [stated condition or event]", etc. Further, as used herein, the terms "has", "having", or "owns", etc. are intended to be open-ended terms. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on".
[0018] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to one of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0019] General Overview
[0020] Generally, aspects of the present disclosure relate to improving object motion prediction by implementing a machine learning model trained using a cycle consistency loss, which is applied in fields such as autonomous vehicles. As described herein, object motion prediction generally describes a mechanism for predicting the future motion (e.g., trajectory) of an object based on the past observed motion of the object or other objects. For example, object motion prediction can be applied in an autonomous vehicle to enable the vehicle to predict how other objects in or near the roadway will move, so that the autonomous vehicle can respond appropriately (e.g., avoid collisions or other safety risks, navigate the roadway legally, etc.). One mechanism for object motion prediction is to implement a machine learning model that can be trained on a historical data set that reflects the observed motion of an object up to a given point in time and the actual motion of the object after that point in time. The historical data set can be used to train the machine learning model to predict the actual motion from the past observed motion, such that the model can be deployed to an autonomous vehicle to predict (unknown) object motion based on the observed motion. In practice, this simplified approach can be inaccurate, which can lead to safety and effectiveness issues in autonomous vehicles. The present disclosure relates to an improved machine learning model for object motion prediction, where the training is modified to include a cycle consistency loss. As described herein, the loss can be calculated based at least in part on an inverse or reverse motion prediction, that is, given the observed motion and the predicted future motion, how likely it is that the predicted future motion will result in the prediction of the observed motion if the predicted future motion is passed through the model as if it were historical data. The model can embody the intuitive understanding that many or most object movements are reversible, such that if the predicted forward motion cannot be reversed, the predicted forward motion is less likely to be an accurate prediction. Thus, training based on inverse or reverse motion prediction can improve the ability of the machine learning model to accurately predict future motion based on the observed motion.
[0021] As will be understood by those skilled in the art in light of the present disclosure, the embodiments disclosed herein improve the ability of a computing system to predict object motion based on historical object motion. In particular, embodiments of the present disclosure train a machine learning model using a cycle consistency loss factor, at least in part based on reverse or backward motion prediction, to provide improved accuracy in predicting object motion. This improved accuracy in turn enables autonomous systems such as autonomous vehicles that rely on object motion prediction to operate more safely and efficiently. Additionally, embodiments of the present disclosure address technical problems inherent within a computing system; specifically, the difficulty of programmatically predicting object motion and the difficulty of safely implementing various automated tasks without accurate object motion prediction. These technical problems are addressed by the various techniques described herein, which include machine learning model architectures that include a cycle consistency loss factor during training. Accordingly, the present disclosure generally represents an improvement to machine learning computing systems and computing systems.
[0022] Now referring to Figure 1 , illustrative example environment 100 is shown, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, area 108, vehicle-to-infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118. Vehicles 102a - 102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 are interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establishing a connection for communication, etc.). In some embodiments, objects 104a - 104n are interconnected with at least one of vehicles 102a - 102n, vehicle-to-infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired and wireless connections.
[0023] Vehicles 102a - 102n (individually referred to as vehicle 102 and collectively as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicle 102 is configured to communicate with V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicle 102 is the same as or similar to vehicle 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, as described herein, vehicle 102 travels along corresponding routes 106a - 106n (individually referred to as route 106 and collectively as routes 106). In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same as or similar to autonomous system 202).
[0024] Objects 104a - 104n (individually referred to as object 104 and collectively as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.), etc. Each object 104 (e.g., located at a fixed location and over a period of time) is stationary or (e.g., having a speed and associated with at least one trajectory) moving. In some embodiments, object 104 is associated with a corresponding location in area 108.
[0025] Routes 106a - 106n (each referred to individually as Route 106 and collectively as Routes 106) are each associated with (e.g., define) a series of actions (also referred to as a trajectory) along which an AV can be navigated. Each Route 106 begins at an initial state (e.g., a state corresponding to a first spatio - temporal location and / or speed, etc.) and ends at a final goal state (e.g., a state corresponding to a second spatio - temporal location different from the first) or a target zone (e.g., a subspace of acceptable states (e.g., a termination state)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, Route 106 includes multiple acceptable sequences of states (e.g., multiple sequences of spatio - temporal locations) that are associated with (e.g., define) multiple trajectories. In an example, Route 106 includes only high - level actions or imprecise state locations, such as a series of connecting roads indicating a direction change at a roadway intersection, etc. Additionally or alternatively, Route 106 can include more precise actions or states, such as, for example, a specific target lane or precise location within a lane area and a target rate at those locations. In an example, Route 106 includes multiple precise state sequences along at least one high - level action with a finite look - ahead horizon to reach an intermediate goal, where the combination of successive iterations of the finite - horizon state sequences cumulatively corresponds to multiple trajectories that together form a high - level route terminating at the final goal state or zone.
[0026] Region 108 includes a physical area (e.g., a geographic area) in which vehicle 102 can be navigated. In an example, Region 108 includes at least one state (e.g., a country, a province, an individual state among multiple states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, Region 108 includes at least one named arterial road (referred to herein as a “road”), such as a highway, an interstate highway, a parkway, an urban street, etc. Additionally or alternatively, in some examples, Region 108 includes at least one unnamed road, such as a lane, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., the portion of the road through which vehicle 102 can pass). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0027] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with a vehicle 102 and / or a V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, a remote AV system 114, a queue management system 116, and / or the V2I system 118 via a network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), lane markings, streetlights, a parking meter, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the queue management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.
[0028] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a Long-Term Evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.
[0029] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.
[0030] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ride-sharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).
[0031] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).
[0032] Provided Figure 1 The number and arrangement of the illustrated elements are by way of example. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements in a different arrangement. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.
[0033] Now referring to Figure 2 , the vehicle 200 (which may be the same as or similar to Figure 1 the vehicle 102) includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208, or is associated with these systems. In some embodiments, the vehicle 200 is the same as the vehicle 102 (see Figure 1)Same or similar. In some embodiments, the autonomous system 202 is configured to endow the vehicle 200 with autonomous driving capabilities (e.g., implement at least one of the following driving automation or maneuver-based functions, features, and / or devices, etc., where the at least one driving automation or maneuver-based function, feature, and / or device enables the vehicle 200 to operate partially or fully without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as level 3 ADS-operated vehicles, etc.), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functionality required to operate the vehicle 200 in road traffic and continuously perform part or all of the dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that incorporates driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.
[0034] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance traveled by the vehicle 200, etc.). In some embodiments, the autonomous system 202 uses one or more devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more devices of the autonomous system 202 may be used by one or more systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle computing 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.
[0035] The camera 202a includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus 302 that is the same or similar to Figure 3 the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) for capturing images including physical objects (e.g., cars, buses, curbs, and / or people, etc.). In some embodiments, the camera 202a generates camera data as an output. In some examples, the camera 202a generates camera data including image data associated with the image. In this example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle for capturing images for the purpose of stereoscopic vision (stereo vision). In some examples, the camera 202a includes generating image data and transmitting the image data to the autonomous vehicle computing 202f and / or a queue management system (e.g., to Figure 1A plurality of cameras of the same or similar queue management system as the queue management system 116. In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the fields of view of at least two of the plurality of cameras based on image data from at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses optimized for sensing objects at one or more distances relative to the camera 202a.
[0036] 35 In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with one or more images. In some examples, the camera 202a generates TLD (Traffic Light Detection) data associated with one or more images including formats (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data is different from the other systems incorporating cameras described herein in that: the camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens with a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.
[0037] The light detection and ranging (LiDAR) sensor 202b includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., with Figure 3at least one device configured to communicate via a bus (e.g., a bus identical or similar to bus 302). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser transmitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b also includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.
[0038] The radio detection and ranging (Radar) sensor 202c includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3 bus 302). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter a physical object and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of the physical object in the field of view of the Radar sensor 202c.
[0039] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus identical or similar to the Figure 3 bus 302). The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein may receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.
[0040] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include a device identical or similar to the Figure 3 communication interface 314. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).
[0041] The autonomous vehicle computing 202f includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the safety controller 202g, and / or the DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as a client device, a mobile device (e.g., a cellular phone and / or a tablet, etc.), and / or a server (e.g., a computing device including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is identical or similar to the autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system identical or similar to the Figure 1 remote AV system 114), a queue management system (e.g., a queue management system identical or similar to the Figure 1 queue management system 116), a V2I device (e.g., a V2I device identical or similar to the Figure 1 V2I device 110), and / or a V2I system (e.g., a V2I system identical or similar to the Figure 1communicate with a V2I system 118 that is the same as or similar to the V2I system).
[0042] The safety controller 202g includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the autonomous vehicle computing 202f, and / or the DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that take precedence over (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.
[0043] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).
[0044] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle movements such as starting to move forward, stopping moving forward, starting to move backward, stopping moving backward, accelerating in a certain direction, decelerating in a certain direction, etc., or causes the vehicle 200 to perform lateral vehicle movements such as making a left turn and / or making a right turn. In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.
[0045] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates two front wheels and / or two rear wheels of the vehicle 200 left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activities required to adjust the y-axis component of the vehicle's movement.
[0046] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.
[0047] In some embodiments, the vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring the nature of the state or condition of the vehicle 200. In some examples, the vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel brake pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor, etc. Although Figure 2 the braking system 208 is shown to be proximal to the vehicle 200, the braking system 208 can be located anywhere in the vehicle 200.
[0048] Now refer to Figure 3 , a schematic diagram of the exemplary device 300. As illustrated, the device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, the device 300 corresponds to: at least one device of the vehicle 102 (e.g., at least one device of the system of the vehicle 102); and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112). In some embodiments, one or more devices of the vehicle 102 (e.g., one or more devices of the system of the vehicle 102), and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112) include at least one device 300 and / or at least one component of the device 300. As Figure 3As shown, device 300 includes bus 302, processor 304, memory 306, storage component 308, input interface 310, output interface 312, and communication interface 314.
[0049] Bus 302 includes components that permit communication between the components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, and / or optical memory, etc.).
[0050] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, RAM, PROM, EPROM, FLASH-EPROM, NV-RAM, and / or another type of computer-readable medium, and corresponding drives.
[0051] Input interface 310 includes components that permit device 300 to receive information such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, input interface 310 includes sensors for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). Output interface 312 includes components for providing output information from device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0052] In some embodiments, communication interface 314 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter, etc.) that permit device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from another device and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, an interface and / or a cellular network interface, etc.
[0053] In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on software instructions stored in a computer-readable medium such as memory 306 and / or storage component 308 by processor 304. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.
[0054] In some embodiments, software instructions are read into memory 306 and / or storage component 308 from another computer-readable medium or from another device via communication interface 314. When the software instructions stored in memory 306 and / or storage component 308 are executed, they cause processor 304 to perform one or more processes described herein. Additionally or alternatively, instead of software instructions or in combination with software instructions, hardwired circuitry is used to perform one or more processes described herein. Thus, unless otherwise explicitly stated, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.
[0055] Memory 306 and / or storage component 308 includes a data storage portion or at least one data structure (e.g., a database, etc.). Device 300 is capable of receiving information from the data storage portion or at least one data structure in memory 306 or storage component 308, storing information in the data storage portion or at least one data structure, communicating information to the data storage portion or at least one data structure, or searching for information stored in the data storage portion or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0056] In some embodiments, device 300 is configured to execute software instructions stored in memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to device 300). As used herein, the term "module" refers to at least one instruction stored in memory 306 and / or the memory of another device, which when executed by processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to device 300), causes device 300 (e.g., at least one component of device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.
[0057] Provide Figure 3 The number and arrangement of the illustrated components are provided as examples. In some embodiments, compared with Figure 3Compared with the illustrated components, the apparatus 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of the apparatus 300 (e.g., one or more than one component) may perform one or more than one function described as being performed by another component or another set of components of the apparatus 300.
[0058] Now refer to Figure 4 A, which illustrates an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack"). As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of the vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more than one independent system (e.g., one or more than one system identical or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more than one independent system located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., via a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that, in some embodiments, the autonomous vehicle computing 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system identical or similar to the remote AV system 114, a queue management system identical or similar to the queue management system 116, and / or a V2I system identical or similar to the V2I system 118, etc.).
[0059] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical objects by the perception system 402, the perception system 402 transmits data associated with the classification of the physical objects to the planning system 404.
[0060] In some embodiments, the planning system 404 receives data associated with a destination, and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical objects described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 101 in road traffic. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with the updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.
[0061] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of roadways (such as traffic speed, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs or various types of other driving signal lights, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.
[0062] In another example, the positioning system 406 receives Global Navigation Satellite System (GNSS) data generated by a Global Positioning System (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.
[0063] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes activities required to adjust the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes activities required to adjust the x-axis component of the vehicle motion. In an example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.
[0064] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model individually or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.). The following are examples regarding Figures 4B to 4D the implementation of machine learning models.
[0065] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component for storing operation-related data and / or software and using the autonomous vehicle computing 400 of at least one system (e.g., associated with Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one region. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., nations), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, backroads, and / or off-road paths, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.
[0066] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116), and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.
[0067] Now refer to Figure 4B , a diagram illustrating the implementation of a machine learning model. More specifically, a diagram illustrating the implementation of a convolutional neural network (CNN) 420. For illustrative purposes, the following description of the CNN 420 will be with respect to implementing the CNN 420 by the perception system 402. However, it will be understood that in some examples, the CNN 420 (e.g., one or more components of the CNN 420) is implemented by other systems different from or in addition to the perception system 402, such as the planning system 404, the positioning system 406, and / or the control system 408, etc. Although the CNN 420 includes certain features as described herein, these features are provided for illustrative purposes and are not intended to limit the present disclosure.
[0068] The CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, the CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, the subsampling layer 428 and / or other subsampling layers have dimensions smaller than the dimensions of the upstream system (i.e., the amount of nodes). By means of the subsampling layer 428 having dimensions smaller than the dimensions of the upstream layer, the CNN 420 combines the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for the CNN 420 to perform downstream convolutional operations. Additionally or alternatively, by means of the subsampling layer 428 being associated with at least one subsampling function (e.g., being configured to perform at least one subsampling function) (as described below with respect to Figure 4C and Figure 4D ), the CNN 420 combines the amount of data associated with the initial input.
[0069] Based on the perception system 402 providing corresponding inputs and / or outputs respectively associated with the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426 to generate corresponding outputs, the perception system 402 performs convolutional operations. In some examples, based on the perception system 402 providing data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426, the perception system 402 implements the CNN 420. In such examples, based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle identical or similar to the vehicle 102, a remote AV system identical or similar to the remote AV system 114, a queue management system identical or similar to the queue management system 116, and / or a V2I system identical or similar to the V2I system 118, etc.), the perception system 402 provides the data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426. The following is a detailed description of Figure 4C including convolutional operations.
[0070] In some embodiments, the perception system 402 provides data associated with an input (referred to as an initial input) to the first convolutional layer 422, and the perception system 402 uses the first convolutional layer 422 to generate data associated with an output. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as an input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as an input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as an upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as downstream layers.
[0071] In some embodiments, before the perception system 402 provides an input to the CNN 420, the perception system 402 processes data associated with the input provided to the CNN 420. For example, based on the perception system 402 normalizing sensor data (such as, for example, image data, LiDAR data, and / or Radar data, etc.), the perception system 402 processes data associated with the input provided to the CNN 420.
[0072] In some embodiments, based on the perception system 402 performing convolution operations associated with each convolutional layer, the CNN 420 generates an output. In some examples, based on the perception system 402 performing convolution operations associated with each convolutional layer and the initial input, the CNN 420 generates an output. In some embodiments, the perception system 402 generates an output and provides the output to the fully connected layer 430. In some examples, the perception system 402 provides the output of the convolutional layer 426 to the fully connected layer 430, where the fully connected layer 430 includes data associated with a plurality of eigenvalues referred to as F1, F2,..., FN. In this example, the output of the convolutional layer 426 includes data associated with a plurality of output eigenvalues representing predictions.
[0073] In some embodiments, based on the perception system 402 identifying an eigenvalue associated with the highest likelihood of being the correct prediction among a plurality of predictions, the perception system 402 identifies a prediction from among the plurality of predictions. For example, in a case where the fully connected layer 430 includes eigenvalues F1, F2, ..., FN and F1 is the largest eigenvalue, the perception system 402 identifies the prediction associated with F1 as the correct prediction among the plurality of predictions. In some embodiments, the perception system 402 trains the CNN 420 to generate predictions. In some examples, based on the perception system 402 providing training data associated with a prediction to the CNN 420, the perception system 402 trains the CNN 420 to generate predictions.
[0074] Now refer Figure 4C and Figure 4D , a diagram illustrating an example operation of the CNN 440 using the perception system 402. In some embodiments, the CNN 440 (e.g., one or more components of the CNN 440) is the same as or similar to the CNN 420 (e.g., one or more components of the CNN 420) (see Figure 4B ).
[0075] In step 450, the perception system 402 provides data associated with an image as an input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with an image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image, which is represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.
[0076] In step 455, the CNN 440 performs a first convolution function. For example, based on the CNN 440 providing the values representing the image as an input to one or more neurons (not explicitly illustrated) included in the first convolutional layer 442, the CNN 440 performs the first convolution function. In this example, the values representing the image may correspond to the values of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values corresponding in size to the values provided as an input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolutional layers, the filters associated with the neurons may be configured to successively identify more complex patterns (e.g., arcs and / or objects, etc.).
[0077] In some embodiments, based on the CNN 440, the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 are multiplied by the values of the filters corresponding to each neuron among the same one or more neurons, and the CNN 440 performs a first convolutional function. For example, the CNN 440 may multiply the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 by the values of the filters corresponding to each neuron among the same one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolutional layer 442 is referred to as the convolutional output. In some embodiments, when each neuron has the same filter, the convolutional output is referred to as a feature map.
[0078] In some embodiments, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of a downstream layer. For clarity, an upstream layer may be a layer that transmits data to a different layer (referred to as a downstream layer). For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of a subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the first subsampling layer 444. In such an example, the CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregated set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.
[0079] In step 460, the CNN 440 performs a first subsampling function. For example, based on the CNN 440 providing the values output by the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444, the CNN 440 may perform a first subsampling function. In some embodiments, the CNN 440 performs the first subsampling function based on an aggregation function. In an example, based on the CNN 440 determining the maximum input among the values provided to a given neuron (referred to as the max pooling function), the CNN 440 performs the first subsampling function. In another example, based on the CNN 440 determining the average input among the values provided to a given neuron (referred to as the average pooling function), the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to each neuron of the first subsampling layer 444, the CNN 440 generates an output, which is sometimes referred to as the subsampled convolutional output.
[0080] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function as described above. In some embodiments, based on the values output by the first subsampling layer 444 being provided as inputs to one or more neurons (not explicitly illustrated) included in the second convolution layer 446, CNN 440 performs the second convolution function. In some embodiments, as described above, each neuron of the second convolution layer 446 is associated with a filter. As described above, the (one or more) filters associated with the second convolution layer 446 may be configured to identify more complex patterns compared to the filters associated with the first convolution layer 442.
[0081] In some embodiments, based on CNN 440 multiplying the values provided as inputs to each of the one or more neurons included in the second convolution layer 446 by the values of the filters corresponding to each of the one or more neurons, CNN 440 performs the second convolution function. For example, CNN 440 may multiply the values provided as inputs to each of the one or more neurons included in the second convolution layer 446 by the values of the filters corresponding to each of the one or more neurons to generate a single value or an array of values as output.
[0082] In some embodiments, CNN 440 provides the output of each neuron of the second convolution layer 446 to the neurons of a downstream layer. For example, CNN 440 may provide the output of each neuron of the first convolution layer 442 to the corresponding neurons of the subsampling layer. In an example, CNN 440 provides the output of each neuron of the first convolution layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the second subsampling layer 448. In such an example, CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregated set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.
[0083] At step 470, the CNN 440 performs a second subsampling function. For example, based on the values output by the second convolutional layer 446 being provided to the respective neurons of the second subsampling layer 448 by the CNN 440, the CNN 440 may perform the second subsampling function. In some embodiments, based on the CNN 440 using an aggregation function, the CNN 440 performs the second subsampling function. In an example, as described above, based on the CNN 440 determining the maximum input or average input among the values provided to a given neuron, the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to the respective neurons of the second subsampling layer 448, the CNN 440 generates an output.
[0084] At step 475, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449. For example, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449 such that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as classification). The prediction may include an indication that the object(s) included in the image provided as input to the CNN 440 includes an object and / or a collection of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to different systems described herein.
[0085] Motion Prediction Using a Machine Learning Model Trained with a Cycle Consistency Loss
[0086] As described above, it may be difficult to train a machine learning model such as the CNN described above to accurately predict future object motion from past observations of motion.
[0087] Reference Figures 5A to 8 , embodiments will be described in which motion prediction using a machine learning model trained with a cycle consistency loss is improved. Specifically, embodiments will be described in which a machine learning model is trained using both forward-time prediction of future motion from past observations of motion and backward (or inverse) time prediction of observed motion from (predicted or known) future motion. By incorporating backward time prediction during training, the accuracy of the resulting model can be improved relative to a model trained only on forward-time prediction. Such a model can then be deployed to enable more accurate motion prediction. For example, such a model can be loaded into an autonomous vehicle (such as Figure 1in a vehicle 102, etc., so that the vehicle can predict future object motion from the past observed motion of an object (e.g., as observed using sensors of the vehicle). In one embodiment, reverse time prediction is implemented during the training of a machine learning model and not during inference. Thus, reverse time prediction can be used to improve the accuracy of the model during inference without otherwise changing how the model is used at deployment.
[0088] In Figure 5A and Figure 5B the visualization of further shows the concept of reverse (or inverted) time prediction, where Figure 5A and Figure 5B visualize both forward time data and reverse time data that can be used as training data when training a machine learning model for object motion prediction. Specifically, Figure 5A depicts forward time data that can be used during a forward pass corresponding to forward time prediction. Figure 5B depicts reverse (or inverted) time data that can be used during a reverse pass corresponding to reverse time prediction.
[0089] In Figure 5A the visualization, multiple objects 502A to 502C are traveling on a set of oriented driving lanes. For example, Figure 5A the visualization can be a top view (or bird's-eye view) of a road intersection, where each traffic lane is indicated as an oriented driving lane. Each object 502 can represent a first-party entity that can be controlled, such as Figure 1 the autonomous vehicle 102, a target object whose motion is desired to be predicted, or another object that is also traveling within the oriented driving lanes (the movement of which may affect the movement of other objects), etc. For example, the visualization can be created based on data collected during the operation of a first-party vehicle (such as from sensors on the vehicle). Those skilled in the art will understand that Figure 5A is a simplified diagram, and the actual training data may be significantly more complex. For example, the actual training data can include more oriented driving lanes, more objects, or more complex movement patterns, etc.
[0090] From Figure 5A it can be seen that forward motion prediction generally involves predicting the future motion of an object from past observed motion. In Figure 5AIn [the figure], the past observed motion is indicated by solid lines, while the future predicted motion is indicated by dashed lines. Thus, a machine learning model can be trained such that when the historical motion of an object 502 is input into the model, the future motion of the object is predicted. For ease of training, a training dataset can include ground truth for future motion prediction (i.e., the observed motion of the object 502 after a given time point), which is used to train the model.
[0091] An object prediction model can utilize only forward-time prediction based on data similar to the data Figure 5A shown. However, to improve the accuracy of such a model, an improved model that utilizes both forward-time prediction and backward-time prediction is disclosed, where the backward-time prediction is visualized Figure 5B in [the figure]. Specifically, Figure 5B depicts Figure 5A a visualization reversal such that the directed driving lanes are reversed in direction and such that the future motion represents the input rather than the output of the machine learning model. Thus, in [the figure], the directed driving lanes are reversed relative to Figure 5B and thus correspond to the reversed driving lanes. Although objects 502A to 502C are in the same location in both Figure 5A and Figure 5A and Figure 5B their driving directions are opposite (e.g., such that object 502A is driving right in Figure 5A and driving left in Figure 5B ). It is noted that this reversal does not reflect different observation scenarios but reflects the same scenario as Figure 5A in the case of time-reversed. In such a time reversal, each object moves from a future position to a past position.
[0092] Thus, when used as part of a training dataset, the future motion of each object 502 (which can be the predicted future motion generated during the forward pass of the machine learning model, the observed future motion reflected in the ground truth data, or a combination thereof) represents the input through the backward pass of the model, while the historical motion of the object 502 is the prediction of the backward pass. As described above, this backward pass can act as a check for the prediction of the forward pass, thus ensuring that the predicted future motion of the forward pass is reasonable. More specifically, since many or most movements in a given scenario (e.g., a vehicle in a regulated roadway) are reversible, the backward pass can be used to represent the intuitiveness that irreversible movements (e.g., such that an object cannot drive back along a reverse path) are less likely to be accurate. Thus, training a machine learning model using reversed-time data (e.g., as the backward pass) can improve the accuracy of the model relative to using only non-reversed (forward) time data.
[0093] Notably, the techniques described herein need not be limited to a particular machine learning model. Instead, these techniques can be applied to a wide variety of machine learning models. An example of such a model is the "Graph-Oriented Heatmap Output for future Motion Estimation" or "GOHOME" model as disclosed in Thomas Gilles et al., "GOHOME: Graph-Oriented Heatmap Output for future Motion", (doi:10.1109 / ICRA46639.2022.9812253) (which is incorporated herein by reference) as published in the 2022 International Conference on Robotics and Automation (ICRA) (pages 9107-9114). Another example of such a model is the "Autobots" model as disclosed in Roger Girgis et al., "Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction" (doi:10.48550 / arXiv.2104.00563) (incorporated herein by reference) as published in the 2022 International Conference on Learning Representations. According to the present disclosure, such models can be modified to incorporate a backward pass using reversed time data during training, thereby improving the accuracy of the trained model.
[0094] Reference Figure 6 , an illustrative interaction in environment 600 will be described that is used to implement an improved machine learning model for object motion prediction by training a model using a backward pass corresponding to reversed time data. Environment 600 illustratively includes a remote AV system 604 and a vehicle 606, where the remote AV system 604 can correspond, for example, to Figure 1 the remote AV system 114, and the vehicle 606 can correspond, for example, to Figure 2 the vehicle 200.
[0095] Figure 6The interaction begins at (1), where the remote AV system 604 obtains a training data set from the data store 602, which can correspond to any persistent or substantially persistent data store. The training data set illustratively includes data corresponding to multiple observed object movements (such as in various locations, etc.), which includes both historical object movements before a given time point (e.g., corresponding to the input to the trained model during inference) and observed object movements after the given time point (e.g., corresponding to the ground truth for training the model for such inference). Illustratively, the data can be generated, in whole or in part, based on the operation of sensors within one or more autonomous vehicles. For example, the data can be obtained by processing sensor data to generate multiple training instances (such as the instances visualized above in Figure 5A etc.). In some embodiments, the training data includes reverse-time instances, each reverse-time instance corresponding to the reverse of a forward-time instance. In other embodiments, only forward-time instances are obtained, and the remote AV system 604 is configured to reverse one or more such instances (e.g., by reversing the traffic lanes and swapping the input and output in each instance so that the observed ground truth movement represents the input and the historically observed movement represents the output, and vice versa) to generate reverse-time instances.
[0096] Thereafter, at (3), the remote AV system 604 trains a machine learning model for object detection based on the training data. Specifically, the model can be trained using a combination of a forward pass corresponding to processing forward-time data and a backward pass corresponding to processing reverse-time data. According to the machine learning algorithm, the error in the predictions of each pass can be captured as a loss value, which is used (e.g., via backpropagation) to modify the weights of the model for subsequent training steps. Since the error in the predictions of the backward pass indicates a lack of consistency between the forward pass and the backward pass, and since the forward pass and the backward pass form a training loop, the loss value of the backward pass can be referred to as the cycle-consistency loss. According to the present disclosure, it can be seen that modifying the weights of the model using the cycle-consistency loss improves the accuracy of the resulting trained model relative to a model trained using only the forward loss. Further discussion related to the training of a machine learning model using the cycle-consistency loss is provided below with reference to Figure 7 and Figure 8 to provide further discussion related to the training of a machine learning model using the cycle-consistency loss.
[0097] At (4), the remote AV system 604 can then send the trained object motion prediction machine learning model to the vehicle 606. Since one or more models have already been trained using the cycle consistency loss, it can be expected that the models will have improved accuracy relative to models trained without the cycle consistency loss. Thus, the vehicle 606 can utilize the model to subsequently predict object motion during the operation of the vehicle 606, thereby providing a safer and more efficient operation.
[0098] The interaction of can be illustratively repeated Figure 6 . For example, during the operation of the vehicle 606, additional sensor data can be collected and used to generate training data stored in the data store 602, which can then be used during Figure 6 subsequent iterations of the interaction of. Thus, Figure 6 the interaction of can be used to iteratively improve the performance of the machine learning model used to interpret sensor data, which in turn can improve various processes that rely on such a model, such as the operation of an autonomous vehicle, etc.
[0099] Figure 7 illustrates an exemplary training sequence for training a machine learning model using the cycle consistency loss. This sequence can be implemented, for example, by the Figure 6 remote AV system 604 during the training of the object motion detection machine learning model.
[0100] Illustratively, the sequence can begin with the Figure 7 forward prediction pass shown at the bottom of. During the forward prediction pass, the forward time training set instance 610 includes, for example, historical motion data and ground truth motion data of a target entity (also referred to as a target agent) moving relative to a lane map (e.g., traffic lanes on a roadway). The forward time training set instance 610 can also include Figure 7 historical motion data of other entities referred to as "background agents" in. For example, background agents can include other vehicles whose motion affects the motion of the target entity. Although shown in Figure 7 for illustrative purposes, in some cases, the forward time training set instance 610 can omit the background agent data. Additionally, although a single target agent and a single background agent are shown in Figure 7 , the instance 610 can include multiple target agents or multiple background agents.
[0101] As Figure 7As shown, the forward prediction pass may include passing the forward-time training set instance 610 through the prediction model 612, which, as described above, may correspond to any kind of machine learning model. For example, the prediction model 612 may be a CNN as described above, configured to generate a predicted motion of the target agent in instance 610. As a result of passing instance 610 through model 612, the model 612 produces a predicted future motion 614 of the target agent, which, for example, represents a predicted trajectory of the agent during a period of time after what is reflected by the agent's historical motion in instance 610. The predicted future motion 614 may then be compared with the ground truth data 618 (i.e., the observed actual motion of the agent during a period of time after what is reflected by the agent's historical motion in instance 610) to calculate the forward loss 616. According to the training algorithm for the machine learning model, the forward loss 616 may be used to modify the weights of model 612 such that subsequent predictions have improved accuracy.
[0102] In addition to the forward pass, Figure 7 the sequence also includes a backward prediction pass, whereby the reversed (or backward) time training set instance 622 is passed through the prediction model. The reversed time training set instance 622 is similar to the forward-time training set instance 610, but reversed. Specifically, the reversed time training set instance 622 includes a lane map that is backward relative to the forward-time training set instance 610. Further, the input to model 612 during the backward prediction pass is not the historical motion as in instance 610, but the backward future path. In one embodiment, the future path is the predicted future path 614 (e.g., the output of model 612 from a previous forward prediction pass). In another embodiment, the future path is the ground truth future path 618.
[0103] In yet another embodiment, as Figure 7 shown, the future path represented (in a backward manner) in the reversed time training set instance 622 is a mixture of the predicted future path 614 and the ground truth future path 618. For example, the ground truth mixing 620 stage may select one of the predicted future path 614 from a past forward prediction and the corresponding ground truth future path 618 to be backward in the reversed time training set instance 622. Illustratively, the ground truth mixing stage 620 may randomly select between the two options according to a predefined weighting, which may be adjusted during the training of model 612. In one embodiment, the predefined weighting is equal, such that there is an equal chance of using the predicted future path 614 from a past forward prediction or the corresponding ground truth future path 618 in each backward prediction.
[0104] As another example, the ground truth mixing 620 stage can combine both the predicted future path 614 from the past forward prediction and the corresponding ground truth future path 618 to produce a fused future that is reversed in the reverse-time training set instance 622. Illustratively, an object path can be represented as a series of waypoints in a coordinate system (such as x, y coordinates, etc.). Each waypoint can reflect the location of the object at a given point in time (e.g., each subsequent second relative to a defined zero point). In one embodiment, the ground truth mixing 620 can include generating a fused path that includes corresponding waypoints selected (with a given probability) from either the predicted future path 614 or the ground truth future path 618 as each waypoint in the sequence. For example, if the predicted future path is a set of waypoints {A 1- A n} and the ground truth future path 618 is a set of waypoints {B 1- B n}, the fused future path can be a set of waypoints {C 1- C n}, where each value C m is selected as A m or B m according to a predefined probability (e.g., equal weighting). In another embodiment, the ground truth mixing 620 can include generating a fused path that includes fused waypoints, each fused waypoint being generated from a combination of corresponding waypoints from the predicted future path 614 or the ground truth future path 618. For example, if a given waypoint has an x value and a y value, the fused waypoint can inherit the x value from the corresponding waypoint of the predicted future path 614 and the y value from the corresponding waypoint of the ground truth future path 618. As another example, the fused waypoint can be an intermediate point (e.g., a center point between two corresponding waypoints) between the corresponding waypoints of the predicted future path 614 and the ground truth future path 618. In some instances, using ground truth mixing can further improve the accuracy of the prediction model when training with reverse prediction passes.
[0105] Then, as Figure 7As shown, the reverse-time training set instance 622 is fed through the model 612 to produce a predicted target agent history 626 (i.e., the predicted (reverse) historical path of the object given the reverse future path). Then, this predicted target agent history 626 can be used to compute a cycle consistency loss 628, which, like the forward loss 616, is used to modify the weights of the prediction model 612 such that the model 612 more accurately reflects the training data. For example, the predicted target agent history 626 can be reversed and compared to the corresponding target agent history in the forward-time training set instance 610, where the cycle consistency loss value indicates how closely the predicted target agent history 626 matches the corresponding target agent history.
[0106] Then, Figure 7 the sequence can continue in the presence of additional training data such that, after processing the training data, the prediction model 612 is trained to accurately predict future object trajectories from observed object trajectories.
[0107] Although Figure 7 described as starting with a forward prediction pass, in some instances, the sequence can start with a reverse prediction pass such as by using only the ground truth future path 618 in the initial reverse-time training set instance 622. Then, the sequence can proceed as described above.
[0108] Referring Figure 8 to, an example routine 800 for motion prediction in an autonomous vehicle using a machine learning model trained with a cycle consistency loss will be described. The routine 800 can be implemented, for example, by Figure 6 the remote AV system 604.
[0109] Routine 800 begins at block 802, where the remote AV system 604 obtains a training data set for object motion prediction. According to the above description, the training data set can reflect multiple forward-time training set instances, each instance reflecting the movement of a given set of objects within the environment. For example, each instance can reflect a specific navigable area of an autonomous vehicle (e.g., a roadway, an intersection, etc.), the location of objects (e.g., vehicles, pedestrians, bicycles, etc.) in that area at a given time point (e.g., time t), the historical movement of the objects before that time point (e.g., at t-1 seconds, t-2 seconds, etc.), and the ground truth movement of at least one object after that time point (e.g., at t+1 seconds, t+2 seconds, etc.). The specific area and objects can vary across instances or be repeated in multiple instances. For example, multiple instances can reflect the movement of different objects in the same area, or can reflect the movement of the same object in different areas. In one embodiment, the instances are generated based on sensor data collected from one or more autonomous vehicles in each respective area. For example, sensor data collected from a vehicle's LiDAR, Radar, or camera can be processed to generate the instances.
[0110] Thereafter, at block 804, the remote AV system 604 trains a machine learning model using a cycle consistency loss. As described above, training the machine learning model can include: passing each instance of the training data set through an initial model (e.g., with randomly initialized weights), comparing the prediction of the model at each pass with the corresponding ground truth data of the instance, and adjusting the weights of the model to make the prediction more reflective of the ground truth.
[0111] According to an embodiment of the present disclosure, each training iteration of block 804 can include both a forward prediction pass and a backward prediction pass. Specifically, at sub-block 806, the remote AV system 604 performs a forward prediction training pass to predict future movement from historical data. Illustratively, the remote AV system 604 can pass an instance through the model to predict the forward movement of a target object based on the historical movement of the target object and / or other objects. Then, the weights of the model can be updated based on a forward loss, where the forward loss compares the predicted forward movement with the ground truth forward movement of the instance.
[0112] Thereafter, at sub - block 808, the remote AV system 604 performs a reverse prediction training pass to predict (reverse) historical data from (reverse) future movements. As described above, the future movement can be the predicted future movement of sub - block 806, the ground truth movement of sub - block 806, or a fusion thereof. The historical data can be the data used to predict the future movement of sub - block 806. Illustratively, both the future movement and the historical movement are reversed (e.g., such that the data at t - 1 becomes the data at t + 1, etc.) along with other relevant data (such as a lane map, etc.). Then, the reversed future movement is passed through the model in the reverse prediction pass such that the model generates a predicted historical movement. Then, the predicted historical movement can be compared to the historical movement corresponding to the forward prediction pass (which represents the "ground truth" data for the reverse prediction path). Specifically, a cycle - consistency loss function can quantify how closely the predicted historical movement matches the historical movement of the training dataset instances. Then, the weights of the model can be updated based on the cycle - consistency loss such that the model in subsequent passes more accurately predicts historical movement from future movement.
[0113] As described above, incorporating the reverse prediction training pass, and in particular the reverse prediction training pass that utilizes ground truth mixing to generate the future movement used as the input for such a reverse prediction training pass, can improve the accuracy of the model as compared to training such a model using only the forward prediction path.
[0114] Thereafter, at block 810, the trained model obtained from block 804 can be applied to predict future object movement from the observed object movement. For example, the model can be loaded into an autonomous vehicle such that the computing device of the vehicle can pass the observed (e.g., derived from the vehicle's sensor data) object movement through the trained model to predict future object movement. Thus, applying the model as in routine 800 can provide for safer and more efficient operation of devices (such as autonomous vehicles, etc.) that rely on accurate motion prediction.
[0115] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary according to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than in a limiting sense. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims that issue from this application in the specific form of the issued claims, including any subsequent amendments. Any definition explicitly set forth herein for terms to be included in such claims shall govern the meaning of such terms as used in the claims. Additionally, when the term "further comprises" is used in a previous specification or the appended claims, the recitation following this phrase may be additional steps or entities, or sub-steps / sub-entities of the previously recited steps or entities.
Claims
1. A method, comprising: Obtaining a training data set for object motion prediction, the training data set including movement data associated with an object, the movement data indicating the historical movement of the object before a certain time point and the ground truth movement of the object after the time point; Training a machine learning model, wherein training the machine learning model includes: Performing a forward-time prediction training pass to predict the predicted future motion of the object after the time point from the historical movement of the object; Performing a backward-time prediction training pass to predict the predicted historical motion of the object before the time point from the predicted future motion; and Modifying the weights of the machine learning model based on a comparison between the predicted historical motion and the historical movement indicated in the movement data of the object; And Providing the trained machine learning model to an autonomous vehicle, wherein the autonomous vehicle is configured to use the trained machine learning model to predict future object motion from observed object motion.
2. The method according to claim 1, wherein The object is a target entity moving relative to a lane map.
3. The method according to claim 1, wherein The training data set includes movement data of one or more additional objects, the movement data of the one or more additional objects indicating the positions of the one or more additional objects relative to the object, the historical movement of the one or more additional objects before the time point, and the ground truth movement of the one or more additional objects after the time point.
4. The method according to claim 1, wherein, Performing a backward-time prediction training pass to predict the predicted historical motion of the object before the time point from the predicted future motion includes: reversing the predicted future motion.
5. The method according to claim 1, wherein, The comparison between the predicted historical motion and the historical movement indicated in the movement data of the object is a comparison between the reverse of the predicted historical motion and the historical movement of the object indicated in the movement data of the object.
6. The method according to claim 1, wherein The backward-time prediction training pass also predicts the predicted historical motion of the object based on the ground truth movement of the object.
7. The method according to claim 1, wherein, The training data set further includes movement data of a second object, the movement data of the second object indicating the historical movement of the second object before a second time point and the ground truth movement of the second object after the second time point, and wherein training the machine learning model further includes: Performing a second forward-time prediction training pass to predict the predicted future motion of the second object after the second time point from the historical movement of the second object; Selecting at least one of the predicted future motion of the second object and the ground truth movement of the second object to be used as an input for a subsequent backward-time prediction training pass; Performing the subsequent backward-time prediction training pass to predict the predicted historical motion of the second object before the second time point from the input; and Modifying the weights of the machine learning model based on a comparison between the predicted historical motion of the second object and the historical movement of the second object indicated in the movement data of the second object.
8. The method according to claim 1, wherein, The data in the training dataset is generated, in whole or in part, based on the operation of sensors within one or more autonomous vehicles.
9. The method according to claim 1, wherein, The training dataset includes data representing object movement in a plurality of physical locations.
10. The method according to claim 1, wherein Training the machine learning model further includes: modifying the weights of the machine learning model based on a comparison between the predicted future movement and the ground truth movement.
11. The method according to claim 1, wherein, The comparison between the predicted historical movement and the historical movement indicated in the movement data of the object includes: calculating the result of a cycle consistency loss function.
12. The method according to claim 11, wherein, The cycle consistency loss function quantifies how closely the predicted historical movement of the object matches the historical movement of the object.
13. A system, comprising: a processor configured to execute computer-executable instructions; and a data store that stores the computer-executable instructions, the computer-executable instructions, when executed by the processor, cause the system to: obtain a training dataset for object movement prediction, the training dataset including movement data associated with an object, the movement data indicating the historical movement of the object before a certain time point and the ground truth movement of the object after the time point; train a machine learning model, wherein training the machine learning model includes: performing a forward-time prediction training pass to predict a predicted future movement of the object after the time point from the historical movement of the object, performing a backward-time prediction training pass to predict a predicted historical movement of the object before the time point from the predicted future movement, and modifying the weights of the machine learning model based on a comparison between the predicted historical movement and the historical movement indicated in the movement data of the object; and providing the trained machine learning model to an autonomous vehicle, wherein the autonomous vehicle is configured to use the trained machine learning model to predict future object movement from observed object movement.
14. The system according to claim 13, wherein, The backward-time prediction training pass also predicts the predicted historical movement of the object based on the ground truth movement of the object.
15. The system according to claim 13, wherein The training dataset further includes movement data of a second object, the movement data of the second object indicating the historical movement of the second object before a second time point and the ground truth movement of the second object after the second time point, and wherein training the machine learning model further includes: performing a second forward-time prediction training pass to predict a predicted future movement of the second object after the second time point from the historical movement of the second object; selecting at least one of the predicted future movement of the second object and the ground truth movement of the second object to be used as an input for a subsequent backward-time prediction training pass; performing the subsequent backward-time prediction training pass to predict a predicted historical movement of the second object before the second time point from the input; and Modifying the weights of the machine learning model based on a comparison between the predicted historical motion of the second object and the historical motion of the second object indicated in the movement data of the second object.
16. The system according to claim 13, wherein, Training the machine learning model further includes: modifying the weights of the machine learning model based on a comparison between the predicted future motion and the ground truth motion.
17. The system according to claim 13, wherein, The comparison between the predicted historical motion and the historical motion indicated in the movement data of the object includes: calculating the result of a cycle consistency loss function.
18. One or more non-transitory computer-readable media, comprising computer-executable instructions that, when executed by a computing system including a processor, cause the computing system to: Obtain a training data set for object motion prediction, the training data set including movement data associated with an object, the movement data indicating the historical movement of the object before a certain time point and the ground truth movement of the object after the time point; Train a machine learning model, where training the machine learning model includes: Performing a forward-time prediction training pass to predict a predicted future motion of the object after the time point from the historical motion of the object, Performing a backward-time prediction training pass to predict a predicted historical motion of the object before the time point from the predicted future motion, and Modifying the weights of the machine learning model based on a comparison between the predicted historical motion and the historical motion indicated in the movement data of the object; And Providing the trained machine learning model to an autonomous vehicle, where the autonomous vehicle is configured to utilize the trained machine learning model to predict future object motion from observed object motion.
19. The one or more non-transitory computer-readable media according to claim 18, wherein, The backward-time prediction training pass also predicts the predicted historical motion of the object based on the ground truth motion of the object.
20. The one or more non-transitory computer-readable media according to claim 18, wherein, The comparison between the predicted historical motion and the historical motion indicated in the movement data of the object includes: calculating the result of a cycle consistency loss function.