Point cloud alignment system for generating high definition map for vehicle navigation
By using machine learning models to determine the relative transformation between point clouds, the problem of difficulty in generating accurate high-definition maps in the prior art is solved, and safe navigation and accurate trajectory determination of autonomous vehicles are realized.
Patent Information
- Application Number
- CN202380071592.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-11
- Filing Date
- 2023-08-08
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to generate accurate and reliable high-definition maps for secure navigation of autonomous vehicles.
By using machine learning models, the relative transformations between point clouds are determined to align the target point clouds with the source point clouds, and an accurate high-definition map is generated. The model includes a convolutional layer, a fully connected layer, and a flattened network, capable of weighted subsampling and feature extraction.
It realizes the generation of accurate and reliable high-definition maps, improves the accuracy of trajectory determination of autonomous vehicles in physical space, and enhances the ability to avoid collisions and comply with traffic rules.
Smart Images

Figure CN120019407A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Patent Application No. 17 / 885,949, filed on August 11, 2022, entitled “POINT CLOUD ALIGNMENT SYSTEMSFOR GENERATING HIGH DEFINITION MAPS FOR VEHICLE NAVIGATION,” the contents of which are incorporated herein by reference in their entirety. Background Art
[0003] An autonomous vehicle is capable of sensing its surroundings and navigating through its surroundings with minimal to no human input. In order to safely navigate the vehicle along a selected path, the vehicle may rely on motion planning processing to generate, update, and execute one or more trajectories through its surroundings. The trajectory of the vehicle may be generated based on the current conditions of the vehicle itself and the conditions present in the vehicle's surroundings, which may include moving objects such as other vehicles and pedestrians, and non-moving objects such as buildings and road poles. For example, a trajectory may be generated to avoid collisions between the vehicle and objects present in its surroundings. In addition, a trajectory may be generated that enables the vehicle to operate according to other desired characteristics, such as path length, ride quality or comfort, required travel time, compliance with traffic regulations, and / or following driving practices. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is an example environment in which a vehicle including one or more components of an autonomous system may be implemented;
[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0006] Figure 3 yes Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;
[0007] Figure 4A is a diagram of some components of an autonomous system;
[0008] Figure 4B is a graph of the implementation of a neural network;
[0009] Figure 4C and Figure 4D is a diagram illustrating an example operation of a CNN;
[0010] Figure 5is a block diagram of a system for determining a trajectory of a vehicle in physical space;
[0011] Figure 6 Describing a system for implementing a process for determining a trajectory of a vehicle within a physical space using a machine learning model for determining relative transformations between point clouds;
[0012] Figure 7 Describing a system for implementing another process of determining a trajectory of a vehicle within a physical space using a machine learning model for determining a plurality of relative transformations between point clouds;
[0013] Fig. 8A Describing a system for performing an implementation of another process for determining a trajectory of a vehicle within a physical space using a machine learning model pre-trained to determine relative transformations;
[0014] Figure 8B shows a block diagram detailing the components and steps included as part of the operation of a reconstruction network for pre-training a machine learning model to determine a relative transformation; and
[0015] Fig. 9 A flow chart illustrating an example of a process for determining a trajectory of a vehicle within a physical space using a machine learning model for determining one or more relative transformations between point clouds is depicted. DETAILED DESCRIPTION
[0016] In the following description, for the purpose of explanation, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described in the present disclosure can be implemented without these specific details. In some instances, well-known configurations and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0017] In the accompanying drawings, for ease of description, the specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that, unless explicitly described, the specific order or arrangement of schematic elements in the accompanying drawings is not intended to mean that a specific processing order or sequence, or separation of processing is required. In addition, unless explicitly described, the inclusion of schematic elements in the accompanying drawings is not intended to mean that such elements are required in all embodiments, nor is it intended to mean that the features represented by such elements cannot be included in some embodiments or cannot be combined with other elements in some embodiments.
[0018] In addition, in the accompanying drawings, connecting elements (such as solid or dotted lines or arrows, etc.) are used to illustrate the connection, relationship or association between or among two or more other schematic elements, and the absence of any such connecting elements is not intended to mean that there can be no connection, relationship or association. In other words, some connections, relationships or associations between elements are not illustrated in the accompanying drawings so as not to obscure the present disclosure. In addition, for ease of illustration, a single connecting element can be used to represent multiple connections, relationships or associations between elements. For example, if the connecting element represents the communication of a signal, data or instruction (e.g., "software instruction"), it will be understood by those skilled in the art that such an element can represent one or more than one signal path (e.g., bus) that may be needed to affect the communication.
[0019] Although the terms "first", "second" and / or "third", etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first", "second" and / or "third" are only used to distinguish one element from another. For example, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact without departing from the scope of the described embodiments. Both the first contact and the second contact are contacts, but they are not the same contacts.
[0020] The terms used in the specification of the various embodiments described herein are included only for the purpose of describing specific embodiments and are not intended to be limiting. As used in the specification of the various embodiments described and the appended claims, the singular forms "a", "an" and "the" are also intended to include plural forms and can be used interchangeably with "one or more than one" or "at least one", unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more than one of the associated listed items. It will also be understood that when the terms "include", "comprise", "have" and / or "have" are used in this specification, the stated features, integers, steps, operations, elements and / or components are specifically stated, but the presence or addition of one or more than one other features, integers, steps, operations, elements, components and / or groups thereof are not excluded.
[0021] As used herein, the terms "communication" and "communicating" refer to at least one of receiving, receiving, transmitting, transmitting and / or providing information (or information represented by, for example, data, signals, messages, instructions and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof) to communicate with another unit, this means that the unit is able to directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is wired and / or wireless in nature. In addition, even if the transmitted information can be modified, processed, relayed and / or routed between the first unit and the second unit, the two units can communicate with each other. For example, even if the first unit passively receives information and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediary unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message may refer to a network packet (eg, a data packet, etc.) that includes data.
[0022] As used herein, the term "if" is optionally interpreted to mean "when," "at," "in response to being determined to be," and / or "in response to being detected," etc., depending on the context. Similarly, the phrases "if it is determined" or "if [the stated condition or event] is detected" are optionally interpreted to mean "when determining," "in response to being determined to be" or "when [the stated condition or event] is detected," and / or "in response to being detected," etc., depending on the context. In addition, as used herein, the terms "have," "have," or "possess," etc. are intended to be open-ended terms. Furthermore, unless expressly stated otherwise, the phrase "based on" is intended to mean "based at least in part on."
[0023] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments described. However, it will be apparent to one of ordinary skill in the art that the various embodiments described may be implemented without these specific details. In other cases, well-known methods, processes, components, circuits, and networks have not yet been described in detail in order not to unnecessarily obscure aspects of the embodiments.
[0024] General Overview
[0025] In some aspects and / or embodiments, the systems, methods, and computer program products described herein include and / or implement a machine learning model for a vehicle that determines a trajectory of the vehicle in a physical space based at least on a first relative transformation (e.g., a target point cloud aligned with a source point cloud). Note that the source point cloud and the target point cloud are three-dimensional point clouds. This alignment enables the generation of accurate and reliable high-definition maps that the vehicle can utilize to safely navigate at least a portion of the physical space. In some instances, the aligned target point cloud can be further aligned with the source point cloud to further improve HD map generation. The vehicle trajectories determined using these maps improve the likelihood of avoiding collisions between the vehicle and other vehicles, pedestrians, and the like. In addition, these determined trajectories ensure that the vehicle operates in accordance with traffic rules and driving best practices, among others. Note that the determined trajectories can be associated with past paths that the vehicle has traveled, and alternatively, can also be associated with future paths that the vehicle can travel.
[0026] By means of the implementation of the systems, methods, and computer program products described herein, techniques for determining one or more relative transformations for point cloud alignment and for training a machine learning model to determine one or more relative transformations are provided. For example, a machine learning model may be trained to determine a first relative transformation for aligning a target point cloud with a source point cloud and a second relative transformation for further aligning the aligned target point cloud with the source point cloud. The alignment may include a translation along at least one of the x-axis, the y-axis, and the z-axis, and additionally or alternatively, the alignment may include a rotation of θ radians about a unit axis (X, Y, Z) around a fixed point. Determining the first relative transformation and the second relative transformation may include: performing weighted subsampling using a convolutional layer (e.g., a 1×1 convolutional layer); flattening the combined feature map into a one-dimensional vector; and processing the one-dimensional vector using a fully connected layer of a machine learning model. In addition, in an embodiment, the machine learning model may be pre-trained to determine the relative transformation. Pre-training may include: reconstructing a point cloud by at least decoding an encoding of a point cloud deformed by one or more offsets; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud. In this way, a machine learning model can be trained to generate one or more relative transformations that are used to align a point cloud for use in generating a high-definition map that facilitates safe vehicle navigation.
[0027] Reference now Figure 1, illustrates an example environment 100 in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a-102n, objects 104a-104n, routes 106a-106n, areas 108, vehicle-to-infrastructure (V2I) devices 110, a network 112, a remote autonomous vehicle (AV) system 114, a fleet management system 116, and a V2I system 118. The vehicles 102a-102n, the vehicle-to-infrastructure (V2I) devices 110, the network 112, the autonomous vehicle (AV) system 114, the fleet management system 116, and the V2I system 118 are interconnected (e.g., establish connections for communication, etc.) via wired connections, wireless connections, or a combination of wired or wireless connections. In some embodiments, objects 104a-104n are interconnected with at least one of vehicles 102a-102n, vehicle-to-infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, fleet management systems 116, and V2I systems 118 via wired connections, wireless connections, or a combination of wired or wireless connections.
[0028] Vehicles 102a-102n (individually referred to as vehicles 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, vehicles 102 are configured to communicate with V2I devices 110, remote AV systems 114, fleet management systems 116, and / or V2I systems 118 via network 112. In some embodiments, vehicles 102 include cars, buses, trucks, and / or trains, etc. In some embodiments, vehicles 102 are similar to vehicles 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, vehicles 102 travel along respective routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) as described herein. In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same as or similar to autonomous system 202).
[0029] Objects 104a-104n (individually referred to as objects 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.), etc. Each object 104 is stationary (e.g., located at a fixed location and over a period of time) or moves (e.g., has a speed and is associated with at least one trajectory). In some embodiments, objects 104 are associated with corresponding locations in area 108.
[0030] Routes 106a-106n (individually referred to as routes 106 and collectively referred to as routes 106) are each associated with (e.g., specify a series of actions (also referred to as trajectories) connecting states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatiotemporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatiotemporal location different from the first spatiotemporal location) or a target zone (e.g., a subspace of acceptable states (e.g., terminal states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, routes 106 include multiple acceptable state sequences (e.g., multiple spatiotemporal location sequences) that are associated with (e.g., define multiple trajectories). In an example, routes 106 include only high-level actions or imprecise state locations, such as a series of connecting roads indicating a change of direction at a roadway intersection, etc. Additionally or alternatively, the route 106 may include more precise actions or states, such as, for example, specific target lanes or precise locations within lane regions and target speeds at those locations, etc. In an example, the route 106 includes a plurality of precise state sequences along at least one high-level action with a limited look-ahead horizon to an intermediate target, wherein a combination of consecutive iterations of the limited horizon state sequences cumulatively correspond to a plurality of trajectories that collectively form a high-level route terminating at a final target state or region.
[0031] The area 108 includes a physical area (e.g., a geographic region) in which the vehicle 102 can navigate. In an example, the area 108 includes at least one state (e.g., a country, a province, a separate state of a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, the area 108 includes at least one named thoroughfare (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, the area 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of an open space and / or undeveloped area, a dirt road, etc. In some embodiments, the road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, the road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0032] The vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the fleet management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), lane markings, street lights, parking meters, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the fleet management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112 .
[0033] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a long-term evolution (LTE) network, a third generation (3G) network, a fourth generation (4G) network, a fifth generation (5G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks, etc.
[0034] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the fleet management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the fleet management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.
[0035] The queue management system 116 includes at least one device configured to communicate with the vehicles 102, the V2I devices 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems)).
[0036] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the fleet management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection other than the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipality or a private agency (e.g., a private agency for maintaining the V2I device 110, etc.).
[0037] supply Figure 1 The number and arrangement of elements illustrated are examples. Figure 1 There may be additional elements, fewer elements, different elements, and / or differently arranged elements than those illustrated. Additionally or alternatively, at least one element of environment 100 may be described as being Figure 1 Additionally or alternatively, at least one set of elements of environment 100 may perform one or more functions described as being performed by at least one different set of elements of environment 100.
[0038] Reference now Figure 2, (can be used with Figure 1 Vehicle 200 includes, or is associated with, autonomous system 202, powertrain control system 204, steering control system 206, and braking system 208. In some embodiments, vehicle 200 is similar to vehicle 102 (see Figure 1 ) is the same or similar. In some embodiments, the autonomous system 202 is configured to give the vehicle 200 autonomous driving capabilities (e.g., implementing at least one of the following driving automatic or maneuver-based functions, features and / or devices, etc., which enable the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as Level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as Level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as Level 3 ADS-operated vehicles, etc.), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functions required to enable the vehicle 200 to operate in road traffic and continuously perform part or all of a dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) including driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., Level 0) to full driving automation (e.g., Level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference may be made to SAE International's standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire contents of which are incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.
[0039] Autonomous system 202 includes a sensor suite that includes one or more devices such as camera 202a, LiDAR sensor 202b, Radar sensor 202c, and microphone 202d. In some embodiments, autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance that vehicle 200 has traveled, etc.). In some embodiments, autonomous system 202 uses one or more devices included in autonomous system 202 to generate data associated with environment 100 described herein. Data generated by one or more devices of autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which vehicle 200 is located. In some embodiments, autonomous system 202 includes communication device 202e, autonomous vehicle computing 202f, drive-by-wire (DBW) system 202h, and safety controller 202g.
[0040] The camera 202a includes a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g configured to communicate with the communication device 202e via a bus (e.g., Figure 3 The camera 202a includes at least one device for communicating with the autonomous vehicle computing 202f (e.g., an image processing unit 202a, a bus ... Figure 1In some embodiments, the autonomous vehicle computing 202f determines a depth to one or more objects in a field of view of at least two of the plurality of cameras based on image data from the at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance relative to the camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, the camera 202a includes features such as sensors and lenses that are optimized for sensing objects at one or more distances relative to the camera 202a.
[0041] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with the one or more images. In some examples, the camera 202a generates TLD (traffic light detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data differs from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras with a wide field of view (e.g., a wide-angle lens, a fisheye lens, and / or a lens with a viewing angle of approximately 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.
[0042] The light detection and ranging (LiDAR) sensor 202b includes a communication device 202e, an autonomous vehicle computing device 202f, and / or a safety controller 202g via a bus (e.g., Figure 3The LiDAR sensor 202b includes at least one device that communicates with a bus (the same or similar bus as the bus 302 of the embodiment of the present invention). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object encountered by the light. The LiDAR sensor 202b also includes at least one light detector that detects the light emitted from the light emitter after encountering the physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundaries of the physical object and / or the surface of the physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundaries of the physical object in the field of view of the LiDAR sensor 202b.
[0043] The radio detection and ranging (Radar) sensor 202c includes a sensor configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with a bus (same or similar bus as bus 302 of the embodiment of the present invention). Radar sensor 202c includes a system configured to transmit (pulsed or continuous) radio waves. The radio waves transmitted by Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves transmitted by Radar sensor 202c encounter physical objects and are reflected back to Radar sensor 202c. In some embodiments, the radio waves transmitted by Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with Radar sensor 202c generates a signal representing an object included in the field of view of Radar sensor 202c. For example, at least one data processing system associated with Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of a physical object in the field of view of Radar sensor 202c.
[0044] The microphone 202d includes a microphone configured to communicate with the communication device 202e, the autonomous vehicle computing device 202f, and / or the safety controller 202g via a bus (e.g., Figure 3 At least one device for communicating with the vehicle 200 (the same or similar bus as bus 302 of FIG. 1 ). Microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, microphone 202d includes a transducer device and / or the like. In some embodiments, one or more systems described herein can receive the data generated by microphone 202d and determine the position (e.g., distance, etc.) of an object relative to vehicle 200 based on an audio signal associated with the data.
[0045] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. For example, the communication device 202e may include at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the drive-by-wire (DBW) system 202h. Figure 3 The communication device 202e may be a device that is the same as or similar to the communication interface 314 of the vehicle. In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (eg, a device for enabling wireless communication of data between vehicles).
[0046] Autonomous vehicle computing 202f includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, safety controller 202g, and / or DBW system 202h. In some examples, autonomous vehicle computing 202f includes devices such as client devices, mobile devices (e.g., cellular phones and / or tablet computers, etc.), and / or servers (e.g., computing devices including one or more central processing units and / or graphics processing units, etc.). In some embodiments, autonomous vehicle computing 202f is the same or similar to autonomous vehicle computing 400 described herein. Additionally or alternatively, in some embodiments, autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., with Figure 1 remote AV system 114 of the same or similar autonomous vehicle system), a fleet management system (e.g., Figure 1 of the same or similar queue management system as the queue management system 116), V2I devices (e.g., Figure 1 V2I device 110 that is the same as or similar to V2I device 110) and / or V2I system (e.g., Figure 1The V2I system 118 may communicate with the same or similar V2I system.
[0047] Safety controller 202g includes at least one device configured to communicate with camera 202a, LiDAR sensor 202b, Radar sensor 202c, microphone 202d, communication device 202e, autonomous vehicle computing 202f, and / or DBW system 202h. In some examples, safety controller 202g includes one or more controllers (electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). In some embodiments, safety controller 202g is configured to generate control signals that take precedence over (e.g., override) control signals generated and / or transmitted by autonomous vehicle computing 202f.
[0048] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (e.g., electrical controllers and / or electromechanical controllers, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (e.g., powertrain control system 204, steering control system 206, and / or braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (e.g., turn signal lights, headlights, door locks, and / or windshield wipers, etc.).
[0049] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to make longitudinal vehicle movements such as starting to move forward, stopping to move forward, starting to move backward, stopping to move backward, accelerating in a certain direction, decelerating in a certain direction, etc. and / or causes the vehicle 200 to make lateral vehicle movements such as turning left, turning right, etc. In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (e.g., fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby rotating or not rotating at least one wheel of the vehicle 200.
[0050] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates the two front wheels and / or the two rear wheels of the vehicle 200 to the left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activity required to adjust the y-axis component of the vehicle's motion.
[0051] Braking system 208 includes at least one device configured to actuate one or more brakes to decelerate and / or hold vehicle 200 stationary. In some examples, braking system 208 includes at least one controller and / or actuator configured to cause one or more calipers associated with one or more wheels of vehicle 200 to close on respective rotors of vehicle 200. Additionally or alternatively, in some examples, braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.
[0052] In some embodiments, vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring a property of a state or condition of vehicle 200. In some examples, vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), wheel rate sensors, wheel brake pressure sensors, wheel torque sensors, engine torque sensors, and / or steering angle sensors. Although braking system 208 is Figure 2 Although illustrated as being located on the proximal side of the vehicle 200 , the braking system 208 may be located anywhere in the vehicle 200 .
[0053] Reference now Figure 3 , a schematic diagram of an exemplary device 300. As illustrated, device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, device 300 corresponds to: at least one device of vehicle 102 (e.g., at least one device of a system of vehicle 102); at least one device of system 500; and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112). In some embodiments, one or more devices of vehicle 102 (e.g., one or more devices of a system of vehicle 102), one or more devices of system 500, and / or one or more devices of network 112 (e.g., one or more devices of a system of network 112) include at least one device 300 and / or at least one component of device 300. As Figure 3 As shown, apparatus 300 includes a bus 302 , a processor 304 , a memory 306 , a storage component 308 , an input interface 310 , an output interface 312 , and a communication interface 314 .
[0054] The bus 302 includes components that permit communication between components of the device 300. In some cases, the processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). The memory 306 includes a random access memory (RAM), a read-only memory (ROM), and / or another type of dynamic and / or static storage device (e.g., flash memory, magnetic memory, and / or optical memory, etc.) that stores data and / or instructions for use by the processor 304.
[0055] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid-state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette, a tape, a CD-ROM, a RAM, a PROM, an EPROM, a FLASH-EPROM, an NV-RAM, and / or another type of computer-readable medium, and a corresponding drive.
[0056] The input interface 310 includes components that permit the device 300 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, the input interface 310 includes a sensor for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). The output interface 312 includes components for providing output information from the device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0057] In some embodiments, communication interface 314 includes a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter, etc.) that permits device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from another device and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface and / or cellular network interface, etc.
[0058] In some embodiments, the device 300 performs one or more processes described herein. The device 300 performs these processes based on the processor 304 executing software instructions stored by a computer-readable medium such as a memory 306 and / or a storage component 308. Computer-readable media (e.g., non-transitory computer-readable media) are defined herein as non-transitory memory devices. Non-transitory memory devices include storage space located within a single physical storage device or storage space distributed across multiple physical storage devices.
[0059] In some embodiments, the software instructions are read into the memory 306 and / or storage component 308 from another computer-readable medium or from another device via the communication interface 314. The software instructions stored in the memory 306 and / or storage component 308, when executed, cause the processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with the software instructions to perform one or more processes described herein. Therefore, unless expressly stated otherwise, the embodiments described herein are not limited to any specific combination of hardware circuitry and software.
[0060] The memory 306 and / or the storage component 308 include a data storage unit or at least one data structure (e.g., a database, etc.). The device 300 can receive information from the data storage unit or at least one data structure in the memory 306 or the storage component 308, store information in the data storage unit or at least one data structure, communicate information to the data storage unit or at least one data structure, or search for information stored in the data storage unit or at least one data structure. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0061] In some embodiments, the device 300 is configured to execute software instructions stored in the memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to the device 300). As used herein, the term "module" refers to at least one instruction stored in the memory 306 and / or the memory of another device, which, when executed by the processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to the device 300), causes the device 300 (e.g., at least one component of the device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.
[0062] supply Figure 3 The number and arrangement of components illustrated are examples. Figure 3 The device 300 may include additional components, fewer components, different components, or differently arranged components than those illustrated. Additionally or alternatively, a set of components (e.g., one or more components) of the device 300 may perform one or more functions described as being performed by another component or set of components of the device 300.
[0063] Reference now Figure 4A, illustrates an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack"). As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of a vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems that are the same or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, planning system 404, positioning system 406, control system 408, and database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., by a microprocessor, microcontroller, application specific integrated circuit (ASIC) and / or field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, the autonomous vehicle computing 400 is configured to communicate with a remote system (e.g., an autonomous vehicle system that is the same or similar to the remote AV system 114, a fleet management system 116 that is the same or similar to the fleet management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.).
[0064] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object) and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with (e.g., representing) one or more physical objects within the field of view of the at least one camera. In such examples, the perception system 402 classifies the at least one physical object based on one or more groups of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical object by the perception system 402, the perception system 402 transmits data associated with the classification of the physical object to the planning system 404.
[0065] In some embodiments, the planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data (e.g., data associated with the classification of physical objects described above) from the perception system 402, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to tactical functions required to operate the vehicle 102 in traffic on the road. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate rate, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with an updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.
[0066] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) a location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In some examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on each point cloud. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated before the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of the roadway (such as traffic speed, traffic volume, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction, or the type and location of lane markings, or a combination thereof), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various types of other driving signals, etc.) In some embodiments, the map is generated in real time based on data received by the perception system.
[0067] In another example, positioning system 406 receives global navigation satellite system (GNSS) data generated by a global positioning system (GPS) receiver. In some examples, positioning system 406 receives GNSS data associated with a location of a vehicle in an area, and positioning system 406 determines the latitude and longitude of the vehicle in the area. In such an example, positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the location of the vehicle determined by positioning system 406, positioning system 406 generates data associated with the position of the vehicle. In such an example, the data associated with the position of the vehicle include data associated with one or more semantic properties corresponding to the position of the vehicle.
[0068] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to operate the powertrain control system (e.g., DBW system 202h and / or powertrain control system 204, etc.), the steering control system (e.g., steering control system 206) and / or the braking system (e.g., braking system 208). For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes the activity required to adjust the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes the activity required to adjust the x-axis component of the vehicle motion. In the example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby turning the vehicle 200 left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (eg, headlights, turn signal lights, door locks, and / or windshield wipers, etc.) to change states.
[0069] In some embodiments, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, perception system 402, planning system 404, positioning system 406, and / or control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in an environment, etc.). The following is about FIG. 4B to FIG. 4D Includes examples of implementations of machine learning models.
[0070] Database 410 stores data transmitted to, received from, and / or updated by perception system 402, planning system 404, positioning system 406, and / or control system 408. In some examples, database 410 includes a storage component (e.g., a storage component) for storing data and / or software related to operations and using at least one system of autonomous vehicle computing 400. Figure 3In some embodiments, database 410 stores data associated with a 2D and / or 3D map of at least one area. In some examples, database 410 stores data associated with a 2D and / or 3D map of a portion of a city, portions of multiple cities, multiple cities, counties, states, and / or countries (State) (e.g., a country), etc. In such an example, a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200) can be driven along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, remote roads, and / or off-road roads, etc.) and cause at least one LiDAR sensor (e.g., a LiDAR sensor that is the same or similar to LiDAR sensor 202b) to generate data associated with an image representing an object included in the field of view of the at least one LiDAR sensor.
[0071] In some embodiments, database 410 can be implemented across multiple devices. In some examples, database 410 includes a vehicle (e.g., a vehicle that is the same or similar to vehicle 102 and / or vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system that is the same or similar to remote AV system 114), a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114), and a fleet management system (e.g., a vehicle that is the same or similar to remote AV system 114). Figure 1 The same or similar queue management system as the queue management system 116 of FIG. 1 and / or the V2I system (e.g., Figure 1 The V2I system 118 is the same as or similar to the V2I system) and the like.
[0072] Reference now Figure 4B , a diagram illustrating an implementation of a machine learning model. More specifically, a diagram illustrating an implementation of a convolutional neural network (CNN) 420. For purposes of illustration, the following description of CNN 420 will be with respect to implementing CNN 420 by perception system 402. However, it will be understood that in some examples, CNN 420 (e.g., one or more components of CNN 420) is implemented by other systems (such as planning system 404, positioning system 406, and / or control system 408, etc.) other than or in addition to perception system 402. Although CNN 420 includes certain features as described herein, these features are provided for purposes of illustration and are not intended to limit the present disclosure.
[0073] CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, subsampling layer 428 and / or other subsampling layers have a dimension that is smaller than the dimension of the upstream system (i.e., the number of nodes). With the subsampling layer 428 having a dimension that is smaller than the dimension of the upstream layer, CNN 420 merges the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for CNN 420 to perform downstream convolution operations. Additionally or alternatively, with the subsampling layer 428 being associated with (e.g., configured to perform) at least one subsampling function (as described below with respect to Figure 4C and Figure 4D As described above, CNN 420 incorporates the amount of data associated with the initial input.
[0074] The perception system 402 performs the convolution operation based on the perception system 402 providing respective inputs and / or outputs associated with each of the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 to generate respective outputs. In some examples, the perception system 402 implements the CNN 420 based on the perception system 402 providing data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426. In such examples, the perception system 402 provides data as input to the first convolution layer 422, the second convolution layer 424, and the convolution layer 426 based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle that is the same or similar to the vehicle 102, a remote AV system that is the same or similar to the remote AV system 114, a queue management system that is the same or similar to the queue management system 116, and / or a V2I system that is the same or similar to the V2I system 118, etc.). The following is about Figure 4C Includes a detailed description of the convolution operation.
[0075] In some embodiments, the perception system 402 provides data associated with the input (referred to as the initial input) to the first convolutional layer 422, and the perception system 402 generates data associated with the output using the first convolutional layer 422. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as the upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as the downstream layer.
[0076] In some embodiments, before the perception system 402 provides the input to the CNN 420, the perception system 402 processes the data associated with the input provided to the CNN 420. For example, the perception system 402 processes the data associated with the input provided to the CNN 420 based on the perception system 402 normalizing the sensor data (e.g., image data, LiDAR data, and / or Radar data, etc.).
[0077] In some embodiments, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer. In some examples, CNN 420 generates an output based on perception system 402 performing convolution operations associated with each convolution layer and the initial input. In some embodiments, perception system 402 generates an output and provides the output to fully connected layer 430. In some examples, perception system 402 provides the output of convolution layer 426 to fully connected layer 430, wherein fully connected layer 430 includes data associated with multiple feature values referred to as F1, F2, ..., FN. In this example, the output of convolution layer 426 includes data associated with multiple output feature values representing predictions.
[0078] In some embodiments, perception system 402 identifies a prediction from the plurality of predictions based on perception system 402 identifying a feature value associated with a highest likelihood of being a correct prediction from the plurality of predictions. For example, where fully connected layer 430 includes feature values F1, F2, ..., FN and F1 is the largest feature value, perception system 402 identifies the prediction associated with F1 as the correct prediction from the plurality of predictions. In some embodiments, perception system 402 trains CNN 420 to generate the predictions. In some examples, perception system 402 trains CNN 420 to generate the predictions based on perception system 402 providing training data associated with the predictions to CNN 420.
[0079] Reference now Figure 4C and Figure 4D , a diagram illustrating an example operation of CNN 440 utilizing perception system 402. In some embodiments, CNN 440 (e.g., one or more components of CNN 440) is coupled to CNN 420 (e.g., one or more components of CNN 420) (see Figure 4B ) are the same or similar.
[0080] At step 450, the perception system 402 provides data associated with the image as input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with the image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.
[0081] At step 455, CNN 440 performs a first convolution function. For example, CNN 440 performs a first convolution function based on CNN 440 providing a value representing an image as an input to one or more neurons (not explicitly illustrated) included in first convolution layer 442. In this example, the value representing the image may correspond to a value of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) may be represented as an array of values corresponding in size to the value provided as input to the neuron. In one example, the filter may be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolution layers, the filters associated with the neurons may be configured to continuously identify more complex patterns (e.g., arcs and / or objects, etc.).
[0082] In some embodiments, CNN 440 performs a first convolution function based on CNN 440 multiplying the values of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values of the filters corresponding to each neuron in the same or more neurons. For example, CNN 440 may multiply the values of each neuron provided as input to one or more neurons included in the first convolution layer 442 by the values of the filters corresponding to each neuron in the one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolution layer 442 is referred to as a convolution output. In some embodiments, when each neuron has the same filter, the convolution output is referred to as a feature map.
[0083] In some embodiments, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of the downstream layer. For clarity, the upstream layer may be a layer that transmits data to a different layer (referred to as the downstream layer). For example, CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In the example, CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the first subsampling layer 444. In such an example, CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.
[0084] At step 460, CNN 440 performs a first subsampling function. For example, based on CNN 440 providing the values output by first convolutional layer 442 to the corresponding neurons of first subsampling layer 444, CNN 440 may perform the first subsampling function. In some embodiments, CNN 440 performs the first subsampling function based on an aggregation function. In an example, CNN 440 performs the first subsampling function based on CNN 440 determining the maximum input (referred to as a maximum pooling function) among the values provided to a given neuron. In another example, CNN 440 performs the first subsampling function based on CNN 440 determining the average input (referred to as an average pooling function) among the values provided to a given neuron. In some embodiments, based on CNN 440 providing values to the respective neurons of first subsampling layer 444, CNN 440 generates an output, which is sometimes referred to as a subsampled convolution output.
[0085] At step 465, CNN 440 performs a second convolution function. In some embodiments, CNN 440 performs the second convolution function in a manner similar to how CNN 440 performs the first convolution function described above. In some embodiments, CNN 440 performs the second convolution function based on CNN 440 providing the value output by first subsampling layer 444 as input to one or more neurons (not explicitly illustrated) included in second convolution layer 446. In some embodiments, as described above, each neuron of second convolution layer 446 is associated with a filter. As described above, the filter (one or more) associated with second convolution layer 446 can be configured to recognize more complex patterns than the filter associated with first convolution layer 442.
[0086] In some embodiments, the CNN 440 performs a second convolution function based on the CNN 440 multiplying the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons. For example, the CNN 440 may multiply the value of each neuron provided as input to the one or more neurons included in the second convolution layer 446 by the value of the filter corresponding to each neuron of the one or more neurons to generate a single value or a value array as an output.
[0087] In some embodiments, the CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of the downstream layer. For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the second subsampling layer 448. In such an example, the CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.
[0088] At step 470, CNN 440 performs a second subsampling function. For example, based on CNN 440 providing the values output by second convolutional layer 446 to corresponding neurons of second subsampling layer 448, CNN 440 may perform a second subsampling function. In some embodiments, based on CNN 440 using an aggregation function, CNN 440 performs a second subsampling function. In an example, as described above, based on CNN 440 determining the maximum input or average input among the values provided to a given neuron, CNN 440 performs a first subsampling function. In some embodiments, based on CNN 440 providing values to respective neurons of second subsampling layer 448, CNN 440 generates an output.
[0089] At step 475, CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449. For example, CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449 so that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as a classification). The prediction may include an indication that the objects included in the image provided as input to CNN 440 include objects and / or sets of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to the various systems described herein.
[0090] Figure 5 is a high-level block diagram of a system for determining a trajectory of a vehicle within a physical space. Note that a trajectory may refer to a path or route that a vehicle has previously traveled, or a predicted path or route that a vehicle may travel in the future. In an embodiment, an autonomous vehicle computing 400 included as part of the vehicle 102 may apply a machine learning model 506 that is trained to determine relative transformations between multiple point clouds. For example, the machine learning model 506 may receive a target point cloud 502 and a source point cloud 504 as inputs, and determine a relative transformation 508 based on these inputs. The relative transformation 508 may be used to align the target point cloud 502 with the source point cloud 504, and may also be used to determine a trajectory of the vehicle, such as within an environment external to the vehicle. The target point cloud 502 and the source point cloud 504 may each be a three-dimensional point cloud that includes a collection of points in a three-dimensional space that are the location of an autonomous vehicle (e.g., Figure 1 The vehicle 102 shown, Figure 2 Representation of various attributes of the physical space outside of the vehicle 200, etc. shown.
[0091] For example, the points within each of the target point cloud 502 and the source point cloud 504 may represent specific attributes of one or more objects located within the vicinity of the vehicle in physical space. A non-limiting example of a physical space may be at least a portion of a street, a highway, a city block, or a city intersection through which the vehicle can navigate. While driving through such a physical space, the LiDAR sensor 202b of the vehicle may capture point cloud information representing one or more objects (such as pedestrians, road signs, traffic signs, buildings, and / or other vehicles, etc.) in real time. The points in the target point cloud 502 and the source point cloud 504 may include various components describing the attributes of the physical space. For example, in addition to information related to intensity and depth, each point in the point cloud may also include coordinate information (e.g., x, y, and z coordinates). The intensity information may be based on the returned intensity of light detected by the LiDAR sensor, and the depth information may be associated with distance information relative to the LiDAR sensor. Autonomous vehicle computing 400 may utilize one or more of the coordinate information, intensity information, and depth information to determine details related to the size of one or more objects present in the physical space and / or the distances between each of these objects (e.g., pedestrians, road signs, traffic signs, buildings, and / or other vehicles, etc.) and the vehicle in the physical space. For example, target point cloud 502 and source point cloud 504 may capture different perspectives of at least a portion of the same physical space based on the position of the LiDAR sensor on the vehicle. As a non-limiting example, both target point cloud 502 and source point cloud 504 may capture different perspectives of the same city block, city street, pedestrians, etc.
[0092] Autonomous vehicle computing 400 may apply a machine learning model to determine a relative transformation for aligning target point cloud 502 with source point cloud 504. For example, target point cloud 502 may be translated and rotated so that target point cloud 502 is aligned with the orientation of source point cloud 504, e.g., target point cloud 502 may be aligned with source point cloud 504 within the coordinate system of source point cloud 504. In some embodiments, additional relative transformations may be determined that may enable further alignment of target point cloud 502 with source point cloud 504.
[0093] Figure 6 Depicts a system 500 for implementing a process for determining a trajectory of a vehicle within a physical space using a machine learning model 506 in accordance with some embodiments of the present subject matter described and illustrated herein. In particular, Figure 6Detailed description of the components and implementation steps associated with the machine learning model 506 to determine the relative transformation used to align the target point cloud 502 with the source point cloud 504. As described above, such alignment enables the generation of aligned point clouds and, in addition, enables the generation of accurate and reliable high-definition maps that can be used by the autonomous vehicle computing 400 to determine the trajectory of the vehicle within the physical space, for example, for navigation on at least a portion of a street, highway, city block, urban intersection, etc.
[0094] In an embodiment, as part of the application of the machine learning model 506, the machine learning model 506 may receive as input each of the target point cloud 502 and the source point cloud 504. As described, the target point cloud 502 and the source point cloud 504 may each be captured by the vehicle's LiDAR sensor 202b and may represent different perspectives of at least a portion of the same physical space (e.g., the same city block, city street, pedestrians, etc.). The received point clouds may be input into the encoding network 602. In some embodiments, the encoding network 602 encodes each of the target point cloud 502 and the source point cloud 504 so that one or more objects present in each of the target point cloud 502 and the source point cloud 504 are represented as one or more columns. The one or more columns are each used to define a plurality of points associated with the corresponding object. For example, the encoding network 602 may receive as input each of the target point cloud 502 and the source point cloud 504 and predict a three-dimensional box (e.g., a "column") of various types and categories of objects present in the physical space and represented by the point cloud. For example, cars, pedestrians, cyclists, traffic lights, road signs, and / or buildings, etc. As part of encoding the point clouds, the encoding network 602 prepares each of the target point cloud 502 and the source point cloud 504 for feature extraction.
[0095] In some embodiments, such an encoding network (e.g., encoding network 602) can correspond to a column feature encoder network and can operate in conjunction with a convolutional backbone and a detection head (not shown). For example, a column feature encoder network can correspond to encoding network 602 and operate to convert a point cloud into a sparse pseudo image. Thereafter, the sparse pseudo image can be input into a two-dimensional convolutional backbone, which processes the pseudo image into a high-level representation of the image. Finally, the detection head can detect and regress a three-dimensional box.
[0096] In some embodiments, the encoding network 602 in the form of a pillar feature encoder network can convert a particular point cloud (e.g., the target point cloud 502 and the source point cloud 504) into a pseudo image in the form of a two-dimensional image embedding (tensor) with more than three channels. Thereafter, as a non-limiting example, each point cloud can be discretized into a uniformly spaced grid with respect to a particular plane (e.g., the xy plane) to create a set of pillars. c ,y c 、z c 、x p and p to enhance each point in each column in the set of columns, where the subscript c represents the distance value relative to the arithmetic mean of all points in the column, and the subscript p represents the offset relative to the center coordinate of the column. Note that the enhanced LiDAR point has nine dimensions. In addition, in some embodiments, a linear layer is applied to each enhanced point in the column to generate a tensor of a specific size (C, P, N). In addition, a maximum operation can be performed on the channel to generate an output tensor of size (C, P). In this way, the encoding network 602 can encode each of the target point cloud 502 and the source point cloud 504, so that one or more objects present in each point cloud can be represented as one or more columns, which define multiple points associated with various objects in physical space (e.g., cars, pedestrians, cyclists, traffic lights, road signs and / or buildings, etc.).
[0097] In some embodiments, the backbone 604 may include two subnetworks. The first subnetwork may be a top-down network that generates features at increasingly smaller spatial resolutions, and the second subnetwork may perform upsampling and cascading of top-down features. For example, a top-down network may be characterized by a series of blocks, each of which operates at a specific stride, wherein the stride may be measured relative to a pseudo image of the original input. Each block may have a specific number of two-dimensional convolutional layers (e.g., 3×3 two-dimensional convolutional layers), which have a specific number of output channels, each of which is followed by BatchNorm and ReLU. The final features from each top-down block may be combined by upsampling and cascading. In particular, a set of features is upsampled using multiple final features using transposed two-dimensional convolutions, and BatchNorm and ReLU may then be applied to the upsampled features. The final feature may be a cascade of all features derived from different strides. Non-limiting examples of the backbone 604 may include ResNet and / or VGG, etc.
[0098] refer to Figure 6, after one or more of the above processes are completed, the output of the backbone 604 is two different feature maps, namely a source feature map 606 and a target feature map 608. The source feature map 606 corresponds to the encoded representation of the source point cloud 504, and the target feature map corresponds to the encoded representation of the target point cloud 502. Thereafter, the cascade 610 receives the source feature map 606 and the target feature map 608 as inputs, and outputs a combined feature map 612. In some embodiments, the combined feature map 612 may be the result of a stacking operation, in which the target feature map 608 may be directly positioned on the source feature map 606, for example, without relying on any calculation or operation. In an embodiment, the result of such a stacking operation may be that the length and width values of the feature maps are retained, while the height value is equal to the combined value of the heights of the feature maps. Thereafter, the combined feature map 612 may be input into a convolutional layer 614, which outputs a fused feature map 616, for example as a result of a 1×1 convolution. For example, as part of the application of the machine learning model 506 , a convolutional layer may weighted sample the combined feature map 612 to generate a fused feature map 616 .
[0099] Still reference Figure 6 , the fused feature map 616 can be input into the flattening network 620. For example, the fused feature map 616 can be in a three-dimensional format, and the flattening network 620 can transform or represent the values in the three-dimensional format into a one-dimensional vector (e.g., a one-dimensional vector array of values) after receiving the fused feature map 616, because such a format can be suitable for processing by other parts of the machine learning model 506. For example, as part of the conversion of the fused feature map 616 to a one-dimensional vector, the flattening network 620 can list the features of the fused feature map 616 in the three-dimensional format as part of an array, so that the length value, width value, and height value of the fused feature map 616 (corresponding to values on different dimensions (e.g., x-axis, y-axis, and z-axis)) can be represented as part of a single one-dimensional vector array of values (e.g., x-value, y-value, z-value, etc.). The one-dimensional vector can be input into the fully connected layer 622, which can only receive input in the one-dimensional vector format.
[0100] In operation, the fully connected layer may receive the fused feature map 616 converted to a one-dimensional vector format and output the relative transformation 508. For example, the relative transformation output by the fully connected layer 622 may be represented in the form of a seven-dimensional vector, such that three values in the seven-dimensional vector correspond to translations along at least one of the x-axis, the y-axis, and the z-axis, and four values in the vector correspond to rotational movements of θ radians around a fixed point about the unit axis (X, Y, Z).
[0101] Figure 7Depicts a system 700 for performing another process for determining a trajectory of a vehicle within a physical space using a machine learning model 701 in accordance with some embodiments of the present subject matter described and illustrated herein. In particular, Figure 7 701 includes a coarse neural network 702 and a fine neural network 706. The machine learning model 701 can receive the target point cloud 501 and the source point cloud 504 as separate inputs. In addition, note that the operation of the coarse neural network 702 is similar to that of the coarse neural network 702. Figure 6 The operation of the machine learning model 506 shown and described in detail above is substantially similar, and thus, the output of the coarse neural network 702 (e.g., the machine learning model 506) can be a first relative transformation 704 that is comparable to the relative transformation 508. The first relative transformation 704 can be used to align the target point cloud 502 with the source point cloud 504 to a certain degree. In embodiments, this preliminary alignment can be considered a coarse point registration of the target point cloud 502 with the source point cloud 504. Further alignment of the target point cloud 502 with the source point cloud 504 is possible and will likely improve the likelihood of generating an accurate and reliable high-definition map, which in turn enables safe operation and navigation of the vehicle.
[0102] Back to Figure 7 After determining the first relative transformation 704, the machine learning model 701 can be applied to the first relative transformation 704 to determine a second relative transformation. The second relative transformation can be used to further align the target point cloud 502 that is aligned with the source point cloud 504. This further alignment corresponds to a fine point registration of the target point cloud 502 relative to the source point cloud 504, for example, within the coordinate system of the source point cloud 504. In this way, the machine learning model 701 can be used to generate Figure 7 An aligned point cloud 710 is shown.
[0103] Fig. 8A Depicts a system 800 for implementing a process for pre-training a machine learning model 506 according to some embodiments of the current subject matter described and illustrated herein. In particular, Fig. 8A Depicts the components and implementation steps associated with pre-training a machine learning model 506 for the purpose of determining a relative transformation 508 for aligning a target point cloud 502 with a source point cloud 504. Doing so can improve the performance of subsequently trained machine learning models 506, and in particular can improve performance when it comes to determining the correct relative transformation 508 when the misalignment between the source point cloud 504 and the target point cloud 502 is large (e.g., large translation and / or large rotation, etc.). Fig. 8A As shown, the pre-trained machine learning model 506 includes using a reconstruction network 820, Figure 8B The reconstruction network 820 is shown in detail in FIG. 8 and described further below.
[0104] refer to Fig. 8A , the example target point cloud 802 and the example source point cloud 804 obtained from the training data for training the machine learning model 506 to determine the first relative transformation for aligning the target point cloud 502 with the source point cloud 504 can be input into the encoding network 602 and the backbone 604, the encoding network 602 encodes the example source point cloud 804 and the example target point cloud 802 respectively, and the backbone 604 processes the encoded example source point cloud 804 and the example target point cloud 802 to generate the example source feature map 806 and the example target feature map 808. Thereafter, the cascader 610 can generate the example combined feature map 810 by, for example, directly stacking the example target feature map 808 onto the example source feature map 806 without relying on any calculation, and the result can be that the length value and width value of each feature map are retained, and the height value is equivalent to the combined value of the heights of the feature maps. Thereafter, the example combined feature map 810 may be input into the convolution layer 614, which in turn outputs the example fused feature map 818 as a result of a 1×1 convolution operation. For example, the convolution layer 614 may perform weighted subsampling on the example combined feature map 810 to generate the example fused feature map 818.
[0105] In an embodiment, the example fused feature map 616 can be input into a flattening network 620, which can transform or represent the feature values of the example fused feature map 818 as a one-dimensional vector. For example, the flattening network 620 can list the features of the example fused feature map 818 as part of an array so that the length value, width value, and height value of the example fused feature map 818 (corresponding to values on different dimensions (e.g., x-axis, y-axis, and z-axis)) can be represented as part of a one-dimensional vector array of values (e.g., x-value, y-value, z-value, etc.). The one-dimensional vector can then be input into a fully connected layer 622, which outputs an example relative transformation 822 in the form of a seven-dimensional vector. However, it is noted that determining the example relative transformation 822 includes using a reconstruction network 820, which receives the example target feature map 808 and the example fused feature map 818, and can feed the offset data to the fully connected layer 622.
[0106] Figure 8BA block diagram is shown detailing components and steps included as part of the operation of a reconstruction network 820 for pre-training a machine learning model 506 to generate a relative transformation to align an example target point cloud 802 with a source point cloud 804. In some embodiments, the reconstruction network 820 includes an offset transformer 823, a deformable convolutional network 826, and a decoding network 828. Prior to completing training the machine learning model 506 to determine the first relative transformation 704, during pre-training the machine learning model 506, the machine learning model 506 can be operated to reconstruct a point cloud using point cloud training data.
[0107] In pre-training, the example target feature map 808 may be input into the deformable convolution network 826, and the example fused feature map 818 may be input into the offset converter 823. The offset converter 823 may receive the example fused feature map 818 and generate an offset 824 that is fed into the deformable convolution network 826. The deformable convolution network 826 applies the received offset 824 to the example target feature map 808 to deform the example target feature map 808, which is generated from the encoded example target point cloud 802. The output of the deformable convolution network 826 (the example target feature map 808 deformed based on the applied offset 824) is input into the decoding network 828, which decodes the deformed example target feature map 808 (the example target feature map 808 is based on the example target point cloud 802). Note that the offset can be used to adjust the operation of the machine learning model 506, namely to minimize the difference between the example target feature map 808 (which is based on the example target point cloud 802) deformed based on the applied offset 824 and the decoded example target feature map 808 (which is associated with the decoded example target point cloud 802) deformed based on the applied offset 824. It should also be noted that the encoder can be pre-trained based on at least the difference between the point cloud (e.g., the example target point cloud 802) and the decoded point cloud (e.g., the example target point cloud 802).
[0108] Fig. 9A flow chart is depicted showing an example of a process 900 for determining a trajectory of a vehicle (e.g., an autonomous vehicle) based on a relative transformation used to further align a target point cloud. In some embodiments, one or more operations described with respect to process 900 are performed by a perception system 402, a planning system 404, and / or a control system 408 (e.g., in whole and / or in part, etc.) of an autonomous vehicle computing 400 of the vehicle, one or more of which may apply a machine learning model 506 trained to determine a first relative transformation, which may be used to align the target point cloud 502 with the source point cloud 504 and determine the trajectory of the vehicle, as described above. Additionally or alternatively, in some embodiments, one or more steps described with respect to process 900 are performed by other devices or groups of devices (e.g., in whole and / or in part, etc.) separate from or including the autonomous vehicle computing 400.
[0109] At 902, a fused feature map may be generated from a plurality of feature maps. For example, the fused feature map may be generated by concatenating a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud. In an embodiment, the target point cloud may represent at least a portion of the same physical space as the source point cloud. For example, the target point cloud may represent a physical space in the form of a city block, where street intersections, traffic lights, stop signs, pedestrians, multiple buildings, etc. may be located on the city block. As a non-limiting example, the target point cloud may represent a physical space relative to a particular direction (e.g., west), while the source point cloud may represent the same physical space (city block) relative to other directions (e.g., north). The first feature map corresponds to an encoded representation of the source point cloud, and the second feature map corresponds to an encoded representation of the target point cloud. As part of the encoding process, one or more objects (e.g., one or more objects in the form of a traffic light, a stop sign, a pedestrian, multiple buildings, etc.) included as part of the source point cloud and the target point cloud may be represented as one or more columns, which serve as a boundary around multiple points associated with a particular object. For example, one or more posts may serve as a boundary around a subset of points associated with, for example, pedestrians, while another set of one or more posts may serve as a boundary around another subset of points associated with, for example, stop signs.
[0110] At 904, a machine learning model may be applied to determine a first relative transformation for aligning the target point cloud with the source point cloud. For example, a machine learning model (e.g., machine learning model 506) may be applied to determine a first relative transformation that may be used to align the target point cloud within the coordinate system of the source point cloud. For example, the first relative transformation of the target point cloud may include a translation along at least one of the x-axis, the y-axis, and the z-axis. Alternatively or additionally, the first relative transformation may include a rotation of θ radians about a unit axis (e.g., the x-axis, the y-axis, or the z-axis) around a fixed point. In operation, the machine learning model (e.g., machine learning model 506) may perform weighted subsampling using a convolutional layer included as part of the machine learning model. Weighted subsampling enables corresponding features from the source point cloud and the target point cloud to be extracted from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map. In addition, the combined feature map may be flattened into a one-dimensional vector to be processed by a fully connected layer of the machine learning model. In an embodiment, the output of the fully connected layer is a first transformation that is used to perform a rough point registration of the target point cloud. For example, the first transformation can be represented as a seven-dimensional vector, where three of the elements correspond to translations along at least one of the x-axis, y-axis, and z-axis, and four of the elements correspond to a rotation of θ radians about a fixed point about a unit axis (e.g., the x-axis, the y-axis, or the z-axis).
[0111] At 906, an aligned point cloud can be generated by transforming the target point cloud at least according to a first relative transformation. For example, the transformation of the target point cloud can include translating the target point cloud along at least one of the x-axis, the y-axis, and the z-axis and / or rotating the target point cloud about a fixed point. In addition, a machine learning model (e.g., the machine learning model 506) can be applied to determine a second relative transformation to further align the aligned target point cloud with the source point cloud. In particular, the further alignment of the target point cloud based on the second relative transformation is a fine point registration of the target point cloud relative to the source point cloud, for example, within the coordinate system of the source point cloud.
[0112] At 908, a trajectory of the vehicle in physical space may be determined based on the relative transformation used to further align the target point cloud. For example, the following path is possible: the vehicle (e.g., Figure 1 The vehicle 102 shown, Figure 2 The vehicle 200 shown, etc.) can travel along the path to navigate within a physical space (e.g., a city block in which street intersections, traffic lights, stop signs, pedestrians, and multiple buildings can be located) so that the vehicle accurately travels from one or more source locations to one or more destination locations while avoiding collisions with other vehicles, pedestrians, and / or buildings, etc.
[0113] According to some non-limiting embodiments or examples, a method is provided, comprising: using at least one data processor to generate a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; using the at least one data processor to apply a machine learning model, which is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; using the at least one data processor to generate an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and using the at least one data processor to determine a trajectory of a vehicle within the physical space based at least on the first relative transformation.
[0114] According to some non-limiting embodiments or examples, a system is provided, comprising: at least one processor; and at least one non-transitory storage medium storing instructions, which, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising: using at least one data processor to generate a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; using the at least one data processor to apply a machine learning model, which is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; using the at least one data processor to generate an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and using the at least one data processor to determine a trajectory of a vehicle within the physical space based at least on the first relative transformation.
[0115] According to some non-limiting embodiments or examples, at least one non-transitory computer-readable medium is provided, which includes one or more instructions, which when executed by at least one processor causes the at least one processor to perform operations, the operations including: using the at least one data processor to generate a fused feature map by cascading at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; using the at least one data processor, applying a machine learning model, which is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; using the at least one data processor, generating an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and using the at least one data processor, determining a trajectory of a vehicle within the physical space based at least on the first relative transformation.
[0116] Further non-limiting aspects or embodiments are set forth in the following numbered clauses:
[0117] Item 1: A method comprising: using at least one data processor to generate a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; using the at least one data processor to apply a machine learning model that is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; using the at least one data processor to generate an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and using the at least one data processor to determine a trajectory of a vehicle within the physical space based at least on the first transformation.
[0118] Clause 2: The method of clause 1, wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
[0119] Clause 3: The method according to Clause 2 further includes: using the at least one data processor to encode the source point cloud and the target point cloud so that one or more objects present in each point cloud are represented as one or more columns, and each column in the one or more columns defines multiple points associated with the corresponding object.
[0120] Clause 4: The method according to any one of clauses 1 to 3 further includes: using the at least one data processor to pre-train the machine learning model before training the machine learning model to determine the first relative transformation, the pre-training including: using the at least one data processor to reconstruct the point cloud by at least decoding the encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud.
[0121] Clause 5: The method of any one of clauses 4, further comprising: pre-training an encoder for generating an encoding of the point cloud based at least on the difference between the point cloud and the decoded point cloud.
[0122] Clause 6: A method according to any one of clauses 1 to 5, wherein the machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
[0123] Clause 7: A method according to clause 6, wherein the machine learning model includes a 1×1 convolutional layer configured to perform the weighted sub-sampling.
[0124] Clause 8: A method according to clause 6 or 7, wherein the machine learning model also determines the first relative transformation by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.
[0125] Clause 9: The method according to any one of clauses 1 to 8 further includes: using the at least one data processor, applying the first relative transformation to perform coarse point alignment of the target point cloud; using the at least one data processor, applying the machine learning model trained to determine a second relative transformation used to further align the aligned target point cloud with the source point cloud; and using the at least one data processor, applying the second relative transformation to perform fine point alignment of the target point cloud.
[0126] Clause 10: The method of any one of clauses 1 to 9, wherein the first relative transformation comprises a translation along at least one of an x-axis, a y-axis, and a z-axis.
[0127] Clause 11: A method according to any one of clauses 1 to 10, wherein the first relative transformation comprises a rotation of θ radians about a unit axis (X, Y, Z) around a fixed point.
[0128] Clause 12: A method according to any one of clauses 1 to 11, wherein the source point cloud and the target point cloud comprise three-dimensional point clouds.
[0129] Item 13: A system comprising: at least one processor; and at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: using at least one data processor to generate a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; using the at least one data processor to apply a machine learning model that is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; using the at least one data processor to generate an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and using the at least one data processor to determine a trajectory of a vehicle within the physical space based at least on the first relative transformation.
[0130] Clause 14: The system of clause 13, wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
[0131] Clause 15: A system according to clause 14, wherein the operation further comprises: using the at least one data processor, encoding the source point cloud and the target point cloud so that one or more objects present in each point cloud are represented as one or more columns, each column in the one or more columns defining multiple points associated with the corresponding object.
[0132] Clause 16: A system according to any one of clauses 13 to 15, wherein the operation further comprises: using the at least one data processor to pre-train the machine learning model before training the machine learning model to determine the first relative transformation, the pre-training comprising: using the at least one data processor to reconstruct the point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud.
[0133] Clause 17: The system of clause 16, wherein the operations further comprise: pre-training an encoder for generating an encoding of the point cloud based at least on the difference between the point cloud and the decoded point cloud.
[0134] Clause 18: A system according to any one of clauses 13 to 17, wherein the machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
[0135] Clause 19: A system according to clause 18, wherein the machine learning model includes a 1×1 convolutional layer configured to perform the weighted sub-sampling.
[0136] Clause 20: A system according to clause 18 or 19, wherein the machine learning model also determines the first relative transformation by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.
[0137] Clause 21: A system according to any one of clauses 13 to 20, wherein the operation further comprises: using the at least one data processor, applying the first relative transformation to perform coarse point alignment of the target point cloud; using the at least one data processor, applying the machine learning model trained to determine a second relative transformation for further aligning the aligned target point cloud with the source point cloud; and using the at least one data processor, applying the second relative transformation to perform fine point alignment of the target point cloud.
[0138] Clause 22: The system of any of clauses 13 to 21, wherein the first relative transformation comprises a translation along at least one of an x-axis, a y-axis, and a z-axis.
[0139] Clause 23: The system of any of clauses 13 to 22, wherein the first relative transformation comprises a rotation of θ radians about a unit axis (X, Y, Z) about a fixed point.
[0140] Clause 24: The system of any one of clauses 13 to 23, wherein the source point cloud and the target point cloud comprise three-dimensional point clouds.
[0141] Item 25: At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: use at least one data processor to generate a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, wherein the target point cloud corresponds to at least a portion of the same physical space as the source point cloud; use the at least one data processor to apply a machine learning model that is trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; use the at least one data processor to generate an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and use the at least one data processor to determine a trajectory of a vehicle within the physical space based at least on the first relative transformation.
[0142] Clause 26: At least one non-transitory storage medium according to clause 25, wherein the first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
[0143] Clause 27: According to the at least one non-transitory storage medium described in Clause 26, the operation also includes: using the at least one data processor, encoding the source point cloud and the target point cloud so that one or more objects present in each point cloud are represented as one or more columns, and each column in the one or more columns defines multiple points associated with the corresponding object.
[0144] Clause 28: At least one non-transitory storage medium according to any one of clauses 25 to 27, wherein the operation further comprises: using the at least one data processor to pre-train the machine learning model before training the machine learning model to determine the first relative transformation, the pre-training comprising: using the at least one data processor to reconstruct the point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud; and pre-training an encoder for generating the encoding of the point cloud based at least on the difference between the point cloud and the decoded point cloud.
[0145] Clause 29: At least one non-transitory storage medium according to any one of clauses 25 to 28, wherein the machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
[0146] Clause 30: At least one non-transitory storage medium according to clause 29, wherein the machine learning model comprises a 1×1 convolutional layer configured to perform the weighted sub-sampling.
[0147] In the previous description, aspects and embodiments of the present disclosure have been described with reference to many specific details, which may vary depending on the implementation. Therefore, the description and the accompanying drawings should be regarded as illustrative, not restrictive. The only and exclusive indication of the scope of the invention, and the applicant's expectation that the content of the scope of the invention is the literal and equivalent scope of the claims issued from this application in the specific form of the claims, including any subsequent amendments. Any definition of the terms used to be included in such claims that are clearly set forth herein should be based on the meaning of such terms as used in the claims. In addition, when the term "also includes" is used in the previous description or the attached claims, the following of the phrase may be an additional step or entity, or a sub-step / sub-entity of the previously described step or entity.
Claims
1. A method comprising: Using at least one data processor, generating a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of the same physical space as the source point cloud; applying, using the at least one data processor, a machine learning model trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; generating, using the at least one data processor, an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; as well as Using the at least one data processor, a trajectory of a vehicle within the physical space is determined based at least on the first relative transformation.
2. The method according to claim 1, wherein: The first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
3. The method according to claim 2, further comprising: Using the at least one data processor, the source point cloud and the target point cloud are encoded such that one or more objects present in each point cloud are represented as one or more columns, each of the one or more columns defining a plurality of points associated with the corresponding object.
4. The method according to any one of claims 1 to 3, further comprising: Using the at least one data processor, before training the machine learning model to determine the first relative transformation, pre-training the machine learning model, the pre-training comprising: using the at least one data processor, reconstructing the point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud.
5. The method according to claim 4, further comprising: An encoder for generating an encoding of the point cloud is pre-trained based at least on the difference between the point cloud and the decoded point cloud.
6. The method according to any one of claims 1 to 5, wherein: The machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
7. The method according to claim 6, wherein: The machine learning model includes a 1×1 convolutional layer configured to perform the weighted sub-sampling.
8. The method according to claim 6 or 7, wherein: The machine learning model also determines the first relative transformation by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.
9. The method according to any one of claims 1 to 8, further comprising: applying, using the at least one data processor, the first relative transformation to perform a coarse point registration of the target point cloud; applying, using the at least one data processor, the machine learning model trained to determine a second relative transformation to further align the aligned target point cloud with the source point cloud; as well as Using the at least one data processor, the second relative transformation is applied to perform fine point registration of the target point cloud.
10. The method according to any one of claims 1 to 9, wherein: The first relative transformation includes a translation along at least one of an x-axis, a y-axis, and a z-axis.
11. The method according to any one of claims 1 to 10, wherein: The first relative transformation includes a rotation of θ radians about a unit axis (X, Y, Z) around a fixed point.
12. The method according to any one of claims 1 to 11, wherein: The source point cloud and the target point cloud include three-dimensional point clouds.
13. A system comprising: at least one processor; as well as at least one non-transitory storage medium storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: Using at least one data processor, generating a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of the same physical space as the source point cloud; applying, using the at least one data processor, a machine learning model trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; generating, using the at least one data processor, an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; and Using the at least one data processor, a trajectory of a vehicle within the physical space is determined based at least on the first relative transformation.
14. The system according to claim 13, wherein: The first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
15. The system of claim 14, wherein: The operations also include: Using the at least one data processor, the source point cloud and the target point cloud are encoded such that one or more objects present in each point cloud are represented as one or more columns, each of the one or more columns defining a plurality of points associated with the corresponding object.
16. A system according to any one of claims 13 to 15, wherein: The operations also include: Using the at least one data processor, before training the machine learning model to determine the first relative transformation, pre-training the machine learning model, the pre-training comprising: using the at least one data processor, reconstructing the point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize the difference between the point cloud and the decoded point cloud.
17. The system of claim 16, wherein: The operations also include: An encoder for generating an encoding of the point cloud is pre-trained based at least on the difference between the point cloud and the decoded point cloud.
18. A system according to any one of claims 13 to 17, wherein: The machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
19. The system of claim 18, wherein: The machine learning model includes a 1×1 convolutional layer configured to perform the weighted sub-sampling.
20. The system according to claim 18 or 19, wherein: The machine learning model also determines the first relative transformation by at least flattening the combined feature map into a one-dimensional vector for processing by a fully connected layer of the machine learning model.
21. A system according to any one of claims 13 to 20, wherein: The operations also include: applying, using the at least one data processor, the first relative transformation to perform a coarse point registration of the target point cloud; applying, using the at least one data processor, the machine learning model trained to determine a second relative transformation to further align the aligned target point cloud with the source point cloud; and Using the at least one data processor, the second relative transformation is applied to perform fine point registration of the target point cloud.
22. A system according to any one of claims 13 to 21, wherein: The first relative transformation includes a translation along at least one of an x-axis, a y-axis, and a z-axis.
23. A system according to any one of claims 13 to 22, wherein: The first relative transformation includes a rotation of θ radians about a unit axis (X, Y, Z) around a fixed point.
24. A system according to any one of claims 13 to 23, wherein: The source point cloud and the target point cloud include three-dimensional point clouds.
25. At least one non-transitory storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: Using at least one data processor, generating a fused feature map by concatenating at least a first feature map corresponding to a source point cloud and a second feature map corresponding to a target point cloud, the target point cloud corresponding to at least a portion of the same physical space as the source point cloud; applying, using the at least one data processor, a machine learning model trained to determine a first relative transformation for aligning the target point cloud with the source point cloud based at least on the fused feature map; generating, using the at least one data processor, an aligned target point cloud by transforming at least the target point cloud according to the first relative transformation; as well as Using the at least one data processor, a trajectory of a vehicle within the physical space is determined based at least on the first relative transformation.
26. The at least one non-transitory storage medium of claim 25, wherein: The first feature map corresponds to an encoded representation of the source point cloud, and wherein the second feature map corresponds to an encoded representation of the target point cloud.
27. The at least one non-transitory storage medium of claim 26, wherein: The operations also include: Using the at least one data processor, the source point cloud and the target point cloud are encoded such that one or more objects present in each point cloud are represented as one or more columns, each of the one or more columns defining a plurality of points associated with the corresponding object.
28. At least one non-transitory storage medium according to any one of claims 25 to 27, wherein: The operations also include: Using the at least one data processor, prior to training the machine learning model to determine the first relative transformation, pre-training the machine learning model, the pre-training comprising: using the at least one data processor, reconstructing the point cloud by at least decoding an encoding of the point cloud deformed by one or more offsets determined by the machine learning model; and adjusting the machine learning model to minimize a difference between the point cloud and the decoded point cloud; and An encoder for generating an encoding of the point cloud is pre-trained based at least on the difference between the point cloud and the decoded point cloud.
29. At least one non-transitory storage medium according to any one of claims 25 to 28, wherein: The machine learning model determines the first relative transformation by performing at least weighted sub-sampling, and the weighted sub-sampling is used to extract corresponding features from the source point cloud and the target point cloud from the fused feature map, so that the first feature map and the second feature map are downsampled into a combined feature map.
30. The at least one non-transitory storage medium of claim 29, wherein: The machine learning model includes a 1×1 convolutional layer configured to perform the weighted sub-sampling.